Qwen3.8-Flash is the latest multi-modal model of the Qwen series, which combines powerful reasoning and generation capabilities with amazing speed. It natively supports context windows with millions of tokens, and can handle long documents, complete code bases, and complex conversations at once. This model excels in code assistance, agent workflow, and visual understanding, whether it is autonomously repairing code, operating desktop applications, or analyzing charts and long videos. Qwen3.8-Flash is fully compatible with OpenAI and Anthropic API protocols, and can be seamlessly integrated with popular development tools such as Claude Code and Codex, helping developers easily build high-concurrency applications and intelligent workflows. With excellent performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and enterprises seeking a balance between performance and cost for AI applications.
- By:
- Qwen
- Input:
- Output:
- Context length:
- 1M
- Max output:
- 1M
- Published:
- 2026-08-27
Providers
Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.
| Provider | ||||
|---|---|---|---|---|
| 10% offalibaba-cloud | $0.16$0.144/1M tokens | $0.47$0.423/1M tokens | Read explicit$0.016$0.0144/1M tokensRead implicit$0.016$0.0144/1M tokensWrite$0.2$0.18/1M tokens | 1M |
Availability
24 hoursSuccess rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
24 hoursAverage output speed on streaming requests, in tokens per second (TPS).
Latency
24 hoursAverage time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
24 hoursToken usage and cost for this model over time, split by the provider that served each request.
Related models
More models from Qwen