Amux
All models

Qwen: Qwen3.8-Max-0902

Active

Qwen3.8-Max-0902 is a snapshot version of Qwen3.8-Max, which has stronger coding performance and is suitable for complex engineering tasks and long-term independent development. It simultaneously improves the agent's tool usage capabilities and visual understanding capabilities for graph reasoning, document parsing and multi-modal perception, and retains the context window of 1 million tokens, reasoning modes and a complete tool ecosystem.

By:
Qwen
Input:
Output:
Context length:
1M
Max output:
128K
Published:
2026-09-02
toolsstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
alibaba-cloud$2/1M tokens$6/1M tokens
Read explicit$0.17/1M tokensRead implicit$0.25/1M tokensWrite$2.5/1M tokens
1M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from Qwen

Qwen: Qwen3.8-Flash1M contextfrom $0.144 / M tokens1 providerQwen3.8-Flash is the latest multi-modal model of the Qwen series, which combines powerful reasoning and generation capabilities with amazing speed. It natively supports context windows with millions of tokens, and can handle long documents, complete code bases, and complex conversations at once. This model excels in code assistance, agent workflow, and visual understanding, whether it is autonomously repairing code, operating desktop applications, or analyzing charts and long videos. Qwen3.8-Flash is fully compatible with OpenAI and Anthropic API protocols, and can be seamlessly integrated with popular development tools such as Claude Code and Codex, helping developers easily build high-concurrency applications and intelligent workflows. With excellent performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and enterprises seeking a balance between performance and cost for AI applications.
Qwen: Qwen3.8-27B1M contextfrom $0.4 / M tokens1 providerThe Qwen3.8-27B native visual language-intensive model has been significantly improved compared to the 3.6-27B version, with key enhancements in programming and office scene capabilities in text and visual modalities. It completes complex tasks end-to-end more reliably and produces trustworthy output.
Qwen: Qwen3.8-2.4T-A95B1M contextfrom $1.8 / M tokens1 providerQwen3.8-2.4T-A95B is the open source version of Tongyi Qianwen’s latest flagship series, released in August 2026. It uses a sparse MoE architecture with a total parameter volume of 2.4 trillion, activating approximately 95 billion parameters at each step, and supports 1 million token context and hybrid attention mechanisms. Core benchmarks include GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, and BabyVision 82.0. Ranked 4th globally on CodeArena.