The Qwen3.8-27B native visual language-intensive model has been significantly improved compared to the 3.6-27B version, with key enhancements in programming and office scene capabilities in text and visual modalities. It completes complex tasks end-to-end more reliably and produces trustworthy output.
Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
24 hours
Average output speed on streaming requests, in tokens per second (TPS).
Latency
24 hours
Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
24 hours
Token usage and cost for this model over time, split by the provider that served each request.
Related models
More models from Qwen
Qwen: Qwen3.8-Max-09021M contextfrom $2 / M tokens1 providerQwen3.8-Max-0902 is a snapshot version of Qwen3.8-Max, which has stronger coding performance and is suitable for complex engineering tasks and long-term independent development. It simultaneously improves the agent's tool usage capabilities and visual understanding capabilities for graph reasoning, document parsing and multi-modal perception, and retains the context window of 1 million tokens, reasoning modes and a complete tool ecosystem.
Qwen: Qwen3.8-Flash1M contextfrom $0.144 / M tokens1 providerQwen3.8-Flash is the latest multi-modal model of the Qwen series, which combines powerful reasoning and generation capabilities with amazing speed. It natively supports context windows with millions of tokens, and can handle long documents, complete code bases, and complex conversations at once. This model excels in code assistance, agent workflow, and visual understanding, whether it is autonomously repairing code, operating desktop applications, or analyzing charts and long videos. Qwen3.8-Flash is fully compatible with OpenAI and Anthropic API protocols, and can be seamlessly integrated with popular development tools such as Claude Code and Codex, helping developers easily build high-concurrency applications and intelligent workflows. With excellent performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and enterprises seeking a balance between performance and cost for AI applications.
Qwen: Qwen3.8-2.4T-A95B1M contextfrom $1.8 / M tokens1 providerQwen3.8-2.4T-A95B is the open source version of Tongyi Qianwen’s latest flagship series, released in August 2026. It uses a sparse MoE architecture with a total parameter volume of 2.4 trillion, activating approximately 95 billion parameters at each step, and supports 1 million token context and hybrid attention mechanisms. Core benchmarks include GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, and BabyVision 82.0. Ranked 4th globally on CodeArena.