Amux
All models

Z.ai: GLM 5.3 FlashX

Active

Z.AI: GLM 5.3 FlashX is a faster, smoother version of GLM-5.3-Flash with inference speeds up to 200 tokens/s. Z.AI: GLM 5.3 FlashX is Z.ai’s native multi-modal model, suitable for efficient programming and long-range agent tasks. Its hybrid sparse and linear attention architecture reduces computational overhead while maintaining accurate long-context performance.

By:
Z.ai
Input:
Output:
Context length:
1M
Max output:
128K
Published:
2026-09-18
toolsreasoningstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
70% offamux-special0.00%$0.375$0.1125/1M tokens$1.25$0.375/1M tokens
Read$0.075$0.0225/1M tokensWrite
1M4.0s117.2 tps

Availability

7 days

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

7 days

Average output speed on streaming requests, in tokens per second (TPS).

Latency

7 days

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

7 days

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from Z.ai

Z.ai: GLM 5.3 Flash1M contextfrom $0.045 / M tokens1 providerGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Z.ai: GLM 5.31M contextfrom $0.42 / M tokens1 providerGLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.
Z.ai: GLM 5.21M contextfrom $0.42 / M tokens1 providerGLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.