Z.AI: GLM 5.3 FlashX is a faster, smoother version of GLM-5.3-Flash with inference speeds up to 200 tokens/s. Z.AI: GLM 5.3 FlashX is Z.ai’s native multi-modal model, suitable for efficient programming and long-range agent tasks. Its hybrid sparse and linear attention architecture reduces computational overhead while maintaining accurate long-context performance.
- By:
- Z.ai
- Input:
- Output:
- Context length:
- 1M
- Max output:
- 128K
- Published:
- 2026-09-18
Providers
Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.
| Provider | |||||||
|---|---|---|---|---|---|---|---|
| 70% offamux-special | 0.00% | $0.375$0.1125/1M tokens | $1.25$0.375/1M tokens | Read$0.075$0.0225/1M tokensWrite— | 1M | 4.0s | 117.2 tps |
Availability
7 daysSuccess rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
7 daysAverage output speed on streaming requests, in tokens per second (TPS).
Latency
7 daysAverage time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
7 daysToken usage and cost for this model over time, split by the provider that served each request.
Related models
More models from Z.ai