Amux
All models

Z.ai: GLM 5.2

Active

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

By:
Z.ai
Input:
Output:
Context length:
1M
Max output:
128K
Published:
2026-06-17
toolsreasoningstructuredOutputstreamingcaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
70% offamux-special$1.4$0.42/1M tokens$4.4$1.32/1M tokens
Read$0.26$0.078/1M tokensWrite
1M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from Z.ai

Z.ai: GLM 5.3 Flash1M contextfrom $0.045 / M tokens1 providerGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Z.ai: GLM 5.31M contextfrom $0.42 / M tokens1 providerGLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.