Amux
All models

MoonshotAI: Kimi K3

Active

MoonshotAI: Kimi K3 is Kimi’s most powerful model to date, with 2.8 trillion parameters. It is built based on the hybrid linear attention mechanism Kimi Delta Attention and attention residual technology. It has native visual understanding capabilities and a context window of 1 million tokens. It is suitable for cutting-edge intelligent scenarios such as software engineering, knowledge work, and deep reasoning.

Input:
Output:
Context length:
1.05M
Max output:
1.05M
Published:
2026-07-17
toolsreasoningstructuredOutputstreamingvisioncaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
5% offalibaba-cloud$3$2.85/1M tokens$15$14.25/1M tokens
Read$0.3$0.285/1M tokensWrite
1.05M
35% offLowestamux-special$3$1.95/1M tokens$15$9.75/1M tokens
Read$0.3$0.195/1M tokensWrite
1.05M

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.