Amux
All models

MinMax: MiniMax M2.1 highspeed

Active

MiniMax-M2.1-Highspeed is a high-speed variant of MiniMax-M2.1, delivering the same state-of-the-art capabilities in coding, agentic workflows, and modern application development with significantly faster inference and greater responsiveness. With only 10 billion activated parameters, it combines strong real-world performance with exceptional latency, scalability, and cost efficiency, making it especially well suited for interactive and latency-sensitive workloads.

Input:
Output:
Context length:
205K
Max output:
131K
Published:
2025-12-23
toolsreasoningstreamingcaching

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
20% offminimax$0.6$0.48/1M tokens$2.4$1.92/1M tokens
Read$0.03$0.024/1M tokensWrite$0.375$0.3/1M tokens
205K

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Throughput

24 hours

Average output speed on streaming requests, in tokens per second (TPS).

Latency

24 hours

Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Related models

More models from MiniMax

MiniMax: MinMax M31M contextfrom $0.3 / M tokens1 providerMiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.
MiniMax: MinMax M2.7205K contextfrom $0.21 / M tokens1 providerMiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.
MinMax: MiniMax M2.7 highspeed205K contextfrom $0.42 / M tokens1 providerMiniMax-M2.7-Highspeed is a high-speed variant of MiniMax-M2.7, delivering the same advanced agentic capabilities and model performance with significantly faster inference and greater responsiveness. Designed for autonomous, real-world productivity, it enables more agile planning, execution, and refinement of complex tasks across dynamic environments.