MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.
Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
24 hours
Average output speed on streaming requests, in tokens per second (TPS).
Latency
24 hours
Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
24 hours
Token usage and cost for this model over time, split by the provider that served each request.
Related models
More models from MiniMax
MiniMax: MinMax M31M contextfrom $0.3 / M tokens1 providerMiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.
MiniMax: MinMax M2.7205K contextfrom $0.21 / M tokens1 providerMiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.
MinMax: MiniMax M2.7 highspeed205K contextfrom $0.42 / M tokens1 providerMiniMax-M2.7-Highspeed is a high-speed variant of MiniMax-M2.7, delivering the same advanced agentic capabilities and model performance with significantly faster inference and greater responsiveness. Designed for autonomous, real-world productivity, it enables more agile planning, execution, and refinement of complex tasks across dynamic environments.