MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.
Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.
Provider
50% offminimax
$0.6-1.2$0.3-0.6/1M tokens
$2.4-4.8$1.2-2.4/1M tokens
Read$0.12-0.24$0.06-0.12/1M tokensWrite—
1M
Availability
24 hours
Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Throughput
24 hours
Average output speed on streaming requests, in tokens per second (TPS).
Latency
24 hours
Average time to first token (TTFT) on streaming requests — how long from sending a request to receiving its first token.
Activity
24 hours
Token usage and cost for this model over time, split by the provider that served each request.
Related models
More models from MiniMax
MiniMax: MinMax M2.7205K contextfrom $0.21 / M tokens1 providerMiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.
MinMax: MiniMax M2.7 highspeed205K contextfrom $0.42 / M tokens1 providerMiniMax-M2.7-Highspeed is a high-speed variant of MiniMax-M2.7, delivering the same advanced agentic capabilities and model performance with significantly faster inference and greater responsiveness. Designed for autonomous, real-world productivity, it enables more agile planning, execution, and refinement of complex tasks across dynamic environments.
MiniMax: MinMax M2.5205K contextfrom $0.24 / M tokens1 providerMiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. Scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, M2.5 is also more token efficient than previous generations, having been trained to optimize its actions and output through planning.