Amux
All models

MiniMax

Models served by MiniMax

Models · 9

  • MiniMax: MiniMax H3
    Chat

    MiniMax: MiniMax H3 is a lightweight open source weighted video generation model developed by MiniMax. It is designed for precise multimodal editing and controlled content generation, including command-driven editing, text and brand rendering, and video-to-video motion transfer. The model is suitable for commercial creative workflows such as advertising, e-commerce, games and interface design, and supports native audiovisual output for reference-driven generation.

    Input:
    Output:
    Input:
    Output:
    $0.08-0.13/second
    Context length:
    13K
    Max output:
    1 provider
  • MiniMax: MiniMax H3 Max
    Chat

    MiniMax: MiniMax H3 Max is a video generation model jointly released by MiniMax and fal.ai. The model was further trained by fal.ai based on MiniMax H3 and optimized for high-speed generation. It supports mainstream 480P and 768P output, generating videos faster than MiniMax H3. Currently, it supports 'text-to-video' and 'image-to-video', and will later support 'reference image generation'.

    Input:
    Output:
    Input:
    Output:
    $0.05-0.08/second
    Context length:
    13K
    Max output:
    1 provider
  • MiniMax: MinMax M2.1
    20% offChat

    MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

    Input:
    Output:
    Input:
    $0.3$0.24/1M
    Output:
    $1.2$0.96/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.1 highspeed
    20% offChat

    MiniMax-M2.1-Highspeed is a high-speed variant of MiniMax-M2.1, delivering the same state-of-the-art capabilities in coding, agentic workflows, and modern application development with significantly faster inference and greater responsiveness. With only 10 billion activated parameters, it combines strong real-world performance with exceptional latency, scalability, and cost efficiency, making it especially well suited for interactive and latency-sensitive workloads.

    Input:
    Output:
    Input:
    $0.6$0.48/1M
    Output:
    $2.4$1.92/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MiniMax: MinMax M2.5
    20% offChat

    MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. Scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, M2.5 is also more token efficient than previous generations, having been trained to optimize its actions and output through planning.

    Input:
    Output:
    Input:
    $0.3$0.24/1M
    Output:
    $1.2$0.96/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.5 highspeed
    20% offChat

    MiniMax-M2.5-Highspeed is a high-speed variant of MiniMax-M2.5, delivering the same SOTA performance and real-world productivity capabilities with significantly faster inference and greater responsiveness. Trained across diverse and complex digital working environments, it extends the coding expertise of M2.1 into general office work, including generating and operating Word, Excel, and PowerPoint files, switching seamlessly across software environments, and collaborating across agent and human teams. With the same strong benchmark performance and token-efficient planning capabilities as M2.5, the Highspeed variant is optimized for faster, more agile execution in latency-sensitive workflows.

    Input:
    Output:
    Input:
    $0.6$0.48/1M
    Output:
    $2.4$1.92/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MiniMax: MinMax M2.7
    30% offChat

    MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

    Input:
    Output:
    Input:
    $0.3$0.21/1M
    Output:
    $1.2$0.84/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.7 highspeed
    30% offChat

    MiniMax-M2.7-Highspeed is a high-speed variant of MiniMax-M2.7, delivering the same advanced agentic capabilities and model performance with significantly faster inference and greater responsiveness. Designed for autonomous, real-world productivity, it enables more agile planning, execution, and refinement of complex tasks across dynamic environments.

    Input:
    Output:
    Input:
    $0.6$0.42/1M
    Output:
    $2.4$1.68/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MiniMax: MinMax M3
    50% offChat

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

    Input:
    Output:
    Input:
    $0.6$0.3/1M
    Output:
    $2.4$1.2/1M
    Context length:
    1M
    Max output:
    512K
    1 provider