Amux

Models

46 models

  • OpenAI: GPT-6 Astra
    95% offChat

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

    Input:
    Output:
    Input:
    $10$0.5/1M
    Output:
    $50$2.5/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • Anthropic: Claude Fable 5.1
    30% offChat

    Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual code generation, and finance and analysis tasks in particular. It also tends to be more concise than Fable 5 in its plans and summaries. We recommend testing it as a direct upgrade wherever you use Fable 5 today, and alongside Opus 5 on reasoning-heavy tasks.

    Input:
    Output:
    Input:
    $10$7/1M
    Output:
    $50$35/1M
    Context length:
    1M
    Max output:
    128K
    2 providers
  • Google: Gemini 3.8 Flash
    50% offChat

    Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

    Input:
    Output:
    Input:
    $1.5$0.75/1M
    Output:
    $7.5$3.75/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • MiniMax: MiniMax H3 Max
    Chat

    MiniMax: MiniMax H3 Max is a video generation model jointly released by MiniMax and fal.ai. The model was further trained by fal.ai based on MiniMax H3 and optimized for high-speed generation. It supports mainstream 480P and 768P output, generating videos faster than MiniMax H3. Currently, it supports 'text-to-video' and 'image-to-video', and will later support 'reference image generation'.

    Input:
    Output:
    Input:
    Output:
    $0.05-0.08/second
    Context length:
    13K
    Max output:
    1 provider
  • Alibaba: Wan3.0-Video
    20% offChat

    Alibaba: Wan3.0-Video is a multi-functional reference video generation model that supports text-to-video, image-to-video (first frame/first and last frame) and reference-based video generation. It is capable of generating videos up to 30 seconds long at 30fps.

    Input:
    Output:
    Input:
    Output:
    $0.04-0.16/second
    Context length:
    20K
    Max output:
    1 provider
  • Alibaba: Wan3.0-Video-Prime
    15% offChat

    Alibaba: Wan3.0-Video-Prime is a high-speed version of the Wan3.0 video generation model. It has the same powerful capabilities as the standard Wan3.0-Video model, supports comprehensive reference input of four modalities, can generate videos up to 30 seconds, and significantly improves the end-to-end generation speed while providing an immersive audio-visual experience.

    Input:
    Output:
    Input:
    Output:
    $0.0578-0.238/second
    Context length:
    20K
    Max output:
    1 provider
  • DeepSeek: DeepSeek V4 Flash Vision Exp
    Chat

    DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731(opens in new tab) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total.

    Input:
    Output:
    Input:
    $0.44/1M
    Output:
    $1.32/1M
    Context length:
    1M
    Max output:
    393K
    1 provider
  • Google: Gemini 3.7 Flash
    15% offChat

    Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

    Input:
    Output:
    Input:
    $1.5$1.275/1M
    Output:
    $7.5$6.375/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • DeepSeek: DeepSeek V4 Pro 0813
    15% offChat

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

    Input:
    Output:
    Input:
    $1.32$1.122/1M
    Output:
    $3.96$3.366/1M
    Context length:
    1M
    Max output:
    393K
    2 providers
  • DeepSeek: DeepSeek V4 Flash 0731
    15% offChat

    DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

    Input:
    Output:
    Input:
    $0.44$0.374/1M
    Output:
    $1.32$1.122/1M
    Context length:
    1M
    Max output:
    393K
    2 providers
  • MiniMax: MiniMax H3
    Chat

    MiniMax: MiniMax H3 is a lightweight open source weighted video generation model developed by MiniMax. It is designed for precise multimodal editing and controlled content generation, including command-driven editing, text and brand rendering, and video-to-video motion transfer. The model is suitable for commercial creative workflows such as advertising, e-commerce, games and interface design, and supports native audiovisual output for reference-driven generation.

    Input:
    Output:
    Input:
    Output:
    $0.08-0.13/second
    Context length:
    13K
    Max output:
    1 provider
  • Anthropic: Claude Opus 5
    90% offChat

    Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents.

    Input:
    Output:
    Input:
    $5$0.5/1M
    Output:
    $25$2.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • Google: Gemini 3.5 Flash Lite
    15% offChat

    Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

    Input:
    Output:
    Input:
    $0.3$0.255/1M
    Output:
    $2.5$2.125/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • Google: Gemini 3.6 Flash
    15% offChat

    Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

    Input:
    Output:
    Input:
    $1.5$1.275/1M
    Output:
    $7.5$6.375/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • OpenAI: GPT-5.6 Luna
    50% offChat

    GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

    Input:
    Output:
    Input:
    $0.2$0.1/1M
    Output:
    $1.2$0.6/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • OpenAI: GPT-5.6 Sol
    95% offChat

    GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

    Input:
    Output:
    Input:
    $5$0.25/1M
    Output:
    $30$1.5/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • OpenAI: GPT-5.6 Terra
    95% offChat

    GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol.

    Input:
    Output:
    Input:
    $2$0.1/1M
    Output:
    $12$0.6/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • 15% offChat

    Google: Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) is Google's fastest and most cost-effective Gemini image model, built for high-frequency development pipelines and rapid visual exploration. It can complete text-to-image generation in about 5 seconds, which is about 2.7 times faster than Gemini 3.1 Flash Image, while maintaining the character consistency, precise editing capabilities and real-world knowledge of the Nano Banana series. A single embedded API handles text-to-image, image editing, and multi-image composition. As a multi-modal model, it can also generate text while returning images. Output resolution is 1K, supports 14 aspect ratios, and comes with an invisible SynthID watermark for easy identification as AI-generated. The best balance of quality and speed in the Nano Banana 2 series, it generates thousands of images at a fraction of the cost of heavyweight production models, making it ideal for prototyping, real-time applications and large-scale visual workflows.

    Input:
    Output:
    Input:
    $0.25$0.2125/1M
    Output:
    $1.5$1.275/1M
    Context length:
    66K
    Max output:
    33K
    1 provider
  • Anthropic: Claude Sonnet 5
    90% offChat

    Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can plan, use tools such as browsers and terminals, and operate autonomously at a level that only recently required larger and more expensive models.

    Input:
    Output:
    Input:
    $3$0.3/1M
    Output:
    $15$1.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • Alibaba: HappyHorse 1.1
    30% offChat

    Alibaba: HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames.

    Input:
    Output:
    Input:
    Output:
    $0.049-0.126/second
    Context length:
    Max output:
    1 provider
  • 15% offChat

    Google: Nano Banana 2 (Gemini 3.1 Flash Image) is designed for speed and efficiency, enabling fast, interactive responses and high throughput.

    Input:
    Output:
    Input:
    $0.5$0.425/1M
    Output:
    $3$2.55/1M
    Context length:
    66K
    Max output:
    33K
    1 provider
  • 15% offChat

    Google: Nano Banana Pro (Gemini 3 Pro Image) is an AI image generation model based on Google Gemini 3 Pro, and is a major upgrade of its previous products. It aims to go beyond simple pattern matching, turn to a more rationally driven system, and make improvements in terms of physical understanding, text rendering, and image consistency. Its main features include faster processing speed, native 2K resolution, and the ability to edit existing images with higher precision, aiming to provide more reliable, professional-level generated results.

    Input:
    Output:
    Input:
    $2$1.7/1M
    Output:
    $12$10.2/1M
    Context length:
    66K
    Max output:
    8K
    1 provider
  • Anthropic: Claude Fable 5
    30% offChat

    Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for long-running, complex, and asynchronous tasks that previously required frequent human check-ins.

    Input:
    Output:
    Input:
    $10$7/1M
    Output:
    $50$35/1M
    Context length:
    1M
    Max output:
    128K
    2 providers
  • MiniMax: MinMax M3
    50% offChat

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

    Input:
    Output:
    Input:
    $0.6$0.3/1M
    Output:
    $2.4$1.2/1M
    Context length:
    1M
    Max output:
    512K
    1 provider
  • Anthropic: Claude Opus 4.8
    90% offChat

    Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for highly autonomous agents, long-horizon agentic work, knowledge work, and memory-driven tasks where coherence over extended sessions matters.

    Input:
    Output:
    Input:
    $5$0.5/1M
    Output:
    $25$2.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • Google: Gemini 3.5 Flash
    15% offChat

    Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs.

    Input:
    Output:
    Input:
    $1.5$1.275/1M
    Output:
    $9$7.65/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • Alibaba: HappyHorse 1.0
    20% offChat

    Alibaba: HappyHorse 1.0 is described as an open source, state-of-the-art AI video generator with native audio and video co-generation capabilities - meaning that the alibaba: HappyHorse 1.0 video model can simultaneously generate video frames and corresponding audio tracks (dialogue, ambience, foley) in a single forward pass, rather than generating silent video first and then dubbing it later. According to an architectural description compiled by the community, the model is built around a 15 billion-parameter unified self-attention Transformer that can process text, image, video, and audio tokens within a single token sequence. It is reportedly built without a dedicated cross-attention branch and without a separate audio module. Combined with DMD-2 distillation technology, its distilled version is said to require only 8 denoising steps on the NVIDIA H100 and requires no classifier-free guidance to generate 1080p video in approximately 38 seconds.

    Input:
    Output:
    Input:
    Output:
    $0.112-0.192/second
    Context length:
    Max output:
    1 provider
  • OpenAI: GPT-5.5
    95% offChat

    GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling large-scale reasoning, coding, and multimodal workflows within a single system.

    Input:
    Output:
    Input:
    $5$0.25/1M
    Output:
    $30$1.5/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • OpenAI: GPT-Image-2
    20% offChat

    OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API.

    Input:
    Output:
    Input:
    $5$4/1M
    Output:
    Context length:
    10K
    Max output:
    1 provider
  • Anthropic: Claude Opus 4.7
    90% offChat

    Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable agentic execution across extended workflows. It is especially effective for asynchronous agent pipelines where tasks unfold over time - large codebases, multi-stage debugging, and end-to-end project orchestration.

    Input:
    Output:
    Input:
    $5$0.5/1M
    Output:
    $25$2.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • MiniMax: MinMax M2.7
    30% offChat

    MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.

    Input:
    Output:
    Input:
    $0.3$0.21/1M
    Output:
    $1.2$0.84/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.7 highspeed
    30% offChat

    MiniMax-M2.7-Highspeed is a high-speed variant of MiniMax-M2.7, delivering the same advanced agentic capabilities and model performance with significantly faster inference and greater responsiveness. Designed for autonomous, real-world productivity, it enables more agile planning, execution, and refinement of complex tasks across dynamic environments.

    Input:
    Output:
    Input:
    $0.6$0.42/1M
    Output:
    $2.4$1.68/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • OpenAI: GPT-5.4 Mini
    50% offChat

    GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments.

    Input:
    Output:
    Input:
    $0.75$0.375/1M
    Output:
    $4.5$2.25/1M
    Context length:
    400K
    Max output:
    128K
    2 providers
  • OpenAI: GPT-5.4 Nano
    20% offChat

    GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution.

    Input:
    Output:
    Input:
    $0.2$0.16/1M
    Output:
    $1.25$1/1M
    Context length:
    400K
    Max output:
    128K
    1 provider
  • OpenAI: GPT-5.4
    50% offChat

    GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, coding, and multimodal analysis within the same workflow.

    Input:
    Output:
    Input:
    $2.5$1.25/1M
    Output:
    $15$7.5/1M
    Context length:
    1.05M
    Max output:
    128K
    2 providers
  • Google: Gemini 3.1 Pro Preview
    15% offChat

    Gemini 3.1 Pro is the next generation in the Gemini series of models, a suite of highly-capable, natively multimodal, reasoning models. Gemini 3 Pro is now Google’s most advanced model for complex tasks, and can comprehend vast datasets, challenging problems from different information sources, including text, audio, images, video, and entire code repositories

    Input:
    Output:
    Input:
    $2$1.7/1M
    Output:
    $12$10.2/1M
    Context length:
    1.05M
    Max output:
    66K
    1 provider
  • Anthropic: Claude Sonnet 4.6
    90% offChat

    Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.

    Input:
    Output:
    Input:
    $3$0.3/1M
    Output:
    $15$1.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • MiniMax: MinMax M2.5
    20% offChat

    MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. Scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, M2.5 is also more token efficient than previous generations, having been trained to optimize its actions and output through planning.

    Input:
    Output:
    Input:
    $0.3$0.24/1M
    Output:
    $1.2$0.96/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.5 highspeed
    20% offChat

    MiniMax-M2.5-Highspeed is a high-speed variant of MiniMax-M2.5, delivering the same SOTA performance and real-world productivity capabilities with significantly faster inference and greater responsiveness. Trained across diverse and complex digital working environments, it extends the coding expertise of M2.1 into general office work, including generating and operating Word, Excel, and PowerPoint files, switching seamlessly across software environments, and collaborating across agent and human teams. With the same strong benchmark performance and token-efficient planning capabilities as M2.5, the Highspeed variant is optimized for faster, more agile execution in latency-sensitive workflows.

    Input:
    Output:
    Input:
    $0.6$0.48/1M
    Output:
    $2.4$1.92/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • Anthropic: Claude Opus 4.6
    90% offChat

    Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations.

    Input:
    Output:
    Input:
    $5$0.5/1M
    Output:
    $25$2.5/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
  • MiniMax: MinMax M2.1
    20% offChat

    MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

    Input:
    Output:
    Input:
    $0.3$0.24/1M
    Output:
    $1.2$0.96/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • MinMax: MiniMax M2.1 highspeed
    20% offChat

    MiniMax-M2.1-Highspeed is a high-speed variant of MiniMax-M2.1, delivering the same state-of-the-art capabilities in coding, agentic workflows, and modern application development with significantly faster inference and greater responsiveness. With only 10 billion activated parameters, it combines strong real-world performance with exceptional latency, scalability, and cost efficiency, making it especially well suited for interactive and latency-sensitive workloads.

    Input:
    Output:
    Input:
    $0.6$0.48/1M
    Output:
    $2.4$1.92/1M
    Context length:
    205K
    Max output:
    131K
    1 provider
  • OpenAI: GPT-Image-1.5
    20% offChat

    OpenAI’s GPT-Image-1.5 is the latest evolution of its AI image generation, offering superior command compliance, fidelity, text rendering and editing control, making it ideal for detailed creative and production work at a faster and lower cost than previous generations such as DALL-E 3. It excels at handling complex requests, maintains character/style consistency, renders clear text visually, and understands subtle cues through built-in reasoning. It has been integrated into ChatGPT and is available through the API.

    Input:
    Output:
    Input:
    $5$4/1M
    Output:
    $10$8/1M
    Context length:
    10K
    Max output:
    1 provider
  • Anthropic: Claude Haiku 4.5
    90% offChat

    Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications.

    Input:
    Output:
    Input:
    $1$0.1/1M
    Output:
    $5$0.5/1M
    Context length:
    200K
    Max output:
    64K
    3 providers
  • 15% offChat

    Gemini 2.5 Flash Image (also known as "Nano Banana") is now generally available. It is a cutting-edge image generation model with context understanding that supports image generation, editing, and multi-turn conversations.

    Input:
    Output:
    Input:
    $0.3$0.255/1M
    Output:
    $2.5$2.125/1M
    Context length:
    33K
    Max output:
    8K
    1 provider
  • OpenAI: GPT-4o
    20% offChat

    GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as fast and 50% more cost-effective. GPT-4o also offers improved performance in processing non-English languages and enhanced visual capabilities.

    Input:
    Output:
    Input:
    $2.5$2/1M
    Output:
    $10$8/1M
    Context length:
    128K
    Max output:
    16K
    1 provider