Amux
All models

Tencent Cloud

Models served by Tencent Cloud

Models · 8

  • MiniMax: MinMax M3
    50% offChat

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

    Input:
    Output:
    Input:
    $0.6$0.3/1M
    Output:
    $2.4$1.2/1M
    Context length:
    1M
    Max output:
    512K
    2 providers
    10/08, 16:00 – 10/11, 16:00—
  • MoonshotAI: Kimi K3
    70% offChat

    MoonshotAI: Kimi K3 is Kimi’s most powerful model to date, with 2.8 trillion parameters. It is built based on the hybrid linear attention mechanism Kimi Delta Attention and attention residual technology. It has native visual understanding capabilities and a context window of 1 million tokens. It is suitable for cutting-edge intelligent scenarios such as software engineering, knowledge work, and deep reasoning.

    Input:
    Output:
    Input:
    $3$0.9/1M
    Output:
    $15$4.5/1M
    Context length:
    1.05M
    Max output:
    1.05M
    4 providers
    10/08, 16:00 – 10/11, 16:00100.00%
  • Tencent: Hy3
    35% offChat

    Tencent: Hy3 is a hybrid expert model launched by Tencent with 770 billion total parameters, of which 49 billion are activation parameters. It is designed for coding agents, complex tool usage workflows, and productivity tasks that require planning, contextual continuity, and sustained multi-step execution.

    Input:
    Output:
    Input:
    $0.132$0.0858/1M
    Output:
    $0.528$0.3432/1M
    Context length:
    262K
    Max output:
    131K
    1 provider
    10/08, 16:00 – 10/11, 16:00—
  • Tencent: Hy4 Preview
    35% offChat

    Tencent: Hy4 Preview is a hybrid expert (MoE) model launched by Tencent, with a total parameter size of 770B and an activation parameter size of 49B. It is designed for coding agents, complex tool usage workflows, and productivity tasks that require planning, contextual continuity, and sustained multi-step execution.

    Input:
    Output:
    Input:
    $0.834$0.5421/1M
    Output:
    $2.501$1.6257/1M
    Context length:
    1.05M
    Max output:
    64K
    1 provider
    10/08, 16:00 – 10/11, 16:00—
  • Z.ai: GLM 5.2
    70% offChat

    GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

    Input:
    Output:
    Input:
    $1.4$0.42/1M
    Output:
    $4.4$1.32/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
    10/08, 16:00 – 10/11, 16:00—
  • Z.ai: GLM 5.3
    70% offChat

    GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

    Input:
    Output:
    Input:
    $1.4$0.42/1M
    Output:
    $4.4$1.32/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
    10/08, 16:00 – 10/11, 16:00100.00%
  • Z.ai: GLM 5.3 Flash
    70% offChat

    GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

    Input:
    Output:
    Input:
    $0.15$0.045/1M
    Output:
    $0.5$0.15/1M
    Context length:
    1M
    Max output:
    128K
    2 providers
    10/08, 16:00 – 10/11, 16:0090.79%
  • Z.ai: GLM 5.3 FlashX
    70% offChat

    Z.AI: GLM 5.3 FlashX is a faster, smoother version of GLM-5.3-Flash with inference speeds up to 200 tokens/s. Z.AI: GLM 5.3 FlashX is Z.ai’s native multi-modal model, suitable for efficient programming and long-range agent tasks. Its hybrid sparse and linear attention architecture reduces computational overhead while maintaining accurate long-context performance.

    Input:
    Output:
    Input:
    $0.375$0.1125/1M
    Output:
    $1.25$0.375/1M
    Context length:
    1M
    Max output:
    128K
    3 providers
    10/08, 16:00 – 10/11, 16:00100.00%