Amux
All models

Alibaba Cloud

Models served by Alibaba Cloud

Models · 6

  • Alibaba: HappyHorse 1.0
    20% offChat

    Alibaba: HappyHorse 1.0 is described as an open source, state-of-the-art AI video generator with native audio and video co-generation capabilities - meaning that the alibaba: HappyHorse 1.0 video model can simultaneously generate video frames and corresponding audio tracks (dialogue, ambience, foley) in a single forward pass, rather than generating silent video first and then dubbing it later. According to an architectural description compiled by the community, the model is built around a 15 billion-parameter unified self-attention Transformer that can process text, image, video, and audio tokens within a single token sequence. It is reportedly built without a dedicated cross-attention branch and without a separate audio module. Combined with DMD-2 distillation technology, its distilled version is said to require only 8 denoising steps on the NVIDIA H100 and requires no classifier-free guidance to generate 1080p video in approximately 38 seconds.

    Input:
    Output:
    Input:
    Output:
    $0.112-0.192/second
    Context length:
    Max output:
    1 provider
  • Alibaba: HappyHorse 1.1
    30% offChat

    Alibaba: HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames.

    Input:
    Output:
    Input:
    Output:
    $0.049-0.126/second
    Context length:
    Max output:
    1 provider
  • Alibaba: Wan3.0-Video
    20% offChat

    Alibaba: Wan3.0-Video is a multi-functional reference video generation model that supports text-to-video, image-to-video (first frame/first and last frame) and reference-based video generation. It is capable of generating videos up to 30 seconds long at 30fps.

    Input:
    Output:
    Input:
    Output:
    $0.04-0.16/second
    Context length:
    20K
    Max output:
    1 provider
  • Alibaba: Wan3.0-Video-Prime
    15% offChat

    Alibaba: Wan3.0-Video-Prime is a high-speed version of the Wan3.0 video generation model. It has the same powerful capabilities as the standard Wan3.0-Video model, supports comprehensive reference input of four modalities, can generate videos up to 30 seconds, and significantly improves the end-to-end generation speed while providing an immersive audio-visual experience.

    Input:
    Output:
    Input:
    Output:
    $0.0578-0.238/second
    Context length:
    20K
    Max output:
    1 provider
  • DeepSeek: DeepSeek V4 Flash 0731
    15% offChat

    DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

    Input:
    Output:
    Input:
    $0.22/1M
    Output:
    $0.66/1M
    Context length:
    1M
    Max output:
    393K
    2 providers
  • DeepSeek: DeepSeek V4 Pro 0813
    15% offChat

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

    Input:
    Output:
    Input:
    $0.66/1M
    Output:
    $1.98/1M
    Context length:
    1M
    Max output:
    393K
    2 providers