Amux
All models

Alibaba: Wan3.0-Video

Active

Alibaba: Wan3.0-Video is a multi-functional reference video generation model that supports text-to-video, image-to-video (first frame/first and last frame) and reference-based video generation. It is capable of generating videos up to 30 seconds long at 30fps.

Input:
Output:
Context length:
20K
Published:
2026-08-24
vision

Providers

Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.

Provider
20% offalibaba-cloud
480P$0.05$0.04/second720P$0.1$0.08/second1080P$0.2$0.16/second
ReadWrite
20K

Availability

24 hours

Success rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.

Activity

24 hours

Token usage and cost for this model over time, split by the provider that served each request.

Split by the provider that served each request. Amounts match what is actually billed — calls that failed upstream are not charged, so they are excluded.

Related models

More models from Alibaba

Alibaba: Wan3.0-Video-Prime20K context1 providerAlibaba: Wan3.0-Video-Prime is a high-speed version of the Wan3.0 video generation model. It has the same powerful capabilities as the standard Wan3.0-Video model, supports comprehensive reference input of four modalities, can generate videos up to 30 seconds, and significantly improves the end-to-end generation speed while providing an immersive audio-visual experience.
Alibaba: HappyHorse 1.11 providerAlibaba: HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up to 1080p and durations of 3 to 15 seconds. It is suited for creative content, social media clips, and image-driven animation, and improves on the prior version with stronger prompt adherence, smoother motion, and more consistent characters across frames.
Alibaba: HappyHorse 1.01 providerAlibaba: HappyHorse 1.0 is described as an open source, state-of-the-art AI video generator with native audio and video co-generation capabilities - meaning that the alibaba: HappyHorse 1.0 video model can simultaneously generate video frames and corresponding audio tracks (dialogue, ambience, foley) in a single forward pass, rather than generating silent video first and then dubbing it later. According to an architectural description compiled by the community, the model is built around a 15 billion-parameter unified self-attention Transformer that can process text, image, video, and audio tokens within a single token sequence. It is reportedly built without a dedicated cross-attention branch and without a separate audio module. Combined with DMD-2 distillation technology, its distilled version is said to require only 8 denoising steps on the NVIDIA H100 and requires no classifier-free guidance to generate 1080p video in approximately 38 seconds.