Alibaba: HappyHorse 1.0 is described as an open source, state-of-the-art AI video generator with native audio and video co-generation capabilities - meaning that the alibaba: HappyHorse 1.0 video model can simultaneously generate video frames and corresponding audio tracks (dialogue, ambience, foley) in a single forward pass, rather than generating silent video first and then dubbing it later. According to an architectural description compiled by the community, the model is built around a 15 billion-parameter unified self-attention Transformer that can process text, image, video, and audio tokens within a single token sequence. It is reportedly built without a dedicated cross-attention branch and without a separate audio module. Combined with DMD-2 distillation technology, its distilled version is said to require only 8 denoising steps on the NVIDIA H100 and requires no classifier-free guidance to generate 1080p video in approximately 38 seconds.
- By:
- Alibaba
- Input:
- Output:
- Published:
- 2026-04-26
Providers
Same model, different providers. Automatic routing picks the cheapest healthy one within the same quality tier.
| Provider | ||||
|---|---|---|---|---|
| 20% offalibaba-cloud | — | 720P$0.14$0.112/second1080P$0.24$0.192/second | Read—Write— | — |
Availability
24 hoursSuccess rate of the requests Amux actually sent to each provider. Client errors are excluded from the denominator — a malformed request isn't the provider's fault.
Activity
24 hoursToken usage and cost for this model over time, split by the provider that served each request.
Split by the provider that served each request. Amounts match what is actually billed — calls that failed upstream are not charged, so they are excluded.
Related models
More models from Alibaba