Alibaba: HappyHorse 1.0 is described as an open source, state-of-the-art AI video generator with native audio and video co-generation capabilities - meaning that the alibaba: HappyHorse 1.0 video model can simultaneously generate video frames and corresponding audio tracks (dialogue, ambience, foley) in a single forward pass, rather than generating silent video first and then dubbing it later. According to an architectural description compiled by the community, the model is built around a 15 billion-parameter unified self-attention Transformer that can process text, image, video, and audio tokens within a single token sequence. It is reportedly built without a dedicated cross-attention branch and without a separate audio module. Combined with DMD-2 distillation technology, its distilled version is said to require only 8 denoising steps on the NVIDIA H100 and requires no classifier-free guidance to generate 1080p video in approximately 38 seconds.
- Input:
- Output:
- Input:
- —
- Output:
- $0.112-0.192/second
- Context length:
- —
- Max output:
- —