Allegro
Apache-licensed diffusion transformer that generates 6-second 720p videos from text prompts.
About
Six seconds of 720p video from a text prompt is what Allegro produces, an open text-to-video release from Rhymes AI built as a diffusion transformer with a 175M-parameter VAE and a 2.8B-parameter DiT. Output is 88 frames at 15 fps in 720x1280 resolution, and the Allegro-TI2V variant adds image-to-video generation conditioned on a first frame and an optional last frame, alongside smaller 40-frame research checkpoints at 720p and 360p. Generation is not fast: around 20 minutes for one clip on a single H100, dropping to roughly 3 minutes across eight H100s, and VRAM needs are 9.3 GB with CPU offloading enabled or 27.5 GB without. Setup requires Python 3.10 or newer, PyTorch 2.4 or newer, and CUDA 12.4 or newer, installed from the GitHub repository with weights on Hugging Face. Both code and weights are Apache-2.0, permitting commercial use, which distinguished it among the fully open video model releases of late 2024. The technical report is on arXiv (2410.15458), and a public gallery shows sample generations.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 10 GB
- Added
- Jul 29, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.