Ovi
Twin-backbone model generating short video clips with synchronized speech and sound in one pass.
About
Ovi answers the Veo 3 question for local hardware: one open model producing video and its soundtrack together. Character AI's release runs an 11B twin-backbone system in which a video branch and a 5B audio branch exchange information during generation, so speech lands on lip movements and effects land on events without a separate foley pass. Clips run 5 or 10 seconds at 24fps in aspect ratios up to a 960x960 pixel budget since the v1.1 update of November 2025, driven by text alone, text plus image, or image-to-video, with spoken lines marked inline in the prompt. Everything is Apache-2.0: weights in safetensors including an FP8 variant, single and multi-GPU inference code, and a Gradio UI. Hardware needs come to 32GB VRAM with CPU offload, dropping to 24GB using FP8 quantization, while 80GB cards run unconstrained; community ComfyUI support arrived through the WanVideoWrapper nodes. For talking characters and sound-on shots generated without cloud APIs, it is currently the leading open option, and the repo stays active with regular updates.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 24 GB
- Added
- Aug 24, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.