Ovi

Twin-backbone model generating short video clips with synchronized speech and sound in one pass.

Open SourceSelf HostedOffline CapableGPU Required (24GB+ VRAM)
0.0 (0)

About

Ovi answers the Veo 3 question for local hardware: one open model producing video and its soundtrack together. Character AI's release runs an 11B twin-backbone system in which a video branch and a 5B audio branch exchange information during generation, so speech lands on lip movements and effects land on events without a separate foley pass. Clips run 5 or 10 seconds at 24fps in aspect ratios up to a 960x960 pixel budget since the v1.1 update of November 2025, driven by text alone, text plus image, or image-to-video, with spoken lines marked inline in the prompt. Everything is Apache-2.0: weights in safetensors including an FP8 variant, single and multi-GPU inference code, and a Gradio UI. Hardware needs come to 32GB VRAM with CPU offload, dropping to 24GB using FP8 quantization, while 80GB cards run unconstrained; community ComfyUI support arrived through the WanVideoWrapper nodes. For talking characters and sound-on shots generated without cloud APIs, it is currently the leading open option, and the repo stays active with regular updates.

Should you use Ovi?

Pick it when

Pick Ovi when you need talking characters or sound-on shots generated locally, with speech marked in the prompt landing on lip movements, Apache-2.0 terms for commercial work, and a 24 GB card for the FP8 build.

Look elsewhere when

Skip it for clips over 10 seconds, output beyond a 960x960 pixel budget, or multi-shot scenes with a recurring character, where LTX-2.5 fits better. For silent footage, Wan 2.1's 1.3B model needs only about 8 GB.

Alternatives to Ovi

  • LTX-2.5

    Larger 22B audio-video model with multishot consistency and NVFP4 quantization, but its community license charges organizations above 10 million dollars in revenue.

  • MiniMax H3

    Native 32 kHz stereo, speech in 11 languages, and clips up to 15 seconds, but the reference deployment spans four GPUs and weights use a community license.

  • SkyReels-V3

    Animates an avatar from audio you supply rather than generating the speech itself, at 720p, under a commercial Skywork license with 24 GB or FP8 needs.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
24 GB
Added
Aug 24, 2026

Tags