Tools/Video Generation/Kandinsky 5.0

Kandinsky 5.0

Open family of flow-matching video and image diffusion models spanning 2B fast to 19B quality tiers.

Open SourceSelf HostedOffline CapableGPU Required (12GB+ VRAM)
0.0 (0)

About

Kandinsky 5.0 covers the whole local video stack in one family: Video Lite 2B checkpoints for 5 and 10 second clips at 768x512, Video Pro at 19B pushing quality up to 1920x1080, plus a 6B Image Lite text-to-image model and a 6B editing variant, all latent diffusion transformers trained with flow matching. Kandinsky Lab published unusually complete assets, including pretrain checkpoints alongside SFT, CFG-free, and distilled versions of the Lite models, so researchers can build on intermediate stages rather than only polished final weights. Speed spans the range too, from 13-second image generation to around 35 seconds for a distilled Lite clip and about 20 minutes for a 5 second 1080p Pro clip on an H100. Everything is MIT licensed with code on GitHub, weights on Hugging Face, Diffusers support, and ComfyUI integration; offloading and quantization options stretch requirements down to roughly 12GB VRAM for the smaller models. The staged release ran September through December 2025 and remains among the most open video model programs anywhere.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Minimum VRAM
12 GB
Added
Aug 24, 2026

Related Tools

Featured

Open-source video generation model by Tencent with text and image conditioning.

Open SourceSelf HostedOfflineGPU 24GB+
Advanced

Image-to-video generation model by Alibaba DAMO Academy.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced

Updated CogVideo model by Zhipu AI with improved video quality.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced

Infinite-length music-driven video generation with visual conditioning.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced

Text-to-video generation framework with cascaded latent diffusion.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced

Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
Browse all Video Generation tools