SkyReels-V2

Diffusion-forcing video model generating effectively unlimited-length clips at 540p and 720p.

Open SourceSelf HostedOffline CapableGPU Required (15GB+ VRAM)
0.0 (0)

About

Billed as the first infinite-length film generative model, SkyReels-V2 from Skywork AI uses a Diffusion Forcing framework that generates video autoregressively, extending clips indefinitely instead of stopping at a fixed frame budget. Released open weights come in 1.3B and 14B sizes, with a 5B series still listed as coming soon, covering text-to-video, image-to-video, video extension, and camera control at 540p (544x960, 97 frames per segment) or 720p (720x1280, 121 frames), and the release includes the SkyCaptioner-V1 video captioning model used to build its training data. Peak VRAM runs from about 14.7 GB for the 1.3B model at 540p up to 51.2 GB for the 14B with diffusion forcing, with xDiT USP multi-GPU inference and Diffusers integration available. Installation is pip install -r requirements.txt on Python 3.10. The models ship under the custom Skywork Community License, which permits commercial use subject to conditions rather than being a standard OSI license. With around 7,300 GitHub stars it is a popular base for long-form open video work, and Skywork shipped the successor SkyReels-V3 in January 2026.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Skywork Community License
Minimum VRAM
15 GB
Added
Jul 29, 2026

Related Tools

Featured

Open-source video generation model by Tencent with text and image conditioning.

Open SourceSelf HostedOfflineGPU 24GB+
Advanced
0.0 (0)

Image-to-video generation model by Alibaba DAMO Academy.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)

Updated CogVideo model by Zhipu AI with improved video quality.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced
0.0 (0)

Infinite-length music-driven video generation with visual conditioning.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)

Text-to-video generation framework with cascaded latent diffusion.

Open SourceSelf HostedOfflineGPU 16GB+
Advanced
0.0 (0)

Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.

Open SourceSelf HostedOfflineGPU 12GB+
Advanced
0.0 (0)
Browse all Video Generation tools