SkyReels-V3

Skywork's open-weight video models for reference-to-video, footage extension, and audio-driven avatars.

Open SourceSelf HostedOffline CapableGPU Required (24GB+ VRAM)
0.0 (0)

About

Skywork's third SkyReels generation splits video work across three open-weight pipelines: a 14B reference-to-video model that preserves the identity of 1 to 4 people or objects from supplied images, a 14B video extension model that continues existing footage with cinematographic transitions, and a 19B audio-driven talking avatar model, all generating at 720p. Inference code landed on GitHub in January 2026 with checkpoints auto-downloaded from Hugging Face or ModelScope; the stack wants Python 3.12 and CUDA 12.8, and cards under 24 GB can still run it by enabling FP8 weight-only quantization with block offload or dropping output to 540p or 480p. The weights ship under the Skywork Community License, which permits commercial use, and the same models back Skywork's paid API for teams that skip local hosting. Multi-subject consistency is the headline capability: reference conditioning makes this one of the few open options for keeping specific faces and products stable across generated shots, a task where prompt-only video models routinely drift.

Should you use SkyReels-V3?

Pick it when

Pick SkyReels-V3 when specific faces or products from 1 to 4 reference images must stay stable across shots, or you need audio-driven talking avatars at 720p, with a commercial-use license and Skywork's paid API as a fallback.

Look elsewhere when

Skip it for prompt-only generation, since all three pipelines start from references, footage, or audio; use SkyReels-V2 or Wan 2.1. Every model is 14B or 19B, so under 24 GB you rely on FP8, offload, or lower resolution.

Alternatives to SkyReels-V3

  • Phantom

    Apache-2.0 subject consistency with a 1.3B model for a single modest GPU, but no audio avatars or footage extension.

  • SkyReels-V2

    Same license family with prompt-driven text and image to video, unlimited length, and a 1.3B model near 15 GB, but no reference conditioning.

  • Ovi

    Generates speech and sound itself from the prompt under Apache-2.0, where V3 needs an audio track, but clips cap at 10 seconds.

  • MiniMax H3

    Reference-to-video with native stereo audio and speech in 11 languages, but 33B weights and a four GPU reference deployment.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Freemium
Platform
Hybrid
Difficulty
Advanced (4/5)
License
Skywork Community License
Minimum VRAM
24 GB
Added
Aug 24, 2026

Tags