LatentSync
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
About
LatentSync approaches lip sync as end-to-end generation in latent space: Whisper turns mel spectrograms into audio embeddings that condition a Stable Diffusion U-Net through cross-attention layers, so the model learns audio-visual correlations directly rather than going through the intermediate motion representations used by Wav2Lip-era tools. ByteDance has iterated on it publicly, with version 1.5 adding temporal layers, better performance on Chinese-language video, and inference in 8 GB of VRAM, and version 1.6 retraining at 512x512 resolution to fix output blurriness at the cost of needing 18 GB. Training scripts are included and require 20 to 55 GB of VRAM depending on stage and resolution. Setup is a single setup_env.sh script that installs dependencies and pulls checkpoints from Hugging Face, and the codebase builds on AnimateDiff with components borrowed from MuseTalk and StyleSync. Released under Apache-2.0 with about 5,900 GitHub stars, it is widely used for video dubbing, translation re-sync, and talking-head production in both photorealistic and anime styles.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 8 GB
- Added
- Jul 29, 2026
Related Tools
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.
Free markerless motion capture system that works with ordinary cameras and no special hardware.
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
Audio-driven talking head animation from a single image.
Real-time high-quality lip-sync model for audio-driven talking face generation.
Effective whole-body pose estimation with few-shot keypoint detection.