LatentSync

ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

LatentSync approaches lip sync as end-to-end generation in latent space: Whisper turns mel spectrograms into audio embeddings that condition a Stable Diffusion U-Net through cross-attention layers, so the model learns audio-visual correlations directly rather than going through the intermediate motion representations used by Wav2Lip-era tools. ByteDance has iterated on it publicly, with version 1.5 adding temporal layers, better performance on Chinese-language video, and inference in 8 GB of VRAM, and version 1.6 retraining at 512x512 resolution to fix output blurriness at the cost of needing 18 GB. Training scripts are included and require 20 to 55 GB of VRAM depending on stage and resolution. Setup is a single setup_env.sh script that installs dependencies and pulls checkpoints from Hugging Face, and the codebase builds on AnimateDiff with components borrowed from MuseTalk and StyleSync. Released under Apache-2.0 with about 5,900 GitHub stars, it is widely used for video dubbing, translation re-sync, and talking-head production in both photorealistic and anime styles.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
8 GB
Added
Jul 29, 2026

Related Tools

Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.

Open SourceSelf HostedOfflineGPU 20GB+
Advanced
0.0 (0)

Free markerless motion capture system that works with ordinary cameras and no special hardware.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.

Open SourceSelf HostedOfflineGPU 10GB+
Advanced
0.0 (0)
Featured

Audio-driven talking head animation from a single image.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Real-time high-quality lip-sync model for audio-driven talking face generation.

Open SourceSelf HostedOfflineGPU 6GB+
Intermediate
0.0 (0)

Effective whole-body pose estimation with few-shot keypoint detection.

Open SourceSelf HostedOfflineGPU 4GB+
Easy
0.0 (0)
Browse all AI Animation & Motion tools