Wan-Dancer
Music-to-dance model generating minute-long 720p dance videos choreographed to a full track.
About
Feed it a song and a reference image of a dancer and Wan-Dancer choreographs a full performance: minute-plus 720p videos at 30fps whose movement stays locked to the beat well past the roughly 20-second wall where ordinary video diffusion loses coherence. The 14B model from Alibaba's Wan team plans hierarchically, first laying out global keyframes against the entire music track, then filling the motion between them with local temporal refinement, and it covers five trained styles: Chinese Classical, K-Pop, Street, Tap, and Latin. Code is Apache-2.0 with weights on Hugging Face and ModelScope plus sample images, music clips, and prompt files, and generation runs through provided shell scripts. Hardware expectations are serious; the reference configuration is eight 80GB A800 GPUs with CUDA 12.4, so this is datacenter or rented-GPU territory rather than a desktop toy. Released July 2026 with an accompanying paper, it stands as the strongest open option for long-form music-conditioned human animation.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026
Related Tools
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.
Free markerless motion capture system that works with ordinary cameras and no special hardware.
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
Audio-driven talking head animation from a single image.
Effective whole-body pose estimation with few-shot keypoint detection.