MimicMotion

Tencent's pose-guided model for generating long, smooth human motion videos from a single image.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

A reference photo plus a pose sequence is all MimicMotion needs to produce a video of that person performing the motion, and its ICML 2025 paper focuses on making such videos long and stable. Confidence-aware pose guidance weights the conditioning by pose-estimation confidence, cutting the distortion that creeps in when keypoints are unreliable, regional loss amplification reduces artifacts in high-confidence regions such as hands, and progressive latent fusion stitches overlapping segments so clips can run to arbitrary length, with segments of up to 72 frames at 576x1024 resolution. The implementation builds on Stable Video Diffusion with DWPose for pose extraction; setup is a conda environment from the provided environment.yaml plus weight downloads from Hugging Face, and generation fits in 8 GB of VRAM for the 16-frame model, with 16 GB recommended for the VAE decoder (a 35-second demo takes about 20 minutes on an RTX 4090). Code is Apache-2.0 licensed with about 2,600 stars, and the model is a staple of ComfyUI dance and motion-transfer workflows.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
8 GB
Added
Jul 29, 2026

Related Tools

Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.

Open SourceSelf HostedOfflineGPU 20GB+
Advanced
0.0 (0)

Free markerless motion capture system that works with ordinary cameras and no special hardware.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.

Open SourceSelf HostedOfflineGPU 10GB+
Advanced
0.0 (0)

ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Featured

Audio-driven talking head animation from a single image.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Effective whole-body pose estimation with few-shot keypoint detection.

Open SourceSelf HostedOfflineGPU 4GB+
Easy
0.0 (0)
Browse all AI Animation & Motion tools