HunyuanVideo-Avatar
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
About
Part of Tencent's HunyuanVideo model family, HunyuanVideo-Avatar turns a character image and an audio track into a talking, emoting video using a multimodal diffusion transformer (MM-DiT). Three modules define the approach: a character image injection mechanism that keeps the avatar consistent across frames, an Audio Emotion Module that transfers emotional cues from a reference image, and a Face-Aware Audio Adapter that isolates audio per face so multiple characters in one scene can speak independently. It handles photorealistic, cartoon, 3D-rendered, and anthropomorphic styles at portrait through full-body scales, targeting e-commerce, streaming, and video editing. Inference officially calls for an NVIDIA GPU with 24 GB of VRAM and 96 GB for best quality, though a June 2025 TeaCache update runs on a single 10 GB card; setup is a conda environment with Python 3.10 and PyTorch 2.4, with Docker images available. Weights ship under the Tencent Hunyuan Community License, which excludes the EU, UK, and South Korea and requires a separate agreement above 100 million monthly users. The repository has about 2,100 stars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Tencent Hunyuan Community License
- Minimum VRAM
- 10 GB
- Added
- Jul 29, 2026
Related Tools
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.
Free markerless motion capture system that works with ordinary cameras and no special hardware.
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
Audio-driven talking head animation from a single image.
Real-time high-quality lip-sync model for audio-driven talking face generation.
Effective whole-body pose estimation with few-shot keypoint detection.