HunyuanImage-3.0
Tencent's 80B MoE autoregressive model, the largest open-weights text-to-image release.
About
At 80 billion total parameters with 13 billion activated per token, HunyuanImage-3.0 is the largest open-weights text-to-image model released to date. Tencent built it as a native multimodal Mixture-of-Experts model with 64 experts that unifies understanding and generation in a single autoregressive framework rather than the usual DiT plus text encoder split, which lets it perform reasoning-based prompt expansion, image editing, and multi-image fusion alongside straight generation. The hardware bill is steep: the README calls for at least three 80 GB GPUs to run the base model and eight for the Instruct variants, on CUDA 12.8 with PyTorch 2.8, though a distilled Instruct release and vLLM acceleration support arrived in January 2026 to ease deployment. Weights are on Hugging Face under the Tencent Hunyuan Community License, which allows commercial use for services below 100 million monthly active users but excludes the EU, UK, and South Korea from its territory. Open-sourced in September 2025 with a technical report on arXiv, the repository has around 3,200 stars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Tencent Hunyuan Community License
- Minimum VRAM
- 240 GB
- Added
- Jul 29, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.