HunyuanWorld 1.0
Generates explorable, mesh-exportable 3D worlds from text or image prompts via panoramic proxies.
About
Text or a single image goes in and an explorable, mesh-exportable 3D world comes out: HunyuanWorld 1.0 was the first open-weight 3D world generation model when Tencent released it in July 2025. The pipeline works in three stages, generating a 360 degree panoramic world proxy with Flux-based diffusion models and LoRA adapters, splitting it into semantic layers, then reconstructing a hierarchical 3D scene whose object layers stay disentangled for interactivity, with export to GLB and Draco mesh formats. The standard pipeline targets A100 or H100 class hardware, though an August 2025 lite release with quantization runs on consumer cards like the RTX 4090, and setup involves conda plus heavyweight dependencies including Real-ESRGAN, ZIM, and GroundingDINO, with a Hugging Face login needed for weight downloads. Follow-up releases added the Voyager RGBD video model, WorldMirror video support, and real-time WorldPlay generation. Distribution falls under the Tencent HunyuanWorld community license: commercial use is allowed below 1 million monthly active users, but the EU, UK, and South Korea are excluded territories.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- 3D Model Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Tencent HunyuanWorld-1.0 Community License
- Minimum VRAM
- 24 GB
- Added
- Jul 29, 2026
Related Tools
CUDA-accelerated Gaussian splatting rasterization library with Python bindings from the Nerfstudio team.
Generates high-resolution textured 3D assets from images or text with separate shape and texture models.
Roblox's research foundation model that generates 3D shapes from text prompts for game assets.
Reconstructs 3D shape, texture, and layout of objects from a single cluttered real-world photo.
Turns a single image into a textured, UV-unwrapped 3D mesh in under a second.
Reconstructs dense 3D scenes from uncalibrated image collections without known camera poses.