Z-Image
6B single-stream diffusion transformer from Alibaba with a fast Turbo variant and bilingual text rendering.
About
Z-Image packs competitive text-to-image quality into a 6B-parameter single-stream diffusion transformer from Alibaba's Tongyi MAI team, whose S3-DiT design concatenates text, visual semantic, and image VAE tokens into one input sequence. The distilled Turbo variant generates in 8 function evaluations without classifier-free guidance, reaching sub-second latency on an H800 and fitting consumer GPUs with 16 GB of VRAM, while the undistilled foundation model runs at 28 to 50 steps with guidance scales of 3.0 to 5.0. Z-Image-Omni-Base covers combined generation and editing, and Z-Image-Edit handles instruction-following image edits. A signature strength is accurate rendering of complex Chinese and English text inside images, which the team reports as comparable to top-tier closed-source models. Everything is Apache-2.0, weights are on Hugging Face and ModelScope with online demos, and a technical report is on arXiv. The authors claim photorealism on par with models an order of magnitude larger, and the repository has climbed to roughly 11.8k stars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 16 GB
- Added
- Jul 29, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.