Z-Image

6B single-stream diffusion transformer from Alibaba with a fast Turbo variant and bilingual text rendering.

Open SourceSelf HostedOffline CapableGPU Required (16GB+ VRAM)
0.0 (0)

About

Z-Image packs competitive text-to-image quality into a 6B-parameter single-stream diffusion transformer from Alibaba's Tongyi MAI team, whose S3-DiT design concatenates text, visual semantic, and image VAE tokens into one input sequence. The distilled Turbo variant generates in 8 function evaluations without classifier-free guidance, reaching sub-second latency on an H800 and fitting consumer GPUs with 16 GB of VRAM, while the undistilled foundation model runs at 28 to 50 steps with guidance scales of 3.0 to 5.0. Z-Image-Omni-Base covers combined generation and editing, and Z-Image-Edit handles instruction-following image edits. A signature strength is accurate rendering of complex Chinese and English text inside images, which the team reports as comparable to top-tier closed-source models. Everything is Apache-2.0, weights are on Hugging Face and ModelScope with online demos, and a technical report is on arXiv. The authors claim photorealism on par with models an order of magnitude larger, and the repository has climbed to roughly 11.8k stars.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
16 GB
Added
Jul 29, 2026

Related Tools

Featured

State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)
Featured

Next-generation image generation model by Black Forest Labs.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)

Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)

Zero-shot identity-preserving image generation from a single face photo.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Image prompt adapter for pre-trained text-to-image diffusion models.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Featured

Neural network architecture for adding spatial control to diffusion models.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Image Generation tools