Mage-Flow
Microsoft's 4B text-to-image and instruction-editing model using rectified flow at native resolution.
About
Microsoft's Mage project holds every model to a fixed 4B parameter budget, and Mage-Flow is its text-to-image and instruction-based editing member, pairing the Mage-VAE tokenizer with a native-resolution multimodal diffusion transformer, NR-MMDiT, trained as a rectified flow. Despite the small budget, the RL-aligned release posts 0.90 on GenEval, ahead of open models five to eight times larger, including FLUX.2-dev at 32B and Qwen-Image at 20B, which both score 0.87. A distilled 4-step Turbo variant renders a 1024px image in about 0.6 seconds on a single A100, and dedicated editing checkpoints handle instruction-driven modification of existing images. Base, RL-aligned, Turbo, and editing weights all sit on Hugging Face under the microsoft organization, released in July 2026 alongside code, with MIT licensing on Mage-Flow leaving commercial use unrestricted. The compact size keeps inference and fine-tuning within a single workstation-class GPU, exactly the niche it targets against far heavier open image models.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MIT
- Added
- Aug 24, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.