SenseNova U1.5
SenseTime's unified 8B model for image understanding, native 4K generation, and multi-image editing.
About
SenseTime's NEO-unify architecture drops the bolted-together design typical of multimodal models: SenseNova U1.5 processes vision and language end to end without a separate visual encoder or VAE, so one checkpoint handles visual question answering, native 4K text-to-image generation, multi-image editing, and interleaved text-image output. The flagship U1.5-8B-MoT pairs roughly 8B understanding parameters with 8B generation parameters, and editing accepts bounding-box and reference-image controls for precise regional changes across several input images. An H100 renders a 2048px image in about 9 seconds; with Q4 quantization and balanced offloading the model squeezes into 10 to 12 GB of VRAM, though 24 GB or more is recommended, and 8-step LoRA variants trade quality for speed. The August 2026 release is Apache 2.0 across GitHub, Hugging Face, and ModelScope, an unusually permissive grant for a unified generation model of this strength, following the original U1 in April and a U1.5 preview in July, with around 5.5k GitHub stars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 12 GB
- Added
- Aug 24, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.