MiniMax M3
MiniMax's open-weight multimodal MoE with sparse attention and a 1M-token context window.
About
Announced June 1, 2026 with weights following on Hugging Face, MiniMax M3 folds frontier coding, million-token context, and native multimodality into one open checkpoint of about 428B total parameters with 23B active. Its defining piece is MiniMax Sparse Attention, an architecture that replaces full attention with learned sparse patterns and delivers a claimed 9x prefill and 15x decode speedup over the previous M2 at 1M-token lengths. The model was trained on mixed modalities from the first step, so image and video understanding are native rather than bolted on; published results include 59.0 on SWE-Bench Pro and 66.0 on Terminal-Bench 2.1. Local deployment works through vLLM, SGLang, Transformers, KTransformers, and Unsloth with recommended sampling of temperature 1.0 and top_p 0.95, though the size still demands a multi-GPU node. Weights are distributed under the MiniMax Community license, with the company's API and Token Plan subscriptions covering hosted use.
Should you use MiniMax M3?
Pick it when
Pick this when one self-hosted checkpoint has to cover agentic coding, image and video understanding, and million-token context, you have a multi-GPU node, and the MiniMax Community license terms work for your use.
Look elsewhere when
Skip it if you need a standard permissive license or single-GPU hardware. GLM-5 and DeepSeek-V4 are MIT, and for vision on one consumer card, Qwen3.8's Apache-2.0 27B or Muse Glimmer fit where M3 cannot.
Alternatives to MiniMax M3
- MiniMax-M1
Apache-2.0 predecessor with 1M-token context and 40K or 80K thinking budgets, but no image input and about twice the active parameters.
- GLM-5
MIT licensed with broad quantized community builds, but built around 200K contexts rather than 1M and focused on code and agents, not vision.
- Kimi K3
Also 1M context with image input and agentic coding focus, but 2.8T total and about 104B active needs a much larger cluster.
- Qwen3.8
Its Apache-2.0 27B reads images and video on one consumer GPU at 4-bit, trading M3's scale and long-context design for cheap hardware.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- MiniMax Community License
- Added
- Aug 24, 2026