Kimi K3
Moonshot AI's 2.8T-parameter open-weight MoE model with 1M context and native image understanding.
About
Moonshot AI pushed open-weight scale to a new ceiling in July 2026 with Kimi K3, a 2.8 trillion parameter mixture-of-experts model that routes each token through 16 of 896 experts for roughly 104B active parameters. The design builds on Kimi Delta Attention, which Moonshot credits with sharply cheaper long-context inference, and a 401M parameter MoonViT-V2 encoder gives it native image understanding alongside text. Context stretches to 1,048,576 tokens, and a configurable thinking-effort control trades latency against deeper reasoning, which pays off in agentic coding and long-horizon knowledge work where K3 posts state-of-the-art open results. Weights ship in MXFP4 with MXFP8 activations for broad hardware compatibility, and recommended serving stacks are vLLM, SGLang, and TokenSpeed; a multi-GPU cluster is still required, so most teams will reach it through Moonshot's OpenAI and Anthropic compatible API endpoints instead. The Kimi K3 License, a modified MIT, keeps commercial use permissive for nearly all deployments.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Kimi K3 License
- Added
- Aug 24, 2026
Related Tools
Lightweight open-weight LLM by Google available in 1B to 27B sizes.
Open-source code LLM family by IBM for enterprise code generation.
Open-weight LLM by Meta in 8B and 70B sizes with strong general capabilities.
High-performance open-weight MoE LLM with 671B total parameters.
Hybrid SSM-Transformer model by AI21 Labs combining Mamba with attention layers.
Open-weight code LLM trained on 2 trillion tokens of code and natural language.