Kimi K3
Moonshot AI's 2.8T-parameter open-weight MoE model with 1M context and native image understanding.
About
Moonshot AI pushed open-weight scale to a new ceiling in July 2026 with Kimi K3, a 2.8 trillion parameter mixture-of-experts model that routes each token through 16 of 896 experts for roughly 104B active parameters. The design builds on Kimi Delta Attention, which Moonshot credits with sharply cheaper long-context inference, and a 401M parameter MoonViT-V2 encoder gives it native image understanding alongside text. Context stretches to 1,048,576 tokens, and a configurable thinking-effort control trades latency against deeper reasoning, which pays off in agentic coding and long-horizon knowledge work where K3 posts state-of-the-art open results. Weights ship in MXFP4 with MXFP8 activations for broad hardware compatibility, and recommended serving stacks are vLLM, SGLang, and TokenSpeed; a multi-GPU cluster is still required, so most teams will reach it through Moonshot's OpenAI and Anthropic compatible API endpoints instead. The Kimi K3 License, a modified MIT, keeps commercial use permissive for nearly all deployments.
Should you use Kimi K3?
Pick it when
Pick this for agentic coding and long-horizon knowledge work over million-token contexts with image input, through Moonshot's compatible APIs or a multi-GPU cluster serving MXFP4 weights in vLLM or SGLang.
Look elsewhere when
Skip it if you need to self-host on modest hardware or keep serving cost low: about 104B parameters activate per token. Kimi K2 is cheaper per token at 32B active, and MiniMax M3 covers 1M context with 23B active.
Alternatives to Kimi K3
- Kimi K2
Predecessor with 32B active and agentic tool use, cheaper per token, but limited to 128K context and text input.
- Qwen3.8
Qwen3.8's 2.4T flagship has 95B active and 262K native context but is text-only and thinking-only; its 27B sibling covers single-GPU use.
- MiniMax M3
About 428B total and 23B active with native image and video and 1M context, far cheaper to serve, under the MiniMax Community license.
- Inkling
Apache-2.0 975B MoE with text, image, and audio input and up to 1M context self-hosted, framed as a fine-tuning base, not a finished model.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Kimi K3 License
- Added
- Aug 24, 2026