Kimi K3

Moonshot AI's 2.8T-parameter open-weight MoE model with 1M context and native image understanding.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

Moonshot AI pushed open-weight scale to a new ceiling in July 2026 with Kimi K3, a 2.8 trillion parameter mixture-of-experts model that routes each token through 16 of 896 experts for roughly 104B active parameters. The design builds on Kimi Delta Attention, which Moonshot credits with sharply cheaper long-context inference, and a 401M parameter MoonViT-V2 encoder gives it native image understanding alongside text. Context stretches to 1,048,576 tokens, and a configurable thinking-effort control trades latency against deeper reasoning, which pays off in agentic coding and long-horizon knowledge work where K3 posts state-of-the-art open results. Weights ship in MXFP4 with MXFP8 activations for broad hardware compatibility, and recommended serving stacks are vLLM, SGLang, and TokenSpeed; a multi-GPU cluster is still required, so most teams will reach it through Moonshot's OpenAI and Anthropic compatible API endpoints instead. The Kimi K3 License, a modified MIT, keeps commercial use permissive for nearly all deployments.

Should you use Kimi K3?

Pick it when

Pick this for agentic coding and long-horizon knowledge work over million-token contexts with image input, through Moonshot's compatible APIs or a multi-GPU cluster serving MXFP4 weights in vLLM or SGLang.

Look elsewhere when

Skip it if you need to self-host on modest hardware or keep serving cost low: about 104B parameters activate per token. Kimi K2 is cheaper per token at 32B active, and MiniMax M3 covers 1M context with 23B active.

Alternatives to Kimi K3

  • Kimi K2

    Predecessor with 32B active and agentic tool use, cheaper per token, but limited to 128K context and text input.

  • Qwen3.8

    Qwen3.8's 2.4T flagship has 95B active and 262K native context but is text-only and thinking-only; its 27B sibling covers single-GPU use.

  • MiniMax M3

    About 428B total and 23B active with native image and video and 1M context, far cheaper to serve, under the MiniMax Community license.

  • Inkling

    Apache-2.0 975B MoE with text, image, and audio input and up to 1M context self-hosted, framed as a fine-tuning base, not a finished model.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Freemium
Platform
Hybrid
Difficulty
Advanced (4/5)
License
Kimi K3 License
Added
Aug 24, 2026

Tags