Kimi K2
Moonshot AI's trillion-parameter MoE model series with 32B active parameters and agentic tool use.
About
Kimi K2 pushed open-weight scale to a trillion parameters: Moonshot AI's MoE series activates 32B parameters per token across 384 experts with 8 selected plus one shared expert, 61 layers, MLA attention, a 160k vocabulary, and a 128K context window. Training stability at that scale came from applying the MuonClip variant of the Muon optimizer. K2-Base serves fine-tuners while K2-Instruct is post-trained for chat and agentic tool use, and the later K2 Thinking release added an open reasoning model that Moonshot reports sustaining 200 to 300 sequential tool calls without human intervention. Weights carry a Modified MIT license: fully commercial, with the sole condition that products exceeding 100 million monthly active users or 20 million US dollars in monthly revenue must prominently display Kimi K2 in their interface. Self-hosting the 1T checkpoint requires a multi-GPU cluster served through vLLM, SGLang, KTransformers, or TensorRT-LLM, so many teams instead use the OpenAI- and Anthropic-compatible API on Moonshot's hosted Kimi platform. It set state-of-the-art open results on agentic and coding benchmarks at release.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Modified MIT
- Added
- Jul 29, 2026
Related Tools
Lightweight open-weight LLM by Google available in 1B to 27B sizes.
Open-source code LLM family by IBM for enterprise code generation.
Open-weight LLM by Meta in 8B and 70B sizes with strong general capabilities.
High-performance open-weight MoE LLM with 671B total parameters.
Hybrid SSM-Transformer model by AI21 Labs combining Mamba with attention layers.
Open-weight code LLM trained on 2 trillion tokens of code and natural language.