Inkling
Thinking Machines Lab's Apache-licensed 975B MoE with 1M context and day-one Tinker fine-tuning.
About
Thinking Machines Lab, the startup Mira Murati founded after leaving OpenAI, shipped its first from-scratch model on July 15, 2026 and put the full weights on Hugging Face under Apache 2.0. Inkling is a 975B parameter mixture-of-experts network with 41B active per token that accepts text, images, and audio, and it debuted as the leading US open-weights model at 41 on the Artificial Analysis Intelligence Index. Context reaches 1M tokens when self-hosting the raw weights, while the company's Tinker fine-tuning platform serves it at 64K and 256K windows with day-one support for post-training, the release being explicitly framed as a base for customization rather than a finished chat product. A smaller sibling, Inkling-Small, followed at the end of July as a 276B MoE with 12B active that matches the flagship on several reasoning and agentic benchmarks at a quarter of the size. Running the full model locally requires a multi-GPU cluster; the permissive license allows unrestricted commercial use, modification, and redistribution.
Should you use Inkling?
Pick it when
Pick this when you plan to post-train your own model on a multimodal base that takes text, images, and audio, using Tinker for managed fine-tuning or self-hosting with up to 1M context under Apache-2.0.
Look elsewhere when
Skip it if you want a finished chat assistant or have no multi-GPU cluster: it is framed as a base for customization. Gemma 4 covers multimodal on modest hardware, and Kimi K3 or DeepSeek-V4 come with general hosted APIs.
Alternatives to Inkling
- Kimi K3
Larger 2.8T MoE with 1M context and image input, offered as a ready model through OpenAI and Anthropic compatible APIs, under modified MIT.
- Tulu 3
If you want the post-training recipe rather than a giant base, it gives open SFT, DPO, and RLVR code and datasets for smaller models.
- MiniMax M3
About 428B with 23B active, native image and video, and 1M context, cheaper per token but under the MiniMax Community license.
- Gemma 4
Apache-2.0 multimodal with audio on E2B and E4B models that run without a discrete GPU, for when cluster-scale capacity is unnecessary.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Freemium
- Platform
- Hybrid
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Added
- Aug 24, 2026