Muse Glimmer
Meta's Apache-licensed 30B agentic model, distilled from Muse Spark to run on one consumer GPU.
About
Meta returned to open weights in August 2026 with Muse Glimmer, a roughly 30B parameter dense transformer distilled from its closed Muse Spark system and licensed Apache 2.0. The 52-layer network mixes sliding-window and global attention in a repeating three-local-one-global pattern with 2,048 token windows, carries a 1.8B parameter vision encoder for image input, covers more than 100 languages, and reads 131K tokens of context. Training centered on end-to-end agent behavior: reliable function calling, multi-step tool workflows, error diagnosis, and recovery from failed actions. What makes it notable is deployment reach rather than raw scale; official K-Quant variants compress the language model below 20GB at 4-bit precision so it runs on a single 24GB consumer card or an Apple Silicon MacBook, and the bundled DFlash drafter adds block-level speculative decoding measured at 3.1x on an RTX 5090 and 1.5x on an Apple M4 Max. That combination puts a full agent loop, including vision, on hardware hobbyists already own, with no usage restrictions attached.
Should you use Muse Glimmer?
Pick it when
Pick this when you want a full agent loop with tool calling, error recovery, and image input running on a single 24 GB consumer GPU or an Apple Silicon MacBook, under Apache-2.0 with no usage restrictions.
Look elsewhere when
Skip it if your GPU has less than 24 GB, where gpt-oss-20b fits in 16 GB, or if you need million-token context or frontier-scale reasoning, since it reads 131K tokens. For pure agentic coding, Devstral is purpose-built.
Alternatives to Muse Glimmer
- Devstral
Apache-2.0 24B also sized for one RTX 4090, specialized for repository-scale coding agents rather than general tool use and vision.
- gpt-oss
Apache-2.0 MoE whose 20B fits 16 GB with configurable reasoning effort, but text-only and tied to the harmony response format.
- Qwen3.8
Its Apache-2.0 27B also fits one consumer GPU at 4-bit and adds video input, with coding as its emphasis instead of agent training.
- Gemma 4
Apache-2.0 26B MoE and 31B dense with 256K context and 140+ languages, longer context but no agent-focused distillation.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Large Language Models (LLMs)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 24 GB
- Added
- Aug 24, 2026