Muse Glimmer

Meta's Apache-licensed 30B agentic model, distilled from Muse Spark to run on one consumer GPU.

Open SourceSelf HostedOffline CapableGPU Required (24GB+ VRAM)
0.0 (0)

About

Meta returned to open weights in August 2026 with Muse Glimmer, a roughly 30B parameter dense transformer distilled from its closed Muse Spark system and licensed Apache 2.0. The 52-layer network mixes sliding-window and global attention in a repeating three-local-one-global pattern with 2,048 token windows, carries a 1.8B parameter vision encoder for image input, covers more than 100 languages, and reads 131K tokens of context. Training centered on end-to-end agent behavior: reliable function calling, multi-step tool workflows, error diagnosis, and recovery from failed actions. What makes it notable is deployment reach rather than raw scale; official K-Quant variants compress the language model below 20GB at 4-bit precision so it runs on a single 24GB consumer card or an Apple Silicon MacBook, and the bundled DFlash drafter adds block-level speculative decoding measured at 3.1x on an RTX 5090 and 1.5x on an Apple M4 Max. That combination puts a full agent loop, including vision, on hardware hobbyists already own, with no usage restrictions attached.

Should you use Muse Glimmer?

Pick it when

Pick this when you want a full agent loop with tool calling, error recovery, and image input running on a single 24 GB consumer GPU or an Apple Silicon MacBook, under Apache-2.0 with no usage restrictions.

Look elsewhere when

Skip it if your GPU has less than 24 GB, where gpt-oss-20b fits in 16 GB, or if you need million-token context or frontier-scale reasoning, since it reads 131K tokens. For pure agentic coding, Devstral is purpose-built.

Alternatives to Muse Glimmer

  • Devstral

    Apache-2.0 24B also sized for one RTX 4090, specialized for repository-scale coding agents rather than general tool use and vision.

  • gpt-oss

    Apache-2.0 MoE whose 20B fits 16 GB with configurable reasoning effort, but text-only and tied to the harmony response format.

  • Qwen3.8

    Its Apache-2.0 27B also fits one consumer GPU at 4-bit and adds video input, with coding as its emphasis instead of agent training.

  • Gemma 4

    Apache-2.0 26B MoE and 31B dense with 256K context and 140+ languages, longer context but no agent-focused distillation.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
24 GB
Added
Aug 24, 2026

Tags