Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
Meta has introduced Muse Glimmer, a new multimodal model specifically designed for local agentic applications. Distilled from the larger Muse model down to 30B parameters and released under the permissive Apache 2.0 license, it is positioned for privacy-aware scenarios such as coding, document analysis, personal assistants, and Claw- or Hermes-like setups. Hugging Face is shipping day-0 support in transformers, llama.cpp, vLLM, Inference Endpoints, and other libraries, making the model immediately accessible.
Benchmarked against Gemma4-31B Thinking Mode and Qwen3.6-27B Thinking Mode, Muse Glimmer leads on many agentic evaluations: MCP Atlas 75.5 vs 54.2/62.5, DeepSearch QA 74.6 vs 61.7/71.1, τ³-Banking 23.5 vs 15.1/16.7, WildClawBench 47.6 vs 37.6/43.2, GAIA2 43.3 vs 36.4/40.0, SWE-Bench Pro 51.2 vs 36.9/50.2, and SciCode 43.6 vs 43.4/39.8. It also tops or ties several multimodal and general reasoning benchmarks, including OmniDocBench v1.5 75.8, MMMU Pro 74 (tied with Gemma), IFBench 77, AIME 2026 94.7, and AA-LCR 80. However, it trails on OSWorld-Verified (65.9 vs 58.5/75.6), TerminalBench 2.1 (51.7 vs 43.4/60.7), SkillsBench (44.3 vs 32.4/46.6), and GDPval-AA (953 vs 811/1141), as well as on GPQA Diamond and Humanity's Last Exam. In safety evaluations, Muse Glimmer shows a CI Memories violation rate of 26.4 (vs 12.1 for Gemma and 53.4 for Qwen) and an AgentDojo attack success rate of 28.4 (vs 25.6/40.3) while retaining high utility.
Architecturally, Muse Glimmer is a dense 30B model consisting of a 2B ViT-style vision encoder (Perception Encoder) and a 28B text decoder. It uses a hybrid attention design that alternates between three sliding-window layers of 2,048 tokens. Additionally, an optional speculative decoding drafter implemented on DFlash can provide much faster generation at the cost of some memory, and is particularly well suited to structured content generation such as coding.
The combination of a permissive license, strong agentic benchmark scores, and local deployability makes Muse Glimmer an attractive foundation for privacy-sensitive agents, cost-efficient inference, and experimentation. The day-0 integrations lower the barrier to adoption in existing tooling.