SGLang
The Muse Glimmer model support pull request was merged on August 11, 2026.
This reference covers the dense open agentic checkpoint distilled from Muse Spark for local execution on a consumer GPU. It keeps model card dates separate from announcement and provider dates.
Muse Glimmer 30B is the vendor reference for meta-models/Muse-Glimmer-30B, as of Aug 12, 2026. Model architecture: Dense Causal Transformer with Perception Encoder; Model card context: 131,072+, as of Aug 12, 2026. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | meta-models/Muse-Glimmer-30B | huggingface.coRetrieved |
| Identity | Authors | Meta Superintelligence Lab | huggingface.coRetrieved |
| Identity | Model card release date | August 2026 | huggingface.coRetrieved |
| Identity | Announcement publication date | 2026-08-10 | research.meta.aiRetrieved |
| Identity | OpenRouter release date | Released Aug 9, 2026 | openrouter.aiRetrieved |
| Architecture | Model architecture | Dense Causal Transformer with Perception Encoder | huggingface.coRetrieved |
| Architecture | Parameter scale | Total ~29.6B parameters including vision encoder | huggingface.coRetrieved |
| Architecture | Transformer configuration | 52 layers; hidden size 6656 | huggingface.coRetrieved |
| Architecture | Attention configuration | 32 Q heads / 2 KV heads (GQA 16:1); sliding window 2048; gated attention | huggingface.coRetrieved |
| Architecture | Perception encoder | ~1.8B parameter ViT-G/14; 50 layers; width 1536; patch size 14 | huggingface.coRetrieved |
| Architecture | DFlash configuration | 5 draft layers; block size 16 | huggingface.coRetrieved |
| Architecture | DFlash behavior | predicts entire blocks of 16 tokens in a single forward pass | huggingface.coRetrieved |
| Architecture | Knowledge cutoff | January 4, 2026 | huggingface.coRetrieved |
| Architecture | Distillation | We trained Muse Glimmer on Muse Spark's outputs using logit distillation, leveraging a similar data mix as the teacher. | research.meta.aiRetrieved |
| Context | Model card context | 131,072+ | huggingface.coRetrieved |
| Context | OpenRouter limits | 131,072 context; 16,384 max output | openrouter.aiRetrieved |
| Modalities | Input and output | Input: text + image, Output: text | huggingface.coRetrieved |
| Modalities | Audio | Audio input/output is not supported. | huggingface.coRetrieved |
| Modalities | Languages | trained on data from more than 100 languages | huggingface.coRetrieved |
| Modalities | Reasoning strength | System-prompt values: low, medium, high, xhigh | huggingface.coRetrieved |
| License | Artifact license | All artifacts are released under Apache 2.0 | huggingface.coRetrieved |
| License | License file | Apache License, Version 2.0 | huggingface.coRetrieved |
| License | Usage policy | Muse Glimmer is not intended for individuals under the age of 18. | huggingface.coRetrieved |
| Pricing | Meta API price: OpenRouter input | $0.30 per million input tokens | openrouter.aiRetrieved |
| Pricing | Meta API price: OpenRouter output | $1.20 per million output tokens | openrouter.aiRetrieved |
| Availability | Repository access | Ungated repository. | huggingface.coRetrieved |
| Availability | Released artifacts | Full-precision weights (BF16), 4-bit quantized weights (2 variants), and DFlash drafter head | huggingface.coRetrieved |
| Availability | Quantized targets | 24 GB, 32 GB, and 64 GB VRAM tiers, shrinking the language model to under 20 GB | huggingface.coRetrieved |
| Availability | Companion repositories | Muse-Glimmer-30B-GGUF, Muse-Glimmer-30B-assistant, and Muse-Glimmer-30B-ExecuTorch-PTE | huggingface.coRetrieved |
| Availability | Vendor serving channels | serve it at scale with vLLM and SGLang | research.meta.aiRetrieved |
| Availability | Main weights | meta-models/Muse-Glimmer-30B | huggingface.coRetrieved |
"Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It's small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation."
research.meta.aiRetrieved
"We trained Muse Glimmer on Muse Spark's outputs using logit distillation, leveraging a similar data mix as the teacher."
research.meta.aiRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| SWE-Bench Verified, Muse Glimmer-30B High Reasoning | 76.0 | huggingface.coRetrieved |
| SWE-Bench Pro, Muse Glimmer-30B High Reasoning | 51.2 | huggingface.coRetrieved |
| AIME 2026, Muse Glimmer-30B High Reasoning | 94.7 | huggingface.coRetrieved |
| GPQA Diamond (AA), Muse Glimmer-30B High Reasoning | 83.5 | huggingface.coRetrieved |
| MCP Atlas (Public), Muse Glimmer-30B High Reasoning | 75.5 | huggingface.coRetrieved |
| OSWorld-Verified, Muse Glimmer-30B High Reasoning | 65.9 | huggingface.coRetrieved |
| MMMU Pro, Muse Glimmer-30B High Reasoning | 74 | huggingface.coRetrieved |
| Terminal-Bench 2.1 with terminus2, Muse Glimmer-30B High Reasoning | 51.7 | huggingface.coRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The Muse Glimmer model support pull request was merged on August 11, 2026.
Support remains a recipe and Docker path because the model code and the muse_glimmer parsers are not in any released vLLM wheel; the main-repository support pull request remains open.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: meta-models/Muse-Glimmer-30B.
As of Aug 12, 2026, Repository access: Ungated repository.
As of Aug 12, 2026, Repository access: Ungated repository.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs