Muse-Glimmer-30B is currently partially supported by llama.cpp and ollama. It might not output to cli at all, use webinterface instead. Make sure to use llama.cpp version greater than b10430, model will fail to load on older version.

This GGUF is pruned, so it only contains latin characters. It might break/die for no reason. Provided by bluevoid-pl.

Pruning reduces size of model by ~30% allowing you to run model on limited VRAM.

We make no guarantees of any kind that this gguf will work at all. Note that pruning process removes emojis, so model is physically incapable of outputting them.(but model still thinks that it can)
llama.cpp detects tokenizer changes and might hang for 5min at first start. Warning W load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect is expected.

Original Overview is below

Muse Glimmer Model Card

Authors: Meta Superintelligence Lab
Model Release Date: August 2026
License: Apache 2.0

Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.

Building effective agents requires key capabilities working together to achieve the user’s goals. Muse Glimmer is trained and evaluated on these capabilities:

  • End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕3-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.
  • Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows.
  • Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows.
  • Failure Recovery. When a tool call fails or returns an unexpected result, the model diagnoses the error and retries rather than halt.
  • Multimodal Input and Reasoning. Through a dedicated perception encoder, the model accepts interleaved text and images. This enables agents to interpret screenshots, charts, and documents alongside conversation.
  • Scaffold Compatibility. Muse Glimmer works across OpenClaw, Hermes Agent, and other agentic orchestration patterns.
  • Controllable Effort. The model supports different reasoning strengths to select the right balance between quality and speed.
  • Multilingual. Muse Glimmer is trained on data from more than 100 languages.
Downloads last month
38
GGUF
Model size
27B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF

Quantized
(130)
this model