Instructions to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Use Docker
docker model run hf.co/bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Ollama:
ollama run hf.co/bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
- Unsloth Studio
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF to start chatting
- Pi
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Docker Model Runner:
docker model run hf.co/bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
- Lemonade
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Muse-Glimmer-30B-pruned-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
Muse-Glimmer-30B is currently partially supported by llama.cpp and ollama. It might not output to cli at all, use webinterface instead. Make sure to use llama.cpp version greater than
b10430, model will fail to load on older version.
This GGUF is pruned, so it only contains latin characters. It might break/die for no reason. Provided by bluevoid-pl.
Pruning reduces size of model by ~30% allowing you to run model on limited VRAM.
We make no guarantees of any kind that this gguf will work at all. Note that pruning process removes emojis, so model is physically incapable of outputting them.(but model still thinks that it can)
llama.cpp detects tokenizer changes and might hang for 5min at first start. Warning W load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect is expected.
Original Overview is below
Muse Glimmer Model Card
Authors: Meta Superintelligence Lab
Model Release Date: August 2026
License: Apache 2.0
Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.
Building effective agents requires key capabilities working together to achieve the user’s goals. Muse Glimmer is trained and evaluated on these capabilities:
- End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕3-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.
- Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows.
- Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows.
- Failure Recovery. When a tool call fails or returns an unexpected result, the model diagnoses the error and retries rather than halt.
- Multimodal Input and Reasoning. Through a dedicated perception encoder, the model accepts interleaved text and images. This enables agents to interpret screenshots, charts, and documents alongside conversation.
- Scaffold Compatibility. Muse Glimmer works across OpenClaw, Hermes Agent, and other agentic orchestration patterns.
- Controllable Effort. The model supports different reasoning strengths to select the right balance between quality and speed.
- Multilingual. Muse Glimmer is trained on data from more than 100 languages.
- Downloads last month
- 38
2-bit
3-bit
4-bit
8-bit
16-bit
Model tree for bluevoid-pl/Muse-Glimmer-30B-pruned-GGUF
Base model
meta-models/Muse-Glimmer-30B