Instructions to use Mapika/decider-4b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Mapika/decider-4b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mapika/decider-4b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mapika/decider-4b-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mapika/decider-4b-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mapika/decider-4b-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mapika/decider-4b-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mapika/decider-4b-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mapika/decider-4b-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mapika/decider-4b-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Mapika/decider-4b-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Mapika/decider-4b-GGUF with Ollama:
ollama run hf.co/Mapika/decider-4b-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Mapika/decider-4b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mapika/decider-4b-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Mapika/decider-4b-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Mapika/decider-4b-GGUF with Docker Model Runner:
docker model run hf.co/Mapika/decider-4b-GGUF:Q4_K_M
- Lemonade
How to use Mapika/decider-4b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mapika/decider-4b-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.decider-4b-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Mapika/decider-4b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mapika/decider-4b-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Mapika/decider-4b-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Mapika/decider-4b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Mapika/decider-4b-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Mapika/decider-4b-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
decider-4b GGUF
GGUF files of Mapika/decider-4b v2.1 (Hub main eb5fbdf), for llama.cpp. decider-4b
does not generate text. It reads a state and one or more questions, each with an explicit option list, and returns a probability
for every option from one forward pass. See the decider-4b card for what the model
is, how it was trained, and where it is weak.
| file | size | use |
|---|---|---|
decider-4b-v2.1-Q4_K_M.gguf |
2.7 GB | smallest; about 0.2 points lower in-task accuracy (table below) |
decider-4b-v2.1-Q8_0.gguf |
4.5 GB | same quality as the bf16 weights |
decider-4b-v2.1-BF16.gguf |
8.4 GB | unquantized, for making other quantizations |
The tokenizer, decider_config.json (temperatures) and decide_gguf.py (the readout on llama.cpp) are in this repository too.
This is not a chat model
Loading a file in llama-cli, llama-server, Ollama or LM Studio gives you a text model that continues prompts. That is not how
decider-4b is used, and its generated text is not its answer. The answer is read from the logits of the option-letter tokens at
each answer slot of a prompt built by decider.prompt, divided by the fitted temperature. Decider (decider-ai 1.6.0) and
decide_gguf.py do this with llama-cpp-python.
Usage
With decider-ai 1.6.0 or newer, Decider loads the GGUF file directly: decide, system_one and the per-type temperatures
of decider_config.json, scored by llama.cpp.
pip install "decider-ai[gguf]" # llama-cpp-python; GPU: CMAKE_ARGS="-DGGML_CUDA=on" (Apple Silicon: -DGGML_METAL=on)
from decider.infer import Decider
d = Decider("Mapika/decider-4b-GGUF", gguf_file="decider-4b-v2.1-Q4_K_M.gguf")
d.decide("My card was charged twice for the same purchase.",
[{"question": "Which department should handle this?", "options": ["billing", "technical", "sales"]}])
gguf_options=dict(n_gpu_layers=0, n_threads=8) runs on the CPU. The standalone script below does the same decide() readout
with decider-ai 1.5.0.
Standalone script (decide_gguf.py)
pip install decider-ai==1.5.0 llama-cpp-python # llama-cpp-python 0.3.35 or newer
# GPU: CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python (Apple Silicon: -DGGML_METAL=on)
hf download Mapika/decider-4b-GGUF --local-dir decider-4b-gguf \
--include "*Q4_K_M.gguf" --include "*.json" --include "*.jinja" --include "*.py"
cd decider-4b-gguf
from decide_gguf import GGUFDecider
d = GGUFDecider("decider-4b-v2.1-Q4_K_M.gguf") # n_gpu_layers=-1 (all on the GPU if the build has one), n_threads=...
d.decide("My card was charged twice for the same purchase.",
[{"question": "Which department should handle this?", "options": ["billing", "technical", "sales"]},
{"question": "How urgent is this?", "options": ["low", "medium", "high"]}])
# [{'choice': 'billing', 'confidence': 0.85, 'probs': {...}}, {'choice': 'medium', 'confidence': 0.43, 'probs': {...}}]
decide_gguf.py covers decide() (choice questions); system_one with score and yes/no answers needs decider-ai 1.6.0 (above).
The HTTP server does not serve GGUF files.
On 8 CPU threads (server CPU), Q4_K_M takes 0.3 to 0.7 s for a request of 40 to 120 tokens; Q8_0 is about 20% slower.
Measured quality
The regression set of the decider-4b card (95 tasks, 144,226 questions, 67 in-task and 28 held-out tasks) at the shipped temperature 1.099, read through llama.cpp (CUDA build, one prompt per decode) and compared row by row with the bf16 weights in PyTorch. Accuracy, NLL and ECE are means over tasks.
| in-task acc / NLL / ECE | held-out acc / NLL / ECE | same answer as bf16 PyTorch | |
|---|---|---|---|
| bf16 weights, PyTorch | 0.8308 / 0.4145 / 0.0308 | 0.7838 / 0.5703 / 0.0781 | |
| BF16 GGUF | 0.8308 / 0.4145 / 0.0308 | 0.7837 / 0.5700 / 0.0782 | 99.45% |
| Q8_0 | 0.8310 / 0.4145 / 0.0309 | 0.7829 / 0.5699 / 0.0778 | 99.31% |
| Q4_K_M | 0.8288 / 0.4194 / 0.0324 | 0.7834 / 0.5691 / 0.0733 | 97.06% |
The BF16 GGUF differs from PyTorch only on near-ties (median probability difference 0.001). Q8_0 is equal to the bf16 weights within that noise; its tasks move up on 31 and down on 39, by at most 0.9 points. Q4_K_M is 0.2 points lower on in-task accuracy with slightly higher NLL; held-out accuracy is unchanged. It is lower than bf16 on 54 tasks and higher on 36; the largest drop is fin_phrasebank (−3.5 points), then medmcqa and mmlu (−1.3). In Q4_K_M the embedding matrix, which is also the output matrix that holds the option-letter rows, is stored in Q6_K.
Notes
- Score one prompt per
llama_decode, asdecide_gguf.pydoes. With several prompts in one decode (as separate sequences), llama.cpp in September 2026 gives probabilities that change with the other prompts in the batch, by up to 0.02 in BF16 and 0.16 in Q4_K_M on this model. One prompt per decode gives the same numbers on every run. - CPU and GPU builds give slightly different probabilities on the same file (for example 0.846 and 0.830 for "billing" above, against 0.844 in PyTorch).
- Conversion: llama.cpp
c9064dded(2026-09-27),convert_hf_to_gguf.py --no-mtp --outtype bf16, thenllama-quantizeto Q8_0 and Q4_K_M.--no-mtpis required: the checkpoint has no multi-token-prediction weights, but its config declares one MTP layer, and without the flag the converter writes a file that llama.cpp cannot load. - The measurements above use the CUDA build. The CPU and Metal builds were not run over the regression set.
License: Apache-2.0, as decider-4b and its base model Qwen/Qwen3.5-4B-Base.
- Downloads last month
- -
4-bit
8-bit
16-bit