Instructions to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S # Run inference directly in the terminal: ./llama-cli -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Use Docker
docker model run hf.co/peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
- LM Studio
- Jan
- vLLM
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
- Ollama
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with Ollama:
ollama run hf.co/peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
- Unsloth Desktop
- Pi
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with Docker Model Runner:
docker model run hf.co/peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
- Lemonade
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Run and chat with the model
lemonade run user.Sharp-MiniCPM5-2B-GGUF-Q4_K_S
List all available models
lemonade list
- Hermes Agent
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q4_K_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sharp-MiniCPM5-2B-GGUF
Dynamic imatrix quants of openbmb/MiniCPM5-2B, a 2.5B dense 131k-context model, carrying our cyber-and-coding-weighted importance matrix, an unsloth-dynamic-style per-tensor quant policy, and the Sharp-MiniCPM chat template with fixes and an improved system prompt.
Which quant should I download?
We recommend the Q6_K_XL for anyone who can run it. Practically lossless, at under half the size of the F16 model, and scoring higher than stock Q8_0 on SWE-bench-Live.
Memory math
99% KL is the 99th percentile of KL divergence against the BF16 weights on held-out code — the rare token where a quant actually diverges. top-1 is how often the quant picks the same next token as BF16. The last three columns are the largest q8_0 KV context that fits beside the weights
on a 4 / 6 / 8 GB card.
| tier | size | 99% KL | top-1 | 4 GB | 6 GB | 8 GB |
|---|---|---|---|---|---|---|
| Q4_K_S | 1.52 GB | 0.333 | 92.3% | 97K | 131K* | 131K* |
| Q4_K_XL | 1.60 GB | 0.274 | 93.0% | 94K | 131K* | 131K* |
| Q5_K_XL | 1.89 GB | 0.086 | 96.3% | 81K | 131K* | 131K* |
| Q6_K_XL | 2.20 GB | 0.024 | 98.1% | 68K | 131K* | 131K* |
| Q8_K_M | 2.93 GB | 0.004 | 99.2% | 36K | 130K | 131K* |
* capped by the model's own 131072 native context, not by VRAM. Context assumes
-ctk q8_0 -ctv q8_0, all layers on the GPU, and 0.5 GiB reserved for compute buffers and runtime;
figures are rounded down to the thousand. Run f16 KV instead and roughly halve them.
On 6 GB and up, the model's own 131K window binds before VRAM does at every tier, so choose on fidelity: Q6_K_XL at 0.024 is near-lossless for 2.20 GB. On 4 GB, Q4_K_S gives you 97K of context, and even the floor tier still agrees with BF16 on 92% of next tokens.
Run it
MiniCPM5-2B is a plain llama architecture, so any recent stock llama.cpp runs it:
llama-server -hf peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF:Q6_K_XL --jinja -ngl 99 \
-c 131072 -ctk q8_0 -ctv q8_0 -fa on \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
- Swap the tag after the colon for another tier (
Q4_K_S,Q4_K_XL,Q5_K_XL,Q8_K_M). --jinjauses the Sharp-MiniCPM template embedded in every file; tool calls come back as standard OpenAItool_calls.- Sampling is openbmb's recommendation (temperature 1.0, top-p 0.95, min-p 0) plus top-k 20, the exact
settings the SWE-bench-Live board above ran with. Keep
--min-p 0.0: llama.cpp's default of 0.05 can trap this model in repetition loops. -ctk q8_0 -ctv q8_0is the KV cache the memory table assumes. With VRAM to spare, drop both flags for anf16cache (about twice the memory per token).
This model takes quantization unusually well
Measured against the same corpus, the same 128 KB held-out slice and the same method as our Sharp-Spark-X2.5-4B ladder:
| tier | MiniCPM5-2B 99% KL | Spark-X2.5-4B 99% KL |
|---|---|---|
| Q4_K_S | 0.333 | 1.156 |
| Q4_K_XL | 0.274 | 0.873 |
| Q5_K_XL | 0.086 | 0.363 |
| Q6_K_XL | 0.024 | 0.178 |
| Q8_K_M | 0.004 | 0.023 |
MiniCPM5-2B loses 3–7× less to quantization at every rung, despite being the smaller model —
the opposite of the usual rule that smaller models have less redundancy to spare. Part of that is
this architecture giving the recipe more to work with: untied embeddings let the output head stay
high while the input lookup drops, and split q/k/v lets the cheap KV projections be raised almost
for free — neither of which Spark's tied head and fused attn_qkv allow.
Read this as quantization robustness, not model quality. Each KL is measured against that model's own BF16, so it says how much the quant damaged the model relative to itself. It says nothing about which model is better at anything.
Architecture notes that shape the ladder
MiniCPM5-2B is a plain llama architecture — stock llama.cpp converts and runs it, no custom build.
Three facts drive the per-tensor policy:
- Embeddings are untied.
output.weightandtoken_embd.weightare separate tensors of 10.6% each.output.weightis the output projection and is pinned high on every tier;token_embdis an input lookup and drops lower, which is where most of the saving at the low tiers comes from. - No sliding-window attention — all 42 layers are full attention, so the bump family is the head and tail blocks rather than a structural minority of layers.
- GQA is 8:1 (16 heads, 2 KV heads, head_dim 128), so
attn_kandattn_vare 0.9% of params each. Every tier raises both; it costs ~11 MB across all 42 layers.
KV cache is 42 KB/token at f16, 22.3 KB/token at q8_0 — larger per token than a bigger model
with windowed attention would need, because every layer here holds a full-length cache.
The Sharp-MiniCPM template
MiniCPM5's own template with six defects repaired and the house terseness prompt spliced in. It is
not a port of another model's template: the control tokens, the <think> convention and the XML
<function>/<param> tool-call format are all MiniCPM5's, untouched, because the model was trained
on them. Every fix was reproduced against the stock template before being written.
- Assistant text after a
<tool_sep>was silently discarded. Stock builds the interleaved content with{% set %}inside a{% for %}; Jinja throws those assignments away when the loop ends, so the content reverted to the text before the first separator and everything after it vanished. Rebuilt on a namespace. - Content blocks were dropped to an empty string — a user turn sent as
[{"type": "text", ...}]arrived completely blank. - Tool results given as content blocks were dumped into the prompt as raw JSON.
- Non-string tool-call arguments rendered as a Python repr — a nested object reached the model
as
{'a': 1, 'b': True}, single quotes and all, instead of JSON. - Stale reasoning was replayed for every historical turn. Stock computes
ns.last_query_indexfor exactly this purpose and then never reads it; we wired it up. - CDATA wrapping did not escape a literal
]]>, which closes the block early and corrupts the rest of the call.
Additions: a terseness instruction force-appended to the system prompt (opt out with
chat_template_kwargs={"terse": false}), and suppress_tool_instructions, which stands the tool
block down when the runtime has injected its own tool protocol.
Verified to render in both jinja2 (transformers) and minja (llama.cpp). It does not render in
minijinja, whose tojson rejects the ensure_ascii keyword — that construct is inherited from the
stock template and is kept deliberately, because dropping it would \u-escape CJK in tool
descriptions and argument values on a bilingual model.
The template ships in this repo as chat_template.jinja.
- Downloads last month
- 3,976
Model tree for peculiar-ragdoll/Sharp-MiniCPM5-2B-GGUF
Base model
openbmb/MiniCPM5-2B