minicpm5-2b-distilled-reasoning GGUF

GGUF quantizations of jigs97022/minicpm5-2b-distilled-reasoning.

A 2B parameter reasoning model distilled from three frontier models — Qwen3.8-Max, GLM-5.2, and Kimi K3 — onto the efficient openbmb/MiniCPM5-2B architecture.
Trained via QLoRA on 10,000 quality-filtered reasoning traces covering math, code, logic puzzles, and structured problem-solving.


Quantization Table

File Bits Size Use Case
model_f16.gguf 16 ~5.0 GB Maximum quality, reference
model_q8_0.gguf 8 ~2.7 GB Best quality / size tradeoff
model_q6_k.gguf 6 ~2.1 GB High quality
model_q5_k_m.gguf 5 ~1.8 GB Recommended
model_q5_k_s.gguf 5 ~1.7 GB Slightly smaller Q5
model_q4_k_m.gguf 4 ~1.5 GB Good balance
model_q4_k_s.gguf 4 ~1.4 GB Smaller Q4
model_q3_k_l.gguf 3 ~1.2 GB Low RAM, large variant
model_q3_k_m.gguf 3 ~1.1 GB Low RAM
model_q3_k_s.gguf 3 ~1.0 GB Minimum RAM Q3
model_q2_k.gguf 2 ~0.8 GB Extreme compression

VRAM Requirements

Quant Min VRAM
F16 8 GB
Q8_0 4 GB
Q4_K_M 2 GB
Q2_K 1.5 GB

Training Details

Parameter Value
Base Model openbmb/MiniCPM5-2B
Teachers Qwen3.8-Max, GLM-5.2, Kimi K3
Dataset r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation
Train Samples 10,000
Method QLoRA (r=64, alpha=32)
Max Context 4,096 tokens
Hardware Kaggle T4 x2

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf \
  --temp 0.7 \
  -n 1024 \
  -p "If 3x + 7 = 22, what is x? Show step-by-step reasoning."

llama-cpp-python

from llama_cpp import Llama

llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=4096)
output = llm(
    "Solve step by step: If 3x + 7 = 22, what is x?",
    max_tokens=1024,
    temperature=0.7,
)
print(output["choices"][0]["text"])

LM Studio

Search tinyopsec/minicpm5-2b-distilled-reasoning-GGUF in the model browser.

Ollama

ollama run hf.co/tinyopsec/minicpm5-2b-distilled-reasoning-GGUF:Q4_K_M

Limitations

  • Reasoning quality may degrade beyond 4K context
  • Trained on 10K samples subset; edge cases may be weaker than full 52K variant
  • English and Chinese only
  • Knowledge cutoff ~2024 (inherited from MiniCPM5-2B base)

Original Model

jigs97022/minicpm5-2b-distilled-reasoning

Downloads last month
1,042
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/minicpm5-2b-distilled-reasoning-GGUF

Quantized
(3)
this model