MiniCPM5-1B-Agent (MLX 4-bit)

4-bit quantized MLX version of Luminia/MiniCPM5-1B-Agent-GGUF. Quantized with mlx_lm.convert using affine mode (group_size=64, 4.5 bits/weight). Runs well on Apple Silicon with ~600 MB memory.

About the model

MiniCPM5-1B-Agent is a tiny agentic coding agent for CPU: a full fine-tune of openbmb/MiniCPM5-1B specialized to reason in <think>, call a small tool set (bash/read/write/edit/glob/grep), and run → read output → debug → patch → verify.

  • Base model: openbmb/MiniCPM5-1B (RL+OPD checkpoint)
  • Architecture: LlamaForCausalLM — 24 layers, 16 attention heads (GQA), 1536 hidden, 130560 vocab
  • Parameters: 1,080,632,832 (quantized to ~4.5 bits/weight)
  • Quantization: affine, group_size=64, 4 bits
  • License: Apache-2.0

How to use

from mlx_lm import load, generate

model, tokenizer = load("MC7ever/MiniCPM5-1B-Agent-mlx-q4")

messages = [
    {"role": "user", "content": "Write a Python function to check if a number is prime."}
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
response = generate(model, tokenizer, prompt=prompt, max_tokens=1024)
print(response)

Credits

Other variants

Variant Size Bits/Weight Repo
Safetensors (fp16) 2.0 GB 16 MC7ever/MiniCPM5-1B-Agent-safetensors
MLX Q4 580 MB 4.5 This repo
MLX Q2 387 MB 3.0 MC7ever/MiniCPM5-1B-Agent-mlx-q2
Downloads last month
19
Safetensors
Model size
0.2B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MC7ever/MiniCPM5-1B-Agent-mlx-q4

Quantized
(2)
this model

Paper for MC7ever/MiniCPM5-1B-Agent-mlx-q4