MiniCPM5-2B — Pollard (MLX)

Pollard shrank this model for Apple Silicon: 5.0 GB (f16) → 1.6 GB68% smaller, 3.1× down.

Three MLX rungs by measured allocation (bits placed per-layer by sensitivity). The mix at the repo root is the recommended default; q8/ and q4/ are the higher- and lower-fidelity rungs.

Pollard MLX builds of openbmb/MiniCPM5-2B made with Pollard Weights. Runs on Apple Silicon via MLX/Metal.

Available rungs

rung bpw size notes path
q8 8.50 2.5 GB near-lossless q8/
mix (recommended) 6.48 1.9 GB measured 4/8 mixed-precision repo root
q4 5.43 1.6 GB smallest q4/

PPL / Mean-KLD benchmarking pending — sizes and allocation are final.

Usage

# recommended (the mix, at the repo root):
mlx_lm.generate --model PollardWeights/MiniCPM5-2B-Pollard-MLX --prompt "Explain why the sky is blue."

# a specific rung (q8 or q4) — fetch the subfolder, then point mlx_lm at it:
huggingface-cli download PollardWeights/MiniCPM5-2B-Pollard-MLX --include "q8/*" --local-dir ./mcpm-mlx
mlx_lm.generate --model ./mcpm-mlx/q8 --prompt "Hello"

Errata

  • MLX quants run on Apple Silicon (Metal) via mlx_lm.
  • Measured allocation places bits by per-layer sensitivity under a size budget; router/embeddings/attention gates kept high.
  • Single machine; replication invited.

Built with Pollard Weights — frontier models, small hardware, no compromise.

Downloads last month
211
Safetensors
Model size
3B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/MiniCPM5-2B-Pollard-MLX

Quantized
(31)
this model