YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

MiniCPM5-1B-NVFP4-W4A4

NVFP4 (W4A4) quantized version of openbmb/MiniCPM5-1B.

Quantized with llm-compressor using the NVFP4 scheme, saved in compressed-tensors format. Both weights and activations are FP4 (W4A4): per-group-16 static fp8 weight scales, and per-group local dynamic fp8 activation scales (static_minmax observer).

Calibration: 128 samples drawn from the canada-quant/hy3-w4a16-mtp-calibration blend (code / math / chat, ~2048 tokens each), formatted through this model's chat template. lm_head is left in bf16.

Loadable directly in vLLM. Needs a Blackwell GPU with CUDA 13 (e.g. RTX 5090) for native FP4 kernels:

vllm serve MaxDevv/MiniCPM5-1B-NVFP4-W4A4 \
  --kv-cache-dtype fp8_e4m3 \
  --attention-backend flashinfer

Original model below:

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support