Smaug-Mini NVFP4

Community NVFP4 quantization of abacusai/Smaug-Mini, with an FP8 KV cache.

The language model is quantized with NVIDIA ModelOpt while the vision tower remains in BF16.

Quantization

  • NVIDIA ModelOpt: 0.47.0
  • Recipe: general/ptq/nvfp4_default-kv_fp8_cast
  • Calibration prompts: 1,024
  • Calibration sequence length: 4,096
  • KV cache: FP8
  • Vision tower: BF16

The checkpoint is stored in ModelOpt's Hugging Face format and includes hf_quant_config.json.

Hugging Face's automated safetensors metadata may display an 8-bit tag and a lower parameter total for this checkpoint because packed FP4 tensors are stored in uint8 containers. The model architecture remains the 27B Smaug-Mini architecture.

Serving

Example vLLM invocation:

vllm serve WiktorMatuszek/smaug-mini-nvfp4 \
  --quantization modelopt_fp4 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

vLLM selects an NVFP4 linear backend according to the GPU and available kernels.

Evaluation

The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.

License and attribution

Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.

Downloads last month
43
Safetensors
Model size
15B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WiktorMatuszek/smaug-mini-nvfp4

Base model

Qwen/Qwen3.8-27B
Quantized
(6)
this model