Smaug-Mini FP8

Community FP8 quantization of abacusai/Smaug-Mini, an agentic fine-tune of Qwen3.8-27B.

The language model is quantized with NVIDIA ModelOpt while the vision tower remains in BF16. The source model's native MTP head is retained.

Quantization

  • NVIDIA ModelOpt: 0.47.0
  • Recipe: general/ptq/fp8_default-kv_fp8_cast
  • Calibration prompts: 1,024
  • Calibration sequence length: 4,096
  • KV cache: FP8
  • Vision tower: BF16

The checkpoint is stored in ModelOpt's Hugging Face format and includes hf_quant_config.json.

Serving

Example vLLM invocation:

vllm serve WiktorMatuszek/smaug-mini-fp8 \
  --quantization modelopt \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The source Smaug-Mini model supports up to a 262,144-token context. Set --max-model-len to a value appropriate for your available memory and workload.

MTP note

Smaug-Mini inherits its MTP head from the underlying Qwen model. The official Smaug-Mini model card states that this head was not retrained after the language-trunk fine-tune and recommends leaving MTP speculative decoding disabled. Standard decoding is unaffected.

Measured native-MTP behavior

In our earlier serving benchmark on one RTX PRO 6000, using greedy decoding over 256 held-out prompts with 512 generated tokens and 16 concurrent requests, the unmodified native MTP head measured:

Verifier Mean accepted length Per-position acceptance
FP8 3.15 0.86 / 0.71 / 0.57
MXFP8 3.14 0.86 / 0.71 / 0.57

At 128 concurrent requests with a 16k maximum output length, the same test setup measured roughly 1,000 tok/s with MTP enabled versus 1,400–1,560 tok/s without MTP. These are serving-performance measurements, not model-quality scores; they are included to document why MTP is not recommended for this Smaug-Mini release.

Evaluation

The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.

License and attribution

Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.

Downloads last month
43
Safetensors
Model size
27B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WiktorMatuszek/smaug-mini-fp8

Base model

Qwen/Qwen3.8-27B
Quantized
(6)
this model