Instructions to use WiktorMatuszek/smaug-mini-fp8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WiktorMatuszek/smaug-mini-fp8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="WiktorMatuszek/smaug-mini-fp8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("WiktorMatuszek/smaug-mini-fp8") model = AutoModelForMultimodalLM.from_pretrained("WiktorMatuszek/smaug-mini-fp8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WiktorMatuszek/smaug-mini-fp8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WiktorMatuszek/smaug-mini-fp8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/WiktorMatuszek/smaug-mini-fp8
- SGLang
How to use WiktorMatuszek/smaug-mini-fp8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-fp8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-fp8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use WiktorMatuszek/smaug-mini-fp8 with Docker Model Runner:
docker model run hf.co/WiktorMatuszek/smaug-mini-fp8
Smaug-Mini FP8
Community FP8 quantization of abacusai/Smaug-Mini, an agentic fine-tune of Qwen3.8-27B.
The language model is quantized with NVIDIA ModelOpt while the vision tower remains in BF16. The source model's native MTP head is retained.
Quantization
- NVIDIA ModelOpt:
0.47.0 - Recipe:
general/ptq/fp8_default-kv_fp8_cast - Calibration prompts: 1,024
- Calibration sequence length: 4,096
- KV cache: FP8
- Vision tower: BF16
The checkpoint is stored in ModelOpt's Hugging Face format and includes hf_quant_config.json.
Serving
Example vLLM invocation:
vllm serve WiktorMatuszek/smaug-mini-fp8 \
--quantization modelopt \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
The source Smaug-Mini model supports up to a 262,144-token context. Set --max-model-len to a value appropriate for your available memory and workload.
MTP note
Smaug-Mini inherits its MTP head from the underlying Qwen model. The official Smaug-Mini model card states that this head was not retrained after the language-trunk fine-tune and recommends leaving MTP speculative decoding disabled. Standard decoding is unaffected.
Measured native-MTP behavior
In our earlier serving benchmark on one RTX PRO 6000, using greedy decoding over 256 held-out prompts with 512 generated tokens and 16 concurrent requests, the unmodified native MTP head measured:
| Verifier | Mean accepted length | Per-position acceptance |
|---|---|---|
| FP8 | 3.15 | 0.86 / 0.71 / 0.57 |
| MXFP8 | 3.14 | 0.86 / 0.71 / 0.57 |
At 128 concurrent requests with a 16k maximum output length, the same test setup measured roughly 1,000 tok/s with MTP enabled versus 1,400–1,560 tok/s without MTP. These are serving-performance measurements, not model-quality scores; they are included to document why MTP is not recommended for this Smaug-Mini release.
Evaluation
The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.
License and attribution
Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.
- Downloads last month
- 43