Instructions to use WiktorMatuszek/smaug-mini-nvfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WiktorMatuszek/smaug-mini-nvfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="WiktorMatuszek/smaug-mini-nvfp4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("WiktorMatuszek/smaug-mini-nvfp4") model = AutoModelForMultimodalLM.from_pretrained("WiktorMatuszek/smaug-mini-nvfp4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WiktorMatuszek/smaug-mini-nvfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WiktorMatuszek/smaug-mini-nvfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-nvfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/WiktorMatuszek/smaug-mini-nvfp4
- SGLang
How to use WiktorMatuszek/smaug-mini-nvfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-nvfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WiktorMatuszek/smaug-mini-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WiktorMatuszek/smaug-mini-nvfp4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use WiktorMatuszek/smaug-mini-nvfp4 with Docker Model Runner:
docker model run hf.co/WiktorMatuszek/smaug-mini-nvfp4
Smaug-Mini NVFP4
Community NVFP4 quantization of abacusai/Smaug-Mini, with an FP8 KV cache.
The language model is quantized with NVIDIA ModelOpt while the vision tower remains in BF16.
Quantization
- NVIDIA ModelOpt:
0.47.0 - Recipe:
general/ptq/nvfp4_default-kv_fp8_cast - Calibration prompts: 1,024
- Calibration sequence length: 4,096
- KV cache: FP8
- Vision tower: BF16
The checkpoint is stored in ModelOpt's Hugging Face format and includes hf_quant_config.json.
Hugging Face's automated safetensors metadata may display an 8-bit tag and a lower parameter total for this checkpoint because packed FP4 tensors are stored in uint8 containers. The model architecture remains the 27B Smaug-Mini architecture.
Serving
Example vLLM invocation:
vllm serve WiktorMatuszek/smaug-mini-nvfp4 \
--quantization modelopt_fp4 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
vLLM selects an NVFP4 linear backend according to the GPU and available kernels.
Evaluation
The full capability evaluation for the unquantized model is published on the Smaug-Mini model card. This repository does not claim an independent rerun of that benchmark suite. Quantization can change outputs, so evaluate the checkpoint on your own workload before deployment.
License and attribution
Apache-2.0, following the source checkpoint. Smaug-Mini is published by Abacus.AI; this quantization is an independent community conversion.
- Downloads last month
- 43