Note on parameter count. Hugging Face reports ~5B parameters from the safetensors metadata, inflated due to how AWQ4 stores each group of 8 int4 weights packed into a single int32 tensor. Actual disk and VRAM footprint of the model is roughly (4.3 GB).
AWQ4-ARG
AWQ4 W4A16 quant of Voxtral-Mini-3B-2507 specialized for Argentine + LatAm Spanish.
When to pick this variant
Use this if your production audio is mostly Spanish (specially Argentine). It matches the bf16 baseline WER on Argentine Spanish and beats plain AWQ4 by 38% on Argentine.
Benchmarks vs bf16 baseline
Data: FLEURS + Common Voice 15 + Multilingual LibriSpeech + ylacombe/google-argentinian-spanish (es-AR). Hardware: 2x RTX 4090, vLLM 0.25. Bold = best 4-bit variant.
| Metric | Voxtral-Mini bf16 (baseline) | AWQ4-ARG | AWQ4-EN | AWQ4-GEN |
|---|---|---|---|---|
| WER es | 5.82 | 4.68 | 4.60 | 4.55 |
| WER de | 6.06 | 7.13 | 6.66 | 8.16 |
| WER fr | 8.01 | 10.59 | 6.70 | 9.33 |
| WER it | 6.50 | 7.02 | 7.01 | 6.94 |
| WER pt | 7.16 | 7.21 | 6.14 | 6.08 |
| WER nl | 9.24 | 7.95 | 11.29 | 10.59 |
| WER en (no MLS) | 6.65 | 7.00 | 6.96 | 7.01 |
| WER avg-6 | 7.13 | 7.43 | 7.07 | 7.61 |
| WER es-AR | 4.69 | 5.00 | 5.11 | 5.04 |
| VRAM (weights) | 8.8 GB | 4.36 GB | 4.36 GB | 4.36 GB |
| Throughput (tok/s, higher=better) | 170.9 | 185.0 | 186.3 | 186.5 |
Usage (vLLM)
# 1. In your vLLM venv, add the deps Voxtral needs
uv pip install --python /path/to/vllm/venv/python "mistral-common[soundfile]"
# 2. Serve
vllm serve JuanGiner/AWQ4-ARG \
--tokenizer-mode mistral \
--config-format mistral \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--port 8000
Then send audio to POST /v1/audio/transcriptions in OpenAI-compatible format.
Calibration
500 audio samples (250 Argentine from ylacombe/google-argentinian-spanish, 250 LatAm from google/fleurs es_419) passed through the audio encoder — captures real prefill activations.
Quantization was done with llmcompressor 0.10 (AWQ modifier, W4A16 scheme, group_size=128, duo_scaling=True),
targeting Linear modules inside LlamaDecoderLayer. The compressed-tensors format matches what vLLM 0.25.x expects
for AWQ models (weights re-fused for the vLLM Voxtral loader).
Base model & license
Base: mistralai/Voxtral-Mini-3B-2507, Apache-2.0.
This quantized derivative is released under the same license.
- Downloads last month
- 526
Model tree for juanginer/Voxtral-Mini-3B-2507-AWQ4-ARG
Base model
mistralai/Voxtral-Mini-3B-2507