Speculative Decoding for Vision-Language Models with Quantized Twin Drafts
An int4 quantization of a vision-language model serves as its own speculative draft inside vLLM. Nothing is trained and no component beyond the published quantized checkpoint is added. The strict tier reproduces the target's greedy output exactly, and the relaxed tier additionally accepts a drafted token whose target logit stays within ln 2 of the argmax, so the target remains the judge of every emitted token. The full experiment report, the failed attempts included, is in VLM_SPEC_DECODING_REPORT.md.
Repository layout
| Path | Content |
|---|---|
VLM_SPEC_DECODING_REPORT.md |
Experiment report E1 to E15, the primary document |
code/ |
Benchmark harness (stage1_spec_smoke.py), acceptance measurement, launch chains, analysis scripts |
vllm_patch/ |
The four patched vLLM 0.27.1 files plus the consolidated diff mrope_draft_spec.patch |
data/ |
MathVista testmini and MM-Vet prompt parquets with the prompt id lists used in every run |
results/ |
Raw per-prompt results of every experiment as jsonl |
models/ |
Unmodified mirrors of the four 8B checkpoints used in the experiments |
Models
The checkpoints under models/ are byte-identical mirrors, packaged here so the whole project resolves from one place. The originals are Qwen/Qwen3-VL-8B-Thinking, Qwen/Qwen3-VL-8B-Instruct, cyankiwi/Qwen3-VL-8B-Thinking-AWQ-4bit, and cyankiwi/Qwen3-VL-8B-Instruct-AWQ-4bit, all released under Apache 2.0, and their pages remain the authoritative source. The 32B experiments in the report use Qwen/Qwen3-VL-32B-Thinking and cyankiwi/Qwen3-VL-32B-Thinking-AWQ-4bit, which are not mirrored here.
Setup
The engine is vLLM 0.27.1, patched in place. Install it, then overwrite four files in site-packages/vllm with the versions in vllm_patch/.
pip install vllm==0.27.1 pyarrow
V=$(python -c "import vllm, os; print(os.path.dirname(vllm.__file__))")
cp vllm_patch/rejection_sampler.py $V/v1/sample/rejection_sampler.py
cp vllm_patch/gpu_model_runner.py $V/v1/worker/gpu_model_runner.py
cp vllm_patch/llm_base_proposer.py $V/v1/spec_decode/llm_base_proposer.py
cp vllm_patch/utils.py $V/model_executor/layers/quantization/compressed_tensors/utils.py
The patch adds draft-model speculation for M-RoPE models, which stock vLLM refuses with a NotImplementedError, and fixes a loader bug that made the quantization path reach into the draft's vision tower. The patched gpu_model_runner.py also captures the drafter's piecewise CUDA graphs during the capture phase, which upstream skips, leaving the drafting loop at eager speed. Note that mrope_draft_spec.patch predates this last change, so prefer the whole files over the diff.
Running
One process measures one leg. The harness decodes each prompt greedily and writes one json record per prompt with the decoded token ids and the decode rate.
# vanilla reference
python code/stage1_spec_smoke.py --mode vanilla \
--model models/Qwen3-VL-8B-Thinking \
--parquet data/testmini.parquet --pids-from data/res05_qwen_thinking.jsonl \
--n 20 --max-tokens 1024 --gpu-mem-util 0.6 --max-model-len 3584 --out van.jsonl
# strict tier, gamma 6
python code/stage1_spec_smoke.py --mode spec --gamma 6 \
--model models/Qwen3-VL-8B-Thinking --draft models/Qwen3-VL-8B-Thinking-AWQ-4bit \
--parquet data/testmini.parquet --pids-from data/res05_qwen_thinking.jsonl \
--n 20 --max-tokens 1024 --gpu-mem-util 0.6 --max-model-len 3584 --out strict.jsonl
# relaxed tier: set VLLM_SPEC_RELAX_LOGTAU=0.693 (tau 0.5) in the environment
Three environment switches control the patched engine. VLLM_SPEC_RELAX_LOGTAU turns on the relaxed acceptance gate, VLLM_SPEC_DRAFT_BLIND_MM lets a draft decode without image embeddings, and VLLM_SPEC_DISABLE_DRAFT_CUDAGRAPH restores the upstream behavior of leaving the drafting loop uncaptured. The launch chains in code/ show the bracketed measurement protocol, with vanilla legs before, between, and after the speculative legs.
Results format
Each line of a results/*.jsonl file is one prompt with fields pid, mode, gamma, in_len, gen_len, t, tok_s, and gen_ids. Speedups in the report are medians over per-prompt paired ratios against the bracketing vanilla legs, with interquartile ranges. File prefixes map to experiments in the report, for example og8_* is the 8B CUDA graphs ladder of E11 and ogi2_* is the Instruct MM-Vet study of E15.