Qwen3-VL-2B-Instruct β ExecuTorch
A vision-language model as one .pte with three entry points. Ask it about a photograph on
the device; nothing leaves it.
- Source: Qwen/Qwen3-VL-2B-Instruct β 2.13B parameters, 28-layer decoder, 24-layer vision tower
- License: apache-2.0
- Input: a 512x512 RGB photograph and a question
- Output: text
Methods
| method | takes | returns |
|---|---|---|
vision_encoder |
pixel_values [1024, 1536] |
merged [256, 2048] and three deepstack maps [256, 2048] |
token_embeddings |
ids [1, seq] |
rows [1, seq, 2048] |
text_model |
rows, positions [3, 1, seq], slots [1, seq], three deepstack maps [1, seq, 2048] |
logits [1, seq, 151936] |
Three methods rather than three files because ExecuTorch copies a mutable buffer into every method that names it, so a prefill method and a decode method would each get their own KV cache and neither would see the other's writes. One graph serves a 277-token prompt and a single token.
What the caller owns
Three things, and getting any of them wrong does not throw β it answers about the wrong thing:
- Splice the 256 vision rows into the prompt embedding at the image token positions.
- Scatter the three deepstack maps to those same positions, zero everywhere else.
Qwen3-VL adds features from vision blocks 5, 11 and 17 into decoder layers 0, 1 and 2,
and only at image positions. Upstream writes that as
hidden_states[visual_pos_masks, :] += embeds, which does not export; here the maps arrive already placed, so the graph just adds them. - Pass Qwen's own 3-D positions from
get_rope_index, notarange. The rope is interleaved M-RoPE withmrope_section [24, 20, 20]β time, height, width. A 1-D position array gives every image patch the same angle and reads as a mild degradation rather than an error.
convert/run_qwen3_vl.py does all three in 40 lines.
Verification (Mac arm64, executorch 1.4.0, torch 2.13.0)
Parity is the last-row logits of a real prompt against Qwen3VLForConditionalGeneration in
eager fp32, per question, because a model of this size can hold a scene and lose the small
print. Both questions are about the same photograph.
| build | file | size (MB) | corr, "name on the sign" | corr, "describe it" | vision (ms) | prefill (ms) | decode |
|---|---|---|---|---|---|---|---|
| 8da8w | qwen3_vl_2b_instruct_xnnpack_8da8w.pte |
2453.1 | 0.993435 | 0.992611 | 568.7 | 590.6 | 17.1 tok/s |
| 8da4w (not shipped) | qwen3_vl_2b_instruct_xnnpack_8da4w.pte |
1697.5 | 0.739964 | 0.948887 | 566.7 | 739.4 | 19.8 tok/s |
The prompt is 277 tokens, of which 256 are the picture. vision_encoder runs
once per image, text_model once per prompt and once per generated token.
What it says
Asked for the name on a pub sign, against the eager model's "Lamb & Flag":
- 8da8w β Based on the image provided, the name of the pub is The Lamb & Flag.
This name is visible on the sign above the entrance to the establishment. The sign also includes the address "
- 8da4w β Based on the text visible on the sign above the entrance of the establishment, the name of the pub is Pine Purpose.
The sign is located on the front of the building, just above
The rewrites, and what each was checked against
Neither this repo's generic VL bundle nor export_llm reaches Qwen3-VL. Five things had to
change, each measured against the module it replaced:
| what | why | checked |
|---|---|---|
| vision constants precomputed | position embeddings are bilinearly interpolated from grid_thw on every call |
part of the 0.000e+00 below |
| variable-length attention β plain SDPA | one image is one sequence, so cu_seqlens is [0, 1024] |
exact |
patch embedding Conv3d β F.linear |
portable convolution takes 3-D or 4-D only: Expect input tensor to be 3-D or 4-D, but got, 5 |
max_abs_diff 4.9e-04 |
| deepstack mask-assignment β add | boolean-mask assignment does not export | first-step logits corr 1.000000 |
| embedding table quantised by hand | see below | corr 0.999965 |
With the first two in place and no quantisation, the vision wrapper reproduces
Qwen3VLVisionModel at max_abs_diff 0.000e+00 across all four outputs, and the decoder
wrapper reproduces the full model's first-step logits at corr 1.000000.
Two traps worth knowing
The token embedding table escapes quantisation. It is an nn.Embedding, so a filter
written as isinstance(module, nn.Linear) skips it β and tie_word_embeddings does not
save it, because quantize_ swaps lm_head.weight for a new tensor while the embedding
keeps pointing at the old fp32 one. At 151,936 x 2048 that is 1.24 GB left in fp32: the
first export of this model came out at 2630 MB. torchao's own IntxWeightOnlyConfig does
reach an nn.Embedding but will not lower β Missing out variants: {'torchao::dequantize_affine'} β so the quantisation is written out by hand, one scale
per row.
512x512 is where small text lives or dies, and the resampling filter decides. Asked for the name on the sign, the eager fp32 model answers:
| resize filter | eager fp32 |
|---|---|
| BICUBIC | "Lamb & Flag" β right |
| BILINEAR | "The Lamb & Flag" β right |
| LANCZOS | "The Pigeon House" β wrong |
| NEAREST | "Pompadour" β wrong |
That is before any quantisation. If fine print matters for your use, the preprocessing is
not a detail; convert/export_qwen3_vl.py puts the resize behind one helper for that reason.
Not shipped
8da4w β the same bundle with a 4-bit decoder β is 1697.5 MB against 2453.1 MB, and it does not hold: logits correlation 0.740 on the sign question against 0.993, and it reads the sign as "Pine Purpose". Measured where the loss is, one arm at a time: the vision tower at int8 costs corr 0.990 against fp32's 0.99982, the hand-quantised table costs 0.99997, and the 4-bit decoder is the rest. Reported here rather than dropped, because a 1.7 GB 2B vision-language model is worth wanting and this is what it costs.
Conversion
python convert/export_qwen3_vl.py # check both halves against eager
python convert/export_qwen3_vl.py --export # write the bundle
python convert/check_qwen3_vl.py 8da8w # parity and speed
python convert/run_qwen3_vl.py photo.jpg "what is on the sign?"
(conversion scripts: executorch-models)
- Downloads last month
- -
Model tree for mlboydaisuke/Qwen3-VL-2B-Instruct-ExecuTorch
Base model
Qwen/Qwen3-VL-2B-Instruct