Qwen3.5-9B — ROCmFP4 / ROCmFPX GGUF

AMD-native FP4 / FP8 GGUF builds of Qwen/Qwen3.5-9B for RDNA3.5 / Strix Halo (gfx1151). Multimodal — mmproj-BF16.gguf included.

Variants

file ftype size decode spread
Qwen3.5-9B-Q4_0_ROCMFP4_COHERENT.gguf 102 5.19 GiB 39.34 t/s 1.0036
Qwen3.5-9B-Q6_0_ROCMFPX_AGENT.gguf 114 7.90 GiB 26.23 t/s 1.0004
Qwen3.5-9B-Q8_0_ROCMFPX.gguf 111 8.67 GiB 23.92 t/s 1.0013
Qwen3.5-9B-Q8_0_ROCMFPX_AGENT.gguf 115 8.77 GiB 23.78 t/s 1.0004

Measured on an idle Ryzen AI MAX+ 395 (Strix Halo, gfx1151, ROCm 7.2.4): -ngl 999 -c 4096 -fa on -fit off -np 1, 300-token generations, 12 samples with two warm-ups on the same prompt as the measurement. Spread = slowest/fastest.

⚠️ An earlier pass of these same files, taken while other jobs shared the GPU, read 20% low with 20%+ spread. On this hardware a co-resident job is the single largest source of benchmark error — measure on an idle box or say what else was resident.

Vision verified 4/4 on a four-quadrant colour image (red / blue / yellow / green) with the bundled mmproj-BF16.gguf. ⛔ Vision needs -fa off.

Head protection

tie_word_embeddings: false — unlike the 0.8B/2B/4B in this family, output.weight is present, so --output-tensor-type does real work here. Both the head and the embedding sit at q6_K on the 4-bit, which is why it lands at 4.97 BPW rather than ~4.7: two 248320x4096 matrices are a large share of a 9B file.

No MTP/draft head ships with the source, so no speculative-decoding numbers are claimed.

Verification

Every artifact was loaded on real hardware and checked for: exact stat bytes vs the --dry-run projection (a constant header delta; a varying one means truncation), the actual token_embd / output.weight types, three correctness answers asserted against content + reasoning with finish_reason recorded, and a decode median.

Credits

FP4/FP8 tensor types from the ROCmFPX fork of llama.cpp. These types do not exist in mainline llama.cpp — a ROCmFPX-capable build is required to load them.

Downloads last month
34
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/Qwen3.5-9B-ROCmFP4-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(460)
this model