Qwen3.8-Flash-Next ROCmFP4-FAST imatrix GGUF

The only quant of this model that fits fully in VRAM on a 128 GB unified-memory device without giving up quality: 87.06 GiB, 4.23 bpw, within 2.5% perplexity of the unquantized model (see below) — and it beats every alternative under 110 GiB we could find. Sized for the 96 GiB VRAM carve-out of a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S). The whole model runs resident on GPU: nothing falls back to host RAM or CPU compute, not the experts, not the 51.2 B-parameter n-gram table.

Two layouts, same weights, same size (87.06 GiB either way — splitting the n-gram table per head is a byte-for-byte restructuring, not a re-quantize):

  • root — n-gram table split per head, fully VRAM-resident.
  • joined/ — n-gram table as one tensor, portable; needs --ngram-on-disk or host RAM for that tensor, since a single tensor that size is past what most Vulkan devices accept as one buffer.

Setup

Qwen3.8-Flash-Next itself is upstream (ggml-org/llama.cpp#27742). This fork is still needed for the ROCmFPx quant types and the per-head PLE layout.

git clone https://github.com/LaurentZuijdwijk/llama.cpp
cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Run

# per-head table, fully on the GPU
./build/bin/llama-cli \
  -m Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix.gguf \
  -ngl 99 -c 32768

# joined table, off the GPU and off host RAM
./build/bin/llama-cli \
  -m joined/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-joined.gguf \
  -ngl 99 --ngram-on-disk --ngram-cache 8192 -c 32768

--ngram-cache defaults to 256 MiB; raise it for long generations or throughput drops off over the course of a conversation.

The two layouts are interchangeable: gguf_split_ple_heads.py converts one to the other by copying quantized bytes verbatim — no dequantize, no requantize, no quality change.

Perplexity

wikitext-2 raw, 145 chunks at -c 2048.

build PPL vs. reference
unquantized reference (as reported in PR 27742) 4.0068 +/- 0.02271 -
this file 4.1062 +/- 0.02329 +2.48%

Against AesSedai's quants, compared the fair way (each build's PPL against its own measured reference, since their test methodology differs from ours):

build size PPL ratio vs. own reference
AesSedai IQ3_S 107.38 GiB +6.10%
this file 87.06 GiB +2.48%
AesSedai IQ4_XS 117.13 GiB +3.12%
AesSedai Q4_K_M 135.38 GiB +0.61%

Beats their IQ4_XS and IQ3_S on quality at a smaller size.

MTP draft head

Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST-GGUF holds the model's own multi-token-prediction head, exported from the same checkpoint. Currently mismatched with this file's expert precision (draft is flat FP4, this file's gate/up are Q4_K) — needs requantizing before it's worth enabling here.

Credits

qwen4exp support is the work of Daniel Han (@danielhanchen), from ggml-org/llama.cpp#27742, merged upstream. This fork is only still needed for what's listed under Setup above.

Quant formats hand-ported from ciru-ai/ROCmFPX. Calibration corpora from bartowski and Thireus, credited above. Base model by the Qwen team.

Not included

Vision tower.

License

Qwen Community License 1.0, included as LICENSE.

Downloads last month
421
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF

Quantized
(126)
this model