Qwen3.8-Flash-Next — W4A16 (Intel Arc Pro B70 / XPU)
W4A16 (int4 group-128, compressed-tensors) quantization of Qwen3.8-Flash-Next (125B MoE / 6B active). Derived from official BF16 weights. Single-quant int4; no FP8 double-quant.
This repository provides a community quantization and XPU serving recipe. The base model, architecture, and license are provided by Qwen (Qwen Community License 1.0). Refer to the base model card for the full architecture description.
Served as the multimodal wrapper Qwen4ExpForConditionalGeneration (text + image input via the Qwen3-VL processor).
Changes from base model
- Expert weights quantized to int4 group-128 (W4A16) via
compressed-tensors. quantization_configignore list maintains these in BF16:lm_head,embed_tokens,mtp.*,ple.*,visual.*,*.gate,hyper_connection*,indexer*,linear_attn.*,shared_expert*.- 17 safetensors shards, ~77GB total.
Hardware requirements (confirmed)
Configuration used for benchmarking:
- 4x Intel Arc Pro B70 (Xe3), 32GB each = 128GB total VRAM
- >=128GB host RAM (247GiB on reference system) — required for the PLE n-gram table
- oneAPI Level-Zero 20.2.0
Performance measured on this hardware (single stream, 512 tokens, greedy):
- 53.4 tok/s decode, TTFT ~0.11s
- 256K context, fp8 KV cache
Memory distribution
Memory is split across two pools for different purposes:
| Pool | Sized for | Contents |
|---|---|---|
| VRAM (4x32GB = 128GB) | weights + KV cache | 77GB W4A16 weights (sharded 4 ways via TP+EP, ~19GB/GPU), fp8 KV cache, activations, and XPU graph workspace |
| Host RAM (>=128GB) | PLE table | 96GB n-gram table (memory-mapped, shared page cache across all 4 ranks) |
- The PLE table resides in host RAM. It is mmap'd from host RAM. Each decode step transfers only 16 rows/token (~few KB) to the GPU.
- A system with 128GB VRAM but only 64GB RAM will hold the weights but fail to map the table. Both ~128GB VRAM and >=128GB RAM are required.
--gpu-memory-utilization 0.85provides VRAM headroom for the 256K context KV cache.
Dependency: patched vLLM
Stock vLLM is not compatible. This model requires a patched vLLM with the qwen4_exp model port and XPU W4A16 / GDN kernels:
- vLLM base:
vllm-project/vllm@c39076fef - torch
2.13.0+xpu - vllm-xpu-kernels
0.1.12(pinned;0.1.13.2is broken)
The port is a 16-file patch, published and verified:
- Fork branch (installable):
devan-carlin/vllm@xpu-qwen4exp—pip install "vllm @ git+https://github.com/devan-carlin/vllm.git@xpu-qwen4exp"(or run the one-shot installer below) - One-shot installer:
setup-vllm-xpu.sh— rustup + venv + torch + kernel pin + build, in one command - Patch file:
qwen4exp-xpu-port.patch— git-apply diff vs the clean base - Operations guide (clean rebuild, bug log, launch config):
qwen4exp-vllm-operations.md
The patch was verified end-to-end: a fresh venv built from the clean base + patch reproduces the reference build (53.4 tok/s on 4x Arc Pro B70).
PLE n-gram table (required, included in repo)
The model uses an N-gram Embedding (PLE) layer backed by a ~96GB hash table. The table is required for correct output. Disabling it (QWEN4EXP_DISABLE_PLE=1) will result in incorrect output.
This repository includes the table (ple_table_qwen4exp.pt) for out-of-the-box use. It is memory-mapped from host RAM and shared across all tensor-parallel ranks via the page cache; it is never loaded onto the GPU.
- Set
PLE_TABLE_PATHto theple_table_qwen4exp.ptfile in this repo. - To rebuild the table (e.g., after a base-model update), run the
phase_b_ple_table_prep.pyscript on a machine with sufficient RAM, then pointPLE_TABLE_PATHto the result.
Launch
python -m vllm.entrypoints.openai.api_server \
--model <this-repo> \
--served-model-name qwen-256k \
--host 0.0.0.0 --port 8000 \
--tensor-parallel-size 4 \
--enable-expert-parallel \
--dtype bfloat16 \
--max-model-len 262144 \
--max-num-seqs 4 \
--gpu-memory-utilization 0.85 \
--kv-cache-dtype fp8 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--generation-config vllm \
--override-generation-config '{"temperature": 0.7}'
Environment: VLLM_XPU_ENABLE_XPU_GRAPH=1, UR_L0_SYNC_MODE=BLOCKING, VLLM_WORKER_MULTIPROC_METHOD=spawn, PLE_TABLE_PATH=/path/to/ple_table.pt.
Notes
- Multimodal. The vision tower (27-block ViT, byte-identical to Qwen3-VL) is included and served via the Qwen3-VL processor. Send images as
image_urlcontent parts (base64 or URL); the model describes them. Text-only prompts work unchanged. - Serving tip — run at temperature 0.7, not 1.0. Qwen's published guidance for Qwen3.8-Flash-Next is temperature 1.0 (the
generation_config.jsonin this repo ships 1.0), but on this checkpoint we observe strange behavior well above 0.7: repetition, incoherent output, and degenerate loops that set in quickly. Greedy (temp 0) is also bad — it tends to lock into repetitive loops. 0.7 is the sweet spot (with high reasoning effort); the launch command above pins it via--override-generation-config. If a client sends its owntemperature, that overrides the server default — keep it at or below ~0.7. - MTP speculative decoding is not wired into the port's forward pass (the
mtp.*weights are present but skipped at load), and is unreliable on XPU due to GDNcausal_conv1dkernel limitations. Keep it disabled.
License
Qwen Community License 1.0 (see LICENSE). Derivative distribution is permitted. Display the model name if MAU >100M or monthly revenue >$20M. MaaS / AI-assistant commercial use requires a separate license from Qwen.
- Downloads last month
- -
Model tree for devan-carlin/Qwen3.8-Flash-Next-W4A16
Base model
Qwen/Qwen3.8-Flash-Next