Qwen3.8-Flash-Next W4A16 Modular (296E)

Expert-pruned, expert-modular repack of the Intel AutoRound W4A16 build of Qwen3.8-Flash-Next: 512 → 296 experts per layer, selected by profiling real agent traffic. Serves on a single 96 GB GPU with a 262K context window.

How this was made

  1. Base: official Qwen/Qwen3.8-Flash-Next → Intel AutoRound W4A16 (calibrated int4).
  2. Profile: ~80 real coding-agent sessions were replayed against a server running SGLang's --expert-distribution-recorder-mode stat, producing per-(layer, expert) route counts. A second profile was taken on anomalous contexts (the failure sessions we wanted the model to keep handling well).
  3. Keep set: per layer, top-294 experts from daily traffic ∪ top-2 from the anomaly profile → 296/layer. Covers 95%+ of routine routing and preserves self-healing on degraded contexts (measured 5/5 recovery vs 0/3 for the daily-only set).
  4. Modular repack: all expert tensors of layer N were re-packed into one experts-L{NN}.safetensors; everything else (dense, GDN linear attention, PLE, embeddings, lm_head) into 14 large backbone shards. No tensor values were modified — this is a pure re-chunking of the same weights, verified key-for-key identical (129,403 tensors) with serving parity (104 tok/s).

Why it works (principle)

MoE layers route each token to only top-k of num_experts experts (here 10 of 512). Routine traffic concentrates on a small subset per layer, so removing never-routed experts is lossless for that workload; the anomaly-profile union buys back robustness for edge contexts. Expert weights live in GPU memory only when kept, so pruning 512→296 frees ~13 GB VRAM → +31% KV pool.

Changes vs the previous release (Pruned-294E)

Pruned-294E Modular (this)
Experts/layer 294 (daily profile only) 296 (+2/layer from anomaly profile → self-heal)
Layout 23 linear shards (~4.5 GB) 48 expert files (~180 MB) + 14 backbone shards
Editing one layer's experts rewrite a 4.5 GB shard + index rewrite one 180 MB file
Quality gates passed passed (incl. anomalous-context suite)

Same base weights, same quantization; day-to-day quality and speed are identical.

Modular layout

Pattern Count Size each Contents
experts-L00..47.safetensors 48 ~180 MB all MoE expert tensors of one layer (w13/w2 qweight/scales/qzeros)
backbone-000XX.safetensors 14 ~4.4 GB dense / GDN linear-attention / attention / PLE / embeddings / lm_head

Swapping, adding or re-pruning experts for layer N only rewrites experts-LNN.safetensors plus config.json (num_experts) and the index weight_map — the backbone never changes.

Want a keep set tuned to your traffic? Profile it in one command and serve pruned without re-exporting: sglang-expert-profile (CLI + community keep-sets; the serving-side keep-mask lives in the sglang fork, ours/main).

Requirements

  • GPU VRAM >= 64 GB (weights ~45 GB; 96 GB recommended for full 262K context)
  • Host RAM >= 64 GB (PLE embedding offload: --ple-offload-embedding)
  • CUDA 13 stack

Serving (SGLang fork with PLE offload + marlin GC fix)

python -m sglang.launch_server \
  --model-path ./Qwen3.8-Flash-Next-W4A16-Modular \
  --chat-template ./qwen3_coder_template.jinja \
  --ple-offload-embedding \
  --moe-a2a-backend none \
  --linear-attn-prefill-backend triton --linear-attn-decode-backend triton \
  --mamba-ssm-dtype bfloat16 \
  --context-length 262144 --mem-fraction-static 0.93 \
  --reasoning-parser qwen3 --tool-call-parser qwen3_coder

Sampling defaults ship in generation_config.json (temp 0.7 / top_p 0.95 / top_k 20, no penalties — Flash-Next is penalty-sensitive, see our notes in the sglang fork). No need to pass them per request.

Note: without the marlin GC patch, loading OOMs at ~91.5 GB on some stacks (gptq_marlin_moe_repack int4→int32 expansion). Patch + details: https://github.com/ranxianglei/sglang (ours branch).

Performance (single RTX Pro 6000 96GB)

  • single stream ~104 tok/s decode @ 262K context
  • aggregate (w48) ~2100 tok/s
  • KV pool: ~856K tokens bf16

Pairs well with billion-context (ACP)

This model's 262K window + huge KV pool makes it an excellent host for our context-compression plugins — long agent sessions stay coherent while effective context grows far beyond the window:

  • billion-context — protocol-level ACP context compression for AI coding agents (drop-in for OpenCode & friends)
  • billion-context-pi — pi/agent integration of the same compression engine

Together: Flash-Next serves the window, billion-context compresses into it — day-long coding agents on one consumer GPU.

Provenance

Base: Qwen/Qwen3.8-Flash-Next → Intel AutoRound W4A16 (calibrated) → expert pruning 512 → 296 by routing-profile keep-set → modular repack (this repo). 294-expert variant (linear shards): ranxianglei/Qwen3.8-Flash-Next-W4A16-Pruned-294E

Downloads last month
28
Safetensors
Model size
66B params
Tensor type
I32
·
F16
·
I64
·
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support