FranQwen3.8-27B-NVFP4-DFlash2
A Frankenstein Qwen3.8-27B. One model's body, another model's head, and a third model's drafter. The name is honest about what it is.
The point of the stitching is speed. Stock NVFP4 checkpoints can't run DFlash2 speculative decoding at all.
| FrankQwen3.8-27B | Base NVFP4 Qwen3.8-27B | |
|---|---|---|
| 1 node | 36.3 tok/s | 24.1 tok/s |
| 2 nodes (TP=2) | 56.3 tok/s | 40.1 tok/s |
Tested on DGX Sparks
Abliterated. Refusal behaviour was removed upstream and this doesn't put it back. No safety tuning. See Safety.
The parts
Three donors, each contributing one thing:
| part | donor | what it brings |
|---|---|---|
| Body (all 48 layers, vision tower) | orcarouter/Qwen3.8-27B-Uncensored-NVFP4 |
the actual model. Uncensored, already quantized to NVFP4 (W4A4, group 16). Also, unexpectedly, the fastest option — see below |
Head (lm_head) |
sakamaki's Qwen3.8-27B-MTP-NVFP4 | a genuine BF16 lm_head, never quantized. This is the piece that makes DFlash2 possible |
| Drafter (separate download) | z-lab/Qwen3.8-27B-DFlash2 |
the block-diffusion draft model that proposes tokens |
Why swap the head
DFlash2 needs an unquantized lm_head. NVFP4 checkpoints quantize theirs, so
vLLM simply refuses to start with a DFlash2 config.
You might think you could just convert the head back to BF16. That doesn't work. Converting back gives you the same rounding damage in twice the storage, so the drafter is still guessing against a degraded output distribution. What's needed is a head that was never quantized in the first place, which is what sakamaki's checkpoint provides.
Is that safe to do?
Checked before merging, not after. The donated head and the body's own head agree to a cosine similarity of 0.999647, with the leftover difference sitting exactly at quantization-rounding size. If abliteration had modified the head you'd see a much bigger, structured difference.
The quality results below confirm it behaviourally too.
Speed
Measured with llama-benchy, the tool
Spark Arena.
Hardware tested is NVIDIA DGX Spark (GB10).
One node
| concurrency | generation (tok/s) | prompt processing (tok/s) |
|---|---|---|
| 1 | 36.34 | 2388 |
| 2 | 57.80 | 2146 |
| 5 | 125.46 | 1444 |
| 10 | 161.43 | 1259 |
Two nodes, TP=2 over ConnectX-7
| concurrency | generation (tok/s) | prompt processing (tok/s) |
|---|---|---|
| 1 | 56.26 | 3887 |
| 2 | 87.03 | 2660 |
| 5 | 118.18 | 1814 |
| 10 | 153.25 | 1544 |
Going from one node to two gives about 1.5x on single-stream generation, and KV capacity grows from 900k to 3.8M tokens.
For comparison, the best published two-node Qwen3.8-27B result on Spark Arena for the same benchmark is 40.09 tok/s at concurrency 1, and 115.04 at concurrency 5.
A few honest notes. Single-stream speed bounces around by roughly 12% run to run, so use an average of several runs rather than your best one. Speed also depends a lot on the workload: code prompts hit 55.3 tok/s on one node because the drafter guesses code far better than prose.
Quality
Run with lm-evaluation-harness on full task sets, speculative decoding off:
| task | this model | body, unmodified | head donor |
|---|---|---|---|
| arc_challenge (acc) | 0.5606 | 0.5572 | 0.5725 |
| arc_challenge (acc_norm) | 0.5853 | 0.5768 | 0.5990 |
| gsm8k 5-shot (flexible) | 0.7672 | 0.7657 | 0.6156 |
| gsm8k 5-shot (strict) | 0.7324 | 0.7437 | 0.5686 |
| truthfulqa_mc2 | 0.5293 | 0.5293 | 0.5595 |
| winogrande | 0.7459 | 0.7498 | 0.7466 |
Two things worth pulling out.
The graft is free. Against the unmodified body, every difference is smaller than the measurement error. Three metrics up, two down, one identical to four decimal places.
It takes the head and nothing else. Look at the head donor's column: that checkpoint is 15 points worse at gsm8k. If grafting dragged across the donor's characteristics, this model's gsm8k would have sagged toward 0.6156. It didn't move at all.
It also scores 44/46 on a private 46-question reasoning set, tying the best of every Qwen3.8-27B checkpoint tested. Vision still works, including with DFlash2 active.
An accidental finding: abliteration makes the drafter's job easier
Worth writing down because it wasn't expected.
The same BF16 head was grafted onto a non-abliterated NVFP4 Qwen3.8-27B and benchmarked identically. Same drafter, same settings, same everything except the body.
| concurrency | abliterated body (this model) | stock body |
|---|---|---|
| 1 | 35.12 | 28.21 |
| 2 | 62.66 | 47.42 |
| 5 | 101.04 | 86.42 |
| 10 | 145.60 | 130.27 |
Generation is 11–24% faster on the abliterated body, while prompt processing is identical to within a rounding error. Identical prefill and different decode points at one thing: the drafter's suggestions get accepted more often. The abliterated model is simply more predictable to guess.
Quality is the same either way, so if you want a non-abliterated build, the same graft works fine on a stock body and costs you roughly 20% throughput.
Running it
Needs vLLM with DFlash2 support (PR #52816, in mainline since 2026-08-21) and the drafter downloaded separately. Only tested on GB10.
vllm serve <this-model> \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.6 \
--max-model-len 262144 \
--max-num-batched-tokens 16384 \
--max-num-seqs 8 \
--trust-remote-code \
--enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml
Use 7 speculative tokens, not 8. Eight is accepted but gave no measurable gain and destabilised things at higher concurrency.
Give it max_tokens of at least 8192 for real work. It's a reasoning model, and
with a small budget the thinking eats the whole allowance and you get empty
output back.
Safety
This model is abliterated. Refusal behaviour was removed upstream, and swapping the head doesn't restore it, because the modified weights are in the body.
It will go along with requests that the official Qwen3.8-27B declines, including harmful ones. It has had no safety tuning, and scoring well on reasoning benchmarks says nothing about how it behaves. If you're putting this anywhere users can reach it, add your own filtering and evaluate it for your own use case.
Credits
Qwen/Qwen3.8-27B, the base modelorcarouter/Qwen3.8-27B-Uncensored-NVFP4, for the abliterated NVFP4 body- sakamaki, for the BF16
lm_head z-lab/dflash, for DFlash2 and the draftereugr/llama-benchyand Spark Arena, for the benchmark harness
Everything upstream is Apache-2.0, and so is this.
- Downloads last month
- 87