- Qwen3.8-Flash-Next NVFP4 on 2Γ DGX Spark β field notes
- Status
- Performance
- Recipe
- Stack (pin these when reproducing)
- The five things this repo adds
- 1.
SGLANG_HOST_IPβ multi-node SGLang deadlocks behind a default-deny firewall - 2. Page cache is half your memory budget β and it is the wedge
- 3. Speculative decoding is at its architectural ceiling β and the drafter saturates it
- 4. Vision works, and the self-review loop closes
- 5. Thinking mode: binary, helps reasoning, and fails catastrophically 30% of the time
- 1.
- Sampling: a 1-in-5 repetition loop at temperature 0 β cause NOT established
- Capability evaluation
- Dead ends β documented so you don't spend the time
- Thermals β much cooler than a comparable dense-ish MoE
- Benchmark discipline
- Files
- Credit
- License
- Status
Qwen3.8-Flash-Next NVFP4 on 2Γ DGX Spark β field notes
Mirror of https://github.com/beastllama/dgx-spark-qwen38-flash-next-recipe β the base recipe is MiaAI-Lab's; this is the config delta plus everything that went wrong on the way and the measurements nobody had published. From the same homelab as the GLM-5.3-Flash + DFlash2 recipe.
Serving Qwen3.8-Flash-Next-NVFP4 across two DGX Sparks (GB10 / sm_121) with SGLang, TP=2 over ConnectX-7 RoCE (single rail β dual-rail is measured fabric capability, untested under SGLang). Vision enabled. 262,144 context. MTP speculative decoding.
Start here, then read the findings. The base stack is MiaAI-Lab's β this repo is the config delta on top of it plus everything that went wrong getting there and how it was fixed (see Credit). What follows was not in either published recipe: repeated node wedges (on unified memory, exhaustion does not error β it takes the whole box), a deadlock that only appears behind a default-deny firewall, the first speculative-decode acceptance measurements we're aware of for this model, and three optimisation avenues that turned out to be dead ends β documented as such, with numbers.
Everything here was measured on real hardware. Where a number is contested or unproven, it says so.
Status
| Serving | β TP=2 across 2 nodes, 262,144 context |
| Vision | β verified end-to-end (see below) |
| Spec decode | β
NEXTN 3/1/4 (num_steps/eagle_topk/num_draft_tokens) β note the engine self-reports speculative_algorithm: EAGLE β the architectural maximum, not a default |
| Decode | ~63 tok/s single-stream on real generation |
| Concurrency | 306 tok/s aggregate at 6 streams |
| Prefill | 3,050 tok/s (cache defeated) |
| Thermals | 52 Β°C / 35 W peak under load |
Conditions, because this repo insists on them: decode ~63 tok/s = real generation of a
10.7k-token HTML file, thinking off, temp 0.3, 2398 MHz idle / 2522 under load. 306 tok/s =
aggregate across 6 concurrent streams, code prompt, 400 max_tokens, ignore_eos. Prefill
3,050 tok/s = ~7,450-token unique prompt per run, cached_tokens=0 asserted, n=6. Thermals =
concurrency 4, 1 Hz sampling. A number without its prompt, token count and clock state is not
comparable to anything β including these.
Performance
Two DGX Sparks. 180B params (125B backbone + 51B PLE), NVFP4, 262k context, vision on.
| tok/s | conditions | |
|---|---|---|
| Single stream, real work | ~63 | 10.7k-token HTML page, natural stop, 2522 MHz |
| Single stream, 400 tok | 63.7 | code prompt, ignore_eos |
| 2 concurrent | 104.9 agg | 52.8/streamΒ² |
| 4 concurrent | 178.7 agg | 45.2/streamΒ² |
| 6 concurrent | 306.6 agg | 51.7/streamΒ² |
| Prefill | 3,050 | ~7,450-token unique prompt, cached_tokens=0 asserted, n=6 |
| Stress floor | 47.6 | ignore_eos + hard prompt + 800 tok β a deliberate FLOOR, see below |
Β² Concurrency measured before the config was pinned β at max_running_requests=12 and an
unpinned KV pool of 850,816 tokens, not the 8 / 600,000 in the recipe below. Under the pinned
config 8 is the cap, so 6 streams is near it rather than "still climbing". Re-measure before
quoting these against the shipped config.
Power and heat, at concurrency 4: 52 Β°C, 35.5 W peak per node; 42 Β°C / 10.4 W idle. That is
roughly half the draw of a comparably-sized dense-ish MoE we previously ran on the same boxes
(88 Β°C / 65 W), at higher clocks. Cause: 6B active params per token (10 of 512 experts) and only
12 of 48 layers are full attention β the rest are linear-attention GDN, so decode waits on memory
rather than burning watts. Practical effect: thermal guard stages sized for the older model are
unreachable by 43 Β°C, and two Sparks serve this at **71 W combined under load**.
Why two numbers for "single stream". ignore_eos benchmarks force generation past the model's
natural stopping point into degenerate text. They're excellent for regression detection and useless
as a headline. Real generation of a complete HTML page runs at ~63 tok/s; the same stack under
ignore_eos on a hard prompt reports 47.6. Both are correct. Quote the one that matches what
you're doing, and say which.
Speculative decoding is doing much of the work, but the delta is not isolated. Our earlier
vLLM deployment of the same checkpoint (no MTP, --enforce-eager) decoded at 20β21 tok/s; this
SGLang stack with MTP runs ~3Γ that. Engine, CUDA-graph mode and MTP all changed together β we
have no SGLang-with-MTP-off measurement, and neither published recipe ships one to compare against.
See Β§3 for why 3/1/4 is the ceiling.
Recipe
1. Base stack. Clone MiaAI-Lab's repo and follow it β it builds the SM121 QSA patch onto the public SGLang image and handles fabric preflight, worker-first ordering and readiness waiting. Everything below is a delta on that.
2. Stage NCCL on both nodes. Both published recipes treat host-staged NCCL as required for GB10 multi-node stability:
mkdir -p ~/nccl-2.30.7
cp /usr/lib/aarch64-linux-gnu/libnccl.so.2.30.7 ~/nccl-2.30.7/
ln -sf libnccl.so.2.30.7 ~/nccl-2.30.7/libnccl.so.2
3. Apply the config delta (table below) to the .env.
4. Pin the control plane to the fabric β the single most important line if you run a default-deny firewall, and the one that cost us the longest debug:
-e SGLANG_HOST_IP=<this node's fabric IP>
5. Evict page cache immediately before launch, and keep it bounded during the load:
# before launch (no root needed β this is what scripts/cache-warden.py automates)
python3 -c "
import os,glob
for p in glob.glob(os.path.expanduser('~/.cache/huggingface/hub/**/*.safetensors'),recursive=True):
rp=os.path.realpath(p)
if os.path.exists(rp):
fd=os.open(rp,os.O_RDONLY); os.posix_fadvise(fd,0,0,os.POSIX_FADV_DONTNEED); os.close(fd)"
# during the load, on BOTH nodes
python3 scripts/cache-warden.py --model-dir ~/.cache/huggingface/hub \
--interval 20 --stop-below-gb 25 --max-runtime 86400 --log ~/warden.jsonl
6. Verify it took β at the point of effect, not in your config file:
docker exec <container> env | grep -E 'SGLANG_HOST_IP|NCCL_IB_HCA|NCHANNELS'
curl -s localhost:8899/get_server_info | python3 -m json.tool | grep -E 'max_total|speculative|context'
A setting you did not confirm arrived is a setting you did not set. Ours silently disagreed with
the .env more than once.
Stack (pin these when reproducing)
| Model | RadixArk/Qwen3.8-Flash-Next-NVFP4 (206 shards, 135.2 GB) |
| Engine | SGLang, image lmsysorg/sglang:qwen38flashnext + MiaAI-Lab's SM121 QSA patch |
| Driver | NVIDIA 580.173.02 (open kernel module, aarch64) |
| NCCL | 2.30.7, host-staged and LD_PRELOADed (both recipes treat this as required for GB10 multi-node) |
| Hardware | 2Γ DGX Spark, GB10 / sm_121, 121 GB unified per node |
Config delta vs. the MiaAI-Lab defaults
Everything else is hers. These are the only changes, and why:
| setting | value | why |
|---|---|---|
SGLANG_HOST_IP |
fabric IP, per node | the deadlock in Β§1 |
NCCL_MIN/MAX_NCHANNELS |
4 | default negotiated 64 channels and hung during init; tonyd2wild pins 4 |
CUDA_GRAPH_BS |
dense 1..8 |
a ladder with gaps forces padding, and padded rows carry decode_len=0, which wedges the sparse indexer under spec decode. Dense inside range means nothing pads. tonyd2wild solves the same problem with --disable-cuda-graph-padding; we do both |
MAX_RUNNING_REQUESTS |
8 | matches the graph ladder β 9β16 ran eager with transient workspace |
--max-total-tokens |
600000 | pinned; unpinned "OOMs under sustained load" (tonyd2wild) |
MEM_FRACTION_STATIC |
0.80 | leaves ~23 GB headroom |
NCCL_CROSS_NIC |
0 | a multi-Spark report had CROSS_NIC=1 wedge within hours under real traffic |
The five things this repo adds
1. SGLANG_HOST_IP β multi-node SGLang deadlocks behind a default-deny firewall
Symptom: both ranks hang forever after CustomAllreduce is disabled. No error, no timeout, no
NCCL warning. NCCL itself completes fine β RoCE connects, channels build.
Diagnosis (py-spy dump on both ranks β this is what made it solvable):
rank 0 (head): wait_until_ready (shm_broadcast.py:333) <- writer awaiting subscriptions
rank 1 (worker): wait_until_ready (shm_broadcast.py:344) <- reader awaiting READY
A ZMQ PUB/SUB deadlock, not an NCCL problem. Root cause: SGLang's get_local_ip_auto() returns the
default-route interface β the management LAN address β and binds the cross-rank XPUB control
socket there. With -P INPUT DROP on that interface, the worker's subscription never arrives.
Fix β pin the control plane to the fabric, per node:
docker run ... -e SGLANG_HOST_IP=10.10.10.1 # head
docker run ... -e SGLANG_HOST_IP=10.10.10.2 # worker
The bug requires two conditions together: the fabric is not the default route, and the default-route
interface drops unsolicited inbound. Absent either, it never appears β which is presumably why the
published recipes don't mention it. vLLM's equivalent is VLLM_HOST_IP, which is what confirmed the fix was legitimate
rather than a workaround.
2. Page cache is half your memory budget β and it is the wedge
On GB10 the GPU and host share one physical pool. safetensors mmaps each shard, so file pages
and resident tensors compete for the same memory. Measured on an idle node:
reading 41.6 GB of shards -> MemFree 103.0 -> 64.1 GB (~1:1)
A full load reads far more than that. When it runs out, the NVIDIA driver fails before the kernel reclaims β from our own kernel log on a hung boot:
18:48:40 NVRM: Out of memory [NV_ERR_NO_MEMORY] .. _memdescAllocInternal @ mem_desc.c:1359
18:48:54 systemd-journald: Under memory pressure, flushing caches.
NVRM failed 14 seconds before the kernel registered any pressure of its own, and that boot
contains no page allocation failure, no order: line, and no OOM-killer invocation. The cache was
resident; the driver could not have it. On unified memory this does not raise β it wedges the
whole box, and recovery is a physical power-cycle. Note: when it wedges this hard the power
button is dead too β no lights, no fans, no response to a long hold. Unplug and replug is the only
recovery we found.
Consequences, all measured:
Gate on
MemFree, neverMemAvailable. MemAvailable counts reclaimable page cache the driver cannot use. Live example from a node in this state:MemFree 1.9 GBvsMemAvailable 17.2 GB.drop_cachesbefore launch is necessary but not sufficient β the cache regrows during the 135 GB read. Seescripts/cache-warden.py, which bounds it during and after load, needs no root, and requires no engine patch.The same pressure silently costs throughput, not just stability:
MemFree median tok/s CV min 1.7 GB 58.77 9.7% 44.96 21 GB 60.69 2.0% 58.74 One root cause, two symptoms. Every benchmark in this repo evicts cache first.
The warden's own A/B, identical 41.6 GB read:
MemFree before β after control (no warden) 103.0 β 63.9 GB (β39.1) with warden 102.8 β 102.6 GB (β0.2) Safe against a running engine.
posix_fadvise(DONTNEED)drops only clean, unmapped pages: pages the live engine has mmap'd are skipped by the kernel, dirty pages are never discarded, and the shards are read-only anyway. Worst case is a re-read from NVMe β latency, never corruption. It carries a hard self-limit, exits when its target process dies, and reports a loudFATALon a bad directory andUNVERIFIEDwhen it cannot evict. All three exit paths were proven by execution, not by reading the code.--max-total-tokensis a ceiling, not a floor. Requesting 600,000 with a warm cache silently yielded 557,120 β the engine under-fills the pool and does not warn.
3. Speculative decoding is at its architectural ceiling β and the drafter saturates it
3/1/4 is not a conservative default. Raising it is refused by the engine:
NotImplementedError: Qwen QSA requires speculative_num_draft_tokens <= the QSA compress ratio (4):
the pending index-key ring holds one group; got 5
Qwen Sparse Attention's pending index-key ring holds exactly one group of 4. num_draft_tokens
can never exceed 4 regardless of acceptance rate.
Per-run acceptance measurements β the first we're aware of for this model β n=16, hard code prompt, per-run:
spec_accept_length 3.300 β 3.850 median 3.500 (HARD CEILING 4.0)
tok/s 52.55 β 62.56
correlation r = +0.786
Two findings:
The throughput variance people see is acceptance, not noise. It is not thermal, not clocks, not page cache β all were flat/controlled. It is inherent and cannot be removed by environmental control.
The drafter reaches 3.85 against a hard maximum of 4.0. Net of the bonus token that is ~95% of draft slots accepted on the best runs. The model would benefit from a larger draft budget and QSA makes that impossible.
Cumulative figure β and the trap we fell into publishing it. Over the eval campaign: accept length 2.99, accept rate 0.66, derived from lifetime counters only:
generation_tokens_total 601,124 + 49,010 = 650,134 (BOTH streaming series, summed by name)
spec_verify_calls_total 217,242
650,134 / 217,242 = 2.99
β οΈ The windowed-gauge trap β we published this mistake before catching it. SGLang's
/metrics gauges sglang:spec_accept_length and spec_accept_rate are recomputed and reset
every decode-log interval. They describe the last few dozen forwards, not the lifetime.
spec_verify_calls_total is lifetime. Pairing them silently labels a window as a campaign.
Watch a single gauge over minutes on an idle-ish engine:
spec_accept_length: 1.45 β 3.425 β 2.05 β 3.325 (window)
spec_verify_calls_total: 211,743 β 217,091 β 217,242 (lifetime, monotonic)
A cumulative average over more calls cannot fall from 3.425 to 2.05. We published "3.425 over 211,743 verify calls" β a gauge read pinned to a counter β and it flattered us by ~15%. The per-run figures above survive, because a per-run gauge read approximates that run's own window. Any dashboard reading of these gauges is a window too.
External anchor. LMSYS reports this model on B200 TP4 NVFP4 at accept length 3.3 (their workload is unstated). Our numbers bracket it β 2.99 cumulative on a mixed workload, 3.50 median on a hard code prompt. Given the Β±40 pp prompt sensitivity below, "the same range" is the most anyone can honestly claim from a cross-workload acceptance comparison.
LMSYS also names the mechanism behind the ceiling: IndexShare MTP reuses QSA selections across draft steps, which is precisely why the pending index-key ring holds a single group.
β οΈ Acceptance swings ~40 pp on prompt alone. Measured on the same engine within one hour:
accept length 3.5β3.7 on one code prompt vs 2.475 on a chat+code mix β i.e. accept rate
~0.84β0.90 vs ~0.49, a ~40 percentage-point swing (length and rate are different units;
rate = (length β 1) / draft_steps). All arithmetically self-consistent β they
measure different prompt mixes. Never quote an acceptance number without naming the prompt set.
Is there a way past the ceiling? Not today. As of 2026-08-27 no DFlash / DSpark / EAGLE3
drafter exists for Flash-Next β z-lab's DFlash repo lists Muse-Glimmer-30B and Qwen3.8-27B
(a different model) and does not mention Flash-Next in supported models, roadmap or TODO. Two
things look like hits and are not: a HuggingFace repo named β¦-MTP-Drafter-GGUF is a repackaging
of the built-in MTP ("extracted β¦ unmodified", 33 tensors), and SGLang's cookbook lists
--speculative-algorithm DFLASH because that's the engine-wide picker on every page β it needs a
--speculative-draft-model-path checkpoint that doesn't exist for this target. The flag being
selectable is a label; the weights are the evidence.
4. Vision works, and the self-review loop closes
Verified with a generated image of known content, not taken from the model card:
224Γ224 PNG, quadrants TL red / TR blue / BL green / BR yellow
answer: all four correct, image_tokens=64, 1.3 s
More useful β the full loop:
model writes HTML β headless Chrome renders at 1280px and 380px
β model reads its own screenshots β critiques its own output
On a run truncated by too small a max_tokens, it reported "the rendering is a complete failureβ¦
just a dark background with a subtle grid pattern" β describing what was on screen, not what it had
intended to write. On a complete run it found a font-size inconsistency and a checkmark-colour
mismatch that required zooming in to confirm. It contradicts its own prior output, which is the
property that makes self-review worth anything.
β οΈ max_tokens β₯ 8000 for a full page. At 2,600 the file truncated mid-CSS and produced a
valid-looking file that rendered blank. No error. Only the screenshot caught it.
5. Thinking mode: binary, helps reasoning, and fails catastrophically 30% of the time
There are no effort levels. The engine reports
ReasoningToggleConfig(toggle_param='enable_thinking', default_enabled=True, effort_kwarg=None).
A/B on 8 reasoning problems with verifiable answers:
| thinking OFF | thinking ON | |
|---|---|---|
| score | 6/8 | 8/8 |
| time | 2.9 s | 14.8 s (5Γ) |
| tokens | 78 | 733, of which 651 thinking (9Γ) |
It fixes exactly the intuition traps: bat-and-ball $0.10 β $0.05, "Sally's sisters" 3 β 2.
β οΈ But do not default it on for code generation. Same task, same config, temperature 0,
max_tokens=14000, n=10 each:
| runaways (empty answer, budget exhausted) | completion tokens | |
|---|---|---|
| thinking ON | 3 / 10 | 1,342 β 14,000 |
| thinking OFF | 0 / 10 | 222 β 287 |
30% of thinking-on requests consumed the entire 14,000-token budget and returned zero characters
of content, with everything in reasoning_content and finish_reason: length. Thinking off
solved the identical task in 222β287 tokens every single time β roughly 50Γ cheaper and
completely stable.
Note the token range under thinking: 1,342 to 14,000, a 10Γ spread at temperature 0. Greedy decoding is not bit-reproducible on this stack (NVFP4 GEMM variance on sm_121 is the usual explanation), and thinking amplifies that divergence into a coin-flip between "fine" and "produces nothing at all".
Practical guidance:
- Reasoning problems, short outputs β thinking ON is a real win (6/8 β 8/8 on classic
intuition traps: bat-and-ball
$0.10 β $0.05, "Sally's sisters"3 β 2). - Code generation, long outputs β thinking OFF. It is faster, ~50Γ cheaper in tokens, and does not silently return nothing.
- If you must run thinking on unattended, you need a guard: treat
finish_reason == "length"or emptycontentas a retryable failure, not as a model answer. A harness without that guard will book 30% of its thinking-arm results as task failures and conclude "thinking hurts on code" β which is not what is happening.
β οΈ Wherever thinking is on, max_tokens must be β₯ 2000 regardless. Thinking consumes the
same budget as the answer, so a small cap guarantees the empty-content outcome rather than
merely risking it.
Tool calling was not harmed by thinking in our testing (correct tool_calls at temp 0.0, 0.7 and
1.0). One recipe reports a token-0 !!!!! repetition loop for thinking+tools; we probed n=6 at
temp 1.0 and saw none, on the riskier configuration (flashinfer sampling, radix cache on).
n=6 cannot prove absence of a rare probabilistic loop. Keep it on the watch list.
That watch-list item now has a confirmed sibling, below β and note why the n=6 probe found nothing: it ran at temperature 1.0, which is precisely the setting that does not loop.
Sampling: a 1-in-5 repetition loop at temperature 0 β cause NOT established
The engine ships sampling_defaults='model', so a request that sends no sampling parameters
gets the checkpoint's own generation_config:
temperature 1.0 top_k 20 top_p 0.95
Passing temperature: 0 overrides that. On this stack, long greedy builds loop
INTERMITTENTLY β measured at 1 of 5 runs.
What was measured (2026-08-27). Identical prompt β rebuild a home page from a structured
brief β thinking off, max_tokens 14000, one run per arm:
| sampling | tokens | finish | outcome | compliance audit |
|---|---|---|---|---|
temperature: 0 |
14,000 | length |
one CSS line emitted 507 times, never escaped | 12/28 |
| (none sent β checkpoint default) | 9,550 | stop |
clean | 27/29 |
temp 0.7 / top_p 0.8 / top_k 20 |
11,831 | stop |
clean | 28/29 |
The greedy run never reached the end of the document, so the page had no <main>, no footer and
no links β which is why the compliance score collapses. At 800 tokens neither config repeats
a line, so whatever this is, it is length-dependent.
It is rare, and our first write-up of it was wrong. The identical greedy build was re-run four more times on the same prompt:
run 1 10,785 tok finish=stop max repeated content line 3 clean
run 2 11,006 tok finish=stop max repeated content line 3 clean
run 3 12,252 tok finish=stop max repeated content line 4 clean
run 4 14,000 tok finish=length max repeated content line 1 clean (long, not looping)
0 of 4. Pooled with the original, the observed rate is 1 in 5 β not something greedy does,
something greedy sometimes does. Greedy decoding is not bit-reproducible on this stack (NVFP4
GEMM variance on sm_121 β the same effect behind the 10Γ token spread documented in the thinking
section above), so temperature: 0 names a distribution, not one trajectory. A small slice of
that distribution lands in a basin greedy cannot leave. A sampled decoder can land in the same
basin and still escape by chance β that asymmetry, not the loop itself, is the finding.
The two sampled arms are one run each. n=1 bounds nothing; treat their rate as unmeasured, merely lower.
Why we are NOT claiming "temperature 0 causes this"
β οΈ There is an uncontrolled confound, and it is a big one. We run
--sampling-backend flashinfer. tonyd2wild's recipe
for the same model and hardware documents a degenerate-output loop and attributes it to that exact
kernel β his fix is --sampling-backend pytorch, described as ruling out "the FlashInfer kernel
arg-maxing a stale row to token 0." With his four-part stack he reports the loop clean at temp
0.0 / 0.2 / 0.7, with a residual edge only at temp 1.0.
We ship two of his four loop-fix elements (enable_thinking: false,
--disable-cuda-graph-padding) and not the other two (--sampling-backend pytorch,
--disable-radix-cache). So the honest statement is:
A long greedy generation looped on a stack missing the sampling-backend fix that a published recipe says prevents exactly this class of failure. Temperature is correlated with the failure in our three runs; it is not established as the cause.
Our manifestation also differs from his β a whole CSS line repeated 507 times, not a token-0 !
loop β so they may be different bugs. Unresolved. Testing it properly means restarting the
engine with --sampling-backend pytorch and re-running all three arms, which we have not done.
(Do not reach for --disable-radix-cache casually as the other half of his stack: his own
2026-08-27 update reports it silently collapses the mamba/SSM state pool to
max_running_requests. Our workload is also prefill-dominated, which is precisely where a prefix
cache pays.)
And no, we cannot tell you greedy is faster
An earlier version of this section claimed temperature 0 was +6.9% faster (48.2 vs 45.1 tok/s), citing higher speculative acceptance (58.3% vs 50.8%) as the mechanism. That claim is withdrawn. It came from n=3 per arm, against a measured inherent CV of ~6.3% on this cluster β the "difference" was the same size as the noise, and the defaults arm contained a 42.3 outlier of exactly the shape this log has previously root-caused to page-cache pressure. This repo's own standard, set after an earlier bad call, is that n=5 is not enough to report a config win. n=3 is not close.
The acceptance figures (2.75 vs 2.525 accept-length) are real per-run gauge reads and the direction is mechanically plausible β greedy tokens are more predictable, so the drafter hits more often. Plausible is not measured. If you want this number, it needs nβ₯16 with page cache evicted and a named prompt.
Practical guidance, as far as it is actually supported:
- Long generation on a flashinfer-sampling stack β send no sampling parameters, or cap temperature at β€0.7 per tonyd2wild. Both completed cleanly here β one run each, so this is a completion, not a rate. The reason to prefer them is the escape asymmetry above, not a measured difference in loop frequency.
- A 1-in-5 chance of losing the whole document is worth engineering around even though it is
rare. If you run greedy on long output, treat
finish_reason == "length"plus a high repeated-line count as a retryable failure. - On the server-side default there is a real trade, and we have not resolved it. We leave
sampling_defaults='model', which serves temp 1.0 to any client that sends nothing β and temp 1.0 is precisely where tonyd2wild reports his residual edge, with an explicit recommendation to cap agent temperature at β€0.7. Keepingmodelpreserves the diagnostic signal and honours the checkpoint's own config; it also defaults silent clients into the one regime the cited source calls risky. Pick deliberately rather than inheriting it, as we did. - Benchmarks β always say which sampling config produced the number. The direction is plausible; the size, and whether it exists at all, is unmeasured.
Capability evaluation
13 tasks across backend Python, backend Node/TS, SQL schema design, debugging, three frontend stacks (vanilla, React, Next.js App Router), Sanity CMS schemas, and multi-file cross-file debugging. Two passes per arm, temp 0.
8 of the 13 are graded by executing held-out tests in a sandbox (backend Python Γ2, Node Γ2, SQL, debugging Γ2, and the cross-file task). The three frontend tasks, the Sanity schema and one large-codebase task are graded by structural checks on the output text β weaker, and the negative controls validate only the executing verifiers.
| arm | scored PASS | INVALID | wall clockΒΉ |
|---|---|---|---|
| thinking OFF, pass 1 | 13 / 13 | 0 | 3 m 44 s |
| thinking OFF, pass 2 | 13 / 13 | 0 | 3 m 45 s |
| thinking ON, pass 1 | 12 / 12 | 1 | 81 min |
| thinking ON, pass 2 | 11 / 11 | 2 | 84 min |
ΒΉ Wall clock between arm-start markers in the run log β this is what you actually wait for.
It is much larger than the sum of per-task elapsed, because retried attempts are not counted in
the per-record figure and one frontend task alone burned ~3 Γ 400 s per thinking-ON arm.
Thinking off is ~22Γ faster in wall clock, with equal correctness. Both INVALIDs in pass 2 were thinking-budget exhaustion at 20,000 tokens on long-output tasks.
A fourth independent measurement of the runaway, from the campaign itself: retries fired in 10 of 30 thinking-ON cells and 0 of 32 thinking-OFF cells. The thinking-ON scoreline is therefore retry-dependent β retries only fire on INVALID (truncation or empty output), never on FAIL, so they cannot turn a wrong answer into a pass, but they do resample a nondeterministic coin-flip. Without the retry policy the thinking arm would show ~30% failures that are not capability failures.
β οΈ A clean sweep measures the suite, not the model. 13/13 bounds the failure rate; it does not locate the ceiling. The 95% Wilson interval on 13/13 is roughly 77β100% β wide, because n is small. The honest reading is "this suite sits below the model's capability", not "this model does not fail". These tasks were written by us and are not a public benchmark.
Disclosure: first-pass results under two buggy verifiers were 12/13. fe-01 (both OFF passes)
and dbg-02 (ON pass 1) were re-run after the verifier fixes described below β fresh
generations, not re-grades. The headline includes those re-run cells.
The hardest task β a four-file service with a cross-file contract bug (a heap negating priority while the constants documented the opposite convention) β passed in 18 s with thinking on (under 4 s with it off), changing only the file that needed changing, fixing the misleading comment that caused it, and satisfying a held-out three-part test covering ordering, FIFO tie-break, and untouched retry semantics.
Verifier validation
Every run includes negative controls whose tests are deliberately unsatisfiable: a Python task
asserting 2+2==5, and a SQL task whose table is pre-created so the model's own DDL must collide.
Both failed correctly in all four arms, and real tasks pass β so the verifiers genuinely execute
and are not merely always-fail.
β οΈ Three of our own checks failed correct output
This is the part worth copying if you build something similar. In the first pass, three verifiers produced confident, specific, false results:
| check | what it did | reality |
|---|---|---|
| mutable-default fix | asserted the caller's list must not be mutated | the prompt never asked for a defensive copy; taking ownership is a normal contract |
| self-contained HTML | banned the substring http:// |
flagged xmlns="http://www.w3.org/2000/svg" β a namespace URI browsers never fetch |
| token budget | 6,000 max_tokens | thinking consumed it, so truncation looked like failure |
Uncorrected, the writeup would have claimed this model fails the classic mutable-default bug and cannot produce self-contained HTML. Both are the opposite of true. A verifier is a claim about the world and needs its own negative controls β ours caught the model's failures fine; what they could not catch was themselves. The tell each time was a surprising failure that turned out, on reading the actual output, to be correct.
Dead ends β documented so you don't spend the time
Clock headroom does not exist. clocks.max.sm reports 3003 MHz; the GPU runs 2528 under load.
Locking -lgc 2800,3003 yields 2528 MHz β the floor does not take β and prefill changes by
0.07%:
| clock | prefill | |
|---|---|---|
| default | 2528 MHz | 3,055 tok/s |
| locked 2800β3003 | 2528 MHz | 3,053 tok/s |
GB10 is memory-bandwidth bound, not clock bound, for both prefill and decode. This also explains the low power draw β the GPU is mostly waiting on memory.
Raising the draft budget is impossible. See Β§3.
--load-format dummy should not be used on GB10 β the rule and the >150 GB transient figure are
Mia's; our contribution is only that it explains one of our own wedges. A "safe rehearsal" is more
dangerous than the real load.
Thermals β much cooler than a comparable dense-ish MoE
Measured at concurrency 4, 1 Hz telemetry:
idle 42.0 Β°C Β· 10.4 W Β· 2398 MHz
peak (load) 52.0 Β°C Β· 35.5 W Β· 2522 MHz
For scale, a previous model on identical hardware peaked at 88 Β°C / 65 W uncapped. Qwen runs ~36 Β°C cooler at ~45% the power β at higher clocks. Cause: ~6B active params/token (10 of 512 experts) and only 12 of 48 layers are full attention; the rest are linear-attention GDN.
Practical effect: thermal guard stages sized for the older model are effectively unreachable (43 Β°C of margin), and a clock cap intended to control thermals has nothing left to control.
Benchmark discipline
Two spectacular false results were produced and caught during this work. Both were prefix-cache artifacts:
"72,000 tok/s prefill" -> identical prompt repeated, radix cache hit
"46,388 tok/s prefill" -> a shell function that never passed its seed argument
true prefill: 3,050 tok/s (unique prompt per run, cached_tokens=0 asserted)
Always assert usage.prompt_tokens_details.cached_tokens == 0 when measuring prefill. A 24Γ
speedup that appears without a config change is a cache hit, not a discovery.
Likewise, ignore_eos benchmarks are a floor, not real-world throughput. They force generation
past the natural stopping point into degenerate text:
| measurement | tok/s |
|---|---|
| real generation (10.7k tokens of HTML) | 62.9 |
harness, ignore_eos, hard prompt, 800 tok |
47.6 |
Both correct; they measure different things. Name the prompt, token count and clock state on every number, or it is not comparable to anything.
Files
scripts/cache-warden.py |
bounds page cache during and after load; no root, no engine patch |
bench/decode-bench.py |
decode benchmark that waits for idle, discards contended runs, reports medians, and names its conditions |
Credit
This work stands on two recipes published first, and would not exist without them:
- MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks β the orchestration and the SM121 QSA Triton fallback kernel that makes this model run on sm_121 at all. Our deployment is this stack.
- tonyd2wild/qwen3.8-flash-next-nvfp4-dgx-spark β
--disable-cuda-graph-padding, NCCL channel pinning (which fixed a 64-channel init hang for us), KV pinning, and the rule that any fix making the model text-only is off the table. - bird/GLM-spark β published the
posix_fadvise(DONTNEED)page-cache mechanism first, as an in-loader vLLM patch. We arrived at it independently and measured it before finding theirs;cache-warden.pyis an out-of-tree, no-root variant that also bounds cache during and after load.
License
MIT.