Qwen3.8-Flash-Next NVFP4 on 2Γ— DGX Spark β€” field notes

Mirror of https://github.com/beastllama/dgx-spark-qwen38-flash-next-recipe β€” the base recipe is MiaAI-Lab's; this is the config delta plus everything that went wrong on the way and the measurements nobody had published. From the same homelab as the GLM-5.3-Flash + DFlash2 recipe.

Serving Qwen3.8-Flash-Next-NVFP4 across two DGX Sparks (GB10 / sm_121) with SGLang, TP=2 over ConnectX-7 RoCE (single rail β€” dual-rail is measured fabric capability, untested under SGLang). Vision enabled. 262,144 context. MTP speculative decoding.

Start here, then read the findings. The base stack is MiaAI-Lab's β€” this repo is the config delta on top of it plus everything that went wrong getting there and how it was fixed (see Credit). What follows was not in either published recipe: repeated node wedges (on unified memory, exhaustion does not error β€” it takes the whole box), a deadlock that only appears behind a default-deny firewall, the first speculative-decode acceptance measurements we're aware of for this model, and three optimisation avenues that turned out to be dead ends β€” documented as such, with numbers.

Everything here was measured on real hardware. Where a number is contested or unproven, it says so.


Status

Serving βœ… TP=2 across 2 nodes, 262,144 context
Vision βœ… verified end-to-end (see below)
Spec decode βœ… NEXTN 3/1/4 (num_steps/eagle_topk/num_draft_tokens) β€” note the engine self-reports speculative_algorithm: EAGLE β€” the architectural maximum, not a default
Decode ~63 tok/s single-stream on real generation
Concurrency 306 tok/s aggregate at 6 streams
Prefill 3,050 tok/s (cache defeated)
Thermals 52 Β°C / 35 W peak under load

Conditions, because this repo insists on them: decode ~63 tok/s = real generation of a 10.7k-token HTML file, thinking off, temp 0.3, 2398 MHz idle / 2522 under load. 306 tok/s = aggregate across 6 concurrent streams, code prompt, 400 max_tokens, ignore_eos. Prefill 3,050 tok/s = ~7,450-token unique prompt per run, cached_tokens=0 asserted, n=6. Thermals = concurrency 4, 1 Hz sampling. A number without its prompt, token count and clock state is not comparable to anything β€” including these.


Performance

Two DGX Sparks. 180B params (125B backbone + 51B PLE), NVFP4, 262k context, vision on.

tok/s conditions
Single stream, real work ~63 10.7k-token HTML page, natural stop, 2522 MHz
Single stream, 400 tok 63.7 code prompt, ignore_eos
2 concurrent 104.9 agg 52.8/streamΒ²
4 concurrent 178.7 agg 45.2/streamΒ²
6 concurrent 306.6 agg 51.7/streamΒ²
Prefill 3,050 ~7,450-token unique prompt, cached_tokens=0 asserted, n=6
Stress floor 47.6 ignore_eos + hard prompt + 800 tok β€” a deliberate FLOOR, see below

Β² Concurrency measured before the config was pinned β€” at max_running_requests=12 and an unpinned KV pool of 850,816 tokens, not the 8 / 600,000 in the recipe below. Under the pinned config 8 is the cap, so 6 streams is near it rather than "still climbing". Re-measure before quoting these against the shipped config.

Power and heat, at concurrency 4: 52 Β°C, 35.5 W peak per node; 42 Β°C / 10.4 W idle. That is roughly half the draw of a comparably-sized dense-ish MoE we previously ran on the same boxes (88 Β°C / 65 W), at higher clocks. Cause: 6B active params per token (10 of 512 experts) and only 12 of 48 layers are full attention β€” the rest are linear-attention GDN, so decode waits on memory rather than burning watts. Practical effect: thermal guard stages sized for the older model are unreachable by 43 Β°C, and two Sparks serve this at **71 W combined under load**.

Why two numbers for "single stream". ignore_eos benchmarks force generation past the model's natural stopping point into degenerate text. They're excellent for regression detection and useless as a headline. Real generation of a complete HTML page runs at ~63 tok/s; the same stack under ignore_eos on a hard prompt reports 47.6. Both are correct. Quote the one that matches what you're doing, and say which.

Speculative decoding is doing much of the work, but the delta is not isolated. Our earlier vLLM deployment of the same checkpoint (no MTP, --enforce-eager) decoded at 20–21 tok/s; this SGLang stack with MTP runs ~3Γ— that. Engine, CUDA-graph mode and MTP all changed together β€” we have no SGLang-with-MTP-off measurement, and neither published recipe ships one to compare against. See Β§3 for why 3/1/4 is the ceiling.


Recipe

1. Base stack. Clone MiaAI-Lab's repo and follow it β€” it builds the SM121 QSA patch onto the public SGLang image and handles fabric preflight, worker-first ordering and readiness waiting. Everything below is a delta on that.

2. Stage NCCL on both nodes. Both published recipes treat host-staged NCCL as required for GB10 multi-node stability:

mkdir -p ~/nccl-2.30.7
cp /usr/lib/aarch64-linux-gnu/libnccl.so.2.30.7 ~/nccl-2.30.7/
ln -sf libnccl.so.2.30.7 ~/nccl-2.30.7/libnccl.so.2

3. Apply the config delta (table below) to the .env.

4. Pin the control plane to the fabric β€” the single most important line if you run a default-deny firewall, and the one that cost us the longest debug:

-e SGLANG_HOST_IP=<this node's fabric IP>

5. Evict page cache immediately before launch, and keep it bounded during the load:

# before launch (no root needed β€” this is what scripts/cache-warden.py automates)
python3 -c "
import os,glob
for p in glob.glob(os.path.expanduser('~/.cache/huggingface/hub/**/*.safetensors'),recursive=True):
    rp=os.path.realpath(p)
    if os.path.exists(rp):
        fd=os.open(rp,os.O_RDONLY); os.posix_fadvise(fd,0,0,os.POSIX_FADV_DONTNEED); os.close(fd)"

# during the load, on BOTH nodes
python3 scripts/cache-warden.py --model-dir ~/.cache/huggingface/hub \
    --interval 20 --stop-below-gb 25 --max-runtime 86400 --log ~/warden.jsonl

6. Verify it took β€” at the point of effect, not in your config file:

docker exec <container> env | grep -E 'SGLANG_HOST_IP|NCCL_IB_HCA|NCHANNELS'
curl -s localhost:8899/get_server_info | python3 -m json.tool | grep -E 'max_total|speculative|context'

A setting you did not confirm arrived is a setting you did not set. Ours silently disagreed with the .env more than once.


Stack (pin these when reproducing)

Model RadixArk/Qwen3.8-Flash-Next-NVFP4 (206 shards, 135.2 GB)
Engine SGLang, image lmsysorg/sglang:qwen38flashnext + MiaAI-Lab's SM121 QSA patch
Driver NVIDIA 580.173.02 (open kernel module, aarch64)
NCCL 2.30.7, host-staged and LD_PRELOADed (both recipes treat this as required for GB10 multi-node)
Hardware 2Γ— DGX Spark, GB10 / sm_121, 121 GB unified per node

Config delta vs. the MiaAI-Lab defaults

Everything else is hers. These are the only changes, and why:

setting value why
SGLANG_HOST_IP fabric IP, per node the deadlock in Β§1
NCCL_MIN/MAX_NCHANNELS 4 default negotiated 64 channels and hung during init; tonyd2wild pins 4
CUDA_GRAPH_BS dense 1..8 a ladder with gaps forces padding, and padded rows carry decode_len=0, which wedges the sparse indexer under spec decode. Dense inside range means nothing pads. tonyd2wild solves the same problem with --disable-cuda-graph-padding; we do both
MAX_RUNNING_REQUESTS 8 matches the graph ladder β€” 9–16 ran eager with transient workspace
--max-total-tokens 600000 pinned; unpinned "OOMs under sustained load" (tonyd2wild)
MEM_FRACTION_STATIC 0.80 leaves ~23 GB headroom
NCCL_CROSS_NIC 0 a multi-Spark report had CROSS_NIC=1 wedge within hours under real traffic

The five things this repo adds

1. SGLANG_HOST_IP β€” multi-node SGLang deadlocks behind a default-deny firewall

Symptom: both ranks hang forever after CustomAllreduce is disabled. No error, no timeout, no NCCL warning. NCCL itself completes fine β€” RoCE connects, channels build.

Diagnosis (py-spy dump on both ranks β€” this is what made it solvable):

rank 0 (head):    wait_until_ready (shm_broadcast.py:333)   <- writer awaiting subscriptions
rank 1 (worker):  wait_until_ready (shm_broadcast.py:344)   <- reader awaiting READY

A ZMQ PUB/SUB deadlock, not an NCCL problem. Root cause: SGLang's get_local_ip_auto() returns the default-route interface β€” the management LAN address β€” and binds the cross-rank XPUB control socket there. With -P INPUT DROP on that interface, the worker's subscription never arrives.

Fix β€” pin the control plane to the fabric, per node:

docker run ... -e SGLANG_HOST_IP=10.10.10.1   # head
docker run ... -e SGLANG_HOST_IP=10.10.10.2   # worker

The bug requires two conditions together: the fabric is not the default route, and the default-route interface drops unsolicited inbound. Absent either, it never appears β€” which is presumably why the published recipes don't mention it. vLLM's equivalent is VLLM_HOST_IP, which is what confirmed the fix was legitimate rather than a workaround.

2. Page cache is half your memory budget β€” and it is the wedge

On GB10 the GPU and host share one physical pool. safetensors mmaps each shard, so file pages and resident tensors compete for the same memory. Measured on an idle node:

reading 41.6 GB of shards  ->  MemFree 103.0 -> 64.1 GB    (~1:1)

A full load reads far more than that. When it runs out, the NVIDIA driver fails before the kernel reclaims β€” from our own kernel log on a hung boot:

18:48:40  NVRM: Out of memory [NV_ERR_NO_MEMORY] .. _memdescAllocInternal @ mem_desc.c:1359
18:48:54  systemd-journald: Under memory pressure, flushing caches.

NVRM failed 14 seconds before the kernel registered any pressure of its own, and that boot contains no page allocation failure, no order: line, and no OOM-killer invocation. The cache was resident; the driver could not have it. On unified memory this does not raise β€” it wedges the whole box, and recovery is a physical power-cycle. Note: when it wedges this hard the power button is dead too β€” no lights, no fans, no response to a long hold. Unplug and replug is the only recovery we found.

Consequences, all measured:

  • Gate on MemFree, never MemAvailable. MemAvailable counts reclaimable page cache the driver cannot use. Live example from a node in this state: MemFree 1.9 GB vs MemAvailable 17.2 GB.

  • drop_caches before launch is necessary but not sufficient β€” the cache regrows during the 135 GB read. See scripts/cache-warden.py, which bounds it during and after load, needs no root, and requires no engine patch.

  • The same pressure silently costs throughput, not just stability:

    MemFree median tok/s CV min
    1.7 GB 58.77 9.7% 44.96
    21 GB 60.69 2.0% 58.74

    One root cause, two symptoms. Every benchmark in this repo evicts cache first.

    The warden's own A/B, identical 41.6 GB read:

    MemFree before β†’ after
    control (no warden) 103.0 β†’ 63.9 GB (βˆ’39.1)
    with warden 102.8 β†’ 102.6 GB (βˆ’0.2)

    Safe against a running engine. posix_fadvise(DONTNEED) drops only clean, unmapped pages: pages the live engine has mmap'd are skipped by the kernel, dirty pages are never discarded, and the shards are read-only anyway. Worst case is a re-read from NVMe β€” latency, never corruption. It carries a hard self-limit, exits when its target process dies, and reports a loud FATAL on a bad directory and UNVERIFIED when it cannot evict. All three exit paths were proven by execution, not by reading the code.

  • --max-total-tokens is a ceiling, not a floor. Requesting 600,000 with a warm cache silently yielded 557,120 β€” the engine under-fills the pool and does not warn.

3. Speculative decoding is at its architectural ceiling β€” and the drafter saturates it

3/1/4 is not a conservative default. Raising it is refused by the engine:

NotImplementedError: Qwen QSA requires speculative_num_draft_tokens <= the QSA compress ratio (4):
the pending index-key ring holds one group; got 5

Qwen Sparse Attention's pending index-key ring holds exactly one group of 4. num_draft_tokens can never exceed 4 regardless of acceptance rate.

Per-run acceptance measurements β€” the first we're aware of for this model β€” n=16, hard code prompt, per-run:

spec_accept_length   3.300 – 3.850   median 3.500     (HARD CEILING 4.0)
tok/s                52.55 – 62.56
correlation                          r = +0.786

Two findings:

The throughput variance people see is acceptance, not noise. It is not thermal, not clocks, not page cache β€” all were flat/controlled. It is inherent and cannot be removed by environmental control.

The drafter reaches 3.85 against a hard maximum of 4.0. Net of the bonus token that is ~95% of draft slots accepted on the best runs. The model would benefit from a larger draft budget and QSA makes that impossible.

Cumulative figure β€” and the trap we fell into publishing it. Over the eval campaign: accept length 2.99, accept rate 0.66, derived from lifetime counters only:

generation_tokens_total  601,124 + 49,010 = 650,134   (BOTH streaming series, summed by name)
spec_verify_calls_total                     217,242
                          650,134 / 217,242 = 2.99

⚠️ The windowed-gauge trap β€” we published this mistake before catching it. SGLang's /metrics gauges sglang:spec_accept_length and spec_accept_rate are recomputed and reset every decode-log interval. They describe the last few dozen forwards, not the lifetime. spec_verify_calls_total is lifetime. Pairing them silently labels a window as a campaign.

Watch a single gauge over minutes on an idle-ish engine:

spec_accept_length:  1.45  β†’  3.425  β†’  2.05  β†’  3.325     (window)
spec_verify_calls_total: 211,743  β†’  217,091  β†’  217,242   (lifetime, monotonic)

A cumulative average over more calls cannot fall from 3.425 to 2.05. We published "3.425 over 211,743 verify calls" β€” a gauge read pinned to a counter β€” and it flattered us by ~15%. The per-run figures above survive, because a per-run gauge read approximates that run's own window. Any dashboard reading of these gauges is a window too.

External anchor. LMSYS reports this model on B200 TP4 NVFP4 at accept length 3.3 (their workload is unstated). Our numbers bracket it β€” 2.99 cumulative on a mixed workload, 3.50 median on a hard code prompt. Given the Β±40 pp prompt sensitivity below, "the same range" is the most anyone can honestly claim from a cross-workload acceptance comparison.

LMSYS also names the mechanism behind the ceiling: IndexShare MTP reuses QSA selections across draft steps, which is precisely why the pending index-key ring holds a single group.

⚠️ Acceptance swings ~40 pp on prompt alone. Measured on the same engine within one hour: accept length 3.5–3.7 on one code prompt vs 2.475 on a chat+code mix β€” i.e. accept rate ~0.84–0.90 vs ~0.49, a ~40 percentage-point swing (length and rate are different units; rate = (length βˆ’ 1) / draft_steps). All arithmetically self-consistent β€” they measure different prompt mixes. Never quote an acceptance number without naming the prompt set.

Is there a way past the ceiling? Not today. As of 2026-08-27 no DFlash / DSpark / EAGLE3 drafter exists for Flash-Next β€” z-lab's DFlash repo lists Muse-Glimmer-30B and Qwen3.8-27B (a different model) and does not mention Flash-Next in supported models, roadmap or TODO. Two things look like hits and are not: a HuggingFace repo named …-MTP-Drafter-GGUF is a repackaging of the built-in MTP ("extracted … unmodified", 33 tensors), and SGLang's cookbook lists --speculative-algorithm DFLASH because that's the engine-wide picker on every page β€” it needs a --speculative-draft-model-path checkpoint that doesn't exist for this target. The flag being selectable is a label; the weights are the evidence.

4. Vision works, and the self-review loop closes

Verified with a generated image of known content, not taken from the model card:

224Γ—224 PNG, quadrants TL red / TR blue / BL green / BR yellow
answer: all four correct, image_tokens=64, 1.3 s

More useful β€” the full loop:

model writes HTML  β†’  headless Chrome renders at 1280px and 380px
                   β†’  model reads its own screenshots  β†’  critiques its own output

On a run truncated by too small a max_tokens, it reported "the rendering is a complete failure… just a dark background with a subtle grid pattern" β€” describing what was on screen, not what it had intended to write. On a complete run it found a font-size inconsistency and a checkmark-colour mismatch that required zooming in to confirm. It contradicts its own prior output, which is the property that makes self-review worth anything.

⚠️ max_tokens β‰₯ 8000 for a full page. At 2,600 the file truncated mid-CSS and produced a valid-looking file that rendered blank. No error. Only the screenshot caught it.

5. Thinking mode: binary, helps reasoning, and fails catastrophically 30% of the time

There are no effort levels. The engine reports ReasoningToggleConfig(toggle_param='enable_thinking', default_enabled=True, effort_kwarg=None).

A/B on 8 reasoning problems with verifiable answers:

thinking OFF thinking ON
score 6/8 8/8
time 2.9 s 14.8 s (5Γ—)
tokens 78 733, of which 651 thinking (9Γ—)

It fixes exactly the intuition traps: bat-and-ball $0.10 β†’ $0.05, "Sally's sisters" 3 β†’ 2.

⚠️ But do not default it on for code generation. Same task, same config, temperature 0, max_tokens=14000, n=10 each:

runaways (empty answer, budget exhausted) completion tokens
thinking ON 3 / 10 1,342 – 14,000
thinking OFF 0 / 10 222 – 287

30% of thinking-on requests consumed the entire 14,000-token budget and returned zero characters of content, with everything in reasoning_content and finish_reason: length. Thinking off solved the identical task in 222–287 tokens every single time β€” roughly 50Γ— cheaper and completely stable.

Note the token range under thinking: 1,342 to 14,000, a 10Γ— spread at temperature 0. Greedy decoding is not bit-reproducible on this stack (NVFP4 GEMM variance on sm_121 is the usual explanation), and thinking amplifies that divergence into a coin-flip between "fine" and "produces nothing at all".

Practical guidance:

  • Reasoning problems, short outputs β†’ thinking ON is a real win (6/8 β†’ 8/8 on classic intuition traps: bat-and-ball $0.10 β†’ $0.05, "Sally's sisters" 3 β†’ 2).
  • Code generation, long outputs β†’ thinking OFF. It is faster, ~50Γ— cheaper in tokens, and does not silently return nothing.
  • If you must run thinking on unattended, you need a guard: treat finish_reason == "length" or empty content as a retryable failure, not as a model answer. A harness without that guard will book 30% of its thinking-arm results as task failures and conclude "thinking hurts on code" β€” which is not what is happening.

⚠️ Wherever thinking is on, max_tokens must be β‰₯ 2000 regardless. Thinking consumes the same budget as the answer, so a small cap guarantees the empty-content outcome rather than merely risking it.

Tool calling was not harmed by thinking in our testing (correct tool_calls at temp 0.0, 0.7 and 1.0). One recipe reports a token-0 !!!!! repetition loop for thinking+tools; we probed n=6 at temp 1.0 and saw none, on the riskier configuration (flashinfer sampling, radix cache on). n=6 cannot prove absence of a rare probabilistic loop. Keep it on the watch list.

That watch-list item now has a confirmed sibling, below β€” and note why the n=6 probe found nothing: it ran at temperature 1.0, which is precisely the setting that does not loop.


Sampling: a 1-in-5 repetition loop at temperature 0 β€” cause NOT established

The engine ships sampling_defaults='model', so a request that sends no sampling parameters gets the checkpoint's own generation_config:

temperature 1.0   top_k 20   top_p 0.95

Passing temperature: 0 overrides that. On this stack, long greedy builds loop INTERMITTENTLY β€” measured at 1 of 5 runs.

What was measured (2026-08-27). Identical prompt β€” rebuild a home page from a structured brief β€” thinking off, max_tokens 14000, one run per arm:

sampling tokens finish outcome compliance audit
temperature: 0 14,000 length one CSS line emitted 507 times, never escaped 12/28
(none sent β€” checkpoint default) 9,550 stop clean 27/29
temp 0.7 / top_p 0.8 / top_k 20 11,831 stop clean 28/29

The greedy run never reached the end of the document, so the page had no <main>, no footer and no links β€” which is why the compliance score collapses. At 800 tokens neither config repeats a line, so whatever this is, it is length-dependent.

It is rare, and our first write-up of it was wrong. The identical greedy build was re-run four more times on the same prompt:

run 1  10,785 tok  finish=stop    max repeated content line 3   clean
run 2  11,006 tok  finish=stop    max repeated content line 3   clean
run 3  12,252 tok  finish=stop    max repeated content line 4   clean
run 4  14,000 tok  finish=length  max repeated content line 1   clean (long, not looping)

0 of 4. Pooled with the original, the observed rate is 1 in 5 β€” not something greedy does, something greedy sometimes does. Greedy decoding is not bit-reproducible on this stack (NVFP4 GEMM variance on sm_121 β€” the same effect behind the 10Γ— token spread documented in the thinking section above), so temperature: 0 names a distribution, not one trajectory. A small slice of that distribution lands in a basin greedy cannot leave. A sampled decoder can land in the same basin and still escape by chance β€” that asymmetry, not the loop itself, is the finding.

The two sampled arms are one run each. n=1 bounds nothing; treat their rate as unmeasured, merely lower.

Why we are NOT claiming "temperature 0 causes this"

⚠️ There is an uncontrolled confound, and it is a big one. We run --sampling-backend flashinfer. tonyd2wild's recipe for the same model and hardware documents a degenerate-output loop and attributes it to that exact kernel β€” his fix is --sampling-backend pytorch, described as ruling out "the FlashInfer kernel arg-maxing a stale row to token 0." With his four-part stack he reports the loop clean at temp 0.0 / 0.2 / 0.7, with a residual edge only at temp 1.0.

We ship two of his four loop-fix elements (enable_thinking: false, --disable-cuda-graph-padding) and not the other two (--sampling-backend pytorch, --disable-radix-cache). So the honest statement is:

A long greedy generation looped on a stack missing the sampling-backend fix that a published recipe says prevents exactly this class of failure. Temperature is correlated with the failure in our three runs; it is not established as the cause.

Our manifestation also differs from his β€” a whole CSS line repeated 507 times, not a token-0 ! loop β€” so they may be different bugs. Unresolved. Testing it properly means restarting the engine with --sampling-backend pytorch and re-running all three arms, which we have not done.

(Do not reach for --disable-radix-cache casually as the other half of his stack: his own 2026-08-27 update reports it silently collapses the mamba/SSM state pool to max_running_requests. Our workload is also prefill-dominated, which is precisely where a prefix cache pays.)

And no, we cannot tell you greedy is faster

An earlier version of this section claimed temperature 0 was +6.9% faster (48.2 vs 45.1 tok/s), citing higher speculative acceptance (58.3% vs 50.8%) as the mechanism. That claim is withdrawn. It came from n=3 per arm, against a measured inherent CV of ~6.3% on this cluster β€” the "difference" was the same size as the noise, and the defaults arm contained a 42.3 outlier of exactly the shape this log has previously root-caused to page-cache pressure. This repo's own standard, set after an earlier bad call, is that n=5 is not enough to report a config win. n=3 is not close.

The acceptance figures (2.75 vs 2.525 accept-length) are real per-run gauge reads and the direction is mechanically plausible β€” greedy tokens are more predictable, so the drafter hits more often. Plausible is not measured. If you want this number, it needs nβ‰₯16 with page cache evicted and a named prompt.

Practical guidance, as far as it is actually supported:

  • Long generation on a flashinfer-sampling stack β†’ send no sampling parameters, or cap temperature at ≀0.7 per tonyd2wild. Both completed cleanly here β€” one run each, so this is a completion, not a rate. The reason to prefer them is the escape asymmetry above, not a measured difference in loop frequency.
  • A 1-in-5 chance of losing the whole document is worth engineering around even though it is rare. If you run greedy on long output, treat finish_reason == "length" plus a high repeated-line count as a retryable failure.
  • On the server-side default there is a real trade, and we have not resolved it. We leave sampling_defaults='model', which serves temp 1.0 to any client that sends nothing β€” and temp 1.0 is precisely where tonyd2wild reports his residual edge, with an explicit recommendation to cap agent temperature at ≀0.7. Keeping model preserves the diagnostic signal and honours the checkpoint's own config; it also defaults silent clients into the one regime the cited source calls risky. Pick deliberately rather than inheriting it, as we did.
  • Benchmarks β†’ always say which sampling config produced the number. The direction is plausible; the size, and whether it exists at all, is unmeasured.

Capability evaluation

13 tasks across backend Python, backend Node/TS, SQL schema design, debugging, three frontend stacks (vanilla, React, Next.js App Router), Sanity CMS schemas, and multi-file cross-file debugging. Two passes per arm, temp 0.

8 of the 13 are graded by executing held-out tests in a sandbox (backend Python Γ—2, Node Γ—2, SQL, debugging Γ—2, and the cross-file task). The three frontend tasks, the Sanity schema and one large-codebase task are graded by structural checks on the output text β€” weaker, and the negative controls validate only the executing verifiers.

arm scored PASS INVALID wall clockΒΉ
thinking OFF, pass 1 13 / 13 0 3 m 44 s
thinking OFF, pass 2 13 / 13 0 3 m 45 s
thinking ON, pass 1 12 / 12 1 81 min
thinking ON, pass 2 11 / 11 2 84 min

ΒΉ Wall clock between arm-start markers in the run log β€” this is what you actually wait for. It is much larger than the sum of per-task elapsed, because retried attempts are not counted in the per-record figure and one frontend task alone burned ~3 Γ— 400 s per thinking-ON arm.

Thinking off is ~22Γ— faster in wall clock, with equal correctness. Both INVALIDs in pass 2 were thinking-budget exhaustion at 20,000 tokens on long-output tasks.

A fourth independent measurement of the runaway, from the campaign itself: retries fired in 10 of 30 thinking-ON cells and 0 of 32 thinking-OFF cells. The thinking-ON scoreline is therefore retry-dependent β€” retries only fire on INVALID (truncation or empty output), never on FAIL, so they cannot turn a wrong answer into a pass, but they do resample a nondeterministic coin-flip. Without the retry policy the thinking arm would show ~30% failures that are not capability failures.

⚠️ A clean sweep measures the suite, not the model. 13/13 bounds the failure rate; it does not locate the ceiling. The 95% Wilson interval on 13/13 is roughly 77–100% β€” wide, because n is small. The honest reading is "this suite sits below the model's capability", not "this model does not fail". These tasks were written by us and are not a public benchmark.

Disclosure: first-pass results under two buggy verifiers were 12/13. fe-01 (both OFF passes) and dbg-02 (ON pass 1) were re-run after the verifier fixes described below β€” fresh generations, not re-grades. The headline includes those re-run cells.

The hardest task β€” a four-file service with a cross-file contract bug (a heap negating priority while the constants documented the opposite convention) β€” passed in 18 s with thinking on (under 4 s with it off), changing only the file that needed changing, fixing the misleading comment that caused it, and satisfying a held-out three-part test covering ordering, FIFO tie-break, and untouched retry semantics.

Verifier validation

Every run includes negative controls whose tests are deliberately unsatisfiable: a Python task asserting 2+2==5, and a SQL task whose table is pre-created so the model's own DDL must collide. Both failed correctly in all four arms, and real tasks pass β€” so the verifiers genuinely execute and are not merely always-fail.

⚠️ Three of our own checks failed correct output

This is the part worth copying if you build something similar. In the first pass, three verifiers produced confident, specific, false results:

check what it did reality
mutable-default fix asserted the caller's list must not be mutated the prompt never asked for a defensive copy; taking ownership is a normal contract
self-contained HTML banned the substring http:// flagged xmlns="http://www.w3.org/2000/svg" β€” a namespace URI browsers never fetch
token budget 6,000 max_tokens thinking consumed it, so truncation looked like failure

Uncorrected, the writeup would have claimed this model fails the classic mutable-default bug and cannot produce self-contained HTML. Both are the opposite of true. A verifier is a claim about the world and needs its own negative controls β€” ours caught the model's failures fine; what they could not catch was themselves. The tell each time was a surprising failure that turned out, on reading the actual output, to be correct.


Dead ends β€” documented so you don't spend the time

Clock headroom does not exist. clocks.max.sm reports 3003 MHz; the GPU runs 2528 under load. Locking -lgc 2800,3003 yields 2528 MHz β€” the floor does not take β€” and prefill changes by 0.07%:

clock prefill
default 2528 MHz 3,055 tok/s
locked 2800–3003 2528 MHz 3,053 tok/s

GB10 is memory-bandwidth bound, not clock bound, for both prefill and decode. This also explains the low power draw β€” the GPU is mostly waiting on memory.

Raising the draft budget is impossible. See Β§3.

--load-format dummy should not be used on GB10 β€” the rule and the >150 GB transient figure are Mia's; our contribution is only that it explains one of our own wedges. A "safe rehearsal" is more dangerous than the real load.


Thermals β€” much cooler than a comparable dense-ish MoE

Measured at concurrency 4, 1 Hz telemetry:

idle           42.0 Β°C Β· 10.4 W Β· 2398 MHz
peak (load)    52.0 Β°C Β· 35.5 W Β· 2522 MHz

For scale, a previous model on identical hardware peaked at 88 Β°C / 65 W uncapped. Qwen runs ~36 Β°C cooler at ~45% the power β€” at higher clocks. Cause: ~6B active params/token (10 of 512 experts) and only 12 of 48 layers are full attention; the rest are linear-attention GDN.

Practical effect: thermal guard stages sized for the older model are effectively unreachable (43 Β°C of margin), and a clock cap intended to control thermals has nothing left to control.


Benchmark discipline

Two spectacular false results were produced and caught during this work. Both were prefix-cache artifacts:

"72,000 tok/s prefill"   -> identical prompt repeated, radix cache hit
"46,388 tok/s prefill"   -> a shell function that never passed its seed argument
true prefill:  3,050 tok/s   (unique prompt per run, cached_tokens=0 asserted)

Always assert usage.prompt_tokens_details.cached_tokens == 0 when measuring prefill. A 24Γ— speedup that appears without a config change is a cache hit, not a discovery.

Likewise, ignore_eos benchmarks are a floor, not real-world throughput. They force generation past the natural stopping point into degenerate text:

measurement tok/s
real generation (10.7k tokens of HTML) 62.9
harness, ignore_eos, hard prompt, 800 tok 47.6

Both correct; they measure different things. Name the prompt, token count and clock state on every number, or it is not comparable to anything.


Files

scripts/cache-warden.py bounds page cache during and after load; no root, no engine patch
bench/decode-bench.py decode benchmark that waits for idle, discards contended runs, reports medians, and names its conditions

Credit

This work stands on two recipes published first, and would not exist without them:

  • MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks β€” the orchestration and the SM121 QSA Triton fallback kernel that makes this model run on sm_121 at all. Our deployment is this stack.
  • tonyd2wild/qwen3.8-flash-next-nvfp4-dgx-spark β€” --disable-cuda-graph-padding, NCCL channel pinning (which fixed a 64-channel init hang for us), KV pinning, and the rule that any fix making the model text-only is off the table.
  • bird/GLM-spark β€” published the posix_fadvise(DONTNEED) page-cache mechanism first, as an in-loader vLLM patch. We arrived at it independently and measured it before finding theirs; cache-warden.py is an out-of-tree, no-root variant that also bounds cache during and after load.

License

MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for randomllama/Qwen3.8-Flash-Next-DGX-Spark-field-notes

Finetuned
(2)
this model