Nanbeige4.2-3B-EAGLE3

EAGLE-3 draft model (speculator) for Nanbeige/Nanbeige4.2-3B, trained against my FP8-Dynamic quant as the verifier. This is the thinking-aware revision: the target defaults to reasoning (enable_thinking: true), so I regenerated the training corpus on-policy with thinking on and retrained. Against the prior head, batch-1 decode acceptance in thinking mode rises from 0.33 to 0.41 (+23%) and thinking-mode throughput from 244 to 272 tok/s (+12%) on an RTX 5090, with non-thinking gaining too. Speculative decoding is lossless: outputs are verified by the target model, so quality is exactly the verifier's. I have not found another published draft model for a looped or weight-shared LLM.

The acceptance gap this fixes

The first head reached val acceptance 0.622 but its 21k-conversation corpus was almost entirely non-thinking generations, while production serves reasoning by default. A drafter learns the distribution it trains on, so the prior head accepted fewer draft tokens in thinking mode (0.33) than in plain mode (0.41): the exact traffic it serves most was the traffic it fit worst. EAGLE-3 acceptance scales with training data that matches the serving distribution, so the fix is data, not architecture.

The looped-arch part (why this needed new wiring)

The target executes 22 layers twice per token (num_loops: 2), so the "layer axis" the draft taps is the 44-step unrolled virtual depth, not the 22 physical layers. Three things follow, all carried by this repo:

  1. Aux feature taps are virtual-depth ids: eagle_aux_hidden_state_layer_ids = [3, 23, 41]: early loop-1, early loop-2, late loop-2. Extraction for training adds the last virtual layer (44) for the loss target.
  2. Serve-time layer naming must offset by virtual depth: vLLM's Eagle3 draft names its layer at num_hidden_layers (22); this collides with the target's loop-2 layer 22, and stamps target_layer_count = 22, which mis-indexes the draft's layer_types. The bundled vllm_plugin/ fixes both (offset 44, stamp 44).
  3. Speculation pays double on this arch: the target reads every weight twice per token, and verification amortizes both passes across the whole draft block.

Training

speculators online training (v0.6.0.dev), hidden states extracted on-the-fly from the FP8 verifier served by vLLM v0.25.1.

The corpus is 28,648 conversations, of which 12,886 (45%) carry reasoning traces:

  • 21,299 workload conversations (my serving mix: summarization, RAG answer generation, dictation rewrite, freeform chat, plus a code slice), replayed to keep the prior head's coverage and prevent forgetting.
  • 7,349 on-policy thinking regenerations: GSM8K-style reasoning and multi-turn chat prompts, answered by the FP8 target itself with thinking on (T=0.6, top-p 0.95), so the draft sees the target's own reasoning tokens, not a different model's.

4 epochs, total-seq-len 4096, draft vocab 32k, single RTX 5090 shared between extract-server and trainer (~29 GB total). This is a full retrain on the combined corpus, not a warm-start on the old head. The continual-learning result is that retraining on D_old plus D_new is the upper bound, and the high replay ratio already protects the prior coverage.

Training curves

epoch val loss val cond_acc (pos 0 / 1 / 2)
1 5.88 0.584 / 0.589 / 0.608
2 4.99 0.625 / 0.632 / 0.650
3 4.29 0.659 / 0.667 / 0.684
4 (this repo) 3.85 0.685 / 0.697 / 0.715

The val split now includes thinking traces, so it's harder than the prior head's split. This head clears 0.622 by epoch 2 anyway, and acceptance was still climbing at epoch 4.

Measured serving results (RTX 5090, vLLM v0.25.1, batch-1, T=0.6)

Both heads measured back-to-back under identical conditions. Real reasoning and code prompts, thinking mode is the production default:

mode metric prior head this head
thinking draft acceptance 0.332 0.409 (+23%)
thinking output tok/s 244 272 (+12%)
non-thinking draft acceptance 0.413 0.445
non-thinking output tok/s 277 289 (+5%)

Synthetic GuideLLM sweep (synchronous, 256 output tokens, both heads re-measured under the same unlocked clocks) shows the same direction across context lengths:

input length prior head tok/s this head tok/s
512 198 280 (+41%)
2048 152 171 (+12%)
8192 75 83 (+11%)

Thinking-mode serve acceptance by draft position: 0.61 / 0.37 / 0.25, about 2.2 tokens per verify cycle. Position 2 (the 3rd draft) still pays at 0.25; a 4th rarely does, which is why num_speculative_tokens stays at 3.

One workload still belongs to ngram speculation, and it is a property of the traffic, not the head: long-document summarization with verbatim copy-spans (2-14k-char inputs) favors prompt-lookup drafting over a learned 3-token draft. For bulk long-doc jobs, run the ngram config; eagle3 stays the default for reasoning, agentic, RAG, chat, and code traffic.

Long-document spec comparison

Unlike ngram speculation on this stack, eagle3 keeps async scheduling, Model Runner V2, and full CUDA-graph decode. One caveat that cost me an afternoon: use TRITON_ATTN as the attention backend: FlashInfer + spec-decode downgrades CUDA graphs to piecewise on SM120 and the head then measures slower than no speculation at all.

Serving (vLLM)

pip install --no-deps ./vllm_plugin   # arch + looped-draft fixes (not upstream yet)
VLLM_ATTENTION_BACKEND=TRITON_ATTN \
vllm serve NullSense/Nanbeige4.2-3B-FP8-Dynamic --trust-remote-code \
  --max-model-len 65536 --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --speculative-config '{"method":"eagle3","model":"NullSense/Nanbeige4.2-3B-EAGLE3","num_speculative_tokens":3}'
  • num_speculative_tokens 3: position-2 acceptance is 0.25, position 3 rarely pays its latency.
  • The head was trained against the FP8-Dynamic verifier. It should transfer to the bf16 original and the other quants in the family (same tokenizer, near-identical distributions) but I have only measured it on FP8-Dynamic.

Download just this artifact:

hf download NullSense/Nanbeige4.2-3B-EAGLE3 --local-dir Nanbeige4.2-3B-EAGLE3

Limitations

  • Acceptance measured at T=0.6 on my workload mix. The corpus is summarize/RAG/rewrite weighted with reasoning and code slices; traffic far from that mix may accept differently.
  • Requires the bundled plugin until the arch lands upstream (vLLM PR #49433); the two looped-draft serve fixes in it are not upstream anywhere yet.
  • Draft vocab is 32k (pruned from 166k); rare-token drafting relies on the fallback mapping.
  • English+Chinese target; training data English-only.

The family

artifact role
FP8-Dynamic verifier / recommended serving quant
NVFP4-FP8-LoopShield 4-bit quant, reasoning at FP8 parity
NVFP4A16 fastest quant, summarization-scoped
EAGLE3 (this repo) draft model: faster decode on top of any of them

Links & provenance

  • Base model: Nanbeige/Nanbeige4.2-3B
  • Trained with speculators (online extract_hidden_states pipeline, vLLM v0.25.1)
  • Arch serving support: vLLM PR #49433 plus this repo's vllm_plugin/ (registers the arch out-of-tree and patches the Eagle3 draft for looped targets)
  • Training config: train_command.txt; per-epoch validation: val_metrics.json

Citation

@misc{peciukonis2026nanbeige42eagle3,
  author       = {Pe{\v{c}}iukonis, Matas (NullSense)},
  title        = {Nanbeige4.2-3B-EAGLE3: an EAGLE-3 draft model for a looped transformer via virtual-depth feature taps},
  year         = {2026},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/NullSense/Nanbeige4.2-3B-EAGLE3},
  note         = {Virtual-depth aux taps (3/23/41 of 44) and looped-target draft-naming fixes; thinking-aware on-policy retrain for reasoning-default serving.}
}
Downloads last month
107
Safetensors
Model size
0.8B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NullSense/Nanbeige4.2-3B-EAGLE3

Collection including NullSense/Nanbeige4.2-3B-EAGLE3