YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

DFlash 路 Qwen3-8B draft 路 step 66604

Device label: RTX PRO5000. Target model: Qwen3-8B.

Core files

File Contents
model.safetensors, config.json Draft weights and inference configuration for the selected checkpoint
training/launch.yaml Core training parameters; supply deployment settings for your environment
training/draft_config.json Model architecture and algorithm parameters used during training
tensorboard/ Logs from the entire training run, including steps beyond the selected checkpoint where applicable
bench/results.json Per-benchmark and aggregate results for both verification widths, at full numerical precision
README.md Training method and benchmark results

Training

Data source: Regenerated PerfectBlend training data (JSONL).

Training uses PerfectBlend data regenerated by Qwen3-8B at temperature 0 with thinking disabled, formatted with the qwen3-instruct template and no system prompt. The target model supplies hidden states and the information required for distillation, and the draft model trains on the extracted features. The 32 training ranks run separately from the 8 target inference services, with Mooncake transferring features between them.

This model was trained with max_length 8192, while DSpine in the same comparison used 3072. Training was planned for 6 epochs; the selected checkpoint steps and training sequence lengths differ across methods.

Parameter Setting
Planned total steps / selected checkpoint 66604 / 66604
block_size / max_length 16 / 8192
Draft layers / hidden size / FFN size 5 / 4096 / 12288
attention heads / KV heads 32 / 8
Target feature layer indices 1, 9, 17, 25, 33
Microbatch per rank / gradient accumulation 2 / 2
Effective global batch size 32 脳 2 脳 2 = 128
Optimizer AdamW with FP32 master parameters, weight_decay=0
Peak learning rate / schedule 6e-4 / cosine
LR warmup / gradient clipping 4% of total steps / max_grad_norm=1.0
num_anchors / loss_decay_gamma 512 / 7.0
objective_chunk_blocks / seed 128 / 42
Parameter dtype / attention BF16 / FlexAttention
Gradient strategy / kernel NO_SHARD / DDP; Liger

DFlash conditions on intermediate features from the target model to predict a masked draft block in parallel. It trains with token cross-entropy weighted by position decay. The fused plain head is enabled, and additional teacher metrics during training are disabled. Training completed at step 66604, and this model is the final checkpoint.

See training/launch.yaml for the core training parameters; supply deployment settings for your environment. The learning rate schedule and training phases defined as fractions of the run use the planned total step count.

Benchmark (temperature=0)

Each mode contains 2400 requests: GSM8K 1319, MATH500 500, HumanEval 164, MBPP 257, and MT-Bench with 80 two-turn conversations (160 requests). Both draft16/verify16 (D16/V16) and draft16/verify8 (D16/V8) use a draft block size of 16; only the target verification width changes. The verification width includes the anchor.

Benchmark Requests D16/V16 request macro 蟿 D16/V16 micro 蟿 D16/V8 request macro 蟿 D16/V8 micro 蟿
GSM8K 1319 8.771741 8.282009 6.153226 6.019889
MATH500 500 8.588024 7.879332 6.081811 5.881190
HumanEval 164 6.520083 6.418145 5.264523 5.258525
MBPP 257 6.077946 5.690416 5.051469 4.879045
MT-Bench 160 4.717996 3.762506 4.005229 3.529633
All requests, pooled 2400 8.020893 7.060505 5.816440 5.512144
Aggregate D16/V16 D16/V8
Equal-weight mean of request macro 蟿 across the five benchmarks 6.935158 5.311252
Equal-weight mean of micro 蟿 across the five benchmarks 6.406482 5.113657
pooled completion_tokens 1030403 1032149
pooled verify_count 145939 187250

See bench/results.json for results and counts at full numerical precision.

For each request, 蟿 = completion_tokens / verify_count, including target correction / bonus tokens. Request macro is the arithmetic mean of request-level 蟿; micro is the total token count divided by the total number of verification rounds. Pooled metrics aggregate all requests. The equal-weight benchmark means first compute each benchmark metric, then average across the five benchmarks.

Evaluation uses SGLang 0.5.21 with the corresponding algorithm adapter, BF16, TP=1, FlashInfer, and context_length=16384; temperature=0, thinking disabled, no system prompt, and max_new_tokens=2048. Each replica uses max_running_requests=1, with CUDA graphs and the Radix cache disabled. Each round generates a draft block of 16, and verification selects 15 or 7 draft candidates plus the anchor. The tables measure acceptance length, not task accuracy or throughput.

Benchmark (temperature=1)

Target temperature is 1, top_p=1 and top_k=-1 (no truncation). Thinking is disabled, the system prompt is empty, and max_new_tokens=2048. Each mode contains 2400 requests on the same five benchmarks as the temperature=0 evaluation.

Native draft policy: greedy. DFlash and Domino keep greedy proposals; DSpark and DSpine retain their native temperature-conditioned proposal sampling. These are native-policy comparisons, not a claim of identical draft distributions across methods.

Both modes compute the complete draft16 before verification. Verify16/8 includes the anchor, so 15/7 candidates are verified. DSpark computes its native 16 candidates before truncation. Sampling verification uses the actual proposal distribution, including the sparse candidate distribution for DSpine.

Benchmark Requests D16/V16 request macro 蟿 D16/V16 micro 蟿 D16/V8 request macro 蟿 D16/V8 micro 蟿
GSM8K 1319 7.367730 6.668662 5.515138 5.271687
MATH-500 500 6.474364 5.419924 5.050207 4.598771
HumanEval 164 5.404718 5.220332 4.526200 4.451852
MBPP sanitized 257 5.276954 4.830809 4.545214 4.233734
MT-Bench 160 4.096092 3.093140 3.531306 2.989503
All requests, pooled 2400 6.605476 5.422073 5.114581 4.582415
Aggregate D16/V16 D16/V8
Equal-weight mean of benchmark macro 蟿 5.723972 4.633613
Equal-weight mean of benchmark micro 蟿 5.046573 4.309110
Pooled completion_tokens 1062439 1061961
Pooled verify_count 195947 231747

Macro averages per-request completion_tokens/verify_count; micro divides total completion tokens by total verification rounds. Correction/bonus tokens are included. Pooled metrics are request-weighted; equal-weight benchmark means are listed separately.

This is one stochastic benchmark run per setting, not a multi-seed confidence interval. Request seeds follow 42 + uint32_le(SHA256(dataset:sample_id:turn)[:4]); RNG consumption and generated histories can differ by model, width, and replica assignment. Small differences should not be interpreted as statistically established improvements.

Runtime: SGLang 0.5.21, BF16, TP=1, two independent replicas per mode, one request per replica, context_length=16384, CUDA graphs and Radix cache disabled. A process-local overlay enables sampled split verification without changing the original installed runtime. Synthetic distribution/acceptance tests and real-model structural smoke audits passed before the full benchmark; these checks are not a proof of bitwise equivalence to target-only decoding.

The original temperature=0 values remain unchanged. Full-precision temperature=1 metrics, counters, model hashes and runtime provenance are under temperature_1 in bench/results.json. These measurements are acceptance lengths, not answer accuracy or throughput.

Downloads last month
34
Safetensors
Model size
1B params
Tensor type
BF16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support