YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
DFlash 路 Qwen3-8B draft 路 step 66604
Device label: RTX PRO5000. Target model: Qwen3-8B.
Core files
| File | Contents |
|---|---|
model.safetensors, config.json |
Draft weights and inference configuration for the selected checkpoint |
training/launch.yaml |
Core training parameters; supply deployment settings for your environment |
training/draft_config.json |
Model architecture and algorithm parameters used during training |
tensorboard/ |
Logs from the entire training run, including steps beyond the selected checkpoint where applicable |
bench/results.json |
Per-benchmark and aggregate results for both verification widths, at full numerical precision |
README.md |
Training method and benchmark results |
Training
Data source: Regenerated PerfectBlend training data (JSONL).
Training uses PerfectBlend data regenerated by Qwen3-8B at temperature 0 with thinking disabled, formatted with the qwen3-instruct template and no system prompt. The target model supplies hidden states and the information required for distillation, and the draft model trains on the extracted features. The 32 training ranks run separately from the 8 target inference services, with Mooncake transferring features between them.
This model was trained with max_length 8192, while DSpine in the same comparison used 3072. Training was planned for 6 epochs; the selected checkpoint steps and training sequence lengths differ across methods.
| Parameter | Setting |
|---|---|
| Planned total steps / selected checkpoint | 66604 / 66604 |
| block_size / max_length | 16 / 8192 |
| Draft layers / hidden size / FFN size | 5 / 4096 / 12288 |
| attention heads / KV heads | 32 / 8 |
| Target feature layer indices | 1, 9, 17, 25, 33 |
| Microbatch per rank / gradient accumulation | 2 / 2 |
| Effective global batch size | 32 脳 2 脳 2 = 128 |
| Optimizer | AdamW with FP32 master parameters, weight_decay=0 |
| Peak learning rate / schedule | 6e-4 / cosine |
| LR warmup / gradient clipping | 4% of total steps / max_grad_norm=1.0 |
| num_anchors / loss_decay_gamma | 512 / 7.0 |
| objective_chunk_blocks / seed | 128 / 42 |
| Parameter dtype / attention | BF16 / FlexAttention |
| Gradient strategy / kernel | NO_SHARD / DDP; Liger |
DFlash conditions on intermediate features from the target model to predict a masked draft block in parallel. It trains with token cross-entropy weighted by position decay. The fused plain head is enabled, and additional teacher metrics during training are disabled. Training completed at step 66604, and this model is the final checkpoint.
See training/launch.yaml for the core training parameters; supply deployment settings for your environment. The learning rate schedule and training phases defined as fractions of the run use the planned total step count.
Benchmark (temperature=0)
Each mode contains 2400 requests: GSM8K 1319, MATH500 500, HumanEval 164, MBPP 257, and MT-Bench with 80 two-turn conversations (160 requests). Both draft16/verify16 (D16/V16) and draft16/verify8 (D16/V8) use a draft block size of 16; only the target verification width changes. The verification width includes the anchor.
| Benchmark | Requests | D16/V16 request macro 蟿 | D16/V16 micro 蟿 | D16/V8 request macro 蟿 | D16/V8 micro 蟿 |
|---|---|---|---|---|---|
| GSM8K | 1319 | 8.771741 | 8.282009 | 6.153226 | 6.019889 |
| MATH500 | 500 | 8.588024 | 7.879332 | 6.081811 | 5.881190 |
| HumanEval | 164 | 6.520083 | 6.418145 | 5.264523 | 5.258525 |
| MBPP | 257 | 6.077946 | 5.690416 | 5.051469 | 4.879045 |
| MT-Bench | 160 | 4.717996 | 3.762506 | 4.005229 | 3.529633 |
| All requests, pooled | 2400 | 8.020893 | 7.060505 | 5.816440 | 5.512144 |
| Aggregate | D16/V16 | D16/V8 |
|---|---|---|
| Equal-weight mean of request macro 蟿 across the five benchmarks | 6.935158 | 5.311252 |
| Equal-weight mean of micro 蟿 across the five benchmarks | 6.406482 | 5.113657 |
| pooled completion_tokens | 1030403 | 1032149 |
| pooled verify_count | 145939 | 187250 |
See bench/results.json for results and counts at full numerical precision.
For each request, 蟿 = completion_tokens / verify_count, including target correction / bonus tokens. Request macro is the arithmetic mean of request-level 蟿; micro is the total token count divided by the total number of verification rounds. Pooled metrics aggregate all requests. The equal-weight benchmark means first compute each benchmark metric, then average across the five benchmarks.
Evaluation uses SGLang 0.5.21 with the corresponding algorithm adapter, BF16, TP=1, FlashInfer, and context_length=16384; temperature=0, thinking disabled, no system prompt, and max_new_tokens=2048. Each replica uses max_running_requests=1, with CUDA graphs and the Radix cache disabled. Each round generates a draft block of 16, and verification selects 15 or 7 draft candidates plus the anchor. The tables measure acceptance length, not task accuracy or throughput.
Benchmark (temperature=1)
Target temperature is 1, top_p=1 and top_k=-1 (no truncation). Thinking is disabled, the system prompt is empty, and max_new_tokens=2048. Each mode contains 2400 requests on the same five benchmarks as the temperature=0 evaluation.
Native draft policy: greedy. DFlash and Domino keep greedy proposals; DSpark and DSpine retain their native temperature-conditioned proposal sampling. These are native-policy comparisons, not a claim of identical draft distributions across methods.
Both modes compute the complete draft16 before verification. Verify16/8 includes the anchor, so 15/7 candidates are verified. DSpark computes its native 16 candidates before truncation. Sampling verification uses the actual proposal distribution, including the sparse candidate distribution for DSpine.
| Benchmark | Requests | D16/V16 request macro 蟿 | D16/V16 micro 蟿 | D16/V8 request macro 蟿 | D16/V8 micro 蟿 |
|---|---|---|---|---|---|
| GSM8K | 1319 | 7.367730 | 6.668662 | 5.515138 | 5.271687 |
| MATH-500 | 500 | 6.474364 | 5.419924 | 5.050207 | 4.598771 |
| HumanEval | 164 | 5.404718 | 5.220332 | 4.526200 | 4.451852 |
| MBPP sanitized | 257 | 5.276954 | 4.830809 | 4.545214 | 4.233734 |
| MT-Bench | 160 | 4.096092 | 3.093140 | 3.531306 | 2.989503 |
| All requests, pooled | 2400 | 6.605476 | 5.422073 | 5.114581 | 4.582415 |
| Aggregate | D16/V16 | D16/V8 |
|---|---|---|
| Equal-weight mean of benchmark macro 蟿 | 5.723972 | 4.633613 |
| Equal-weight mean of benchmark micro 蟿 | 5.046573 | 4.309110 |
| Pooled completion_tokens | 1062439 | 1061961 |
| Pooled verify_count | 195947 | 231747 |
Macro averages per-request completion_tokens/verify_count; micro divides total completion tokens by total verification rounds. Correction/bonus tokens are included. Pooled metrics are request-weighted; equal-weight benchmark means are listed separately.
This is one stochastic benchmark run per setting, not a multi-seed confidence interval. Request seeds follow 42 + uint32_le(SHA256(dataset:sample_id:turn)[:4]); RNG consumption and generated histories can differ by model, width, and replica assignment. Small differences should not be interpreted as statistically established improvements.
Runtime: SGLang 0.5.21, BF16, TP=1, two independent replicas per mode, one request per replica, context_length=16384, CUDA graphs and Radix cache disabled. A process-local overlay enables sampled split verification without changing the original installed runtime. Synthetic distribution/acceptance tests and real-model structural smoke audits passed before the full benchmark; these checks are not a proof of bitwise equivalence to target-only decoding.
The original temperature=0 values remain unchanged. Full-precision temperature=1 metrics, counters, model hashes and runtime provenance are under temperature_1 in bench/results.json. These measurements are acceptance lengths, not answer accuracy or throughput.
- Downloads last month
- 34