Baseline drafter for the RSI Bench task speculative-decoder-discovery

A DSpark speculative-decoding drafter for Qwen/Qwen3-8B, with the architecture of DeepSeek's dspark_qwen3_8b_block7 (5 layers, block size 7, target layers 1/9/17/25/33, Markov head rank 256, confidence head, full vocabulary), trained from scratch with speculators 0.8.0 on yemara/specdec-discovery-starter-data for 12 H100-hours (hidden-state server on one H100, trainer on another, 6 hours, 44,346 steps of 8,192 packed tokens, seed 0). train.sh is the exact recipe. No pretrained drafter weights were used. It is the equal-compute baseline of the task.

Serve with vLLM 0.30.0: vllm serve Qwen/Qwen3-8B --speculative-config '{"method": "dspark", "model": "<this repo>", "num_speculative_tokens": 7}'

Downloads last month
35
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yemara/specdec-discovery-baseline

Finetuned
Qwen/Qwen3-8B
Finetuned
(2171)
this model