ctest: Qwen3-8B DSpark speculator โ FP8 hidden-states ablation (bf16)
Test checkpoint โ not a production release. Trained to validate the FP8-quantized hidden-states transfer backend added in vllm-project/speculators#1028 (draft, reviving #491).
- Verifier:
Qwen/Qwen3-8B - Algorithm: dspark
- Hidden-states precision during training: bf16 (
existing bf16 file backend) - Training data: 5K magpie + 5K ultrachat samples from
inference-optimization/Dataset-Qwen3-235B-Instruct
This is one of a matched pair (...-fp8ablation-bf16 / ...-fp8ablation-fp8) trained
identically except for the hidden-states transfer precision, to isolate the effect of
FP8 quantization on speculator quality. See the sibling repo for the other precision.
Validation metrics (final epoch)
confidence_loss_epoch: 0.252827confidence_abs_error_epoch: 0.224563confidence_pred_mean_epoch: 0.499095confidence_cumprod_bias_epoch: 0.021723loss_epoch: 0.583228ce_loss_epoch: 1.192882tv_loss_epoch: 0.234570accept_rate_epoch: 0.455785accept_len_epoch: 2.873075full_acc_epoch: 0.482406position_0_acc_epoch: 0.698634position_1_acc_epoch: 0.592615position_2_acc_epoch: 0.525559position_3_acc_epoch: 0.472596position_4_acc_epoch: 0.433272position_5_acc_epoch: 0.400534position_6_acc_epoch: 0.376377position_7_acc_epoch: 0.358044
guidellm serving eval
See acceptance.csv in this repo for the full per-subset guidellm breakdown (9 subset rows).
Full ablation writeup, all 6 checkpoints' results, exact commands, and code: see the results package referenced from PR #1028.
- Downloads last month
- 29
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support