YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
MoS-27B-Checkpoints
Epoch checkpoints of the Qwen3.8-27B DFlash2 MoS line (project MoS, branch ryan/mos-improve,
experiments/dflash2/qwen3.8-27b-100k/). Started 2026-09-22, when the general archive repo
ryan-0608/MoS-Aurora-Experiment-Archive reached the Hugging Face 20,000-file limit.
Layout
dflash2_27b_traj_20260915/main-800k/<arm>/epoch<N>/
model.safetensors drafter weights (speculators DFlash2MoSDraftModel)
optimizer_state_dict.pt full AdamW state -> exact resume
scheduler_state_dict.pt, training_state.json, config.json, config.py, train_command.txt
val_metrics.json trainer held-out 10% validation at the end of the epoch
serving/ fixed-500 serving result for this checkpoint
acceptance_summary.json pooled AL = sum completion tokens / sum verify calls
acceptance_trace.jsonl per-request counts
client.log, server.log
Arms
| arm (path) | Ryan's name | recipe |
|---|---|---|
dense4 |
27B dense (4-epoch baseline) | dense drafter warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4; cosine over 4 epochs, warmup 0.005; same corpus, 3 verifier + 5 trainer layout, global batch and 73,670 steps per epoch as x4rand; training (started 2026-09-25) |
x4rand |
27B 4expert rand LR6 | K=4 full-width expert MLPs per draft layer, random init (gate/up N(0,0.02), down N(0,1e-3)); shared MLP, attention and the rest warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4, expert LR 6e-4, router LR 5e-4; top-2, tau 0.9, balance 0.01; cosine over 4 epochs, warmup 0.005; corpus Current-800K-Self-27B-Traj (719,455 train rows); 73,670 steps per epoch |
Results (target Qwen/Qwen3.8-27B @1d4bf0f2)
Serving protocol: sglang main f5866545, TP2, DFLASH 16 draft tokens, fixed-500 prompts, max 128 new tokens, greedy, concurrency 1, 500/500 completed. Same protocol as the dense and K4-jitter references.
| arm | epoch | global_step | train-side val AL | serving pooled AL | vs dense same epoch |
|---|---|---|---|---|---|
| dense4 | 1 | 73,670 | 5.0262 | 5.0571 | โ |
| dense4 | 2 | 147,335 | 5.0261 | 5.0510 | โ |
| dense4 | 3 | 221,001 | 5.026 (from train.log) | 5.0558 | โ |
| x4rand | 1 | 73,670 | 5.1485 | 5.2314 | +3.45% (dense4 5.0571) |
| x4rand | 2 | 147,335 | 5.2056 | 5.2922 | +4.78% (dense4 5.0510) |
| x4rand | 3 | 221,001 | 5.2202 | 5.3261 | +5.34% (dense4 5.0558) |
| x4rand | 4 | 294,682 | 5.2201 | 5.3113 | +5.17% vs dense E2 5.0502; best epoch = 3 |
Reference arms (model-only, no optimizer) are in ryan-0608/MoS-Aurora-Experiment-Archive under
dflash2_27b_traj_20260915/main-800k/: dense/epoch1, dense/epoch2, mos/epoch1 (K4 shared+1% jitter,
expert LR 1e-4; serving 5.1501, +1.83%).
Use
Resume training: pass the epoch directory as the checkpoint dir to speculators train.py (same train_command.txt,
--epochs 4). Serve: export_d2_for_sglang.py <epoch dir> <out> then sglang --speculative-algorithm DFLASH --speculative-draft-model-path <out>; the MoS architecture needs apply_serving_dflash2_mos.py on sglang main.