qwen3-8B-async-iter179
GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page /
finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/baseline of the
length-penalty experiment matrix, checkpoint at training iteration 179.
- Code: https://github.com/ys-2020/miles (branch
browsecomp-rl, seedocs/experiments/browsecomp-length-penalty-results.mdfor the full study) - WandB: project
browsecomp-b300 - Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.
Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)
| iter | accuracy | mean response len (tokens) | truncated ratio |
|---|---|---|---|
| 19 | 0.273 | 2381 | 0.24 |
| 39 | 0.293 | 2584 | 0.26 |
| 59 | 0.313 | 2598 | 0.33 |
| 79 | 0.273 | 2683 | 0.30 |
| 99 | 0.367 | 2646 | 0.39 |
| 119 | 0.307 | 2856 | 0.43 |
| 139 | 0.353 | 2766 | 0.44 |
| 159 | 0.373 | 2751 | 0.39 |
| 179 | 0.360 | 2895 | 0.37 |
Resuming training in miles
Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to
iter_0000179, write 179 into latest_checkpointed_iteration.txt, then
launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).
training_metadata.json in this repo records provenance.
- Downloads last month
- 27
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support