qwen3-8B-async-iter179

GRPO-trained deep-research agent on BrowseComp-Plus (search / open_page / finish tools over a fixed 100K-doc corpus). Cell qwen3-8B/baseline of the length-penalty experiment matrix, checkpoint at training iteration 179.

  • Code: https://github.com/ys-2020/miles (branch browsecomp-rl, see docs/experiments/browsecomp-length-penalty-results.md for the full study)
  • WandB: project browsecomp-b300
  • Trained with miles fully-async GRPO: group size 8, global batch 256, temp 1.0, lr 1e-6, KL 0.001, ~100-turn ReAct rollouts, 40960-token context.

Offline eval (150 BrowseComp-Plus test questions, temp 0.6, 1 sample/q)

iter accuracy mean response len (tokens) truncated ratio
19 0.273 2381 0.24
39 0.293 2584 0.26
59 0.313 2598 0.33
79 0.273 2683 0.30
99 0.367 2646 0.39
119 0.307 2856 0.43
139 0.353 2766 0.44
159 0.373 2751 0.39
179 0.360 2895 0.37

Resuming training in miles

Convert back with tools/convert_hf_to_torch_dist.py, rename the saved dir to iter_0000179, write 179 into latest_checkpointed_iteration.txt, then launch with RESUME=1 (see examples/browsecomp/slurm_gb300/).

training_metadata.json in this repo records provenance.

Downloads last month
27
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shangy/browsecomp-qwen3-8B-async-iter179

Finetuned
Qwen/Qwen3-8B
Finetuned
(1964)
this model