BrowseComp perception e2e RL

Private archival release of the Qwen3-8B perception-surrogate checkpoint from the BrowseComp collaboration campaign.

The small model compresses long search observations for a frozen large-model agent. It was trained with complete-episode reward and faithfulness penalties, without length shaping. Revision iter149 is the recommended campaign endpoint: held-out accuracy was 0.780 versus a 0.717 zero-shot perception baseline, while the output-cap rate fell from 82% at iter9 to 60% at iter149.

Provenance

  • Revision: iter149
  • Source export: perc-e2e-rl-gb300-iter149
  • W&B project/run: perception-rl/qtizfliz
  • Experiment documentation: docs/experiments/perception-collab-gb300-campaign.md
  • Emergency source snapshot: ys-2020/miles@5ed731544
  • Archived on 2026-08-21 before cluster checkout

The reported score requires the matching perception harness and is not a standalone language-model benchmark.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shangy/browsecomp-perception-e2e-rl

Finetuned
Qwen/Qwen3-8B
Finetuned
(1973)
this model