profbench9b_specgap_ship_websearch (global step 120)
Merged HF weights (bf16 safetensors) of the verl FSDP checkpoint profbench9b_specgap_ship_websearch/global_step_120.
- Base model: Qwen/Qwen3.5-9B
- Experiment: RRIMed (self-evolving rewards), held-out arm on ProfBench, ARM 20
- Score: ProfBench rubric accuracy 0.392 on the 40 tasks (run best 0.405 at step 170, rotated away; untrained 0.356)
- Note: Web search only (no domain knowledge base).
This is the best checkpoint of the run whose weights survived checkpoint rotation; the run's best-validation step is stated above when it differs.
Research artifact trained on model-written tasks and evaluated on one benchmark family. Not for clinical use.
- Downloads last month
- -