profbench9b_specgap_ship_websearch (global step 120)

Merged HF weights (bf16 safetensors) of the verl FSDP checkpoint profbench9b_specgap_ship_websearch/global_step_120.

  • Base model: Qwen/Qwen3.5-9B
  • Experiment: RRIMed (self-evolving rewards), held-out arm on ProfBench, ARM 20
  • Score: ProfBench rubric accuracy 0.392 on the 40 tasks (run best 0.405 at step 170, rotated away; untrained 0.356)
  • Note: Web search only (no domain knowledge base).

This is the best checkpoint of the run whose weights survived checkpoint rotation; the run's best-validation step is stated above when it differs.

Research artifact trained on model-written tasks and evaluated on one benchmark family. Not for clinical use.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Video Preview
loading

Model tree for ddvd233/profbench9b_specgap_ship_websearch_global_step_120

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(776)
this model