sol-high.grpo-verifier-replay.run_default.broadcasts.step_1
AgentPTB sweep checkpoint. Cell sol-high — Codex / gpt-5.6-sol @ effort high.
| field | value |
|---|---|
| plot cell | sol-high |
| driver | Codex / gpt-5.6-sol |
| reasoning effort | high |
| run boot (UTC) | 2026-08-08T07:28:19Z |
| role | intermediate |
| checkpoint path in run | outputs/grpo-verifier-replay/run_default/broadcasts/step_1 |
| shards | 4 |
| size | 18.8 GB |
| base model | Qwen/Qwen3.5-9B-Base |
| eos_token_id | None ⚠️ MISSING 248046 |
Reading the eos field
248046 is <|im_end|>, the token the Qwen3.5 chat template ends every assistant turn with.
Checkpoints missing it do not stop at end-of-turn and overrun the context window, so their
eval numbers are a floor, not a measurement — compare them only against other checkpoints
with the same eos status, or re-package before evaluating.
Cell note: best cell in the sweep
Mapping back to the figures
Join on plot cell above. The sweep figures are keyed by these same cell names.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for agentic-ptb/sol-high.h048.grpo-verifier-replay.run_default.broadcasts.step_1
Base model
Qwen/Qwen3.5-9B-Base