olmo32b-cue โ€” GRPO with forced cue opener

LoRA adapters from the cue-forced GRPO run on Olmo-3-1125-32B (MATH training set, 300 steps; every rollout prefilled with the ".\n\nOkay," opener). Companion to the vanilla run at ReasoningRegisters/olmo32b.

Checkpoints 50-300 (every 50 steps): adapter weights + tokenizer + trainer state. Training rollout logs: ReasoningRegisters/olmo32b-cue-rollouts (dataset).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ReasoningRegisters/olmo32b-cue

Adapter
(5)
this model