answer_alright_v2 (corrected v2)

Start from duck_all_v2; paragraph-initial Answer: -> Alright, in reddit_to_flashcards only; preserve standalone capital letters.

Corrected checkpoint for Sophie’s September 13 experiment. Use this v2 model for the corrected experiment; legacy v1 checkpoints are retained separately and differ in capitalization and/or label matching.

Training: 10B-token budget, seed 1337, final step 4769, original recipe commit a21aad24c5f3873d4fdc74f9ef145f978801131b. Export: BF16 safetensors.

Experiment specification. Validated results.

MATH-500, RL-Zero prompt, 500 problems, 7,168 generated-token limit. Named p2 prefixes are a period, two newlines, then the capitalized word. Accuracy is mean individual-rollout correctness (pass@1); screen and 32-rollout runs are kept separate.

Prefix Rollouts/problem Pass@1
none 32 25.02%
p2_alright 32 5.65%
p2_chicken 4 2.85%
p2_duck 32 32.17%
p2_hmm 4 35.00%
p2_okay 32 33.64%

See run_manifest.json for exact edits, prefix strings, seeds, and grader hash.

Downloads last month
220
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support