Intermediate RL checkpoints (training trajectory), epoch 1.
- base: SmolLM3-3B, pretraining rung
661Btokens - 31 checkpoints under
step-XXXX/, bf16, inference only - step spacing widens with training: 20 to step 200, then 40, 80, 120
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support