Intermediate RL checkpoints (training trajectory), epoch 1.

  • base: SmolLM3-3B, pretraining rung 283B tokens
  • 31 checkpoints under step-XXXX/, bf16, inference only
  • step spacing widens with training: 20 to step 200, then 40, 80, 120
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support