SmolLM2-135M Reasoning Recovery 30K — 4096

Continued full-parameter supervised fine-tuning from Ma7ee7/SmolLM2-135M-Reasoning-5K.

Reasoning-token format

<think> and </think> are atomic non-special added tokens. This keeps the markers visible to normal decoders and reasoning-aware chat interfaces.

Expected assistant response:

<think>
Short, relevant reasoning.
</think>
Final answer

Training

  • Context length: 4096
  • Train rows: 28,786
  • Validation rows: 1,000
  • Learning rate: 1.5e-05
  • Epochs: 2.0
  • Reasoning loss weight: 1.0
  • Final-answer loss weight: 4.0
  • Seed: 3407
Downloads last month
11
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ma7ee7/SmolLM2-135M-Reasoning-Recovery-30K-4096

Datasets used to train Ma7ee7/SmolLM2-135M-Reasoning-Recovery-30K-4096