Small-1B Pretraining Model

Pretraining model (base, no SFT/alignment) — 1B Transformer++ trained from scratch (next-token prediction on quality-filtered tiered data: math, web, code, synth, reformat + gold set).

Spec Value
Params 1,031,898,624 (~1.03B)
Hidden 1536 · 32 layers · 12 attn heads · 4 KV (GQA) · head_dim 128
FFN 4608 SwiGLU
Seq len 2048 (train) / 8192 (config)
Vocab 49152 (SmolLM2-135M tokenizer, BPE)
Optimizer CautiousAdamW, bf16, cosine LR
Best loss 1.7888 @ step 66253 (115K-cycle hot state: stopped at step 68.8K)

This is a raw pretrained base model — no instruction tuning, no chat template. Suitable for continued pretraining, SFT, or DPO. Weight tying: off (untied embeddings).

Downloads last month
341
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kenpeter123/small-1b-pretrain

Finetuned
(932)
this model