mini-llm-tinystories
A 12.2M-parameter Llama-style language model trained from scratch on TinyStories, following the Stanford CS229 lecture Building Large Language Models: BPE tokenizer β pretraining β scaling laws β SFT β DPO β evaluation, all in plain PyTorch. It writes simple children's stories; it is an educational model, not a general assistant.
Variants
| Folder | What it is |
|---|---|
base/ |
Pretrained only β continues text, ignores instructions |
sft/ |
Supervised fine-tuned on synthetic "write a story using these words" instructions |
dpo/ |
DPO on top of SFT (Ξ²=0.1, 3 epochs, reward v1) β the original recipe; shows keyword stuffing |
dpo-best/ |
Best variant from the DPO sweep (selected with fluency/stuffing guardrails) (b0.3-e3-v2) |
Architecture
Decoder-only Transformer: 6 layers, d_model 384, 6 heads, context 256, vocab 4096 (byte-level BPE), RMSNorm, RoPE, SwiGLU, tied embeddings.
Evaluation
200 held-out prompts of the form "Write a short story that uses the words: X, Y, Z." Fluency ppl = perplexity of the generated story under the frozen base model.
| Model | Val ppl β | Keyword hit β | All 3 words β | Finished β | Fluency ppl β | Mentions/word |
|---|---|---|---|---|---|---|
| base | 4.60 | 7.7% | 0.0% | 95.0% | 7.37 | 0.14 |
| sft | 5.12 | 31.8% | 3.0% | 83.0% | 2.43 | 0.74 |
| dpo | 5.81 | 61.2% | 25.5% | 99.5% | 3.40 | 1.83 |
DPO sweep
Best: b0.3-e3-v2
| Variant | Ξ² | Epochs | Reward | All 3 words β | Fluency ppl β | Mentions/word | Guardrails |
|---|---|---|---|---|---|---|---|
| b0.1-e3-v1 | 0.1 | 3 | v1 | 25.5% | 3.40 | 1.83 | fail |
| b0.1-e1-v1 | 0.1 | 1 | v1 | 14.5% | 2.85 | 1.40 | fail |
| b0.3-e3-v1 | 0.3 | 3 | v1 | 14.5% | 2.82 | 1.51 | fail |
| b0.5-e3-v1 | 0.5 | 3 | v1 | 13.0% | 2.83 | 1.34 | fail |
| b0.1-e3-v2 | 0.1 | 3 | v2 | 16.5% | 3.06 | 1.10 | fail |
| b0.3-e3-v2 | 0.3 | 3 | v2 | 10.5% | 2.72 | 0.95 | pass |
Scaling law
L(N) = 1.42 + 3201Β·N^-0.60 fitted on 4 small models at 20 tokens/param.
Predicted loss for the main model: 1.611; actual: 1.526.
Usage
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("nandutt/mini-llm-tinystories")
sys.path.insert(0, path) # the repo ships its own code
import torch
from common import build_prompt, load_model_dir
from tokenizer import BPETokenizer
model, cfg = load_model_dir(f"{path}/dpo-best", "cpu")
tok = BPETokenizer.load(f"{path}/tokenizer.json")
ids = tok.encode(build_prompt(["dragon", "cookie", "rain"]))
out, _ = model.generate(torch.tensor([ids]), 200, temperature=0.7, top_k=50,
eos_id=tok.eos_id, valid_vocab=tok.vocab_size)
print(tok.decode([t for t in out[0].tolist() if t != tok.eos_id]))
Limitations
- Trained only on TinyStories: simple vocabulary, children's-story domain, English only.
- The DPO reward is rule-based (keyword presence + finishing), so it does not measure story quality; watch for keyword stuffing.
- Some training stories contain mojibake (e.g.
Γ’β¬β’) that the model reproduces.