mini-llm-tinystories

A 12.2M-parameter Llama-style language model trained from scratch on TinyStories, following the Stanford CS229 lecture Building Large Language Models: BPE tokenizer β†’ pretraining β†’ scaling laws β†’ SFT β†’ DPO β†’ evaluation, all in plain PyTorch. It writes simple children's stories; it is an educational model, not a general assistant.

Variants

Folder What it is
base/ Pretrained only β€” continues text, ignores instructions
sft/ Supervised fine-tuned on synthetic "write a story using these words" instructions
dpo/ DPO on top of SFT (Ξ²=0.1, 3 epochs, reward v1) β€” the original recipe; shows keyword stuffing
dpo-best/ Best variant from the DPO sweep (selected with fluency/stuffing guardrails) (b0.3-e3-v2)

Architecture

Decoder-only Transformer: 6 layers, d_model 384, 6 heads, context 256, vocab 4096 (byte-level BPE), RMSNorm, RoPE, SwiGLU, tied embeddings.

Evaluation

200 held-out prompts of the form "Write a short story that uses the words: X, Y, Z." Fluency ppl = perplexity of the generated story under the frozen base model.

Model Val ppl ↓ Keyword hit ↑ All 3 words ↑ Finished ↑ Fluency ppl ↓ Mentions/word
base 4.60 7.7% 0.0% 95.0% 7.37 0.14
sft 5.12 31.8% 3.0% 83.0% 2.43 0.74
dpo 5.81 61.2% 25.5% 99.5% 3.40 1.83

DPO sweep

Best: b0.3-e3-v2

Variant Ξ² Epochs Reward All 3 words ↑ Fluency ppl ↓ Mentions/word Guardrails
b0.1-e3-v1 0.1 3 v1 25.5% 3.40 1.83 fail
b0.1-e1-v1 0.1 1 v1 14.5% 2.85 1.40 fail
b0.3-e3-v1 0.3 3 v1 14.5% 2.82 1.51 fail
b0.5-e3-v1 0.5 3 v1 13.0% 2.83 1.34 fail
b0.1-e3-v2 0.1 3 v2 16.5% 3.06 1.10 fail
b0.3-e3-v2 0.3 3 v2 10.5% 2.72 0.95 pass

Scaling law

L(N) = 1.42 + 3201Β·N^-0.60 fitted on 4 small models at 20 tokens/param. Predicted loss for the main model: 1.611; actual: 1.526.

Usage

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("nandutt/mini-llm-tinystories")
sys.path.insert(0, path)  # the repo ships its own code

import torch
from common import build_prompt, load_model_dir
from tokenizer import BPETokenizer

model, cfg = load_model_dir(f"{path}/dpo-best", "cpu")
tok = BPETokenizer.load(f"{path}/tokenizer.json")
ids = tok.encode(build_prompt(["dragon", "cookie", "rain"]))
out, _ = model.generate(torch.tensor([ids]), 200, temperature=0.7, top_k=50,
                        eos_id=tok.eos_id, valid_vocab=tok.vocab_size)
print(tok.decode([t for t in out[0].tolist() if t != tok.eos_id]))

Limitations

  • Trained only on TinyStories: simple vocabulary, children's-story domain, English only.
  • The DPO reward is rule-based (keyword presence + finishing), so it does not measure story quality; watch for keyword stuffing.
  • Some training stories contain mojibake (e.g. Ò€ℒ) that the model reproduces.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train nandutt/mini-llm-tinystories