TinyStories 22M

A 22.7M parameter language model trained from scratch on TinyStories (V2, GPT-4-generated short children's stories), as a learning project. It is a small educational model: useful for learning how language models work and for quick experiments, not for any real application. Its stories are fluent and usually complete, with simple plots and morals.

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("tegancao/tinystories-22m")
model = AutoModelForCausalLM.from_pretrained("tegancao/tinystories-22m")

torch.manual_seed(0)
ids = tok("<|endoftext|>" + "Once upon a time, a little fox", return_tensors="pt").input_ids   # start like a new document
out = model.generate(ids, max_new_tokens=100, do_sample=True, temperature=0.8, top_p=0.9)
print(tok.decode(out[0, 1:], skip_special_tokens=True))

Sample output (the code above, as is):

Once upon a time, a little fox named Tim went to the woods to find food. He was very hungry and wanted to eat. He walked and walked until he saw a big, red apple. Tim thought, "This apple looks delicious!"
Tim tried to take the apple, but it was too heavy. He tried and tried, but he could not move it. Then, a friendly bird named Sue saw Tim. Sue flew down and said, "I can help you!"
Sue helped Tim get the apple. They took it

It is a base model: it continues text and does not follow instructions or answer questions. Prompts that look like the start of its training documents work best.

Training

Training data TinyStories (V2, GPT-4-generated short children's stories)
Tokenizer byte-level BPE trained on the same data, 10,000-token vocabulary, GPT-2 pre-tokenization pattern; `<
Architecture Llama-style decoder: 4 layers, d_model 512, 16 heads (head dim 32), SwiGLU FFN (d_ff 1344), RMSNorm (pre-norm), RoPE (θ = 10,000), no biases, untied input/output embeddings
Context length 256 tokens
Optimizer AdamW (β₁ 0.9, β₂ 0.95, ε 1e-8, weight decay 0.1), gradient clipping at 1.0
Learning rate 3e-3 peak, 1,000 linear warm-up steps, cosine decay to 3e-4
Batch 32 sequences × 256 tokens = 8,192 tokens per step
Steps / tokens 40,000 steps = 327.68M tokens
Precision / hardware float32, one NVIDIA L4 GPU
Training time ~3.1 hours

The model was implemented and trained with my own from-scratch PyTorch code, then converted to the standard LlamaForCausalLM format (RoPE pairs reordered in the Q/K projections). The conversion was verified: the converted model's logits match the original within ~2 × 10⁻⁵, greedy generations are identical, and the tokenizer produces identical token ids on 2 MB of validation text.

Evaluation

Metric Value
Validation loss (cross-entropy, nats per token) 1.382
Validation perplexity 4.0

Measured on held-out validation text at the end of training, averaged over 20 random batches (about 164K tokens). The model saw about 61% of its training data (one partial pass).

Evaluation and experiments

In a benchmark comparison with the original TinyStories models, Pythia and GPT-2, this model scores 59.8% on the BLiMP grammar test and is near chance on reasoning benchmarks (HellaSwag, PIQA, ARC), as expected for a model of children's stories.

Intended use

  • Learning how language models work: inspecting a small model, experimenting with sampling (temperature, top-p), seeing how training data shapes behavior.
  • Fast experiments that need a tiny model: decoding methods, interpretability, testing inference code.
  • Demos of what a model trained from scratch on a small budget can and cannot do.

Out of scope

  • Any factual use: the model makes things up.
  • Anything user-facing or in production: there is no instruction tuning and no safety tuning.
  • Decisions about people, or any application where mistakes matter.

Limitations

  • Continues text instead of answering it (base model, no instruction tuning).
  • Loose logic and weak tracking of who is who over longer passages; falls into repetition loops, especially with greedy decoding.
  • Sees at most the last 256 tokens.
  • Knows only the world of simple children's stories: no general knowledge.

Acknowledgments

The architecture and training setup follow the publicly available materials of Stanford CS336 (Language Modeling from Scratch). Training data: TinyStories (V2, GPT-4-generated short children's stories). Please see the dataset's own card for its license and terms.

Downloads last month
-
Safetensors
Model size
22.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train tegancao/tinystories-22m