ternary15M
A 15.19M-parameter Llama-style language model trained from scratch on TinyStories where all 42 attention and FFN linear layers are ternary {−1, 0, +1} with per-output-channel absmean scales, in the style of BitNet b1.58. Embeddings (tied input/output) and RMSNorm gains remain FP32.
Training code: github.com/brianbell-x/ternary15M
Architecture
Mirrors karpathy's stories15M: dim 288 · 6 layers · 6 heads · 6 KV heads · SwiGLU hidden 768 · vocab 32,000 · context 256 · RMSNorm · RoPE · tied embeddings. 15,191,712 parameters.
Files
ternary.pt— deployable hard-ternary checkpoint (43 MB). No latent weights: the 42 linear layers store only int8 values in {−1, 0, +1} plus one FP32 scale per output channel. Verified by reloading from disk and generating text.ckpt.pt— full latent training checkpoint (182 MB, FP32 latent weights + optimizer + RNG, resumable, step 10,000).tokenizer.model— the Llama 2 SentencePiece tokenizer from karpathy/llama2.c (vocab 32,000).
Training
- Data: TinyStories (~470M tokens), packed to 256-token sequences with the Llama 2 tokenizer
- 10,000 steps × 65,536 tokens/step = 655M tokens, bf16 autocast, AdamW lr 6e-3 cosine to 10%, warmup 2%, weight decay 0.1, single L40S GPU (~50 min)
- Straight-through estimator: FP32 latent weights updated by the optimizer; forward passes always use the ternarized weights
Evaluation
| Metric | Value |
|---|---|
| Final validation loss (training) | 1.5895 |
| Validation loss, latent STE | 1.5970 |
| Validation loss, hard ternary (deployed form) | 1.6074 |
Hard-ternary inference costs only +0.01 loss over the latent model.
Usage
Requires the training repo for the model definition:
import sys, torch
sys.path.insert(0, "ternary15M") # clone of github.com/brianbell-x/ternary15M
from ternary15m.checkpoint import load_hard_model
from ternary15m.runtime import generate_stories, load_tokenizer
model, meta = load_hard_model("ternary.pt", torch.device("cpu"))
tokenizer = load_tokenizer("tokenizer.model")
print(generate_stories(model, tokenizer, torch.device("cpu"), count=1)[0])
Sample output (hard-ternary, CPU)
Once upon a time, there was a little boy named Tim. Tim loved to bake with his mom. One day, they wanted to make cookies for Mom. Tim was very happy.
Tim's mom said, "Let's mix them together." Tim and his mom worked together to mix the cookies. Tim was very good at this job. They made the cookies pretty and fun. Tim's mom said, "You are a great helper, Tim!"
Citations
@article{ma2024bitnet158,
title = {The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits},
author = {Ma, Shuming and Wang, Hongyu and Ma, Lingxiao and Wang, Lei and Wang, Wenhui and Huang, Shaohan and Dong, Li and Wang, Ruiping and Xue, Jilong and Wei, Furu},
journal = {arXiv preprint arXiv:2402.17764},
year = {2024}
}
@article{eldan2023tinystories,
title = {TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
author = {Eldan, Ronen and Li, Yuanzhi},
journal = {arXiv preprint arXiv:2305.07759},
year = {2023}
}
@misc{llama2c,
title = {llama2.c: Inference Llama 2 in one file of pure C},
author = {Karpathy, Andrej},
howpublished = {\url{https://github.com/karpathy/llama2.c}},
year = {2023}
}
Trained weights released under MIT. TinyStories dataset (Eldan & Li) is MIT-licensed; the tokenizer file is redistributed from llama2.c under its MIT license.