Kat

A ~5.2M-parameter decoder-only transformer trained from scratch on DailyDialog. It makes short, grammatical small talk. It knows no facts and has no memory beyond a 128-token context window — that is by design.

Best validation loss 3.346 (perplexity 28.4), trained in 37 minutes on 12 CPU cores, no GPU.

Setup

uv venv --python 3.12 .venv
VIRTUAL_ENV=.venv uv pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
VIRTUAL_ENV=.venv uv pip install -r requirements.txt

torch must be installed separately from the CPU index: that index also serves an old requests which conflicts with datasets.

Build it

.venv/bin/python data.py        # download + clean -> data/{train,val}.txt
.venv/bin/python tokenizer.py   # 8k BPE -> tokenizer.json, data/*.bin
.venv/bin/python check_model.py # sanity gates; run before training
.venv/bin/python train.py       # ~37 min -> kat.pt, trainlog.csv

Talk to it

.venv/bin/python chat.py
.venv/bin/python chat.py --say "hello, how are you?"
.venv/bin/python evaluate.py    # fixed probe prompts + val perplexity

Layout

file role
config.py every hyperparameter
data.py download, clean, format dialogues
tokenizer.py byte-level BPE, encodes splits to uint16
model.py the transformer, hand-written
check_model.py init-loss, causality and generation gates
train.py training loop, saves best-by-val checkpoint
chat.py REPL and the KatChat helper
evaluate.py fixed probes for comparing runs
PLAN.md the plan, plus a log of what actually differed
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support