Kat
A ~5.2M-parameter decoder-only transformer trained from scratch on DailyDialog. It makes short, grammatical small talk. It knows no facts and has no memory beyond a 128-token context window — that is by design.
Best validation loss 3.346 (perplexity 28.4), trained in 37 minutes on 12 CPU cores, no GPU.
Setup
uv venv --python 3.12 .venv
VIRTUAL_ENV=.venv uv pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
VIRTUAL_ENV=.venv uv pip install -r requirements.txt
torch must be installed separately from the CPU index: that index also serves an
old requests which conflicts with datasets.
Build it
.venv/bin/python data.py # download + clean -> data/{train,val}.txt
.venv/bin/python tokenizer.py # 8k BPE -> tokenizer.json, data/*.bin
.venv/bin/python check_model.py # sanity gates; run before training
.venv/bin/python train.py # ~37 min -> kat.pt, trainlog.csv
Talk to it
.venv/bin/python chat.py
.venv/bin/python chat.py --say "hello, how are you?"
.venv/bin/python evaluate.py # fixed probe prompts + val perplexity
Layout
| file | role |
|---|---|
config.py |
every hyperparameter |
data.py |
download, clean, format dialogues |
tokenizer.py |
byte-level BPE, encodes splits to uint16 |
model.py |
the transformer, hand-written |
check_model.py |
init-loss, causality and generation gates |
train.py |
training loop, saves best-by-val checkpoint |
chat.py |
REPL and the KatChat helper |
evaluate.py |
fixed probes for comparing runs |
PLAN.md |
the plan, plus a log of what actually differed |
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support