Koachie 312M — a language model built from scratch for £60
Koachie is a 312M-parameter GPT-style language model built entirely from scratch by Baller Knowledge in Glasgow, Scotland — own tokenizer, own architecture implementation, own training loop. It is not a fine-tune of any existing model and no external model's weights were used at any point.
Total pretraining cost was roughly £60 of rented GPU time. It serves real users today at koachiefitai.co.uk.
Architecture
| Parameters | 312.06M |
| Layers / width / heads | 24 × 1024d × 16h (head_dim 64) |
| Positions | RoPE |
| Norm / MLP | RMSNorm (pre-norm) / SwiGLU |
| Attention | fused flash attention (PyTorch SDPA) |
| Context | 512 tokens |
| Vocabulary | 8,000 (custom BPE, trained from the corpus) |
| Embeddings | tied |
Training
- Data: ~40B tokens, predominantly FineWeb-Edu, with Wikipedia, Project Gutenberg, PubMed Central open-access and OASST/Dolly dialogue. Licence-aware curation throughout (Reddit-derived data deliberately excluded).
- Optimiser: Muon (Newton–Schulz orthogonalised momentum) for hidden matrices, AdamW for embeddings/norms; WSD (warmup–stable–decay) schedule.
- Hardware: single-GPU rentals (RTX 5090 class) plus a desktop RTX 3070; DDP path written and measured at 137k tokens/sec on 4× RTX 5090.
- A finding worth sharing: Muon learning-rate stability followed a reproducible law across two machines — each halving of LR extended the stable training horizon roughly 8× (0.006 diverged near 2k steps, 0.003 near 25k, 0.0015 near 40k, 0.00075 past 180k).
Evaluation
EleutherAI lm-evaluation-harness, full test sets, no sampling. Two checkpoints: the pretrained base and the conversational SFT that serves the live product.
| task (acc, %) | base | SFT |
|---|---|---|
| SciQ | 71.9 | 66.2 |
| PIQA | 62.0 | 61.7 |
| WinoGrande | 50.4 | 51.4 |
| ARC-Easy | 45.7 | 45.2 |
| OpenBookQA (norm) | 31.8 | 31.8 |
| HellaSwag (norm) | 29.4 | 29.6 |
| ARC-Challenge (norm) | 22.9 | 25.3 |
| LAMBADA (ppl ↓) | 176.1 | 241.0 |
Roughly GPT-2-class performance. The base is the stronger language model (note the perplexity gap); the SFT trades some of that for conversational behaviour, which is the expected pattern.
Honest limitations
This is a small model and it behaves like one. Closed-book, it will confidently state wrong facts; ARC-Challenge is near chance. The production system around it does not rely on the model's memory: factual questions are resolved against Wikidata's structured records or answered by quoting a retrieved source verbatim, generated answers are marked by a judge before they ship, and medical red flags are intercepted and referred to a professional. The model is the voice; verification is the product.
Files
koachie-312m-base.pt— pretrained base checkpoint (step 304,000, ~40B tokens)tokenizer.json— the 8k custom BPE tokenizerlm_eval_base.json/lm_eval_sft.json— raw harness output for the table above
Weights are a plain PyTorch state_dict under model with model_config
alongside; the reference implementation (model code, training loop, evaluation
adapters) lives in the project repository.
Intended use
Research into small-model training economics, retrieval-first product architectures, and as evidence that a working language model is buildable at consumer cost. Not a general-purpose assistant on its own.
Built by Steven McKinney, Baller Knowledge UK Ltd, Glasgow.