🧬 NanoDex
Train your first decoder-only language model from scratch, in the cloud.
Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss curve you watch fall in real time.
Pick a size (492k / 1.06M / 5.17M / 8.06M parameters), pick a token budget (200M – 1.5B tokens of fineweb-edu), and queue it. A worker picks it up, you watch the loss fall live, and when it's done the model is pushed to your Hugging Face account with a generated model card.
Pages
| Route | What it is |
|---|---|
/ |
Landing page when signed out; your dashboard when signed in |
/train |
Four-step run wizard — size, tokens, name, review |
/models |
Everything you've trained: training, queued (with position), ready |
/models/{id} |
One run in detail — live loss chart, metrics, logs, publish, playground |
/queue |
Global queue and worker status |
/profile |
Account, totals, permissions |
/about |
The full recipe |
Architecture
A standard modern decoder-only transformer (LlamaForCausalLM): SiLU MLP,
RMSNorm (ε=1e-5), rotary position embeddings (θ=10000), grouped-query attention,
tied input/output embeddings, no biases. The four tiers are that same recipe
scaled down in width and depth.
| Tier | Parameters | Layers | Hidden | Heads (KV) | FFN | Tokens / step |
|---|---|---|---|---|---|---|
| NanoDex-500k | 492,192 | 3 | 96 | 6 (2) | 256 | 131,072 |
| NanoDex-1M | 1,062,272 | 5 | 128 | 8 (4) | 288 | 262,144 |
| NanoDex-5M | 5,172,384 | 9 | 224 | 8 (2) | 592 | 393,216 |
| NanoDex-8M | 8,060,256 | 9 | 288 | 9 (3) | 704 | 524,288 |
A typical small-LM vocabulary (49k tokens) would be 28M embedding parameters on
its own — more than three times the biggest model here. So NanoDex ships its own
2,048-token byte-level BPE trained on fineweb-edu (2.7 chars/token). The
parameter counts above are real totals, embeddings included.
Data
HuggingFaceFW/fineweb-edu,
config sample-10BT. A 1.505B-token slice is tokenized once into a flat
uint16 file (~2.8 GB) that every run reads from, so workers are never
data-bound. Each run starts at its own random offset. Building the cache is a
one-time cost per container; a run asking for more tokens than are cached yet
waits for the writer and reports it in the run log.
Training
AdamW (β=0.9/0.95, wd=0.1, grad-clip 1.0) · 2% warmup then cosine decay to 10% of peak · bf16 autocast · 512-token context · 128–256 sequences per forward pass · 131k–524k tokens per optimizer step, with automatic micro-batch backoff if CUDA runs out of memory. One worker thread per GPU claims the oldest queued run atomically, so a 4-GPU Space trains four models at once.
Layout
main.py FastAPI app — pages, JSON API, Hugging Face OAuth
Dockerfile
web/templates/ landing, home, train wizard, models, detail, queue,
profile, about, error
web/static/ app.css, app.js (dependency-free charts + polling)
nanodex/config.py tier definitions + exact parameter counting
nanodex/data.py shared fineweb-edu token cache
nanodex/trainer.py the training loop
nanodex/worker.py one worker per GPU
nanodex/hub.py push a finished run to the user's account
nanodex/estimate.py ETA, self-calibrating from completed runs
nanodex/db.py SQLite job store
tokenizer/ the 2,048-token BPE
A note on scale
1.5 billion tokens into an 8M-parameter model is ~9× past Chinchilla-optimal, and into a 500k-parameter model over 150× past it — on purpose. Tiny models keep improving long past the compute-optimal point, and the goal here isn't FLOP efficiency — it's watching cross-entropy fall from 7.6 (uniform over 2,048 tokens) toward 4, and seeing a network that started as pure noise begin to emit English words.
Nothing trained here is a useful assistant. That was never the point.
Built with 🤗 by hugging-science.