🧬 NanoDex

Train your first decoder-only language model from scratch, in the cloud.

Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss curve you watch fall in real time.

Pick a size (492k / 1.06M / 5.17M / 8.06M parameters), pick a token budget (200M – 1.5B tokens of fineweb-edu), and queue it. A worker picks it up, you watch the loss fall live, and when it's done the model is pushed to your Hugging Face account with a generated model card.

Pages

Route What it is
/ Landing page when signed out; your dashboard when signed in
/train Four-step run wizard — size, tokens, name, review
/models Everything you've trained: training, queued (with position), ready
/models/{id} One run in detail — live loss chart, metrics, logs, publish, playground
/queue Global queue and worker status
/profile Account, totals, permissions
/about The full recipe

Architecture

A standard modern decoder-only transformer (LlamaForCausalLM): SiLU MLP, RMSNorm (ε=1e-5), rotary position embeddings (θ=10000), grouped-query attention, tied input/output embeddings, no biases. The four tiers are that same recipe scaled down in width and depth.

Tier Parameters Layers Hidden Heads (KV) FFN Tokens / step
NanoDex-500k 492,192 3 96 6 (2) 256 131,072
NanoDex-1M 1,062,272 5 128 8 (4) 288 262,144
NanoDex-5M 5,172,384 9 224 8 (2) 592 393,216
NanoDex-8M 8,060,256 9 288 9 (3) 704 524,288

A typical small-LM vocabulary (49k tokens) would be 28M embedding parameters on its own — more than three times the biggest model here. So NanoDex ships its own 2,048-token byte-level BPE trained on fineweb-edu (2.7 chars/token). The parameter counts above are real totals, embeddings included.

Data

HuggingFaceFW/fineweb-edu, config sample-10BT. A 1.505B-token slice is tokenized once into a flat uint16 file (~2.8 GB) that every run reads from, so workers are never data-bound. Each run starts at its own random offset. Building the cache is a one-time cost per container; a run asking for more tokens than are cached yet waits for the writer and reports it in the run log.

Training

AdamW (β=0.9/0.95, wd=0.1, grad-clip 1.0) · 2% warmup then cosine decay to 10% of peak · bf16 autocast · 512-token context · 128–256 sequences per forward pass · 131k–524k tokens per optimizer step, with automatic micro-batch backoff if CUDA runs out of memory. One worker thread per GPU claims the oldest queued run atomically, so a 4-GPU Space trains four models at once.

Layout

main.py                  FastAPI app — pages, JSON API, Hugging Face OAuth
Dockerfile
web/templates/           landing, home, train wizard, models, detail, queue,
                         profile, about, error
web/static/              app.css, app.js (dependency-free charts + polling)
nanodex/config.py        tier definitions + exact parameter counting
nanodex/data.py          shared fineweb-edu token cache
nanodex/trainer.py       the training loop
nanodex/worker.py        one worker per GPU
nanodex/hub.py           push a finished run to the user's account
nanodex/estimate.py      ETA, self-calibrating from completed runs
nanodex/db.py            SQLite job store
tokenizer/               the 2,048-token BPE

A note on scale

1.5 billion tokens into an 8M-parameter model is ~9× past Chinchilla-optimal, and into a 500k-parameter model over 150× past it — on purpose. Tiny models keep improving long past the compute-optimal point, and the goal here isn't FLOP efficiency — it's watching cross-entropy fall from 7.6 (uniform over 2,048 tokens) toward 4, and seeing a network that started as pure noise begin to emit English words.

Nothing trained here is a useful assistant. That was never the point.

Built with 🤗 by hugging-science.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train SLM-Archive/hyperdex-trainer