YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ο»Ώ# Neo β Custom LLM (500M)
A from-scratch ~500M parameter decoder-only transformer, pre-trained on FineWeb-Edu and fine-tuned for math reasoning and coding. Includes its own SentencePiece tokenizer (50k vocab) and a cached-KV chat generator.
Status β read
TRAINING_REPORT.md. The current checkpoint is undertrained: pretraining stopped at ~1.28B of the 3B-token target (held-out loss 3.93, perplexity 51). That is why chat output looks like noise β it is not a bug. Finishtrain.py(option 4, it resumes automatically) and only then re-run SFT. Usediagnose.py(option 9) to check progress at any time.
Project structure
CustomLLM/
βββ run.bat # Launcher (menu) β always uses the venv Python
βββ Generate.py # Chat with Neo (auto-learns every 5 replies)
βββ Learn.py # Manual catch-up fine-tune on chat history
βββ learn_core.py # Shared self-learning engine (both scripts use it)
βββ train.py # Pre-training (500M, ~100M tokens)
βββ train_sft.py # SFT fine-tuning (reasoning + coding, loss masking)
βββ eval.py # Validation perplexity / loss
βββ diagnose.py # Checkpoint progress + real loss (read-only)
βββ push_to_hub.py # Back up weights + code to Hugging Face
βββ pull_from_hub.py # Restore them on a new machine / OS
βββ setup_linux.sh # Rebuild .venv + deps on Linux
βββ run.sh # Linux launcher (same menu as run.bat)
βββ prepare_data.sh # Linux one-shot data prep
βββ TRAINING_REPORT.md # Why Neo is weak + the fix, in order
βββ build_coding_data.py # Generates coding CoT data -> data/neo_sft_coding.txt
βββ build_reasoning_data.py # Generates math reasoning data -> data/neo_sft_reasoning.txt
βββ model/
β βββ config.py # Model hyperparameters (neo_500m())
β βββ transformer.py # Transformer + KV cache implementation
βββ data/
β βββ build_corpus.py # Stream-download corpus (5 sources)
β βββ prepare.py # Clean + dedup corpus -> data/training.txt
β βββ build_tokens.py # Tokenize pre-training corpus -> data/training.bin (+.meta.json)
β βββ prepare_sft.py # Build SFT tokens + loss masks -> data/sft_*.npy
β βββ optimize_storage.py # uint16 token migration + disk reclaim
β βββ dataset.py # Dataset classes (TextDataset, SFTDataset)
βββ tokenizer/
β βββ neo_final.model # Trained SentencePiece model (50k vocab)
βββ checkpoints/ # Trained weights (.pt β not tracked by git)
Requirements
Run everything through the project virtualenv β never the bare python
command (the Windows Store Python alias hangs at startup due to a broken
gstreamer_bundle.pth in its site-packages):
run.bat :: interactive menu
run.bat Generate.py :: run any script directly
.venv\Scripts\python.exe Generate.py :: equivalent
Dependencies (already installed in .venv): torch (CUDA), sentencepiece,
numpy, datasets.
Pipeline (in order)
| Step | Script | Output | Time (RTX 4070) |
|---|---|---|---|
| 1. Pre-training corpus | data/build_corpus.py |
data/raw/*.txt |
download-bound (resumable) |
| 2. Clean + dedup corpus | data/prepare.py |
data/training.txt |
~10-30 min |
| 3. Tokenize corpus | data/build_tokens.py |
data/training.bin (uint16, ~2.37B tokens, 4.4 GB) |
~1β2 h |
| 4. Pre-train | train.py |
checkpoints/neo_500m_latest.pt |
~9β12 h, resumable |
| 5. SFT data | build_coding_data.py, build_reasoning_data.py |
data/neo_sft_*.txt |
~10 s each |
| 6. Tokenize SFT | data/prepare_sft.py |
data/sft_*.npy (uint16 + loss masks) |
~2β4 min |
| 7. Fine-tune | train_sft.py |
checkpoints/neo_500m_sft_final.pt |
~2β4 h |
| 8. Chat (auto-learns) | Generate.py |
updated neo_500m_sft_final.pt |
interactive |
| 9. Catch-up learn | Learn.py |
updated neo_500m_sft_final.pt |
~1β2 min |
| (opt) Evaluate | eval.py |
val loss/perplexity | ~1 min |
| (opt) Diagnose | diagnose.py |
checkpoint steps + held-out loss | ~1 min |
| (opt) Storage | data/optimize_storage.py |
uint16 tokens + reclaimed disk | seconds |
Learning from conversations (continuous)
Generate.py logs every exchange to data/chat_logs.jsonl and auto-fine-tunes
on new conversations every 5 replies (plus a final catch-up when you exit) β
no manual step needed. Learn.py (or run.bat β option 7) runs the same
process on demand for anything left over. Either way, Neo learns on exactly
what you talked about:
- Deduplicated β conversations already learned (tracked by hash in
data/learned_hashes.txt) are skipped; re-running is a safe no-op. - Properly masked β reuses
data/prepare_sft.pytokenization, so loss is computed on Neo's replies only (user prompts and padding are masked out). - Gentle β low learning rate (2e-5, warmup + cosine decay) so it picks up your material without forgetting its reasoning/coding skills.
- Penalises mistakes β tokens Neo currently gets wrong contribute up to
2.5x more to the update (
wrong_penalty=1.5), so hard errors dominate learning. If the loss is still above 1.0 after a pass, extra remedial rounds run with an escalating penalty (up to 3 extra rounds) until the conversation is actually learned. - Reversible β the previous weights are always backed up to
checkpoints/neo_500m_sft_final_backup.ptbefore any update. - Persistent β learned conversations are also appended to
data/neo_sft_chat.txt, whichdata/prepare_sft.pyincludes, so the knowledge survives any future full SFT re-train.
Use /learn in chat to fine-tune immediately, /forget to delete the log
(not already-learned knowledge); restore the backup checkpoint to undo the
last learning session.
Warm start: removed. The legacy checkpoints/neo_282m_latest.pt (an older
32k-vocab architecture) has been deleted by data/optimize_storage.py β
it was incompatible with the current 500M model.
Checkpoints
| File | Size | Description |
|---|---|---|
neo_500m_sft_final.pt |
1.0 GB | Fine-tuned final (used by Generate.py) |
neo_500m_sft_latest.pt |
3.0 GB | Fine-tuned latest (optimizer state for resume) |
neo_500m_latest.pt |
3.0 GB | Pre-trained latest (optimizer state for resume) |
*_latest.pt files contain optimizer state for resuming interrupted runs;
*_final.pt / inference checkpoints are inference-only and much smaller.
Note: neo_500m_sft_final.pt is rewritten by chat self-learning
(Generate.py / Learn.py), so its step counts SFT/learning steps, not
pretraining steps. neo_500m_sft_final_backup.pt is the undo copy of the
previous state. neo_500m_latest.pt is the only pretraining checkpoint.
Chatting with a specific checkpoint
Generate.py picks the best available checkpoint by default (chat final β
chat latest β pretrained base) β that default is why a plain
run.bat Generate.py reports the ~7k step count. To load a different one
at startup:
run.bat Generate.py --base :: pretrained base
run.bat Generate.py --checkpoint PATH :: any checkpoint file
run.bat Generate.py --no-learn :: no self-learning this session
run.bat Generate.py --base --learn :: force learning anyway (see below)
Or switch without restarting, mid-chat:
/model :: list every checkpoint with its step
/model base :: load checkpoints/neo_500m_latest.pt (the pretrain one)
/model chat :: load checkpoints/neo_500m_sft_final.pt (fine-tuned)
/model backup :: load the pre-self-learning chat backup
/model checkpoints/<file>.pt
Switching clears the conversation so the next reply starts clean under the newly loaded weights.
Remember the two step counters mean different things: neo_500m_latest.pt
(step ~156k, ~1.28B tokens) is pretraining progress, while
neo_500m_sft_final.pt (step ~7.3k) is the SFT/learning counter β a
60M-token SFT run is only ~7.3k steps long, so a small number there is
expected, not a mistake.
--base prints the pretraining progress (step β tokens β % of the 3B target)
and how to resume. The base is not instruction-tuned, so expect raw text
continuation rather than a chat reply β the pretrained checkpoint exists to
show how much pretraining has actually been learned, and to keep pretraining
runs resumable.
Self-learning only ever writes checkpoints/neo_500m_sft_final.pt, so it is
disabled automatically whenever you chat from any other checkpoint (a chat
on the base would otherwise overwrite the fine-tuned model with base weights +
a few updates). While it is disabled nothing is logged to
data/chat_logs.jsonl, so a later Learn.py run cannot fine-tune the chat
checkpoint on raw base output. --learn forces both back on.
Backing up / moving to another machine (Hugging Face)
Everything except the trained weights is regenerable, so the only thing
that must be backed up is checkpoints/. push_to_hub.py uploads the
weights, the tokenizer, all the code and your chat-learned
conversations to a Hugging Face repo, so you can rebuild the project
anywhere β including on Linux after switching OS.
Neo is a hand-rolled nn.Module, not a transformers model, so
push_to_hub() does not exist for it; these scripts use the Hub API
directly and mirror the project layout.
One-time setup
pip install huggingface_hub
hf auth login :: token with "write" permission
Upload (on Windows, before switching)
run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --dry-run :: preview
run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --private :: ~7 GB
run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --with-tokens ;; + data, ~12 GB
| Flag | Adds | Why |
|---|---|---|
--private |
β | keep the weights to yourself (recommended) |
--with-tokens |
data/training.bin, data/sft_*.npy (~5 GB) |
needed to keep training without re-downloading ~12 GB of raw text and re-tokenizing |
--with-raw |
data/raw/*.txt (~12 GB) |
rarely needed; build_corpus.py can re-download |
Default upload = code + docs + tokenizer + all four checkpoints + your
chat history (chat_logs.jsonl, learned_hashes.txt,
neo_sft_chat.txt). That last group is tiny but is the record of what
Neo has actually learned from chatting with you β always include it.
Re-running is cheap: files already on the Hub at the same size are skipped, so after a new checkpoint only that file is pushed.
Note on
data/training.bin: it is a 4.4 GB binary blob that can be rebuilt withdata/build_corpus.pyβdata/prepare.pyβdata/build_tokens.py. Use--with-tokensif you want to skip that (it is many hours of downloading plus ~1β2 h of tokenizing).
Download (on the new machine)
git lfs install
git clone https://huggingface.co/YOUR_USER/neo-500m
cd neo-500m
./setup_linux.sh :: rebuilds .venv, installs deps
./run.sh diagnose.py :: confirm the restored step count
./run.sh train.py :: resumes where it left off
Prefer not to clone? pull_from_hub.py does the same job (and resumes
after interruptions):
python pull_from_hub.py --repo-id YOUR_USER/neo-500m --with-tokens
Switching from Windows to Linux specifically
.venv cannot be copied β Windows and Linux need different torch
wheels. Delete it and rebuild with setup_linux.sh. The Hub repo already
has everything that matters (code, tokenizer, checkpoints, learned
conversations).
Differences on Linux:
| Windows | Linux | |
|---|---|---|
| launcher | run.bat |
./run.sh |
| venv python | .venv\Scripts\python.exe |
.venv/bin/python |
| setup | β | ./setup_linux.sh |
| shell scripts | prepare_data.bat |
./prepare_data.sh |
torch.save checkpoints are portable across OSes, so neo_500m_latest.pt
restores and resumes normally on Linux. If you have an NVIDIA GPU,
install the CUDA build of torch in setup_linux.sh (it is commented out
by default so the script works on any machine).
Storage notes
Token ids are all < 50,000, so corpus/SFT tokens are stored as uint16
(2 bytes/token) instead of int32. data/training.bin carries a
data/training.bin.meta.json sidecar recording the dtype; data/dataset.py
reads it automatically (and still loads legacy int32 files). Run
data/optimize_storage.py to (re)shrink token files or reclaim redundant
disk space β see TRAINING_REPORT.md for the measured savings (~27 GB).
data/raw/ starts at ~12 GB and is never re-read once training.bin
exists, so it can be gzip-compressed (prepare.py/build_corpus.py read
*.txt.gz transparently):
run.bat data\optimize_storage.py --compress-raw :: list candidates
run.bat data\optimize_storage.py --compress-raw --yes :: compress (~several GB)
Only compress once you are done downloading β a *.txt.gz cannot be
appended to, so build_corpus.py would re-download that source.