YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ο»Ώ# Neo β€” Custom LLM (500M)

A from-scratch ~500M parameter decoder-only transformer, pre-trained on FineWeb-Edu and fine-tuned for math reasoning and coding. Includes its own SentencePiece tokenizer (50k vocab) and a cached-KV chat generator.

Status β€” read TRAINING_REPORT.md. The current checkpoint is undertrained: pretraining stopped at ~1.28B of the 3B-token target (held-out loss 3.93, perplexity 51). That is why chat output looks like noise β€” it is not a bug. Finish train.py (option 4, it resumes automatically) and only then re-run SFT. Use diagnose.py (option 9) to check progress at any time.

Project structure

CustomLLM/
β”œβ”€β”€ run.bat                  # Launcher (menu) β€” always uses the venv Python
β”œβ”€β”€ Generate.py              # Chat with Neo (auto-learns every 5 replies)
β”œβ”€β”€ Learn.py                 # Manual catch-up fine-tune on chat history
β”œβ”€β”€ learn_core.py            # Shared self-learning engine (both scripts use it)
β”œβ”€β”€ train.py                 # Pre-training (500M, ~100M tokens)
β”œβ”€β”€ train_sft.py             # SFT fine-tuning (reasoning + coding, loss masking)
β”œβ”€β”€ eval.py                  # Validation perplexity / loss
β”œβ”€β”€ diagnose.py              # Checkpoint progress + real loss (read-only)
β”œβ”€β”€ push_to_hub.py           # Back up weights + code to Hugging Face
β”œβ”€β”€ pull_from_hub.py         # Restore them on a new machine / OS
β”œβ”€β”€ setup_linux.sh           # Rebuild .venv + deps on Linux
β”œβ”€β”€ run.sh                   # Linux launcher (same menu as run.bat)
β”œβ”€β”€ prepare_data.sh          # Linux one-shot data prep
β”œβ”€β”€ TRAINING_REPORT.md       # Why Neo is weak + the fix, in order
β”œβ”€β”€ build_coding_data.py     # Generates coding CoT data        -> data/neo_sft_coding.txt
β”œβ”€β”€ build_reasoning_data.py  # Generates math reasoning data    -> data/neo_sft_reasoning.txt
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ config.py            # Model hyperparameters (neo_500m())
β”‚   └── transformer.py       # Transformer + KV cache implementation
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ build_corpus.py     # Stream-download corpus (5 sources)
β”‚   β”œβ”€β”€ prepare.py           # Clean + dedup corpus              -> data/training.txt
β”‚   β”œβ”€β”€ build_tokens.py      # Tokenize pre-training corpus     -> data/training.bin (+.meta.json)
β”‚   β”œβ”€β”€ prepare_sft.py       # Build SFT tokens + loss masks    -> data/sft_*.npy
β”‚   β”œβ”€β”€ optimize_storage.py  # uint16 token migration + disk reclaim
β”‚   └── dataset.py           # Dataset classes (TextDataset, SFTDataset)
β”œβ”€β”€ tokenizer/
β”‚   └── neo_final.model      # Trained SentencePiece model (50k vocab)
└── checkpoints/             # Trained weights (.pt β€” not tracked by git)

Requirements

Run everything through the project virtualenv β€” never the bare python command (the Windows Store Python alias hangs at startup due to a broken gstreamer_bundle.pth in its site-packages):

run.bat                          :: interactive menu
run.bat Generate.py              :: run any script directly
.venv\Scripts\python.exe Generate.py   :: equivalent

Dependencies (already installed in .venv): torch (CUDA), sentencepiece, numpy, datasets.

Pipeline (in order)

Step Script Output Time (RTX 4070)
1. Pre-training corpus data/build_corpus.py data/raw/*.txt download-bound (resumable)
2. Clean + dedup corpus data/prepare.py data/training.txt ~10-30 min
3. Tokenize corpus data/build_tokens.py data/training.bin (uint16, ~2.37B tokens, 4.4 GB) ~1–2 h
4. Pre-train train.py checkpoints/neo_500m_latest.pt ~9–12 h, resumable
5. SFT data build_coding_data.py, build_reasoning_data.py data/neo_sft_*.txt ~10 s each
6. Tokenize SFT data/prepare_sft.py data/sft_*.npy (uint16 + loss masks) ~2–4 min
7. Fine-tune train_sft.py checkpoints/neo_500m_sft_final.pt ~2–4 h
8. Chat (auto-learns) Generate.py updated neo_500m_sft_final.pt interactive
9. Catch-up learn Learn.py updated neo_500m_sft_final.pt ~1–2 min
(opt) Evaluate eval.py val loss/perplexity ~1 min
(opt) Diagnose diagnose.py checkpoint steps + held-out loss ~1 min
(opt) Storage data/optimize_storage.py uint16 tokens + reclaimed disk seconds

Learning from conversations (continuous)

Generate.py logs every exchange to data/chat_logs.jsonl and auto-fine-tunes on new conversations every 5 replies (plus a final catch-up when you exit) β€” no manual step needed. Learn.py (or run.bat β†’ option 7) runs the same process on demand for anything left over. Either way, Neo learns on exactly what you talked about:

  • Deduplicated β€” conversations already learned (tracked by hash in data/learned_hashes.txt) are skipped; re-running is a safe no-op.
  • Properly masked β€” reuses data/prepare_sft.py tokenization, so loss is computed on Neo's replies only (user prompts and padding are masked out).
  • Gentle β€” low learning rate (2e-5, warmup + cosine decay) so it picks up your material without forgetting its reasoning/coding skills.
  • Penalises mistakes β€” tokens Neo currently gets wrong contribute up to 2.5x more to the update (wrong_penalty=1.5), so hard errors dominate learning. If the loss is still above 1.0 after a pass, extra remedial rounds run with an escalating penalty (up to 3 extra rounds) until the conversation is actually learned.
  • Reversible β€” the previous weights are always backed up to checkpoints/neo_500m_sft_final_backup.pt before any update.
  • Persistent β€” learned conversations are also appended to data/neo_sft_chat.txt, which data/prepare_sft.py includes, so the knowledge survives any future full SFT re-train.

Use /learn in chat to fine-tune immediately, /forget to delete the log (not already-learned knowledge); restore the backup checkpoint to undo the last learning session.

Warm start: removed. The legacy checkpoints/neo_282m_latest.pt (an older 32k-vocab architecture) has been deleted by data/optimize_storage.py β€” it was incompatible with the current 500M model.

Checkpoints

File Size Description
neo_500m_sft_final.pt 1.0 GB Fine-tuned final (used by Generate.py)
neo_500m_sft_latest.pt 3.0 GB Fine-tuned latest (optimizer state for resume)
neo_500m_latest.pt 3.0 GB Pre-trained latest (optimizer state for resume)

*_latest.pt files contain optimizer state for resuming interrupted runs; *_final.pt / inference checkpoints are inference-only and much smaller.

Note: neo_500m_sft_final.pt is rewritten by chat self-learning (Generate.py / Learn.py), so its step counts SFT/learning steps, not pretraining steps. neo_500m_sft_final_backup.pt is the undo copy of the previous state. neo_500m_latest.pt is the only pretraining checkpoint.

Chatting with a specific checkpoint

Generate.py picks the best available checkpoint by default (chat final β†’ chat latest β†’ pretrained base) β€” that default is why a plain run.bat Generate.py reports the ~7k step count. To load a different one at startup:

run.bat Generate.py --base                      :: pretrained base
run.bat Generate.py --checkpoint PATH           :: any checkpoint file
run.bat Generate.py --no-learn                  :: no self-learning this session
run.bat Generate.py --base --learn              :: force learning anyway (see below)

Or switch without restarting, mid-chat:

/model                 :: list every checkpoint with its step
/model base            :: load checkpoints/neo_500m_latest.pt  (the pretrain one)
/model chat            :: load checkpoints/neo_500m_sft_final.pt (fine-tuned)
/model backup          :: load the pre-self-learning chat backup
/model checkpoints/<file>.pt

Switching clears the conversation so the next reply starts clean under the newly loaded weights.

Remember the two step counters mean different things: neo_500m_latest.pt (step ~156k, ~1.28B tokens) is pretraining progress, while neo_500m_sft_final.pt (step ~7.3k) is the SFT/learning counter β€” a 60M-token SFT run is only ~7.3k steps long, so a small number there is expected, not a mistake.

--base prints the pretraining progress (step β†’ tokens β†’ % of the 3B target) and how to resume. The base is not instruction-tuned, so expect raw text continuation rather than a chat reply β€” the pretrained checkpoint exists to show how much pretraining has actually been learned, and to keep pretraining runs resumable.

Self-learning only ever writes checkpoints/neo_500m_sft_final.pt, so it is disabled automatically whenever you chat from any other checkpoint (a chat on the base would otherwise overwrite the fine-tuned model with base weights + a few updates). While it is disabled nothing is logged to data/chat_logs.jsonl, so a later Learn.py run cannot fine-tune the chat checkpoint on raw base output. --learn forces both back on.

Backing up / moving to another machine (Hugging Face)

Everything except the trained weights is regenerable, so the only thing that must be backed up is checkpoints/. push_to_hub.py uploads the weights, the tokenizer, all the code and your chat-learned conversations to a Hugging Face repo, so you can rebuild the project anywhere β€” including on Linux after switching OS.

Neo is a hand-rolled nn.Module, not a transformers model, so push_to_hub() does not exist for it; these scripts use the Hub API directly and mirror the project layout.

One-time setup

pip install huggingface_hub
hf auth login                          :: token with "write" permission

Upload (on Windows, before switching)

run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --dry-run   :: preview
run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --private   :: ~7 GB
run.bat push_to_hub.py --repo-id YOUR_USER/neo-500m --with-tokens  ;; + data, ~12 GB
Flag Adds Why
--private β€” keep the weights to yourself (recommended)
--with-tokens data/training.bin, data/sft_*.npy (~5 GB) needed to keep training without re-downloading ~12 GB of raw text and re-tokenizing
--with-raw data/raw/*.txt (~12 GB) rarely needed; build_corpus.py can re-download

Default upload = code + docs + tokenizer + all four checkpoints + your chat history (chat_logs.jsonl, learned_hashes.txt, neo_sft_chat.txt). That last group is tiny but is the record of what Neo has actually learned from chatting with you β€” always include it.

Re-running is cheap: files already on the Hub at the same size are skipped, so after a new checkpoint only that file is pushed.

Note on data/training.bin: it is a 4.4 GB binary blob that can be rebuilt with data/build_corpus.py β†’ data/prepare.py β†’ data/build_tokens.py. Use --with-tokens if you want to skip that (it is many hours of downloading plus ~1–2 h of tokenizing).

Download (on the new machine)

git lfs install
git clone https://huggingface.co/YOUR_USER/neo-500m
cd neo-500m
./setup_linux.sh                          :: rebuilds .venv, installs deps
./run.sh diagnose.py                      :: confirm the restored step count
./run.sh train.py                         :: resumes where it left off

Prefer not to clone? pull_from_hub.py does the same job (and resumes after interruptions):

python pull_from_hub.py --repo-id YOUR_USER/neo-500m --with-tokens

Switching from Windows to Linux specifically

.venv cannot be copied β€” Windows and Linux need different torch wheels. Delete it and rebuild with setup_linux.sh. The Hub repo already has everything that matters (code, tokenizer, checkpoints, learned conversations).

Differences on Linux:

Windows Linux
launcher run.bat ./run.sh
venv python .venv\Scripts\python.exe .venv/bin/python
setup β€” ./setup_linux.sh
shell scripts prepare_data.bat ./prepare_data.sh

torch.save checkpoints are portable across OSes, so neo_500m_latest.pt restores and resumes normally on Linux. If you have an NVIDIA GPU, install the CUDA build of torch in setup_linux.sh (it is commented out by default so the script works on any machine).

Storage notes

Token ids are all < 50,000, so corpus/SFT tokens are stored as uint16 (2 bytes/token) instead of int32. data/training.bin carries a data/training.bin.meta.json sidecar recording the dtype; data/dataset.py reads it automatically (and still loads legacy int32 files). Run data/optimize_storage.py to (re)shrink token files or reclaim redundant disk space β€” see TRAINING_REPORT.md for the measured savings (~27 GB).

data/raw/ starts at ~12 GB and is never re-read once training.bin exists, so it can be gzip-compressed (prepare.py/build_corpus.py read *.txt.gz transparently):

run.bat data\optimize_storage.py --compress-raw           :: list candidates
run.bat data\optimize_storage.py --compress-raw --yes     :: compress (~several GB)

Only compress once you are done downloading β€” a *.txt.gz cannot be appended to, so build_corpus.py would re-download that source.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support