Small language models trained with nanochat (Assignment 1, Agentic LLMs, Leiden University)
Small language models trained with nanochat on one NVIDIA RTX 2070 SUPER (8 GB), for Assignment 1 of the Agentic LLMs course. Authors: Kevin Bretz and An Nguyen. Code, setup and all commands: https://github.com/kevin-bretz/Small-Language-Model
| Folder | Model | Training |
|---|---|---|
checkpoints/d2_pretraining |
depth 2, 13.0M parameters | pre-training on ClimbMix, 420 steps, 55M tokens |
checkpoints/d4_pretraining |
depth 4, 36.7M parameters | pre-training on ClimbMix, 528 steps, 138M tokens |
checkpoints/d2_mid, checkpoints/d4_mid |
stage 1, mid-training on MMLU + GSM8K, 310 steps | |
checkpoints/d2_sft, checkpoints/d4_sft |
stage 2, SFT on SmolTalk, 500 steps | |
tokenizer |
BPE tokenizer with 32,768 tokens, used by all models | 500M characters of ClimbMix |
tokenizer_8192 |
BPE tokenizer with 8,192 tokens (tokenizer comparison only) | 500M characters of ClimbMix |
Every checkpoint folder contains the weights (model_*.pt), nanochat's metadata (meta_*.json) and the optimizer
state (optim_*_rank0.pt).
Using the checkpoints
Set up the code and the virtual environment as described in the README of our GitHub repository.
Copy the folders of this repository into nanochat's cache folder
~/.cache/nanochat/:This repository ~/.cache/nanochat/checkpoints/d2_pretraining,checkpoints/d4_pretrainingbase_checkpoints/d2,base_checkpoints/d4checkpoints/d2_mid,d2_sft,d4_mid,d4_sftchatsft_checkpoints/d2_mid,d2_sft,d4_mid,d4_sfttokenizertokenizerandtokenizer_32768(the same files)tokenizer_8192tokenizer_8192From the repository root, run for example:
python -m assignment.task3 eval d4 # ARC-Easy, ARC-Challenge and GSM8K after each training stage python -m assignment.task4 d4_sft # our temperature study through nanochat's scripts/chat_cli.pyor chat with a model yourself:
python -m scripts.chat_cli -i sft -g d4_sft
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support