karpathy/climbmix-400b-shuffle
Preview • Updated • 43.9k • 71
The pretrained foundation of WesleyGPT: a 286-million-parameter language model trained from scratch on one RTX 3060 in Wesley Mangum's home, in about 14 hours. Built on nanochat; code at WasueM/WesleyGPT.
This is a text continuer, not a chat model: give it the start of a document and it writes more. For conversation, use WesleyGPT.
git clone https://github.com/WasueM/WesleyGPT && cd WesleyGPT
uv sync --extra cpu --extra release
from wesleygpt.release import fetch_release, load_release
model, tokenizer = load_release(fetch_release("Wasue/WesleyGPT-Base", "wesleygpt-base"))
It is not a transformers model, so AutoModel will not load it.
| Parameters | 286,261,730 (12 layers, width 768, 6 heads) |
| Context | 2,048 tokens, sliding-window attention |
| Tokenizer | 32,768-token byte-level BPE |
| Data | 1.32 billion tokens of ClimbMix (2,520 steps × 524,288 tokens) |
| Compute | ~14 h on 1× RTX 3060 12 GB |
| Validation loss | 0.844 bits per byte |
| CORE (nanochat's 22-task in-context benchmark) | 0.151 |
CC-BY-NC-4.0, matching NVIDIA's Nemotron-ClimbMix, from which the pretraining data is derived. Free to use, share, and adapt with credit, not commercially. The code is MIT.