GPT-2 (124M), trained from scratch: base model

The pretrained model: 2.5B tokens of FineWeb-Edu, validation loss 3.300. It continues text; it was not trained to follow instructions. Part of github.com/zuurashad/gpt2-from-scratch, a PyTorch reproduction of GPT-2 small trained from scratch on a single 6GB laptop GPU. Try it in your browser.

The architecture is exactly GPT-2 small: 12 layers, 12 heads, width 768, a 1024-token context and the GPT-2 BPE tokenizer. The weights load into Hugging Face's GPT2LMHeadModel unchanged.

Files

file what it is
model.safetensors full-precision (fp32) weights
onnx/model_quantized.onnx per-channel int8 ONNX for the browser demo. Validation loss vs fp32: +0.0018 generating token by token (8,192 FineWeb-Edu tokens), +0.0048 over one forward pass (51,200 tokens)
training.json training step, validation loss and the training arguments

Use it

from transformers import pipeline

generate = pipeline("text-generation", model="zuu007/gpt2-from-scratch")
print(generate("Photosynthesis is the process by which", max_new_tokens=60)[0]["generated_text"])

In the browser, with transformers.js: await AutoModelForCausalLM.from_pretrained("zuu007/gpt2-from-scratch", { dtype: "q8" }).

Evaluation

Zero-shot, in fp32, scored by per-choice log-likelihood (acc, and length-normalised acc_norm as in lm-evaluation-harness). The baseline is OpenAI's GPT-2 124M through the same harness.

task metric this model OpenAI GPT-2 124M
hellaswag acc_norm 27.03% 29.55%
arc_easy acc_norm 42.30% 38.17%
arc_challenge acc_norm 23.21% 22.95%
piqa acc_norm 60.45% 61.81%
openbookqa acc_norm 27.00% 27.20%
winogrande acc_norm 52.96% 51.62%
boolq acc_norm 52.19% 48.64%
lambada acc 20.80% 32.56%
lambada perplexity 81.1 18.0

Limitations

This is a 124M-parameter model trained on 2.5B tokens. It writes fluent English, but it is frequently wrong, especially about facts, arithmetic and recent events. Its only alignment is supervised fine-tuning on a small instruction dataset, so it can produce incorrect or inappropriate text.

Data and licences

  • Pretraining: FineWeb-Edu (ODC-By 1.0)
  • Fine-tuning (chat model only): databricks-dolly-15k, licensed CC BY-SA 3.0. Treat the chat weights as share-alike.
  • Code: MIT, derived in part from Andrej Karpathy's build-nanogpt (MIT)
Downloads last month
16
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train zuu007/gpt2-from-scratch

Space using zuu007/gpt2-from-scratch 1