word

A GPT-2 language model trained from random initialisation on WikiText-2, rather than adapted from OpenAI's released weights. The architecture is GPT-2 small (twelve layers, 768 embedding dimensions, 50257-token vocabulary), which works out at about 124M parameters; the previous version of this card put the figure at 117M, which does not match the 496 MB float32 checkpoint. config.json carries no _name_or_path, confirming nothing was fine-tuned. One structural difference from stock GPT-2 is worth knowing: n_positions is 512 rather than 1024, and the position embedding really is 512 rows, so this model accepts half the context a standard GPT-2 does.

The tokenizer is GPT-2's own byte-pair vocabulary, taken from openai-community/gpt2, because the checkpoint was trained from scratch and therefore ships no vocabulary of its own. It matches: both are 50257 entries, so the token ids in the weights line up with this tokenizer exactly. An earlier version of this repository shipped no tokenizer at all, and AutoTokenizer then returned an empty vocabulary of length one rather than raising, which meant every input collapsed to the unknown token and generation returned noise without any error. That is fixed. GPT-2 uses byte-pair encoding without a distinct unknown token, so rare inputs are segmented rather than discarded, which is one reason a model of this size produces fluent-looking nonsense.

Usage

from transformers import pipeline

generator = pipeline("text-generation", model="harpertoken/word")
print(generator("The quick brown fox", max_new_tokens=60)[0]["generated_text"])

Limitations

WikiText-2 is about two million tokens drawn from Wikipedia verification articles. Training on it for a few epochs produces a model that reproduces that register, namely encyclopaedic and declarative English, and little else. The previous card reported three epochs at batch size one completed in roughly ten minutes on CPU or MPS; for a 124M-parameter model over that corpus at that batch size the arithmetic does not work, so I have left the training time unstated rather than repeat it.

There are no evaluation figures. Perplexity on WikiText-2 would be the obvious measurement and has not been recorded here. Treat this as an architectural and training demonstration: it is a working GPT-2 implementation at small scale, not a language model to build on. Concretely, generation collapses to end-of-text: the end token is the top prediction on ordinary prompts, with measured probability 0.31 to 0.93, so greedy decoding returns the prompt unchanged and sampling usually stops within a token or two. The usage example above will print the prompt back. It has no instruction tuning, no dialogue behaviour, and no pretraining beyond a narrow slice of one corpus.

Attribution

The architecture follows Radford et al., Improving Language Understanding by Generative Pre-Training (2018). WikiText-2 is described in Merity et al., Pointer Sentinel Mixture Models (2016).

Downloads last month
553
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support