MA Experiment 001
A small GPT language model trained from scratch for a Swiss Matura research project on language-model efficiency. This repository contains the FP16 inference weights, the matching tokenizer, and the model implementation.
Model details
- Architecture: decoder-only GPT, based on the nanoGPT implementation
- Parameters: 30,384,000
- Configuration: 10 layers, 6 attention heads, hidden width 384, context length 1,024
- Training: 18,310 steps on FineWeb-Edu, seed 1337
- Weights: FP16 reference export of the final pretrained checkpoint
- Tokenizer: custom 32,000-token BPE tokenizer; use the included
tokenizer.json - Intended use: educational experiments and research on compact language models
- Limitations: this is a small base language model, not an instruction-tuned assistant. Outputs can be incoherent, biased, or factually incorrect. The model has not been evaluated for production or safety-critical use.
The weights are saved in the project's PyTorch model format and are not directly loadable with transformers.AutoModel. Use the included model.py implementation.
Generate text
Install PyTorch and the Hugging Face Tokenizers library:
pip install torch tokenizers
Then run this from the directory containing the downloaded files:
import torch
from tokenizers import Tokenizer
from model import GPT, GPTConfig
weights = torch.load("model.fp16.pt", map_location="cpu", weights_only=True)
model = GPT(GPTConfig(**weights["model_args"])).eval().half()
model.load_state_dict(weights["model"])
tokenizer = Tokenizer.from_file("tokenizer.json")
prompt = "Das Modell"
eos_id = tokenizer.token_to_id("<|endoftext|>")
ids = [eos_id] + tokenizer.encode(prompt).ids
input_ids = torch.tensor([ids], dtype=torch.long)
with torch.no_grad():
output = model.generate(input_ids, max_new_tokens=100, temperature=0.8, top_k=40)
print(tokenizer.decode(output[0].tolist()))
CPU inference is supported. For CUDA inference, move both the model and input tensor to a CUDA device.
Provenance
The FP16 export was produced from the final base checkpoint. Source checkpoint SHA-256: 161cfbe96ca708db2986f1edc10293ddf516cb41820c7dc4e36c295c20f86cf9.
- FP16 weights SHA-256:
bd79619ab0b61541b2130ea0e140dcc432e1d3f54c30c1ac95c4431d97ca78c2 - Tokenizer SHA-256:
5b714d3a5ff1aa068038083a48c79e89b55f78c5645857e920ccb0e7fca354d3
The project implementation follows Andrej Karpathy's nanoGPT and retains its MIT license; see LICENSE. Training data source: FineWeb-Edu.