GTM-3-base by OpenGCM
A ~75M parameter, decoder-only GPT-style base language model, trained from scratch on a single RTX Pro 6000, using RoPE (rotary position embeddings)
This is a base (pretrained) model, not an instruction-tuned or chat model. It completes/continues text; it does not reliably follow instructions or answer questions directly. It also has no fine-tuning for factual accuracy β it can and does confidently generate fluent, plausible-sounding but factually incorrect content. Treat it as a small, from-scratch research/hobby model, not a production system.
Model details
Architecture: nanoGPT-style decoder-only transformer with RoPE (rotary position embeddings) instead of learned absolute position embeddings
- 10 layers, 8 attention heads, 608 embedding dim
- ~74.9M parameters (weight-tied embeddings/output head)
- Context length: 1024 tokens
- Uses PyTorch's fused
scaled_dot_product_attention(flash-attention kernel) - MLP, LayerNorm, and overall block structure otherwise match the previous GTM-v2-base recipe β RoPE is the one architectural change in this version
Tokenizer:
tiktokenGPT-2 BPE encoding (tiktoken.get_encoding("gpt2")), vocab size 50,257. No custom tokenizer was trained.Optimizer: Muon (for 2D weight matrices) + AdamW (for embeddings, layernorms, biases) β a hybrid setup, following the approach popularized in recent efficient-pretraining recipes.
Precision: trained with bf16 autocast; released weights are fp32.
Training data: streamed via HuggingFace
datasets, mixed and tokenized on the fly:- FineWeb-Edu (
HuggingFaceFW/fineweb-edu,sample-10BT) β 35% - Cosmopedia-v2 (
HuggingFaceTB/cosmopedia-v2) β 25% - FineMath (
HuggingFaceTB/finemath,finemath-4plus) β 15% - TxT360 (
LLM360/TxT360,defaultconfig) β 15% - Wikipedia (
wikimedia/structured-wikipedia,enwiki_namespace_0config) β 10%, flattened from its nested article/section JSON schema into plain text
This mix replaces the raw FineWeb slice used in earlier GTM releases with TxT360 (a more heavily deduplicated/curated web corpus) and adds a dedicated Wikipedia slice for denser factual content β both aimed at known weak points from earlier GTM models (noisy web text, unreliable factual recall).
Note: code data (The Stack v2, StarCoderData, the-stack-smol) was again left out of this base-pretraining mix β every BigCode-hosted code corpus checked so far is gated behind a terms-of-use click-through on HuggingFace. Code is planned as a separate SFT-stage release (GTM-3-coder) rather than being forced into this mix.
- FineWeb-Edu (
Known limitations
- No reliable factual recall. E.g. prompted with "The capital of France is", the model may not consistently produce "Paris" β it can produce fluent but factually invented content instead. The added Wikipedia data is intended to help this relative to earlier GTM releases, but this has not yet been verified.
- No code capability. No code-specific training data was included (see above). The model can produce code-shaped text (recognizing "write code" prompts and using plausible syntax) but not functionally correct code.
- Not instruction-tuned. Does not follow instructions or answer questions in a chat-like way β it continues text.
- Repetition tendency. Base generation (greedy or low-temperature sampling) can fall into repetition loops or lock onto structural templates when it lacks a confident continuation. A repetition penalty (see usage example) substantially reduces this.
- Smaller than prior GTM releases. This was a deliberate choice for this run, not a regression, but it means raw capability is likely lower in absolute terms even where the architecture/data improvements help.
Usage
Requires model.py (included in this repo) alongside the checkpoint β
this is a plain PyTorch model, not a transformers AutoModel.
pip install torch safetensors tiktoken
import json
import torch
from safetensors.torch import load_file
from model import GPT, GPTConfig
with open("config.json") as f:
cfg_dict = json.load(f)
config = GPTConfig(
vocab_size=cfg_dict["vocab_size"], block_size=cfg_dict["block_size"],
n_layer=cfg_dict["n_layer"], n_head=cfg_dict["n_head"],
n_embd=cfg_dict["n_embd"], dropout=cfg_dict["dropout"], bias=cfg_dict["bias"],
rope_theta=cfg_dict.get("rope_theta", 10000.0),
)
model = GPT(config)
state_dict = load_file("model.safetensors")
model.load_state_dict(state_dict)
model.eval()
import tiktoken
enc = tiktoken.get_encoding("gpt2")
prompt = "Once upon a time,"
ids = enc.encode_ordinary(prompt)
x = torch.tensor([ids], dtype=torch.long)
with torch.no_grad():
out = model.generate(
x, max_new_tokens=128, temperature=0.8, top_k=50,
eot_token=enc.eot_token, repetition_penalty=1.3,
)
print(enc.decode(out[0].tolist()))
License
Apache 2.0 for this repo's contents (model weights, model.py, and this
README). The underlying training data retains its own licenses regardless
of the license on this trained model β FineWeb-Edu, Cosmopedia-v2, and
FineMath are each ODC-BY-1.0; TxT360 is ODC-BY-1.0; Wikipedia content via
wikimedia/structured-wikipedia is CC-BY-SA-4.0 + GFDL (see each
dataset's HF card for full terms and attribution requirements). This repo
distributes model weights, not the training data itself.
- Downloads last month
- 120
