PAT-32
Model weights accompanying the PAT training and inference supplement. PAT-32 has 32 layers, width 2048, 32 attention heads, and 1,715,105,280 unique parameters, with tied input/output embeddings. It was trained to a target of 9,520,000,000 lexical tokens and reached 9,520,047,061 tokens after optimizer step 95,473. The document context length is 16,384 tokens.
These are pretrained text continuation weights, with no instruction tuning or
chat template. Use the accompanying pat Python package and its bundled exact
tokenizer. Weights are stored as FP32 safetensors and inference uses CUDA BF16
autocast. The files are a model-weight export; they do not include the original
optimizer state. Starting training from these weights creates a new optimizer.
The artifact manifest lists SHA-256 hashes for every required file. Download using an immutable repository commit and verify the bundle before loading:
from pat.checkpoint import load_checkpoint
from pat.generate import generate
model = load_checkpoint("checkpoints/PAT-32", device="cuda")
print(generate(model, "A useful way to understand attention is", max_new_tokens=64)["text"])
The code supplement documents the architecture, preparation of JSONL documents, training and resume commands, generation, serving, measured verification, and limitations. The package provides one fixed CUDA attention path per architecture.
- Downloads last month
- 274