Alloy 312M 2K

Alloy is an experimental 311.9M-parameter English causal language model with a 2,048-token context window. It is a base model for completing English text, not an instruction-following assistant. It was trained from scratch on 5.0B tokens.

Try it in the free Alloy 312M Demo.

Train a model on a normal PC

Alloy is a starter project for training a language model on accessible hardware. Its 312M-parameter, 2K-context model was trained and evaluated on an NVIDIA RTX 3060 Ti with 8 GB VRAM.

Training and data

  • Architecture: decoder-only Transformer, 22 layers, 1,024 hidden size, 16 attention heads, and a 32,768-token byte-level BPE tokenizer.
  • Training: three 1B-token stages followed by a 2B-token continuation.
  • Data: a filtered English mix of FineWeb-Edu, English Wikipedia, and TinyStories.

Evaluation

Zero-shot evaluation used lm-evaluation-harness 0.4.13. standard-eval.json contains the full Alloy result.

Task Bronze Silver Gold Alloy
Best validation loss ↓ 2.998 2.695 2.632 2.609
ARC-Easy 42.93% 47.98% 49.75% 49.92%
ARC-Challenge 18.77% 21.25% 20.99% 21.42%
HellaSwag 27.39% 27.95% 28.63% 29.87%
PIQA 59.09% 59.96% 60.66% 62.95%
SciQ 64.70% 68.80% 69.50% 71.90%
WinoGrande 50.36% 49.41% 49.64% 50.67%

Bronze, Silver, and Gold are successive 1B-token stages; Alloy adds the final 2B-token continuation. Alloy improves on Gold on every listed task.

Use

pip install -r requirements.txt
python -m tiny_english.generate --model-dir . --device cuda --prompt "The invention of the printing press changed" --max-new-tokens 128 --temperature 0.7 --top-k 40 --seed 42

Generation can repeat phrases, make factual errors, or produce incoherent continuations.

Next

The next planned direction is supervised fine-tuning on curated instruction data.

License

The model weights, tokenizer, and bundled inference code are released under MIT. Training sources remain subject to their own terms.

Downloads last month
478
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train helloboy91/tiny-base-300m

Space using helloboy91/tiny-base-300m 1