ULTRON-128M Foundation Model (Generation 2)

ULTRON-128M is a high-density foundation model pre-trained from scratch on 2x Tesla T4 GPUs across 65.5 Million clean tokens.

Specifications

  • Parameters: 151,020,288 (~125.8M non-embedding)
  • Hidden Size ($d_{model}$): 768
  • Layers ($L$): 16
  • Attention Heads ($n_q / n_{kv}$): 12 / 4 (Grouped-Query Attention 3:1)
  • Intermediate Size ($d_{ff}$): 2048 (SwiGLU)
  • Vocabulary Size ($V$): 32,768 (Byte-Level BPE)
  • Context Length ($T$): 2048 tokens
  • Positional Embeddings: RoPE ($\theta = 500,000$)
  • Tokens Trained: 65,536,000 clean tokens
Downloads last month
10
Safetensors
Model size
0.2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support