ULTRON-128M Foundation Model (Generation 2)
ULTRON-128M is a high-density foundation model pre-trained from scratch on 2x Tesla T4 GPUs across 65.5 Million clean tokens.
Specifications
- Parameters: 151,020,288 (~125.8M non-embedding)
- Hidden Size ($d_{model}$): 768
- Layers ($L$): 16
- Attention Heads ($n_q / n_{kv}$): 12 / 4 (Grouped-Query Attention 3:1)
- Intermediate Size ($d_{ff}$): 2048 (SwiGLU)
- Vocabulary Size ($V$): 32,768 (Byte-Level BPE)
- Context Length ($T$): 2048 tokens
- Positional Embeddings: RoPE ($\theta = 500,000$)
- Tokens Trained: 65,536,000 clean tokens
- Downloads last month
- 10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support