Armenian-MiniLlama 100M (Chinchilla 2.1B Optimal)
The first dedicated Armenian foundation language model trained strictly from scratch (ab initio) on a 2.56 Billion Token multi-domain Armenian mega-corpus using an AMD Instinct MI300X accelerator.
Architecture
- Parameters: 100,682,496 (~100.7M)
- Attention: Grouped-Query Attention (GQA, 12 query heads, 4 KV heads)
- Feed-Forward: SwiGLU ($d_{\text{ff}} = 2048$)
- Positional Encoding: Rotary Embeddings (RoPE, $\theta = 10000.0$)
- Context Length: 512 tokens
- Precision:
bfloat16with SDPA FlashAttention - Tokenizer: Custom Byte-Level BPE (16,384 vocab) with full Armenian alphabet (
ิฑ-ี), archaic characters, punctuation (ึ,ี,ี,ี), and Latin acronym coverage.
Training Details
- Training Tokens: 2,097,152,000 (~2.10 Billion tokens, strictly Chinchilla compute-optimal)
- Batch Size: 256 ($131,072$ tokens per step)
- Steps: 16,000 steps
- Learning Rate: Cosine decay from $5 \cdot 10^{-4}$ to $5 \cdot 10^{-5}$
- Hardware: AMD Instinct MI300X (192 GB HBM3, 750W TDP)
- Final Validation Perplexity: 12.57
- Final Cross-Entropy Loss: 2.5312
Multi-Domain Benchmark Highlights
- Distinct-3 Diversity: 0.991
- Demonstrates zero repetition stutter, proper Armenian declension cases, verb aspect agreements, and punctuation placement.
- Downloads last month
- 343