π¦π² Armenian-SmolLlama-116M-Instruct (30 Layers Deep & Thin)
Armenian-SmolLlama-116M is a state-of-the-art compact Armenian language model built on the modern Deep & Thin paradigm pioneered by Meta's MobileLLM (ICML 2024) and Hugging Face's SmolLM (135M).
Key Architectural Specifications:
- Parameters: 115,640,640 (115.6M)
- Transformer Layers: 30 Layers (2.5x deeper than legacy 12-layer models)
- Hidden Dimension: 576
- Attention: 9 Query Heads, 3 Key-Value Heads (3:1 Grouped-Query Attention)
- Feed-Forward: SwiGLU with intermediate dimension 1536
- Embedding Sharing: Weight tying (
lm_head.weight == tok_embeddings.weight) - Pre-training: 2.1 Billion tokens on AMD Instinct MI300X accelerator.
- Alignment: Case-Augmented LIMA Gold alignment (100% case-invariant).
- Downloads last month
- -