Arya-3B-MoE: Next-Gen Hybrid Architecture
Arya-3B is a research foundation model combining:
- 6:1:1 Hybrid Sequence Mixer: Gated DeltaNet (O(1) memory recurrence) + Sparse Window Attention + Global Multi-Head Latent Attention (MLA).
- Fine-Grained MoE: 16 Routed Experts (Top-4 Active) + 1 Isolated Shared Expert.
- Multi-Token Prediction (MTP): Built-in 2-token speculative decoding head.
- Anti-NaN Numerical Stability: RMSNorm QK-Norm across all attention and recurrent mixer heads.
- Distilled from Frontier Models: DeepSeek-R1 / Qwen reasoning traces.
Organization
Trained by Chanakya Labs / Aniket Jha.
- Downloads last month
- 127
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support