Arya-3B-MoE: Next-Gen Hybrid Architecture

Arya-3B is a research foundation model combining:

  • 6:1:1 Hybrid Sequence Mixer: Gated DeltaNet (O(1) memory recurrence) + Sparse Window Attention + Global Multi-Head Latent Attention (MLA).
  • Fine-Grained MoE: 16 Routed Experts (Top-4 Active) + 1 Isolated Shared Expert.
  • Multi-Token Prediction (MTP): Built-in 2-token speculative decoding head.
  • Anti-NaN Numerical Stability: RMSNorm QK-Norm across all attention and recurrent mixer heads.
  • Distilled from Frontier Models: DeepSeek-R1 / Qwen reasoning traces.

Organization

Trained by Chanakya Labs / Aniket Jha.

Downloads last month
127
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support