Navdyut-60M
This is the 60M parameter model from the Navdyut Foundational Suite, trained by Dicom Pathak at Navdyut AI Labs.
Model Architecture
- Hidden size: 768
- Intermediate size: 3,072
- Number of attention heads: 12
- Number of hidden layers: 8
- Number of key-value heads: 4
- Maximum position embeddings: 256
- Activation function: SwiGLU
- Positional embeddings: Rotary (RoPE) with theta=10,000
- Training: Grouped-query attention and bfloat16 mixed-precision.
Training
This model was trained from scratch using Maximal Update Parametrization ($\mu$P) and Square Root Batch Sizing. It was trained exclusively on a synthetic, highly-curated dataset of 14.4 Billion tokens containing Cosmopedia and CodeSearchNet traces. It is optimized for deterministic structural logic and Python syntax execution.
You can view the full training logs and metrics for this model run here: Weights & Biases Training Log
Note: Proprietary hyperparameter configurations (such as peak learning rate scalars) have been redacted.
- Downloads last month
- 2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support