Safetensors
llama

Navdyut-60M

This is the 60M parameter model from the Navdyut Foundational Suite, trained by Dicom Pathak at Navdyut AI Labs.

Model Architecture

  • Hidden size: 768
  • Intermediate size: 3,072
  • Number of attention heads: 12
  • Number of hidden layers: 8
  • Number of key-value heads: 4
  • Maximum position embeddings: 256
  • Activation function: SwiGLU
  • Positional embeddings: Rotary (RoPE) with theta=10,000
  • Training: Grouped-query attention and bfloat16 mixed-precision.

Training

This model was trained from scratch using Maximal Update Parametrization ($\mu$P) and Square Root Batch Sizing. It was trained exclusively on a synthetic, highly-curated dataset of 14.4 Billion tokens containing Cosmopedia and CodeSearchNet traces. It is optimized for deterministic structural logic and Python syntax execution.

You can view the full training logs and metrics for this model run here: Weights & Biases Training Log

Note: Proprietary hyperparameter configurations (such as peak learning rate scalars) have been redacted.

Downloads last month
2
Safetensors
Model size
81.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support