Vayu-ML (94M Parameters)

Vayu-ML is a compact, custom decoder-only Transformer pre-trained from scratch on 500M tokens and aligned via Supervised Fine-Tuning (SFT) on ML/DL theoretical concepts.

Model Architecture

  • Parameters: ~94 Million
  • Layers: 12
  • Hidden Dimension: 768
  • Intermediate Dimension (SwiGLU): 2048
  • Attention Heads: 12
  • Normalization: RMSNorm
  • Vocabulary Size: 32,768 (Custom BPE Tokenizer)
  • Context Length: 512 tokens

Training Details

  • Pretraining: ~500M tokens in native bfloat16 using AdamW.
  • Alignment: SFT on curated ML/DL conceptual Q&A with causal response masking.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support