Vayu-ML (94M Parameters)
Vayu-ML is a compact, custom decoder-only Transformer pre-trained from scratch on 500M tokens and aligned via Supervised Fine-Tuning (SFT) on ML/DL theoretical concepts.
Model Architecture
- Parameters: ~94 Million
- Layers: 12
- Hidden Dimension: 768
- Intermediate Dimension (SwiGLU): 2048
- Attention Heads: 12
- Normalization: RMSNorm
- Vocabulary Size: 32,768 (Custom BPE Tokenizer)
- Context Length: 512 tokens
Training Details
- Pretraining: ~500M tokens in native
bfloat16using AdamW. - Alignment: SFT on curated ML/DL conceptual Q&A with causal response masking.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support