VegaLM1-42M-Base

Note: AI was used in the creation of this project. But that's why you're here, isn't it?


This is the base model for VegaLM1-42M, an experimental SLM trained on various corpi... corpuses... datasets. It's not Fable 6, but it's good enough, ok?

The first 262M tokens the model saw came from a 90M selection of fineweb-edu The next 688M was from a 50/25/15/10 split of fineweb-edu, fineweb, wikipedia, and a replay respectively, totaling around 420M tokens The final 2B tokens came from a 60/25/15 split of stack-v3-train, dclm-baseline-1.0, and smollm-corpus, trained for 1 epoch because I'm lazy

  • Layers: 12
  • Hidden size: 512
  • Attention heads: 8
  • KV heads: 4
  • Context length: 2,048 tokens
  • Intermediate size: 880
  • Vocabulary: 32,000
  • Parameters: 42,054,144
Downloads last month
14
Safetensors
Model size
42.1M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CNWPlayer/VegaLM1-42M-Base

Finetunes
1 model

Datasets used to train CNWPlayer/VegaLM1-42M-Base