SupraGDN-5M • GatedDeltaNet • Tiny-SOTA
We are introducing SupraGDN-5M, a GatedDeltaNet-architecture-based model with 5 million parameters, pretrained from scratch as a base model on a single RTX Pro 4500 SE on 5B Fineweb-Edu tokens in about ~2.5 hours.
This is the base work for future models, as the upcoming Supra3-family with our most capable tiny SOTA SLM models.
Please note, that this is an undertrained, experimental base model with no high capabilities.
Pretraining
The pretraining ran for exactly 1 epoch on the first 5B tokens of Fineweb-Edu sample-10BT on a single RTX Pro 4500 SE rented from Runpod.
- Batch Size: 1024
- Context: 256 tokens
- Pretraining tokens: 5B
- Vocab size of custom tokenizer: 6000 tokens
- GPU: RTX Pro 4500 SE on Runpod
- Total time: ~2.5 hours
- Total cost: ~$2
Final Loss
The Val Loss dropped from ~7.1 to ~3.5251 and a Val-PPL of 33.96. Same goes for the Train Loss.
Samples
PROMPT: The history of
OUTPUT: The history of Native Americans during the Civil War was much worse during the Civil War. Since then, and the culture of North America has changed, the history of Native Americans from the beginning, the story of the revolutionary colonies of America, and the history of America has changed dramatically. The history of North America in 1876 marks the twentieth-morning of America’
--------------------------------------------------------------------------------
PROMPT: In order to understand how
OUTPUT: In order to understand how the problem may interfere, most importantly, in a way that can be applied to other things in a way that is useful to society.
--------------------------------------------------------------------------------
PROMPT: The most important thing about science is
OUTPUT: The most important thing about science is that the first step in the scientific profession is to study more and more science projects on the planet. Some people may believe that science projects are a way of life but others do not realize what is going on in their entire lives.
--------------------------------------------------------------------------------
Benchmarks
| Metric | acc_norm |
|---|---|
| ARC-Easy | 33.59 |
| ARC-Challenge | 23.21 |
| HellaSwag | 26.74 |
| PIQA | 52.88 |
How this model compares to other SOTA mini ~5M parameter models
| Model | Pretraining tokens | ARC-Easy | ARC-Challenge | HellaSwag | PIQA |
|---|---|---|---|---|---|
| Supra-5M-GatedDeltaNet | 5B tokens | 33.59% | 23.21% | 26.74% | 52.88% |
| fromziro/Qana-mini-5M | 21B tokens | 34.97% | 23.21% | 27.60% | 57.18% |
| AxiomicLabs/GPT-S2-5M | 75B tokens | 33.92% | 22.87% | 27.87% | 57.56% |
| User01110/CMA-8M | 21B tokens | 35.35% | 23.29% | 28.19% | 58.22% |
Verdict: We can see clearly, that our GDN model is highly competitive in hard tasks like ARC-Easy, ARC-Challenge and HellaSwag, while in PIQA it's not very good. Future work - to improve.
How to run the model
The full model config can be found in gatedflow_model.py. Have fun.
