SupraGDN-5M • GatedDeltaNet • Tiny-SOTA

image

We are introducing SupraGDN-5M, a GatedDeltaNet-architecture-based model with 5 million parameters, pretrained from scratch as a base model on a single RTX Pro 4500 SE on 5B Fineweb-Edu tokens in about ~2.5 hours.
This is the base work for future models, as the upcoming Supra3-family with our most capable tiny SOTA SLM models.
Please note, that this is an undertrained, experimental base model with no high capabilities.

Pretraining

The pretraining ran for exactly 1 epoch on the first 5B tokens of Fineweb-Edu sample-10BT on a single RTX Pro 4500 SE rented from Runpod.

  • Batch Size: 1024
  • Context: 256 tokens
  • Pretraining tokens: 5B
  • Vocab size of custom tokenizer: 6000 tokens
  • GPU: RTX Pro 4500 SE on Runpod
  • Total time: ~2.5 hours
  • Total cost: ~$2

Final Loss

The Val Loss dropped from ~7.1 to ~3.5251 and a Val-PPL of 33.96. Same goes for the Train Loss.

Samples


PROMPT: The history of
OUTPUT: The history of Native Americans during the Civil War was much worse during the Civil War. Since then, and the culture of North America has changed, the history of Native Americans from the beginning, the story of the revolutionary colonies of America, and the history of America has changed dramatically. The history of North America in 1876 marks the twentieth-morning of America’
--------------------------------------------------------------------------------
PROMPT: In order to understand how
OUTPUT: In order to understand how the problem may interfere, most importantly, in a way that can be applied to other things in a way that is useful to society.

--------------------------------------------------------------------------------
PROMPT: The most important thing about science is
OUTPUT: The most important thing about science is that the first step in the scientific profession is to study more and more science projects on the planet. Some people may believe that science projects are a way of life but others do not realize what is going on in their entire lives.

--------------------------------------------------------------------------------

Benchmarks

Metric acc_norm
ARC-Easy 33.59
ARC-Challenge 23.21
HellaSwag 26.74
PIQA 52.88

How this model compares to other SOTA mini ~5M parameter models

Model Pretraining tokens ARC-Easy ARC-Challenge HellaSwag PIQA
Supra-5M-GatedDeltaNet 5B tokens 33.59% 23.21% 26.74% 52.88%
fromziro/Qana-mini-5M 21B tokens 34.97% 23.21% 27.60% 57.18%
AxiomicLabs/GPT-S2-5M 75B tokens 33.92% 22.87% 27.87% 57.56%
User01110/CMA-8M 21B tokens 35.35% 23.29% 28.19% 58.22%

Verdict: We can see clearly, that our GDN model is highly competitive in hard tasks like ARC-Easy, ARC-Challenge and HellaSwag, while in PIQA it's not very good. Future work - to improve.

How to run the model

The full model config can be found in gatedflow_model.py. Have fun.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using SupraLabs/SupraGDN-5M 1

Collection including SupraLabs/SupraGDN-5M