OrangeFantasy
Orae128
ยท
AI & ML interests
None yet
Recent Activity
published a Space 3 days ago
Orae128/Test reacted to Banaxi-Tech's post with ๐ฅ 15 days ago
We're excited to release BananaMind 2 Micro, our smallest model yet.
It fits a compact architecture in only 2.9M parameters achieving the highest parameter efficiency on BananaMind Base Bench against comparable models.
It achieves comparable performance to GPT S2 5M and GPT S 5M at almost half the size while beating CMA 1M Mini.
BananaMind 2 Micro achieved the #1 spot on the Open SLM Leaderboard for the sub 3M category (not added yet but it achieves #1)
For the training we used Muon + the XSA Refresh Gate with a 5e-2 lr for Muon and 4e-3 for the 1D weights.
Its score on our efficiency measure is 0.326 getting the first place with Syn 2.6M on the second place scoring 0.291 and GPT S 5M at 0.235*
Check it out at https://huggingface.co/BananaMind/BananaMind-2-Micro and follow us at:
@vovaRL
@DedeProGames
@Banaxi-Tech
https://huggingface.co/BananaMind
Our new releases aren't stopping ๐ August 13-14 BananaMind 2 Pro reacted to KlondikeDev's post with ๐ฅ 16 days ago
Boris-2 coming soon!
The Models:
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th.
Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th.
Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We hope to end up in the ballpark of https://huggingface.co/AxiomicLabs/GPT-X2.5-135M or https://huggingface.co/BananaMind/BananaMind-2-Pro-Preview
Following this, we will release the Pro, Instruct and Pro-Instruct variants. More info will be coming soon!Organizations
None yet