The repo for the first in the Biggerbrain 3 lineup, a 151m parameter model.

This model uses a Recurrent transformer(looping the central 5 layers), and soft averaged MoE architecture to achieve new levels of reasoning(reletive to the Biggerbrain lineup). This model used an improved training pipeline*, and a substantial parameter increase over V2, and a great boost to intelligence, focus, and overall usefullness.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Skull18500/Biggerbrain3_151m