Amazing Work!

#7
by Quazim0t0 - opened

Wow! This is exactly what I have been waiting to see from the SLM community!

It is what Richard Sutton was pointing at when he criticized the field. He said to stop treating every generation as a bigger training bill. Beat the training it used to take and get to the same point, or a better one, with less of it. Amazing job!

Bench Labs org

Really appreciate this. Efficiency was one of the main things we cared about with Cagliostro-v3, so seeing someone recognize that means a lot. We wanted to see how far we could push a small model without just relying on more tokens or more compute every generation. Still plenty to improve, but this is exactly the direction we want to keep exploring.

Sign up or log in to comment