A very impressive result
Hi, congratulations -- do I get it right that you did this based on an Ornith -10.-35B checkpoint in about a month, give or take?
This is quite an achievement; I am surprised that your model does not get the amount of attention I would expect it to get.
In fact, I was so puzzled of why it does not receive the expected amount of attention, that I ran some very simple perplexity benchmarks for a few recent Qwen 35B - ish models, including yours, just to check if I am missing something -- and the tests very well correlate with your own benchmark reports:
I really wonder how you did -- if you care to share the details on training, personally I'd be very interested.
Thanks for the feedback on our model! Unfortunately, we won't be sharing the training data, but I'm really grateful for how well my model performs in terms of perplexity. Thanks for the benchmarks!