Bench maxed -- don't be fooled by the table presented in the model card

#84
by pathosethoslogos - opened

They set the thinking default to 'xhigh' in the chat template Jinja file and makes it think for a very long time. This is something not expected from any normal user until 3.6. This is an unexpected change and not communicated.

This makes a lot of intelligence metrics look very good, as they do on the model card and elsewhere. But it doesn't make compute time, number of tokens used, etc. metrics look good, but they don't present those.

Who cares, it's better! The last 12 hours for me have been life changing. Most benchmarks require substantial generalization ability to achieve the numbers. While it is possible to benchmax, I don't believe that's happened here.

This is a fair matchup, after all, Opus 4.6 is also at max reasoning effort.

They had to compare it to opus 4.6 because opus 4.7, 4.8 and 5 would obliterate it.
Also all the benchmarks it beats opus 4.6 on are alibaba benchmarks. It loses on all the external benchmarks

They had to compare it to opus 4.6 because opus 4.7, 4.8 and 5 would obliterate it.
Also all the benchmarks it beats opus 4.6 on are alibaba benchmarks. It loses on all the external benchmarks

So what?

Sign up or log in to comment