Some thoughts

#5
by NikiKrutan - opened

Hey!
I've seen your post in DavidAU's repo about benches.
I fully agree with you since I am interested in agentic performance. Their benches say little about that.
Still the hype is fantastically extreme.
What I believe you should do: run ARC-Challenge and post results if they are better than theirs. Just another one bench.
They specifically tune to that bench, so that is not honest comparison. That's why post only if yours is better.

Orion LLM Labs org
edited 4 days ago

Yes, people generally look for benchmarks like Terminal Bench 2.1, which specifically shows the model's Agentic Performance. Even if the hype is huge and it might even beat Fable 5 on the ARC-Challenge, that says nothing about whether it will outperform Fable 5 in Agentic Performance.

Regarding benchmarking my model, I'm not sure if I should benchmark it on ARC, because nobody really cares about that. There are many other benchmarks that are much better at showing whether a model has a high capacity for organized logical reasoning. If it lacks modern benchmarks used by state-of-the-art models (and instead relies on classic benchmarks), there's no way to confidently tell the public if that model truly has the incredible reasoning performance that Fable demonstrates. This just leaves people more in doubt about whether the model has actually exceeded the 'OpenAI, Claude and Gemini zone of intelligence' or if it doesn't even know how to use tools.

Furthermore, DavidAU never takes his models seriously. He presents a model 'as good as Fable-5', but it comes with an informal, poorly organized README.md with random GIFs.

In contrast, my models feature organized releases, a clean and easy-to-read README.md, demonstrate solid performance, and use the same benchmarks as frontier models.

Also check out this i opened => https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF/discussions/83

Agreed. But somehow people download his models much more. That's why I've proposed to beat him on his bench instead of making him doing right benches (which he refuses from what I've seen). Just for marketing.
Most people don't understand all these benches and stuff. They install LM Studio (or Ollama, or whatever), download most popular gguf and live with it for the next few months. Some go little further: most popular has arc=711, let's check other guys, oh, they don't have arc at all and they don't have millions of downloads, obviously 711 is better. I would like to be wrong about people, but everything points in that direction.

Anyway I clearly understand why you don't do classic benchmarks. And I appreciate a lot your benchmark selection and effort you put in benchmarking (I've tried to do some benches myself and it was real pain). That's why I use GRM: quality is proven. And I see that improvement over original Qwen in my sessions.

BTW your post is deleted.

Orion LLM Labs org

I know! My post was deleted BY David himself! He wants to censor me; here is an article with all the information from the post => https://huggingface.co/blog/DedeProGames/davidau-deleted-my-post-and-banned-me-for-asking-a
Also, I'm now banned from posting any Discussions on his models.
If I could ask you for one thing: could you help me spread the word about this?

Upvoted and left a comment.
I've properly evaluated his quants of 711. Some are good, most are not: https://huggingface.co/NikiKrutan/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GGUF

Now I begin to understand this "market". Everything as usual: louder = better. Sadly.

Just FYI.
I've tested both GRM-2.6-Plus-0628 and 711 in exact same agentic task. I gave them crude draft of design architecture doc with not obvious logic decisions. All in Russian, that adds up complexity as well. The task was to refine architecture decisions, flag potential problems, recommend on forks. All in tight dialogue with the user.

711 failed to find some existing code (wrong tool call), even didn't retried after failure. That was main code that was going to be refactored btw! Finished with recommendations that mostly were just rephrased design doc. Almost no new information. Obvious recommendations that were already clear enough. Something like what you get if you ask Google in their free online chat (exactly what they call Gemini level of intelligence). BUT! That all was in near perfect Russian, much better language than design doc. Beautiful but unusable for any hard design job.

GRM read all relevant existing code, spawned 3 (!!) sub-agents to inspect artifacts produced by old code (1000's of files), came back with thorough analysis and valid recommendations. Made some mistakes in Russian words, even added some Chinese, partly English answer, though. BUT! After some discussion we went to real coding finally. It suits the job well.

Final thoughts: 711 is really good at writing, that's it. If agentic/coding/design decisions are needed it is much worse than original Qwen.
GRM on the other hand inherited tongue-tiedness original Qwen shipped, but obviously improved on intellect and agentic skills.

This is subjective and without hard proof ("take my word on this"), but difference is so striking, that I don't need any benches of 711 anymore to make my choice.

Sign up or log in to comment