Why? What? TL;DR?
Various tests on, well, various LLM.
Available Tests
LLM Drawing Test (2026 ongoing)
This test is meant to evaluate models in a difficult task requiring a competency in spatial awareness, image recognition, tools calls, and creativity. In a new empty chat, the model is asked to copy the provided image to the best of its abilities. Model is then evaluated on accuracy of function calls (fail rate) and the resemblance to the original drawing.
DoggoEval (Done)
The goal of this test, featuring a dog (Rex) and his owner (EsKa), is to determine if a model is good at obeying a system prompt and character card. The trick being that dogs can't talk, but LLM love to.
- Results and discussions are hosted in this thread (old thread here)
- Files, cards and settings can be found here
- TODO: Charts and screenshots
Limitations
I'm testing for things I'm interested in. I do not pretend any of this is very scientific or accurate: as much as I try to reduce the amount of variables, a small LLM is still a small LLM at the end of the day. The results for other seeds, or with the smallest of change, are bound to give very different results.
I usually give the different models I'm testing a fair shake in a more casual settings. I regen tons of outputs with random seeds, and while there are (large) variations, it tends to even out to the results shown in testing. Otherwise I'll make a note of it.