Leaderboard

#5
by Banaxi-Tech - opened

Hi @GODELEV
We have decided this model at this time cannot be added to the BananaMind Base Bench Leaderboard.
The score variations between the Public and the Private set and our other sets are too big.
We may be able to add it in the future but not right now.
But still great model!

I appreciate it
No problem Man , Take your time

GODELEV changed discussion status to closed

@GODELEV We want to inform you, we are currently training our v6 version of our benchmark contamination classifier. This means if it doesent detect anything odd, Rose-Medium is going to be added to the BananaMind Base Bench Leaderboard soon!

its should not detect any odd

GODELEV changed discussion status to open

its should not detect any odd

ik its just that my previous v4 found something was wrong with it so ill run it trough v6, but v6 still currently being built

its should not detect any odd

ik its just that my previous v4 found something was wrong with it so ill run it trough v6, but v6 still currently being built

Why are you using a detector to discern if a model was trained on a benchmark or not? The detector can make mistakes.

I also doubt that @GODELEV would train on benchmarks.

Trust your eyes, not an ML model. If the score is obviously too high for that size of the model, than you know it was trained on benchmarks.

Furthermore, is the private set harder than the public one? These are all things you have to take into consideration when making your own benchmark, eval harness, and leaderboard.

Ik, im training a detector because i want to see how good it is and mainly because models like the Lumen.

Ik, im training a detector because i want to see how good it is and mainly because models like the Lumen.

Well I guess that makes sense. Sorry for being so blunt earlier.

besides that, great work!

Sign up or log in to comment