Add .eval_results for benchmarks with public HF datasets

#8
Mind Lab org

Add structured evaluation results for benchmarks with public Hugging Face datasets so scores link to dataset pages and leaderboard buttons.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment