Add ResearchClawBench evaluation result

#17

by CoCoOne - opened 2 days ago

base: refs/heads/main

←

from: refs/pr/17

Discussion Files changed

-0

CoCoOne

2 days ago

•

edited 2 days ago

Hi Xiaomi MiMo team,

This PR adds the ResearchClawBench overall evaluation result for MiMo-V2.5.

ResearchClawBench is an end-to-end scientific research benchmark for evaluating AI agents and LLMs on tasks that require reading task data and related work, writing and executing code, producing figures, and generating publication-style reports. Final reports are scored against expert checklists derived from human-authored target papers.

The run was executed with ResearchHarness, using tools enabled, code execution, and a file-system workspace. The submitted value is the overall mean score out of 100 over completed ResearchClawBench tasks:

Model: MiMo-V2.5
Score: 16.91 / 100
Completed tasks: 39/40
Run date: 2026-05-15
Benchmark task id: overall

The detailed leaderboard is available here: https://internscience.github.io/ResearchClawBench-Home/

Thank you!

Add ResearchClawBench evaluation result905cf418

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Ready to merge

This branch is ready to get merged automatically.

· Sign up or log in to comment