Update MDPBench evaluation results

#170
No description provided.

Hi Kimi-K3 team,

We evaluated Kimi-K3 on MDPBench, our multilingual document parsing benchmark covering 17 languages with both digital-born and photographed documents.

Kimi-K3 is currently the strongest open-source model on MDPBench:

  • Overall: 83.6
  • Digital: 90.8
  • Photo: 81.2
  • Latin Avg.: 86.2
  • Non-Latin Avg.: 80.7
  • Private: 85.6

In our evaluation, Kimi-K3 outperforms strong closed-source baselines including GPT-5.2, Claude-Sonnet-4.6, and Doubao-2.0-pro across the main aggregate metrics. It also surpasses Gemini-3-pro-preview on Digital documents and several language/script subsets, including IT, KO, and ZH.

Leaderboard:
https://huggingface.co/datasets/Delores-Lin/MDPBench

Dataset:
https://huggingface.co/datasets/Delores-Lin/MDPBench

GitHub:
https://github.com/Yuliang-Liu/MultimodalOCR/tree/main/MDPBench

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment