Paimon
This repository contains the evaluation results for Paimon, a vision-language-action policy evaluated on the VLA-Arena benchmark. This is a results-only submission for the checkpoint at training step 40,000。
The evaluated system uses a two-stage pipeline: a VLM-Agent first generates or decomposes the subtask, and the VLA then executes the resulting subtask in the simulator.
VLA-Arena evaluation
Paimon was evaluated on all 33 VLA-Arena task/level cells using the official-style evaluation setup.
- Episodes: 1,700
- Successes: 1,202
- Episode-weighted success rate: 70.71%
- Cell-mean success rate: 69.82%
- L0 / L1 / L2 cell-mean success rate: 88.00% / 67.27% / 54.18%