RoboJudge-9B

RoboJudge-9B is a multimodal judge for generated embodied-manipulation videos. It scores two axes on a 1–5 scale:

  • Physical Adherence (PA): agent consistency, scene consistency, and interaction realism.
  • Instruction Alignment (IA): agent match, object correctness, and goal completion.

The model is trained with full-parameter supervised fine-tuning followed by 200 steps of reinforcement learning. The release code, exact prompts, and evaluation scripts are available at https://github.com/SiyuanMaCS/robojudge-iclr2027.

Benchmark results

Metric Pearson r with human ratings
PA 0.645
IA 0.775
Overall 0.719

The exact released predictions reproduce PA 0.644617, IA 0.775028, and pooled Overall 0.719211 with the repository evaluator.

Download and inference

git clone https://github.com/SiyuanMaCS/robojudge-iclr2027.git
cd robojudge-iclr2027
pip install -r requirements.txt

hf download HuggingFriends/RoboJudge-9B \
  --local-dir checkpoints_hf/RoboJudge-9B

CKPT=checkpoints_hf/RoboJudge-9B \
DATA_ROOT=/path/to/mllm-as-embodied-world-judge \
OUT=outputs/robojudge_predictions.jsonl \
bash code/run_inference.sh

A smaller BF16 sharded copy and the exact one-epoch SFT checkpoint are available at https://huggingface.co/HuggingFriends/robojudge-iclr2027-checkpoints.

Checkpoint identity

  • Weight format: safetensors
  • Original training dtype: FP32
  • model.safetensors SHA256: 34e17912b402182ab8588494d2715461c7120761e98f8b055b1adbbee400e103

Data

License

The RoboJudge release code and model weights are provided under Apache-2.0. Use of the associated datasets is additionally subject to the terms of their upstream sources.

Downloads last month
12
Safetensors
Model size
9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HuggingFriends/RoboJudge-9B

Quantizations
1 model