so101_chess_molmoact2_dagger

The so101_chess_molmoact2 policy, continued on corrections recorded from its own mistakes. Told pick up the piece on e2 and place it on e4 or take the piece on d5 off the board, it carries the move out on an SO-101 arm from two camera images and six joint angles. Same inputs, same outputs, same usage as the model it starts from; what changed is the data.

What was fixed. The base model predicts 30 actions per call and, as shipped, executes all 30 - three seconds of motion decided before it starts. Captures, the long trajectory out to the tray, drifted over that stretch and failed a third of the time. More clean demonstrations did not help (two attempts, both parity). What did: letting the policy play, handing over to the scripted expert the moment a grasp went wrong, and recording only the expert's way out. A thousand of those corrections, weighted three times in the training set, took captures at the shipped horizon from 68 to 78 of 100 on the same positions, and cost nothing anywhere else.

Results

Scored on the same 300 positions as the base model, paired, 450 control steps each, with the board shifted up to 10 mm and the arm starting up to 0.1 rad off its parked pose:

so101_chess_molmoact2 this model
as shipped, 30 actions per call - moves (200) 159 (80%) 155 (78%)
as shipped, 30 actions per call - captures (100) 68 (68%) 78 (78%)
5 actions per call - moves (200) 156 (78%) 157 (79%)
5 actions per call - captures (100) 85 (85%) 85 (85%)

Overall at the shipped setting: 233 against 227 of 300. On captures the corrected model won 21 positions the base lost and lost 11 the other way (paired McNemar p = 0.11); the failure the corrections were collected for - the piece never lifted - fell from 24 to 13. Moves are within noise (p = 0.69). At five actions per call the two are the same model, which is what a fix for open-loop drift should look like: it matters where the loop is open and nowhere else.

The five-action setting is still the better one in simulation, where re-planning is free:

policy.config.n_action_steps = 5   # then policy.reset(): the action queue is rebuilt from it

Use this model

Everything on the base model's card applies with this Hub id in place of that one - installation, the one-command demo, the Python and raw-LeRobot calls, fine-tuning and the real-arm command. The short form:

pip install "lerobot[molmoact2] @ git+https://github.com/huggingface/lerobot"
git clone https://github.com/qalby-tech/embryo_lab_chess_so101 && pip install -e embryo_lab_chess_so101
cd embryo_lab_chess_so101
MUJOCO_GL=egl python examples/evaluate_policy.py --checkpoint XvKuoMing/so101_chess_molmoact2_dagger \
    --moves 4 --captures 2 --video so101_chess_dagger.mp4 --video-cameras top wrist
from chess_sim import LeRobotPolicy, LeRobotPolicyConfig
policy = LeRobotPolicy.load(LeRobotPolicyConfig(checkpoint="XvKuoMing/so101_chess_molmoact2_dagger"))
key shape what
task string the instruction, in the two forms above
observation.state (6,) joint angles, SO-101 degrees, shoulder pan first, gripper last
observation.images.top (3, 480, 640) overhead camera, RGB in [0, 1]
observation.images.wrist (3, 480, 640) wrist camera, RGB in [0, 1]
action (6,) joint targets, SO-101 degrees, one every 0.1 s

Training

Continued from so101_chess_molmoact2 (itself MolmoAct2 with LoRA rank 64, 70,000 steps on 6,095 demonstrations) for 15,000 steps at a tenth of every learning rate, batch 8, bf16 with gradient checkpointing, on one RTX 5090. The rate matters: restarting a converged policy at its full peak rate made it drop pieces; a tenth left it undamaged.

The training set, 11,083 episodes and 941,592 frames at 10 Hz:

part episodes what
the published demonstrations 6,095 XvKuoMing/so101_chess
more capture demonstrations 1,982 same expert, captures only; on their own they changed nothing
expert corrections, counted three times 3 × 1,000 see below
earlier corrections 6

How a correction is recorded. The base policy plays a capture at its shipped 30-action horizon. Every 15 control steps a rule checks the board: the wrong piece engaged, a neighbour disturbed, the named piece shoved across its square without being lifted, or not lifted by 40% of the budget. Any of those hands the episode to the scripted expert, which finishes the capture from wherever the arm is, and only the expert's part is recorded. A piece the policy has already knocked over ends the episode with nothing recorded. 6,056 policy episodes gave 1,524 hand-overs and 1,000 corrections the expert completed; 378 episodes ended at a toppled piece. That is recover() in the framework and examples/collect_recoveries.py.

docs/EXPERIMENTS.md in this repository carries the full record, including the two attempts that scored parity (§4.7-4.8), the check that the hand-over fires where captures fail (§4.9) and this run (§4.10).

Limitations

Those of the base model: simulation only, two instruction families, no grasp for a piece on its side, 28 mm squares and one camera geometry baked in. Captures at the shipped horizon are better, not solved: 22 of 100 still fail, 13 of them at the grasp.

Roadmap

  • Reinforcement learning on top of this checkpoint: the policy's own rollouts, scored in simulation and weighted by return, folded back into training.
  • A wider training distribution - many piece sets and board styles - so that the policy works on a board it has not seen.

Licensing follows the base model, allenai/MolmoAct2-SO100_101.

Downloads last month
11
Safetensors
Model size
6B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for XvKuoMing/so101_chess_molmoact2_dagger

Finetuned
(1)
this model

Dataset used to train XvKuoMing/so101_chess_molmoact2_dagger

Collection including XvKuoMing/so101_chess_molmoact2_dagger