ACT · SO-101 · 16 teleop + 16 handheld demonstrations

An ACT policy for the SO-ARM101 arm, trained on 16 teleoperated episodes plus 16 handheld demonstrations — recorded with a GoPro on a hand-held gripper mock-up, never touching the robot, then retargeted into the arm's joint space.

This is the arm the comparison is about: it is the only one of the three whose success rate was measured on this exact checkpoint.

The comparison

Three ACT policies share a training recipe exactly — 50,000 steps, batch 8, seed 1000, one wrist camera at 10 fps — and differ only in what they were trained on. Task: pick a small box from a 20-square grid and place it on a target.

training data checkpoint frames held-out action error measured success
16 teleop act_so101_t16b 5,600 3.07° 17% (2/12) †
16 teleop + 16 handheld act_so101_t16b_u16 9,376 2.82° 65% (13/20)
32 teleop act_so101_t16b_t16 11,200 2.20° 70% (14/20) †

† Measured on a separately trained arm of the same teleop episode count, not on this exact checkpoint. Only the middle row was scored on the checkpoint published here. A 12-attempt session on the 32-teleop condition scored 75%; both are on the record. Rates come from the operator's raw per-attempt records and are re-derived by verify_results.py.

Doubling teleoperation moved success from 17% to ~70%. Replacing that second half with handheld demonstrations reached 65% — about nine tenths of the gain, on 16% fewer training frames, from data collected with a $0 gripper mock-up and a GoPro instead of a second robot session.

A later 10-placement session on a rebuilt scene put the handheld arm at 60% (6/10), consistent with the 65%.

Do not rank these policies by action error

Held-out action error is in the table because it is cheap and reproducible, and because it is wrong here. Calibrated on the teleop pair, the 0.25° gap for the handheld arm predicts roughly 33% success. It measured 65%. A policy can track a demonstrator's joint trajectory closely and still fail to close the gripper, and the converse. Rank on the bench, not on this number.

Use it

pip install lerobot==0.6.1
from lerobot.policies.act.modeling_act import ACTPolicy

policy = ACTPolicy.from_pretrained("robotfuel/act_so101_t16b_u16")
policy.eval()

# observation.state: (6,) joint positions, degrees
# observation.images.wrist: (3, 480, 640) uint8, wrist camera
action = policy.select_action(observation)   # (6,) target joint positions

The policy expects a single wrist camera and 6 joint positions, and emits absolute joint targets at 10 Hz. It was evaluated with n_action_steps=30.

Training

architecture ACT (LeRobot implementation)
dataset robotfuel/so101_t16b_u16
steps 50,000
batch size 8
seed 1000
camera wrist only, 640×480, 10 fps
action space 6 absolute joint positions

Every checkpoint in this set uses these settings unchanged. The only variable is the dataset.

Evaluation protocol

Attempts are run on a printed 20-square grid. The box is placed by hand at a prescribed coordinate, the arm is homed and held under torque, and the episode runs 35 s to completion without intervention. Scene brightness is measured before each attempt and attempts outside the trained range are recorded as skipped rather than failed, so they leave the denominator instead of deflating the rate.

An attempt counts as a success if the box ends on the target, including a scrappy recovery. Attempts that were skipped, voided by a rig fault, or interrupted are not scored.

Limitations

One task, one arm, one operator, one training seed per condition. Sample sizes are 12–20 attempts, which separates 17% from 65% comfortably but does not separate 65% from 70%. The policy is sensitive to lighting outside the trained range and to box placements at the far corners of the grid. Nothing here has been tested on a second robot, a second task, or a second scene.

Downloads last month
26
Safetensors
Model size
51.6M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train robotfuel/act_so101_t16b_u16