Instructions to use Harrysunshine/so101-pi0-9task with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Harrysunshine/so101-pi0-9task with LeRobot:
- Notebooks
- Google Colab
- Kaggle
SO-ARM101 pick-and-place — π0 full fine-tune (9 tasks)
A π0 flow-matching VLA policy fine-tuned on a 9-task SO-ARM101 tabletop pick-and-place dataset. Language-conditioned: one checkpoint serves all nine tasks, selected by the instruction string you pass at inference.
Trained with LeRobot (lerobot-train), stock recipe,
full-parameter fine-tuning (vision encoder not frozen, action expert not trained in isolation).
The base-model tag may render as
lerobot/pi0_old: upstream renamed that repo andlerobot/pi0now redirects to it. It is the same weights this was fine-tuned from.
Tasks
Nine tabletop tasks. The instruction must be passed verbatim — the strings below are exactly the
ones present in the training data, including the article oddity in a eraser and the missing it
in place in the bin.
| Instruction | Episodes | Frames |
|---|---|---|
Pick up a battery and place in the bin |
200 | 56,295 |
Pick up a can and place in the bin |
250 | 82,792 |
Pick up a cube and place in the bin |
300 | 106,085 |
Pick up a eraser and place in the bin |
300 | 109,899 |
Pick up a golf and place in the bin |
250 | 77,409 |
Pick up a medicine bottle and place in the bin |
300 | 116,342 |
Pick up a plush toy and place in the bin |
200 | 80,811 |
Stack the cube on the can |
200 | 74,425 |
Stack the smaller cube on the larger one |
200 | 80,905 |
| Total | 2,200 | 784,963 |
All nine were merged into a single training set with no re-balancing / no up-sampling — episode counts are already roughly uniform, so the merge is balanced as-is.
Data
- Source:
zhuzhuangtian/so101-pick-place-taskson ModelScope (Apache-2.0), LeRobot v3.0 format. - Robot: SO-ARM101 follower arm, 6 DoF.
- Cameras: two RGB streams,
observation.images.topandobservation.images.wrist, 480×640 each, 30 fps. observation.stateandactionare bothfloat32(6):shoulder_pan,shoulder_lift,elbow_flex,wrist_flex,wrist_roll,gripper(joint positions, absolute targets — not deltas).
Training
| Base | lerobot/pi0 (3.3 B params), full fine-tune |
| Steps | 30,000 (≈ 2.4 epochs at effective batch 64) |
| Batch | 8 per device × 8 devices = 64 effective |
| Precision | bfloat16 |
| Optimizer | AdamW, lr 2.5e-5, wd 0.01, betas (0.9, 0.95) |
| Schedule | cosine decay, 2,000 warmup steps, decay to 2.5e-6 |
| Action chunk | 50 (chunk_size = n_action_steps = 50) |
| Augmentation | LeRobot stock image transforms, strengthened photometric range: brightness/contrast 0.7–1.3, hue ±0.05, saturation 0.5–1.5, sharpness 0.5–1.5, affine ±5° / translate 0.05; up to 5 of 6 sampled per frame |
| Normalization | state/action MEAN_STD from real dataset statistics; visual IDENTITY |
Published checkpoint = step 25,000, not the final step. All 12 saved checkpoints were evaluated offline and step 25,000 was best on both metrics below; the curve is U-shaped (degrades after 25k), so training longer would have been worse.
Offline evaluation
There is no simulator for this rig, so evaluation is offline only, against a copy-the-current-joint-angles baseline:
ratio = MAE(prediction, action) / MAE(current_state, action)— must be **< 1**, i.e. the policy must beat "just hold still". This is a much harder bar than raw correlation: the naive baseline already scores 0.93–0.997 correlation on every joint, socorr > 0.9has no discriminative power on this data and is not used here.delta_corr = corr(prediction − state, action − state)— must be > 0 (direction of motion).delta_std— guards against a policy that collapses to shrinking outputs.
Published checkpoint (200 frames sampled uniformly across the whole dataset):
| Joint | ratio ↓ | delta_corr ↑ |
|---|---|---|
| shoulder_pan | 0.548 | 0.880 |
| shoulder_lift | 0.497 | 0.898 |
| elbow_flex | 0.394 | 0.923 |
| wrist_flex | 0.775 | 0.627 |
| wrist_roll | 0.906 | 0.796 |
| gripper | 0.368 | 0.864 |
| worst | 0.906 | 0.627 |
All six joints beat the baseline. Weakest axes are the two wrist DoF (wrist_roll, wrist_flex)
— those are the most likely to misbehave first on hardware.
Real-robot status
One closed-loop run has been done: the policy drove a physical SO-ARM101 through
Pick up a plush toy and place in the bin with both cameras live and inference served remotely from
a GPU host, visually confirmed correct by a human observer. No quantified success rate exists.
Treat this as "the checkpoint loads, the scale is right, and it moves sensibly", not as a validated
success rate.
Usage
from lerobot.policies.pi0.modeling_pi0 import PI0Policy
policy = PI0Policy.from_pretrained("Harrysunshine/so101-pi0-9task")
# batch: observation.state (6,), observation.images.top (3,480,640),
# observation.images.wrist (3,480,640), task = one instruction string from the table above
action = policy.select_action(batch)
Or as a LeRobot policy server for a low-power robot host:
# on the GPU host
lerobot-policy-server --host=0.0.0.0 --port=8080
# on the robot host
lerobot-robot-client \
--policy_type=pi0 \
--pretrained_name_or_path=Harrysunshine/so101-pi0-9task \
--task="Pick up a plush toy and place in the bin" \
--actions_per_chunk=50
The two *normalizer_processor.safetensors files are required. Without them LeRobot silently
falls back to identity normalization: loss-free, but every action comes out at the wrong scale and
the arm will fly off. Load the repo as a whole, do not cherry-pick model.safetensors.
Only inference assets are published here (weights, config, pre/post-processors, training config). Optimizer/scheduler/RNG state is not included, so this repo cannot be used to resume training.
Limitations
- Rig-specific top camera. The training data comes from a different physical setup than ours: the wrist view matches pixel-for-pixel (same SO-ARM101 follower hardware), but the top camera extrinsics differ — the same joint configuration images the arm at a different position in frame. Dropping this checkpoint onto another SO-ARM101 will not work without camera adaptation or a short fine-tune on data from the target rig.
- Offline metrics do not predict real-world success. Related work on other arms has produced checkpoints with 0.995 offline correlation that still failed to grasp on hardware.
- Absolute joint-position action space. Sending these as deltas, or recording target data with a per-session origin offset baked in, will diverge in closed loop.
- Fixed 480×640 dual-camera observation. Single-camera or differently-cropped inputs are out of distribution.
- Tabletop objects and bin from the training scene only; no generalization claims beyond them.
Related
Harrysunshine/so101-smolvla-9task— SmolVLA on the same merged dataset, same evaluation harness (cost/accuracy comparison point).Harrysunshine/so101-act-cube— ACT on the single cube task.
License
Apache-2.0, following the lerobot/pi0 base model and the source dataset.
中文说明
SO-ARM101 桌面抓放 π0 全参微调权重,9 个任务合并训练,语言条件——一个 checkpoint 覆盖全部九个任务,
靠推理时传入的指令字符串区分。指令必须与上表逐字一致(注意 a eraser 和 place in the bin 里没有 it)。
- 数据:ModelScope
zhuzhuangtian/so101-pick-place-tasks(Apache-2.0,LeRobot v3.0), 2200 集 / 784,963 帧 / 双相机 480×640 / 6 维绝对关节位置。九任务全量合并,不做偏斜上采样。 - 训练:
lerobot-train官方入口,全参微调,30k 步,有效 batch 64,bf16,余弦退火 + 2000 步 warmup。 - 发布的是第 25,000 步,不是末点:12 个存点逐点离线验收,25k 双指标同时最优,曲线 U 形(25k 后退化)。
- 验收判据不用
corr > 0.9:「照抄当前关节角」这个恒等基线的逐关节 corr 就有 0.93–0.997,全部过门、零筛选力。 改用ratio = MAE(预测) / MAE(照抄)必须 < 1,配delta_corr与防坍缩的delta_std。 - 六关节 ratio 全部 < 1;短板是腕部两个自由度(
wrist_roll0.906、wrist_flex0.775)。 - 真机只做过一次闭环目检通过(抓毛绒玩具放进筐,双相机 + 远程 GPU 推理),没有量化成功率。
- 两个
*normalizer_processor.safetensors必须一起加载,缺了会静默退回恒等归一化,动作尺度整体错、机械臂跑飞。 - 边界:这批数据的 top 相机外参与我们自己的臂不同(wrist 视图逐像素同款,top 不同), 直接换机器用不了,需要相机适配或用目标机器的少量数据再微调。
- Downloads last month
- -
Model tree for Harrysunshine/so101-pi0-9task
Base model
lerobot/pi0_old