YAM Duster-in-Box β Diffusion Policy (baseline)
A UNet Diffusion Policy trained on
Dimios45/yam_duster_in_box
for a single-arm I2RT YAM robot: "pick up the duster and put the duster in the box."
This is the baseline for the B-spline Policy comparison. It was trained on byte-identical
data with the same backbone, optimizer, schedule, and epoch count as
Dimios45/yam-duster-bspline-dp β
the only difference is the action representation. Use it to measure what the B-spline
parameterization actually buys in success rate and completion time.
Action space
Joint-space targets, not end-effector poses:
[right_joint_1..6 (radians), right_gripper (0 = open, 1 = closed)]
No IK is involved at deploy. The policy predicts a plain (16, 7) trajectory of future joint
targets on a fixed 10 Hz grid and executes the first 8 (n_action_steps: 8) before replanning β
roughly 0.8 s per prediction, versus the ~1.1 s continuous, rescalable curve the B-spline
variant emits.
Observations
| key | shape | notes |
|---|---|---|
top_image |
(3, 128, 128) |
RGB, resized from 640Γ480 (plain squash, no crop) |
wrist_image |
(3, 128, 128) |
RGB, right wrist camera |
joint_pos |
(7,) |
measured joints + gripper |
Two observation steps. Images normalized to [0,1], random crop to 116Γ116 in training.
Two naming constraints in this codebase: the lowdim key must be named
joint_pos(get_normalizeronly accepts keys containingpos/quat/qpos), and the task config must setabs_action: True. At 7Dβ7Dabs_actionis a no-op for the data, but it also selects the action normalizer βFalseinstalls an identity normalizer, and joints 2/3 (up to 2.77 rad) then fall outside the sampler'sclip_samplerange of [-1, 1], silently breaking training.
Files
| file | size | use |
|---|---|---|
deploy_ema.ckpt |
426 MB | Inference. EMA weights only. |
epoch0600_full.ckpt |
1.5 GB | model + ema_model + optimizer, for resuming or fine-tuning. |
Both embed the Hydra config (pickled with dill); diffusion_policy must be importable on load.
Training
| base | B-spline-policy/bspline-policy UNet DP |
| data | 50 episodes, 37,986 frames @ 30 Hz β resampled to 10 Hz (12,677 steps) |
| hardware | 1Γ RTX 4090, 2 h 00 m, ~16 it/s |
| epochs / batch | 601 / 64 |
| optimizer | AdamW, lr 1e-4, cosine, 500 warmup steps, EMA |
| params | 66.9M diffusion + 22.4M vision (ResNet18 Γ2) |
| scheduler | DDIM, 100 train timesteps, 16 inference steps, epsilon prediction |
| horizon / obs / action steps | 16 / 2 / 8 |
| final train_loss | 0.001 |
Measured behavior
Over 200 held-out samples, predicted chunks vs the recorded demos:
| metric | this model | B-spline variant |
|---|---|---|
| open-loop arm error | median 0.48Β°, p90 1.20Β° | median 0.81Β°, p90 1.68Β° |
| gripper predicted range | [β0.012, 1.000] | [β0.31, 1.31] |
| peak arm velocity | 46 Β°/s p95, 110 Β°/s max | 48 Β°/s p95, 112 Β°/s max |
| inference latency | 44 ms (RTX 4090) | 46 ms |
Velocity was sampled at this model's native 10 Hz action rate versus 100 Hz for the B-spline curve, so those two rows are not strictly equivalent β the 10 Hz figure cannot see sub-100 ms spikes.
Because DP predicts the trajectory directly rather than control points, its outputs stay inside the demonstrated range, so it does not need the gripper clamp the B-spline variant requires.
Deployment
Environment setup, camera configuration, arm bring-up, and the five repo patches for joint-space
actions are identical to the
B-spline model card β follow steps 1β4
there. Those patches are applied and verified in the working tree these models were trained
from, but are not in upstream B-spline-policy/bspline-policy; apply them yourself if you
start from upstream. The gripper clamp in decode_action_vector is unnecessary for this model
(see below) but harmless and worth keeping.
They are needed because the upstream rollout path decodes end-effector actions (its iPhone
teleop records EE poses), while this model predicts joints β a difference in what was
recorded, not in the LeRobot file format. Two things carry over unchanged: a missing top_image
does not raise (policy_local_bspline.py:625-629 feeds a black frame, so a dead camera yields a
half-blind policy rather than an error), and image resolution needs no change because the rollout
resizes from the checkpoint's shape_meta. The top camera is an Intel RealSense; use
TOP_CAMERA_TYPE = 'realsense' so the RGB pipeline matches training.
The gripper-convention check in the B-spline card applies here too β verify it before the first rollout, or the policy will open to grasp and close to release.
hf download Dimios45/yam-duster-dp deploy_ema.ckpt --local-dir ./ckpt
sudo ip link set can_follower_r up type can bitrate 1000000
python real_env/yam_teleop/yam_server.py --channel can_follower_r
python real_env/yam_teleop/rollout_local_policy.py \
--env yam --policy dp \
--ckpt-path ./ckpt/deploy_ema.ckpt \
--diffusion-policy-dir $PWD/diffusion_policy \
--control-freq 100 --data-freq 10 \
--save --output-dir data/local_policy_rollouts
Note --policy dp (not bspline), and that --origin-time-scale / --predict-before-end /
--speed-up-times do not apply β this model emits a fixed-rate action grid with no temporal
rescaling. That is precisely the capability the B-spline variant adds.
Reproducing
python tools/lerobot_v3_to_robomimic.py \
--repo-id Dimios45/yam_duster_in_box \
--output diffusion_policy/data/yam_duster_in_box.hdf5 \
--target-fps 10 --image-size 128
cd diffusion_policy && python train.py \
--config-name=yam_duster_unet_dp \
hydra.run.dir=../outputs/yam_duster_dp_full \
training.resume=false logging.mode=offline \
checkpoint.topk.k=601 dataloader.persistent_workers=True
Citation
@inproceedings{chi2023diffusionpolicy,
title={Diffusion Policy: Visuomotor Policy Learning via Action Diffusion},
author={Chi, Cheng and Feng, Siyuan and Du, Yilun and Xu, Zhenjia and
Cousineau, Eric and Burchfiel, Benjamin and Song, Shuran},
booktitle={Proceedings of Robotics: Science and Systems (RSS)},
year={2023}
}