Robotics
LeRobot
Safetensors
bimanual
openarm
imitation-learning
act

Model Card for act

Action Chunking with Transformers (ACT) is an imitation-learning method that predicts short action chunks instead of single steps. It learns from teleoperated data and often achieves high success rates.

This policy has been trained and pushed to the Hub using LeRobot. See the full documentation at LeRobot Docs.


How to Get Started with the Model

For a complete walkthrough, see the training guide. Below is the short version on how to train and run inference/eval:

Train from scratch

lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --policy.type=act \
  --output_dir=outputs/train/<desired_policy_repo_id> \
  --job_name=lerobot_training \
  --policy.device=cuda \
  --policy.repo_id=${HF_USER}/<desired_policy_repo_id>
  --wandb.enable=true

Writes checkpoints to outputs/train/<desired_policy_repo_id>/checkpoints/.

Evaluate the policy/run inference

lerobot-record \
  --robot.type=so100_follower \
  --dataset.repo_id=<hf_user>/eval_<dataset> \
  --policy.path=<hf_user>/<desired_policy_repo_id> \
  --episodes=10

Prefix the dataset repo with eval_ and supply --policy.path pointing to a local or hub checkpoint.


Model Details

  • License: apache-2.0

Observation / action scheme (important)

This policy follows the LeRobot ACT interface, but the robot-state vector is a concatenation of three proprioceptive signals, because upstream ACT reads only a single observation.state key:

Slice Dims Signal Source key in the dataset
[0:16] 16 joint position (rad) observation.state
[16:32] 16 joint velocity observation.velocity
[32:48] 16 joint effort / torque observation.effort

So observation.state passed to this policy must be 48-dimensional, in that exact order.

Inputs

  • observation.state: (48,) float32 — position | velocity | torque (order above)
  • observation.images.front: RGB image
  • observation.images.top: RGB image

Output

  • action: (16,) float32 — target joint positions (rad). The policy predicts a chunk of chunk_size future position targets.

Joint order (16 DOF, bimanual): right joint_1..joint_7 + right gripper, then left joint_1..joint_7 + left gripper.

Building the 48-dim state at inference

import torch

state = torch.cat([position, velocity, effort], dim=-1)  # (..., 48)
batch = {
    "observation.state": state,
    "observation.images.front": front_rgb,
    "observation.images.top": top_rgb,
}

Loading this policy

from lerobot.policies.act.modeling_act import ACTPolicy
from lerobot.policies.factory import make_pre_post_processors

repo = "tamkohlaboratory/act_openarm_headrest_torque"
policy = ACTPolicy.from_pretrained(repo)
preprocessor, postprocessor = make_pre_post_processors(policy.config, pretrained_path=repo)

policy.reset()
obs = preprocessor(batch)              # normalizes state + images
action = postprocessor(policy.select_action(obs))   # -> joint positions (rad)

Normalization statistics are bundled in the processor files, so no dataset access is needed for inference.

Downloads last month
25
Safetensors
Model size
51.7M params
Tensor type
F32
·
Video Preview
loading

Paper for tamkohlaboratory/act_openarm_headrest_torque