Robotics
LeRobot
Safetensors
act
openarm
bimanual
imitation-learning

ACT β€” OpenArm headrest, position-only baseline

This is an Action Chunking Transformer (ACT) imitation-learning policy trained with LeRobot. It predicts short chunks of future joint-position commands from two camera views and the current joint positions.

This repository contains the position-only baseline. It is separate from the position+velocity+effort model at act_openarm_headrest_torque.

Model inputs and outputs

The policy uses the standard LeRobot ACT state interface:

Input/output Shape Description
observation.state (16,) Joint positions in radians; position only
observation.images.front RGB Front camera image, 480Γ—640
observation.images.top RGB Top camera image, 480Γ—640
action (16,) Target joint positions in radians

The 16-joint order is: right arm joint_1..joint_7 plus right gripper, followed by left arm joint_1..joint_7 plus left gripper.

Important: observation.state must contain only the 16 joint positions. Do not concatenate velocity or effort/torque. The policy does not expect a 48-dimensional state.

ACT predicts a chunk of 100 future position targets (chunk_size=100, n_action_steps=100). The bundled preprocessor normalizes the state and images, while the bundled postprocessor converts predicted actions back to radians.

Training details

  • Dataset: tamkohlaboratory/openarm_session_20260924_01
  • Task: Pull out the headrest
  • Dataset size: 30 episodes, 15,667 frames, 30 FPS
  • Training steps: 100,000
  • Batch size: 8
  • Input state: 16-dimensional joint position only
  • Action: 16-dimensional target joint position
  • Cameras: front and top
  • Vision backbone: ResNet-18 with ImageNet-pretrained weights
  • ACT model: CVAE enabled, kl_weight=10
  • Optimizer: AdamW, learning rate 1e-5, weight decay 1e-4
  • Normalization: mean/std normalization for state, images, and actions
  • Hardware: NVIDIA RTX 3080 with CUDA and AMP enabled
  • LeRobot version: 0.4.4

The final checkpoint was saved after step 100,000. The final training log reported loss=0.0335, l1=0.0274, and kld=0.0006 at step 100,000.

Loading and inference

import torch
from lerobot.policies.act.modeling_act import ACTPolicy
from lerobot.policies.factory import make_pre_post_processors

repo = "tamkohlaboratory/act_openarm_headrest_position_only"
policy = ACTPolicy.from_pretrained(repo)
preprocessor, postprocessor = make_pre_post_processors(
    policy.config, pretrained_path=repo
)

# position is (..., 16), in the joint order described above
batch = {
    "observation.state": position,
    "observation.images.front": front_rgb,
    "observation.images.top": top_rgb,
}

policy.reset()
obs = preprocessor(batch)
action = postprocessor(policy.select_action(obs))
# action: 16-dim target joint positions in radians

Normalization statistics are included in the processor files, so dataset access is not required for inference.

Files included

  • model.safetensors β€” ACT policy weights
  • config.json β€” model architecture and feature shapes
  • policy_preprocessor.json plus state β€” input normalization pipeline
  • policy_postprocessor.json plus state β€” action unnormalization pipeline
  • train_metadata.json β€” dataset and training provenance
Downloads last month
-
Safetensors
Model size
51.7M params
Tensor type
F32
Β·
Video Preview
loading

Paper for tamkohlaboratory/act_openarm_headrest_position_only