Instructions to use tamkohlaboratory/act_openarm_headrest_position_only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use tamkohlaboratory/act_openarm_headrest_position_only with LeRobot:
- Notebooks
- Google Colab
- Kaggle
ACT β OpenArm headrest, position-only baseline
This is an Action Chunking Transformer (ACT) imitation-learning policy trained with LeRobot. It predicts short chunks of future joint-position commands from two camera views and the current joint positions.
This repository contains the position-only baseline. It is separate from the position+velocity+effort model at act_openarm_headrest_torque.
Model inputs and outputs
The policy uses the standard LeRobot ACT state interface:
| Input/output | Shape | Description |
|---|---|---|
observation.state |
(16,) |
Joint positions in radians; position only |
observation.images.front |
RGB | Front camera image, 480Γ640 |
observation.images.top |
RGB | Top camera image, 480Γ640 |
action |
(16,) |
Target joint positions in radians |
The 16-joint order is: right arm joint_1..joint_7 plus right gripper, followed by left arm joint_1..joint_7 plus left gripper.
Important: observation.state must contain only the 16 joint positions. Do not concatenate velocity or effort/torque. The policy does not expect a 48-dimensional state.
ACT predicts a chunk of 100 future position targets (chunk_size=100, n_action_steps=100). The bundled preprocessor normalizes the state and images, while the bundled postprocessor converts predicted actions back to radians.
Training details
- Dataset:
tamkohlaboratory/openarm_session_20260924_01 - Task: Pull out the headrest
- Dataset size: 30 episodes, 15,667 frames, 30 FPS
- Training steps: 100,000
- Batch size: 8
- Input state: 16-dimensional joint position only
- Action: 16-dimensional target joint position
- Cameras: front and top
- Vision backbone: ResNet-18 with ImageNet-pretrained weights
- ACT model: CVAE enabled,
kl_weight=10 - Optimizer: AdamW, learning rate
1e-5, weight decay1e-4 - Normalization: mean/std normalization for state, images, and actions
- Hardware: NVIDIA RTX 3080 with CUDA and AMP enabled
- LeRobot version: 0.4.4
The final checkpoint was saved after step 100,000. The final training log reported loss=0.0335, l1=0.0274, and kld=0.0006 at step 100,000.
Loading and inference
import torch
from lerobot.policies.act.modeling_act import ACTPolicy
from lerobot.policies.factory import make_pre_post_processors
repo = "tamkohlaboratory/act_openarm_headrest_position_only"
policy = ACTPolicy.from_pretrained(repo)
preprocessor, postprocessor = make_pre_post_processors(
policy.config, pretrained_path=repo
)
# position is (..., 16), in the joint order described above
batch = {
"observation.state": position,
"observation.images.front": front_rgb,
"observation.images.top": top_rgb,
}
policy.reset()
obs = preprocessor(batch)
action = postprocessor(policy.select_action(obs))
# action: 16-dim target joint positions in radians
Normalization statistics are included in the processor files, so dataset access is not required for inference.
Files included
model.safetensorsβ ACT policy weightsconfig.jsonβ model architecture and feature shapespolicy_preprocessor.jsonplus state β input normalization pipelinepolicy_postprocessor.jsonplus state β action unnormalization pipelinetrain_metadata.jsonβ dataset and training provenance
- Downloads last month
- -