YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ACT Policy for SO-101 Robotic Manipulation

Vision-based imitation learning policy for autonomous object pick-and-place with the SO-101 robotic arm.

This repository provides a trained Action Chunking with Transformers (ACT) policy for the SO-101 robotic manipulation platform. The policy is trained entirely from human teleoperation demonstrations collected with the LeRobot framework. Given visual observations and robot proprioceptive states, the policy predicts a sequence of future robot actions and executes the learned manipulation behavior autonomously.


πŸ€– Overview

The goal of this project is to learn a manipulation policy directly from human demonstrations, without explicitly programming the robot trajectory.

Task

The trained policy learns the following behavior:

Detect β†’ Approach β†’ Grasp β†’ Transport β†’ Place

Specifically, the robot is trained to:

  1. Visually identify a target square object.
  2. Move the end-effector toward the object.
  3. Grasp the object.
  4. Transport it to the target region.
  5. Release the object at the designated location.

The policy demonstrates that manipulation behaviors can be learned directly from relatively small-scale real-world demonstrations.


🧠 Method

The policy is based on Action Chunking with Transformers (ACT). Instead of predicting only the next robot action, ACT predicts a chunk of future actions, allowing the robot to execute smoother and more temporally consistent motion. The overall learning pipeline is:

Human Teleoperation
        β”‚
        β–Ό
SO-101 Demonstrations
        β”‚
        β–Ό
     Dataset
        β”‚
        β–Ό
Vision + Robot State
        β”‚
        β–Ό
ACT Policy
        β”‚
        β–Ό
Action Chunk Prediction
        β”‚
        β–Ό
SO-101 Robot

The model learns the mapping:

Observation_t
      ↓
[Camera Images + Robot State]
      ↓
ACT Transformer
      ↓
Action Chunk
      ↓
Robot Execution

πŸ“Š Dataset

The training dataset was collected directly from a physical SO-101 robotic system using human teleoperation.

Property Value
Robot SO-101
Data collection Human teleoperation
Framework LeRobot
Demonstrations 98 episodes
Environment Real-world
Task Square object pick-and-place
Supervision Demonstration trajectories

The demonstrations contain synchronized robot states, actions, and visual observations collected during teleoperation.


πŸ‹οΈ Training

The model was trained using the ACT imitation-learning pipeline provided by LeRobot.

Training paradigm

Demonstration Dataset
        β”‚
        β”œβ”€β”€ Images
        β”œβ”€β”€ Robot State
        └── Actions
              β”‚
              β–Ό
        ACT Training
              β”‚
              β–Ό
      Learned Policy

The training objective is to reproduce the demonstrated manipulation behavior while learning a temporally coherent sequence of actions.

Key components

  • Policy: ACT
  • Learning paradigm: Behavior Cloning / Imitation Learning
  • Backbone: Transformer-based policy
  • Action prediction: Chunked future actions
  • Data source: Real-world teleoperation
  • Robot platform: SO-101
  • Training framework: LeRobot

πŸŽ₯ Observation Space

The policy receives multimodal robot observations consisting of visual information and robot proprioception.

Visual observations

Camera observations provide information about:

  • Target object position
  • Object appearance
  • Workspace geometry
  • Relative position between the robot and target

Robot state

Robot proprioceptive information provides the current configuration of the manipulator.

Conceptually:

Camera Observation
        +
Robot Proprioception
        β”‚
        β–Ό
     ACT Policy
        β”‚
        β–Ό
 Future Action Sequence

πŸ“ˆ Results

The trained policy was evaluated on the physical SO-101 robot.

Qualitative Result

The learned policy successfully demonstrated the ability to:

  • Recognize the target square object
  • Reach toward the object
  • Execute grasping behavior
  • Transport the object
  • Place the object into the designated region

This demonstrates that the policy can acquire a meaningful manipulation behavior from a relatively small number of human demonstrations.

Demonstration-to-Execution

Human Demonstration
        ↓
98 Teleoperated Episodes
        ↓
ACT Imitation Learning
        ↓
Learned Manipulation Policy
        ↓
Autonomous Robot Execution
        ↓
Square Object Pick-and-Place βœ“

πŸ”¬ Why ACT?

Traditional behavior cloning typically predicts an individual action from the current observation.

ACT instead predicts a sequence of future actions, which provides several advantages for manipulation:

  • More temporally consistent motion
  • Reduced action jitter
  • Better modeling of short-term motion trajectories
  • More natural execution of multi-step manipulation behaviors

For contact-rich robotic manipulation, predicting a short horizon of future actions can be particularly useful because individual actions are highly correlated over time.


βš™οΈ Hardware

Robot

SO-101

The system consists of a physical SO-101 robotic manipulation platform with a leader-follower teleoperation setup.

Computing

The model can be trained and deployed using a GPU-enabled environment supported by PyTorch and LeRobot.

Software

  • Python
  • PyTorch
  • Hugging Face Hub
  • LeRobot
  • ACT

πŸš€ Quick Start

1. Install LeRobot

pip install lerobot

2. Download the model

hf download <USERNAME>/<MODEL_NAME>

or clone the repository:

git clone https://huggingface.co/<USERNAME>/<MODEL_NAME>

3. Configure the SO-101

Connect the follower arm and camera, then verify the corresponding device ports.

4. Run the policy

Load the trained ACT policy through the LeRobot inference pipeline.


Downloads last month
44
Safetensors
Model size
51.7M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support