YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
ACT Policy for SO-101 Robotic Manipulation
Vision-based imitation learning policy for autonomous object pick-and-place with the SO-101 robotic arm.
This repository provides a trained Action Chunking with Transformers (ACT) policy for the SO-101 robotic manipulation platform. The policy is trained entirely from human teleoperation demonstrations collected with the LeRobot framework. Given visual observations and robot proprioceptive states, the policy predicts a sequence of future robot actions and executes the learned manipulation behavior autonomously.
π€ Overview
The goal of this project is to learn a manipulation policy directly from human demonstrations, without explicitly programming the robot trajectory.
Task
The trained policy learns the following behavior:
Detect β Approach β Grasp β Transport β Place
Specifically, the robot is trained to:
- Visually identify a target square object.
- Move the end-effector toward the object.
- Grasp the object.
- Transport it to the target region.
- Release the object at the designated location.
The policy demonstrates that manipulation behaviors can be learned directly from relatively small-scale real-world demonstrations.
π§ Method
The policy is based on Action Chunking with Transformers (ACT). Instead of predicting only the next robot action, ACT predicts a chunk of future actions, allowing the robot to execute smoother and more temporally consistent motion. The overall learning pipeline is:
Human Teleoperation
β
βΌ
SO-101 Demonstrations
β
βΌ
Dataset
β
βΌ
Vision + Robot State
β
βΌ
ACT Policy
β
βΌ
Action Chunk Prediction
β
βΌ
SO-101 Robot
The model learns the mapping:
Observation_t
β
[Camera Images + Robot State]
β
ACT Transformer
β
Action Chunk
β
Robot Execution
π Dataset
The training dataset was collected directly from a physical SO-101 robotic system using human teleoperation.
| Property | Value |
|---|---|
| Robot | SO-101 |
| Data collection | Human teleoperation |
| Framework | LeRobot |
| Demonstrations | 98 episodes |
| Environment | Real-world |
| Task | Square object pick-and-place |
| Supervision | Demonstration trajectories |
The demonstrations contain synchronized robot states, actions, and visual observations collected during teleoperation.
ποΈ Training
The model was trained using the ACT imitation-learning pipeline provided by LeRobot.
Training paradigm
Demonstration Dataset
β
βββ Images
βββ Robot State
βββ Actions
β
βΌ
ACT Training
β
βΌ
Learned Policy
The training objective is to reproduce the demonstrated manipulation behavior while learning a temporally coherent sequence of actions.
Key components
- Policy: ACT
- Learning paradigm: Behavior Cloning / Imitation Learning
- Backbone: Transformer-based policy
- Action prediction: Chunked future actions
- Data source: Real-world teleoperation
- Robot platform: SO-101
- Training framework: LeRobot
π₯ Observation Space
The policy receives multimodal robot observations consisting of visual information and robot proprioception.
Visual observations
Camera observations provide information about:
- Target object position
- Object appearance
- Workspace geometry
- Relative position between the robot and target
Robot state
Robot proprioceptive information provides the current configuration of the manipulator.
Conceptually:
Camera Observation
+
Robot Proprioception
β
βΌ
ACT Policy
β
βΌ
Future Action Sequence
π Results
The trained policy was evaluated on the physical SO-101 robot.
Qualitative Result
The learned policy successfully demonstrated the ability to:
- Recognize the target square object
- Reach toward the object
- Execute grasping behavior
- Transport the object
- Place the object into the designated region
This demonstrates that the policy can acquire a meaningful manipulation behavior from a relatively small number of human demonstrations.
Demonstration-to-Execution
Human Demonstration
β
98 Teleoperated Episodes
β
ACT Imitation Learning
β
Learned Manipulation Policy
β
Autonomous Robot Execution
β
Square Object Pick-and-Place β
π¬ Why ACT?
Traditional behavior cloning typically predicts an individual action from the current observation.
ACT instead predicts a sequence of future actions, which provides several advantages for manipulation:
- More temporally consistent motion
- Reduced action jitter
- Better modeling of short-term motion trajectories
- More natural execution of multi-step manipulation behaviors
For contact-rich robotic manipulation, predicting a short horizon of future actions can be particularly useful because individual actions are highly correlated over time.
βοΈ Hardware
Robot
SO-101
The system consists of a physical SO-101 robotic manipulation platform with a leader-follower teleoperation setup.
Computing
The model can be trained and deployed using a GPU-enabled environment supported by PyTorch and LeRobot.
Software
- Python
- PyTorch
- Hugging Face Hub
- LeRobot
- ACT
π Quick Start
1. Install LeRobot
pip install lerobot
2. Download the model
hf download <USERNAME>/<MODEL_NAME>
or clone the repository:
git clone https://huggingface.co/<USERNAME>/<MODEL_NAME>
3. Configure the SO-101
Connect the follower arm and camera, then verify the corresponding device ports.
4. Run the policy
Load the trained ACT policy through the LeRobot inference pipeline.
- Downloads last month
- 44