Instructions to use jstm/molmoact2_mpc_100 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use jstm/molmoact2_mpc_100 with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Model Card for molmoact2_mpc_100
MolmoAct2 is an open robotics foundation model from the Allen Institute for AI (Ai2) that maps camera images and language instructions to robot action chunks.
This is a continuous-action finetune of allenai/MolmoAct2-LIBERO on N=100 motion-planning demonstrations per MPC tabletop task (384 episodes after merge), with absolute end-effector pose control (pd_ee_pose) and 378×378 images from camera_center, camera_left, and camera_wrist.
This policy has been trained and pushed to the Hub using LeRobot.
Learn how to train and run it in the LeRobot molmoact2 guide, or browse the full documentation.
Model Details
- License: apache-2.0
- Base model:
allenai/MolmoAct2-LIBERO - Robot type:
panda - Control mode: absolute end-effector pose (
pd_ee_pose) - Cameras:
camera_center,camera_left,camera_wrist
Inputs & Outputs
The policy consumes these observation features and produces these action features.
Inputs
| Feature | Type | Shape |
|---|---|---|
observation.state |
STATE | (9,) |
observation.images.camera_center |
VISUAL | (3, 378, 378) |
observation.images.camera_left |
VISUAL | (3, 378, 378) |
observation.images.camera_wrist |
VISUAL | (3, 378, 378) |
Outputs
| Feature | Type | Shape |
|---|---|---|
action |
ACTION | (7,) |
Training Dataset
- Repository: jstm/mpc_lerobot_pd_ee_pose_100
- Episodes: 384
- Frames: 97753
- Frame rate: 30 FPS
| Env ID | Task language | Demonstrations |
|---|---|---|
PickCube-v2-wrist |
lift the red cube | 96 |
PushCube-v2 |
push the cube to the goal | 96 |
LiftPegUpright-v2 |
lift the peg upright | 96 |
PullCubeTool-v2 |
grasp the red L-shaped hook and use it to pull the blue cube closer to the robot | 96 |
| Total | 384 |
Training Configuration
| Setting | Value |
|---|---|
| Training steps | 10000 |
| Batch size | 32 |
| Optimizer | adamw |
| Learning rate | 1e-05 |
| Seed | 1000 |
| LeRobot version | 0.6.0 |
How to Get Started with the Model
lerobot-eval \
--policy.path=jstm/molmoact2_mpc_100 \
--policy.device=cuda \
--policy.inference_action_mode=continuous \
--trust_remote_code=true \
--env.type=maniskill \
--env.control_mode=pd_ee_pose \
--env.task="PickCube-v2-wrist::lift the red cube" \
--env.episode_length=100 \
--env.observation_height=378 \
--env.observation_width=378 \
--env.goal_height=0.1 \
--eval.n_episodes=50 \
--eval.batch_size=20
Train your own policy
lerobot-train \
--dataset.repo_id=${HF_USER}/<dataset> \
--policy.type=molmoact2 \
--output_dir=outputs/train/<policy_repo_id> \
--job_name=lerobot_training \
--policy.device=cuda \
--policy.repo_id=${HF_USER}/<policy_repo_id> \
--wandb.enable=true
Citation
If you use this policy, please cite the method linked in the description above, along with LeRobot:
@misc{cadene2024lerobot,
author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
howpublished = "\url{https://github.com/huggingface/lerobot}",
year = {2024}
}
- Downloads last month
- 11