Robotics
LeRobot
Safetensors
molmoact2
maniskill

Model Card for molmoact2_mpc_100

MolmoAct2 is an open robotics foundation model from the Allen Institute for AI (Ai2) that maps camera images and language instructions to robot action chunks.

This is a continuous-action finetune of allenai/MolmoAct2-LIBERO on N=100 motion-planning demonstrations per MPC tabletop task (384 episodes after merge), with absolute end-effector pose control (pd_ee_pose) and 378×378 images from camera_center, camera_left, and camera_wrist.

This policy has been trained and pushed to the Hub using LeRobot.

Learn how to train and run it in the LeRobot molmoact2 guide, or browse the full documentation.


Model Details

  • License: apache-2.0
  • Base model: allenai/MolmoAct2-LIBERO
  • Robot type: panda
  • Control mode: absolute end-effector pose (pd_ee_pose)
  • Cameras: camera_center, camera_left, camera_wrist

Inputs & Outputs

The policy consumes these observation features and produces these action features.

Inputs

Feature Type Shape
observation.state STATE (9,)
observation.images.camera_center VISUAL (3, 378, 378)
observation.images.camera_left VISUAL (3, 378, 378)
observation.images.camera_wrist VISUAL (3, 378, 378)

Outputs

Feature Type Shape
action ACTION (7,)

Training Dataset

Env ID Task language Demonstrations
PickCube-v2-wrist lift the red cube 96
PushCube-v2 push the cube to the goal 96
LiftPegUpright-v2 lift the peg upright 96
PullCubeTool-v2 grasp the red L-shaped hook and use it to pull the blue cube closer to the robot 96
Total 384

Training Configuration

Setting Value
Training steps 10000
Batch size 32
Optimizer adamw
Learning rate 1e-05
Seed 1000
LeRobot version 0.6.0

How to Get Started with the Model

lerobot-eval \
  --policy.path=jstm/molmoact2_mpc_100 \
  --policy.device=cuda \
  --policy.inference_action_mode=continuous \
  --trust_remote_code=true \
  --env.type=maniskill \
  --env.control_mode=pd_ee_pose \
  --env.task="PickCube-v2-wrist::lift the red cube" \
  --env.episode_length=100 \
  --env.observation_height=378 \
  --env.observation_width=378 \
  --env.goal_height=0.1 \
  --eval.n_episodes=50 \
  --eval.batch_size=20

Train your own policy

lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --policy.type=molmoact2 \
  --output_dir=outputs/train/<policy_repo_id> \
  --job_name=lerobot_training \
  --policy.device=cuda \
  --policy.repo_id=${HF_USER}/<policy_repo_id> \
  --wandb.enable=true

Citation

If you use this policy, please cite the method linked in the description above, along with LeRobot:

@misc{cadene2024lerobot,
    author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
    title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
    howpublished = "\url{https://github.com/huggingface/lerobot}",
    year = {2024}
}
Downloads last month
11
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Video Preview
loading