ERVLA Pretrained

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation · Project page

ERVLA is a vision-language-action model for robot manipulation. It learns from embodied chain-of-thought that links task understanding to end-effector movements and image-space trajectories. With reasoning dropout, the model can generate actions directly at inference without decoding a chain of thought.

This is the pretrained ERVLA model. Our embodied reasoning corpus contains 978,743 trajectories, 226.3 million samples, and 2,592.5 hours of robot data.

Downloads last month
17
Safetensors
Model size
5B params
Tensor type
BF16
·
Video Preview
loading

Paper for ERVLA/pretrain