Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Paper • 2606.03784 • Published
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation · Project page
ERVLA is a vision-language-action model for robot manipulation. It learns from embodied chain-of-thought that links task understanding to end-effector movements and image-space trajectories. With reasoning dropout, the model can generate actions directly at inference without decoding a chain of thought.
This is the pretrained ERVLA model. Our embodied reasoning corpus contains 978,743 trajectories, 226.3 million samples, and 2,592.5 hours of robot data.