Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Paper • 2606.03784 • Published
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation · Project page
ERVLA is a vision-language-action model for robot manipulation. It learns from embodied chain-of-thought that links task understanding to end-effector movements and image-space trajectories. With reasoning dropout, the model can generate actions directly at inference without decoding a chain of thought.
This model is post-trained on VLABench and achieves 53.2% average success rate across its five evaluation tracks.