Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Paper • 2606.03784 • Published
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation · Project page
ERVLA is a vision-language-action model for robot manipulation. It learns from embodied chain-of-thought that links task understanding to end-effector movements and image-space trajectories. With reasoning dropout, the model can generate actions directly at inference without decoding a chain of thought.
This model is post-trained on LIBERO and achieves 86.9% overall success rate on LIBERO-Plus through zero-shot transfer.