SimpleMemVLA
Collection
A Simple but Effective Native-Video Memory for Vision-Language-Action Models • 11 items • Updated • 1
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Paper · Project Page · Code
SimpleMemVLA is a vision-language-action (VLA) model that keeps the sampled history intact and feeds it to the backbone in a timestamped video format, using self-attention as the only memory interface. It achieves state-of-the-art results on four memory benchmarks and matches the best general-purpose control performance on LIBERO.
Checkpoints and datasets are released for five benchmarks (RMBench, RoboMME, MIKASA-Robo, RoboMemArena, LIBERO). See the GitHub repository for training, evaluation, and deployment instructions.