SimpleMemVLA

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Paper · Project Page · Code

SimpleMemVLA is a vision-language-action (VLA) model that keeps the sampled history intact and feeds it to the backbone in a timestamped video format, using self-attention as the only memory interface. It achieves state-of-the-art results on four memory benchmarks and matches the best general-purpose control performance on LIBERO.

Checkpoints and datasets are released for five benchmarks (RMBench, RoboMME, MIKASA-Robo, RoboMemArena, LIBERO). See the GitHub repository for training, evaluation, and deployment instructions.

Downloads last month
3
Safetensors
Model size
6B params
Tensor type
BF16
·
Video Preview
loading

Collection including yinchenghust/simplememvla_robomme

Paper for yinchenghust/simplememvla_robomme