SimpleMemVLA

SimpleMemVLA is a vision-language-action (VLA) model for long-horizon robotic manipulation. It does not rely on a dedicated memory module; instead, it keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process. The model is introduced in the paper SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models.

Downloads last month
2
Safetensors
Model size
6B params
Tensor type
BF16
·
Video Preview
loading

Collection including yinchenghust/simplememvla_robomemarena

Paper for yinchenghust/simplememvla_robomemarena