SimpleMemVLA

SimpleMemVLA is a Vision-Language-Action (VLA) model for long-horizon robot manipulation with a native-video memory mechanism. It keeps the sampled history intact and processes it as timestamped video through the backbone, using the hidden states of a generated sub-task to condition a flow-matching action head.

Paper: SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Project page: https://huggingface.co/collections/yinchenghust/simplememvla

Code: https://github.com/wadeKeith/SimpleMemVLA

Downloads last month
23
Safetensors
Model size
6B params
Tensor type
BF16
·
Video Preview
loading

Collection including yinchenghust/simplememvla_libero

Paper for yinchenghust/simplememvla_libero