MemoryVLA
MemoryVLA checkpoint for real-robot action prediction. This repository is
organized so that additional real-world task checkpoints can be added under
checkpoints/<task-name>/ without replacing the current model.
Latest checkpoint
The current default checkpoint is:
checkpoints/frozen-ema-step-20000.pt
It is the EMA action-diffusion checkpoint from the Frozen+EMA training run. The matching action normalization statistics are stored at:
configs/piper-color-sorting/dataset_statistics.json
The exact runtime configuration for this run is stored at:
configs/piper-color-sorting/config.yaml
configs/piper-color-sorting/config.json
The checkpoint is intended to be loaded with the MemoryVLA codebase and
use_ema=True. The model uses the custom_finetuning normalization key.
For direct loading from the repository root, a server-side LFS alias is also provided at:
checkpoints/frozen-ema-step-20000.pt
Set MEMORYVLA_SKIP_BASE_WEIGHTS=1 for inference. The task checkpoint already
contains the trained LLM, vision backbone, projector, action model, and EMA
weights. The required Llama-2 tokenizer/config metadata is mirrored at:
base_models/llama-2-7b-hf/
It does not need to download the full Llama or vision weight files again.
Point MEMORYVLA_LLAMA2_7B_PATH at that downloaded metadata directory for
offline inference. The metadata is provided for loader compatibility; the
gated Llama model license still applies to the underlying model family.
For future real-task training, the matching CogACT-Large initialization checkpoint is available at:
base_models/CogACT-Large/CogACT-Large.pt
The CogACT base checkpoint is only needed when initializing a new training run. It is not needed when loading the Frozen+EMA task checkpoint for inference.
The original Llama model weights are not duplicated in this repository because they are gated third-party weights. Use the original Hugging Face model with the appropriate access terms and token.
Training configuration
- Task: Piper color sorting
- Data: 21 valid real-world episodes, 14,300 frames
- Observation: front camera only in this dataset; no wrist-camera stream was available in the training data
- Base VLM:
prism-dinosiglip-224px+7b - Vision backbone: frozen
- LLM backbone: frozen
- Last LLM layer: frozen
- Trainable components: action diffusion model and MemoryVLA trainable modules
- Action dimension: 7 (
x, y, z, roll, pitch, yaw, gripper) - Action model:
DiT-L - Future action window: 15
- Memory length: 256
- Retrieval layers: 2
- Fusion:
gate - Consolidation:
tome - Per-device batch size: 2
- Global batch size: 32
- Optimizer learning rate:
2e-5 - Scheduler: linear warmup + cosine decay
- Warmup ratio: 0.03
- EMA: enabled, decay
0.999 - Mixed precision: enabled
- Training strategy: FSDP full shard
- Maximum training steps: 20,000
- Image augmentation: disabled
- Random seed: 42
Open-loop evaluation
On the same 21 training episodes, using frame-by-frame memory-aware inference:
- Normalized overall action RMSE:
0.2124 - Gripper accuracy:
98.64% - Memory reset: at the first frame of every episode
- Cognitive and perception memory banks: capped at 256 entries
These are training-set open-loop results and should not be interpreted as unseen-task generalization.
Loading outline
The exact loader depends on the MemoryVLA code revision. The essential settings for this checkpoint are:
model = load_vla(
model_id_or_path="Haongchen/MemoryVLA",
load_for_training=False,
action_dim=7,
future_action_window_size=15,
action_model_type="DiT-L",
mem_length=256,
retrieval_layers=2,
use_timestep_pe=True,
fusion_type="gate",
consolidate_type="tome",
update_fused=False,
use_ema=True,
)
For real deployment, reset the episode memory before the first observation of
each task and pass the matching dataset_statistics.json when unnormalizing
actions. The gripper action is sign-encoded (-1 / +1) by the current
pipeline.
Repository layout
checkpoints/
frozen-ema-step-20000.pt
configs/
piper-color-sorting/
config.yaml
config.json
dataset_statistics.json
Future tasks should use a separate directory, for example:
checkpoints/
drawer-opening/
peg-insertion/
configs/
piper-color-sorting/
drawer-opening/
peg-insertion/
Keep the checkpoint, action statistics, task name, camera convention, and training configuration together for every task.
Code and reproducibility
The training, inference, real-robot wrapper, preprocessing, and evaluation code are available at:
https://github.com/Laryderr/MemoryVLA
The checkpoint in this repository was produced with code commit:
dd9b4709db9acb48b4e074f71632cb6cc3570662
The Hugging Face repository stores model artifacts and deployment metadata, while GitHub stores source code and reproducibility files. The GitHub repository is not required to download the weights, but it should be kept in sync for future task training and real-robot deployment.
Intended use
Research and development for real-robot manipulation. Validate workspace limits, action scaling, emergency stop behavior, camera calibration, and gripper sign conventions before sending actions to hardware.
- Downloads last month
- -