Instructions to use ukcastle/MolmoAct2-SO101-10L-VTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ukcastle/MolmoAct2-SO101-10L-VTP with Transformers:
# Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("ukcastle/MolmoAct2-SO101-10L-VTP", trust_remote_code=True, device_map="auto") - LeRobot
How to use ukcastle/MolmoAct2-SO101-10L-VTP with LeRobot:
- Notebooks
- Google Colab
- Kaggle
MolmoAct2-SO101 · 10-layer ViT + token pruning
A compressed allenai/MolmoAct2-SO100_101,
fine-tuned on an SO-101 cube pick-and-place dataset and built for on-device deployment
(Qualcomm IQ-9075, 2× Hexagon v73 NSP).
Two compressions are applied on top of the base checkpoint:
| what | effect | |
|---|---|---|
| layer prune | vision tower 25 → 10 blocks | 5.44B → 5.22B parameters |
| token prune | grid token sampler, 16 tokens per crop | LLM prompt 496 → 136 tokens |
Both are healed by training — the tower by embedding distillation, the token sampler by the task fine-tune (it ships with no pretrained weights and is useless without one).
How it was built
Stage 1 — depth prune + embedding distillation. The vision tower keeps blocks
[0,1,2,3,5,6,7,8,9,24], chosen by per-layer SNR (10·log10(E‖in‖² / E‖out−in‖²); a high value
means the block barely changes the signal). Student block j inherits teacher block keep[j].
The connector taps move from [-3,-9] to [-1,-4] so the deep tap stays on the same teacher
layer (24) — which is why 24 is pinned in the keep set.
The pruned tower is then distilled against the teacher's connector output with
relative MSE + (1 − cosine). The MSE is normalised by target energy because these embeddings
carry massive activations (1e2–1e3) and a raw MSE (1e5) would drown the cosine term.
Held-out cosine, measured on an episode-level holdout rather than the training batch:
warm start 0.13441 → step 3000 0.95149 (12 evaluations, monotone)
Stage 2 — task fine-tune. 6000 steps, effective batch 28 (4 × 7 GPUs), LoRA r=64 on the
VLM with a fully trainable action expert, use_grid_token_sampler=true so the sampler is
trained from its random initialisation. LoRA is merged into the base weights in this checkpoint.
Measured results
Open-loop action-chunk prediction on a held-out episode tail: 200 frames, identical frames for
every row (the hold baseline is 5.968 across all four, which is how you can tell). Frames whose
ground-truth chunk barely moves — episode-end padding — are excluded, since "hold still" scores
them perfectly and they say nothing about a policy.
| checkpoint | MAE ↓ | vs hold-still | step 0 | step 29 |
|---|---|---|---|---|
| base teacher, recalibrated, no fine-tune | 10.125 | 1.696 | 4.333 | 14.828 |
| 10L student, distilled only | 10.550 | 1.768 | 4.494 | 15.434 |
| teacher fine-tuned — reference ceiling | 4.737 | 0.794 | 2.813 | 6.200 |
| this model | 5.389 | 0.903 | 2.985 | 7.514 |
Task fine-tuning is what matters: it takes the teacher from 10.125 to 4.737 (−53%). Compression then costs +13.8% on top of that, and the compressed model still beats holding still (0.903 < 1.0).
Error grows along the chunk, so the cost depends on how much of the 30-step chunk you execute before re-inferring:
| steps executed | teacher | this model | gap |
|---|---|---|---|
| 5 of 30 | 3.26 | 3.51 | +7.7% |
| 10 of 30 | 3.68 | 4.01 | +9.1% |
| 30 of 30 | 4.74 | 5.39 | +13.8% |
At 30 fps a 30-step chunk is a 1-second budget and the deployment pipeline runs well inside it,
so a shorter n_action_steps is a real lever rather than a compromise.
Limitations — read these
- Open-loop MAE is a proxy, not a success rate. There is no simulator for this robot, so no closed-loop number exists. The upstream compression study on the LIBERO track explicitly found that its distillation cosine did not predict downstream success (two 0.970 configurations scored worse than a 0.956 one) — treat any single scalar here with the same caution.
- The two compressions are not separated. The +13.8% is layer pruning and token pruning together against a reference that has neither. Which dominates has not been measured.
norm_stats.jsonwas regenerated from the training dataset. The base checkpoint's joint calibration is offset from this dataset by roughly 153° onshoulder_liftand 103° onelbow_flex; feeding SO-101 states through the original statistics saturates the normaliser and makes the model return its ownq01vector regardless of the image. If you use this model on a different SO-101 setup, regenerate the statistics again.- Trained for one task family — 10 cube pick-and-place instructions, 624 episodes, 0.94 epochs. The reference recipe's own data shows fine-tuning had not plateaued at that budget (its 12-layer student went 88.5% → 94.5% between 2k and 5k steps), so this is likely under-trained.
Usage
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
repo = "ukcastle/MolmoAct2-SO101-10L-VTP"
processor = AutoProcessor.from_pretrained(repo, trust_remote_code=True, extra_special_tokens={})
model = AutoModelForImageTextToText.from_pretrained(
repo, trust_remote_code=True, dtype=torch.bfloat16
).to("cuda").eval()
out = model.predict_action(
processor=processor,
images=[top_rgb, wrist_rgb], # PIL or RGB arrays
task="Pick up the large blue cube and put it in the box.",
state=joint_positions, # raw 6-dim, degrees
norm_tag="so100_so101_molmoact2",
inference_action_mode="continuous",
enable_depth_reasoning=False,
num_steps=10,
)
actions = out.actions # [1, 30, 6], robot scale
use_grid_token_sampler is baked into config.json and must stay on. The token sampler is
inlined into this checkpoint's own modeling_molmoact2.py and processing_molmoact2.py, so
trust_remote_code=True is enough and lerobot does not need to be installed. A checkpoint that
carries sampler weights next to base-model remote code instead loads without any error and is
silently wrong: the 8 sampler tensors are dropped as UNEXPECTED and a 32-token-trained model
runs at 392 tokens.
Licence
The base model allenai/MolmoAct2-SO100_101
carries no declared licence, on its model card or as a hub tag, so no licence is asserted here
either. The training data is Apache-2.0. Confirm the upstream terms with AllenAI before
redistributing this checkpoint or using it commercially.
- Downloads last month
- 21
Model tree for ukcastle/MolmoAct2-SO101-10L-VTP
Base model
allenai/MolmoAct2-SO100_101