Robotics
LeRobot
Safetensors
smolvla
so101

smolvla_pick_pen_v2_frozen

SmolVLA (450M) fine-tuned from lerobot/smolvla_base on an SO-101 teleop dataset for language-conditioned pen selection: "Pick up the blue pen" / "Pick up the pink pen" / "Pick up the grey pen" (all three pens present in every episode).

This variant keeps SmolVLA's defaults: only the action expert trains, the VLM and vision encoder are frozen, so fine-tuning physically cannot overwrite the pretrained language grounding. Part of a 3-way sweep — see also smolvla_pick_pen_v2_lr1e4 and smolvla_pick_pen_v2_lr5e5, which both unfreeze the vision encoder.

Training

Base model lerobot/smolvla_base
Dataset bklassen3434/pick_pen_v2_20260920_124400
Steps 10,000 (batch size 64, ~24 epochs)
Learning rate 1e-4 (default), cosine decay over 10,000 steps
Trainable action expert only (freeze_vision_encoder=true, train_expert_only=true)
Hardware 1x A100-80GB (Modal), 1h23m, 15.5 GB peak
Final loss 0.043

Cameras are recorded as observation.images.top / observation.images.wrist and remapped onto SmolVLA's baked-in camera keys at train time:

--rename_map='{"observation.images.top": "observation.images.camera1",
               "observation.images.wrist": "observation.images.camera2"}'

The same rename map must be passed at eval time (lerobot-rollout / lerobot-record).

Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for bklassen3434/smolvla_pick_pen_v2_frozen

Finetuned
(7829)
this model

Dataset used to train bklassen3434/smolvla_pick_pen_v2_frozen