Robotics
LeRobot
Safetensors
smolvla
so101

smolvla_pick_pen_v2_lr5e5

SmolVLA (450M) fine-tuned from lerobot/smolvla_base on an SO-101 teleop dataset for language-conditioned pen selection: "Pick up the blue pen" / "Pick up the pink pen" / "Pick up the grey pen" (all three pens present in every episode).

Training

Base model lerobot/smolvla_base
Dataset bklassen3434/pick_pen_v2_20260920_124400
Steps 10,000 (batch size 64, ~24 epochs)
Learning rate 5e-5, cosine decay over 10,000 steps
Vision encoder unfrozen (freeze_vision_encoder=false, train_expert_only=false)
Hardware 1x A100-80GB (Modal), 2h41m
Final loss 0.022

Cameras are recorded as observation.images.top / observation.images.wrist and remapped onto SmolVLA's baked-in camera keys at train time:

--rename_map='{"observation.images.top": "observation.images.camera1",
               "observation.images.wrist": "observation.images.camera2"}'

The same rename map must be passed at eval time (lerobot-rollout / lerobot-record).

Downloads last month
10
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for bklassen3434/smolvla_pick_pen_v2_lr5e5

Finetuned
(7863)
this model

Dataset used to train bklassen3434/smolvla_pick_pen_v2_lr5e5