Instructions to use Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full with LeRobot:
# See https://github.com/huggingface/lerobot?tab=readme-ov-file#installation for more details git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e .[smolvla]
# Launch finetuning on your dataset python lerobot/scripts/train.py \ --policy.path=Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full \ --dataset.repo_id=lerobot/svla_so101_pickplace \ --batch_size=64 \ --steps=20000 \ --output_dir=outputs/train/my_smolvla \ --job_name=my_smolvla_training \ --policy.device=cuda \ --wandb.enable=true
# Run the policy using the record function python -m lerobot.record \ --robot.type=so101_follower \ --robot.port=/dev/ttyACM0 \ # <- Use your port --robot.id=my_blue_follower_arm \ # <- Use your robot id --robot.cameras="{ front: {type: opencv, index_or_path: 8, width: 640, height: 480, fps: 30}}" \ # <- Use your cameras --dataset.single_task="Grasp a lego block and put it in the bin." \ # <- Use the same task description you used in your dataset recording --dataset.repo_id=HF_USER/dataset_name \ # <- This will be the dataset name on HF Hub --dataset.episode_time_s=50 \ --dataset.num_episodes=10 \ --policy.path=Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full - Notebooks
- Google Colab
- Kaggle
SmolVLA SO-101 Multi-Task β V6 Full
SmolVLA (~450M) fine-tuned on all four tasks, with the vision encoder unfrozen and image augmentation enabled. The best SmolVLA model in the project β though still meaningfully behind Pi0.5.
Part of Project-IRA β Interactive Robotic Arm. Code: https://github.com/Project-IRA/interactive-robotic-arm
| Base model | lerobot/smolvla_base |
| Robot | SO-101 follower (6-DOF) |
| Training data | Project-IRA/TPSoSe2026_Dataset_Full_Merged_Final_LeRobot_SO101_V1 |
| Recommended checkpoint | around 150000 (not systematically tested) |
| Inputs | camera1 (wrist) + camera2 (desk) images β renamed keys, see Usage β 6-dim joint state, English instruction |
| Outputs | 6-dim continuous action chunks |
Quality
OK-ish. Roughly the same quality as the single-task V2 Lego model, but across most tasks rather than just Lego β so the multi-task generalisation worked, at the same per-task quality level.
Quality was assessed around checkpoint 150000, but checkpoints were not compared systematically for this run β 150000 is a rough indication, not a validated optimum. Checkpoints exist every 10000 steps to 200000 plus last, so it is worth trying several.
For comparison, Pi0.5 on the same dataset works really well at checkpoint 8000. SmolVLA at ~450M parameters appears capacity-limited for this four-task set.
Training
SLURM job 2164957. Single L40S GPU.
| Setting | Value |
|---|---|
| Base | lerobot/smolvla_base |
| Dataset | 930-episode merged set |
| Steps | 200000, completed (--save_freq=10000); assessed around 150000 |
| Batch size | 64 |
| Vision encoder | unfrozen (--policy.freeze_vision_encoder=false) |
| Image augmentation | on (--dataset.image_transforms.enable=true) |
| AMP | on (--policy.use_amp=true) |
| Parameters | 99,880,992 trainable of 450,046,176 total |
| Camera keys | renamed β wrist_left->camera1, desk_view->camera2 |
| SLURM | --gres=gpu:L40S:1 --cpus-per-task=16 --mem=92G --time=52:00:00 |
lerobot-train \
--policy.path=lerobot/smolvla_base \
--dataset.repo_id=TPSoSe2026_Dataset_Full_Merged_Final_LeRobot_SO101 \
--dataset.image_transforms.enable=true \
--policy.freeze_vision_encoder=false \
--policy.use_amp=true \
--batch_size=64 --steps=200000 --save_freq=10000 \
--num_workers=16 --tolerance_s=0.01 \
--output_dir=outputs/train/smolvla_full_v2 \
--job_name=smolvla_full_v2 \
--policy.device=cuda --policy.push_to_hub=false --wandb.enable=false \
--rename_map='{"observation.images.wrist_left": "observation.images.camera1", "observation.images.desk_view": "observation.images.camera2"}'
What changed vs V5
V5 used SmolVLA's defaults, where the vision encoder is frozen β the model's "eyes" stayed locked to their pretrained state and could not adapt to our bricks, lighting and camera angles. V6 unfroze it, added image augmentation and AMP, and doubled the step count from 100k to 200k.
A caveat on freeze_vision_encoder=false
Unfreezing raised trainable parameters only modestly: the run logs report
num_learnable_params=99880992 of num_total_params=450046176 β the same ~100M as the
frozen V5 run. SmolVLA is a SmolVLM2-500M backbone plus a smaller action expert, and
most of the backbone stays frozen regardless of this flag. Do not expect this setting
alone to make all 450M parameters trainable.
Usage
This model expects renamed camera keys. Training used
--rename_mapto remap the dataset's camera features:
Dataset feature What the policy expects Physical camera observation.images.wrist_leftobservation.images.camera1wrist observation.images.desk_viewobservation.images.camera2desk If you feed this policy
wrist_left/desk_viewit will fail or silently misbehave. Name your camerascamera1(wrist) andcamera2(desk) at inference time, or apply the same--rename_mapwhen re-training. The Pi0.5 models do not do this β they use the nativewrist_left/desk_viewnames.
Model files are nested under
outputs_V6/, sofrom_pretrainedon the repo ID will not work:outputs_V6/train/smolvla_full_v2/checkpoints/<step>/pretrained_model/Checkpoints present: every 10000 steps from
010000to200000, pluslast. Repo total ~27.7 GB.hf download Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full \ --include 'outputs_V6/train/smolvla_full_v2/checkpoints/150000/pretrained_model/*' \ --local-dir ./smolvla_v6
import torch
from lerobot.policies.smolvla.modeling_smolvla import SmolVLAPolicy
policy = SmolVLAPolicy.from_pretrained("Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full")
policy = policy.to("cuda").eval()
On-robot rollout:
lerobot-record \
--robot.type=so101_follower \
--robot.port=/dev/ttyACM0 \
--robot.id=$ROBOT_ID \
--robot.cameras='{
camera1: {type: opencv, index_or_path: /dev/v4l/by-path/$WRIST_PATH, width: 640, height: 480, fps: 30},
camera2: {type: opencv, index_or_path: /dev/v4l/by-path/$DESK_PATH, width: 640, height: 480, fps: 30}
}' \
--policy.path=Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full \
--dataset.repo_id=$HF_USER/eval_run \
--dataset.single_task="Sort the lego by color" \
--episodes=10
Use one of the exact training prompts (see the dataset card) as the task string. Both cameras run at 640x480 at inference time even though
desk_viewwas recorded at 800x600.
Robot setup
| Robot | SO-101 follower arm (6-DOF), robot_type: so_follower |
| Teleoperation | SO-101 leader arm |
| Control frequency | 30 fps |
| State / action space | 6-dim: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_roll.pos, gripper.pos |
Camera observation.images.desk_view |
800x600, h264 (recording) |
Camera observation.images.wrist_left |
640x480, h264 (recording) |
Inference note: both cameras are run at 640x480 during inference, not at their recording resolutions, to reduce the payload sent to the inference server.
Environment notes
All training ran on a SLURM cluster with L40S GPUs. Two environment details were required and are easy to miss when reproducing:
- ffmpeg libraries for torchcodec. A minimal conda env supplies the shared libraries
that
torchcodecdiscovers at runtime:export LD_LIBRARY_PATH=$CONDA_PREFIX/envs/ffmpeg_libs_v8/lib:<venv>/lib/python3.12/site-packages/nvidia/npp/lib:$LD_LIBRARY_PATH --tolerance_s=0.01on every run, to accommodate timestamp jitter in the recorded episodes.
Multi-GPU runs additionally set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.
Datasets and the virtualenv were copied to node-local /scratch before training rather
than read from shared storage.
No Weights & Biases logging was enabled for any run (--wandb.enable=false), so there are
no public training curves β the job.*.err SLURM logs are the record.
Tasks and prompts
The model is conditioned on English natural-language instructions. Prompt phrasing was varied roughly every 10 episodes during recording, giving 93 distinct prompts in the merged dataset. Use one of the training prompts verbatim for best results β the full lists are on the dataset card.
Limitations
- Behaviour cloning. The policy imitates teleoperated demonstrations and has no notion of recovery beyond what was demonstrated. It is susceptible to covariate shift and can fail to recover from states outside the demonstration distribution.
- Recovery data is incidental, not systematic. Recovery behaviour appears in the data only where the operator happened to make and correct a mistake during recording; no recovery episodes were scripted deliberately.
- Single environment. All data comes from one lab desk with one lighting setup, one camera geometry, and one set of physical objects. Expect degradation elsewhere.
- Prompt sensitivity. Language conditioning was trained on a fixed set of phrasings (listed in the dataset card). Prompts far from those phrasings may behave unpredictably.
- No formal evaluation. Quality assessments below are qualitative, from operators observing rollouts on the physical arm. There are no success-rate numbers.
- Not safety-rated. Supervise all physical execution and keep the workspace clear.
Upstream licensing & attribution
This model is a derivative work of Apache-2.0 licensed components:
| Component | Upstream | License |
|---|---|---|
| LeRobot framework | https://github.com/huggingface/lerobot | Apache-2.0 |
lerobot/smolvla_base |
https://huggingface.co/lerobot/smolvla_base | Apache-2.0 |
Apache-2.0 permits relicensing derivative works. We retain the upstream copyright notices, license text, and NOTICE files for the incorporated material, as Apache-2.0 Section 4 requires. The upstream components remain under Apache-2.0 β only this project's own contributions (the fine-tuned weights and training configuration) are offered under CC BY-SA 4.0.
CC BY-SA 4.0 was chosen because it is share-alike: derivatives must be released under the same licence, so this work cannot be taken closed-source. The project's source code lives in a separate repository under its own licence β see https://github.com/Project-IRA/interactive-robotic-arm.
Citation
@misc{project_ira_2026,
title = {Project-IRA: Interactive Robotic Arm},
author = {Baten, Cleo and Keppler, Bela and Sapper, Jonas},
year = {2026},
howpublished = {\url{https://huggingface.co/Project-IRA}},
note = {Code: \url{https://github.com/Project-IRA/interactive-robotic-arm}}
}
Model tree for Project-IRA/TPSoSe2026_SmolVLA_LeRobot_SO101_Finetuning_V6_Full
Base model
lerobot/smolvla_base