Instructions to use IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam with LeRobot:
# See https://github.com/huggingface/lerobot?tab=readme-ov-file#installation for more details git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e .[smolvla]
# Launch finetuning on your dataset python lerobot/scripts/train.py \ --policy.path=IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam \ --dataset.repo_id=lerobot/svla_so101_pickplace \ --batch_size=64 \ --steps=20000 \ --output_dir=outputs/train/my_smolvla \ --job_name=my_smolvla_training \ --policy.device=cuda \ --wandb.enable=true
# Run the policy using the record function python -m lerobot.record \ --robot.type=so101_follower \ --robot.port=/dev/ttyACM0 \ # <- Use your port --robot.id=my_blue_follower_arm \ # <- Use your robot id --robot.cameras="{ front: {type: opencv, index_or_path: 8, width: 640, height: 480, fps: 30}}" \ # <- Use your cameras --dataset.single_task="Grasp a lego block and put it in the bin." \ # <- Use the same task description you used in your dataset recording --dataset.repo_id=HF_USER/dataset_name \ # <- This will be the dataset name on HF Hub --dataset.episode_time_s=50 \ --dataset.num_episodes=10 \ --policy.path=IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam - Notebooks
- Google Colab
- Kaggle
SmolVLA - Stack (joint, 2-cam)
HuggingFace LeRobot SmolVLA fine-tuned on the ManiGuard stack base task (sim Franka Panda). Part of the ManiGuard VLA benchmark - SmolVLA vs pi0.5 vs GR00T on the same task families with identical data, cameras, and controller.
Model
- Base: lerobot/smolvla_base - SmolVLM2 vision-language backbone + flow-matching action expert
- Embodiment: Franka Panda, 8-D joint state/action (7 arm joints + 1 gripper), padded to SmolVLA's 32-D
- Cameras (2):
observation.images.top(overview =leftview) +observation.images.wrist(256x256) - Action: absolute joint targets (fed straight to a JointController at eval); 50-step chunk
- Tuning: SmolVLA default - vision encoder frozen, train the action expert (no LoRA)
Training
- 8-GPU accelerate DDP, global batch 512, 10360 steps (~2 epochs over 2,652,083 frames)
- Data: IDEAS-Lab-Northwestern/datagen-stack-v1-joint-5cam - the 5-cam datagen dataset, prepared to a 2-cam standard-keyed copy (H.264 videos)
Usage
Load with SmolVLAPolicy.from_pretrained("IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam-yanZ") from LeRobot. The checkpoint carries the normalization stats.
WARNING - Convention (must match at eval): joint-space JointController (absolute joint targets) + 2 cameras (
observation.images.top= theleftoverview,observation.images.wrist). A mismatched controller, camera set, or overview view silently feeds an out-of-distribution input.
Paper & Citation
Part of ManiGuard: paper (arXiv:2608.17386) · code · docs
@misc{peng2026maniguard,
title = {{MANIGUARD}: A Benchmark and Data Suite for Specification-Grounded
Safety Evaluation and Improvement of Robotic Manipulation},
author = {Peng, Yiyan and Wang, Philip and Zhan, Simon Sinong and Lyu, Yiqi
and Ni, Zhenyang and Yan, Jixin and Wong, Fiorelli and Jiao, Ruochen
and Yin, Hang and Cao, Xinyu and Shao, Huajie and Li, Manling
and Zhang, Ruohan and Zhu, Qi},
year = {2026},
eprint = {2608.17386},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2608.17386},
}
License
The base model lerobot/smolvla_base
does not declare an explicit license for its weights (its SmolVLM2-500M backbone and
the lerobot training code are both Apache-2.0). ManiGuard's fine-tuning contributions
are released under Apache-2.0; the licensing status of the combined weights follows
the upstream base model. This tag will be aligned if upstream declares a license.
- Downloads last month
- 45
Model tree for IDEAS-Lab-Northwestern/smolvla-base-datagen-v1-stack-joint-2cam
Base model
lerobot/smolvla_base