Robotics
LeRobot
Safetensors
smolvla

Model Card for smolvla

SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.

This policy has been trained and pushed to the Hub using LeRobot. See the full documentation at LeRobot Docs.


How to Get Started with the Model

For a complete walkthrough, see the training guide. Below is the short version on how to train and run inference/eval:

Train from scratch

lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --policy.type=act \
  --output_dir=outputs/train/<desired_policy_repo_id> \
  --job_name=lerobot_training \
  --policy.device=cuda \
  --policy.repo_id=${HF_USER}/<desired_policy_repo_id>
  --wandb.enable=true

Writes checkpoints to outputs/train/<desired_policy_repo_id>/checkpoints/.

Evaluate the policy/run inference

lerobot-record \
  --robot.type=so100_follower \
  --dataset.repo_id=<hf_user>/eval_<dataset> \
  --policy.path=<hf_user>/<desired_policy_repo_id> \
  --episodes=10

Prefix the dataset repo with eval_ and supply --policy.path pointing to a local or hub checkpoint.


Model Details

  • License: apache-2.0

Comments

  • This smolvla policy completes the pick and place task but not always consistent. If the object is placed at a different spot each episode (left, right, near, far) the policy picks up the object correctly.
  • Checked with changing phone camera zoom, night mode one/off and it does completes the task. However when the phone camera is changed from the one it's trained (changes the zoom, and lighting) it does not work consistently.
  • Solved the jittering problem upto good extent by makeing policy.num_steps = 20 from earlier 10. Reducing policy.n_action_steps = 50 to 30, 20 does make the task completion worse and don't solve the jittering either.
Downloads last month
90
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for subhodipsaha/smolvla_pick_place_07_16

Finetuned
(6985)
this model

Dataset used to train subhodipsaha/smolvla_pick_place_07_16

Paper for subhodipsaha/smolvla_pick_place_07_16