LingBot-VLA 2.0 (robotwin)

LingBot-VLA 2.0 converted for LeRobot (policy.type=lingbot_vla_v2): a Qwen3-VL-4B backbone with a sparse-MoE Qwen2 action expert, trained with flow matching.

Checkpoint fine-tuned by the authors on RoboTwin 2.0 (checkpoints/global_step_50000).

The weights are the upstream robbyant/lingbot-vla-v2-6b-robotwin weights, unchanged (fp32, every tensor checked against upstream), with a LeRobot config and processors.

Usage

lerobot-eval \
  --policy.path=lerobot/lingbot_vla_v2_robotwin \
  --env.type=robotwin \
  --env.task=beat_block_hammer \
  --eval.batch_size=1 \
  --eval.n_episodes=100 \
  --rename_map='{"observation.images.head_camera": "observation.images.cam_high", "observation.images.left_camera": "observation.images.cam_left_wrist", "observation.images.right_camera": "observation.images.cam_right_wrist"}'

License

Apache-2.0, as the upstream release. Fine-tuning with the dual-query distillation (on by default) downloads frozen teachers under their own licenses, including DINO-Video weights under the DINOv3 License. Inference does not use them.

Citation

@article{lingbotvla2,
  title={From Foundation to Application: Improving VLA Models in Practice},
  author={Wei Wu and others},
  journal={arXiv preprint arXiv:2607.06403},
  year={2026}
}
Downloads last month
48
Safetensors
Model size
6B params
Tensor type
F32
·
Video Preview
loading

Collection including lerobot/lingbot_vla_v2_robotwin

Paper for lerobot/lingbot_vla_v2_robotwin