Robotics
LeRobot
Safetensors
groot
so101
dagger

SO-101 vials to rack: GR00T N1.7 after two DAgger rounds

A GR00T N1.7 (3B) policy for the SO-101 arm and the task "Pick up the vial and place it in the rack". It is the DR500 simulation checkpoint fine-tuned on 70 human corrections collected on the real robot with HG-DAgger: 40 grasp-and-lift corrections (round 1) and 30 placement corrections (round 2).

On our real SO-101 it completed the full task in 5 of 5 trials. The DR500 checkpoint it starts from never grasped the vial on the same setup.

Results on the real SO-101

Five trials per model, vial at nearby positions:

Model Training data Result
DR500 checkpoint 500 simulated demonstrations Never grasps
Round 1 50 simulated episodes + 40 grasp-and-lift corrections Grasps and lifts; no placement
Round 2 + 20 placement corrections Full task 2/5
This model + 30 placement corrections Full task 5/5

Limitations: five trials at nearby vial positions only. Far positions still fail at reach or grasp, because no correction covers the approach. All corrections come from one robot, camera setup and rack position; on another setup, expect to collect your own corrections.

Training

  • Starting checkpoint: sreetz-nv/so101_newton_dr500_vision_h32_b32_50k_20260925, GR00T N1.7 trained on 500 simulated demonstrations.
  • Data: 120 episodes.
  • Recipe: LeRobot 0.6 lerobot-train, 5,000 steps, batch 32, AdamW at lr 1e-4 with 500 warmup steps and cosine decay, seed 42, image augmentation. Action chunks of 32 steps, with relative arm actions and an absolute gripper. The vision encoder, projector and action head are trained; the language model is frozen. The checkpoint keeps the DR500 normalization statistics.
  • About 1.4 h on one RTX PRO 6000, with the Qwen3-VL image preprocessing moved to the GPU. The stock CPU pipeline gives the same losses and robot results.

Usage

Inputs: a wrist and an external RGB camera (640×480), the six joint positions (degrees, gripper in percent) and the task text. Start each trial from the DR500 start pose [0, -34.88, -3.91, 86.86, -91.87, 1.18]. Our robot trials ran with LeRobot 0.6:

lerobot-rollout --strategy.type=episodic --policy.path=knpt/round2_fast_30ep_5k --policy.base_model_path=nvidia/GR00T-N1.7-3B --policy.n_action_steps=16 --robot.type=so101_follower --robot.port=<port> --robot.id=<id> --robot.use_degrees=true --robot.cameras="<wrist and external cameras>" --task="Pick up the vial and place it in the rack" --dataset.single_task="Pick up the vial and place it in the rack" --dataset.repo_id=<user>/<eval_dataset> --dataset.num_episodes=5 --dataset.episode_time_s=30 --dataset.push_to_hub=false --device=cuda --inference.type=rtc --inference.rtc.enabled=false --inference.queue_threshold=0 --interpolation_multiplier=2

License

Derived from NVIDIA GR00T N1.7. Licensed by NVIDIA Corporation under the NVIDIA Open Model License Agreement.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
F32
·
Video Preview
loading

Model tree for knpt/round2_fast_30ep_5k

Datasets used to train knpt/round2_fast_30ep_5k