ICRA-S
GR00T-N1.6-3B policy fine-tuned on a Stretch cafe-serving corpus with dual-camera observations (head + gripper).
Training
| Base model | GR00T-N1.6-3B |
| Embodiment tag | NEW_EMBODIMENT |
| Corpus | 120 episodes — 40 GT + 20 relit + 60 pseudo-labeled |
| Frames | 214,566 |
| Epochs | 27 |
| Global batch size | 64 |
| Steps | 90,520 |
| Hardware | 2 x A100-PCIE-40GB |
| Wall time | 31 h 20 m |
| Final loss | 0.0060 (mean of last 50 steps) |
Tuned modules: projector + diffusion head. Vision tower and LLM stay frozen.
Preprocessing
- Letterbox padding to 320 x 320
- Relative action representation, normalized with
relative_stats.json - Mu-law companding (mu = 3) on the wrist and gripper dimensions
- q99 percentile anchors for normalization
- Arm dimension clipped at q01 = 0.0
Pseudo-labels
The 60 pseudo-labeled episodes come from an inverse dynamics model trained on dual-camera serve data with mu-law (mu = 3). Episodes were admitted by a gate combining a global correlation threshold and a windowed NMAE threshold.
Contents
Weights and configuration only. Optimizer state (global_step90520/) is not
included, so this checkpoint is for inference and evaluation rather than
resuming training.
- Downloads last month
- -
Model tree for Gom-sy/ICRA-S
Base model
nvidia/GR00T-N1.6-3B