HumanoidToolBench DreamZero5B: Real G1 additional training, 91 demonstrations
This training stage has not started. No trained checkpoint is available in this repository.
This is a HumanoidToolBench training-result repository. It does not substitute an upstream pretrained policy for a HumanoidToolBench-trained checkpoint.
| Setting | Value |
|---|---|
| Training stage | Real G1 additional training, 91 demonstrations |
| Target optimizer updates | 5000 |
| Per-GPU batch / GPUs / global batch | 16 / 8 / 128 |
| Gradient accumulation | 1 |
| Conditions per global batch | 4 |
| Dataset revision | 47eca9322bb53fa1c685363271a87d2e414cb0e8 |
Initialization: snupilab/humanoidtoolbench-dreamzero5b-sim-3003 after 40,000 simulation updates, followed by 5,000 new updates using the 91 real G1 demonstrations.
The four real conditions are StickMove Standard/Reasoning and HookRetrieve Standard/Reasoning. Recordings are 20 Hz. Hardware executed joint-target actions require the matching real G1 control adapter. A simulation adapter must not be assumed compatible.
Use the model's native HumanoidToolBench adapter and model-specific dependencies. This repository does not claim compatibility with arbitrary Transformers or simulation loaders. No evaluation score is claimed by checkpoint publication.
Training uses independent model optimizers and shared GPU execution through MPS. Publication is performed by a CPU uploader after final checkpoint validation.