pi0 drawer_open_place (RoboSynChallenge), LoRA fine-tune of pi0_base

What this is

  • A π0 policy (openpi, JAX) fine-tuned with LoRA (gemma_2b_lora + gemma_300m_lora) from the openpi pi0_base checkpoint, on the RoboSynChallenge/cobotmagic_Sim_drawer_open_place dataset (LeRobot v2.1).
  • Checkpoint: step 16000 (directory label 15999). Params manifest sha256 1232ab36376e95323ee9d106f246a370787c6ad0d86e46665513d90beea5eca5; normalization statistics 8af9ac2030aca1fcae58f72545a27d5adea11da5cec335e988687a118e34fe96.
  • Training recipe: cosine schedule (warmup 1000, peak 2.5e-5, decay over 30000 steps to 2.5e-6), AdamW, batch 32, seed 42, 16000 steps; action offset +1; prompt = the dataset's task text.
  • Use: the RoboSynChallenge simulator (EmbodiChain, CobotMagic dual arm), through the repository's policy/smolvla_multitask adapter with backend: pi0 (policy server in policy/pi05's environment).

Modification notice

This model was modified from pi0_base, which contains PaliGemma / Gemma components. Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms. A copy of the Gemma Terms of Use is included (GEMMA_TERMS_OF_USE.txt, snapshot of the version last modified 2026-04-01), together with the NOTICE file.

Use restrictions (binding on every user and redistributor)

You must not use this model, or any model derived from it, for the restricted uses set forth in the Gemma Prohibited Use Policy at ai.google.dev/gemma/prohibited_use_policy, or in violation of applicable laws and regulations. These restrictions (Section 3.2 of the Gemma Terms of Use) are part of the terms under which this model is distributed, and anyone who redistributes this model or a derivative must pass them on, include a copy of the Gemma Terms of Use and the NOTICE file, and mark modified files as modified.

Evaluation (internal, simulator only)

  • Independent test on 300 frozen drawer_open_place scenes (zero overlap with every earlier scene list), paired against the published SmolVLA drawer checkpoint run unchanged: 203/300 vs 131/300 successes (+24.0 points; exact McNemar p = 1.1e-9). Both systems are complete pipelines (different action execution); the gain is not attributed to the base model alone. This is our evaluation, not the organisers'.
  • Inference: the first call of a process includes JAX compilation (74.7 s on an RTX 4080 SUPER); later calls about 0.15 s.

Limitations

  • Simulation only; trained and tested on one task.
  • Model outputs on GPU are not bit-for-bit reproducible across processes (bf16-level differences observed).

Licences of the inputs -- TODO before release

  • pi0_base weights (openpi): TODO record the licence / terms under which Physical Intelligence distributes them.
  • PaliGemma tokenizer (big_vision/paligemma_tokenizer.model): not included unless its redistribution terms are confirmed; the adapter can obtain it from the original source.
  • Training data (RoboSynChallenge/cobotmagic_Sim_*): no licence tag on Hugging Face; terms for derived weights not confirmed by the organisers.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading