Controlled multi-task pi0 SFT

A: D0; B: D0+D1; C: D0+D1+D2. Independent official pi0_base starts, LoRA rank/alpha32, batch8, seed0, 3001 updates. 50/100/150 demonstrations; corpus-specific normalization.

Final D0/D1/D2/H1/H2 success (%), N50 paired seeds10000–10049: A=70/0/0/0/0; B=56/18/0/0/0; C=42/38/62/0/0. No observed common held-out gain. Training diversity, unique data, per-task exposure and normalization are confounded.

All15 checkpoints at0/500/1000/2000/3001 include parameters, normalization assets and training state. Final checkpoints are at checkpoints/{A,B,C}/3001. Update0 is saved LoRA initialization, not dense zero-shot pi0.

User requested remaining intermediate evaluation be stopped on2026-09-23. Twelveof36 intermediate N50cells completed; interrupted B500 partial episodes are raw evidence only. All15final N50cells and image sensitivity completed. Missing points are not zeroes.

Code, report, notebook and videos: https://github.com/qrvmil/maniskill_myws/tree/exp/pi0-multitask-sft-generalization

OpenPI981483dca0fd9acba698fea00aa6e52d56a66c58; LIBERO8f1084e3132a39270c3a13ebe37270a43ece2a01. Follow repository methods and evaluation API with each checkpoint's own assets/norm_stats.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading