MolmoBot-Pi0 iThor FP13, full FT with trainable vision tower (step 5000)
Companion to the Stage-II adapter ablation: same data/batch/schedule as the
plain arm, but the SigLIP vision tower is trained freely (no freeze, no L2-SP).
| base | MolmoBot-Pi0-DROID (plain, no reasoning adapter) |
| config | molmobot_pi0_droid, action data pos_data/ithor-fp13 |
| batch | 256 micro x 2 accum = 512 effective |
| lr | warmup 1000, then constant 5e-5 |
| freeze | none (vision tower, LLM, action expert, projections all trainable) |
| vision regularizer | none |
| step | 5000 (final) |
| val_action_loss | 0.0062 |
Val curve: 500:0.0110 -> 1000:0.0114 -> 1500:0.0099 -> 2000:0.0089 -> 2500:0.0082 -> 3000:0.0076 -> 3500:0.0069 -> 4000:0.0066 -> 4500:0.0067 -> 5000:0.0062
Flat PI0Pytorch state dict (model.safetensors); optimizer state not included.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support