MolmoBot-Pi0 iThor FP13, full FT with trainable vision tower (step 5000)

Companion to the Stage-II adapter ablation: same data/batch/schedule as the plain arm, but the SigLIP vision tower is trained freely (no freeze, no L2-SP).

base MolmoBot-Pi0-DROID (plain, no reasoning adapter)
config molmobot_pi0_droid, action data pos_data/ithor-fp13
batch 256 micro x 2 accum = 512 effective
lr warmup 1000, then constant 5e-5
freeze none (vision tower, LLM, action expert, projections all trainable)
vision regularizer none
step 5000 (final)
val_action_loss 0.0062

Val curve: 500:0.0110 -> 1000:0.0114 -> 1500:0.0099 -> 2000:0.0089 -> 2500:0.0082 -> 3000:0.0076 -> 3500:0.0069 -> 4000:0.0066 -> 4500:0.0067 -> 5000:0.0062

Flat PI0Pytorch state dict (model.safetensors); optimizer state not included.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support