AirVLA FT-D-KI (pi0, step 5000, frozen backbone)

Knowledge-insulated variant of FT-D: the VLM backbone is frozen during fine-tuning. Recovers approach precision and navigation but only part of FT-C's grounding.

Part of the AirVLA MSc dissertation campaign: adapting a pretrained pi0 vision-language-action policy to a quadcopter with a 2-DoF arm and parallel gripper, evaluated in MuJoCo 3.3.4.

Property Value
Base model lerobot/pi0
Training data hanapasta/airvla_v22 (v2 + F1 + F2)
Initialised from FT-C, VLM backbone frozen
Training steps 5,000
Checkpoint selection lowest pinned-noise validation MSE over the saved ladder

Frozen-protocol result (n = 60 pick + 20 navigation, paired scenes)

0/60 grasped - 43/60 correct-target - 117.8 mm median (88.2 mm target-true, best pure-policy precision at the time) - 12/20 navigation.

Usage

Evaluated with the frozen harness eval_v2.py from the code repository:

python eval_v2.py <this_checkpoint> 60 20 --torchseed 1000 --tag run --video

Code, reproduction guide and the full experimental ledger: https://github.com/robotics-hana/drone-version2

Evaluation uses MuJoCo 3.3.4; a different simulator version changes contact behaviour enough to invalidate comparison with the banked results.

Downloads last month
16
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading