pi0.5 β€” full fine-tune, end-effector space (13-dim pose, 6D rotation) β€” step 15000

This is not the best checkpoint.

It is an intermediate point on the overfitting curve. Held-out loss here is 0.0312, 9% worse than the same run's own minimum of 0.0286 at step 5000.

For the model you should actually deploy, see maskjp/pi05-relative-eef-full-ft.

Published for step-matched comparison. It is 24% better than the 30K endpoint but still worse than the best. It is a research artifact.

Fine-tuned from lerobot/pi05_base on the 949-episode base4 multi-task mixture (900 train / 49 held out, 8 tasks, 3 cameras, 50 Hz), with every parameter trainable β€” 4,143,404,816: vision tower, language model and action expert.

step held-out loss
5K 0.0286
10K 0.0293
15K 0.0312 best
20K 0.0346
25K 0.0375
30K 0.0388

Train loss falls to ~0.008 throughout while held-out loss climbs from step 5000 onward. 4.14B parameters memorise 900 episodes in well under half an epoch. The lesson is about the schedule, not the architecture: 30K steps is wrong for full fine-tuning on this dataset.

For contrast, the frozen-VLM arm (maskjp/pi05-relative-eef-frozen-vlm, 693M trainable) ran the same 30K schedule and never overfit β€” its held-out loss fell monotonically to 0.0283 and was still improving at the end.

config.json differs from training

Training set vision_encoder_lr_multiplier=0.1 (vision tower and multi_modal_projector at 2.5e-6 against 2.5e-5 elsewhere). That field does not exist in stock lerobot 0.6.2 and draccus rejects unknown config keys, so it has been removed from the uploaded config.json. It is read only by get_optim_params and has no effect at inference. train_config.json retains it.

Camera sensitivity

Not measured for this arm. The matched frozen-VLM models score 0.088-0.091 on the camera-swap sensitivity ratio, far below the 0.5 grounding threshold.

Deploying without the training dataset

Normalisation statistics are baked into policy_preprocessor_step_3_normalizer_processor.safetensors and action_feature_names lives in config.json, so no dataset is needed:

from lerobot.configs.policies import PreTrainedConfig
from lerobot.policies.factory import get_policy_class, make_pre_post_processors

repo = "maskjp/pi05-relative-eef-full-ft-15k"
cfg = PreTrainedConfig.from_pretrained(repo)
cfg.device = "cuda"
policy = get_policy_class(cfg.type).from_pretrained(repo, config=cfg)
pre, post = make_pre_post_processors(cfg, pretrained_path=repo)

Avoid make_policy(cfg, ds_meta=...): a non-None ds_meta overwrites action_feature_names from the dataset. This repo also ships meta/ (~1.8 MB) if your deployment insists on it.

Action representation

Targets are relative: action[t+k] - state[anchor], one anchor per chunk, added back after inference. The gripper stays absolute (relative_exclude_joints=['gripper']).

Configuration

pretrained_path lerobot/pi05_base
chunk_size 50   n_action_steps 10   n_obs_steps 1
freeze_vision_encoder false   train_expert_only false   (all 4.14B trainable)
gradient_checkpointing true   compile_model true   dtype bfloat16
optimizer_lr 2.5e-5, warmup 1000, cosine to 2.5e-6 over 30K
vision tower + projector at 0.1x that (2.5e-6)
norm  VISUAL IDENTITY | STATE QUANTILES | ACTION QUANTILES
batch 21/rank x 3 GPUs = 63 effective, DDP, seed 1000

Apache-2.0, inherited from LeRobot.

Downloads last month
7
Safetensors
Model size
4B params
Tensor type
F32
Β·
BF16
Β·
Video Preview
loading