Instructions to use maskjp/pi05-relative-eef-full-ft-15k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use maskjp/pi05-relative-eef-full-ft-15k with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
pi0.5 β full fine-tune, end-effector space (13-dim pose, 6D rotation) β step 15000
This is not the best checkpoint.
It is an intermediate point on the overfitting curve. Held-out loss here is
0.0312, 9% worse than the same run's own minimum of0.0286at step 5000.For the model you should actually deploy, see
maskjp/pi05-relative-eef-full-ft.Published for step-matched comparison. It is 24% better than the 30K endpoint but still worse than the best. It is a research artifact.
Fine-tuned from lerobot/pi05_base on the
949-episode base4 multi-task mixture (900 train / 49 held out, 8 tasks, 3 cameras, 50 Hz),
with every parameter trainable β 4,143,404,816: vision tower, language model and action
expert.
| step | held-out loss |
|---|---|
| 5K | 0.0286 |
| 10K | 0.0293 |
| 15K | 0.0312 best |
| 20K | 0.0346 |
| 25K | 0.0375 |
| 30K | 0.0388 |
Train loss falls to ~0.008 throughout while held-out loss climbs from step 5000 onward. 4.14B parameters memorise 900 episodes in well under half an epoch. The lesson is about the schedule, not the architecture: 30K steps is wrong for full fine-tuning on this dataset.
For contrast, the frozen-VLM arm
(maskjp/pi05-relative-eef-frozen-vlm, 693M trainable) ran the same
30K schedule and never overfit β its held-out loss fell monotonically to 0.0283 and was
still improving at the end.
config.json differs from training
Training set vision_encoder_lr_multiplier=0.1 (vision tower and multi_modal_projector at
2.5e-6 against 2.5e-5 elsewhere). That field does not exist in stock lerobot 0.6.2 and draccus
rejects unknown config keys, so it has been removed from the uploaded config.json. It is
read only by get_optim_params and has no effect at inference. train_config.json retains it.
Camera sensitivity
Not measured for this arm. The matched frozen-VLM models score 0.088-0.091 on the camera-swap sensitivity ratio, far below the 0.5 grounding threshold.
Deploying without the training dataset
Normalisation statistics are baked into
policy_preprocessor_step_3_normalizer_processor.safetensors and action_feature_names lives in
config.json, so no dataset is needed:
from lerobot.configs.policies import PreTrainedConfig
from lerobot.policies.factory import get_policy_class, make_pre_post_processors
repo = "maskjp/pi05-relative-eef-full-ft-15k"
cfg = PreTrainedConfig.from_pretrained(repo)
cfg.device = "cuda"
policy = get_policy_class(cfg.type).from_pretrained(repo, config=cfg)
pre, post = make_pre_post_processors(cfg, pretrained_path=repo)
Avoid make_policy(cfg, ds_meta=...): a non-None ds_meta overwrites action_feature_names
from the dataset. This repo also ships meta/ (~1.8 MB) if your deployment insists on it.
Action representation
Targets are relative: action[t+k] - state[anchor], one anchor per chunk, added back after
inference. The gripper stays absolute (relative_exclude_joints=['gripper']).
Configuration
pretrained_path lerobot/pi05_base
chunk_size 50 n_action_steps 10 n_obs_steps 1
freeze_vision_encoder false train_expert_only false (all 4.14B trainable)
gradient_checkpointing true compile_model true dtype bfloat16
optimizer_lr 2.5e-5, warmup 1000, cosine to 2.5e-6 over 30K
vision tower + projector at 0.1x that (2.5e-6)
norm VISUAL IDENTITY | STATE QUANTILES | ACTION QUANTILES
batch 21/rank x 3 GPUs = 63 effective, DDP, seed 1000
Apache-2.0, inherited from LeRobot.
- Downloads last month
- 7