Robotics
behavior-1k
pi0.5
vla

BEHAVIOR-1K task 37 clean_a_trumpet: fine-tuned from the 2025 1st-place checkpoint

"In the bedroom, pick up the scrub brush from the desk and scrub the cornet (trumpet) on the desk until it's no longer covered in dust."

These weights fine-tune checkpoint_3 of the 1st-place 2025 BEHAVIOR Challenge solution (Larchenko, Zarin, Karnatak; paper) on task 37 alone. The model is Pi0.5-based, with 3.4B parameters. The original checkpoint_3 was trained on 13 tasks (4, 27, 31–33, 35–39, 41, 46, 49).

Status: training only. None of these checkpoints has been evaluated in the BEHAVIOR simulator yet. Lower training loss does not guarantee a higher task success rate.

loss curve

Which checkpoint to use

Folder What it is
lr5e-6/step_09999/ Final checkpoint of the completed run (recommended starting point)
lr5e-6/step_01000/ … step_09000/ Intermediate checkpoints from the same run, every 1k steps
step_01000/ (repo root) From an abandoned first run (peak LR 2.5e-5). Its loss rose after warmup. Kept for reference only.

Each folder contains params/ (orbax, ~12.6 GB) and assets/ (norm stats + FAST tokenizer, copied unchanged from checkpoint_3). Optimizer state is not included.

Results (training metrics only)

start (steps 0–450 avg) end (last 1k steps avg)
action_loss 0.102 0.075

Final logged step (9950): total loss 0.099, subtask accuracy 0.998, FAST token accuracy 0.841, grad norm 1.6. These are single batches of 16 with no held-out validation split.

Training setup

Init IliaLarchenko/behavior_submission/checkpoint_3/params (all weights, incl. task embeddings)
Data Task 37 only: 200 episodes, ~1.06M frames, from IliaLarchenko/behavior_224_rgb (224×224 RGB, 3 cameras)
Steps / batch 10,000 steps × batch 16 (~15% of one epoch). Batch 32 runs out of memory on one 80 GB A100.
LR cosine: warmup 200 → peak 5e-6 → 5e-7 at 10k
Model config Same as the winners' pi_behavior_b1k_fast (FAST aux loss, correlated noise, 15 flow samples, frozen vision backbone)
Norm stats / tokenizer Reused from checkpoint_3 (not recomputed on task 37)
Hardware / time 1× A100-SXM4-80GB, 3.3 s/step, 9 h 08 min
Code behavior-1k-solution ca556f7 + changes.patch; BEHAVIOR-1K 684a83050 + behavior1k_decode.patch

Deviations from the original data pipeline (read before comparing)

  • Videos re-encoded. Task 37's HEVC videos (a keyframe every ~240 frames) were re-encoded to H.264, CRF 18, -g 4, with frame counts verified identical, because random-frame decoding was the bottleneck. The re-encode is lossy but visually near-identical at 224px.
  • Decoder seek backoff reduced from 5 s to 0.5 s (behavior1k_decode.patch). This is safe only with the short-GOP videos above.
  • Added a tasks filter to DataConfig. A weight-loader fix was needed for orbax 0.11.13 (changes.patch).

The abandoned first run

The first attempt used peak LR 2.5e-5. The action loss fell to ~0.10 during warmup, then climbed to ~0.19 once the LR peaked and never recovered. That suggests the LR was too high for a checkpoint already adapted to this task. The run later crashed at step 3000 when a checkpoint save hit a full disk. Only its step_01000 survives.

Usage

Use with the winners' repo and config pi_behavior_b1k_fast (same architecture):

huggingface-cli download jermo1120/b1k-task37-clean-a-trumpet --include "lr5e-6/step_09999/*" --local-dir ./ckpt
uv run scripts/serve_b1k.py policy:checkpoint --policy.config pi_behavior_b1k_fast --policy.dir ./ckpt/lr5e-6/step_09999

Then run the BEHAVIOR eval with task.name=clean_a_trumpet.

Credits

Base model and training code: Ilia Larchenko, Gleb Zarin, Akash Karnatak (repo, checkpoints), built on Physical Intelligence's Pi0.5 / openpi. Benchmark: BEHAVIOR-1K, Stanford Vision and Learning Lab.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for jermo1120/b1k-task37-clean-a-trumpet

Finetuned
(1)
this model

Datasets used to train jermo1120/b1k-task37-clean-a-trumpet

Paper for jermo1120/b1k-task37-clean-a-trumpet