BEHAVIOR-1K task 37 clean_a_trumpet: fine-tuned from the 2025 1st-place checkpoint
"In the bedroom, pick up the scrub brush from the desk and scrub the cornet (trumpet) on the desk until it's no longer covered in dust."
These weights fine-tune checkpoint_3 of the 1st-place 2025 BEHAVIOR Challenge solution (Larchenko, Zarin, Karnatak; paper) on task 37 alone. The model is Pi0.5-based, with 3.4B parameters. The original checkpoint_3 was trained on 13 tasks (4, 27, 31–33, 35–39, 41, 46, 49).
Status: training only. None of these checkpoints has been evaluated in the BEHAVIOR simulator yet. Lower training loss does not guarantee a higher task success rate.
Which checkpoint to use
| Folder | What it is |
|---|---|
lr5e-6/step_09999/ |
Final checkpoint of the completed run (recommended starting point) |
lr5e-6/step_01000/ … step_09000/ |
Intermediate checkpoints from the same run, every 1k steps |
step_01000/ (repo root) |
From an abandoned first run (peak LR 2.5e-5). Its loss rose after warmup. Kept for reference only. |
Each folder contains params/ (orbax, ~12.6 GB) and assets/ (norm stats + FAST tokenizer, copied unchanged from checkpoint_3). Optimizer state is not included.
Results (training metrics only)
| start (steps 0–450 avg) | end (last 1k steps avg) | |
|---|---|---|
| action_loss | 0.102 | 0.075 |
Final logged step (9950): total loss 0.099, subtask accuracy 0.998, FAST token accuracy 0.841, grad norm 1.6. These are single batches of 16 with no held-out validation split.
Training setup
| Init | IliaLarchenko/behavior_submission/checkpoint_3/params (all weights, incl. task embeddings) |
| Data | Task 37 only: 200 episodes, ~1.06M frames, from IliaLarchenko/behavior_224_rgb (224×224 RGB, 3 cameras) |
| Steps / batch | 10,000 steps × batch 16 (~15% of one epoch). Batch 32 runs out of memory on one 80 GB A100. |
| LR | cosine: warmup 200 → peak 5e-6 → 5e-7 at 10k |
| Model config | Same as the winners' pi_behavior_b1k_fast (FAST aux loss, correlated noise, 15 flow samples, frozen vision backbone) |
| Norm stats / tokenizer | Reused from checkpoint_3 (not recomputed on task 37) |
| Hardware / time | 1× A100-SXM4-80GB, 3.3 s/step, 9 h 08 min |
| Code | behavior-1k-solution ca556f7 + changes.patch; BEHAVIOR-1K 684a83050 + behavior1k_decode.patch |
Deviations from the original data pipeline (read before comparing)
- Videos re-encoded. Task 37's HEVC videos (a keyframe every ~240 frames) were re-encoded to H.264, CRF 18,
-g 4, with frame counts verified identical, because random-frame decoding was the bottleneck. The re-encode is lossy but visually near-identical at 224px. - Decoder seek backoff reduced from 5 s to 0.5 s (
behavior1k_decode.patch). This is safe only with the short-GOP videos above. - Added a
tasksfilter toDataConfig. A weight-loader fix was needed for orbax 0.11.13 (changes.patch).
The abandoned first run
The first attempt used peak LR 2.5e-5. The action loss fell to ~0.10 during warmup, then climbed to ~0.19 once the LR peaked and never recovered. That suggests the LR was too high for a checkpoint already adapted to this task. The run later crashed at step 3000 when a checkpoint save hit a full disk. Only its step_01000 survives.
Usage
Use with the winners' repo and config pi_behavior_b1k_fast (same architecture):
huggingface-cli download jermo1120/b1k-task37-clean-a-trumpet --include "lr5e-6/step_09999/*" --local-dir ./ckpt
uv run scripts/serve_b1k.py policy:checkpoint --policy.config pi_behavior_b1k_fast --policy.dir ./ckpt/lr5e-6/step_09999
Then run the BEHAVIOR eval with task.name=clean_a_trumpet.
Credits
Base model and training code: Ilia Larchenko, Gleb Zarin, Akash Karnatak (repo, checkpoints), built on Physical Intelligence's Pi0.5 / openpi. Benchmark: BEHAVIOR-1K, Stanford Vision and Learning Lab.
Model tree for jermo1120/b1k-task37-clean-a-trumpet
Base model
IliaLarchenko/behavior_submission