b1k single-task checkpoints
Params-only checkpoints for offline evaluation. train_state is not included, so these
cannot be resumed from โ inference/eval only.
task5_human_only_21k
Single-task fine-tune of the 50-task meta checkpoint on task 5 = setting_mousetraps,
using the 200 human demonstrations from the 2026 BEHAVIOR-1K challenge set and nothing else.
| init | IliaLarchenko/behavior_50t_checkpoint (params) |
| data | 2026-challenge-demos-224, activity setting_mousetraps, 200 episodes / 2,039,222 frames |
| step | 21,000 of a planned 30,000 (run stopped early; resumable copy retained locally) |
| batch | 224 (32/GPU ร 7ร H200), FSDP over 7 devices |
| LR | 8.75e-7 โ 8.75e-5 โ 1.75e-6, 2,000-step warmup (linear rule off BS-128 / 5e-5) |
| norm stats | norm-stats-fixed (sha256 2b42bdfbโฆ), i.e. 2025 stats with the robot-frame base_qvel correction on state dims 0โ2 |
| task space | B1K_TASK_SPACE=100; setting_mousetraps is index 5 in both the 2025 and 2026 tables, so its task/stage embeddings are pretrained, not randomly initialised |
| final logged | action_loss = 0.0209 at step 21,775 |
task10_taskfinetune
Rollout evaluation: see
task10_taskfinetune/EVAL_RESULTS.mdโ 0.0 % success / avg q 0.1917 over 20public_testinstances, vs 5.0 % / 0.1750 for the officialcheckpoint_1. The difference is not statistically significant (sign test p = 1.000).
Single-task fine-tune of the 50-task meta checkpoint on task 10 =
set_up_a_coffee_station_in_your_kitchen, using the 200 human demonstrations from the
2026 BEHAVIOR-1K challenge set and nothing else. Run completed in full.
| init | IliaLarchenko/behavior_50t_checkpoint (params) |
| data | 2026-challenge-demos-224, activity set_up_a_coffee_station_in_your_kitchen, 200 episodes / 1,253,243 frames |
| step | 29,999 of 30,000 (completed, 16 h 54 m on 8ร H200) |
| batch | 256 (32/GPU ร 8ร H200), FSDP over 8 devices |
| LR | flat 5e-6 โ init = peak = decay = 5e-6, so no warmup ramp and no cosine decay |
| norm stats | 2025 stats with the robot-frame base_qvel correction on state dims 0โ2 (same qvelfix stats as the task-5 runs) |
| task space | B1K_TASK_SPACE=100; set_up_a_coffee_station_in_your_kitchen is index 10 in both the 2025 and 2026 tables, so its task/stage embeddings are pretrained, not randomly initialised |
| final logged | action_loss = 0.0294, total_loss = 0.0565, fast_accuracy = 0.8208 at step 29,975 |
Because the LR is flat with no annealing, the final step is not necessarily the best checkpoint โ there is no decay phase to settle into a minimum. Steps 4000/8000/.../28000 were retained locally and can be uploaded if you want to sweep across them.
Video encoding caveat for this checkpoint
Trained on the 2026-challenge-demos-224 build, whose bitstream reports x264 crf=20.0,
keyint=250. Measured against b1k-224x224-gop8-fixed (x265 crf=28.0, keyint=8) on
identical frames, this build is slightly softer: 0.886ร Laplacian variance, 0.933ร HF
spectral energy. The downscaling filter is not recorded in the bitstream and is unverified.
No task-success delta has been measured for this difference โ it is an input-statistics
observation only.
One video file in the source data (right_realsense .../chunk-010/file-002.mp4, backing 46 of
the 200 episodes) had no moov atom and was unreadable; it was re-encoded from the intact
480ร480 HEVC original at x264 CRF 23 / GOP 250 to match its neighbours, verified at 43โ50 dB
PSNR against the source.
Reading the loss
action_loss here is training-set loss. There is no validation split and no rollout
evaluation has been run, so it says nothing about task success or generalisation. It is also
not comparable across tasks โ episode length varies ~6ร between activities.
Reproducing the input pipeline
The model consumes a 23-dim state extracted from the 61-dim observation.state
(base_qvel 0:3, arm_left 3:10, gripper_left 24:26, arm_right 28:35,
gripper_right 49:51, trunk 53:57), and predicts 23-dim actions with the trunk and both
arms as deltas relative to the current state (use_delta_joint_actions=True); base
velocity and both grippers stay absolute. Action normalisation is per-timestamp
(use_per_timestamp_norm=True) over a 30-step horizon โ the scalar mean/std in
norm_stats.json are not what the pipeline divides by for actions.
Cameras are 224ร224 h264: zed_link_camera_0 (head), left_realsense_link_camera_0,
right_realsense_link_camera_0.
Model tree for fastwalker1118/b1kcheckpoints
Base model
IliaLarchenko/behavior_50t_checkpoint