WorldToken Checkpoints
Paper | Code | Experiment records | Checkpoints | Hugging Face paper page
Model card and index for the 26 selected RoboCasa checkpoints accompanying WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning.
WorldToken maps RGB observations, proprioception and task language into tokens, models their history with a causal Transformer, and predicts action chunks with a diffusion head. The BC-Transformer entry uses the native robomimic baseline.
| Paper section | Assets | Checkpoints |
|---|---|---|
| §4 — Multitask control and scaling | N1–N3 scaling grid, seed 0, D50/D100/D300/D1000/D2900; BC-Transformer at D300, seed 123 | 15 + 1 |
| §5 — Token interface | N2, K=4 and K=50, seed 0, all five data settings | 10 |
Each run has its own directory and run_info.json with its parameters,
original training name and final step. The section indexes list the expected
weight filenames; BC-Transformer uses models/model_epoch_1000.pth.
Each directory also includes a config.json, copied from the matching
training record with HDF5 data paths normalized to
${DATA_ROOT}/robocasa/mg_im/v0.1/.... The 25 WorldToken configs also carry
the matching public recipe's eval_window_spec for offline holdout RMSE;
the BC-Transformer uses its separate native evaluation pipeline.
Release availability
All 26 selected checkpoints are available: 15 scaling-grid policies, the BC-Transformer baseline, and 10 token-interface policies.
Training and evaluation records
These policies were trained by imitation learning on the paper's 23-task RoboCasa MG demonstration subsets and evaluated on 23 tasks × 50 episodes, with three full-history repeats per checkpoint.
Full details are in the companion WorldToken Experiment Records package
(WorldToken_Experiment_Records), under the same <section>/<run>/ path:
config.jsonand split files: architecture, training settings and data selection.- Training logs: optimization and holdout metrics.
rollouts/: episode outcomes, summary statistics and videos.
Use the matching run's results when assessing a checkpoint; paper averages over multiple training seeds are not the score of a single released weight file.
Usage
Download all currently available weights and their configurations with the Hugging Face CLI:
hf download mepi31415/WorldToken_Checkpoints --local-dir ./WorldToken_Checkpoints
For only the N2/D300 example below, append
--include "04_robocasa_scaling/scaling_n2_d300_seed0/*".
Set CHECKPOINT_ROOT to the absolute path of the downloaded directory.
Follow the code repository's environment setup
and data preparation.
Evaluation requires RoboCasa assets and demonstration HDF5 metadata as well as
the weight file. From the WorldToken code repository, after setting
DATA_ROOT and CHECKPOINT_ROOT to your local downloads:
RUN=scaling_n2_d300_seed0
SECTION=04_robocasa_scaling
EVAL_RUN="$PWD/runs/checkpoint_eval/$RUN"
mkdir -p "$EVAL_RUN"
cp "$CHECKPOINT_ROOT/$SECTION/$RUN/config.json" "$EVAL_RUN/config.json"
python -m experiments.common.rollout \
--config "experiments/$SECTION/configs/$RUN.json" \
--run-dir "$EVAL_RUN" --data-root "$DATA_ROOT" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt"
This runs one full-history evaluation and saves new outputs under EVAL_RUN.
The section READMEs in the code repository describe other variants and the baseline.
Offline holdout RMSE (WorldToken)
Use the code repository's Linux/CUDA model environment and prepare the frozen CLIP text encoder as described in its environment instructions. Download the matching experiment-records package as well as the checkpoint. From the code repository root, set these paths to your local copies:
export CODE_ROOT="$PWD"
export DATA_ROOT=/absolute/path/to/data
export CHECKPOINT_ROOT=/absolute/path/to/WorldToken_Checkpoints
export RECORDS_ROOT=/absolute/path/to/WorldToken_Experiment_Records
export RUN=scaling_n2_d300_seed0
export SECTION=04_robocasa_scaling
export EVAL_RUN="$CODE_ROOT/runs/checkpoint_eval/$RUN"
# Optional: an existing CLIP cache from the matching training environment.
# export LANG_EMB_CACHE=/absolute/path/to/robocasa_clip.npz
Prepare an evaluation copy with the snippet below. Historical split files may
use robocasa/v0.1/...; it maps each HDF5 path to the published config by its
portable single_stage/... identity, preserving the demonstration list.
It copies an optional existing language cache, or builds one from the holdout
language strings using the pinned robomimic CLIP wrapper. ROBOMIMIC_SRC is
optional when that source is already installed in the environment. The
original records and source cache are not modified.
python - <<'PY'
import json
import os
import shutil
from pathlib import Path
from worldtoken.data import build_lang_embeddings, portable_episode_key
from worldtoken.eval_holdout_rmse import load_persisted_holdout_refs
from worldtoken.train_utils import expand_path_placeholders
run = Path(os.environ["EVAL_RUN"]).resolve()
relative = Path(os.environ["SECTION"]) / os.environ["RUN"]
records = Path(os.environ["RECORDS_ROOT"]) / relative
config_path = Path(os.environ["CHECKPOINT_ROOT"]) / relative / "config.json"
config = json.loads(config_path.read_text(encoding="utf-8"))
run.mkdir(parents=True, exist_ok=True)
paths = config.get("hdf5_paths") or config["dataset"]
by_file = {portable_episode_key(p): str(expand_path_placeholders(p).resolve()) for p in paths}
if len(by_file) != len(paths):
raise ValueError("Duplicate portable HDF5 identities in config")
split = json.loads((records / "holdout_demos.json").read_text(encoding="utf-8"))
for demo in split["demos"]:
demo["hdf5_path"] = by_file[portable_episode_key(demo["hdf5_path"])]
(run / "holdout_demos.json").write_text(json.dumps(split, indent=2) + "\n", encoding="utf-8")
shutil.copy2(records / "metrics.jsonl.gz", run / "metrics.jsonl.gz")
cache = run / "holdout_clip.npz"
if os.environ.get("LANG_EMB_CACHE"):
source = expand_path_placeholders(os.environ["LANG_EMB_CACHE"]).resolve()
if source != cache:
shutil.copy2(source, cache)
src = os.environ.get("ROBOMIMIC_SRC") # Use the local installation, not an archived source path.
src = expand_path_placeholders(src).resolve() if src else None
config["lang_emb_cache"] = str(cache)
config["robomimic_src"] = str(src) if src else None
spec = expand_path_placeholders(config["eval_window_spec"]).resolve()
if not spec.is_file():
raise FileNotFoundError(spec)
config["eval_window_spec"] = str(spec)
(run / "config.json").write_text(json.dumps(config, indent=2) + "\n", encoding="utf-8")
build_lang_embeddings(
load_persisted_holdout_refs(run), device_arg="cpu", cache_path=cache,
robomimic_src=src, mode=config["lang_emb_mode"], write_cache=True,
)
print(f"Prepared {run}; fixed windows: {spec}")
PY
python -m worldtoken.eval_holdout_rmse \
--run-dir "$EVAL_RUN" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" \
--require-fixed-windows
For another WorldToken run, change SECTION, RUN and the weight filename
together. Always use the window spec from its matching recipe; the two C=10
files have different historical start frames and cannot be substituted based
on history length alone. If evaluating an already prepared historical run
whose config lacks this field, obtain the path from that run's public recipe
and pass --eval-window-spec PATH with --require-fixed-windows. The latter
prevents silent window reselection.
The result under EVAL_RUN/holdout_grouped_rmse_v2/ includes the actual window
path, evaluation protocol, and legacy_v1_parity against the
final-step row in the copied metrics.jsonl.gz. An existing metrics.jsonl
takes precedence, so use a dedicated evaluation directory for each run.
Add --require-legacy-parity to fail when no reference is available or the
comparison exceeds --legacy-parity-atol (default 5e-6); the result is saved
for inspection before a parity failure. Use --force to recompute an existing
result after changing inputs.
Keep eval_batch_size, eval_seed, holdout_rmse_samplers, denoising settings
and precision unchanged for a paper comparison. Diffusion samples depend on
batch boundaries, so reducing batch size can change the answer despite fixed
windows. Use a complete matching language cache when available; rebuilding it
or changing hardware/software can introduce numerical differences, which
should be reported with the parity result. The fixed-window check validates
sample selection, not numerical equivalence of the complete evaluator.
The --debug-max-demos and --debug-crops-per-demo subset options cannot be
combined with fixed windows, which require the full saved split and crop count.
Scope and license
These are simulation research policies; real-robot performance and transfer to other tasks or observation/action interfaces have not been established. The weights and documentation in this package are licensed under the MIT License. Third-party materials retain their own licenses. The separate experiment-records package specifies its own license.