WorldToken Checkpoints

Paper | Code | Experiment records | Checkpoints | Hugging Face paper page

Model card and index for the 26 selected RoboCasa checkpoints accompanying WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning.

WorldToken maps RGB observations, proprioception and task language into tokens, models their history with a causal Transformer, and predicts action chunks with a diffusion head. The BC-Transformer entry uses the native robomimic baseline.

Paper section Assets Checkpoints
§4 — Multitask control and scaling N1–N3 scaling grid, seed 0, D50/D100/D300/D1000/D2900; BC-Transformer at D300, seed 123 15 + 1
§5 — Token interface N2, K=4 and K=50, seed 0, all five data settings 10

Each run has its own directory and run_info.json with its parameters, original training name and final step. The section indexes list the expected weight filenames; BC-Transformer uses models/model_epoch_1000.pth.

Each directory also includes a config.json, copied from the matching training record with HDF5 data paths normalized to ${DATA_ROOT}/robocasa/mg_im/v0.1/.... The 25 WorldToken configs also carry the matching public recipe's eval_window_spec for offline holdout RMSE; the BC-Transformer uses its separate native evaluation pipeline.

Release availability

All 26 selected checkpoints are available: 15 scaling-grid policies, the BC-Transformer baseline, and 10 token-interface policies.

Training and evaluation records

These policies were trained by imitation learning on the paper's 23-task RoboCasa MG demonstration subsets and evaluated on 23 tasks × 50 episodes, with three full-history repeats per checkpoint.

Full details are in the companion WorldToken Experiment Records package (WorldToken_Experiment_Records), under the same <section>/<run>/ path:

  • config.json and split files: architecture, training settings and data selection.
  • Training logs: optimization and holdout metrics.
  • rollouts/: episode outcomes, summary statistics and videos.

Use the matching run's results when assessing a checkpoint; paper averages over multiple training seeds are not the score of a single released weight file.

Usage

Download all currently available weights and their configurations with the Hugging Face CLI:

hf download mepi31415/WorldToken_Checkpoints --local-dir ./WorldToken_Checkpoints

For only the N2/D300 example below, append --include "04_robocasa_scaling/scaling_n2_d300_seed0/*". Set CHECKPOINT_ROOT to the absolute path of the downloaded directory.

Follow the code repository's environment setup and data preparation. Evaluation requires RoboCasa assets and demonstration HDF5 metadata as well as the weight file. From the WorldToken code repository, after setting DATA_ROOT and CHECKPOINT_ROOT to your local downloads:

RUN=scaling_n2_d300_seed0
SECTION=04_robocasa_scaling
EVAL_RUN="$PWD/runs/checkpoint_eval/$RUN"
mkdir -p "$EVAL_RUN"
cp "$CHECKPOINT_ROOT/$SECTION/$RUN/config.json" "$EVAL_RUN/config.json"
python -m experiments.common.rollout \
  --config "experiments/$SECTION/configs/$RUN.json" \
  --run-dir "$EVAL_RUN" --data-root "$DATA_ROOT" \
  --checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt"

This runs one full-history evaluation and saves new outputs under EVAL_RUN. The section READMEs in the code repository describe other variants and the baseline.

Offline holdout RMSE (WorldToken)

Use the code repository's Linux/CUDA model environment and prepare the frozen CLIP text encoder as described in its environment instructions. Download the matching experiment-records package as well as the checkpoint. From the code repository root, set these paths to your local copies:

export CODE_ROOT="$PWD"
export DATA_ROOT=/absolute/path/to/data
export CHECKPOINT_ROOT=/absolute/path/to/WorldToken_Checkpoints
export RECORDS_ROOT=/absolute/path/to/WorldToken_Experiment_Records
export RUN=scaling_n2_d300_seed0
export SECTION=04_robocasa_scaling
export EVAL_RUN="$CODE_ROOT/runs/checkpoint_eval/$RUN"
# Optional: an existing CLIP cache from the matching training environment.
# export LANG_EMB_CACHE=/absolute/path/to/robocasa_clip.npz

Prepare an evaluation copy with the snippet below. Historical split files may use robocasa/v0.1/...; it maps each HDF5 path to the published config by its portable single_stage/... identity, preserving the demonstration list. It copies an optional existing language cache, or builds one from the holdout language strings using the pinned robomimic CLIP wrapper. ROBOMIMIC_SRC is optional when that source is already installed in the environment. The original records and source cache are not modified.

python - <<'PY'
import json
import os
import shutil
from pathlib import Path

from worldtoken.data import build_lang_embeddings, portable_episode_key
from worldtoken.eval_holdout_rmse import load_persisted_holdout_refs
from worldtoken.train_utils import expand_path_placeholders

run = Path(os.environ["EVAL_RUN"]).resolve()
relative = Path(os.environ["SECTION"]) / os.environ["RUN"]
records = Path(os.environ["RECORDS_ROOT"]) / relative
config_path = Path(os.environ["CHECKPOINT_ROOT"]) / relative / "config.json"
config = json.loads(config_path.read_text(encoding="utf-8"))
run.mkdir(parents=True, exist_ok=True)

paths = config.get("hdf5_paths") or config["dataset"]
by_file = {portable_episode_key(p): str(expand_path_placeholders(p).resolve()) for p in paths}
if len(by_file) != len(paths):
    raise ValueError("Duplicate portable HDF5 identities in config")
split = json.loads((records / "holdout_demos.json").read_text(encoding="utf-8"))
for demo in split["demos"]:
    demo["hdf5_path"] = by_file[portable_episode_key(demo["hdf5_path"])]
(run / "holdout_demos.json").write_text(json.dumps(split, indent=2) + "\n", encoding="utf-8")
shutil.copy2(records / "metrics.jsonl.gz", run / "metrics.jsonl.gz")

cache = run / "holdout_clip.npz"
if os.environ.get("LANG_EMB_CACHE"):
    source = expand_path_placeholders(os.environ["LANG_EMB_CACHE"]).resolve()
    if source != cache:
        shutil.copy2(source, cache)
src = os.environ.get("ROBOMIMIC_SRC")  # Use the local installation, not an archived source path.
src = expand_path_placeholders(src).resolve() if src else None
config["lang_emb_cache"] = str(cache)
config["robomimic_src"] = str(src) if src else None
spec = expand_path_placeholders(config["eval_window_spec"]).resolve()
if not spec.is_file():
    raise FileNotFoundError(spec)
config["eval_window_spec"] = str(spec)
(run / "config.json").write_text(json.dumps(config, indent=2) + "\n", encoding="utf-8")
build_lang_embeddings(
    load_persisted_holdout_refs(run), device_arg="cpu", cache_path=cache,
    robomimic_src=src, mode=config["lang_emb_mode"], write_cache=True,
)
print(f"Prepared {run}; fixed windows: {spec}")
PY

python -m worldtoken.eval_holdout_rmse \
  --run-dir "$EVAL_RUN" \
  --checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" \
  --require-fixed-windows

For another WorldToken run, change SECTION, RUN and the weight filename together. Always use the window spec from its matching recipe; the two C=10 files have different historical start frames and cannot be substituted based on history length alone. If evaluating an already prepared historical run whose config lacks this field, obtain the path from that run's public recipe and pass --eval-window-spec PATH with --require-fixed-windows. The latter prevents silent window reselection.

The result under EVAL_RUN/holdout_grouped_rmse_v2/ includes the actual window path, evaluation protocol, and legacy_v1_parity against the final-step row in the copied metrics.jsonl.gz. An existing metrics.jsonl takes precedence, so use a dedicated evaluation directory for each run. Add --require-legacy-parity to fail when no reference is available or the comparison exceeds --legacy-parity-atol (default 5e-6); the result is saved for inspection before a parity failure. Use --force to recompute an existing result after changing inputs.

Keep eval_batch_size, eval_seed, holdout_rmse_samplers, denoising settings and precision unchanged for a paper comparison. Diffusion samples depend on batch boundaries, so reducing batch size can change the answer despite fixed windows. Use a complete matching language cache when available; rebuilding it or changing hardware/software can introduce numerical differences, which should be reported with the parity result. The fixed-window check validates sample selection, not numerical equivalence of the complete evaluator. The --debug-max-demos and --debug-crops-per-demo subset options cannot be combined with fixed windows, which require the full saved split and crop count.

Scope and license

These are simulation research policies; real-robot performance and transfer to other tasks or observation/action interfaces have not been established. The weights and documentation in this package are licensed under the MIT License. Third-party materials retain their own licenses. The separate experiment-records package specifies its own license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for mepi31415/WorldToken_Checkpoints