CompACT — 16-token checkpoints

Final 16-token checkpoints for Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model (CVPR 2026).

Code and setup instructions · Paper · Project

Included models

Directory Model Resolution Checkpoint
tokenizer-16-224 CompACT, 16 tokens 224 × 224 checkpoints/epoch=24-step=500000.ckpt
tokenizer-16-256 CompACT, 16 tokens 256 × 256 checkpoints/epoch=24-step=500000.ckpt
cdit-b-16 CDiT-B world model 224 × 224 checkpoints/latest.pth.tar
cdit-l-16 CDiT-L world model 224 × 224 checkpoints/latest.pth.tar

Both world models use tokenizer-16-224. The 256-resolution tokenizer is provided separately. These are the original full training checkpoint files, including training state; world-model inference uses the ema weights. Only the final checkpoint for each variant is included. Exact training steps, original experiment names, file sizes, and SHA-256 checksums are in manifest.json.

Download

Install the environment following the code repository. Run from its root:

from huggingface_hub import snapshot_download

snapshot_download(repo_id="kdwon/CompACT", local_dir="checkpoints/CompACT")

To download just one variant, include its config:

snapshot_download(
    repo_id="kdwon/CompACT",
    local_dir="checkpoints/CompACT",
    allow_patterns=["tokenizer-16-224/**", "manifest.json", "SHA256SUMS"],
)

Configuration

Each variant includes .hydra/config.yaml, in the format expected by the code repository. Set these environment variables to your local paths:

export COMPACT_CKPT_ROOT="$(pwd)/checkpoints/CompACT"
export BASE_TOKENIZER_CKPT=/absolute/path/to/base-checkpoints
export DATASET_PREFIX=/absolute/path/to/datasets

The current constructors also require these initialization files:

  • $BASE_TOKENIZER_CKPT/mage_vqgan/vqgan_jax_strongaug.ckpt
  • $BASE_TOKENIZER_CKPT/dinov3/dinov3_vitb16_pretrain_lvd1689m-73cec8be.pth

See the code repository for obtaining the MAGE and DINOv3 files. DATASET_PREFIX must be defined even for standalone tokenizer loading because the world-model loader resolves the complete saved tokenizer configuration.

World-model configs use ${oc.env:COMPACT_CKPT_ROOT}/tokenizer-16-224 and ${oc.env:DATASET_PREFIX}/nwm/{recon,sacson,scand} instead of the original machine's paths. Adjust dataset overrides to match your local layout.

Load a tokenizer

uv run load_tokenizer_checkpoint.py "$COMPACT_CKPT_ROOT/tokenizer-16-224" --no-test

Use tokenizer-16-256 for the 256-resolution variant. For image encoding, use the saved DINO normalization (dinov2_mean and dinov2_std); the tokenizer's output normalization buffers describe decoded images.

Planning evaluation

With the required navigation datasets installed:

uv run bash scripts/plan.sh --nproc=1 -- \
  ++exp_dir="$COMPACT_CKPT_ROOT/cdit-b-16" \
  ++tokenizer_path="$COMPACT_CKPT_ROOT/tokenizer-16-224" \
  ++ckp=latest

Replace cdit-b-16 with cdit-l-16 to use the larger world model. Dataset setup and evaluation options are documented in the code repository.

Validation

These files passed strict weight-loading checks in the public CompACT codebase. Both tokenizers encoded and reconstructed a real image. Both the regular and EMA weights of the world models passed forward checks, and EMA models completed image-to-predicted-image inference. Validation used PyTorch 2.6.0+cu124 and bfloat16 inference on an RTX 6000 Ada GPU. These smoke checks do not constitute a rerun of the paper's evaluation benchmarks.

Verify downloaded checkpoint bytes from the download directory:

sha256sum -c SHA256SUMS

Citation

@inproceedings{kim2026planning,
  title={Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model},
  author={Kim, Dongwon and Seo, Gawon and Lee, Jinsung and Cho, Minsu and Kwak, Suha},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for kdwon/CompACT