--- language: - en tags: - compact - image-tokenization - world-model - robotics - pytorch - image-to-image --- # CompACT — 16-token checkpoints Final 16-token checkpoints for **Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model** (CVPR 2026). [Code and setup instructions](https://github.com/kdwonn/CompACT) · [Paper](https://arxiv.org/abs/2603.05438) · [Project](https://kdwonn.github.io/CompACT) ## Included models | Directory | Model | Resolution | Checkpoint | |---|---|---|---| | `tokenizer-16-224` | CompACT, 16 tokens | 224 × 224 | `checkpoints/epoch=24-step=500000.ckpt` | | `tokenizer-16-256` | CompACT, 16 tokens | 256 × 256 | `checkpoints/epoch=24-step=500000.ckpt` | | `cdit-b-16` | CDiT-B world model | 224 × 224 | `checkpoints/latest.pth.tar` | | `cdit-l-16` | CDiT-L world model | 224 × 224 | `checkpoints/latest.pth.tar` | Both world models use **`tokenizer-16-224`**. The 256-resolution tokenizer is provided separately. These are the original full training checkpoint files, including training state; world-model inference uses the `ema` weights. Only the final checkpoint for each variant is included. Exact training steps, original experiment names, file sizes, and SHA-256 checksums are in [manifest.json](manifest.json). ## Download Install the environment following the [code repository](https://github.com/kdwonn/CompACT). Run from its root: ```python from huggingface_hub import snapshot_download snapshot_download(repo_id="kdwon/CompACT", local_dir="checkpoints/CompACT") ``` To download just one variant, include its config: ```python snapshot_download( repo_id="kdwon/CompACT", local_dir="checkpoints/CompACT", allow_patterns=["tokenizer-16-224/**", "manifest.json", "SHA256SUMS"], ) ``` ## Configuration Each variant includes `.hydra/config.yaml`, in the format expected by the code repository. Set these environment variables to your local paths: ```bash export COMPACT_CKPT_ROOT="$(pwd)/checkpoints/CompACT" export BASE_TOKENIZER_CKPT=/absolute/path/to/base-checkpoints export DATASET_PREFIX=/absolute/path/to/datasets ``` The current constructors also require these initialization files: - `$BASE_TOKENIZER_CKPT/mage_vqgan/vqgan_jax_strongaug.ckpt` - `$BASE_TOKENIZER_CKPT/dinov3/dinov3_vitb16_pretrain_lvd1689m-73cec8be.pth` See the code repository for obtaining the MAGE and DINOv3 files. `DATASET_PREFIX` must be defined even for standalone tokenizer loading because the world-model loader resolves the complete saved tokenizer configuration. World-model configs use `${oc.env:COMPACT_CKPT_ROOT}/tokenizer-16-224` and `${oc.env:DATASET_PREFIX}/nwm/{recon,sacson,scand}` instead of the original machine's paths. Adjust dataset overrides to match your local layout. ## Load a tokenizer ```bash uv run load_tokenizer_checkpoint.py "$COMPACT_CKPT_ROOT/tokenizer-16-224" --no-test ``` Use `tokenizer-16-256` for the 256-resolution variant. For image encoding, use the saved DINO normalization (`dinov2_mean` and `dinov2_std`); the tokenizer's output normalization buffers describe decoded images. ## Planning evaluation With the required navigation datasets installed: ```bash uv run bash scripts/plan.sh --nproc=1 -- \ ++exp_dir="$COMPACT_CKPT_ROOT/cdit-b-16" \ ++tokenizer_path="$COMPACT_CKPT_ROOT/tokenizer-16-224" \ ++ckp=latest ``` Replace `cdit-b-16` with `cdit-l-16` to use the larger world model. Dataset setup and evaluation options are documented in the code repository. ## Validation These files passed strict weight-loading checks in the public CompACT codebase. Both tokenizers encoded and reconstructed a real image. Both the regular and EMA weights of the world models passed forward checks, and EMA models completed image-to-predicted-image inference. Validation used PyTorch 2.6.0+cu124 and bfloat16 inference on an RTX 6000 Ada GPU. These smoke checks do not constitute a rerun of the paper's evaluation benchmarks. Verify downloaded checkpoint bytes from the download directory: ```bash sha256sum -c SHA256SUMS ``` ## Citation ```bibtex @inproceedings{kim2026planning, title={Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model}, author={Kim, Dongwon and Seo, Gawon and Lee, Jinsung and Cho, Minsu and Kwak, Suha}, booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year={2026} } ```