DepthWorld: 3D World Model for Robot Manipulation

Jai Bardhan, Josef Šivic, Vladimír Petrík

Czech Institute of Informatics, Robotics and Cybernetics (CIIRC), Czech Technical University in Prague

🎉 Accepted to CoRL 2026 🎉

TL;DR

DepthWorld is an action-conditioned world model for DROID-style robot manipulation that jointly predicts multi-view RGB and metric depth, producing rollouts that compose into a single, geometrically consistent 3D world. RGB and depth are packed side-by-side along width per view and modeled by a single Stable Video Diffusion UNet; an optional DPT head additionally regresses per-pixel world-frame point maps.

DepthWorld architecture

Checkpoints

Both are consolidated 90k-step EMA weights (fp32 .pt state dicts), trained on the raw DROID dataset:

file contents description license
horiz_90k_ema.pt UNet + action encoder RGB+depth horiz world model Stability AI Community
horiz_dpt_vggt_90k_ema.pt + DPT pointmap head same, plus a DPT head initialized from facebook/VGGT-1B predicting world-frame XYZ + confidence, trained with a MapAnything-style robust loss Stability AI Community + CC-BY-NC 4.0 (non-commercial)

See License for details.

Usage

hf download jaibrdhn/depthworld horiz_90k_ema.pt config.json --local-dir checkpoints
hf download jaibrdhn/depthworld horiz_dpt_vggt_90k_ema.pt config.json --local-dir checkpoints

Refer to the official GitHub repository for setup, data preparation, training, and evaluation instructions. The eval scripts consume these files directly, e.g.:

python scripts/eval_horiz_chunk.py \
    --ckpt_path checkpoints/horiz_dpt_vggt_90k_ema.pt --width 640 ...

Acknowledgements

Built upon Ctrl-World and Stable Video Diffusion. The pointmap head is warm-started from VGGT and trained with a loss following MapAnything. Trained on the DROID dataset.

This work was supported by the European Union's Horizon Europe projects AGIMUS (No. 101070165), euROBIN (No. 101070596), ERC FRONTIER (No. 101097822), ELIAS (No. 101120237), ELLIOT (No. 101214398), ČVUT Starting grant "DREAM-ACT" (Project ID CVUT-StG-26-089), and CTU Future Fund (Project ID: CVUT-BrF-26-22825M). This work was also supported by the EU’s Horizon Europe Programme under the Grant agreement No. 101136607 (CLARA Project), and was co-funded by the EU from the Operational Programme Jan Amos Komenský (OP JAK) (project "Center for Artificial Intelligence and Quantum Computing in System Brain Research", reg. no. CZ.02.01.01/00/23_029/0008437). Compute resources and infrastructure were supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90254).

Citation

@inproceedings{bardhan2026depthworld,
  title     = {DepthWorld: 3D World Model for Robot Manipulation},
  author    = {Bardhan, Jai and \v{S}ivic, Josef and Petr\'{i}k, Vladim\'{i}r},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}

License

Original code and contributions are released under the MIT license (see the GitHub repository). The checkpoints themselves are derivative works of upstream models and remain subject to their terms:

Downloads last month
5
Video Preview
loading

Model tree for jaibrdhn/depthworld

Finetuned
(7)
this model

Paper for jaibrdhn/depthworld