DepthWorld: 3D World Model for Robot Manipulation
Jai Bardhan, Josef Šivic, Vladimír Petrík
Czech Institute of Informatics, Robotics and Cybernetics (CIIRC), Czech Technical University in Prague
TL;DR
DepthWorld is an action-conditioned world model for DROID-style robot manipulation that jointly predicts multi-view RGB and metric depth, producing rollouts that compose into a single, geometrically consistent 3D world. RGB and depth are packed side-by-side along width per view and modeled by a single Stable Video Diffusion UNet; an optional DPT head additionally regresses per-pixel world-frame point maps.
Checkpoints
Both are consolidated 90k-step EMA weights (fp32 .pt state dicts), trained
on the raw DROID dataset:
| file | contents | description | license |
|---|---|---|---|
horiz_90k_ema.pt |
UNet + action encoder | RGB+depth horiz world model | Stability AI Community |
horiz_dpt_vggt_90k_ema.pt |
+ DPT pointmap head | same, plus a DPT head initialized from facebook/VGGT-1B predicting world-frame XYZ + confidence, trained with a MapAnything-style robust loss |
Stability AI Community + CC-BY-NC 4.0 (non-commercial) |
See License for details.
Usage
hf download jaibrdhn/depthworld horiz_90k_ema.pt config.json --local-dir checkpoints
hf download jaibrdhn/depthworld horiz_dpt_vggt_90k_ema.pt config.json --local-dir checkpoints
Refer to the official GitHub repository for setup, data preparation, training, and evaluation instructions. The eval scripts consume these files directly, e.g.:
python scripts/eval_horiz_chunk.py \
--ckpt_path checkpoints/horiz_dpt_vggt_90k_ema.pt --width 640 ...
Acknowledgements
Built upon Ctrl-World and Stable Video Diffusion. The pointmap head is warm-started from VGGT and trained with a loss following MapAnything. Trained on the DROID dataset.
This work was supported by the European Union's Horizon Europe projects AGIMUS (No. 101070165), euROBIN (No. 101070596), ERC FRONTIER (No. 101097822), ELIAS (No. 101120237), ELLIOT (No. 101214398), ČVUT Starting grant "DREAM-ACT" (Project ID CVUT-StG-26-089), and CTU Future Fund (Project ID: CVUT-BrF-26-22825M). This work was also supported by the EU’s Horizon Europe Programme under the Grant agreement No. 101136607 (CLARA Project), and was co-funded by the EU from the Operational Programme Jan Amos Komenský (OP JAK) (project "Center for Artificial Intelligence and Quantum Computing in System Brain Research", reg. no. CZ.02.01.01/00/23_029/0008437). Compute resources and infrastructure were supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90254).
Citation
@inproceedings{bardhan2026depthworld,
title = {DepthWorld: 3D World Model for Robot Manipulation},
author = {Bardhan, Jai and \v{S}ivic, Josef and Petr\'{i}k, Vladim\'{i}r},
booktitle = {Conference on Robot Learning (CoRL)},
year = {2026}
}
License
Original code and contributions are released under the MIT license (see the GitHub repository). The checkpoints themselves are derivative works of upstream models and remain subject to their terms:
- Both checkpoints are fine-tuned from
stabilityai/stable-video-diffusion-img2vidand are therefore subject to the Stability AI Community License. - The DPT pointmap head in
horiz_dpt_vggt_90k_ema.ptwas initialized fromfacebook/VGGT-1B(CC-BY-NC 4.0); treat that checkpoint as non-commercial research use.
- Downloads last month
- 5
Model tree for jaibrdhn/depthworld
Base model
stabilityai/stable-video-diffusion-img2vid