WorldCast: Distributed Multiplayer World Models
Ziyang Ye1, Junchao Huang1,2, Evelyn Zhang2, Zhihao Xie1, Ruicheng Zhang3, Boyao Han1, Litao Ban4, Ziye Wang4, Xinting Hu5, Shaoshuai Shi4, Zhuotao Tian2, Li Jiang1,2†
1CUHK-Shenzhen 2SLAI 3Tsinghua SIGS 4Voyager Research, Didi Chuxing 5USTC †Corresponding author
Project page | Code | Paper (coming soon)
WorldCast is a distributed multiplayer world model. Each player runs a local client, a video generator fine-tuned from Wan2.2-TI2V-5B, on its own GPU. Clients exchange only player states, which each client projects into a camera-aligned player state field, and a shared scene state of generated blocks. This repository holds the weights of the release.
Files
| file | stage | steps | dtype | size | sha256 | use |
|---|---|---|---|---|---|---|
worldcast_4step_bf16.safetensors |
4: 4-step student (distribution matching distillation) | 600 | bf16 | 10.2 GB | 8737c86b94659ba08469778fc20a5e309d9b96b6e9b4bfb3cb816d538db21f17 |
the generator the inference code and the demo run (paper Table 3) |
worldcast_stage3_ar_fp32.safetensors |
3: block-causal model with scene state | 5,000 | fp32 | 20.4 GB | e510c06a4040ee401ca1dde5f04005ad1052182f3a5b33ceca7bc1c1d1306dca |
training checkpoint (initialises stage 4) |
worldcast_stage2s_bidirectional_fp32.safetensors |
2s: bidirectional model with player state field and scene state | 25,000 (20,000 + 5,000) | fp32 | 20.4 GB | 68883cf28a38b0cde657aa5864df453a419ed35797f09b6283468402765f9c94 |
training checkpoint (initialises stage 3) |
depth_head.safetensors |
- | - | fp32 | 176.8 MB | ab599fcd68115142cf3b946e147f3cb89465d0b389371cc3501e1eb3101f6e76 |
picture depth head of the scene state (depth of each generated block) |
depth_readout.safetensors |
- | - | fp32 | 833.4 kB | 6994480d726b7e958f519e835fa306b7045cfb19171ade4aebcb594a67ccd873 |
read-out of the depth head |
fixed_prompt_umt5xxl_bf16.safetensors |
- | - | bf16 | 4.2 MB | a4157803a2c381835b219c079b7d811b7950c3fb2c47741d4552b7bba37cbd0a |
umT5-XXL embedding of the fixed prompt (the client then skips the 11 GB text encoder) |
examples/ |
- | - | - | 201.6 MB | listed in examples/manifest.json of the code |
inputs and expected videos of six recorded rounds (examples/run.sh of the code) |
All generator files hold EMA weights under the parameter names of the release model (WorldCastGenerator), so a
strict load_state_dict works. The training checkpoints also contain the visibility probe used during training
(8 tensors). The training code is released later; the inference code needs only the 4-step model, the depth head,
the read-out and the prompt embedding.
Quick start
git clone https://github.com/Ziyang-Ye/WorldCast && cd WorldCast
pip install -e .
python tools/download_weights.py --out-dir weights
This fetches the four inference files above, plus the Wan2.2 VAE, tokenizer and config.json from
Wan-AI/Wan2.2-TI2V-5B, and writes weights/paths.yaml. --with-training-checkpoints also fetches the two training
checkpoints; --with-taehv fetches the TAEHV tiny decoder (MIT, from its upstream repository) used by the
low-latency demo. Running a session, the examples and the demo: see the
code repository.
License and attribution
- The WorldCast weights are released under the Apache License 2.0.
- They are fine-tuned from Wan2.2-TI2V-5B (Apache License 2.0).
- Training data: the OpenCS2 dataset
(CC BY 4.0). The files under
examples/are derived from it. - The model renders Counter-Strike 2 game content and is intended for research. Counter-Strike 2 is a trademark of Valve Corporation; this work is not affiliated with or endorsed by Valve.
Citation
@article{ye2026worldcast,
title = {WorldCast: Distributed Multiplayer World Models},
author = {Ye, Ziyang and Huang, Junchao and Zhang, Evelyn and Xie, Zhihao and Zhang, Ruicheng and Han, Boyao and Ban, Litao and Wang, Ziye and Hu, Xinting and Shi, Shaoshuai and Tian, Zhuotao and Jiang, Li},
journal = {arXiv preprint},
year = {2026}
}
Model tree for ZiyangYe/WorldCast
Base model
Wan-AI/Wan2.2-TI2V-5B