Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

GitHub Project Page arXiv

Model weights for EVOKE, a 3-step, CFG-free interactive world model that generates 384 Γ— 640 @ 24 fps video and stays coherent over 30 s rollouts. Code, docs and demos live in the GitHub repository β€” this repository holds weights only.

  • ⚑ 3 steps, zero CFG β€” 1.5 s of video every 2.11 s on one H200, one forward per step.
  • 🌍 Endless, not windowed β€” scene geometry lives in an external camera-indexed world state bank, so the denoiser context stays bounded however long the session runs.
  • πŸŽ›οΈ Re-promptable mid-flight β€” change the prompt while the rollout is running, no cut, no restart.

Contents

Every EVOKE directory is the parent of a transformer/, because it loads as from_pretrained(path, subfolder="transformer").

evoke-base/                                   vae / text_encoder / tokenizer / scheduler only
evoke/
β”œβ”€β”€ stage1_camera_control/transformer/        multi-step camera-controllable model
β”œβ”€β”€ stage2_few_step_training/transformer/     few-step distillation (3-step pyramid)
β”œβ”€β”€ stage3_long_distillation/transformer/     30 s long-video distillation (post-distill init)
β”œβ”€β”€ stage3_post_distillation/transformer/     the shipped model
└── evoke_teacher/{high,low}_noise/           the two DMD teacher experts -- training only

Usage

git clone https://github.com/SII-YuanyangYin/Evoke && cd Evoke
pip install -r requirements.txt

huggingface-cli download SII-YuanyangYin/Evoke --local-dir models
huggingface-cli download pkqbajng/ViGeo       --local-dir models/ViGeo1.1   # REQUIRED depth backend

MODE=t2v NUM_CHUNKS=20 bash scripts/inference/infer_post_distill.sh

ViGeo is a separate download and is required β€” every shipped recipe uses it as the depth backend behind the world state bank. Depth-Anything-3 is optional. Both ship under CC-BY-NC-4.0, which is more restrictive than this repository's Apache-2.0; check their licences before any commercial use.

Inference modes, the mode Γ— model matrix, hour-scale rollouts and training are documented in the GitHub repository.

Notes

The distilled models were trained on v2v conditioning alone, so MODE=i2v|t2v on them is zero-shot; only stage1_camera_control has all three modes in distribution.

The vae / text encoder / tokenizer / scheduler in evoke-base/ come from the released Helios base, which traces them to Wan. The EVOKE teacher is built on LingBot-World.

Citation

@article{evoke2026,
  title = {Alaya-EVOKE: From Linear-Scaling Supervision to Endless World},
  year  = {2026},
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support