Memorizon: Training World Models Beyond Their Context Window
Tingting Liao ยท
Xuezhi Liang ยท
Hao Li ยท
Guangyi Liu
Institute of Foundation Models (IFM), MBZUAI
โจ This checkpoint
| Model | Memorizon, 4-step distilled (Self-Forcing + DMD, CFG distilled in) |
| Base | Wan2.2-TI2V-5B |
| Input | one image + keyboard actions or a camera trajectory |
| Output | 864ร480, 16 fps, generated one second at a time |
| Sampling | 4 steps, no CFG (read from the config) |
๐ Usage
git clone https://github.com/TingtingLiao/memorizon.git && cd memorizon && pip install -e .
python scripts/generate.py --checkpoint Luffuly/memorizon --image photo.jpg \
--actions "w*16 l*24 w*16 l*24" --output out.mp4
from memorizon import MemorizonPipeline, actions_to_c2w, save_video
pipe = MemorizonPipeline.from_pretrained("Luffuly/memorizon")
video = pipe("photo.jpg", actions_to_c2w("w*16 l*24 w*16 l*24"))
save_video(video, "out.mp4")
๐ฎ Keyboard actions โ one per 0.25 s
| Key | Action | Key | Action |
|---|---|---|---|
w s |
forward / backward ยท 0.25 m | j l |
turn left / right ยท 7.5ยฐ |
a d |
left / right ยท 0.25 m | i k |
look up / down ยท 7.5ยฐ |
. |
stay | *N |
repeat N times |
Prompts โ captioned automatically
Without a prompt, the image is captioned by Qwen3-VL-8B in the training format
(memorizon/caption.py). Custom prompts should follow it โ a perspective prefix,
then one paragraph:
First-person perspective โ character not visible. A winding paved path curves beneath a canopy of vibrant pink cherry blossoms, โฆ
Camera trajectories
Instead of actions, pass --trajectory a [T, 4, 4] array of camera-to-world poses,
T = 1 + 4 ร seconds, in the first frame's coordinates (OpenCV axes), translations
in metres / 4.
๐ Citation
@article{memorizon2026,
title = {Memorizon: Training World Models Beyond Their Context Window},
author = {Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu},
year = {2026},
note = {arXiv preprint}
}
- Downloads last month
- 22
Model tree for Luffuly/memorizon
Base model
Wan-AI/Wan2.2-TI2V-5B-Diffusers