All-in-One Unified Egocentric World Simulator
EgoU-5B is a unified ego–exo video world simulator: one diffusion transformer that generates synchronized egocentric and exocentric videos from any combination of observed clips, seed frames, text, and the wearer's head-pose trajectory. It is built on Wan2.2-TI2V-5B and trained on EgoExoVideo-600K (Ego-Exo4D and Nymeria, undistorted, 480 px, 37 frames at 15 fps).
egou_5b.safetensors is the checkpoint reported in the paper (the model of the main table, the last row of the
ablation table, and the qualitative figures), evaluated on the 420-group EgoU-Benchmark. It is a full-weight
fine-tune: every base-model tensor is present, plus the Ω geometry cross-attention with its scene-token and
head-pose camera-token inputs, the VGGT-Ω token projection, the head-pose ray embedding and the view-type
embedding (1,263 tensors, 6.15 B parameters, 12.3 GB). No LoRA merge is needed.
Dependencies not in this file
- The Wan2.2-TI2V-5B VAE and umT5 text encoder (unchanged, from
Wan-AI/Wan2.2-TI2V-5B). - The frozen VGGT-Ω backbone that produces the geometry tokens (only its projection is trained and stored here).
- The ego head-pose track of a clip (6-DoF Aria RGB-camera pose per frame) if the ray conditioning is to be used; the same weights run without it.
Loading
# DiffSynth-Studio fork with the EgoU pipeline (diffsynth/pipelines/egou_video_vggt.py)
python tools/eval_egou_multiview.py --finetuned_ckpt egou_5b.safetensors \
--geom_backbone omega --vggt_model_path <VGGT-Omega> --vggt_condition_frames 2 \
--egou_omega_cross_attn --egou_omega_time_bias exp --egou_type_embed \
--egou_camera_tokens --egou_pose_encoder_no_time \
--egou_pose_ray_embed --egou_pose_ray_focal 300 --egou_pose_track_pack <pose pack> \
--frame_grid train --cfg_scale 5.0 ...
The --egou_* enable flags must be passed: the loader creates the modules first and then loads the
tensors; without them the 431 branch tensors are dropped silently (unexpected=431 in the log) and the
model runs without its geometry branch. --egou_camera_tokens --egou_pose_encoder_no_time add the head-pose
camera tokens to the Ω cross-attention (the egou_pose_token_encoder.* and blocks.*.omega_cross_attn.cam_*
tensors). Drop --egou_pose_ray_embed/--egou_pose_track_pack (or pass --no_pose_track) to run without a
head-pose trajectory; the camera tokens are then empty and the same weights apply.
Tasks (one set of weights)
Exo2Ego (with / without an ego seed frame), Ego2Exo, joint cross-view generation, temporal extension, interpolation, Text2Ego. The task is specified by the canvas layout (panels = views, ego first) and a per-panel conditioning mask; observed latents are clamped at every denoising step.
Model tree for EgoVideo/EgoU-5B
Base model
Wan-AI/Wan2.2-TI2V-5B