All-in-One Unified Egocentric World Simulator

EgoU-5B is a unified ego–exo video world simulator: one diffusion transformer that generates synchronized egocentric and exocentric videos from any combination of observed clips, seed frames, text, and the wearer's head-pose trajectory. It is built on Wan2.2-TI2V-5B and trained on EgoExoVideo-600K (Ego-Exo4D and Nymeria, undistorted, 480 px, 37 frames at 15 fps).

egou_5b.safetensors is the checkpoint reported in the paper (the model of the main table, the last row of the ablation table, and the qualitative figures), evaluated on the 420-group EgoU-Benchmark. It is a full-weight fine-tune: every base-model tensor is present, plus the Ω geometry cross-attention with its scene-token and head-pose camera-token inputs, the VGGT-Ω token projection, the head-pose ray embedding and the view-type embedding (1,263 tensors, 6.15 B parameters, 12.3 GB). No LoRA merge is needed.

Dependencies not in this file

  • The Wan2.2-TI2V-5B VAE and umT5 text encoder (unchanged, from Wan-AI/Wan2.2-TI2V-5B).
  • The frozen VGGT-Ω backbone that produces the geometry tokens (only its projection is trained and stored here).
  • The ego head-pose track of a clip (6-DoF Aria RGB-camera pose per frame) if the ray conditioning is to be used; the same weights run without it.

Loading

# DiffSynth-Studio fork with the EgoU pipeline (diffsynth/pipelines/egou_video_vggt.py)
python tools/eval_egou_multiview.py --finetuned_ckpt egou_5b.safetensors \
    --geom_backbone omega --vggt_model_path <VGGT-Omega> --vggt_condition_frames 2 \
    --egou_omega_cross_attn --egou_omega_time_bias exp --egou_type_embed \
    --egou_camera_tokens --egou_pose_encoder_no_time \
    --egou_pose_ray_embed --egou_pose_ray_focal 300 --egou_pose_track_pack <pose pack> \
    --frame_grid train --cfg_scale 5.0 ...

The --egou_* enable flags must be passed: the loader creates the modules first and then loads the tensors; without them the 431 branch tensors are dropped silently (unexpected=431 in the log) and the model runs without its geometry branch. --egou_camera_tokens --egou_pose_encoder_no_time add the head-pose camera tokens to the Ω cross-attention (the egou_pose_token_encoder.* and blocks.*.omega_cross_attn.cam_* tensors). Drop --egou_pose_ray_embed/--egou_pose_track_pack (or pass --no_pose_track) to run without a head-pose trajectory; the camera tokens are then empty and the same weights apply.

Tasks (one set of weights)

Exo2Ego (with / without an ego seed frame), Ego2Exo, joint cross-view generation, temporal extension, interpolation, Text2Ego. The task is specified by the canvas layout (panels = views, ego first) and a per-panel conditioning mask; observed latents are clamped at every denoising step.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EgoVideo/EgoU-5B

Finetuned
(103)
this model