BadWAM: When World-Action Models Dream Right but Act Wrong
Paper • 2607.15207 • Published • 54
Fast-WAM joint variant (create_fastwam_joint: action attends to ALL video latent tokens),
built on Wan2.2-TI2V-5B video DiT + ActionDiT (MoT, ~6B params), trained on RoboTwin 2.0.
Clean 50,000-step run on 8 GPUs (effective batch 128), lr 1e-4 cosine (warmup 5%), AdamW(0.9,0.95), weight_decay 1e-2, bf16, ZeRO-1. 3 cams (cam_high + L/R wrist, 240x320 each, tiled), num_frames 33, action_video_freq_ratio 4 (32 actions / 9 video frames), action&state dim 14.
This matches the BadWAM (arXiv:2607.15207) reproduction recipe (8xH100 / 50k / batch16 / lr1e-4 cosine), whose reported clean RoboTwin success rates are joint 90.9% / idm 91.4% / action-only 92.1%.
Notes:
Base model
Wan-AI/Wan2.2-TI2V-5B