VAM-Cross MimicVideo World2Action decoder

This public repository contains the World2Action decoder checkpoint from iteration 900 of w2a_panda_robosuite_level2_2cam_wrist_ablation_v2_action_iter2374_videolora_iter200_widowx_teleop_recording_frame_v1. The run stopped because of completed. The newest complete model, optimizer, scheduler, and trainer checkpoint set was verified before selecting the uploaded model weight.

Required frozen inputs

  • MimicVideo commit: e3355dbc93132b576c02f920a59b4fc18a4f5906
  • Initial Video2World backbone: dreamdifferent/widowx250-wrist-ablation-v1-video-fused@8e39d96344dea0a82ae673874a38206a6c412948
  • Initial action decoder: dreamdifferent/widowx250-wrist-ablation-v1-action-decoder-iter2374@0ea6db91672db019d9d6dc9a6c9bd1ecc0504001
  • Frozen Video LoRA: dreamdifferent/panda-robosuite-level2-wrist-ablation-v2-video-lora-iter200@399d3773289e26a0e063aa7418f05368306d065a

Action and data contract

  • Dataset: dreamdifferent/vam-cross-level2-panda-robosuite-widowx-texture-corner-wrist-v2@2a165cf153fca01ad80cab3689d778e34b63ff38
  • Episodes/frames: 280 / 54426
  • Cameras: observation.images.corner_cam, observation.images.wrist_cam
  • Target: 15 achieved-EE/gripper actions at 5 Hz
  • Pose target: relative_to_current_achieved_pose in widowx_reference_base/teleop_aligned_tool
  • Rotation: rotation_6d

The dataset and frozen inputs are not included. Use the pinned JSON and effective config.yaml included in this repository.

Downloads last month
6
Video Preview
loading