VAM-Cross MimicVideo fused Video2World DiT

This public repository contains the Bridge-initialized Video2World DiT after fusing the trainable-only LoRA from iteration 1060 of v2w_widowx_target_bridgeinit_lora_r256. The fused checkpoint is checkpoints/video_backbone/iter_000001060_fused.pt.

The Bridge backbone and LoRA adapter are no longer separate at load time. This is a fused DiT checkpoint, not a complete inference bundle: the compatible MimicVideo code, video tokenizer, and T5 encoder are still required.

Provenance

  • MimicVideo commit: e3355dbc93132b576c02f920a59b4fc18a4f5906
  • Bridge checkpoint repository: jonpai/mimic-video
  • Bridge checkpoint revision: f28339034831e3c2374be075e622e1ff38ebe0f8
  • Source dataset revision: dreamdifferent/vam-cross-target-widowx250-native@44c28a3bada463f06845daee8a1dd87607d32956
  • Cameras: observation.images.corner_cam, observation.images.front_cam
  • Video rate: 5 Hz

The private training dataset is not included. Users must comply with the upstream MimicVideo, NVIDIA Cosmos, and base-checkpoint terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading