MiniMax-H3 — Ref2VA LoRA: CrossView Warp v1
Give it a video and a camera offset — azimuth, elevation — and it generates the same scene from that new viewpoint.
This is the CrossView-Warp idea ported to MiniMax-H3. The model reads two videos: a depth-warp of your clip, which carries the geometry, and the clip itself, which carries identity and appearance. The warp comes from the CrossViewWarp ComfyUI node.
Examples
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
- Prompt
- crossview
The compare_* clips show the warp, the original and the result side by side.
Usage (ComfyUI)
LoRA strength: 0.8 - 1.0 (0.8 is recommended)
Trigger word: crossview (note: You can also influence the generation by prompting what you want to see in the unseen (magenta) areas.)
Example ComfyUI workflow: crossview-warp-h3.json
The CrossViewWarp node has a walkthrough video that covers the camera controls; everything there applies unchanged:
The steps below are what it wires up.
- Install the ComfyUI-CrossViewWarp node.
- Feed your clip to Run MoGe Inference, and its geometry to CrossView Warp.
- Load
MiniMax-H3_Ref2VA-LoRA-CrossView-Warp_v1_3500.safetensorsas a LoRA on the ref2va checkpoint. - Warp -> MiniMax H3 Add Guide,
frame_idx = 0. Source -> MiniMax H3 Reference To Video - Resize warp and source to the same size before both nodes.
- The trigger word:
crossview. You can also influence the generation by prompting what you want to see in the unseen (magenta) areas.
Settings the released examples were generated with:
| Frames | 124 |
| Resolution (1st pass) | 0.5 MP, 16:9 |
| Resolution (2st pass) | 1.5 MP, 16:9 |
| Sampler / steps | res_multistep, 8 steps with the DMD 8-step turbo LoRA |
| Sigma shift | 12 video / 3 audio |
| LoRA strength | 0.8 |
Dataset details
The same 719 Blender scenes as the LTX v2 release. If you want to know more about the dataset you can find more information in my LTX crossview-warp repo:
Third-party assets in the renders (CC-BY 3D models, CC0 HDRIs and textures, CMU motion capture) are listed in ATTRIBUTION.md.
Training
Trained on RunPod — NVIDIA RTX PRO 6000 Blackwell, 96 GB.
| Base model | MiniMax-H3, ref2va |
| Framework | AkaneTendo25/musubi-tuner, branch minimax-h3 |
| Strategy | Ref2VA, aligned_guide_indices = [0] |
| Released checkpoint | step 3,500 of 6,000 |
| LoRA rank / alpha | 32 / 32 |
| Optimizer | adamw8bit 1e-4 |
| Base precision | INT8 ConvRot |
| Resolution | 512×512 × 73 frames |
| Schedule | 6,000 steps, batch 1, no gradient accumulation |
| Measured | 20.35 s/step, 33 h 55 min, 38.4 of 95.6 GB peak |
License
The LoRA weights are a Model Derivative of MiniMax H3 and are distributed under the MiniMax H3 Community License Agreement
Support
Everything here is open, and the GPUs behind it are rented. If this was useful, please consider supporting my work:
- Downloads last month
- 152
Model tree for Cseti/MiniMax-H3_Ref2VA-LoRA-CrossView-Warp_v1
Base model
MiniMaxAI/MiniMax-H3