CoaG: Cylinders on a Grid โ€” LoRA for Wan2.2-Fun-A14B-Control

Coarse 3D layout control for video generation: a control video made of a ground grid and one solid cylinder per person (81 frames, 16 fps) plus a background reference image and a text prompt produce a video in which people stand where the cylinders stand, move as the cylinders move, and the camera follows the drawn camera path.

Files

file expert notes
coag_lora_low_noise_r64_e1.safetensors low-noise rank 64, alpha 32, targets q,k,v,ffn.0,ffn.2; 1 epoch (241 steps x 8 GPUs)
coag_lora_high_noise_r64_e1.safetensors high-noise same recipe

Trained 2026-09-13 with VideoX-Fun (commit 968f0e2 + the small dataset patch in the GitHub repo) on 1935 tuples (control video, LaMa background reference, caption, target clip) at the 480p bucket (token length 640), 81 frames.

Use

Load both files with VideoX-Fun's examples/wan2.2_fun/predict_v2v_control_ref.py (lora_path = low-noise file, lora_high_path = high-noise file, weight 0.55 each), give it the control video, a background image as ref_image, and a prompt; 50 steps, guidance 6, 480x832x81. The GitHub repo has an environment-variable driven version (train/patched/infer_control_ref.py, train/run_infer_case.sh) and an editor that produces compatible control videos.

License

The LoRA weights are derived from Wan2.2-Fun-A14B-Control and follow its Apache-2.0 license; the code that trained them is MIT (GitHub).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for zshyang1106/CoaG-Wan2.2-Fun-A14B-Control-LoRA

Adapter
(1)
this model