Instructions to use hmkang/wam_3latin_joint_mildrot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Wan2.2
How to use hmkang/wam_3latin_joint_mildrot with Wan2.2:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
WAM_DIT4DIT — 3-latin JOINT (video + action), mild + rotate augmentation, RoboCasa kitchen
Wan2.2-TI2V-5B (diffusers) video DiT + DiT-B action head trained jointly on RoboCasa kitchen with the base recipe (effective batch 64: 4 GPU × per-device 8 × grad-accum 2, fm shift 5.0, bell time-weight, EMA warmup on video + action head, 4-knob A joint setup).
Geometry: 3-latin history — nin=17 / nout=33, frame-downsample 2 (hist 9 → 3 cond latent
- 4 future latent). Image augmentation: mild color/crop jitter + rotation ±5° on all views
(
WAM_IMAGE_AUG=1 WAM_IMAGE_AUG_STRONGER=0 WAM_IMAGE_AUG_MILD_ROTATE=1).
| folder | steps |
|---|---|
checkpoint-<step> |
every 20k (80k first) |
Weights (*.safetensors) + config.json + processor / experiment config only; optimizer state and
training_args.bin are not included. Training code: https://github.com/HEMMO0208/wam
(gr00t/model/wam_dit4dit/image_augmentations.py, launcher finetune_wam_dit4dit_robocasa_kitchen_DiT4DiTjoint_wan22.sh).
- Downloads last month
- -