Instructions to use hmkang/wam_ctxpool_avg with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Wan2.2
How to use hmkang/wam_ctxpool_avg with Wan2.2:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
WAM_DIT4DIT โ context pooling (avg), RoboCasa kitchen, video-only
Wan2.2-TI2V-5B video DiT fine-tuned on RoboCasa (base recipe, effective batch 64:
4 GPU ร per-device 8 ร grad-accum 2), 4-latin history (nin=25 / nout=41, fdf=2),
with average context pooling of the past latent frames inserted before block L.
| folder | pooling | layer | steps |
|---|---|---|---|
L12/checkpoint-<step> |
avg | 12 | every 20k |
L3/checkpoint-<step> |
avg | 3 | every 20k |
Weights (*.safetensors) + config.json + processor / experiment config only; optimizer
state and training_args.bin are not included. Training code: https://github.com/HEMMO0208/wam
(run_scripts/train/wam_dit4dit/compression/finetune_wam_dit4dit_robocasa_kitchen_ctxpool.sh).
Note: an earlier version of this repo held effective-batch-32 runs; those were removed. All checkpoints here are eff-64.
- Downloads last month
- 184