Aligning Forest and Trees in Images and Long Captions for Visually Grounded Understanding
Paper • 2602.02977 • Published
Pretrained checkpoints for CAFT (Cross-domain Alignment of Forests and Trees), from "Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding".
| File | Pre-training data |
|---|---|
CAFT-3M.pt |
CC3M-recap (DreamLIP-3M) |
CAFT-12M.pt |
CC12M-recap (DreamLIP-12M) |
CAFT-15M.pt |
YFCC15M-recap (DreamLIP-15M) |
CAFT-30M.pt |
DreamLIP-30M (merged DreamLIP-3M + DreamLIP-12M + DreamLIP-15M) — default |
Load with the CAFT codebase (model config CAFT-B in src/flair/model_configs/):
torchrun --nproc_per_node 1 -m main \
--model CAFT-B \
--pretrained /path/to/CAFT-30M.pt \
--inference-mode caft \
--alpha 0.3
See the GitHub repo for full training/inference instructions and evaluation setup.
@article{woo2026caft,
title={Aligning Forest and Trees in Images \& Long Captions for Visually Grounded Understanding},
author={Woo, Byeongju and Wang, Zilin and Pak, Byeonghyun and Mo, Sangwoo and Yu, Stella X.},
journal={arXiv preprint arXiv:2602.02977},
year={2026}
}