CAFT

Pretrained checkpoints for CAFT (Cross-domain Alignment of Forests and Trees), from "Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding".

Checkpoints

File Pre-training data
CAFT-3M.pt CC3M-recap (DreamLIP-3M)
CAFT-12M.pt CC12M-recap (DreamLIP-12M)
CAFT-15M.pt YFCC15M-recap (DreamLIP-15M)
CAFT-30M.pt DreamLIP-30M (merged DreamLIP-3M + DreamLIP-12M + DreamLIP-15M) — default

Usage

Load with the CAFT codebase (model config CAFT-B in src/flair/model_configs/):

torchrun --nproc_per_node 1 -m main \
    --model CAFT-B \
    --pretrained /path/to/CAFT-30M.pt \
    --inference-mode caft \
    --alpha 0.3

See the GitHub repo for full training/inference instructions and evaluation setup.

Citation

@article{woo2026caft,
  title={Aligning Forest and Trees in Images \& Long Captions for Visually Grounded Understanding},
  author={Woo, Byeongju and Wang, Zilin and Pak, Byeonghyun and Mo, Sangwoo and Yu, Stella X.},
  journal={arXiv preprint arXiv:2602.02977},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for byeongju-woo/CAFT