Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models
Paper • 2609.06967 • Published
Checkpoints for Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models (BMVC 2026). Paper · Code
Each folder holds model.safetensors (the full model state dict, fp32) and config.json, which names the released
config to evaluate it with.
hf download SoongE/ReCalCon --include "distribution_shift/imagenet_b16/*" --local-dir checkpoints
python -m scripts.eval --config imagenet_b16 \
--eval.checkpoint checkpoints/distribution_shift/imagenet_b16/model.safetensors
Distribution shift (top-1 accuracy, %)
| Checkpoint | Row | ImageNet | -R | -A | -V2 | -Sketch | ObjectNet |
|---|---|---|---|---|---|---|---|
distribution_shift/imagenet_b16 |
ViT-B/16, Update, ImageNet | 83.284 | 71.607 | 54.893 | 74.1 | 51.781 | 58.388 |
distribution_shift/imagenet_b32 |
ViT-B/32, Update, ImageNet | 79.354 | 62.41 | 32.573 | 68.15 | 44.503 | 49.203 |
distribution_shift/imagenet_petl_b16 |
ViT-B/16, Freeze (PETL), ImageNet | 80.206 | 73.163 | 50.867 | 70.13 | 49.614 | 55.691 |
distribution_shift/imagenet_petl_b32 |
ViT-B/32, Freeze (PETL), ImageNet | 76.12 | 63.563 | 30.453 | 64.52 | 42.267 | 46.662 |
iWildCam (macro-F1, %). The released evaluation reports the OOD split.
| Checkpoint | Row | ID (macro-F1) | OOD (macro-F1) |
|---|---|---|---|
distribution_shift/iwildcam_b16 |
ViT-B/16, Update, iWildCam | 51.253 | 37.803 |
distribution_shift/iwildcam_b32 |
ViT-B/32, Update, iWildCam | 42.116 | 29.275 |
distribution_shift/iwildcam_petl_b16 |
ViT-B/16, Freeze (PETL), iWildCam | 45.49 | 30.327 |
distribution_shift/iwildcam_petl_b32 |
ViT-B/32, Freeze (PETL), iWildCam | 34.709 | 24.502 |
Transfer learning (top-1 accuracy, %)
| Checkpoint | Dataset | Accuracy |
|---|---|---|
transfer/caltech101 |
Caltech101 | 97.85 |
transfer/flowers102 |
Flowers102 | 98.894 |
transfer/pcam |
PCam | 89.413 |
transfer/stanfordcars |
StanfordCars | 91.419 |
Apache License 2.0. The models are fine-tuned from OpenAI CLIP (MIT License); each evaluation dataset keeps its own terms.
@inproceedings{oh2026recalibrated,
title = {Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models},
author = {Oh, Seungmin and Kang, Seunghun and Ryu, Jongbin},
booktitle = {Proceedings of the British Machine Vision Conference},
year = {2026},
publisher = {BMVA},
url = {https://arxiv.org/abs/2609.06967}
}
Base model
openai/clip-vit-base-patch16