🌦️ [COLM 2026] All-Weather VLM

LoRA adapter for All-Weather VLM: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions, accepted at the Conference on Language Modeling (COLM) 2026.

πŸ‘₯ Authors: Tianfu Wang1,* , Mingyang Xie1,* , Haoming Cai1, Tianyi Xiong1, Xiyao Wang1, Dongdong Fu2, Guan-Ming Su2, Paola Cascante-Bonilla1,3,†, Christopher Metzler1,†

1University of Maryland Β· 2Dolby Laboratories Inc. Β· 3Stony Brook University
*Equal contribution Β· †Equal advising

🌐 Project Website Β· πŸ“„ Paper Β· πŸ“š Citation

πŸ“– Model summary

All-Weather VLM makes Qwen2.5-VL-7B-Instruct robust to adverse imaging conditions: rain, fog, snow, haze, motion blur, low light, turbulence, and adherent raindrops. The degradation type does not need to be known in advance. Before answering, the model reasons inside <think> tags about the degradation type and its severity, then gives its answer. It is trained with a multi-stage degradation-aware fine-tuning recipe followed by Direct Preference Optimization (DPO), which treats responses from clean images as the more reliable reference for scene content.

This repository contains a LoRA adapter only. It requires the base model Qwen/Qwen2.5-VL-7B-Instruct, which also provides the tokenizer and image processor.

Base model Qwen/Qwen2.5-VL-7B-Instruct
Adapter type LoRA (PEFT)
Rank / alpha / dropout 256 / 512 / 0.05
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj (language model attention and MLP; vision-encoder MLP)
Adapter weights 584 tensors, float32

πŸ”Ž Output format

The model first writes a degradation diagnosis, then its answer:

<think>
1. Degradation: fog.
2. Severity: approximately 5 out of 10.
</think>
B

To keep only the answer, remove the <think>...</think> block. The code release does this automatically before scoring.

πŸš€ Usage

With the code release

Download the adapter:

huggingface-cli download Tianfwang/All-Weather-VLM --local-dir checkpoints/all-weather-vlm

Then run:

vlm-degradation infer --config examples/inference/single_image.yaml \
  --image path/to/image.png \
  --question 'Describe the scene.' \
  --output outputs/prediction.json

With Transformers and PEFT

Tested with transformers==4.50.0 and peft==0.17.0.

import re

import torch
from peft import PeftModel
from PIL import Image
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration

base = "Qwen/Qwen2.5-VL-7B-Instruct"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    base, torch_dtype=torch.bfloat16, device_map="cuda"
)
model = PeftModel.from_pretrained(model, "Tianfwang/All-Weather-VLM").eval()
processor = AutoProcessor.from_pretrained(base, min_pixels=256 * 28 * 28, max_pixels=3072 * 28 * 28)

system = (
    "You are an expert in analyzing degraded images. Upon receiving a new image, "
    "you must first reason inside <think> tags:\n"
    "1. Identify degradation type.\n2. Rate severity (1–10 from light to severe).\n"
    "For a clean image, use no degradation and severity 0. "
    "Output your response after </think>."
)
image = Image.open("image.png").convert("RGB")
messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "How many pedestrians are visible?"}]},
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
response = processor.batch_decode(output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]
print(response)                                                    # diagnosis + answer
print(re.sub(r"<think>.*?</think>", "", response, flags=re.DOTALL).strip())  # answer only

πŸ§ͺ Training

The recipe follows the paper.

  1. Degradation encoder: fine-tune the vision encoder to predict degradation type and severity.
  2. Degradation-conditioned VLM: fine-tune the language model to answer conditioned on degradation information.
  3. Unified degradation CoT: diagnose the degradation inside <think> tags, then answer, in one pass.
  4. DPO: preferred responses come from clean images and rejected responses from their degraded counterparts; the model always sees the degraded image.

Data. Clean image–question pairs from the LLaVA-NeXT training data. For supervised fine-tuning, each sampled pair is corrupted with a randomly selected degradation (fog, haze, rain, adherent raindrops, snow, low light, motion blur, or atmospheric turbulence) at a randomly sampled severity from 1 to 10; clean images are included as a no degradation class. For DPO, 20,000 caption pairs are generated from clean and degraded views of LLaVA-NeXT images with randomly sampled degradations.

Hyperparameters.

Stage Steps Batch size Learning rate Other
Stages 1–3 (SFT) 20,000 each 16 1e-6 LoRA on the vision encoder and language model
Stage 4 (DPO) 20,000 preference pairs 4 1e-6 Ξ² = 0.3

Training used 4Γ— NVIDIA RTX A6000 GPUs.

πŸ“Š Evaluation

Results reported in the paper:

Benchmark Qwen2.5-VL-7B All-Weather VLM
RealWorldQA, clean (accuracy %) 69.0 70.7
RealWorldQA, average over 8 degradations (accuracy %) 57.0 61.0
Seeing Through Fog (real adverse-weather driving scenes) β€” ~10% gain over base
R-Bench (real in-the-wild corruptions) β€” ~4% gain over base

⚠️ Limitations

  • Trained on synthetic degradations; real-world corruptions outside these eight families may not be diagnosed precisely.
  • Severity ratings are approximate and not calibrated measurements.
  • Inherits the limitations and biases of Qwen2.5-VL-7B-Instruct.

πŸ“š Citation

If you use this work, please cite:

@inproceedings{wang2026allweather,
  title     = {{All-Weather VLM}: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions},
  author    = {Wang, Tianfu and Xie, Mingyang and Cai, Haoming and Xiong, Tianyi and Wang, Xiyao and Fu, Dongdong and Su, Guan-Ming and Cascante-Bonilla, Paola and Metzler, Christopher},
  booktitle = {Conference on Language Modeling},
  year      = {2026},
  url       = {https://openreview.net/pdf?id=oHJGJS4rLD}
}

βš–οΈ License

The adapter is released under the Apache License 2.0. Use of the base model is governed by the Qwen2.5-VL-7B-Instruct license.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Tianfwang/All-Weather-VLM

Adapter
(361)
this model