Instructions to use Tianfwang/All-Weather-VLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Tianfwang/All-Weather-VLM with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-VL-7B-Instruct") model = PeftModel.from_pretrained(base_model, "Tianfwang/All-Weather-VLM") - Notebooks
- Google Colab
- Kaggle
π¦οΈ [COLM 2026] All-Weather VLM
LoRA adapter for All-Weather VLM: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions, accepted at the Conference on Language Modeling (COLM) 2026.
π₯ Authors: Tianfu Wang1,* , Mingyang Xie1,* , Haoming Cai1, Tianyi Xiong1, Xiyao Wang1, Dongdong Fu2, Guan-Ming Su2, Paola Cascante-Bonilla1,3,β , Christopher Metzler1,β
1University of Maryland Β· 2Dolby Laboratories Inc. Β· 3Stony Brook University
*Equal contribution Β· β Equal advising
π Project Website Β· π Paper Β· π Citation
π Model summary
All-Weather VLM makes Qwen2.5-VL-7B-Instruct robust to adverse imaging conditions: rain, fog, snow, haze, motion blur, low light, turbulence, and adherent raindrops. The degradation type does not need to be known in advance. Before answering, the model reasons inside <think> tags about the degradation type and its severity, then gives its answer. It is trained with a multi-stage degradation-aware fine-tuning recipe followed by Direct Preference Optimization (DPO), which treats responses from clean images as the more reliable reference for scene content.
This repository contains a LoRA adapter only. It requires the base model Qwen/Qwen2.5-VL-7B-Instruct, which also provides the tokenizer and image processor.
| Base model | Qwen/Qwen2.5-VL-7B-Instruct |
| Adapter type | LoRA (PEFT) |
| Rank / alpha / dropout | 256 / 512 / 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj (language model attention and MLP; vision-encoder MLP) |
| Adapter weights | 584 tensors, float32 |
π Output format
The model first writes a degradation diagnosis, then its answer:
<think>
1. Degradation: fog.
2. Severity: approximately 5 out of 10.
</think>
B
To keep only the answer, remove the <think>...</think> block. The code release does this automatically before scoring.
π Usage
With the code release
Download the adapter:
huggingface-cli download Tianfwang/All-Weather-VLM --local-dir checkpoints/all-weather-vlm
Then run:
vlm-degradation infer --config examples/inference/single_image.yaml \
--image path/to/image.png \
--question 'Describe the scene.' \
--output outputs/prediction.json
With Transformers and PEFT
Tested with transformers==4.50.0 and peft==0.17.0.
import re
import torch
from peft import PeftModel
from PIL import Image
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
base = "Qwen/Qwen2.5-VL-7B-Instruct"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
base, torch_dtype=torch.bfloat16, device_map="cuda"
)
model = PeftModel.from_pretrained(model, "Tianfwang/All-Weather-VLM").eval()
processor = AutoProcessor.from_pretrained(base, min_pixels=256 * 28 * 28, max_pixels=3072 * 28 * 28)
system = (
"You are an expert in analyzing degraded images. Upon receiving a new image, "
"you must first reason inside <think> tags:\n"
"1. Identify degradation type.\n2. Rate severity (1β10 from light to severe).\n"
"For a clean image, use no degradation and severity 0. "
"Output your response after </think>."
)
image = Image.open("image.png").convert("RGB")
messages = [
{"role": "system", "content": system},
{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "How many pedestrians are visible?"}]},
]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
response = processor.batch_decode(output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]
print(response) # diagnosis + answer
print(re.sub(r"<think>.*?</think>", "", response, flags=re.DOTALL).strip()) # answer only
π§ͺ Training
The recipe follows the paper.
- Degradation encoder: fine-tune the vision encoder to predict degradation type and severity.
- Degradation-conditioned VLM: fine-tune the language model to answer conditioned on degradation information.
- Unified degradation CoT: diagnose the degradation inside
<think>tags, then answer, in one pass. - DPO: preferred responses come from clean images and rejected responses from their degraded counterparts; the model always sees the degraded image.
Data. Clean imageβquestion pairs from the LLaVA-NeXT training data. For supervised fine-tuning, each sampled pair is corrupted with a randomly selected degradation (fog, haze, rain, adherent raindrops, snow, low light, motion blur, or atmospheric turbulence) at a randomly sampled severity from 1 to 10; clean images are included as a no degradation class. For DPO, 20,000 caption pairs are generated from clean and degraded views of LLaVA-NeXT images with randomly sampled degradations.
Hyperparameters.
| Stage | Steps | Batch size | Learning rate | Other |
|---|---|---|---|---|
| Stages 1β3 (SFT) | 20,000 each | 16 | 1e-6 | LoRA on the vision encoder and language model |
| Stage 4 (DPO) | 20,000 preference pairs | 4 | 1e-6 | Ξ² = 0.3 |
Training used 4Γ NVIDIA RTX A6000 GPUs.
π Evaluation
Results reported in the paper:
| Benchmark | Qwen2.5-VL-7B | All-Weather VLM |
|---|---|---|
| RealWorldQA, clean (accuracy %) | 69.0 | 70.7 |
| RealWorldQA, average over 8 degradations (accuracy %) | 57.0 | 61.0 |
| Seeing Through Fog (real adverse-weather driving scenes) | β | ~10% gain over base |
| R-Bench (real in-the-wild corruptions) | β | ~4% gain over base |
β οΈ Limitations
- Trained on synthetic degradations; real-world corruptions outside these eight families may not be diagnosed precisely.
- Severity ratings are approximate and not calibrated measurements.
- Inherits the limitations and biases of Qwen2.5-VL-7B-Instruct.
π Citation
If you use this work, please cite:
@inproceedings{wang2026allweather,
title = {{All-Weather VLM}: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions},
author = {Wang, Tianfu and Xie, Mingyang and Cai, Haoming and Xiong, Tianyi and Wang, Xiyao and Fu, Dongdong and Su, Guan-Ming and Cascante-Bonilla, Paola and Metzler, Christopher},
booktitle = {Conference on Language Modeling},
year = {2026},
url = {https://openreview.net/pdf?id=oHJGJS4rLD}
}
βοΈ License
The adapter is released under the Apache License 2.0. Use of the base model is governed by the Qwen2.5-VL-7B-Instruct license.
- Downloads last month
- 10
Model tree for Tianfwang/All-Weather-VLM
Base model
Qwen/Qwen2.5-VL-7B-Instruct