Vision-OPD Qwen3.5-4B Reproduction

This repository contains an unofficial reproduction of Vision-OPD based on Qwen/Qwen3.5-4B. It is not an official model release from the Vision-OPD authors.

Vision-OPD applies regional-to-global on-policy self-distillation to improve fine-grained visual perception. This checkpoint was produced with the public Vision-OPD implementation and the public yuanqianhao/Vision-OPD-6K training set.

Training

Setting Value
Base model Qwen/Qwen3.5-4B
Training data yuanqianhao/Vision-OPD-6K (6,241 samples)
Epochs 1
Final step 65
Batch size 96
Rollouts per prompt 8
Learning rate 2e-6
Hardware 8 x NVIDIA B200

The uploaded artifact is the merged Transformers checkpoint, not an FSDP training shard.

Evaluation

Generation used non-thinking mode with a maximum output length of 32,768 tokens. The answer judge was openai/gpt-oss-120b, served locally with a 65,536-token context window and a 2,048-token judge output limit.

Benchmark Qwen3.5-4B baseline This reproduction Paper Vision-OPD-4B
V* 84.29 90.05 92.15
ZoomBench 47.69 59.64 59.76
HRBench-4K 84.38 81.75 84.50
HRBench-8K 80.13 80.00 80.38
MME-RealWorld-EN 63.86 71.96 74.88
MME-RealWorld-CN 63.70 69.56 70.76
Average 70.68 75.49 77.07

The reproduction improves the six-benchmark average by 4.81 points over the reported Qwen3.5-4B baseline and remains 1.58 points below the paper result.

Usage

Use a recent Transformers release with Qwen3.5 support (transformers>=5.5.0).

from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "Hugo0713/vision-opd-qwen3.5-4b-reproduction"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

The model can also be served with vLLM:

vllm serve Hugo0713/vision-opd-qwen3.5-4b-reproduction \
  --served-model-name Vision-OPD-4B \
  --gpu-memory-utilization 0.85

Limitations

This is a single reproduction run. The reported results depend on the local inference and judge configuration described above and should not be treated as an official Vision-OPD release.

Citation

@article{yuan2026vision,
  title={Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation},
  author={Yuan, Qianhao and Lou, Jie and Yu, Xing and Lin, Hongyu and Sun, Le and Han, Xianpei and Lu, Yaojie},
  journal={arXiv preprint arXiv:2605.18740},
  year={2026}
}

License

Apache-2.0. See LICENSE for details.

Downloads last month
10
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Hugo0713/vision-opd-qwen3.5-4b-reproduction

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(536)
this model

Dataset used to train Hugo0713/vision-opd-qwen3.5-4b-reproduction

Paper for Hugo0713/vision-opd-qwen3.5-4b-reproduction