S²VOPD-Qwen3.5-4B

Official model release for Self-Supervised Visual On-Policy Distillation.

Paper · Project page · Code

A fine-grained visual understanding model built on Qwen/Qwen3.5-4B and trained with on-policy self-distillation -- no reward model, no human annotation, and no larger teacher.

Results

Fine-grained perception

Greedy decoding, 4096 tokens, seed 42.

Benchmark Score
V* Bench 88.48
ZoomBench 56.57
HR-Bench 4K 84.75
HR-Bench 8K 84.62
MME-RealWorld 75.06
MME-RealWorld-CN 72.18
Mean 76.94

Mathematical reasoning

24576 tokens, T=0.3, top-p 0.95, top-k 20, presence penalty 1.5, seed 42; judge = Qwen2.5-72B-Instruct.

Benchmark Score
MathVista 80.80
MathVerse 73.83
MathVision 64.67
WeMath 85.98
Mean 76.32

The long generation budget is required for the math suite: at 16k, 8-11% of generations still hit the cap, so math numbers produced under a greedy/4096 protocol are not comparable to these.

Method

On-policy self-distillation. The student generates on a degraded view of an image while an EMA teacher scores the same prefix on the clean original. The disagreement between the two is the entire training signal -- it is large exactly on the tokens that depend on fine visual detail, and near zero elsewhere. The loss is a generalized JSD. No reward model, no human annotation, and no larger teacher.

Hyperparameter Value
Training steps 65
Batch / rollouts 96 prompts x n=8
Learning rate 2e-6, 10 warmup steps
EMA teacher update rate 0.05
Distillation generalized JSD, alpha=0.5, top-k=100 (renormalized inside the teacher's top-k), is_clip=2.0
Lengths prompt <= 8192, response <= 1024

Augmentation

Every training image receives exactly one degradation, drawn from three families at 30% each; the remaining 10% is left untouched.

Degradation Probability What it does
Non-uniform resolution 0.30 The image is laid on a canvas shrunk to 50-70% of its size; the background is rendered at 20-35% resolution while a single region covering 15-35% of the canvas stays sharp
Zoom-out 0.30 The image is scaled to 40-80% and padded with black
Downscale + photometric jitter 0.30 Downscaled to 30-60% without being resized back, then brightness, contrast and saturation each perturbed with probability 0.5
None 0.10 Unchanged

One degradation per image is deliberate. Stacking them sequentially destroys the image, and once the student can no longer recover the content the teacher-student gap stops carrying information about visual detail.

Training data

Vision-OPD-6K -- 6,241 samples.

Usage

from transformers import AutoProcessor, AutoModelForImageTextToText
import torch

model = AutoModelForImageTextToText.from_pretrained(
    "s2vopd/S2VOPD-Qwen3.5-4B", dtype=torch.bfloat16,
    attn_implementation="flash_attention_2", device_map="cuda")
processor = AutoProcessor.from_pretrained("s2vopd/S2VOPD-Qwen3.5-4B")

Citation

@misc{li2026selfsupervisedvisualonpolicydistillation,
  title={Self-Supervised Visual On-Policy Distillation},
  author={Yijiang Li and Yijun Liang and Yunjie Tian and Bingyang Wang and Ke Zhang and Zhenfei Yin and Di Fu and Philip Torr and Nuno Vasconcelos},
  year={2026},
  eprint={2608.14144},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2608.14144},
}
Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for s2vopd/S2VOPD-Qwen3.5-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(675)
this model

Paper for s2vopd/S2VOPD-Qwen3.5-4B