A4: PRPO+RVD on ViRL39K from Qwen2.5-VL-3B-Instruct

PRPO with RVD (relative value decomposition).

One arm of a controlled study on what reinforcement learning does to a vision-language model's use of visual evidence. Every arm starts from the same base, sees the same data for the same number of steps, and differs only in the RL objective and whether the vision tower is trained. That is what makes the arms comparable; it also means a checkpoint on its own is not the result -- the finding lives in the differences between arms.

Training

base model Qwen/Qwen2.5-VL-3B-Instruct
objective PRPO+RVD
vision tower trained
data ViRL39K, 31,629 training prompts after dedup (1,000 held out)
steps 162
seed 1
optimizer AdamW, lr 1e-6 constant, no warmup
global batch / rollout batch 128 / 384
rollouts per prompt 8
rollout top-p 0.99
clip low / high 0.2 / 0.28 (DAPO asymmetric)
KL penalty none in the loss; KL-to-base logged as a readout
max response length 2048
precision bf16

Trained with a fork of EasyR1 (veRL).

Use

from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration

model = Qwen2_5_VLForConditionalGeneration.from_pretrained("Ahsanz/e2-virl39k-3b-a4", dtype="bfloat16")
processor = AutoProcessor.from_pretrained("Ahsanz/e2-virl39k-3b-a4")

Greedy decoding was used for every evaluation in the study.

Intended use and limits

Research only, on the terms of the Qwen Research License. This is a 3B model trained for 162 RL steps on one dataset with one seed; it is an experimental artefact for studying RL's effect on grounding, not a model intended for deployment. Accuracy on general benchmarks was not a training target and is not reported here.

Evaluation of these arms found that RL changes how the model uses the image in ways that accuracy alone does not reveal, including a drop in binding a numeral to the element it labels. Treat outputs on figure-reading tasks with that in mind.

Licence

Inherited from the base model: Qwen Research License (research-only, non-commercial). The full text ships beside the weights as LICENSE.

Downloads last month
12
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ahsanz/e2-virl39k-3b-a4

Finetuned
(878)
this model