Instructions to use Ahsanz/e2-virl39k-3b-a4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ahsanz/e2-virl39k-3b-a4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Ahsanz/e2-virl39k-3b-a4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Ahsanz/e2-virl39k-3b-a4") model = AutoModelForMultimodalLM.from_pretrained("Ahsanz/e2-virl39k-3b-a4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Ahsanz/e2-virl39k-3b-a4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ahsanz/e2-virl39k-3b-a4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ahsanz/e2-virl39k-3b-a4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Ahsanz/e2-virl39k-3b-a4
- SGLang
How to use Ahsanz/e2-virl39k-3b-a4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ahsanz/e2-virl39k-3b-a4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ahsanz/e2-virl39k-3b-a4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ahsanz/e2-virl39k-3b-a4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ahsanz/e2-virl39k-3b-a4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Ahsanz/e2-virl39k-3b-a4 with Docker Model Runner:
docker model run hf.co/Ahsanz/e2-virl39k-3b-a4
A4: PRPO+RVD on ViRL39K from Qwen2.5-VL-3B-Instruct
PRPO with RVD (relative value decomposition).
One arm of a controlled study on what reinforcement learning does to a vision-language model's use of visual evidence. Every arm starts from the same base, sees the same data for the same number of steps, and differs only in the RL objective and whether the vision tower is trained. That is what makes the arms comparable; it also means a checkpoint on its own is not the result -- the finding lives in the differences between arms.
Training
| base model | Qwen/Qwen2.5-VL-3B-Instruct |
| objective | PRPO+RVD |
| vision tower | trained |
| data | ViRL39K, 31,629 training prompts after dedup (1,000 held out) |
| steps | 162 |
| seed | 1 |
| optimizer | AdamW, lr 1e-6 constant, no warmup |
| global batch / rollout batch | 128 / 384 |
| rollouts per prompt | 8 |
| rollout top-p | 0.99 |
| clip low / high | 0.2 / 0.28 (DAPO asymmetric) |
| KL penalty | none in the loss; KL-to-base logged as a readout |
| max response length | 2048 |
| precision | bf16 |
Trained with a fork of EasyR1 (veRL).
Use
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
model = Qwen2_5_VLForConditionalGeneration.from_pretrained("Ahsanz/e2-virl39k-3b-a4", dtype="bfloat16")
processor = AutoProcessor.from_pretrained("Ahsanz/e2-virl39k-3b-a4")
Greedy decoding was used for every evaluation in the study.
Intended use and limits
Research only, on the terms of the Qwen Research License. This is a 3B model trained for 162 RL steps on one dataset with one seed; it is an experimental artefact for studying RL's effect on grounding, not a model intended for deployment. Accuracy on general benchmarks was not a training target and is not reported here.
Evaluation of these arms found that RL changes how the model uses the image in ways that accuracy alone does not reveal, including a drop in binding a numeral to the element it labels. Treat outputs on figure-reading tasks with that in mind.
Licence
Inherited from the base model: Qwen Research License (research-only, non-commercial). The full
text ships beside the weights as LICENSE.
- Downloads last month
- 12
Model tree for Ahsanz/e2-virl39k-3b-a4
Base model
Qwen/Qwen2.5-VL-3B-Instruct