Visual Jev β€” Browser Use (4B)

LoRA adapters for Qwen3-VL-4B-Instruct that act on a web page: read a screenshot and a DOM element table once, answer each decision from that shared prefill, and continue the same KV cache to write a field value when the chosen operation needs one.

Code, scored outputs and the measurements behind every number: guanxuyu-sv/Visual-Jev-Browser-Use

This extends Visual Jev (paper) from answering many questions about one image to acting on a page, where the questions are which operation and which element, and the answers have to survive an executor.

What is here

Four arms of one ablation, plus a second training stage. They are only meaningful as a set: the comparison is the result.

path observation output mechanism
(repository root) screenshot + DOM branch readout + conditional text on reused KV
arm-a-dom-ar/ DOM only compact autoregressive action line
arm-b-screenshot-ar/ screenshot + DOM compact autoregressive action line
arm-c-dom-branch/ DOM only branch readout + conditional text on reused KV
terminal-trained/arm-{a,b,c,d}/ as above as above, plus 400 steps of terminal-action supervision

All four train on the same 3327 viewport-reduced Mind2Web steps with the same candidate lists and the same 2000-step budget. Only the modality and the output mechanism differ, which is what makes the differences attributable. LoRA r=16, alpha=32 on the language tower; vision tower frozen. 33.0M trainable of 4.47B (0.74%). One seed.

What they score

Held-out Mind2Web websites, n=234, never seen in training:

arm joint target CLICK operation
A β€” DOM, AR 54.7% 58.5% 43.4% 94.0%
B β€” screenshot, AR 88.9% 92.3% 88.4% 95.7%
C β€” DOM, branch 56.8% 59.4% 43.9% 95.7%
D β€” screenshot, branch (root) 91.0% 94.4% 90.2% 96.6%

MiniWoB++ closed loop, success decided by the environment's own reward, 224 tasks per arm: 38.4 / 45.5 / 43.8 / 49.1%.

Operation accuracy is 94–97% in every arm. The screenshot cannot move it and has no room to. The entire gain is target selection. Deciding what to do needs only the DOM text; deciding which element to do it to needs the picture.

The arms are not interchangeable

An adapter trained for the branch readout expects candidate symbols at an Answer: position and is read from one logit; one trained for the compact line expects to generate CLICK 4. Loading a branch adapter and then sampling text, or the reverse, produces nonsense. vjb/model/policy.py in the code repository selects the matching path by arm.

from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from peft import PeftModel

base = "Qwen/Qwen3-VL-4B-Instruct"
model = Qwen3VLForConditionalGeneration.from_pretrained(base, dtype="bfloat16").to("cuda")
model = PeftModel.from_pretrained(model, "guanxuyu/visual-jev-browser-4b")          # arm D
# model = PeftModel.from_pretrained(model, "guanxuyu/visual-jev-browser-4b",
#                                   subfolder="arm-a-dom-ar")                        # arm A
processor = AutoProcessor.from_pretrained(base)

The observation format matters as much as the weights: the element table, the candidate symbols and the numbered screenshot overlay all have to be built the way training built them. Use the code repository rather than reconstructing the prompt.

Which to use

terminal-trained/ if you are running a closed loop. The base four were never shown a terminal action β€” Mind2Web records one action per step and stops, so all 6454 of its steps are CLICK, TYPE_TEXT or SELECT and DONE never appears as a label. The consequence is that 105 of 106 successful runs kept acting on a finished page until the step budget ran out. The terminal-trained variants stop, at a median of 4 steps and 4.5–5.3 s per successful task.

They carry their own problem. The terminal observations were collected only at the instant each task completed, when MiniWoB has already emptied the page, so what the model learned is closer to "the page looks empty" than "the goal is met". False DONE went from 0 to 21–62 per 216 tasks, which costs success rate. Negative examples from mid-trajectory states are the fix and are not in this release.

Limitations

One training seed. One backbone. No VisualWebArena β€” its site images are hosted where the training network could not reach them. The paired visual diagnostic, which carries the strongest claim in the accompanying analysis, has five pairs. The unified decision-and-generation mechanism these adapters implement does not improve accuracy over a fair autoregressive baseline: four of five comparisons span zero. Its measurable benefits are decision latency (145 ms against 351 ms) and field text (+9.1%, 95% CI [+2.1, +15.7]). That negative result is reported in full in the code repository rather than omitted here.

Page content is untrusted data. These adapters act on a page; an executor that runs their output should check legality itself, as the reference one does.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for guanxuyu/visual-jev-browser-4b

Adapter
(189)
this model

Paper for guanxuyu/visual-jev-browser-4b