Instructions to use guanxuyu/visual-jev-browser-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use guanxuyu/visual-jev-browser-4b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/mnt/local_data/vjb/models/Qwen3-VL-4B-Instruct") model = PeftModel.from_pretrained(base_model, "guanxuyu/visual-jev-browser-4b") - Notebooks
- Google Colab
- Kaggle
Visual Jev β Browser Use (4B)
LoRA adapters for Qwen3-VL-4B-Instruct that act on a web page: read a
screenshot and a DOM element table once, answer each decision from that shared
prefill, and continue the same KV cache to write a field value when the chosen
operation needs one.
Code, scored outputs and the measurements behind every number: guanxuyu-sv/Visual-Jev-Browser-Use
This extends Visual Jev (paper) from answering many questions about one image to acting on a page, where the questions are which operation and which element, and the answers have to survive an executor.
What is here
Four arms of one ablation, plus a second training stage. They are only meaningful as a set: the comparison is the result.
| path | observation | output mechanism |
|---|---|---|
| (repository root) | screenshot + DOM | branch readout + conditional text on reused KV |
arm-a-dom-ar/ |
DOM only | compact autoregressive action line |
arm-b-screenshot-ar/ |
screenshot + DOM | compact autoregressive action line |
arm-c-dom-branch/ |
DOM only | branch readout + conditional text on reused KV |
terminal-trained/arm-{a,b,c,d}/ |
as above | as above, plus 400 steps of terminal-action supervision |
All four train on the same 3327 viewport-reduced Mind2Web steps with the same candidate lists and the same 2000-step budget. Only the modality and the output mechanism differ, which is what makes the differences attributable. LoRA r=16, alpha=32 on the language tower; vision tower frozen. 33.0M trainable of 4.47B (0.74%). One seed.
What they score
Held-out Mind2Web websites, n=234, never seen in training:
| arm | joint | target | CLICK | operation |
|---|---|---|---|---|
| A β DOM, AR | 54.7% | 58.5% | 43.4% | 94.0% |
| B β screenshot, AR | 88.9% | 92.3% | 88.4% | 95.7% |
| C β DOM, branch | 56.8% | 59.4% | 43.9% | 95.7% |
| D β screenshot, branch (root) | 91.0% | 94.4% | 90.2% | 96.6% |
MiniWoB++ closed loop, success decided by the environment's own reward, 224 tasks per arm: 38.4 / 45.5 / 43.8 / 49.1%.
Operation accuracy is 94β97% in every arm. The screenshot cannot move it and has no room to. The entire gain is target selection. Deciding what to do needs only the DOM text; deciding which element to do it to needs the picture.
The arms are not interchangeable
An adapter trained for the branch readout expects candidate symbols at an
Answer: position and is read from one logit; one trained for the compact line
expects to generate CLICK 4. Loading a branch adapter and then sampling text,
or the reverse, produces nonsense. vjb/model/policy.py in the code repository
selects the matching path by arm.
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from peft import PeftModel
base = "Qwen/Qwen3-VL-4B-Instruct"
model = Qwen3VLForConditionalGeneration.from_pretrained(base, dtype="bfloat16").to("cuda")
model = PeftModel.from_pretrained(model, "guanxuyu/visual-jev-browser-4b") # arm D
# model = PeftModel.from_pretrained(model, "guanxuyu/visual-jev-browser-4b",
# subfolder="arm-a-dom-ar") # arm A
processor = AutoProcessor.from_pretrained(base)
The observation format matters as much as the weights: the element table, the candidate symbols and the numbered screenshot overlay all have to be built the way training built them. Use the code repository rather than reconstructing the prompt.
Which to use
terminal-trained/ if you are running a closed loop. The base four were never
shown a terminal action β Mind2Web records one action per step and stops, so all
6454 of its steps are CLICK, TYPE_TEXT or SELECT and DONE never appears as a
label. The consequence is that 105 of 106 successful runs kept acting on a
finished page until the step budget ran out. The terminal-trained variants stop,
at a median of 4 steps and 4.5β5.3 s per successful task.
They carry their own problem. The terminal observations were collected only at the instant each task completed, when MiniWoB has already emptied the page, so what the model learned is closer to "the page looks empty" than "the goal is met". False DONE went from 0 to 21β62 per 216 tasks, which costs success rate. Negative examples from mid-trajectory states are the fix and are not in this release.
Limitations
One training seed. One backbone. No VisualWebArena β its site images are hosted where the training network could not reach them. The paired visual diagnostic, which carries the strongest claim in the accompanying analysis, has five pairs. The unified decision-and-generation mechanism these adapters implement does not improve accuracy over a fair autoregressive baseline: four of five comparisons span zero. Its measurable benefits are decision latency (145 ms against 351 ms) and field text (+9.1%, 95% CI [+2.1, +15.7]). That negative result is reported in full in the code repository rather than omitted here.
Page content is untrusted data. These adapters act on a page; an executor that runs their output should check legality itself, as the reference one does.
- Downloads last month
- 9
Model tree for guanxuyu/visual-jev-browser-4b
Base model
Qwen/Qwen3-VL-4B-Instruct