diffcider-browser

A typed-decision model made with sysone: the encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 (lora) under a 0-layer yesno head, trained on osunlp/Multimodal-Mind2Web at 1b4c6a8c.

The adapter makes the masked diffusion language model Qwen3-0.6B-diffusion-mdlm-v0.1 answer a browser agent's questions, in the way of laya-ultrafast: at each step, which operation comes next (CLICK, TYPE_TEXT, SELECT, or a control such as DONE) and, for an operation, which element of the page it acts on, in one forward pass, reading the model's own Yes against its No at a mask after each option. When the agent types, the same model writes the text with the adapter switched off (dec.generate). It trained on 1,500 steps of Multimodal-Mind2Web's train split, read without the screenshots, with every action offered on every step, so that the actions a step offers never give its answer away, and with the rare operations oversampled to equal shares, since 85% of the steps are clicks. On 300 steps from 9 websites it never saw, each offering the actions its candidate elements allow, it picks the right operation 85% of the time, where always clicking is right 78% of the time and choosing the rarest action offered 80%; the right element among up to 45 65% of the time (25% untrained); and the whole step, the typed text included, 47% of the time. Offered every action on every step, it leans toward typing and selecting (57% right), having trained on the operations in equal shares: offer it the actions the page's elements allow. With the adapter off, on 120 typing steps whose value the task's wording contains, the model writes exactly what the person typed 49% of the time (8 unmasking steps, 372 ms on a T4). It has never seen a DONE page: Mind2Web records only the steps people took. An earlier version, in this repo's history, trained on the steps as they come and always answered CLICK. sysone's mdlm module loads the checkpoint's a2d-qwen3 model type with its own classes, so no remote code runs.

The repo holds the head and a LoRA adapter, not the encoder's weights: Decider.load downloads dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1 at commit c8d24a3f4a from the Hub and puts the adapter on it.

Use it

Install sysone from GitHub, in Python 3.11 or newer:

pip install git+https://github.com/sgaseretto/sysonelib

A state is whatever the decision is about, as JSON-like data. Each question has a type: a choice among named options, a score on an ordered scale, or a noul, a statement that is true or false. The options are named when asking, so they can be new ones.

from sysone.inference import Decider

dec = Decider.load("sgaseretto/diffcider-browser")

from sysone.datasets import CONTROLS, NEXT_ACTION, OPERATIONS, TARGET

goal = "Find one-way flights from New York to Chicago on April 12 for one adult."
state = {
    "page": {"url": "https://united.com/", "title": "united",
             "text": "Book Flight status Check-in My trips Roundtrip One-way From New York To Depart Travelers 1 Adult Find flights"},
    "recent_actions": [{"action": "[radio]  One-way", "kind": "click", "text": None, "page_changed": True},
                       {"action": "[textbox]  From", "kind": "fill", "text": "New York", "page_changed": False},
                       {"action": "[textbox]  To", "kind": "click", "text": None, "page_changed": False}],
}
fields = {"1": "[1] From (textbox) = 'New York'", "2": "[2] To (textbox)", "3": "[3] Depart (textbox)"}
buttons = {"4": "[4] Roundtrip (radio) checked=False", "5": "[5] One-way (radio) checked=True", "6": "[6] Find flights (button)"}
def target(op, elements): return {"type": "choice", "instructions": {"goal": goal, "operation": op, "rules": [NEXT_ACTION, TARGET]}, "criteria": elements}
questions = {
    "operation": {"type": "choice", "instructions": {"goal": goal, "rules": NEXT_ACTION},
                  "criteria": {k: OPERATIONS[k] for k in ("CLICK", "TYPE_TEXT")} | CONTROLS},
    "type_text_target": target("TYPE_TEXT", fields),
    "click_target": target("CLICK", buttons),
}
answers = dec.predict(state, questions)     # System One: the next operation, and an element for each, in one pass with the adapter
# on this made-up page it hesitates, CLICK 0.51 against TYPE_TEXT 0.49, and would type into To (0.87)
field = fields[answers["type_text_target"]["choice"]].split("] ", 1)[1].split(" = ")[0]
text = dec.generate(f"The user's goal: {goal}\nThe page: {state['page']['title']}\nThe field: {field}\n"
                    "What should be typed into this field?", max_new_tokens=16, steps=8,
                    system="You fill in web forms for a user. Answer with the exact text to type, nothing else.")
                                            # System Two: the text for that field, written by the same model without the adapter
print(text)                                 # Chicago

answers holds an answer per question, in the Jev schema: a choice names the likeliest option and gives every option's probability, a score gives its expected level on the scale (0 for the first) and every level's probability, and a noul gives the probability that its statement is true; each says how confident it is. This model's answers to the example:

{
    "operation": {
        "type": "choice",
        "choice": "CLICK",
        "probabilities": {
            "CLICK": 0.5107,
            "TYPE_TEXT": 0.4856,
            "WAIT": 0.0009,
            "SCROLL_DOWN": 0.0007,
            "SCROLL_UP": 0.0007,
            "DONE": 0.0007,
            "BLOCKED": 0.0007,
        },
        "confidence": 0.6297,
        "answer_confidence": 0.5107,
    },
    "type_text_target": {
        "type": "choice",
        "choice": "2",
        "probabilities": {"1": 0.0555, "2": 0.8736, "3": 0.0709},
        "confidence": 0.5758,
        "answer_confidence": 0.8736,
    },
    "click_target": {
        "type": "choice",
        "choice": "5",
        "probabilities": {"4": 0.0165, "5": 0.6087, "6": 0.3748},
        "confidence": 0.3285,
        "answer_confidence": 0.6087,
    },
}

Results

On osunlp/Multimodal-Mind2Web's test_website split (600 decisions), its first 300 steps, from 9 websites none of the training steps came from, two decisions each, each step offering the actions its candidate elements allow (the last metric: every action offered); calibrated, on a Kaggle T4:

metric value
accuracy_operation 0.8533
accuracy_element 0.6467
ece_operation 0.0542
ece_element 0.0826
accuracy_step 0.47
accuracy_operation_every_action_offered 0.5733

Calibration temperatures: choice 1.3, choice:11+ 1.27, choice:2 5, choice:6-10 1.36.

Training

Data osunlp/Multimodal-Mind2Web at 1b4c6a8c; 1255 train, 89 valid, 156 calib cases
Encoder dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1, lora (LoRA r=16, α=32, on down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj)
Head 0 layers, yesno scorer, readout anchor
Loss soft_ce_rps=0.25, preset t4, seed 0
Fit 1 1 epoch, cosine, lr 0.001 (encoder 0.0002), batch 4×4, fp16 on cuda: 150 steps in 89.4 min, peak memory 7.8 GB, final loss 1.129
Machine Intel(R) Xeon(R) CPU @ 2.00GHz, 4 cores, 31.3 GB; accelerator cuda (Tesla T4, Tesla T4); a kaggle run (job mdlm-browser-e4)

Reproduce

  • sysone: sysonelib at commit 123ad8bac6 (main), with uncommitted changes in 30 files (diff SHA-256 b81003e7b8e8)
  • Ran: run_f.py
  • Python 3.13.15, sysone 0.3.0, torch 2.11.0+cu128, transformers 5.16.1, peft 0.20.0, accelerate 1.14.0, datasets 4.8.5, huggingface_hub 1.29.0, tokenizers 0.23.1, safetensors 0.8.0, numpy 2.1.3, fastcore 2.2.32, plum-dispatch 2.10.1; environment.txt lists every package
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sgaseretto/diffcider-browser

Finetuned
Qwen/Qwen3-0.6B
Adapter
(5)
this model

Dataset used to train sgaseretto/diffcider-browser

Evaluation results