SmolVLM-500M-WindowsAgent โ€” sft_v1_step2800

Author: Gabriel Schwarzbauer

A small research model that operates a simulated Windows 11 desktop. It gets a screenshot, a task and its own previous actions, and answers with one mouse or keyboard action per step (for example click 14 977). It is a fine-tune of SmolVLM-500M-Instruct (~507M parameters, bf16, ~1 GB).

Read this first.

  • This is an early research checkpoint. On the benchmark below it solves 35.6 % of the Level 0 tasks and 43.4 % of the Level 1 tasks. It fails more often than it succeeds and is not reliable.
  • It was trained and evaluated only in a simulator of Windows 11 (a web re-creation built for this project). It has never been tested on a real Windows machine and is expected to transfer poorly.
  • It has no safety mechanism. Do not connect it to a real computer with access to anything you care about.
  • Not affiliated with, endorsed by or derived from Microsoft. "Windows" is a trademark of Microsoft.

What it does

Each step is one closed-loop call:

screenshot (1280x800) + "Task: โ€ฆ" + own previous actions  โ”€โ”€โ–บ  model  โ”€โ”€โ–บ  one action line

The model sees only pixels (no accessibility tree, no element list). Example, task "Show the Start menu.":

click 14 977
done

Action space

The answer is exactly one line in this grammar (everything else counts as an invalid action):

Action Meaning
click X Y / double_click X Y / right_click X Y mouse click at bin coordinates
type "text" type text (\n = Enter, \" and \\ escaped)
key <key> one key or a combination with +, e.g. key enter, key ctrl+s, key win
scroll X Y N mouse wheel by N ticks at (X, Y); positive = down, range โˆ’20โ€ฆ20
drag X1 Y1 X2 Y2 press at (X1, Y1), release at (X2, Y2)
wait do nothing for a moment
done declare the task finished

X, Y are integers 0โ€ฆ999, independent of the real resolution: (0, 0) is the top-left corner, (999, 999) the bottom-right. Pixel = (bin + 0.5) / 1000 * width (or height).

The simulator additionally knows remember "text" (a note to self) and finished_with_error "text". This checkpoint was not trained on either and should not be expected to use them.

Not possible: middle/triple click, hover without clicking, holding keys, right-button drag.

Prompt format (must match exactly)

Task: <task text>
Actions so far: <"none" or "1) click 14 977; 2) type \"start\"; โ€ฆ" (last 12 actions)>
Next action:

The image goes first in the user turn (SmolVLM chat template, see the quickstart). The screenshot must be 1280ร—800 and the image processor must use longest_edge = 1536 (the processor_config.json in this repo is set accordingly; the model was trained with 7 tiles / 448 image tokens per screenshot). Decode greedily, max_new_tokens=28, stop at <end_of_utterance>, use the first line of the output.

Quickstart

import re, torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText

repo = "PATH_OR_HF_REPO_ID"          # this model
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32

processor = AutoProcessor.from_pretrained(repo)
processor.image_processor.size = {"longest_edge": 1536}          # as in training (already set in this repo's config)
processor.image_processor.max_image_size = {"longest_edge": 512}
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=dtype).to(device).eval()

eos = [processor.tokenizer.convert_tokens_to_ids("<end_of_utterance>"), processor.tokenizer.eos_token_id]

def next_action(screenshot: Image.Image, task: str, past_actions: list[str]) -> str:
    screenshot = screenshot.convert("RGB").resize((1280, 800))
    shown = past_actions[-12:]
    start = len(past_actions) - len(shown)
    history = "; ".join(f"{start + i + 1}) {a}" for i, a in enumerate(shown)) or "none"
    prompt = f"Task: {task}\nActions so far: {history}\nNext action:"
    messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": prompt}]}]
    text = processor.apply_chat_template(messages, add_generation_prompt=True)
    inputs = processor(text=[text], images=[[screenshot]], return_tensors="pt").to(device)
    inputs["pixel_values"] = inputs["pixel_values"].to(dtype)
    with torch.inference_mode():
        out = model.generate(**inputs, max_new_tokens=28, do_sample=False, eos_token_id=eos)
    answer = processor.tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
    return answer.strip().split("\n")[0].strip()

def to_pixels(x_bin, y_bin, width, height):
    return (x_bin + 0.5) / 1000 * width, (y_bin + 0.5) / 1000 * height

history = []
action = next_action(Image.open("screenshot.png"), "Show the Start menu.", history)
print(action)                                   # e.g. click 14 977
m = re.match(r"click (\d+) (\d+)$", action)
if m:
    print(to_pixels(int(m[1]), int(m[2]), 2560, 1600))
history.append(action)

The snippet was tested with transformers 5.17 / torch 2.14 on a GPU (CPU path untested); on older transformers versions use torch_dtype= instead of dtype=. It builds exactly the same input tokens and pixel values as the training pipeline.

You have to provide the environment yourself: take the screenshot, execute the action, append it to history, repeat until the model says done (or stop after a step budget). The simulator used for training is not part of this release.

Results

All numbers come from the project's own closed-loop benchmark in the simulator (one fresh simulator per episode, greedy decoding, base seed 777, 10 task variations per family). Success means all expected UI events happened in the right order; the score is the fraction of the reference event sequence that was reached (weighted longest common subsequence). The step budget is twice the length of the reference solution.

Benchmark step1600 (earlier checkpoint) step2800 (this model)
Level 0 โ€“ 9 one-click shell tasks (Start, Run, tray, Win+X, โ€ฆ) 65.6 % (21/32), score 0.74 35.6 % (32/90), score 0.47
Level 1 โ€“ 32 of 75 tasks (Explorer, desktop, windows, Notepad, โ€ฆ) 28.1 % (9/32), score 0.64 43.4 % (139/320), score 0.81

95 % Wilson intervals: Level 0 26โ€“46 %, Level 1 38โ€“49 % (step2800); Level 0 48โ€“80 %, Level 1 16โ€“45 % (step1600).

How to read this:

  • Level 0 got worse (65.6 % โ†’ 35.6 %). The step1600 number is based on only 32 episodes, so the exact size of the drop is uncertain, but it is very likely real. Most Level 0 failures in the training-time evaluation were budget overruns (10 of 32 episodes at step2800 vs. 3 of 32 at step1800): the Level 0 budget is only 2ร— a 1โ€“3 action reference, so a model that takes a different but valid route (e.g. searching for "start" instead of clicking the Start button) is scored as a failure. Why the model regressed is not established; the likely reason is that Level 1 dominates the training mix after step 1800, but this was not tested.
  • Level 1 improved (28.1 % โ†’ 43.4 %). Step1600 was not a zero-shot baseline: its training data already contained 3 demonstrations of every Level 1 family (see below). The step1600 numbers use 1 episode per family, the step2800 numbers 10, so the comparison is rough.
  • The benchmark covers the first 32 of 75 Level 1 families in registry order, not all of them. The 43.4 % is therefore not a statement about all of Level 1.
  • Not an out-of-distribution test. Benchmark tasks come from the same generators and value pools as the training demonstrations, only with different random seeds. Overlap in concrete task instances (file names, folders, wording) is possible and was not excluded.
  • Level 2 (multi-step, two-subgoal tasks, not trained): 2/32 successes (score 0.56) in an in-training evaluation at step2800 (seed 7777, n = 32).
  • Per-family results are not published for this release: the raw per-family output of the final 410-episode run was not saved. They will be added after a re-run.
  • Action syntax validity was not re-measured for this release.

Training

Base model SmolVLM-500M-Instruct (SigLIP vision encoder + SmolLM2-360M language model)
Method Behavior cloning / supervised fine-tuning; loss only on the answer tokens (the action line)
Trainable LoRA r = 64, ฮฑ = 128, dropout 0.05 on the language model's attention and MLP projections (q, k, v, o, gate, up, down); the connector fully trained; vision encoder frozen
Release weights LoRA merged into the base weights (merge_and_unload), bf16
Optimizer AdamW (ฮฒ = 0.9/0.99, no weight decay), peak LR 2e-4, 50 warm-up steps, cosine decay to 0.35ร— peak, gradient clipping 1.0
Batch 16 (2 ร— 8 gradient accumulation), 2800 optimizer steps
Precision / hardware bf16, gradient checkpointing, one RTX 4060 Laptop GPU (8 GB), roughly 4โ€“6 optimizer steps/min
Data, steps 0โ€“1800 demos_v0: 252 demonstrations (1,414 steps): 3 per task family for all 84 families โ€“ 27 Level 0 episodes and 225 Level 1 episodes
Data, steps 1800โ€“2800 roughly equal share of demos_v0 and demos_v1 (7,496 demonstrations, 44,900 steps, 75 families ร— 100)

Training data is fully synthetic. The demonstrations were produced by a scripted oracle that solves the tasks in the simulator; screenshots are rendered from the simulator (2560ร—1600, downscaled to 1280ร—800) and paired with the oracle's action. There are no human demonstrations, no data from real Windows sessions and no personal data. The dataset is not part of this release.

Curriculum / what is planned

The project trains in levels; only Levels 0 and 1 were used for this checkpoint.

Level Content State
0 9 one-click shell primitives in this release
1 75 families, 3โ€“12 actions, 8 apps (Explorer, Notepad, Calculator, Settings, Task Manager, Terminal, Run dialog, shell) in this release
2 45 families, two dependent subgoals, up to 22 actions data exists, not trained in this checkpoint
3 20 long multi-app tasks (20โ€“40 actions) that need the remember note task generators exist, no training yet
4 deliberately unsolvable tasks (finished_with_error) planned

Training is being continued from this checkpoint with a larger adapter (rank 256, initialised from these weights so nothing is lost); that model will be a separate release.

Limitations

  • Simulator only; one fixed layout: 2560ร—1600 desktop at 150 % scaling, English UI, default light/dark themes of the simulator. No real Windows, no other resolutions, languages or DPI settings were tested.
  • Fails most tasks that need more than a few steps; the fine coordinate targeting of small icons (tray, taskbar) is weak.
  • Can loop or take unnecessary steps; no built-in detection of failure.
  • Only the actions listed above; no browser, no web content, no text reading beyond what a 500M model can extract from a downscaled screenshot.
  • English task descriptions only.

Intended use

Research and education on small GUI-agent models, and as a starting point for further fine-tuning. Not intended for automating real systems, and not evaluated for safety or misuse.

Citation

@misc{schwarzbauer2026smolvlmwindowsagent,
  author = {Schwarzbauer, Gabriel},
  title  = {SmolVLM-500M-WindowsAgent (sft_v1_step2800)},
  year   = {2026},
  note   = {Fine-tune of SmolVLM-500M-Instruct for a simulated Windows 11 desktop}
}

Support

This is an independent one-person research project trained on a single laptop GPU. If you find it useful and want to support further work: ko-fi.com/windowsagentresearch. Donations do not buy support, features or any guarantee.

Acknowledgements

Built on SmolVLM by Hugging Face (Apache-2.0).

Downloads last month
42
Safetensors
Model size
0.5B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Gabriel8495839/AgentforWindows