Instructions to use Gabriel8495839/AgentforWindows with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Gabriel8495839/AgentforWindows with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Gabriel8495839/AgentforWindows") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Gabriel8495839/AgentforWindows") model = AutoModelForMultimodalLM.from_pretrained("Gabriel8495839/AgentforWindows", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Gabriel8495839/AgentforWindows with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Gabriel8495839/AgentforWindows" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Gabriel8495839/AgentforWindows", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Gabriel8495839/AgentforWindows
- SGLang
How to use Gabriel8495839/AgentforWindows with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Gabriel8495839/AgentforWindows" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Gabriel8495839/AgentforWindows", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Gabriel8495839/AgentforWindows" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Gabriel8495839/AgentforWindows", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Gabriel8495839/AgentforWindows with Docker Model Runner:
docker model run hf.co/Gabriel8495839/AgentforWindows
SmolVLM-500M-WindowsAgent โ sft_v1_step2800
Author: Gabriel Schwarzbauer
A small research model that operates a simulated Windows 11 desktop. It gets a screenshot, a task and its own previous actions, and answers with one mouse or keyboard action per step (for example click 14 977). It is a fine-tune of SmolVLM-500M-Instruct (~507M parameters, bf16, ~1 GB).
Read this first.
- This is an early research checkpoint. On the benchmark below it solves 35.6 % of the Level 0 tasks and 43.4 % of the Level 1 tasks. It fails more often than it succeeds and is not reliable.
- It was trained and evaluated only in a simulator of Windows 11 (a web re-creation built for this project). It has never been tested on a real Windows machine and is expected to transfer poorly.
- It has no safety mechanism. Do not connect it to a real computer with access to anything you care about.
- Not affiliated with, endorsed by or derived from Microsoft. "Windows" is a trademark of Microsoft.
What it does
Each step is one closed-loop call:
screenshot (1280x800) + "Task: โฆ" + own previous actions โโโบ model โโโบ one action line
The model sees only pixels (no accessibility tree, no element list). Example, task "Show the Start menu.":
click 14 977
done
Action space
The answer is exactly one line in this grammar (everything else counts as an invalid action):
| Action | Meaning |
|---|---|
click X Y / double_click X Y / right_click X Y |
mouse click at bin coordinates |
type "text" |
type text (\n = Enter, \" and \\ escaped) |
key <key> |
one key or a combination with +, e.g. key enter, key ctrl+s, key win |
scroll X Y N |
mouse wheel by N ticks at (X, Y); positive = down, range โ20โฆ20 |
drag X1 Y1 X2 Y2 |
press at (X1, Y1), release at (X2, Y2) |
wait |
do nothing for a moment |
done |
declare the task finished |
X, Y are integers 0โฆ999, independent of the real resolution: (0, 0) is the top-left corner, (999, 999) the bottom-right. Pixel = (bin + 0.5) / 1000 * width (or height).
The simulator additionally knows remember "text" (a note to self) and finished_with_error "text". This checkpoint was not trained on either and should not be expected to use them.
Not possible: middle/triple click, hover without clicking, holding keys, right-button drag.
Prompt format (must match exactly)
Task: <task text>
Actions so far: <"none" or "1) click 14 977; 2) type \"start\"; โฆ" (last 12 actions)>
Next action:
The image goes first in the user turn (SmolVLM chat template, see the quickstart). The screenshot must be 1280ร800 and the image processor must use longest_edge = 1536 (the processor_config.json in this repo is set accordingly; the model was trained with 7 tiles / 448 image tokens per screenshot). Decode greedily, max_new_tokens=28, stop at <end_of_utterance>, use the first line of the output.
Quickstart
import re, torch
from PIL import Image
from transformers import AutoProcessor, AutoModelForImageTextToText
repo = "PATH_OR_HF_REPO_ID" # this model
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
processor = AutoProcessor.from_pretrained(repo)
processor.image_processor.size = {"longest_edge": 1536} # as in training (already set in this repo's config)
processor.image_processor.max_image_size = {"longest_edge": 512}
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=dtype).to(device).eval()
eos = [processor.tokenizer.convert_tokens_to_ids("<end_of_utterance>"), processor.tokenizer.eos_token_id]
def next_action(screenshot: Image.Image, task: str, past_actions: list[str]) -> str:
screenshot = screenshot.convert("RGB").resize((1280, 800))
shown = past_actions[-12:]
start = len(past_actions) - len(shown)
history = "; ".join(f"{start + i + 1}) {a}" for i, a in enumerate(shown)) or "none"
prompt = f"Task: {task}\nActions so far: {history}\nNext action:"
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": prompt}]}]
text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(text=[text], images=[[screenshot]], return_tensors="pt").to(device)
inputs["pixel_values"] = inputs["pixel_values"].to(dtype)
with torch.inference_mode():
out = model.generate(**inputs, max_new_tokens=28, do_sample=False, eos_token_id=eos)
answer = processor.tokenizer.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
return answer.strip().split("\n")[0].strip()
def to_pixels(x_bin, y_bin, width, height):
return (x_bin + 0.5) / 1000 * width, (y_bin + 0.5) / 1000 * height
history = []
action = next_action(Image.open("screenshot.png"), "Show the Start menu.", history)
print(action) # e.g. click 14 977
m = re.match(r"click (\d+) (\d+)$", action)
if m:
print(to_pixels(int(m[1]), int(m[2]), 2560, 1600))
history.append(action)
The snippet was tested with transformers 5.17 / torch 2.14 on a GPU (CPU path untested); on older transformers versions use torch_dtype= instead of dtype=. It builds exactly the same input tokens and pixel values as the training pipeline.
You have to provide the environment yourself: take the screenshot, execute the action, append it to history, repeat until the model says done (or stop after a step budget). The simulator used for training is not part of this release.
Results
All numbers come from the project's own closed-loop benchmark in the simulator (one fresh simulator per episode, greedy decoding, base seed 777, 10 task variations per family). Success means all expected UI events happened in the right order; the score is the fraction of the reference event sequence that was reached (weighted longest common subsequence). The step budget is twice the length of the reference solution.
| Benchmark | step1600 (earlier checkpoint) |
step2800 (this model) |
|---|---|---|
| Level 0 โ 9 one-click shell tasks (Start, Run, tray, Win+X, โฆ) | 65.6 % (21/32), score 0.74 | 35.6 % (32/90), score 0.47 |
| Level 1 โ 32 of 75 tasks (Explorer, desktop, windows, Notepad, โฆ) | 28.1 % (9/32), score 0.64 | 43.4 % (139/320), score 0.81 |
95 % Wilson intervals: Level 0 26โ46 %, Level 1 38โ49 % (step2800); Level 0 48โ80 %, Level 1 16โ45 % (step1600).
How to read this:
- Level 0 got worse (65.6 % โ 35.6 %). The step1600 number is based on only 32 episodes, so the exact size of the drop is uncertain, but it is very likely real. Most Level 0 failures in the training-time evaluation were budget overruns (10 of 32 episodes at step2800 vs. 3 of 32 at step1800): the Level 0 budget is only 2ร a 1โ3 action reference, so a model that takes a different but valid route (e.g. searching for "start" instead of clicking the Start button) is scored as a failure. Why the model regressed is not established; the likely reason is that Level 1 dominates the training mix after step 1800, but this was not tested.
- Level 1 improved (28.1 % โ 43.4 %). Step1600 was not a zero-shot baseline: its training data already contained 3 demonstrations of every Level 1 family (see below). The step1600 numbers use 1 episode per family, the step2800 numbers 10, so the comparison is rough.
- The benchmark covers the first 32 of 75 Level 1 families in registry order, not all of them. The 43.4 % is therefore not a statement about all of Level 1.
- Not an out-of-distribution test. Benchmark tasks come from the same generators and value pools as the training demonstrations, only with different random seeds. Overlap in concrete task instances (file names, folders, wording) is possible and was not excluded.
- Level 2 (multi-step, two-subgoal tasks, not trained): 2/32 successes (score 0.56) in an in-training evaluation at step2800 (seed 7777, n = 32).
- Per-family results are not published for this release: the raw per-family output of the final 410-episode run was not saved. They will be added after a re-run.
- Action syntax validity was not re-measured for this release.
Training
| Base model | SmolVLM-500M-Instruct (SigLIP vision encoder + SmolLM2-360M language model) |
| Method | Behavior cloning / supervised fine-tuning; loss only on the answer tokens (the action line) |
| Trainable | LoRA r = 64, ฮฑ = 128, dropout 0.05 on the language model's attention and MLP projections (q, k, v, o, gate, up, down); the connector fully trained; vision encoder frozen |
| Release weights | LoRA merged into the base weights (merge_and_unload), bf16 |
| Optimizer | AdamW (ฮฒ = 0.9/0.99, no weight decay), peak LR 2e-4, 50 warm-up steps, cosine decay to 0.35ร peak, gradient clipping 1.0 |
| Batch | 16 (2 ร 8 gradient accumulation), 2800 optimizer steps |
| Precision / hardware | bf16, gradient checkpointing, one RTX 4060 Laptop GPU (8 GB), roughly 4โ6 optimizer steps/min |
| Data, steps 0โ1800 | demos_v0: 252 demonstrations (1,414 steps): 3 per task family for all 84 families โ 27 Level 0 episodes and 225 Level 1 episodes |
| Data, steps 1800โ2800 | roughly equal share of demos_v0 and demos_v1 (7,496 demonstrations, 44,900 steps, 75 families ร 100) |
Training data is fully synthetic. The demonstrations were produced by a scripted oracle that solves the tasks in the simulator; screenshots are rendered from the simulator (2560ร1600, downscaled to 1280ร800) and paired with the oracle's action. There are no human demonstrations, no data from real Windows sessions and no personal data. The dataset is not part of this release.
Curriculum / what is planned
The project trains in levels; only Levels 0 and 1 were used for this checkpoint.
| Level | Content | State |
|---|---|---|
| 0 | 9 one-click shell primitives | in this release |
| 1 | 75 families, 3โ12 actions, 8 apps (Explorer, Notepad, Calculator, Settings, Task Manager, Terminal, Run dialog, shell) | in this release |
| 2 | 45 families, two dependent subgoals, up to 22 actions | data exists, not trained in this checkpoint |
| 3 | 20 long multi-app tasks (20โ40 actions) that need the remember note |
task generators exist, no training yet |
| 4 | deliberately unsolvable tasks (finished_with_error) |
planned |
Training is being continued from this checkpoint with a larger adapter (rank 256, initialised from these weights so nothing is lost); that model will be a separate release.
Limitations
- Simulator only; one fixed layout: 2560ร1600 desktop at 150 % scaling, English UI, default light/dark themes of the simulator. No real Windows, no other resolutions, languages or DPI settings were tested.
- Fails most tasks that need more than a few steps; the fine coordinate targeting of small icons (tray, taskbar) is weak.
- Can loop or take unnecessary steps; no built-in detection of failure.
- Only the actions listed above; no browser, no web content, no text reading beyond what a 500M model can extract from a downscaled screenshot.
- English task descriptions only.
Intended use
Research and education on small GUI-agent models, and as a starting point for further fine-tuning. Not intended for automating real systems, and not evaluated for safety or misuse.
Citation
@misc{schwarzbauer2026smolvlmwindowsagent,
author = {Schwarzbauer, Gabriel},
title = {SmolVLM-500M-WindowsAgent (sft_v1_step2800)},
year = {2026},
note = {Fine-tune of SmolVLM-500M-Instruct for a simulated Windows 11 desktop}
}
Support
This is an independent one-person research project trained on a single laptop GPU. If you find it useful and want to support further work: ko-fi.com/windowsagentresearch. Donations do not buy support, features or any guarantee.
Acknowledgements
Built on SmolVLM by Hugging Face (Apache-2.0).
- Downloads last month
- 42
Model tree for Gabriel8495839/AgentforWindows
Base model
HuggingFaceTB/SmolLM2-360M