Instructions to use infinitylogesh/GUI-Decisions-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use infinitylogesh/GUI-Decisions-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="infinitylogesh/GUI-Decisions-2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("infinitylogesh/GUI-Decisions-2B") model = AutoModelForMultimodalLM.from_pretrained("infinitylogesh/GUI-Decisions-2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use infinitylogesh/GUI-Decisions-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "infinitylogesh/GUI-Decisions-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/infinitylogesh/GUI-Decisions-2B
- SGLang
How to use infinitylogesh/GUI-Decisions-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "infinitylogesh/GUI-Decisions-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "infinitylogesh/GUI-Decisions-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use infinitylogesh/GUI-Decisions-2B with Docker Model Runner:
docker model run hf.co/infinitylogesh/GUI-Decisions-2B
GUI-Decisions-2B
Slot reading with the 31B model (from the blog post): reading slots (left, 158 ms per decision) against writing the step as JSON (right, 520 ms). Try the 31B model in the Space.
mPLUG/GUI-Owl-1.5-2B-Instruct with the slot-reading LoRA merged in (bf16), which makes a computer-use step readable from one forward pass instead of decoded: the action, the point (256 bins per axis), a swipe's end, the key and the scroll, as next-token distributions at fixed slots. The small-model result of the blog post Stop Decoding Coordinates: 3× Faster Computer-Use Grounding. The LoRA alone: GUI-Decisions-2B-LoRA; the 31B version: GUI-Decisions-31B.
Results
Accuracy (blog post):
| computer-use step (held-out AGUVIS, 1,600) | ScreenSpot clicks (1,272) | |
|---|---|---|
| GUI-Owl-2B + the LoRA | 0.626 | 0.770 |
| GUI-Owl-2B (its own tool-call format), native resolution | 0.626 | 0.645 |
| GUI-Owl-2B, same 1 MP images | 0.610 | 0.626 |
Vanilla GUI-Owl's click numbers use its agent prompt, not its dedicated grounding prompt.
This merged model, read in-process with transformers (bf16): ScreenSpot clicks 0.763 against the published LoRA's 0.770 on the same cases (paired difference −0.006 [−0.013, +0.001], not significant; the unmerged LoRA read the same way scores 0.764), and computer-use steps 0.627 on the held-out AGUVIS set.
Latency of one computer-use step, measured with the LoRA on the base model (this merged model reads at
the same speed: ~83 against ~87 ms end to end): vLLM 0.30, one RTX PRO 6000, cold screenshots, one
request at a time, all conditions in one session (scripts/gui_owl.py bench 80):
| how the step is decided | median (p90) |
|---|---|
slots, lazy: action, x, y, then only the slots the action needs (what decide does) |
65 ms (71) |
| slots, the whole template (8 slots) | 93 ms (104) |
| GUI-Owl-2B writing its tool call (~50 tokens), same 1 MP images | 194 ms (217) |
| GUI-Owl-2B writing its tool call (~50 tokens), native resolution | 323 ms (391) |
The blog's earlier session measured 133, 225 and 444 ms for the last three rows on another box; the
ratios agree (whole-template slots at ~30% of native-resolution GUI-Owl). End to end through decide,
including the client's JPEG encoding, a step takes ~85 ms.
Use
With vLLM (recommended)
Served on one RTX PRO 6000 a computer-use step takes ~83 ms end to end through the code below (median over 60 cold screenshots, client JPEG encoding included; the LoRA on the base model: ~87 ms).
1. Install vLLM 0.30 and the reader (slotread, from the code repository):
git clone https://github.com/infinitylogesh/gui-decisions && cd gui-decisions
pip install "vllm==0.30.*" -e .
bash serving/patch_vllm.sh # recommended: 256 label ids per request
2. Serve:
vllm serve infinitylogesh/GUI-Decisions-2B --served-model-name gui-decisions-2b --port 8001 \
--max-model-len 16384 --gpu-memory-utilization 0.6 --max-logprobs 300 --enable-prefix-caching \
--mm-processor-cache-gb 0 --api-server-count 8 --limit-mm-per-prompt '{"image":1}'
(MODEL=infinitylogesh/GUI-Decisions-2B NAME=gui-decisions-2b MAX_LEN=16384 GPU_UTIL=0.6 EXTRA_ARGS='--limit-mm-per-prompt {"image":1}' bash serving/serve_vllm.sh runs the same command.)
3. Read steps:
from PIL import Image
from slotread.client import SlotModel
m = SlotModel("gui-decisions-2b", "infinitylogesh/GUI-Decisions-2B") # server: VLLM_URL, default http://127.0.0.1:8001
step = m.decide(Image.open("screenshot.png"), "open the settings") # {'action': 'click', 'x': ..., 'y': ..., 'ms': ...}
decide reads action, x and y first, then only the slots the action needs. Screenshots are
downscaled to at most 1,048,576 pixels, as in training (image_cap in slotread_config.json; SlotModel
applies it).
What a read is. For each slot, SlotModel sends one chat request whose assistant turn is prefilled up
to that slot, and asks vLLM for the log-probabilities of exactly that slot's labels; nothing is generated
beyond one token. The requests for one step run in parallel and share the image through vLLM's prefix cache:
{"model": "gui-decisions-2b",
"messages": [{"role": "system", "content": [{"type": "text", "text": "<the task's system prompt, from slotread_config.json>"}]},
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
{"type": "text", "text": "Please generate the next move according to the UI screenshot, instruction and previous actions.\n\nInstruction: turn on dark mode\n\nPrevious actions:\nNone"}]},
{"role": "assistant", "content": [{"type": "text", "text": "action: ___ x:"}]}],
"add_generation_prompt": false, "continue_final_message": true, "chat_template_kwargs": {"enable_thinking": false},
"max_tokens": 1, "temperature": 0, "logprobs": true, "top_logprobs": 1, "return_tokens_as_token_ids": true,
"logprob_token_ids": ["<the 256 label ids of slot x, from slotread_config.json>"]}
The slot's answer is the softmax over those labels: the argmax for action; for a coordinate, the argmax
bin refined by its ±2 neighbours (bin k of 256 is the value (k + 0.5) / 256 across the screen).
Notes:
patch_vllm.shraises vLLM's per-requestlogprob_token_idscap from 128 to 256, so a 256-bin slot is one request. Without it, setMAX_LABEL_IDS=128for the client (each coordinate then takes two requests).--api-server-count 8: a single API-server process serializes a step's parallel slot requests.--mm-processor-cache-gb 0: vLLM 0.30's media cache raced on concurrent requests for the same image.previoustakes the earlier steps as numbered lines, e.g."1. click(One way)\n2. click(From)".
In-process with transformers (no server)
from PIL import Image
from slotread.local import LocalSlotModel
m = LocalSlotModel.from_pretrained("infinitylogesh/GUI-Decisions-2B")
print(m.decide(Image.open("screenshot.png"), "open the settings"))
Training
LoRA r 16, alpha 32, on the language model's attention and MLP projections; AGUVIS computer-use steps (5,000) and Wave-UI element clicks (2,000); one RTX PRO 6000, about 40 minutes.
Limits
One step at a time, no memory and no acting; computer use only (no boxes or masks). Text arguments (what to type, the final answer) need another writer: asked with our plain prompt, base GUI-Owl answered a city to type with the field's placeholder text. Trained on AGUVIS and Wave-UI; check their terms for your use.
Citation
@article{umapathi2026slotreading,
title = "Stop Decoding Coordinates: 3.5× Faster Computer-Use Grounding",
author = "Umapathi, Logesh Kumar",
journal = "logeshumapathi.com",
year = "2026",
month = "Oct",
url = "https://logeshumapathi.com/blog/2026/10/06/slot-reading.html"
}
- Downloads last month
- 37
