Instructions to use infinitylogesh/GUI-Decisions-31B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use infinitylogesh/GUI-Decisions-31B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="infinitylogesh/GUI-Decisions-31B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("infinitylogesh/GUI-Decisions-31B") model = AutoModelForMultimodalLM.from_pretrained("infinitylogesh/GUI-Decisions-31B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use infinitylogesh/GUI-Decisions-31B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "infinitylogesh/GUI-Decisions-31B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/infinitylogesh/GUI-Decisions-31B
- SGLang
How to use infinitylogesh/GUI-Decisions-31B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "infinitylogesh/GUI-Decisions-31B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "infinitylogesh/GUI-Decisions-31B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "infinitylogesh/GUI-Decisions-31B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use infinitylogesh/GUI-Decisions-31B with Docker Model Runner:
docker model run hf.co/infinitylogesh/GUI-Decisions-31B
GUI-Decisions-31B
Same model, same screenshots: reading slots (left, 158 ms per decision) against writing the step as JSON (right, 520 ms). Try it in the Space.
google/gemma-4-31B-it with the slot-reading LoRA merged in (bf16), which makes computer-use decisions readable from one forward pass instead of decoded: the action, the point (256 bins per axis), a swipe's end, the key and the scroll, all as next-token distributions at fixed slots. It also reads boxes (centre + log size) and coarse masks (centre + 24 rays) for general objects.
From the blog post Stop Decoding Coordinates: 3× Faster Computer-Use Grounding (run E3). The LoRA alone: GUI-Decisions-31B-LoRA. Try it: GUI-Decisions Space.
How it works
The assistant turn is a fixed template; every slot ends in a placeholder ( ___):
action: ___ x: ___ y: ___ x2: ___ y2: ___ key: ___ sdir: ___ samt: ___
A slot's answer is the next-token distribution at the token before its placeholder, restricted to the
slot's labels (slotread_config.json lists them: action words, and 256 single-token "byte" labels per
coordinate). Training used cross-entropy against soft ordinal (Gaussian) targets, so the bins form a
real distribution over the screen.
Use
With vLLM
Served in bf16 on one RTX PRO 6000 a computer-use step takes 190 ms (median over 60 cold screenshots). The LoRA on the NVFP4 base is faster (152 ms) and keeps the base model for text steps; use this merged model when you want one plain checkpoint.
1. Install vLLM 0.30 and the reader (slotread, from the code repository):
git clone https://github.com/infinitylogesh/gui-decisions && cd gui-decisions
pip install "vllm==0.30.*" -e .
bash serving/patch_vllm.sh # recommended: 256 label ids per request
2. Serve (bf16, 62 GB of weights: one 80–96 GB GPU):
vllm serve infinitylogesh/GUI-Decisions-31B --served-model-name gui-decisions --port 8001 \
--max-model-len 8192 --gpu-memory-utilization 0.85 --max-logprobs 300 --enable-prefix-caching \
--mm-processor-cache-gb 0 --api-server-count 8
3. Read steps and boxes:
from PIL import Image
from slotread.client import SlotModel
m = SlotModel("gui-decisions", "infinitylogesh/GUI-Decisions-31B") # server: VLLM_URL, default http://127.0.0.1:8001
step = m.decide(Image.open("screenshot.png"), "turn on dark mode") # {'action': 'click', 'x': 0.928, 'y': 0.247, 'ms': 190.0}
box = m.box(Image.open("photo.jpg"), "the red frisbee") # {'box': [x1, y1, x2, y2], 'ms': ...}, in [0, 1]
decide reads action, x and y first, then only the slots the action needs (key for press, sdir
and samt for scroll, x2 and y2 for swipe). x and y are fractions of the screenshot's width and
height. For type / answer steps, serve the base model as well and pass its name as text_model= (the
merged weights only answer slots).
What a read is. For each slot, SlotModel sends one chat request whose assistant turn is prefilled up
to that slot, and asks vLLM for the log-probabilities of exactly that slot's labels; nothing is generated
beyond one token. The requests for one step run in parallel and share the image through vLLM's prefix cache:
{"model": "gui-decisions",
"messages": [{"role": "system", "content": [{"type": "text", "text": "<the task's system prompt, from slotread_config.json>"}]},
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
{"type": "text", "text": "Please generate the next move according to the UI screenshot, instruction and previous actions.\n\nInstruction: turn on dark mode\n\nPrevious actions:\nNone"}]},
{"role": "assistant", "content": [{"type": "text", "text": "action: ___ x:"}]}],
"add_generation_prompt": false, "continue_final_message": true, "chat_template_kwargs": {"enable_thinking": false},
"max_tokens": 1, "temperature": 0, "logprobs": true, "top_logprobs": 1, "return_tokens_as_token_ids": true,
"logprob_token_ids": ["<the 256 label ids of slot x, from slotread_config.json>"]}
The slot's answer is the softmax over those labels: the argmax for action; for a coordinate, the argmax
bin refined by its ±2 neighbours (bin k of 256 is the value (k + 0.5) / 256 across the screen).
Notes:
patch_vllm.shraises vLLM's per-requestlogprob_token_idscap from 128 to 256, so a 256-bin slot is one request. Without it, setMAX_LABEL_IDS=128for the client (each coordinate then takes two requests).--api-server-count 8: a single API-server process serializes a step's parallel slot requests.--mm-processor-cache-gb 0: vLLM 0.30's media cache raced on concurrent requests for the same image.- Screenshots go as JPEG at the model's input size (896 × 896 pixels' worth), the fast path. For tiny
targets send lossless full-size images instead (
IMG_FMT=png IMG_MAX_PX=0), as the evaluations did. previoustakes the earlier steps as numbered lines, e.g."1. click(One way)\n2. click(From)".
In-process with transformers (no server)
from PIL import Image
from slotread.local import LocalSlotModel
m = LocalSlotModel.from_pretrained("infinitylogesh/GUI-Decisions-31B")
step = m.decide(Image.open("screenshot.png"), "turn on dark mode")
box = m.box(Image.open("photo.jpg"), "the red frisbee")
Results
This merged model in bf16 with transformers (slotread.local), against the published evaluation of the
LoRA (NVFP4 base, vLLM), case by case on the same sets (paired bootstrap 95% interval on the difference):
| evaluation | this model | published LoRA | difference |
|---|---|---|---|
| ScreenSpot clicks (1,272): point inside the element | 0.869 | 0.866 | +0.003 [−0.006, +0.013] |
| computer-use steps, held-out AGUVIS (1,600) | 0.650 | 0.659 | −0.009 [−0.021, +0.003] |
| RefCOCO val photo boxes (3,811): mean IoU | 0.770 | 0.770 | +0.000 [−0.003, +0.003] |
| COCO val2017 boxes (1,000): mean IoU | 0.744 | 0.742 | +0.002 [−0.002, +0.006] |
No difference is significant. Merging rounds the LoRA's small weight changes into bf16: against the unmerged LoRA on the same bf16 base (0.659 on computer-use steps, exactly the published score) the merged model trends lower on steps (−0.009 [−0.019, +0.001], McNemar p 0.08; 38 of 1,600 actions differ), while clicks and boxes are unchanged. For the closest match to the published numbers, or to write text steps, use the LoRA on the base model. A forward pass reading every slot takes 150–210 ms in plain transformers (bf16, one RTX PRO 6000); served on vLLM with NVFP4 weights a step takes ~146 ms, against 472–546 ms for the base model writing its native JSON.
Training
LoRA r 16, alpha 32, on the language model's attention and MLP projections. Warm-started from the computer-use run (E) and trained further on a mix of AGUVIS computer-use steps (5,000), Wave-UI element clicks (1,000) and boxes (4,000), RefCOCO photo boxes (5,000) and RefCOCO masks (6,000). One RTX PRO 6000, about 4 hours.
Limits
- One step at a time from the screenshot, the instruction and the previous actions: no memory and no acting. Text arguments (what to type, the final answer) should come from the base model: the merged weights cannot switch the adapter off, so for agents that type, use the LoRA on the base model instead.
- Trained on AGUVIS, Wave-UI and RefCOCO; check their terms for your use.
- Points can miss small targets; boxes on UI elements are weaker (mean IoU 0.467 on ScreenSpot) than on photos.
Citation
@article{umapathi2026slotreading,
title = "Stop Decoding Coordinates: 3.5× Faster Computer-Use Grounding",
author = "Umapathi, Logesh Kumar",
journal = "logeshumapathi.com",
year = "2026",
month = "Oct",
url = "https://logeshumapathi.com/blog/2026/10/06/slot-reading.html"
}
- Downloads last month
- 28
