gua

gua is a computer-use decision model for macOS. Given a screenshot, the task and the last few actions, it chooses what kind of action comes next and which on-screen element it acts on, and can also ground an element from a description. It answers as typed multiple-choice decisions and returns a probability for every option, so an agent can act on confident answers and defer on the rest.

It is Qwen/Qwen3.5-4B fine-tuned (supervised fine-tuning with LoRA, rank 8 on every linear layer of the language model) on GUI trajectories extracted from screen-recorded tutorials of seven kinds of software: macOS system apps, video editing, Blender, Godot, Logic Pro, OBS Studio and Plasticity. This repository holds the LoRA merged into the base weights (MLX, bf16) and, in adapter/, the LoRA itself.

Research model. The training data was derived from public YouTube tutorials and is not released. See Training data and Limitations.

How it answers

The model does not generate free text. Each step is asked as up to three questions, all sharing one prompt prefix (system prompt, screenshot, task, past actions, numbered element list), and the probability of every option is read from the model's next-token distribution:

Question Options Use
action click, double-click, right-click, drag, type, press, shortcut, scroll, hover, wait, finish the kind of the next action
next_target the numbered elements on screen where the next action goes
grounding the numbered elements on screen which element matches a description you give

Elements are candidates found on the screenshot: text lines from the macOS Vision framework, and, during training and evaluation, icons from the OmniParser v2 icon detector. The model chooses among candidates; it cannot point at something no candidate covers. When running on a live Mac, the accessibility tree is a better candidate source than detection.

The prompt format is fixed by training. Use inference/cua_decide.py, which reproduces it exactly (checked to give identical option probabilities to the evaluation code).

Usage

Apple Silicon with mlx-vlm 0.7:

pip install "mlx-vlm>=0.7,<0.8" pyobjc-framework-vision pyobjc-framework-quartz
python inference/cua_decide.py --model <user>/gua \
    --image screen.png --task "Turn on Dark Mode" \
    --past "Click the Apple menu" --past "Click System Settings" \
    --describe "the Appearance item in the sidebar" \
    --temperatures temperatures.json

Each question prints its choice and calibrated confidence, for example:

{"question": "action", "choice": "click", "confidence": 0.94, "action": "click"}
{"question": "next_target", "choice": "15", "confidence": 0.60, "element": {"text": "...", "centre": [0.033, 0.136], "box": [...]}}

To use the LoRA instead of the merged weights, pass the base model and the adapter: --model mlx-community/Qwen3.5-4B-MLX-bf16 --adapter adapter.

OmniParser's icon detector improves candidate recall but its weights are AGPL-3.0 and are not included; pass --icon-weights icon_detect/model.pt if you download them yourself. Without it, only text candidates are offered, which lowers recall on icon-only controls.

Training

gua v2 continues the v1 adapter on a larger, more varied dataset.

Method Supervised fine-tuning with LoRA: rank 8, alpha 16, on every linear layer of the language model (q/k/v/o_proj, linear-attention in_proj_*/out_proj, MLP gate/up/down_proj); vision tower frozen
Objective Log loss of the right option label (and end-of-turn token) only, never of the prompt
Stage 1 (v1) 8,000 questions from macos-cua v2 (macOS system, video editing, Blender, Godot)
Stage 2 (v2) 12,000 further questions from macos-cua v3, shared evenly across its seven topics (each topic at most about 1,940; smaller topics give all they have)
Optimiser Adam, learning rate 5e-5, batch 1, gradient clipping 1.0
Hardware One Apple Silicon Mac (MLX), gradient checkpointing, about 110 h in total, peak 182 GB

The recipe was chosen on held-out dev splits, never on an eval split. The released weights are the end of stage 2, which scored higher on the full v3 dev split (mean accuracy 0.453) than the checkpoint with the best interim dev check (0.442).

Training data

Trajectories were extracted automatically from screen-recorded software tutorials on YouTube: videos were searched and screened (recording quality, platform), segmented and annotated into step-by-step actions by vision-language models, each click was located on the frame, and before/after screenshots were cut. The training split of macos-cua v3, and the questions stage 2 drew from it:

Topic Videos Steps Questions used in stage 2
Plasticity (CAD) 162 16,656 1,937
Godot 4 (game engine) 79 6,271 1,937
Blender (3D) 40 3,833 1,937
OBS Studio (recording) 37 2,890 1,937
Logic Pro (audio) 43 2,020 1,937
macOS video editing (Final Cut Pro, Premiere Pro, CapCut, iMovie) 7 640 1,468
macOS system and productivity apps 6 420 844

Splits are by video (dev 44 videos, eval 99 videos, 38 applications appear only in eval); every video of the earlier v2 release kept its split, so v2's eval is part of v3's eval and was never trained on. Godot typing steps were left out of training: an audit found their labels mostly wrong. Many tutorials were recorded on Windows or full screen; steps that would be wrong on a Mac (Windows shortcuts, taskbar, native dialogs) were removed.

Label quality. Labels are model-generated, not human-verified. A stronger model audited a random sample per topic:

Topic Audited steps Element located correctly Action correct
Logic Pro 150 94.0% 84.0%
OBS Studio 150 92.0% 81.3%
macOS video editing 395 92.4% 78.2%
macOS system and productivity 200 91.5% 75.5%
Blender 180 89.4% 71.7%
Plasticity 570 79.6% 71.4%
Godot 270 87.0% 54.8%

The dataset itself is not released.

Evaluation

macos-cua v3 eval split (9,972 steps from 99 videos never seen in training, 6,509 with an element target), gua v2 against gua v1 with the same prompt and candidates, 95% CI by bootstrap paired by trajectory:

Question gua v1 gua v2 Difference
action accuracy 0.507 0.537 +0.030 (+0.021, +0.040)
grounding accuracy 0.516 0.524 +0.008 (+0.002, +0.015)
next_target accuracy 0.303 0.323 +0.020 (+0.012, +0.029)
action and next_target both right 0.182 0.216

On the 1,767 steps from applications never seen in training: action 0.602, grounding 0.628, next target 0.357. By topic (gua v2):

Topic action grounding next_target
OBS Studio 0.806 0.803 0.453
macOS system and productivity 0.720 0.643 0.465
Logic Pro 0.687 0.566 0.341
macOS video editing 0.592 0.644 0.388
Blender 0.539 0.503 0.253
Godot 0.519 0.657 0.275
Plasticity 0.462 0.381 0.308

macos-cua v2 eval split (3,442 steps, the eval gua v1 was published with), against the untrained base:

Question Base gua v1 gua v2
action accuracy 0.481 0.525 0.542
grounding accuracy 0.443 0.603 0.605
next_target accuracy 0.155 0.261 0.289

A candidate covers the target in 70.2% of v3's element steps (83.4% in v2's), which bounds the element questions; Plasticity's small, dense CAD icons are often missed. Excluding the Godot typing steps, whose labels are mostly wrong, v3 action accuracy is 0.563. Because those steps were left out of training, the model rarely answers type.

Selective accuracy (v3 eval). Acting only on the most confident half of the answers: grounding 0.711, action 0.665, next target 0.444.

Calibration. temperatures.json holds one temperature per question, fitted on the v3 dev split by log loss and applied as p^(1/T) renormalised (inference/cua_decide.py --temperatures). Expected calibration error on the v3 eval split, which the fit never saw:

Question Temperature ECE before ECE after
action 2.15 0.294 0.060
grounding 1.25 0.259 0.173
next_target 1.20 0.253 0.141

Without scaling the model is overconfident on every question; after it, element answers remain somewhat overconfident.

Limitations

  • macOS screenshots only, English UI and tasks.
  • Candidate-bound. The model picks among candidates; if no candidate covers the target (about 30% of element steps in v3's eval, most often in CAD), the element questions have no right answer. On a live Mac, the accessibility tree is a better candidate source.
  • Noisy labels, as audited above; action-type labels are the weakest, and Plasticity's element labels the least reliable.
  • Weak on dense professional UIs: Plasticity grounding is 0.38.
  • Not an agent by itself. It decides the action kind and target; the text to type, the keys of a shortcut and drag end points are not predicted. It rarely predicts type.
  • Merged weights round to bf16. Against base plus adapter, option probabilities differ by at most 0.035 in total variation on our check; use adapter/ for exact reproduction.
  • Confidence is only meaningful after temperature scaling, and only for the question types and candidate sources it was fitted with.

Versions

Version Training Notes
v2 (this) v1 continued on 12,000 balanced questions from macos-cua v3 adds Logic Pro, OBS Studio, Plasticity; better on every question
v1 8,000 questions from macos-cua v2 first release; tagged v1 in this repository

Licence and provenance

The weights are a derivative of Qwen3.5-4B (Apache-2.0) and are released under Apache-2.0 (LICENSE); the modification is the LoRA fine-tuning described above. The training data was derived from public YouTube videos for research; no video frames are distributed here. inference/cua_decide.py is released under the same licence. OmniParser v2 (optional at inference) is licensed separately by its authors.

Downloads last month
24
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LaplAI/gua

Finetuned
Qwen/Qwen3.5-4B
Adapter
(715)
this model