OmniParser v2 for Core ML
Screen parsing for macOS on the Apple Neural Engine. Given a screenshot, the detector finds interactive elements (buttons, icons, fields) and the captioner describes what each one does, so an agent can refer to and act on the screen by element instead of by pixel.
This is Microsoft's OmniParser v2 converted to Core ML:
- Detector:
icon_detect_v3(YOLOv9-E), the MIT-licensed detector checkpoint. - Captioner: the OmniParser v2 fine-tune of Florence-2-base, split into an encoder and a decoder.
Every executable op in all three graphs is planned for the Neural Engine (zero CPU or GPU fallbacks).
Requirements
- macOS 15 or later on Apple Silicon
- Core ML must be able to write its user cache; the first load compiles the packages
Download
hf download 1of2/omniparser-v2-coreml --local-dir omniparser-v2-coreml
cd omniparser-v2-coreml && shasum -a 256 -c SHA256SUMS
Contents
| Path | Purpose | Size |
|---|---|---|
models/detector/IconDetectV3.mlpackage |
element detector, two functions | 117 MB |
models/caption/FlorenceEncoderContextEmbeds.mlpackage |
caption encoder (DaViT vision + BART encoder) | 270 MB |
models/caption/FlorenceDecoderPrefixEmbeds.mlpackage |
caption decoder | 192 MB |
models/caption/shared_embedding_fp16.npy |
token embedding table [51289, 768] fp16; memory-map read-only | 79 MB |
models/caption/tokenizer/ |
BART tokenizer and Florence-2 processor config | 2.5 MB |
runtime_config.json |
every shape, threshold and generation setting below, machine-readable | |
models/*/LICENSE, models/detector/provenance.json |
licences and detector source hash | |
SHA256SUMS |
checksums for every file |
Detector
IconDetectV3.mlpackage is a multifunction package:
| Function | Input image |
Output predictions |
|---|---|---|
detect_640 (default) |
fp16 [1, 3, 640, 640] | [1, 5, 8400] |
detect_1280 |
fp16 [1, 3, 1280, 1280] | [1, 5, 33600] |
Preprocessing: RGB in 0..1, letterboxed to a centred square with pad value 114, resampled with Lanczos.
Each anchor carries one class score and a box in grid units ([1, 5, anchors]).
Postprocessing: confidence threshold 0.05, non-maximum suppression at IoU 0.1, at most 300 detections, then boxes overlapping by more than IoU 0.7 are merged so each element appears once. Boxes map back to the original image by undoing the letterbox.
Use detect_1280 for dense or high-resolution screens; detect_640 is faster and enough for most windows.
Captioner
Each detected element is cropped, resized to 64x64, then processed like Florence-2-base input: resized to
768x768, rescaled by 1/255 and normalised with ImageNet mean [0.485, 0.456, 0.406] and std
[0.229, 0.224, 0.225].
| Graph | Inputs | Output |
|---|---|---|
| Encoder | pixel_values fp16 [1, 3, 768, 768], prompt_embeds fp16 [1, 8, 768] |
encoder_hidden_states [1, 585, 768] |
| Decoder | decoder_inputs_embeds fp16 [1, 20, 768], encoder_hidden_states [1, 585, 768] |
logits [1, 20, 51289] |
The prompt is <CAPTION>, tokenised to [0, 2264, 473, 5, 2274, 6190, 116, 2]; look its ids up in
shared_embedding_fp16.npy to form prompt_embeds. Run the encoder once per crop.
Generation is greedy over a fixed 20-token decoder prefix: start token 2, forced BOS 0, EOS and forced EOS 2, pad 1, no repeated 3-grams, at most 20 new tokens. At each step embed the tokens so far into the padded prefix, run the decoder, and take the argmax at the current position after applying those constraints. Decode with special tokens skipped and whitespace trimmed.
Faithfulness
The conversion targets parity with the upstream PyTorch pipeline at three points: the detector's input pixels, the decoded boxes, and the caption text. fp16 weights and activations can still move a borderline detection or change a caption token in rare cases.
Limitations
- Captions are short functional descriptions ("settings", "close window"), not full sentences.
- Text on screen is not read by these models; pair them with an OCR engine for labels and content.
- Trained mostly on desktop and web UIs; unusual custom-drawn interfaces detect less reliably.
License and attribution
Released under the MIT License.
- Detector:
microsoft/OmniParser-v2.0icon_detect_v3/model.ptat revisionf55d0750e5b94db2125ef0b45b0fa4a85ddc59b4(sha25611c6cbb77f22569fab22d86c76407a83ec81ab89dbfe28279854822d6e3fb00c), MIT. Seemodels/detector/LICENSE. - Captioner:
microsoft/OmniParser-v2.0icon_captionat the same revision, fine-tuned frommicrosoft/Florence-2-base, MIT. Seemodels/caption/LICENSE.
@misc{lu2024omniparserpurevisionbased,
title={OmniParser for Pure Vision Based GUI Agent},
author={Yadong Lu and Jianwei Yang and Yelong Shen and Ahmed Awadallah},
year={2024},
eprint={2408.00203},
archivePrefix={arXiv},
}
- Downloads last month
- -
Model tree for 1of2/omniparser-v2-coreml
Base model
microsoft/Florence-2-base