OmniParser v2 for Core ML

Screen parsing for macOS on the Apple Neural Engine. Given a screenshot, the detector finds interactive elements (buttons, icons, fields) and the captioner describes what each one does, so an agent can refer to and act on the screen by element instead of by pixel.

This is Microsoft's OmniParser v2 converted to Core ML:

  • Detector: icon_detect_v3 (YOLOv9-E), the MIT-licensed detector checkpoint.
  • Captioner: the OmniParser v2 fine-tune of Florence-2-base, split into an encoder and a decoder.

Every executable op in all three graphs is planned for the Neural Engine (zero CPU or GPU fallbacks).

Requirements

  • macOS 15 or later on Apple Silicon
  • Core ML must be able to write its user cache; the first load compiles the packages

Download

hf download 1of2/omniparser-v2-coreml --local-dir omniparser-v2-coreml
cd omniparser-v2-coreml && shasum -a 256 -c SHA256SUMS

Contents

Path Purpose Size
models/detector/IconDetectV3.mlpackage element detector, two functions 117 MB
models/caption/FlorenceEncoderContextEmbeds.mlpackage caption encoder (DaViT vision + BART encoder) 270 MB
models/caption/FlorenceDecoderPrefixEmbeds.mlpackage caption decoder 192 MB
models/caption/shared_embedding_fp16.npy token embedding table [51289, 768] fp16; memory-map read-only 79 MB
models/caption/tokenizer/ BART tokenizer and Florence-2 processor config 2.5 MB
runtime_config.json every shape, threshold and generation setting below, machine-readable
models/*/LICENSE, models/detector/provenance.json licences and detector source hash
SHA256SUMS checksums for every file

Detector

IconDetectV3.mlpackage is a multifunction package:

Function Input image Output predictions
detect_640 (default) fp16 [1, 3, 640, 640] [1, 5, 8400]
detect_1280 fp16 [1, 3, 1280, 1280] [1, 5, 33600]

Preprocessing: RGB in 0..1, letterboxed to a centred square with pad value 114, resampled with Lanczos. Each anchor carries one class score and a box in grid units ([1, 5, anchors]).

Postprocessing: confidence threshold 0.05, non-maximum suppression at IoU 0.1, at most 300 detections, then boxes overlapping by more than IoU 0.7 are merged so each element appears once. Boxes map back to the original image by undoing the letterbox.

Use detect_1280 for dense or high-resolution screens; detect_640 is faster and enough for most windows.

Captioner

Each detected element is cropped, resized to 64x64, then processed like Florence-2-base input: resized to 768x768, rescaled by 1/255 and normalised with ImageNet mean [0.485, 0.456, 0.406] and std [0.229, 0.224, 0.225].

Graph Inputs Output
Encoder pixel_values fp16 [1, 3, 768, 768], prompt_embeds fp16 [1, 8, 768] encoder_hidden_states [1, 585, 768]
Decoder decoder_inputs_embeds fp16 [1, 20, 768], encoder_hidden_states [1, 585, 768] logits [1, 20, 51289]

The prompt is <CAPTION>, tokenised to [0, 2264, 473, 5, 2274, 6190, 116, 2]; look its ids up in shared_embedding_fp16.npy to form prompt_embeds. Run the encoder once per crop.

Generation is greedy over a fixed 20-token decoder prefix: start token 2, forced BOS 0, EOS and forced EOS 2, pad 1, no repeated 3-grams, at most 20 new tokens. At each step embed the tokens so far into the padded prefix, run the decoder, and take the argmax at the current position after applying those constraints. Decode with special tokens skipped and whitespace trimmed.

Faithfulness

The conversion targets parity with the upstream PyTorch pipeline at three points: the detector's input pixels, the decoded boxes, and the caption text. fp16 weights and activations can still move a borderline detection or change a caption token in rare cases.

Limitations

  • Captions are short functional descriptions ("settings", "close window"), not full sentences.
  • Text on screen is not read by these models; pair them with an OCR engine for labels and content.
  • Trained mostly on desktop and web UIs; unusual custom-drawn interfaces detect less reliably.

License and attribution

Released under the MIT License.

  • Detector: microsoft/OmniParser-v2.0 icon_detect_v3/model.pt at revision f55d0750e5b94db2125ef0b45b0fa4a85ddc59b4 (sha256 11c6cbb77f22569fab22d86c76407a83ec81ab89dbfe28279854822d6e3fb00c), MIT. See models/detector/LICENSE.
  • Captioner: microsoft/OmniParser-v2.0 icon_caption at the same revision, fine-tuned from microsoft/Florence-2-base, MIT. See models/caption/LICENSE.
@misc{lu2024omniparserpurevisionbased,
  title={OmniParser for Pure Vision Based GUI Agent},
  author={Yadong Lu and Jianwei Yang and Yelong Shen and Ahmed Awadallah},
  year={2024},
  eprint={2408.00203},
  archivePrefix={arXiv},
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1of2/omniparser-v2-coreml

Quantized
(5)
this model

Paper for 1of2/omniparser-v2-coreml