Grounding DINO tiny β€” ExecuTorch (open-vocabulary detection)

You name the thing in text and it finds it. No fixed class list: the prompt is "a person. a car. a bicycle. a dog." today and "a fire hydrant. a traffic cone." on the next call, with the same weights.

(pixel_values (1,3,640,640), input_ids (1,32), token_type_ids (1,32),
 attention_mask (1,32), text_masks (1,32,32), position_ids (1,32))
   -> scores (1,900,256), boxes (1,900,4)
build file MB scores corr boxes corr detections agreeing Mac ms*
fp32 grounding_dino_tiny_xnnpack_fp32.pte 690.3 0.999833 0.998992 26/26 584.1
int8 (dynamic) grounding_dino_tiny_xnnpack_int8.pte 254.4 0.804530 0.777331 26/27 738.7

*Mac arm64, single process, median of 5 β€” a reference point for relative cost, not a device number. Torch eager fp32 on the same machine: 376.7 ms. XNNPACK delegate coverage 63.0% (3503/5559 ops) for fp32, 71.5% for int8.

Read that int8 row carefully. Its correlations are 0.80 and 0.78, which on any other model on this shelf would withdraw the build. Here they are dominated by the queries that detect nothing, whose coordinates are free to move without changing an answer β€” and the detections themselves hold: all 26 that eager finds are found, at IoU > 0.5 with the same prompt token, plus one box eager does not have. It is 37% of the fp32 file and it is what you would put on a phone, but the evidence behind it is 26 detections over five photographs, not a COCO run, and one of those 27 boxes already disagrees. It is also slower than fp32 on this Mac (738.7 ms against 584.1).

There is no Core ML build. The graph reshapes to rank 6 and Core ML's ceiling is 5 β€” the same wall RAFT-small hits with its correlation volume. This one is XNNPACK only, and the same file runs on Android and on the Mac.

Running it

1. The image. RGB, divide by 255, ImageNet normalise (mean .485/.456/.406, std .229/.224/.225), resize to 640Γ—640.

2. The prompt. Lower-case phrases separated by ". ", ending in a period. Tokenise with the repo's tokenizer at padding="max_length", max_length=32. input_ids, token_type_ids and attention_mask are int64 (1,32).

3. The two text tensors that are not the tokenizer's. text_masks and position_ids come from transformers' generate_masks_with_special_tokens_and_transfer_map(input_ids) β€” the mask that keeps each phrase's tokens attending only to their own phrase. It is built with cummax/cummin, which are not in the Core ATen opset, so it is lifted out of the graph and passed in β€” text_masks is bool (1,32,32) and position_ids is int64 (1,32), and ExecuTorch will not coerce either. Both depend on input_ids alone, never on the image, so compute them once when the user types the prompt. The model stays open-vocabulary; only the token count is fixed.

4. Reading the answer. boxes is (cx, cy, w, h) normalised to the input square. scores[0, q, t] is how much query q points at text position t. A detection is a query whose best score clears your threshold (0.3 is a reasonable default), and its label is the input_ids token at that argmax.

Two things about that text axis. It is always 256 wide β€” the model's max_text_len β€” no matter how long your prompt is, and every column past your prompt's tokens is exactly 0. And the peak often lands on an article rather than the noun: a query for a person may argmax on the a of "a person". Map the winning token back to the phrase it belongs to, the way transformers' own post-processor does, rather than reading the single token as the label.

What is measured

The number that decides is detections, not correlation. Against eager PyTorch on five street photographs at threshold 0.3:

fp32 reproduces 26 of 26, each at IoU > 0.5 with the same prompt token, and invents nothing. That includes a photo with 17 detections in it. int8 scores 26 of 27 β€” it finds all of eager's and one box besides. The denominator is whichever build found more, so an invented detection costs the same as a missed one; counting only the reference's detections would give a build that returns every box in the image a perfect score.

python convert/audit_int8.py grounding_dino_tiny --variant fp32   # or int8
python convert/verify_cards.py grounding_dino_tiny

Correlation is in the table above for completeness, and the boxes figure (0.998992) is below this shelf's usual fp32 bar of 0.999. It is not measuring what it looks like: the largest box differences are on queries that fire at nothing, where a coordinate is free to wander without changing any answer. The detection count is what was gated on.

The number that had this model filed as broken

Grounding DINO writes -inf into every text column the prompt does not use. The score tensor is 900 queries Γ— 256 columns; a four-phrase prompt uses 14 of them, so 94.5% of the raw logits are -inf. Correlation over that tensor is undefined, and subtracting two builds gives NaN rather than a difference. Measured that way this conversion read "corr 0.857, returns the stopword 'a' where eager finds four people" and sat parked. Both halves were artefacts: eager returns that same a query at the same 0.385, and the detections were identical all along.

The shipped graph ends in a sigmoid, so those positions are exactly 0, the output is in the 0..1 an app wants, and the gate measures the model instead of its padding.

torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)

Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlboydaisuke/GroundingDINO-tiny-ExecuTorch

Quantized
(2)
this model