Florence-2-base-ft vision encoder — Core ML (fp16)

The vision encoder (DaViT) of Florence-2-base-ft, converted to Core ML so the image side of captioning runs on Apple GPUs. Built for Retriever, a local semantic search over personal video archives on macOS.

What was modified: format only. The ONNX export from onnx-community/Florence-2-base-ft (vision_encoder.onnx, fp32) was simplified with fixed input shape (onnx-simplifier), loaded with onnx2torch, wrapped with ImageNet normalization and converted with coremltools 9 to an ML Program with float16 weights.

Input and output

Input image 768×768 RGB image, values 0–255 (the model scales to [0, 1] and applies mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225] itself)
Output embedding image_features, shape (1, 577, 768), float16; feed it to the text encoder exactly like the ONNX output

Fidelity against the ONNX original

40 real frames, per-token cosine against the ONNX features and the generated caption compared verbatim:

Compute unit ms / image Min cosine Identical captions
GPU 128 0.9933 38 / 40
Neural Engine 1189 (compilation fails, falls back) 0.9660 35 / 40
CPU (Core ML) 188 0.2808 8 / 40

ONNX Runtime on CPU: 744 ms per image. Run it on the GPU (MLComputeUnits.cpuAndGPU); the Neural Engine cannot compile the full graph and the Core ML CPU path is numerically off.

License

Florence-2 is released by Microsoft under the MIT license; this conversion keeps it.

Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for antonlnz/florence2-base-vision-coreml

Quantized
(7)
this model