OpenSfM models

Models used by OpenSfM at run time, downloaded on demand. Each family lives in its own folder with its licence; manifest.json at the root lists every file with its size, SHA-256 and interface, and OpenSfM pins a revision of this repository.

sam3/: text-prompted semantic segmentation

Meta's SAM 3 (Segment Anything with Concepts), exported for OpenSfM's segment step: images, from street level to nadir and oblique aerial, are labelled with the classes of a taxonomy, each class being a list of text prompts ("building", "roof", ...). The models are independent of the taxonomy: the text embeddings are an input, computed once per prompt.

Derived from the official SAM 3 checkpoint (sam3.pt) and distributed under the SAM License (sam3/LICENSE), which must accompany any redistribution.

Files

File Role
onnx/sam3-image-encoder-1008-fp16.onnx (+ .onnx.data) Image encoder. pixel_values [1,3,1008,1008] fp16 (RGB / 255, mean 0.5, std 0.5) → fpn_feat_0..2 [1,256,288²/144²/72²] (+ constant fpn_pos_0..2)
onnx/sam3-text-encoder-ctx32-fp16.onnx Text encoder. input_ids [N,32] int64 (CLIP BPE, start 49406, end 49407, padded with 0), attention_mask [N,32] int64 → text_features [32,N,256], text_mask [N,32] bool (True = padding)
onnx/sam3-decoder-1008-p16-top64-fp16.onnx Prompted decoder, 16 prompt slots (pad by repeating a prompt). fpn_feat_0..2, text_features [32,16,256], text_mask [16,32] bool → prompt_scores [16,288,288] (per prompt: max over the detections of score × mask, and the semantic map × presence), scores [16,200], boxes [16,200,4] (cx, cy, w, h normalised)
onnx/sam3-decoder-1008-p16-top64-inst-fp16.onnx Prompted decoder with instances (used by OpenSfM since this revision): as above, but → prompt_scores [16,288,288], instance_index [16,288,288] (per prompt and pixel, the detection 0..63 whose mask wins the pixel, -1: none), instance_scores [16,64] (the 64 best detections' scores, 0 below 0.3), instance_boxes [16,64,4] (cx, cy, w, h normalised)
coreml/sam3-image-encoder-1008-fp16.mlpackage Image encoder, Core ML ML program (macOS 15+); same interface, outputs fpn_feat_0..2
coreml/sam3-decoder-1008-p16-top64-fp16.mlpackage Prompted decoder, Core ML; as the ONNX one with text_mask as fp16 (1 = padding)
coreml/sam3-decoder-1008-p16-top64-inst-fp16.mlpackage Prompted decoder with instances, Core ML; as the ONNX one with text_mask as fp16
embeddings/sam3-text-ctx32-fp16.npz Precomputed prompts: prompts [N], features [N,32,256], mask [N,32]
tokenizer/sam3-clip-bpe-tokenizer.json CLIP BPE tokenizer (Hugging Face tokenizers)

Naming: <model>-<component>-<input size>[-<variant>]-<precision>.<format> (p16: 16 prompt slots, top64: masks computed for the 64 best detections per prompt, inst: instance outputs, ctx32: 32 text tokens).

Export

Exported with the Ultralytics SAM 3 implementation and the export patches of greenjava/sam3-onnx, plus: 1008 px input, the per-prompt score maps merged in the graph, masks for the 64 best detections only, window partitioning and mask products rewritten for Core ML (rank ≤ 5, matmul), opset 18. The Core ML encoder is ONNX Runtime's ML-program conversion with its constants moved to the weight blob; the Core ML decoder is converted from PyTorch with coremltools. The inst decoders select the 64 best detections with a one-hot matmul over the detections' ranks rather than topk, whose indices come out wrong on Apple GPUs (the older decoders lose detections there).

On an Apple M5 Pro (Core ML, CPU + GPU): encoder 0.42 s and decoder 0.8 s per image; label maps agree with the PyTorch reference (transformers Sam3Model, 1008 px) on ~97% of the pixels.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support