OpenSfM models
Models used by OpenSfM at run time,
downloaded on demand. Each family lives in its own folder with its licence;
manifest.json at the root lists every file with its size, SHA-256 and
interface, and OpenSfM pins a revision of this repository.
sam3/: text-prompted semantic segmentation
Meta's SAM 3 (Segment Anything
with Concepts), exported for OpenSfM's segment step: images, from street
level to nadir and oblique aerial, are labelled with the classes of a
taxonomy, each class being a list of text prompts ("building", "roof", ...).
The models are independent of the taxonomy: the text embeddings are an
input, computed once per prompt.
Derived from the official SAM 3 checkpoint (sam3.pt) and distributed under
the SAM License (sam3/LICENSE), which must accompany any redistribution.
Files
| File | Role |
|---|---|
onnx/sam3-image-encoder-1008-fp16.onnx (+ .onnx.data) |
Image encoder. pixel_values [1,3,1008,1008] fp16 (RGB / 255, mean 0.5, std 0.5) → fpn_feat_0..2 [1,256,288²/144²/72²] (+ constant fpn_pos_0..2) |
onnx/sam3-text-encoder-ctx32-fp16.onnx |
Text encoder. input_ids [N,32] int64 (CLIP BPE, start 49406, end 49407, padded with 0), attention_mask [N,32] int64 → text_features [32,N,256], text_mask [N,32] bool (True = padding) |
onnx/sam3-decoder-1008-p16-top64-fp16.onnx |
Prompted decoder, 16 prompt slots (pad by repeating a prompt). fpn_feat_0..2, text_features [32,16,256], text_mask [16,32] bool → prompt_scores [16,288,288] (per prompt: max over the detections of score × mask, and the semantic map × presence), scores [16,200], boxes [16,200,4] (cx, cy, w, h normalised) |
onnx/sam3-decoder-1008-p16-top64-inst-fp16.onnx |
Prompted decoder with instances (used by OpenSfM since this revision): as above, but → prompt_scores [16,288,288], instance_index [16,288,288] (per prompt and pixel, the detection 0..63 whose mask wins the pixel, -1: none), instance_scores [16,64] (the 64 best detections' scores, 0 below 0.3), instance_boxes [16,64,4] (cx, cy, w, h normalised) |
coreml/sam3-image-encoder-1008-fp16.mlpackage |
Image encoder, Core ML ML program (macOS 15+); same interface, outputs fpn_feat_0..2 |
coreml/sam3-decoder-1008-p16-top64-fp16.mlpackage |
Prompted decoder, Core ML; as the ONNX one with text_mask as fp16 (1 = padding) |
coreml/sam3-decoder-1008-p16-top64-inst-fp16.mlpackage |
Prompted decoder with instances, Core ML; as the ONNX one with text_mask as fp16 |
embeddings/sam3-text-ctx32-fp16.npz |
Precomputed prompts: prompts [N], features [N,32,256], mask [N,32] |
tokenizer/sam3-clip-bpe-tokenizer.json |
CLIP BPE tokenizer (Hugging Face tokenizers) |
Naming: <model>-<component>-<input size>[-<variant>]-<precision>.<format>
(p16: 16 prompt slots, top64: masks computed for the 64 best detections
per prompt, inst: instance outputs, ctx32: 32 text tokens).
Export
Exported with the Ultralytics SAM 3 implementation and the export patches of
greenjava/sam3-onnx, plus:
1008 px input, the per-prompt score maps merged in the graph, masks for the
64 best detections only, window partitioning and mask products rewritten for
Core ML (rank ≤ 5, matmul), opset 18. The Core ML encoder is ONNX Runtime's
ML-program conversion with its constants moved to the weight blob; the Core
ML decoder is converted from PyTorch with coremltools. The inst decoders
select the 64 best detections with a one-hot matmul over the detections'
ranks rather than topk, whose indices come out wrong on Apple GPUs (the
older decoders lose detections there).
On an Apple M5 Pro (Core ML, CPU + GPU): encoder 0.42 s and decoder 0.8 s
per image; label maps agree with the PyTorch reference (transformers
Sam3Model, 1008 px) on ~97% of the pixels.
- Downloads last month
- 24