Instructions to use cstr/sam2.1-hiera-tiny-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sam2
How to use cstr/sam2.1-hiera-tiny-ONNX with sam2:
# Use SAM2 with images import torch from sam2.sam2_image_predictor import SAM2ImagePredictor predictor = SAM2ImagePredictor.from_pretrained("cstr/sam2.1-hiera-tiny-ONNX") with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): predictor.set_image(<your_image>) masks, _, _ = predictor.predict(<input_prompts>)# Use SAM2 with videos import torch from sam2.sam2_video_predictor import SAM2VideoPredictor predictor = SAM2VideoPredictor.from_pretrained("cstr/sam2.1-hiera-tiny-ONNX") with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): state = predictor.init_state(<your_video>) # add new prompts and instantly get the output on the same frame frame_idx, object_ids, masks = predictor.add_new_points(state, <your_prompts>) # propagate the prompts to get masklets throughout the video for frame_idx, object_ids, masks in predictor.propagate_in_video(state): ... - Notebooks
- Google Colab
- Kaggle
SAM 2.1 Hiera-tiny, ONNX (image mode)
ONNX export of Meta's SAM 2.1 Hiera-tiny
for single-image segmentation with point and box prompts. It is the model behind
the sam mask provider of Crisp3DS / Crisp 3D Studio,
which runs it from Rust through ONNX Runtime. The video-memory parts of SAM 2 are not included.
| File | Size | SHA-256 |
|---|---|---|
encoder.onnx |
109.5 MB | 4fda6e68561e27808cedffa8689749a622bba5f11ba66f4ca4e3a8a9f6360db5 |
decoder.onnx |
20.7 MB | 0ae944e08e55814c19afa4d143ef07b1a3dbad2decd77c15ec3ab220b9f57247 |
model.json |
sizes, normalisation, hashes, operator counts |
Interface
- Encoder input
image, float32[1, 3, 1024, 1024]: RGB / 255, antialiased bilinear resize to 1024×1024 without keeping the aspect ratio, then (x − mean) / std with mean[0.485, 0.456, 0.406]and std[0.229, 0.224, 0.225]. Outputs, float32:image_embed[1, 256, 64, 64],high_res_0[1, 32, 256, 256],high_res_1[1, 64, 128, 128]. - Decoder inputs: those three tensors,
point_coordsfloat32[1, N, 2]andpoint_labelsint64[1, N](N variable). Coordinates are pixels scaled to the 1024 frame (x · 1024 / W, y · 1024 / H), no half-pixel shift. Labels: 1 object, 0 background; a box is given as its two corners with labels 2 and 3, placed before the click points. No padding point is needed (the graph appends one). Outputs:mask_logitsfloat32[1, 4, 256, 256]andioufloat32[1, 4]; index 0 is the single-mask output, 1–3 the multimask proposals. For masks, enlarge the logits bilinearly (align_corners=False) to the photo size and threshold at 0.
Opset 17, exported with PyTorch 2.7 by
crates/dense/tools/sam2_export_onnx.py.
The encoder's bicubic position-embedding resize is rewritten as two matrix products
(equal to 1e-5), which removes the cubic resize and makes the file smaller.
Verification
Against PyTorch on the CPU, 8 photos at 1749×1155 with real prompts, ONNX Runtime 1.30: image embedding max difference 1.5e-5, mask logits 7.2e-5, scores 6.6e-7, selected mask IoU 1.0 on 8 of 8.
Note on PyTorch MPS: PyTorch 2.7 on Apple's MPS backend computes the strided
query max_pool2d in the Hiera encoder incorrectly (making the input contiguous fixes it).
Compare against PyTorch on the CPU, not MPS.
License
Apache-2.0, as the original model: SAM 2 is Copyright Meta Platforms, Inc. and affiliates.
This repository redistributes converted weights under the same license (see LICENSE).
The export is a format conversion; no weights were retrained or changed.
Model tree for cstr/sam2.1-hiera-tiny-ONNX
Base model
facebook/sam2.1-hiera-tiny