RF-DETR Seg Medium — FP32 ONNX (COCO instance segmentation)

An independent, unofficial FP32 ONNX conversion of Roboflow's RF-DETR Seg Medium pretrained checkpoint, produced and validated for AnnotateIt's local, in-browser auto-annotation. Exported with the official rfdetr exporter — not a re-implementation.

  • Upstream: roboflow/rf-detr tag 1.9.4 (commit 9b009fa928d6218320439803d1da01869a85c072), checkpoint Roboflow/rf-detr-seg-medium — Apache-2.0.
  • Task: instance segmentation, 80 COCO thing classes in a sparse COCO-91 logits layout.
  • Precision: FP32 · opset 17 · fixed batch 1 · input 432×432 · 35,694,259 parameters.
  • model.onnx: 130,098,762 bytes · SHA-256 29f12eb98f8c8143145501fcb6e2fbe1043dc242155afff90227b7e9aed3768e.

Checkpoint loading

Every learned tensor is verified identical to the original pinned checkpoint.

Official upstream metrics (as reported by Roboflow)

Metric Value
COCO mask AP 50:95 45.3
COCO mask AP50 68.4
Parameters 35.7M
Latency 5.9 ms — measured on NVIDIA T4, TensorRT 10.4, CUDA 12.4, FP16, batch 1 (not a browser/CPU/WASM figure)

This conversion's measured results

Full COCO val2017 (5000 images) was re-evaluated for BOTH the pinned PyTorch checkpoint and this ONNX graph, with model inference on CUDA — NVIDIA GeForce RTX 3090, FP32, CPU preprocessing/RLE/COCO scoring, and unchanged official PostProcess on CUDA (see validation-report.json). When CUDA is used, TF32 is disabled for both backends. Structural, determinism, and real-image Torch↔ORT smoke checks still run on CPU; browser validation remains WASM.

CPU real-image smoke requires identical selected counts/classes, score Δ ≤ 0.001, box Δ ≤ 0.1 pixel, and mask IoU > 0.99.

Backend bbox AP segm AP segm AP50
PyTorch 2.13.0+cu130 (CUDA — NVIDIA GeForce RTX 3090) 0.5342 0.4499 0.6865
ONNX Runtime 1.29.0 (CUDA — NVIDIA GeForce RTX 3090) 0.5337 0.4495 0.6860

Torch↔ONNX segm-AP agreement is 0.044 AP points. Comparison with the official 45.3 uses a different evaluation protocol (this evaluation pins num_select=200 and the deploy preprocessing); no metric here is claimed as an improvement or replacement of the upstream number.

The recorded browser gate covers Chromium only (Chromium 151 (launch channel external), onnxruntime-web 1.24.3, WASM EP, single thread): session init 1623 ms, first run 5028 ms, warm p50 5001 ms — with selected-query parity vs Python ORT of max score Δ 1.13e-06, max box Δ 1.34e-06 (normalized), min mask IoU 1.0000 at the logit-0 cut. The browser-canvas compatibility proxy matched Python's detection count and classes; it uses canvas.drawImage, not the product's OpenCV INTER_LINEAR preprocessing. WKWebView and WebView2 were not exercised. An optional V8 JS-heap diagnostic is recorded when available; its gc_forced field states whether garbage collection could be forced. It excludes WASM linear memory and is not a memory-leak gate.

Contract

Input input: float32 [1, 3, 432, 432], RGB, NCHW. Preprocessing: stretch-resize to 432×432 (bilinear, half-pixel, antialias offF.resize(antialias=False)cv2.INTER_LINEAR), scale to [0,1], normalize with ImageNet mean [0.485, 0.456, 0.406] / std [0.229, 0.224, 0.225]. No letterbox, no padding.

Outputs (the official exporter's names are counterintuitive — labels holds the logits, dets the boxes):

Name Shape Meaning
dets [1, 200, 4] cxcywh boxes, normalized 0..1 (regression is unbounded — clamp after conversion)
labels [1, 200, 91] raw class logits (sigmoid per cell; sparse COCO-91 columns)
masks [1, 200, 108, 108] per-query mask logits

Postprocessing (the official PostProcess semantics): sigmoid(labels) → stable global top-K over the flattened 200×91 scores (num_select=200; ties broken by ascending flattened index) → queryIndex = flat // 91, classIndex = flat % 91 → keep score > threshold (default 0.5, strict) → gather dets[queryIndex], convert cxcywh→xyxy, scale, clamp → gather masks[queryIndex] for kept rows only, resize bilinear (align_corners=False) to the image size, binarize at logit > 0.0. No NMS. A query may be selected more than once for different classes.

Class ids are native sparse COCO-91 (person=1, car=3, …, toothbrush=90; index 0 is the background slot and ten N/A holes exist) — never compact them to contiguous COCO-80 and never shift by one.

Files

model.onnx · config.json · manifest.json · validation-report.json · annotateit-compatibility.md · LICENSE · NOTICE · SHA256SUMS · export/ (reproducible export + verification scripts and the exact dependency freeze).

Reproducing

./export/reproduce.sh

On Windows:

export\reproduce.cmd

Both launchers call the portable export/reproduce.py runner using paths relative to the package, independent of the caller's current directory. It creates a platform-native venv, rebuilds twice, and fails unless the rebuilt bytes have release SHA-256 29f12eb98f8c8143145501fcb6e2fbe1043dc242155afff90227b7e9aed3768e.

License

Apache-2.0, matching the upstream model and code. Full Apache-2.0 text included as LICENSE. This is an independent conversion — not endorsed by or affiliated with Roboflow.

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AnnotateIt/rfdetr-seg-medium-coco-onnx

Quantized
(1)
this model