relsgg-convnext-base

Open-vocabulary relation prediction from any boxes or masks. Give the model an image and regions from any source (a detector, a segmenter, ground truth); it returns ranked relations over a predicate vocabulary supplied at inference, and optionally two graphs (spatial + semantic) from the same forward pass. Object class labels are never an input.

Part of RelateAnything (code · paper: RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs). Trained on RA-4M; evaluated with OV-SGG-Bench.

Use it

pip install git+https://github.com/Maelic/RelateAnything
hf download maelic/relsgg-convnext-base          # optional; the API fetches on first use
from relsgg import RelateAnything

# Regions come from any detector, any segmenter, or your own annotation.
# Object class labels are never an input.
model = RelateAnything.from_pretrained("maelic/relsgg-convnext-base", device="cuda")
for t in model.predict(image, boxes_xyxy, topk=20):    # PIL/ndarray, boxes [N, 4] in pixels
    print(t)                                           # (person) --riding [0.67]--> (horse)

# Masks instead of boxes: pass the [N, H, W] binary masks beside their extents.
triplets = model.predict(image, boxes_xyxy, masks=masks, topk=20)

# The vocabulary is an input. Any strings, at any time, without retraining.
model.set_vocabulary(["about to collide with", "reflected in"])

# Or answer from the whole training vocabulary, 19,103 strings, read from the weights.
model = RelateAnything.from_pretrained("maelic/relsgg-convnext-base", full_vocabulary=True, device="cuda")

# Two graphs from one forward pass.
graphs = model.predict(image, boxes_xyxy, decompose=True)   # {"spatial": [...], "semantic": [...]}

Every vocabulary is encoded once by the text student shipped beside the weights, and the head is reparameterized onto it; scoring afterwards is vision only. full_vocabulary=True reads predicate_embeddings.npz instead of encoding, which turns a minute and a half of CPU work into a download. model.pth embeds the backbone configuration, so running these weights needs no gated DINOv3 login.

Files: model.pth (torch, EMA weights), text_student.pt + tokenizer, predicate_embeddings.npz (the training vocabulary, encoded), relateanything.onnx, predicate_bank.npz, thresholds.json, calibration.json, README.md.

Every number below is generated from measured eval artifacts (release/make_model_cards.py); none is hand-typed.

Closed-vocabulary transfer (reparameterized, TEST, graph-constrained)

source R@50 mR@50 F1@50
vg150 0.523 0.285 0.369
psg 0.381 0.297 0.333
indoorvg 0.528 0.295 0.378
hicodet 0.466 0.312 0.374

Open-vocabulary, NO reparameterization (all 19,103 predicates deployed)

Synonym-matched at the calibrated tau (see provenance). This is the honest "the model never saw your label set" protocol.

source SoftR@50 SoftmR@50 SoftF1@50
vg150 0.546 0.339 0.419
psg 0.280 0.285 0.283
indoorvg 0.534 0.332 0.410

Spatial reasoning (SpatialSense, adversarial true/false; chance = 0.5)

Macro AUC over predicates: 0.7122

Two-graph decomposition (spatial / semantic, type-stratified protocol)

source spatial R@50 / mR@50 semantic R@50 / mR@50
vg150 0.628 / 0.313 0.491 / 0.304
psg 0.613 / 0.560 0.420 / 0.320
indoorvg 0.611 / 0.361 0.425 / 0.296

Deployment thresholds (per-predicate best-F1, measured on THIS checkpoint)

Score scales are checkpoint-specific (the output head is rank-trained), so these thresholds transfer to no other model. Regime: gt boxes, pair_weight=0, 5000 val images. Top predicates by support:

predicate threshold best F1 GT support
behind 0.885 0.328 3591
in front of 0.860 0.338 3578
wearing 0.980 0.688 3416
to the right of 0.855 0.390 3196
to the left of 0.855 0.374 3095
resting on 0.975 0.576 2166
on 0.930 0.454 2041
holding 0.980 0.462 1551
beside 0.980 0.198 1403
next to 0.950 0.239 1353
above 0.870 0.319 1281
below 0.895 0.315 1243
part of 0.885 0.488 1135
supporting 0.980 0.187 939
looking at 0.970 0.257 873

Provenance

run relsgg-convnext-base
git e10f70a7256936058af4e3cc6d1f3fe1e69c5aac
backbone facebook/dinov3-convnext-base-pretrain-lvd1689m
text student runs/packed/text_student_v2_512/student.pt sha256 e0317830b68ea51e...
ONNX opset / parity 17 / max
torch / transformers 2.14.1+cu130 / 5.19.0
training mixture megasg_clean + vg_raw + hicodet, per-image 0.727/0.063/0.210; source-aware negatives: ['hicodet']

License and data notices

Weights are a derivative of Meta DINOv3 pretrained weights and are distributed under the DINOv3 license. Training annotations (RA-4M) were generated by gemma-4-26B and carry the Gemma Terms of Use notice; images are referenced by identifier only (Objects365/COCO/OpenImages). The vg_raw subset derives from Visual Genome (CC BY 4.0). Predicate synonyms are deliberately never collapsed — surface-form diversity is part of the label space. Full notices: THIRD_PARTY_NOTICES.md in the code repository.

Citation

@article{neau2026relateanything,
  title   = {RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs},
  author  = {Neau, Ma"elic},
  journal = {arXiv preprint arXiv:2609.12552},
  eprint  = {2609.12552},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url     = {https://arxiv.org/abs/2609.12552},
  year    = {2026}
}
Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including maelic/relsgg-convnext-base

Paper for maelic/relsgg-convnext-base

Evaluation results

  • F1@50 (vg150 test, graph-constrained) on Visual Genome 150 (test)
    self-reported
    0.369