GraphMatch Bioactivity: models and LOPO evaluation data

This repository is the versioned artifact store for graphmatch-bioactivity. It contains pretrained GraphMatch checkpoints and the frozen evaluation sets needed to reproduce the reported leave-one-protein-out (LOPO) experiments.

Contents

  • graphmatch-bioactivity-models-v1.tar.gz: versioned model registry.
  • evaluation/v1/manifest.json: families, proteins, hashes, class counts and compatible model identifiers.
  • evaluation/v1/PFxxxxx/UNIPROT/test_balanced.parquet: 1:1 evaluation.
  • evaluation/v1/PFxxxxx/UNIPROT/test_imbalanced_1to10.parquet: 1:10 evaluation for PR-AUC and early enrichment.

The v1 evaluation release contains 25 held-out proteins from five Pfam families. Evaluation files are copied from the frozen experimental campaign without resampling. They must not be merged into training data when reporting LOPO performance.

Download through the package

pip install "graphmatch-bioactivity[models]"
graphmatch-bioactivity models download
graphmatch-bioactivity datasets list
graphmatch-bioactivity datasets download --family PF00069 --protein P28482

Use the model whose identifier has the same family and held-out protein as the test set. Thresholds stored in LOPO checkpoints were selected on validation, not on these test labels.

Interpretation and limitations

GraphMatch estimates shared bioactivity for molecular pairs; it does not receive a protein sequence during inference. Pseudo-negatives, bioactivity records and derived pairs inherit the biases and licensing conditions of their upstream sources. Attention and occlusion explain model behavior and are not proof of a pharmacophore or physical binding mode.

The source code is MIT-licensed. Model weights and evaluation data retain the terms and attribution requirements of their upstream data sources; consult the project documentation before redistribution or commercial use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support