QIME: Ontology-Grounded Interpretable Biomedical Embeddings
This repository releases the main training-free QIME method (TF+MMR) from Asking the Right Questions: Ontology-Grounded Interpretable Embeddings for Biomedical Text. The official GitHub repository contains question generation, experiments, and the separate QIME_CLS classifier comparison.
QIME represents each text using a fixed bank of biomedical yes/no questions. It encodes the text and questions with MedEmbed, computes cosine relevance, and applies Maximal Marginal Relevance (MMR) to activate 256 out of 8,855 question coordinates. The resulting embedding is binary and directly inspectable. There are no additional trained QIME weights and no inference-time LLM calls.
| Setting | Released default |
|---|---|
| Method | QIME (TF+MMR), main method reported in paper |
| Backbone | abhinand/MedEmbed-large-v0.1 |
| Backbone revision | 963121bfb9c625475f65b08fb54990ce9c4e7a1a |
| Question bank | 8,855 questions, fixed order |
| Active coordinates per text | k = 256 |
| MMR relevance weight | lambda = 0.7 |
| Output | NumPy float32 array, shape (num_texts, 8855), values 0 or 1 |
Download and run
Use Python 3.9-3.12. A CUDA GPU is recommended; CPU inference is supported but can be slow. The repository contains code and the question bank, not MedEmbed weights. On first initialization, the loader automatically downloads the pinned MedEmbed revision from its upstream repository and caches it on the machine performing inference. Later runs reuse that cache.
python3 -m venv qime-env
source qime-env/bin/activate
python -m pip install "huggingface-hub==0.36.0"
hf download tyxqiean/QIME --local-dir ./QIME
python -m pip install -r QIME/requirements.txt
python QIME/example.py
For a particular GPU:
python QIME/example.py --device cuda:0
The example checks that both vectors have 8,855 coordinates and exactly 256 ones, then prints ten active questions in coordinate order.
Python usage
After downloading the release and installing its requirements:
import sys
import numpy as np
sys.path.insert(0, "./QIME")
from qime import QIMEModel
# Load the downloaded question bank/configuration and the pinned backbone.
model = QIMEModel.from_pretrained("./QIME")
texts = [
"Patient presents with severe chest pain.",
"Treatment involves daily insulin injections.",
]
embeddings = model.encode(texts, batch_size=32)
print(embeddings.shape) # (2, 8855)
print(embeddings.sum(axis=1)) # [256. 256.]
for index in np.flatnonzero(embeddings[0]):
print(index, model.questions[index])
Once the qime.py module is available locally, the same loader can fetch the
bank/configuration from the Hub:
model = QIMEModel.from_pretrained("tyxqiean/QIME", device="cuda:0")
# To pin a QIME release, additionally pass revision="<QIME commit SHA>".
The Hub loader downloads data/configuration; it does not dynamically download
or execute Python code. Download this release first to obtain qime.py.
This is a custom Python encoder, not a standard Transformers or
SentenceTransformer checkpoint. Use QIMEModel, rather than
SentenceTransformer("tyxqiean/QIME") or a generic hosted inference widget.
For an offline installation, first provision both the QIME files and the matching MedEmbed snapshot. Then use:
model = QIMEModel.from_pretrained(
"./QIME",
backbone_path="/path/to/local/MedEmbed-large-v0.1/snapshot",
local_files_only=True,
)
An explicit backbone_path should contain the pinned revision shown above.
cache_dir can be supplied to control cache placement. local_files_only=True
prevents new downloads; missing files raise an error.
Representation and interpretation
- Each coordinate is defined by the question at the same JSON index. The loader verifies the question-bank SHA-256 before constructing the encoder.
- The vector is sparse in its values but returned in dense array storage. Its dimension is 8,855, not 256.
- A value of 1 denotes similarity-based semantic activation, not an LLM answer or a calibrated probability that the question is true.
- Selected coordinates all have value 1; sorting them by value does not
provide a relevance ranking.
model.one_indicesrecords MMR selection order for the latest encode call, which also is not a pure cosine-relevance ranking. - Cosine similarity can compare vectors encoded with the same question bank and settings. For fixed k, it equals the number of shared active coordinates divided by k.
- As in the GitHub implementation, MMR uses
lambda * relevance - (1 - lambda) * max(0, redundancy). Lambda weights relevance; increasing it reduces the diversity penalty.
You may override topk and mmr_lambda when initializing the loader:
model = QIMEModel.from_pretrained("./QIME", topk=128, mmr_lambda=0.7)
These overrides change the experiment configuration. The main paper setting
is k=256 and lambda=0.7. The supported lambda range is (0, 1].
Provenance and validation
The question bank is copied byte-for-byte from TF/data/questions.json at
GitHub commit 67bb43ad0264763270c9d2d6a7a527016ec7dc09. Its SHA-256 is:
79b00c0675cdf38a0a5de0bbae7c055a0c3072bf8a868ae362d739fe85ea5e74
The release wrapper preserves the native MMR selection rule and adds
local/Hub loading, an explicit backbone revision, input and bank validation,
and bounded document batches. provenance.json records the source paths,
revision, and packaging changes. validation.json records release smoke
checks and their actual status. These smoke checks validate the release's
behavior; benchmark reproduction is reported separately below.
The pinned requirements.txt is an inference environment, not the full
Linux/CUDA question-generation and experiment environment on GitHub.
Question generation and QIME_CLS checkpoints are not included in this release.
Benchmark reproduction
The main QIME (TF+MMR) method has been reproduced on all twelve of the paper's benchmarks, with scores close to Tables 1 and 2. The six-task retrieval average is 41.12, compared with 41.10 reported in the paper.
The evaluation setup uses this release at revision 18c5f7b62a8c392d6cdf1bc6a9e76e567883f980, the pinned MedEmbed backbone, k=256, lambda=0.7, MTEB 1.39.7, datasets 3.6.0, seed 42, and batch size 128. PublicHealthQA uses only its English subset. Scores below are scaled by 100 and rounded to two decimal places.
| Dataset | Metric | Paper reported | Reproduced |
|---|---|---|---|
| BIOSSES | cosine_spearman | 79.66 | 79.66 |
| NFCorpus | ndcg_at_10 | 25.09 | 25.08 |
| PublicHealthQA | ndcg_at_10 | 75.64 | 75.78 |
| MedicalQARetrieval | ndcg_at_10 | 62.36 | 62.36 |
| TRECCOVID | ndcg_at_10 | 64.65 | 64.65 |
| R2MEDIIYiClinicalRetrieval | ndcg_at_10 | 11.79 | 11.79 |
| R2MEDPMCClinicalRetrieval | ndcg_at_10 | 7.08 | 7.08 |
| BiorxivClusteringP2P | v_measure | 40.37 | 40.43 |
| BiorxivClusteringS2S | v_measure | 36.78 | 36.77 |
| MedrxivClusteringP2P | v_measure | 33.92 | 33.92 |
| MedrxivClusteringS2S | v_measure | 31.44 | 31.44 |
| ClusTREC-Covid | v_measure | 81.99 | 81.99 |
For the pinned 12-task evaluation configuration and validated environment, see the GitHub evaluation instructions and requirements-eval.txt. benchmark_reproduction.json records scores, result sources, dependency versions, data revisions, and completion status.
Intended use and limitations
QIME supports non-commercial research on biomedical retrieval, clustering, semantic similarity, and interpretable representations. The question bank is in English. Similarity-based activations may miss negation or fine-grained relations and should not be treated as factual clinical decisions.
License
Original QIME materials in this repository are licensed under CC BY-NC 4.0. See LICENSE. Attribution is required and commercial use requires separate permission. Non-commercial adaptations are permitted subject to the license terms.
MedEmbed is downloaded separately from its upstream repository and remains under its original Apache-2.0 license. This repository does not include or relicense its weights. Other third-party dependencies retain their respective licenses.
Please cite the QIME paper when using this release in research.