Instructions to use NYSgpt/44b-reranker-gte-modernbert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use NYSgpt/44b-reranker-gte-modernbert with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("NYSgpt/44b-reranker-gte-modernbert") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
π Headline results β specialized beats generalist at reranking
On a held-out test of 3,157 queries over 1,534 scientific papers (9-way reranking with hard negatives), this 149M-parameter model picks the correct document first 93% of the time β and lands it in the top 3 a remarkable 98.2% of the time β outscoring OpenAI's general embeddings, keyword search, and hybrid retrieval, at 41 ms/query on hardware you can self-host.
| Method | Precision@1 | Precision@3 | MRR |
|---|---|---|---|
Dense embeddings β OpenAI text-embedding-3-small |
0.681 | β | 0.745 |
| Hybrid retrieval (RRF, k=60) | 0.728 | β | 0.788 |
| BM25 (keyword search) | 0.856 | β | 0.905 |
| 44b-reranker-minilm-l6 (22.7M, fast tier) | 0.865 | 0.954 | 0.914 |
| 44b-reranker-gte-modernbert (149.6M β this model) | 0.929 | 0.982 | 0.956 |
What that means, all measured on the same test set:
- 93% vs 68% over OpenAI's general embeddings on top-1 accuracy (+25 points); 95.6% vs 74.5% on MRR.
- 93% vs 86% over BM25 keyword search; 93% vs 73% over hybrid retrieval.
- Right document in the top 3: 98.2% β up from 92.6% for the general-purpose base.
- 57% fewer ranking errors than the base it is fine-tuned from (16.7% β 7.1%). Mean rank 1.148 of 9.
Scope: reranking numbers β selecting the right document from a candidate set β not open-corpus search across millions. P@3 for the first-stage retrieval arms was not reported. A companion 22.7M model (44b-reranker-minilm-l6) trades ~6 points of accuracy for ~6Γ the speed.
44b-reranker-gte-modernbert
A cross-encoder reranker fine-tuned on peer review β paper pairs mined from the 44B corpus.
The supervision is unusual and it is the point: the query is a peer reviewer's own summary of a paper, written by a domain expert who had never heard of this model, and the positive is the paper that review is attached to. The relevance label is not a heuristic, not a click, and not LLM-generated.
Fine-tuning on it cut the base model's error rate by 57% on a held-out, paper-disjoint evaluation.
Results
Evaluated on 3,157 held-out queries over 1,534 papers, none of which appear anywhere in training β not as positives and not as negatives. Each query is scored against 9 candidates: 1 positive and 8 mined hard negatives. Random chance is P@1 = 0.1111.
| model | P@1 | P@3 | MRR | mean rank | ms/query |
|---|---|---|---|---|---|
| this model | 0.9290 | 0.9816 | 0.9563 | 1.148 | 41.1 |
gte-reranker-modernbert-base (base) |
0.8334 | 0.9262 | 0.8888 | 1.464 | 53.1 |
| chance | 0.1111 | 0.3333 | β | 5.000 | β |
Error rate 16.66% β 7.10%, a 57.4% relative reduction. At n = 3,157 the gap is far beyond noise (z β 11.9 on an unpaired standard error, which is conservative β these are paired measurements on identical queries).
For reference, the first-stage retrieval arms scored on the same 3,157 queries:
| first stage | P@1 | MRR |
|---|---|---|
| BM25 (lexical) | 0.8562 | 0.9046 |
| Hybrid RRF (k=60) | 0.7282 | 0.7881 |
Dense (text-embedding-3-small) |
0.6807 | 0.7450 |
A sibling model, NYSgpt/44b-reranker-minilm-l6,
was trained on the identical data and reaches P@1 0.8654 at 7.4 ms/query β 1/7th the
parameters and ~6Γ the speed. If latency matters more than the last three points of
accuracy, use that one.
Read the eval shape before you read the numbers
This is a 9-way ranking task, not open-corpus retrieval. A 0.9290 here means "picks the right paper out of nine confusable ones," not "finds the right paper in three million." The negatives are hard by construction β nearest neighbours in embedding space plus lexical near-matches β but the pool is small and the task is closed.
What it was trained on
| training groups | 24,177 |
| negatives per group | 7 |
| distinct papers in train | 12,058 |
| held out (val + test) | 3,099 papers |
| query | an official review's summary, or a meta-review's metareview |
| positive | the paper the review is attached to |
| negatives | mined per query β semantic (HNSW over text-embedding-3-small, ef_search=160) and lexical (Postgres FTS over abstracts) |
| source venues | ICLR 2024, ICLR 2025, NeurIPS 2024 |
Splits are disjoint by paper, not by pair. Papers carry ~4.5 reviews each, so splitting
by pair would place the same paper in train and test and inflate every number downstream.
Proven at build time rather than asserted: |train β© val| = 0, |train β© test| = 0,
|val β© test| = 0.
Held-out papers are excluded from the negative pool as well as the positive one. Label leakage was measured after the fact, not assumed: 0 of 1,534 test positives appear as training positives; 23 (1.5%) appear as a negative for some other query, which biases mildly against this model rather than for it.
Training
| base | Alibaba-NLP/gte-reranker-modernbert-base (150M params) |
| objective | grouped cross-entropy β one positive against 7 hard negatives per group |
| epochs | 2 |
| max sequence length | 512 |
| query / document truncation | 600 / 900 characters |
| learning rate | 2e-5 |
| hardware | 1Γ NVIDIA A10G (g5.xlarge) |
| wall clock | 83 minutes |
| training loss | 0.4734 β 0.0883 |
Limitations
Stated plainly, because most of these are the kind of thing a reader would otherwise have to discover.
- Query register is long-form prose. Training and evaluation queries average
700 characters (174 tokens) of fluent reviewer writing. This model has never seen a short keyword query. Do not assume the gains transfer to a two-word search box β that is untested. - One discipline, one document type. Every review comes from top-tier machine learning conferences (ICLR, NeurIPS; 2024β2025). No other field, venue tier, or document class is represented.
- Context is 512, not 8192. The base supports long context; this fine-tune only ever saw 512 tokens, so the extra capacity is untrained.
- Negatives skew semantic. 92% of mined negatives came from the embedding arm and 8% from the lexical arm β the lexical miner conjoined a review's rarest terms and returned nothing on most queries, so the semantic arm backfilled. A balanced mine is the obvious next version.
- Trained on 7 negatives, evaluated against 8. Evaluation is marginally harder than training was.
- The evaluation shares a pipeline with the training data. Same source, same construction, disjoint only by paper. It measures the reranker's job honestly; it is not an independent measure of end-to-end answer quality in a live system.
- CPU latency is unmeasured. Every timing here is A10G at batch 128. No CPU serving number exists yet.
- False-negative screening may under-catch. Candidate negatives were rejected for sharing the positive's Semantic Scholar id or a near-identical title; the rate was 0.118%, and the S2-duplicate check fired zero times across 28,798 candidates. That is either a clean corpus or a check that is not working, and we have not yet established which.
- The supervision is not exhausted. This model trained on 30,509 of 77,147 available pairs. It is a first cut, not a ceiling.
Intended use
Second-stage reranking of scientific-document retrieval β reordering a first-stage candidate list for queries written in the register of a paragraph of technical prose.
Out of scope
Short keyword queries Β· non-ML disciplines Β· any setting requiring a measured CPU latency guarantee Β· any claim of quality on open-corpus retrieval, which this evaluation does not test.
Provenance
Reviews from OpenReview; papers from arXiv, OpenAlex and Semantic Scholar, unified in the 44B corpus. Trained 2026-07-31.
- Downloads last month
- 5
Model tree for NYSgpt/44b-reranker-gte-modernbert
Base model
answerdotai/ModernBERT-base