Instructions to use AmritJain/Medallion with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use AmritJain/Medallion with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("AmritJain/Medallion") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Medallion: business entity resolution models
Trained weights of a multi-stage business entity resolution pipeline. It matches business records from several noisy sources (names and addresses with abbreviations, typos, transliteration and missing fields) to a deduplicated reference set, using dense retrieval, pair cross-encoders, a LoRA-tuned 4B reranker and XGBoost stackers with a final calibrator.
All models were trained only on the provided training data; no external data was used.
Contents
| Folder | Base model (licence) | What it is |
|---|---|---|
biencoder_qwen3_embedding_0.6b/ |
Qwen/Qwen3-Embedding-0.6B (Apache-2.0) | Fine-tuned bi-encoder for candidate retrieval (sentence-transformers format) |
cross_encoder_1_bge_reranker_v2_m3/ |
BAAI/bge-reranker-v2-m3 (Apache-2.0) | Pair cross-encoder "ce1" |
cross_encoder_2_bge_reranker_v2_m3/ |
BAAI/bge-reranker-v2-m3 (Apache-2.0) | Pair cross-encoder "ce3", trained on ce1's hard negatives |
judge_qwen3_reranker_4b_lora_round{1,2,3}/ |
Qwen/Qwen3-Reranker-4B (Apache-2.0) | LoRA adapters (rank 16) for the 4B pair judge, three hard-pair rounds; round 3 is the one used in the final calibrator |
mmbert_base_cross_encoder_seed{0,1}/ |
jhu-clsp/mmBERT-base (MIT) | Pair cross-encoders on raw text, two seeds |
xgboost/ranker*/ |
XGBoost | Candidate ranker (5 folds each, features.csv gives the feature order) |
xgboost/stack*/ |
XGBoost | Stackers stack1 to stack8b (5 folds each) |
xgboost/final_calibrator/sibcal6.json |
XGBoost | Final sibling-context residual calibrator (applied on top of the stacker logit via base_margin) |
The 4B judge adapters need the base model Qwen/Qwen3-Reranker-4B, which is downloaded from Hugging Face. The empty-address owner model and the base-pipeline stacker whose predictions are used as t7 train and predict in a single script, so they have no separate weight files; the code regenerates them.
Usage
from huggingface_hub import snapshot_download
path = snapshot_download("AmritJain/Medallion") # everything (~10 GB)
path = snapshot_download("AmritJain/Medallion", allow_patterns=["xgboost/*"]) # tabular models only
- Bi-encoder:
SentenceTransformer(path + "/biencoder_qwen3_embedding_0.6b") - Cross-encoders / mmBERT:
AutoModelForSequenceClassification.from_pretrained(<folder>), input"name | address"pairs - Judge:
PeftModel.from_pretrained(AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-Reranker-4B"), <adapter folder>) - XGBoost:
xgboost.Booster(model_file=<fold json>)
Licence
Released under Apache-2.0. The mmBERT-based cross-encoders derive from an MIT-licensed base model.