Medallion: business entity resolution models

Trained weights of a multi-stage business entity resolution pipeline. It matches business records from several noisy sources (names and addresses with abbreviations, typos, transliteration and missing fields) to a deduplicated reference set, using dense retrieval, pair cross-encoders, a LoRA-tuned 4B reranker and XGBoost stackers with a final calibrator.

All models were trained only on the provided training data; no external data was used.

Contents

Folder Base model (licence) What it is
biencoder_qwen3_embedding_0.6b/ Qwen/Qwen3-Embedding-0.6B (Apache-2.0) Fine-tuned bi-encoder for candidate retrieval (sentence-transformers format)
cross_encoder_1_bge_reranker_v2_m3/ BAAI/bge-reranker-v2-m3 (Apache-2.0) Pair cross-encoder "ce1"
cross_encoder_2_bge_reranker_v2_m3/ BAAI/bge-reranker-v2-m3 (Apache-2.0) Pair cross-encoder "ce3", trained on ce1's hard negatives
judge_qwen3_reranker_4b_lora_round{1,2,3}/ Qwen/Qwen3-Reranker-4B (Apache-2.0) LoRA adapters (rank 16) for the 4B pair judge, three hard-pair rounds; round 3 is the one used in the final calibrator
mmbert_base_cross_encoder_seed{0,1}/ jhu-clsp/mmBERT-base (MIT) Pair cross-encoders on raw text, two seeds
xgboost/ranker*/ XGBoost Candidate ranker (5 folds each, features.csv gives the feature order)
xgboost/stack*/ XGBoost Stackers stack1 to stack8b (5 folds each)
xgboost/final_calibrator/sibcal6.json XGBoost Final sibling-context residual calibrator (applied on top of the stacker logit via base_margin)

The 4B judge adapters need the base model Qwen/Qwen3-Reranker-4B, which is downloaded from Hugging Face. The empty-address owner model and the base-pipeline stacker whose predictions are used as t7 train and predict in a single script, so they have no separate weight files; the code regenerates them.

Usage

from huggingface_hub import snapshot_download
path = snapshot_download("AmritJain/Medallion")                      # everything (~10 GB)
path = snapshot_download("AmritJain/Medallion", allow_patterns=["xgboost/*"])   # tabular models only
  • Bi-encoder: SentenceTransformer(path + "/biencoder_qwen3_embedding_0.6b")
  • Cross-encoders / mmBERT: AutoModelForSequenceClassification.from_pretrained(<folder>), input "name | address" pairs
  • Judge: PeftModel.from_pretrained(AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-Reranker-4B"), <adapter folder>)
  • XGBoost: xgboost.Booster(model_file=<fold json>)

Licence

Released under Apache-2.0. The mmBERT-based cross-encoders derive from an MIT-licensed base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support