OpenX: trained weights for business entity resolution

Trained weights of Team OpenX's solution to the Amazon ML Challenge 2026 (business entity resolution: finding the copies of each clean business record among messy records from two other sources). The code that uses them ships in our submission package under code/business_entity_resolution; python src/download_weights.py there puts these folders in place under artifacts/experiments/.

Folder What it is
v7f/model, v7j/model, v7k/model the three trained runs: LightGBM filter, matcher and final scorer, plus the fine-tuned mDeBERTa cross-encoder
sfadapt/cross_encoder_v2, sfadapt/cross_encoder_v3 copies of the v7f cross-encoder tuned on synthetic French pairs
analysis/dense_retrieval/e5s_ft2 multilingual-e5-small fine-tuned on normalized name and address keys
analysis/dense_retrieval_raw/e5s_raw multilingual-e5-small fine-tuned on raw record text

The ensemble has 1,630,699,080 parameters combined (1.63B), all under the MIT licence.

The cross-encoders are fine-tuned from microsoft/mdeberta-v3-base (MIT) and the retrieval encoders from intfloat/multilingual-e5-small (MIT); the other models are LightGBM (MIT). They were trained on the challenge's training data. The two French cross-encoders were also tuned on synthetic pairs generated from the French Source 1 records of the test set, without labels. No challenge data is stored in this repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Heman1324/openx-business-entity-resolution

Finetuned
(198)
this model