OpenX: trained weights for business entity resolution
Trained weights of Team OpenX's solution to the Amazon ML Challenge 2026 (business entity resolution: finding the
copies of each clean business record among messy records from two other sources). The code that uses them ships in
our submission package under code/business_entity_resolution; python src/download_weights.py there puts these
folders in place under artifacts/experiments/.
| Folder | What it is |
|---|---|
v7f/model, v7j/model, v7k/model |
the three trained runs: LightGBM filter, matcher and final scorer, plus the fine-tuned mDeBERTa cross-encoder |
sfadapt/cross_encoder_v2, sfadapt/cross_encoder_v3 |
copies of the v7f cross-encoder tuned on synthetic French pairs |
analysis/dense_retrieval/e5s_ft2 |
multilingual-e5-small fine-tuned on normalized name and address keys |
analysis/dense_retrieval_raw/e5s_raw |
multilingual-e5-small fine-tuned on raw record text |
The ensemble has 1,630,699,080 parameters combined (1.63B), all under the MIT licence.
The cross-encoders are fine-tuned from microsoft/mdeberta-v3-base (MIT) and the retrieval encoders from
intfloat/multilingual-e5-small (MIT); the other models are LightGBM (MIT). They were trained on the challenge's
training data. The two French cross-encoders were also tuned on synthetic pairs generated from the French Source 1
records of the test set, without labels. No challenge data is stored in this repository.
Model tree for Heman1324/openx-business-entity-resolution
Base model
intfloat/multilingual-e5-small