CodeMixESC cross-lingual retriever

paraphrase-multilingual-mpnet-base-v2 fine-tuned with the Multiple Negatives Ranking (in-batch contrastive) loss on 16,951 Hinglish-English sentence pairs (PHINC + Roman-script Hinglish rewrites of ESConv training utterances), so that a Roman-script Hinglish query retrieves the same English ESConv cases as its English version. Part of CodeMixESC (course project, IIT Bhilai); used by its live demo.

Dev retrieval (151 turns, k = 10) Overlap@10 Light Overlap@10 Heavy
all-roberta-large-v1 (MultiAgentESC) 56.0 22.1
multilingual mpnet (base) 69.9 29.6
this model 72.9 54.0

Overlap@10 = share of the cases retrieved for a Hinglish post that are also retrieved for its English original. Training: 1 epoch, batch 32, lr 2e-5, scale 20, word embeddings frozen; best checkpoint by dev retrieval. esconv_bank_embeddings.npy holds this model's embeddings of the 13,484 ESConv case-bank posts (no text). For research use.

from sentence_transformers import SentenceTransformer
m = SentenceTransformer("roshan9136/codemix-retriever")
m.encode(["yaar bahut tension ho rahi hai", "I am very stressed"])
Downloads last month
14
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for roshan9136/codemix-retriever