Hindi fork: Bharat-Embed 270M (measured numbers inside)

#10
by GautamKishore - opened

We forked the text-only 270M path for Hinglish/Hindi retrieval over at eulogik/bharat-embed-270m-gemma2 (Apache-2.0). Measured on Hindi IndicQA: base 0.7279 vs ours 0.7324 NDCG@10, STS tie. Truncation table measured too: 512d -0.008, 256d -0.030, 128d -0.103, so we default 256d. ONNX INT8 parity drift 0.0066, GGUF Q4 cosine 0.75 on a Hindi spot check. Misses stated plainly on the card. Feedback welcome, especially Tamil/Telugu eval help.

Sign up or log in to comment