NGA-KR/ko-embed-nli

Korean text embedding model (568M) fine-tuned from BAAI/bge-m3 for sentence-pair classification (NLI / paraphrase).

Training data (train splits only)

  • KLUE-NLI, KorNLI (MNLI + SNLI translations): premise → entailment hypothesis as positive, contradiction hypotheses as hard negatives
  • PAWS-X (ko): sentence1 → paraphrase (label 1) as positive, non-paraphrase (label 0) of the same sentence1 as hard negatives

Declared as KLUE-NLI, KorNLI, PawsXPairClassification in the MTEB model metadata; base-model training data is inherited from bge-m3.

Evaluation (MTEB(kor, v2), PairClassification)

Task AP
KorNLI 84.13
KLUE-NLI 77.74
PawsXPairClassification 60.10
Average 73.99

Training

CachedMultipleNegativesRankingLoss · batch 512 · lr 1.5e-5 · 2 epochs · max_len 128 · CLS pooling · bf16

License

MIT (same as BAAI/bge-m3).

Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NGA-KR/ko-embed-nli

Base model

BAAI/bge-m3
Finetuned
(557)
this model

Space using NGA-KR/ko-embed-nli 1