e5-tr-nli — Turkish Sentence Embedding Model (Bi-Encoder)

A Turkish sentence-embedding (bi-encoder) model for retrieval and semantic search, fine-tuned from intfloat/multilingual-e5-base with contrastive learning (MultipleNegativesRankingLoss) on Turkish NLI triplets.

  • Embedding dimension: 768
  • Max sequence length: 256 tokens (trained); backbone supports 512
  • Pooling: mean pooling
  • Similarity function: cosine
  • Parameters: ~278M (XLM-RoBERTa-base backbone)
  • Prefix convention: e5-style query: / passage: (see Usagerequired)

Table of Contents

  1. Intended Use
  2. Usage
  3. Training Data
  4. Training Procedure
  5. Evaluation
  6. TR-MTEB Results
  7. Limitations & Biases
  8. Compute & Environmental Impact
  9. Licensing
  10. Citation
  11. Reproduction

Intended Use

In scope

  • Turkish semantic search / passage retrieval (dense retrieval, RAG)
  • Semantic textual similarity, paraphrase & duplicate detection
  • Clustering and topic grouping of Turkish text
  • Feature extraction for downstream Turkish NLP classifiers

Out of scope

  • Re-ranking with query–document cross-attention (use a cross-encoder instead)
  • Long-document embedding beyond 256 tokens without chunking
  • High-stakes decisions (legal, medical, hiring) without human oversight
  • Non-Turkish text (backbone is multilingual, but this model is tuned for Turkish)

Usage

⚠️ e5 prefixes are mandatory

The model was trained with the e5 asymmetric prefixes and must be used with them, consistently at training / evaluation / inference:

Role Prefix
Query / anchor / any symmetric-task sentence (STS, classification, clustering) query:
Passage / document / positive & negative candidates passage:

Omitting the prefixes degrades quality noticeably.

With sentence-transformers

from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim

model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli")

# Asymmetric retrieval: query vs. candidate passages
query = "query: Türkiye'nin başkenti neresidir?"
passages = [
    "passage: Ankara, Türkiye'nin başkentidir.",
    "passage: İstanbul Türkiye'nin en kalabalık şehridir.",
    "passage: Muz tropikal bir meyvedir.",
]
q = model.encode(query, normalize_embeddings=True)
p = model.encode(passages, normalize_embeddings=True)
print(cos_sim(q, p))   # highest score -> the Ankara passage

Semantic similarity (symmetric — use query: on both sides)

a = model.encode("query: Bugün hava çok güzel.", normalize_embeddings=True)
b = model.encode("query: Hava bugün oldukça güzel.", normalize_embeddings=True)
print(float(cos_sim(a, b)))   # ~0.9

With 🤗 Transformers (mean pooling)

import torch, torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel

tok = AutoTokenizer.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")
mdl = AutoModel.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli")

def encode(texts):
    batch = tok(texts, padding=True, truncation=True, max_length=256, return_tensors="pt")
    with torch.no_grad():
        out = mdl(**batch)
    mask = batch["attention_mask"].unsqueeze(-1).float()
    emb = (out.last_hidden_state * mask).sum(1) / mask.sum(1)   # mean pooling
    return F.normalize(emb, p=2, dim=1)

emb = encode(["query: örnek cümle", "passage: örnek pasaj"])

Training Data

  • Source: mertcobanov/all-nli-triplets-turkish — a machine-translated Turkish version of the AllNLI (SNLI + MultiNLI) triplet set.

  • Format: (anchor, positive, negative) triplets, where the negative acts as a hard negative for the contrastive objective.

  • Column handling: Only the Turkish columns (anchor_translated, positive_translated, negative_translated) were used and renamed to anchor / positive / negative. All English columns were discarded.

  • Cleaning: rows with None / empty / whitespace-only fields were filtered out.

  • Resulting sizes (after filtering):

    Split Triplets
    train 277,167
    dev 6,584
    test 6,609

Training Procedure

Objective

  • Loss: CachedMultipleNegativesRankingLoss (MNRL with in-batch + hard negatives; the cached variant is used to simulate a large effective batch on limited VRAM).
  • Batch sampler: NO_DUPLICATES (required for MNRL to avoid trivial in-batch collisions).
  • In-batch negatives: every other positive/negative in the batch serves as a negative for a given anchor, so larger batches yield a stronger contrastive signal.

Hyperparameter search

An Optuna search (TPE sampler) was run before the final training:

Setting Value
Trials 10
Subset 40,000 random training triplets
Steps / trial 600
Search space lr ∈ {1e-5, 2e-5, 3e-5}, batch ∈ {32, 64, 128}, warmup_ratio ∈ {0.0, 0.1}
Selection metric stsb-tr dev spearman_cosine
Best config lr = 2e-5, batch = 32, warmup_ratio = 0.1 (dev spearman = 0.8285)

Final training

Setting Value
Base model intfloat/multilingual-e5-base
Epochs 1
Learning rate 2e-5
Batch size 32 (cached MNRL)
Warmup ratio 0.1
Max sequence length 256
Precision bf16
Seed 42
Eval strategy steps, monitoring eval_stsb-tr-dev_spearman_cosine
Checkpoint selection load_best_model_at_end=True (best dev spearman)

Frameworks & hardware

  • sentence-transformers 5.2.2, transformers 5.0.0, torch 2.11.0+cu128, datasets 4.4.1, Python 3.12
  • 1× NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), CUDA 12.8

Evaluation

Baseline (base model) vs. Final (this model)

Measured on the dataset's own test triplets (TripletEvaluator, cosine accuracy) and on the external Turkish STS set emrecan/stsb-mt-turkish (EmbeddingSimilarityEvaluator, Spearman; scores normalised 0–5 → 0–1):

Metric Baseline (e5-base) Final Δ
stsb-tr dev spearman_cosine 0.8092 0.8312 +0.0220
stsb-tr test spearman_cosine 0.7761 0.7814 +0.0053
nli-tr test cosine_accuracy 0.8968 0.9260 +0.0292

Methodology notes for TR-MTEB

  • Datasets pulled directly from HuggingFace (trmteb/*); metrics computed with a standalone script (no mteb package dependency), following standard MTEB protocols:
    • STS → Spearman of cosine vs. gold; Retrieval → nDCG@10 / Recall@10 / MAP@10 via exact cosine search; Classification → logistic regression on frozen embeddings (accuracy + macro-F1); PairClassification → average precision of cosine; Clustering → V-measure (MiniBatchKMeans, k = #labels); BitextMining → top-1 nearest-neighbour accuracy.
  • e5 prefixes applied throughout (query: for queries/symmetric sides, passage: for corpus).
  • Retrieval search runs on-GPU in fp16.
  • 75haber, thy_sa, irony-tr have no train split in the trmteb org, so a seeded 70/30 stratified self-split was used for those classification tasks (marked note in trmteb_results.json). All other classification sets use their official train/test.

Category summary (macro averages)

Task type # datasets Avg. main metric
STS (spearman_cosine) 1 0.791
BitextMining (accuracy) 1 0.975
Classification (accuracy) 6 0.817
PairClassification (AP) 3 0.573
Clustering (v-measure) 2 0.487
Retrieval (nDCG@10) 10 0.407
Overall (mean of category means) 23 0.675

TR-MTEB Results (full)

Evaluated on the TR-MTEB datasets (Baysan & Güngör, TR-MTEB, Findings of EMNLP 2025).

Task type Dataset Main metric Value
STS stsb-tr spearman_cosine 0.7907
BitextMining wmt16_en_tr accuracy 0.9753
Classification news-cat accuracy 0.9600
Classification 75haber ᵃ accuracy 0.8696
Classification ts_timeline_news_category accuracy 0.7250
Classification thy_sa ᵃ accuracy 0.8386
Classification offenseval accuracy 0.8427
Classification irony-tr ᵃ accuracy 0.6667
PairClassification snli_tr ap_cosine 0.5280
PairClassification multinli_tr ap_cosine 0.6324
PairClassification xnli_tr ap_cosine 0.5591
Clustering ts_abstract_corpus v_measure 0.5762
Clustering 630koseyazisi v_measure 0.3978
Retrieval tquad ndcg@10 0.8222
Retrieval quora-tr ndcg@10 0.7323
Retrieval squad-tr ndcg@10 0.6584
Retrieval scifact-tr ndcg@10 0.5309
Retrieval cqadupstack-gaming-tr ndcg@10 0.3584
Retrieval msmarco-tr ndcg@10 0.2734
Retrieval arguana-tr ndcg@10 0.2583
Retrieval nfcorpus-tr ndcg@10 0.2272
Retrieval fiqa-tr ndcg@10 0.1237
Retrieval scidocs-tr ndcg@10 0.0887

Limitations & Biases

  • Translated training data. The NLI triplets are machine-translated; translation artefacts and noise are inherited, which can cap fine-grained semantic precision.
  • Short-text bias. Trained on NLI-style short sentences at 256 tokens; long-document retrieval requires chunking and may underperform (see the lower fiqa/scidocs scores).
  • Prefix sensitivity. Using the model without query:/passage: prefixes degrades results.
  • Domain gaps. Scientific/financial retrieval (scidocs, fiqa) is weak; the model is strongest on general-domain QA-style retrieval (tquad, squad-tr, quora-tr).
  • Societal bias. The backbone and NLI data may encode social biases; validate before deployment in sensitive settings.
  • Single epoch. Trained for 1 epoch; longer schedules or larger effective batches may improve retrieval further.

Compute & Environmental Impact

  • Single consumer laptop GPU (RTX 5050, 8 GB). Full pipeline (Optuna sweep + 1-epoch training on 277k triplets + evaluation) ran in a few GPU-hours. No multi-GPU/cluster training was used.

Licensing

  • Backbone intfloat/multilingual-e5-base: MIT.
  • Training data derives from AllNLI (SNLI is CC BY-SA 4.0; MultiNLI has mixed source licenses), translated to Turkish. Because the most restrictive component is CC BY-SA 4.0, this model card is released under CC BY-SA 4.0; verify that this suits your use case and comply with the ShareAlike terms.
  • Evaluation datasets belong to their respective authors (see the trmteb org and emrecan/stsb-mt-turkish).

Citation

This model:

@misc{e5-tr-nli,
  title  = {e5-tr-nli: A Turkish Sentence Embedding Model},
  note   = {Fine-tuned from intfloat/multilingual-e5-base on Turkish NLI triplets},
  year   = {2026}
}

TR-MTEB benchmark:

@inproceedings{baysan-gungor-2025-trmteb,
  title     = {{TR-MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations},
  author    = {Baysan, Mehmet Selman and G{\"u}ng{\"o}r, Tunga},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
  year      = {2025}
}

Base model (E5):

@article{wang2024multilingual,
  title   = {Multilingual E5 Text Embeddings: A Technical Report},
  author  = {Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
  journal = {arXiv preprint arXiv:2402.05672},
  year    = {2024}
}

Reproduction

The full pipeline (data prep, Optuna sweep, training, and TR-MTEB evaluation) is scripted:

  • train_embedding_tr.py--mode sweep | train | eval | mteb
  • eval_trmteb_hf.py — TR-MTEB evaluation straight from HuggingFace trmteb/* datasets
  • config.py, Makefile, requirements.txt

To publish (disabled by default — no secrets are used):

from sentence_transformers import SentenceTransformer
SentenceTransformer("models/e5-tr-nli-final").push_to_hub("thealper2/intfloat-multilingual-e5-base-tr-nli")
Downloads last month
11
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thealper2/intfloat-multilingual-e5-base-tr-nli

Finetuned
(147)
this model

Datasets used to train thealper2/intfloat-multilingual-e5-base-tr-nli

Paper for thealper2/intfloat-multilingual-e5-base-tr-nli

Evaluation results