Latensis — Turkish RoBERTa Base

Understanding Beyond the Visible

Latensis is a Turkish-specific RoBERTa base model trained from scratch on 1GB of curated Turkish text using the Hecemen Unigram 128k tokenizer — a tokenizer designed specifically for Turkish morphology.

Why "Latensis"?

The name is derived from the Latin root latent (hidden, concealed) — a direct reference to the latent representations that embedding models learn. The suffix -ensis (of, belonging to) completes the name: Latensis roughly means "of the latent space" or "that which belongs to what is hidden."

Model Details

Property Value
Architecture RoBERTa base
Parameters 184.5M
Vocabulary size 128,000
Tokenizer Hecemen Unigram 128k
Training data ~1GB Turkish text
Training steps 500,000
Val loss 3.21
Perplexity 25

Training Data

The model was trained on a carefully curated ~1GB Turkish corpus covering:

  • News articles
  • Books and literature
  • Web content
  • Legal and administrative texts
  • Agricultural and domain-specific texts

Benchmark Results

All fine-tuning experiments were conducted starting from this base model. Comparisons are against BERTurk (dbmdz/bert-base-turkish-cased).

Sentiment Analysis

Dataset Our Model BERTurk Training Data
TRSAv1 (Accuracy) 0.9027 0.8372 19k examples
TRSAv1 (F1-macro) 0.9026 0.8350 19k examples
Winvoker (Accuracy) 0.8534 0.6944 19k examples

Natural Language Inference → RTE

Dataset Our Model BERTurk Training Data
RTE (Accuracy) 0.7960 0.7780 50k NLI examples
RTE (F1-macro) 0.7959 0.7770 50k NLI examples

BERTurk NLI model was trained on 482k NLI examples. Our model achieves better results with only 50k examples (~10x less data).

Named Entity Recognition

Dataset Our Model BERTurk Training Data
WikiANN-TR (F1-macro) 0.9429 0.9522 20k examples

Within BERTurk's reported variance range (93.92 ± 0.07).

Semantic Textual Similarity

Metric Our Model Emrecan Training Data
Pearson 0.7687 0.8340 50k NLI + TrGLUE STS
Spearman 0.7976 0.8300 50k NLI + TrGLUE STS

Emrecan model uses 482k NLI + full STS-b-TR. Our model uses 50k NLI (~10x less).

QNLI

Dataset Our Model Training Data
TrGLUE QNLI (Accuracy) 0.8749 120k examples

Key Findings

✅ Outperforms BERTurk on Sentiment (TRSAv1 +6pp, Winvoker +16pp) ✅ Outperforms BERTurk on RTE with 10x less NLI training data ✅ Competitive NER performance within BERTurk's variance range ✅ Competitive STS with 10x less training data than Emrecan ✅ All results achieved with only 1GB training data

Usage

As a base model for fine-tuning

from transformers import RobertaModel, AutoTokenizer
import sentencepiece as spm
from huggingface_hub import hf_hub_download

# Load tokenizer
model_path = hf_hub_download(
    repo_id="mursideaki/hecemen-tokenizer-unigram-128k",
    filename="tr_unigram_tokenizer.model"
)
sp = spm.SentencePieceProcessor()
sp.load(model_path)

# Load model
model = RobertaModel.from_pretrained("mursideaki/latensis-roberta-base-tr")

Masked Language Modeling

from transformers import pipeline

mlm = pipeline(
    "fill-mask",
    model="mursideaki/latensis-roberta-base-tr"
)
# Note: Use with Hecemen tokenizer for best results

Companion Models & Tokenizers

Citation

@misc{latensis2026,
  author    = {Mürşide Aki},
  title     = {Latensis: A Turkish RoBERTa Base Model},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/mursideaki/latensis-roberta-base-tr}
}

License

MIT

Downloads last month
17
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support