You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Chagatai NER — mBERT

The first named entity recognition model for Chagatai, a historical Turkic language written in Arabic script, used across Central Asia from the 13th to the 20th century and the literary ancestor of modern Uzbek and Uyghur.

Fine-tuned from bert-base-multilingual-cased on a semi-automatically annotated corpus of 181 manuscript pages from Shajarei Türki (شجرة ترکی) by Abu al-Ghazi Bahadur Khan (17th c.).

This model accompanies the paper "Towards Named Entity Recognition for Chagatai: First Annotated Corpus and Multilingual BERT Baselines" (Omirbay & Imangaliyev, IEEE 3rd International Student Conference on Digital Generation, 2026). See Citation below.

Model description

  • Base model: bert-base-multilingual-cased (179M params, 12 layers, WordPiece tokenizer)
  • Task: Token classification (NER), 7 entity types, IOB2 scheme
  • Why mBERT: of four multilingual models tested (mBERT, XLM-R base, XLM-R large, Turkish BERT), mBERT performed best — attributed to WordPiece fragmenting Arabic-script tokens less than XLM-R's SentencePiece.
  • Entity types: PER (person), LOC (location), ORG (dynasty / political entity), DATE, EVENT, MISC (ethnic/religious groups, honorifics), WORK (titles of written works)

Intended uses & limitations

Intended for NER on Chagatai (and possibly closely related historical Turkic/Arabic-script text) for digital humanities and historical linguistics research. Not intended for production or commercial use (see License).

Known limitations:

  • Trained on ~2,200 sentences from a single source text — generalization to other authors/genres/periods is untested.
  • WORK (3 test examples) and EVENT are not reliably learnable at this corpus size; expect near-zero recall on these two labels.
  • No corpus-wide inter-annotator agreement metric is available; only a 30-page subset (~16.6%) was checked by a second, expert annotator.
  • No tokenizer used here was trained on Arabic-script Turkic text, so subword fragmentation still limits input quality.

Training data

  • Source: Shajarei Türki, a 17th-century genealogical chronicle of the Turkic and Mongol peoples by Abu al-Ghazi Bahadur Khan, digitized via the Kazakh National Corpus (Institute of Linguistics named after Akhmet Baitursynuly), Historical Subcorpus.
  • Scale: 181 pages, ~66,059 tokens, ~2,200 sentences, ~8,000 entity tokens (12% of the corpus).
  • Annotation: semi-automatic — a Gemini vision-language model produced initial IOB2 tags per page from the manuscript image plus Cyrillic transliteration and Kazakh translation as weak reference signals; the first author reviewed/corrected all 181 pages, and a Chagatai linguist (co-author) additionally checked a 30-page subset.
  • Split: 70% train / 15% validation / 15% test, sentence-level random split (train ≈1,395–1,400, val ≈300, test ≈300 sentences; test-set entity support = 1,568 tokens).
  • Dataset release: the annotated corpus itself will be published on the Hugging Face Hub separately.

Training procedure

Best configuration ("V3" in the paper):

  • Standard cross-entropy loss (no class weighting — weighted CE was tried and hurt performance, collapsing precision to ~0.15)
  • 3× oversampling of sentences containing rare entities (DATE/MISC/WORK/EVENT)
  • 2-phase training: bottom 8 layers frozen for the first 2 of 15 epochs, then unfrozen
  • Learning rate 5e-6, batch size 32 (effective 64 with gradient accumulation ×2), 15 epochs, bf16 on an NVIDIA A100
  • Inference uses aggregation_strategy="first" (not "simple"), which fixes Arabic subword fragmentation during entity aggregation

Evaluation results

Test-set F1/Precision/Recall across all four models and three configs:

Config Model F1 P R
V1 (baseline) mBERT 0.5232 0.5875 0.4716
V1 XLM-R base 0.3666 0.4581 0.3056
V1 Turkish BERT 0.3214 0.4355 0.2546
V2 (augmented, worse) mBERT 0.2306 0.1531 0.4666
V2 XLM-R large 0.2331 0.1510 0.5107
V2 XLM-R base 0.1978 0.1257 0.4641
V2 Turkish BERT 0.1264 0.0762 0.3702
V3 (this model) mBERT 0.5947 0.5901 0.5995
V3 XLM-R large 0.5528 0.5996 0.5128
V3 XLM-R base 0.4189 0.5120 0.3544
V3 Turkish BERT 0.3936 0.4283 0.3640

Per-entity results, this model (V3 mBERT, test set):

Entity F1 P R Support
PER 0.71 0.65 0.78 774
LOC 0.57 0.52 0.63 327
ORG 0.51 0.52 0.50 177
DATE 0.26 0.34 0.21 56
MISC 0.22 0.46 0.14 231
WORK 0.00 0.00 0.00 3
Micro avg 0.5947 0.5901 0.5995 1,568

Usage

from transformers import pipeline

ner = pipeline(
    "ner",
    model="Darkhan04/chagatai-ner-mbert",
    tokenizer="Darkhan04/chagatai-ner-mbert",
    aggregation_strategy="first",  # required — see Training procedure
)

text = "چنکز خان نینک ایکنچی اوغلی چغتای خان نینک اولادیندین ماور النهر دا"
print(ner(text))

Citation

If you use this model, please cite the accompanying paper:

@inproceedings{omirbay2026chagatai,
  author    = {Omirbay, Darkhan and Imangaliyev, Almas},
  title     = {Towards Named Entity Recognition for {Chagatai}: First Annotated Corpus and Multilingual {BERT} Baselines},
  booktitle = {2026 IEEE 3rd International Student Conference on Digital Generation},
  year      = {2026},
  address   = {Astana, Kazakhstan}
}

Note: BibTeX key/venue metadata above is a placeholder pending official proceedings publication — update year/pages/DOI once available.

Please also cite the base model and the source corpus:

@inproceedings{pires2019multilingual,
  title     = {How multilingual is Multilingual {BERT}?},
  author    = {Pires, Telmo and Schlinger, Eva and Garrette, Dan},
  booktitle = {Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics},
  year      = {2019}
}

@misc{qazcorpus2025,
  title        = {Kazakh National Corpus, Historical Subcorpus},
  author       = {{Institute of Linguistics named after Akhmet Baitursynuly}},
  year         = {2025},
  howpublished = {\url{https://qazcorpus.kz/}}
}

License

Released under CC-BY-NC 4.0 (non-commercial), consistent with the non-commercial access terms of the source manuscript corpus (Kazakh National Corpus). See LICENSE.

Acknowledgments

Prepared within Project BR28712621, "The Manuscript Heritage of Kazakhstan: Registration, Restoration, Scientific Cataloguing, Digitization, and Comprehensive Codicological Studies."

Contact

Darkhan Omirbay — School of Artificial Intelligence and Data Science, Astana IT University — domirbai04@gmail.com

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Darkhan04/chagatai-ner-mbert

Finetuned
(1002)
this model

Evaluation results

  • Micro F1 (test) on Chagatai NER (Shajarei Türki, 181 pages)
    self-reported
    0.595
  • Micro Precision (test) on Chagatai NER (Shajarei Türki, 181 pages)
    self-reported
    0.590
  • Micro Recall (test) on Chagatai NER (Shajarei Türki, 181 pages)
    self-reported
    0.600