BashkirRoBERTa NER

ONNX models for finding person names in Bashkir text, with a compact CPU default.

Overview

BashkirRoBERTa NER is a person-name tagger fine-tuned from the BashkirRoBERTa masked-language encoder. It recognizes people (PER) and returns spans in the original text. The release includes FP32 and dynamically quantized INT8 ONNX graphs, the original SentencePiece tokenizer, and a small standalone runtime. The INT8 model is the default for CPU use.

At a glance
Task Bashkir person named entity recognition (PER)
Default artifact onnx/model_int8.onnx
Source A Bashkir-language encoder and person-entity training annotations
Version / license v1 / Apache-2.0

Contents

Files and Configurations

File Purpose Size
onnx/model_int8.onnx CPU default, dynamic INT8 82.4 MB
onnx/model_fp32.onnx FP32 reference with PyTorch tag parity 200.3 MB
spm_bashkir_bert_16k.model SentencePiece tokenizer 0.59 MB
runtime.py Word alignment, BIO spans, character offsets —
config.json Architecture, label order, runtime contract —
benchmark_summary.json Aggregate evaluation without source text —
example.py Download and run the default model —
META.json Release passport and artifact hashes —
LICENSE Full license text —
SHA256SUMS Checksums of the public files —

Model Architecture

Property Value
Encoder BashkirRoBERTa, eight Pre-LayerNorm Transformer blocks
Head Linear word-level classifier
Labels O, B-PER, I-PER
Tokenization SentencePiece, first piece of each original word is labeled
Input input_ids, int64, dynamic batch and sequence dimensions
Output logits, float, [batch, sequence, 3]
Maximum context 256 SentencePiece tokens, including boundary tokens

Examples

Outputs below were produced by the included INT8 runtime:

Input Found people
Рәми Ғарипов һәм Мостай Кәрим тураһында һөйләштек. Рәми Ғарипов; Мостай Кәрим
Мәрйәм менән Илнур мәктәпкә китте. Мәрйәм; Илнур
Өфө ҡалаһында яңы мәктәп асылды. None

Method

BashkirRoBERTa base encoder
  ├── 1. tokenization with first-subword alignment mapping
  ├── 2. fine-tune token classification head on Bashkir person-entity corpus
  ├── 3. export FP32 ONNX inference graph
  └── 4. dynamic INT8 quantization & word-level BIO decoding

The encoder was fine-tuned with Bashkir person-entity annotations; its encoder layers and classification head changed. The original masked-language ONNX graph cannot be combined with this head to reproduce the release. Training data and evaluation text are not included. The tokenizer and first-subword alignment are provided so ONNX scores can be mapped back to complete words.

Benchmark

On a held-out, human-verified person-entity set, exact person-span micro F1 is 0.8594 for FP32 and 0.8589 for INT8 when run one sentence at a time through the included adapter. FP32 makes no word-tag changes relative to the PyTorch checkpoint. The separately authored 100-sentence stress suite gives span F1 0.9294 (FP32) and 0.9231 (INT8), with 92 sentences entirely correct in each format. The stress suite is a targeted diagnostic, not an estimate of population accuracy. Dynamic INT8 outputs can vary slightly with batch composition; the figures above use the public single-sentence runtime. Counts and the BIO scoring convention are in benchmark_summary.json.

Quality and Use

Use this model to find mentions of people in Bashkir text, including inflected names and multiword names. The runtime.py adapter returns each span with character offsets. predict_tokens also accepts pre-tokenized words for evaluation workflows. The default INT8 graph is smaller while preserving nearly all of the observed gold-set score; FP32 is available when exact checkpoint parity matters.

Limitations

  • Only person entities are labeled. Places and organizations have no output class.
  • Rare names can be missed. Long names and adjacent people can be split or merged.
  • Names inside institution names can be mistaken for person mentions.
  • Inputs beyond the context limit are truncated; the runtime reports truncated: true.
  • Word boundaries from noisy OCR, unusual punctuation, and mixed writing systems can differ from the training tokenization.
  • The authored stress suite is too small and structured to replace the held-out human-verified benchmark.

Related Resources

  • BashkirRoBERTa is the underlying masked-language model; use it for fill-mask work and new fine-tuning tasks rather than person-name extraction.

Usage

pip install onnxruntime sentencepiece numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download

model_dir = snapshot_download(
    "failed09/bashkir-roberta-ner",
    allow_patterns=["onnx/model_int8.onnx", "spm_bashkir_bert_16k.model",
                    "config.json", "runtime.py"],
)
sys.path.insert(0, model_dir)
from runtime import BashkirPersonNER

ner = BashkirPersonNER(model_dir)
print(ner.predict("Рәми Ғарипов һәм Мостай Кәрим тураһында һөйләштек."))

For the FP32 graph, download onnx/model_fp32.onnx as well and construct BashkirPersonNER(model_dir, precision="fp32"). The included example.py runs the default configuration.

License

The model weights, tokenizer and runtime are distributed under the Apache-2.0 license. The model is fine-tuned from BashkirRoBERTa, which is also released under Apache-2.0. Training texts and benchmark sentences are not included; rights in the original texts remain with their respective owners. Contact the maintainer through the Hub for provenance or removal requests.

Citation

@software{failed09_bashkir_roberta_ner_2026,
  title = {BashkirRoBERTa NER},
  author = {failed09},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/failed09/bashkir-roberta-ner},
  note = {ONNX person named entity recognition for Bashkir}
}

Open Bashkir Data and Sources 🐝

This release is part of an open-source effort to support the development, preservation and practical use of the Bashkir language. Other related models, datasets and tools are available on the author's Hugging Face profile.

The author does not claim ownership or authorship of the source texts or other materials used to derive this release; rights and licensing remain with the original authors, publishers and dataset providers. Source texts are not redistributed in this repository, so users should follow the licenses and attribution requirements of the relevant upstream resources.

Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support