BashkirRoBERTa NER
ONNX models for finding person names in Bashkir text, with a compact CPU default.
Overview
BashkirRoBERTa NER is a person-name tagger fine-tuned from the BashkirRoBERTa
masked-language encoder. It recognizes people (PER) and returns spans in the
original text. The release includes FP32 and dynamically quantized INT8 ONNX
graphs, the original SentencePiece tokenizer, and a small standalone runtime.
The INT8 model is the default for CPU use.
| At a glance | |
|---|---|
| Task | Bashkir person named entity recognition (PER) |
| Default artifact | onnx/model_int8.onnx |
| Source | A Bashkir-language encoder and person-entity training annotations |
| Version / license | v1 / Apache-2.0 |
Contents
Files and Configurations
| File | Purpose | Size |
|---|---|---|
onnx/model_int8.onnx |
CPU default, dynamic INT8 | 82.4 MB |
onnx/model_fp32.onnx |
FP32 reference with PyTorch tag parity | 200.3 MB |
spm_bashkir_bert_16k.model |
SentencePiece tokenizer | 0.59 MB |
runtime.py |
Word alignment, BIO spans, character offsets | — |
config.json |
Architecture, label order, runtime contract | — |
benchmark_summary.json |
Aggregate evaluation without source text | — |
example.py |
Download and run the default model | — |
META.json |
Release passport and artifact hashes | — |
LICENSE |
Full license text | — |
SHA256SUMS |
Checksums of the public files | — |
Model Architecture
| Property | Value |
|---|---|
| Encoder | BashkirRoBERTa, eight Pre-LayerNorm Transformer blocks |
| Head | Linear word-level classifier |
| Labels | O, B-PER, I-PER |
| Tokenization | SentencePiece, first piece of each original word is labeled |
| Input | input_ids, int64, dynamic batch and sequence dimensions |
| Output | logits, float, [batch, sequence, 3] |
| Maximum context | 256 SentencePiece tokens, including boundary tokens |
Examples
Outputs below were produced by the included INT8 runtime:
| Input | Found people |
|---|---|
Рәми Ғарипов һәм Мостай Кәрим тураһында һөйләштек. |
Рәми Ғарипов; Мостай Кәрим |
Мәрйәм менән Илнур мәктәпкә китте. |
Мәрйәм; Илнур |
Өфө ҡалаһында яңы мәктәп асылды. |
None |
Method
BashkirRoBERTa base encoder
├── 1. tokenization with first-subword alignment mapping
├── 2. fine-tune token classification head on Bashkir person-entity corpus
├── 3. export FP32 ONNX inference graph
└── 4. dynamic INT8 quantization & word-level BIO decoding
The encoder was fine-tuned with Bashkir person-entity annotations; its encoder layers and classification head changed. The original masked-language ONNX graph cannot be combined with this head to reproduce the release. Training data and evaluation text are not included. The tokenizer and first-subword alignment are provided so ONNX scores can be mapped back to complete words.
Benchmark
On a held-out, human-verified person-entity set, exact person-span micro F1 is
0.8594 for FP32 and 0.8589 for INT8 when run one sentence at a time
through the included adapter. FP32 makes no word-tag changes relative to the
PyTorch checkpoint. The separately authored 100-sentence stress suite gives span F1
0.9294 (FP32) and 0.9231 (INT8), with 92 sentences entirely correct in
each format. The stress suite is a targeted diagnostic, not an estimate of
population accuracy. Dynamic INT8 outputs can vary slightly with batch
composition; the figures above use the public single-sentence runtime. Counts
and the BIO scoring convention are in
benchmark_summary.json.
Quality and Use
Use this model to find mentions of people in Bashkir text, including inflected
names and multiword names. The runtime.py adapter returns each span with
character offsets. predict_tokens also accepts pre-tokenized words for
evaluation workflows. The default INT8 graph is smaller while preserving nearly
all of the observed gold-set score; FP32 is available when exact checkpoint
parity matters.
Limitations
- Only person entities are labeled. Places and organizations have no output class.
- Rare names can be missed. Long names and adjacent people can be split or merged.
- Names inside institution names can be mistaken for person mentions.
- Inputs beyond the context limit are truncated; the runtime reports
truncated: true. - Word boundaries from noisy OCR, unusual punctuation, and mixed writing systems can differ from the training tokenization.
- The authored stress suite is too small and structured to replace the held-out human-verified benchmark.
Related Resources
- BashkirRoBERTa is the underlying masked-language model; use it for fill-mask work and new fine-tuning tasks rather than person-name extraction.
Usage
pip install onnxruntime sentencepiece numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download
model_dir = snapshot_download(
"failed09/bashkir-roberta-ner",
allow_patterns=["onnx/model_int8.onnx", "spm_bashkir_bert_16k.model",
"config.json", "runtime.py"],
)
sys.path.insert(0, model_dir)
from runtime import BashkirPersonNER
ner = BashkirPersonNER(model_dir)
print(ner.predict("Рәми Ғарипов һәм Мостай Кәрим тураһында һөйләштек."))
For the FP32 graph, download onnx/model_fp32.onnx as well and construct
BashkirPersonNER(model_dir, precision="fp32"). The included example.py
runs the default configuration.
License
The model weights, tokenizer and runtime are distributed under the Apache-2.0 license. The model is fine-tuned from BashkirRoBERTa, which is also released under Apache-2.0. Training texts and benchmark sentences are not included; rights in the original texts remain with their respective owners. Contact the maintainer through the Hub for provenance or removal requests.
Citation
@software{failed09_bashkir_roberta_ner_2026,
title = {BashkirRoBERTa NER},
author = {failed09},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/failed09/bashkir-roberta-ner},
note = {ONNX person named entity recognition for Bashkir}
}
Open Bashkir Data and Sources 🐝
This release is part of an open-source effort to support the development, preservation and practical use of the Bashkir language. Other related models, datasets and tools are available on the author's Hugging Face profile.
The author does not claim ownership or authorship of the source texts or other materials used to derive this release; rights and licensing remain with the original authors, publishers and dataset providers. Source texts are not redistributed in this repository, so users should follow the licenses and attribution requirements of the relevant upstream resources.
- Downloads last month
- 25