YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other

marian-uk-verbalizer

A 52.7M-parameter Marian sequence-to-sequence model that converts written Ukrainian text into the form it should be spoken — the text-normalization front end for a Ukrainian TTS pipeline. It expands Roman numerals (with correct case/gender agreement), dates, currencies, coordinates, phone numbers, IBANs, card numbers, math expressions, symbols, abbreviations and acronyms into spoken words, while leaving ordinary prose untouched.

Потрібно прочитати XX розділ до понеділка.
  -> Потрібно прочитати двадцятий розділ до понеділка.

Я пропрацював у ФБР двадцять років.
  -> Я пропрацював у еф-бе-ер двадцять років.

Full training/evaluation code, the router this model is meant to sit behind, and per-domain evaluation scripts: GitHub repository.

A quantized int8 CTranslate2 export of this same checkpoint (52 MB, ~99.1% output-identical) is published separately at aloudreader/marian-uk-verbalizer-ct2-int8.

Model description

  • Architecture: MarianMTModel, 6 encoder + 6 decoder layers, d_model=512, 8 attention heads, 2048 FFN dim — 52.7M parameters.
  • Tokenizer: shared 16k-token SentencePiece vocabulary (source and target).
  • Decoder start token is <s> (token id 0), not the usual pad token — this matters if you re-export the model (see the GitHub repo's CTranslate2 conversion notes).
  • Trained with input_preprocessing: model-routing-v1: the model always receives the original, unmodified sentence and its output is used raw — no rule-based pre- or post-processing anywhere in the pipeline. This is a hard project constraint, not an implementation detail: see AGENT.md in the GitHub repository.

Intended use

Sits behind a lightweight shape-based router that sends plain Cyrillic prose straight through and everything else (digits, Latin letters, symbols, all-caps runs, …) to this model. Feeding the model plain prose it doesn't need to touch is safe — the training data includes plenty of identity (unchanged) targets — but the router avoids the extra inference cost. The router implementation (regex/shape only, no word lists) is in the GitHub repository at src/verbalizer/text.py.

This model is Ukrainian-specific and was trained and evaluated on general prose plus template-generated numeric/structured text. It is not intended for other languages or for tasks other than pre-TTS text normalization.

How to use

from transformers import MarianMTModel, MarianTokenizer

tokenizer = MarianTokenizer.from_pretrained("aloudreader/marian-uk-verbalizer")
model = MarianMTModel.from_pretrained("aloudreader/marian-uk-verbalizer")

text = "Потрібно прочитати XX розділ до понеділка."
batch = tokenizer([text], return_tensors="pt", padding=True)
out = model.generate(**batch, max_new_tokens=128, num_beams=1)
print(tokenizer.batch_decode(out, skip_special_tokens=True)[0])

Greedy decoding (num_beams=1) is what this checkpoint was evaluated with; beam search is supported but was not found to improve the benchmark scores enough to justify the extra latency.

Training data

  • The full skypro1111/uk-text-normalization dataset (a 3,597-row slice is frozen out as an evaluation-only benchmark and never trained on).
  • Synthetic rows generated offline with num2words and pymorphy3 for domains and grammatical cases the base dataset covers thinly — most recently, Roman-numeral-plus-noun case/gender agreement across many governing verbs, prepositions and nouns per case, so the model learns to read the noun's ending rather than memorizing one preposition per case. Every generated target is cross-checked by parsing it back to a numeric value and comparing it against the source before it enters training.

No text extracted from books is included in or was used to build this model's public evaluation data (see the GitHub repository's data policy).

Evaluation

Frozen 3,597-row benchmark (held out of training, no data leakage):

Metric Score
Exact string match 85.2%
Word error rate 3.0%
Character error rate 1.2%
Numeric value fidelity (every digit sequence read back and compared to source) 100.0%
Digit sequence fidelity 100.0%
Identity accuracy (plain-prose rows left untouched) 100.0%

Generated per-domain evaluation (lenient frame+slot scoring — the fixed part of each template sentence must match exactly, and each slot must be read as one of its accepted spoken forms):

Domain Correct Rows
Coordinates 100.0% 500
Phone numbers 100.0% 500
Roman numeral + noun case agreement 99.0% 301
Context (held-out book sentences, not published) 97.4% 500
English words embedded in Ukrainian 88.4% 500
Card numbers / IBAN 86.8% 500
Acronyms (letter-by-letter reading) 85.5% 207

Limitations

  • Acronym and card/IBAN reading are the weakest domains (~85–87%); most failures are single-letter substitutions in acronym spelling or a single misread digit group in long card numbers.
  • The model has only ever seen Roman numerals I–XXXIX (by project convention, L/C/D/M are treated as ordinary letters — vitamin C, size L, flight D12 — since they collide with common non-numeral usage in Ukrainian text).
  • Like any seq2seq model, it can occasionally hallucinate or drop a token on out-of-distribution input; there is no rule-based safety net downstream — by design, its raw output is what gets spoken.

License

MIT. See the GitHub repository for training/evaluation code under the same license. The skypro1111/uk-text-normalization training dataset has its own license — check it before redistributing data derived from this model.

Downloads last month
130
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aloudreader/marian-uk-verbalizer

Finetunes
1 model

Dataset used to train aloudreader/marian-uk-verbalizer

Evaluation results

  • Exact match on frozen 3,597-row benchmark slice
    self-reported
    85.210
  • Word error rate on frozen 3,597-row benchmark slice
    self-reported
    2.980
  • Character error rate on frozen 3,597-row benchmark slice
    self-reported
    1.200