Instructions to use aloudreader/marian-uk-verbalizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aloudreader/marian-uk-verbalizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="aloudreader/marian-uk-verbalizer")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("aloudreader/marian-uk-verbalizer") model = AutoModelForSeq2SeqLM.from_pretrained("aloudreader/marian-uk-verbalizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use aloudreader/marian-uk-verbalizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aloudreader/marian-uk-verbalizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aloudreader/marian-uk-verbalizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/aloudreader/marian-uk-verbalizer
- SGLang
How to use aloudreader/marian-uk-verbalizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "aloudreader/marian-uk-verbalizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aloudreader/marian-uk-verbalizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "aloudreader/marian-uk-verbalizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aloudreader/marian-uk-verbalizer", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use aloudreader/marian-uk-verbalizer with Docker Model Runner:
docker model run hf.co/aloudreader/marian-uk-verbalizer
YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other
marian-uk-verbalizer
A 52.7M-parameter Marian sequence-to-sequence model that converts written Ukrainian text into the form it should be spoken — the text-normalization front end for a Ukrainian TTS pipeline. It expands Roman numerals (with correct case/gender agreement), dates, currencies, coordinates, phone numbers, IBANs, card numbers, math expressions, symbols, abbreviations and acronyms into spoken words, while leaving ordinary prose untouched.
Потрібно прочитати XX розділ до понеділка.
-> Потрібно прочитати двадцятий розділ до понеділка.
Я пропрацював у ФБР двадцять років.
-> Я пропрацював у еф-бе-ер двадцять років.
Full training/evaluation code, the router this model is meant to sit behind, and per-domain evaluation scripts: GitHub repository.
A quantized int8 CTranslate2 export of this same checkpoint (52 MB, ~99.1%
output-identical) is published separately at
aloudreader/marian-uk-verbalizer-ct2-int8.
Model description
- Architecture: MarianMTModel, 6 encoder + 6 decoder layers,
d_model=512, 8 attention heads, 2048 FFN dim — 52.7M parameters. - Tokenizer: shared 16k-token SentencePiece vocabulary (source and target).
- Decoder start token is
<s>(token id 0), not the usual pad token — this matters if you re-export the model (see the GitHub repo's CTranslate2 conversion notes). - Trained with
input_preprocessing: model-routing-v1: the model always receives the original, unmodified sentence and its output is used raw — no rule-based pre- or post-processing anywhere in the pipeline. This is a hard project constraint, not an implementation detail: seeAGENT.mdin the GitHub repository.
Intended use
Sits behind a lightweight shape-based router that sends plain Cyrillic prose
straight through and everything else (digits, Latin letters, symbols,
all-caps runs, …) to this model. Feeding the model plain prose it doesn't
need to touch is safe — the training data includes plenty of identity
(unchanged) targets — but the router avoids the extra inference cost. The
router implementation (regex/shape only, no word lists) is in the GitHub
repository at src/verbalizer/text.py.
This model is Ukrainian-specific and was trained and evaluated on general prose plus template-generated numeric/structured text. It is not intended for other languages or for tasks other than pre-TTS text normalization.
How to use
from transformers import MarianMTModel, MarianTokenizer
tokenizer = MarianTokenizer.from_pretrained("aloudreader/marian-uk-verbalizer")
model = MarianMTModel.from_pretrained("aloudreader/marian-uk-verbalizer")
text = "Потрібно прочитати XX розділ до понеділка."
batch = tokenizer([text], return_tensors="pt", padding=True)
out = model.generate(**batch, max_new_tokens=128, num_beams=1)
print(tokenizer.batch_decode(out, skip_special_tokens=True)[0])
Greedy decoding (num_beams=1) is what this checkpoint was evaluated with;
beam search is supported but was not found to improve the benchmark scores
enough to justify the extra latency.
Training data
- The full
skypro1111/uk-text-normalizationdataset (a 3,597-row slice is frozen out as an evaluation-only benchmark and never trained on). - Synthetic rows generated offline with
num2wordsandpymorphy3for domains and grammatical cases the base dataset covers thinly — most recently, Roman-numeral-plus-noun case/gender agreement across many governing verbs, prepositions and nouns per case, so the model learns to read the noun's ending rather than memorizing one preposition per case. Every generated target is cross-checked by parsing it back to a numeric value and comparing it against the source before it enters training.
No text extracted from books is included in or was used to build this model's public evaluation data (see the GitHub repository's data policy).
Evaluation
Frozen 3,597-row benchmark (held out of training, no data leakage):
| Metric | Score |
|---|---|
| Exact string match | 85.2% |
| Word error rate | 3.0% |
| Character error rate | 1.2% |
| Numeric value fidelity (every digit sequence read back and compared to source) | 100.0% |
| Digit sequence fidelity | 100.0% |
| Identity accuracy (plain-prose rows left untouched) | 100.0% |
Generated per-domain evaluation (lenient frame+slot scoring — the fixed part of each template sentence must match exactly, and each slot must be read as one of its accepted spoken forms):
| Domain | Correct | Rows |
|---|---|---|
| Coordinates | 100.0% | 500 |
| Phone numbers | 100.0% | 500 |
| Roman numeral + noun case agreement | 99.0% | 301 |
| Context (held-out book sentences, not published) | 97.4% | 500 |
| English words embedded in Ukrainian | 88.4% | 500 |
| Card numbers / IBAN | 86.8% | 500 |
| Acronyms (letter-by-letter reading) | 85.5% | 207 |
Limitations
- Acronym and card/IBAN reading are the weakest domains (~85–87%); most failures are single-letter substitutions in acronym spelling or a single misread digit group in long card numbers.
- The model has only ever seen Roman numerals I–XXXIX (by project convention,
L/C/D/Mare treated as ordinary letters — vitamin C, size L, flight D12 — since they collide with common non-numeral usage in Ukrainian text). - Like any seq2seq model, it can occasionally hallucinate or drop a token on out-of-distribution input; there is no rule-based safety net downstream — by design, its raw output is what gets spoken.
License
MIT. See the GitHub repository
for training/evaluation code under the same license. The
skypro1111/uk-text-normalization training dataset has its own license —
check it before redistributing data derived from this model.
- Downloads last month
- 130
Model tree for aloudreader/marian-uk-verbalizer
Dataset used to train aloudreader/marian-uk-verbalizer
Evaluation results
- Exact match on frozen 3,597-row benchmark sliceself-reported85.210
- Word error rate on frozen 3,597-row benchmark sliceself-reported2.980
- Character error rate on frozen 3,597-row benchmark sliceself-reported1.200