ReNikud — armROBUST6clean (blind)

Hebrew text in, IPA out, with stress marks — the part unpointed Hebrew leaves to context. For TTS front-ends, lexicon work and pronunciation research.

This is the honest public model. It is a low-LR touch-up of armCLEAN on a corpus-derived payload that never opened test.tsv for selection. Decode is the published narrow armCLEAN path (no W4 wall cells). Numbers below are gold v2 and carry no †‡.

The shipped rescoring store is the CONS-scrubbed extended attested-base table (v2 corpus + commission_blind injects). Raw datastore_ext_v6blind is not what is in this repo: that file inherits 88 CONS/test.tsv-derived readings from datastore v5, and those keys are the entire 9.87 test.tsv movement. They are stripped here.

word accuracy by category

What this is, and what it is not

this repo (armROBUST6clean) do not confuse with
Init armCLEAN CONS / wall-trained checkpoints
Decode narrow legality, scrubbed extended store, τ = 0.7 W4 widened stack, datastore v5, τ = 0.5
test.tsv mark blind — 10.64 7.23 †‡ / 6.91 †‡ / 9.87 (v6blind with CONS keys)
Ranking axis pooled ILSpeech 9.07 a test.tsv number from the injection lineage

If a card, chart or Space still quotes 7.23 or 6.91 as “RenikudPlus”, that is the benchmark-shaped CONS stack, not these weights. Do not quote 9.02 / 9.87 either — those are v6blind before CONS keys were removed.

Results

Pooled ILSpeech (15,300 words, exact-MAP) — the ranking axis:

line WER_strict
this ONNX, bare 9.31 (tie with armCLEAN bare, 87f/87b, p = 1.000)
+ original 130k store, τ = 0.2 9.23 (tie with armCLEAN+rescA, 67f/67b, p = 1.000)
+ scrubbed extended store, τ = 0.7 9.07 (51f/27b vs orig-store 9.23, p = 0.0088)

9.07 vs armCLEAN+rescA is 97f/73b, p = 0.077 — not claimed as a ranking win over the champion. The 0.16 pp is a decode-layer gain on this checkpoint's original store.

test.tsv (3,110 targets, gold v2, 2026-08-21 owner patch):

config WER accuracy mark
armCLEAN bare 13.86 / 13.67* ~86.3% previous champion
armCLEAN + original store, τ = 0.2 13.54 / 13.34* ~86.7% deployment point #1
this ONNX, bare 10.80 89.2% blind
this ONNX + original store, τ = 0.2 10.68 89.3% blind
this ONNX + scrubbed store, τ = 0.7 10.64 89.4% blind

* Two gold-v2 dumps of armCLEAN exist (13.86 certified vs 13.67 current frozen dump). The −2.73 pp held-out figure below uses the dumps that produced 10.68.

The store upgrade does not move test.tsv (332 → 331 errors). A raw v6blind decode scores 9.87; 24 of those 25 words are CONS-store snaps and are not on this card.

Held-out test.tsv (surfaces never in the payload): 13.09% → 10.30% (−2.73 pp, −85 words on 3,049 / 3,110 positions, orig-store τ = 0.2). 91% of the net gain is off commissioned surfaces. Payload selection did not use test.tsv (selection_used_test = 0).

Per-category WER, gold v2, original store τ = 0.2 (diagnostics only — a category cannot rank models; this is the checkpoint/test.tsv story, not the store upgrade):

category armCLEAN bare this model
Gender 0.66 0.66
Min. Stress Pairs 7.33 3.33
Stress Homographs 7.88 2.96
Names 21.33 11.33
Acronyms 21.05 13.82
Slang 23.72 17.95
Penult. Stress 12.58 5.30
Foreign 19.35 14.19
Rare Phonemes 43.05 38.41
Colloquial 49.67 47.68
ILSpeech-test 6.95 6.11
OVERALL 13.67 10.68

Usage

pip install onnxruntime numpy

Download renikud_onnx.py and a model file from this repo into the same folder:

from renikud_onnx import G2P

g2p = G2P("model_int8.onnx")

g2p.phonemize("שלום, מה נשמע?")
# ʃalˈom, mˈa niʃmˈa?

g2p.phonemize("הלכתי לספר וקראתי ספר בזמן שהוא ספר כמה אצבעות יש לו")
# halˈaχti lasapˈaʁ vekaʁˈati sˈefeʁ bizmˈan ʃehˈu safˈaʁ kˈama ʔetsbaʔˈot jˈeʃ lˈo
#            barber        book                    counted

Four occurrences of ספר, four different words, each resolved from context.

The graph alone is the bare decode (pooled 9.31 / test.tsv 10.80). The published 9.07 / 10.64 add a rescoring pass against corpus_datastore.json at τ = 0.7. That layer is corpus-side Python over candidate energies; it is not inside the ONNX file. ONNX weights are unchanged from the previous card (original-store 9.23 / 10.68).

Speaker and addressee

Hebrew inflects for who is speaking and who is addressed; the text usually shows neither. Both controls default to 0 (unknown), 1 = male, 2 = female.

g2p.phonemize("אני רוצה להגיד לך משהו חשוב", speaker=2, target_speaker=1)
speaking → addressing אני רוצה להגיד לך משהו חשוב
man → man ʔanˈi ʁotsˈe lehaɡˈid leχˈa mˈaʃu χaʃˈuv
man → woman ʔanˈi ʁotsˈe lehaɡˈid lˈaχ mˈaʃu χaʃˈuv
woman → man ʔanˈi ʁotsˈa lehaɡˈid leχˈa mˈaʃu χaʃˈuv
woman → woman ʔanˈi ʁotsˈa lehaɡˈid lˈaχ mˈaʃu χaʃˈuv

speaker sets רוצה, target_speaker sets לך. Four readings of one unchanged sentence.

Numbers and digits

Digits are expanded to Hebrew words before phonemization, with gender agreement taken from the counted noun:

g2p.phonemize("יש לי 3 ילדים")        # jˈeʃ lˈi ʃloʃˈa jeladˈim
g2p.phonemize("יש לי 3 בנות")         # jˈeʃ lˈi ʃalˈoʃ banˈot
g2p.phonemize("המחיר 1250 שקלים")     # hameχˈiʁ ʔˈelef matˈajim veχamiʃˈim ʃkalˈim

Files

file what
model.onnx fp32, exact-MAP metadata, narrow legality (no W4). Bit-identical to PyTorch on all 15,787 pooled words and 1,653 test.tsv sentences.
model_int8.onnx dynamic int8, ~4× smaller. Pooled 9.35 (+0.04 pp), 99.19% word agreement vs fp32.
model.safetensors PyTorch weights (armROBUST6clean/last)
corpus_datastore.json scrubbed extended attested-base store (v2 corpus + 121 commission_blind injects; 88 CONS-gained keys removed). τ = 0.7
renikud_onnx.py wrapper — decoding, long-input windowing, number expansion
arch.json / tokenizer*.json checkpoint sidecar

Recipe (short)

  • Checkpoint: armROBUST6clean/last — 4,000-step run-to-cap touch-up of armCLEAN
  • Diet: stage3_v4robust6blind_diet (corpus-derived; selection_used_test = 0)
  • Decode: narrow src/phonology table; not the CONS5/W4 stack
  • Rescoring: scrubbed extended inventory (corpus_datastore.json), τ = 0.7 (val249 unique min on this store+base pair; val249 is a selector, not a result)
  • Store honesty: v6blind = v5 (CONS5) + 121 blind injects. The 121 are commission_blind; the 88 inherited CONS-gained readings are not in this file. See STORE_HONESTY.md in the training repo.

Citation

@misc{melichov2026renikud,
  title={ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion},
  author={Maxim Melichov and Yakov Kolani and Morris Alper},
  year={2026},
  note={Code and models forthcoming},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support