kokoro-g2p
These are the English text front-end models that voxr's Kokoro-82M pipeline reads at run time. They mirror the g2p/ folder voxr expects beside a Kokoro checkpoint:
<kokoro dir>/g2p/tagger/ tagger.safetensors, tagger.json
<kokoro dir>/g2p/oov/ config.json, model.safetensors, vocab.json
tagger/
This is spaCy's en_core_web_sm 3.8.0 part-of-speech tagger, converted to safetensors.
- Contents: the embedding, encoder and softmax weights.
tagger.json: the feature attributes, the hash seeds and the label set.- Use: misaki's pipeline needs the tags to choose between POS-dependent pronunciations.
- License: the weights derive from
en_core_web_sm(MIT, Explosion AI).
oov/
This is a small Llama-style character model that predicts Kokoro phonemes for words missing from misaki's lexicons.
- Input:
<bos> <dialect> graphemes <sep>. - Output: phonemes until
<eos>. - Training data: misaki 0.9.4's US and GB gold and silver lexicons (Apache-2.0, https://github.com/hexgrad/misaki).
- Vocab:
vocab.jsonholds the token ids. Phoneme symbols are Kokoro-82M's.
License
- Repository: Apache-2.0.
tagger/: also subject to the MIT license of spaCy'sen_core_web_sm.