Instructions to use micahchoo/piper-kn with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Piper
How to use micahchoo/piper-kn with Piper:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Piper voice: Kannada, SYSPIN male (medium)
A Piper voice that reads Kannada in
one male voice. It is a VITS model in ONNX, 22,050 Hz, single speaker, driven
by espeak-ng's kn phonemes like every Piper voice. It reads aloud in the
md-translator.
| File | What it is |
|---|---|
kn_IN-syspin_male-medium.onnx |
The model, 63.5 MB |
kn_IN-syspin_male-medium.onnx.json |
Its config: phoneme map, sample rate, inference scales |
python3 -m piper -m kn_IN-syspin_male-medium.onnx -f hello.wav -- 'ನಮಸ್ಕಾರ, ನೀವು ಹೇಗಿದ್ದೀರಿ?'
How it was made
Fine-tuned from the English en_US-lessac-medium checkpoint
(rhasspy/piper-checkpoints)
on 3,157 clips (7.29 h) of the SYSPIN Kannada male speaker, for 75 epochs
(about 14,500 steps, batch 16, bf16) on one AMD Radeon 8060S. This file is
epoch 2239 of that run.
How good it is
Measured on 63 held-out clips the voice never trained on, against the real recordings of the same sentences. UTMOS predicts naturalness from 1 to 5; it learned from English speech, so read it as a comparison with this speaker's own recordings, not as an absolute score.
| UTMOS | |
|---|---|
| The real recordings | 3.76 |
| This voice | 3.20 |
It speaks about 4% faster than the speaker (length ratio 0.96, range 0.79 to 1.14). Listeners hear a few misplaced pauses and some roughness in the voice. 53 of the training transcripts (1.7%) carried SYSPIN's number markup, the digits and then the spoken words, so text with that markup reads badly.
Licence and credits
- Recordings: SYSPIN, Indian Institute of Science, Bengaluru, under CC BY 4.0. https://syspin.iisc.ac.in/datasets
- Starting weights: the lessac voice, trained on the Blizzard Challenge 2013 Lessac recordings, which are under a research licence. Whether that licence binds weights fine-tuned from them is unsettled. Treat this voice as free for non-commercial use only.
- Training code: piper1-gpl, GPL 3.0.
- Downloads last month
- 8