SupraTTS-0.1-Beta
About
SupraTTS-0.1-Beta is an English single-speaker TTS model. Glow-TTS, trained from scratch on LJSpeech in Coqui-TTS.
Successor to Flare-TTS-v1.5. Same dataset, almost the same size. The only goal was to sound better.
- Fixed a learning-rate scheduler bug (the LR scheduler was stepping per-epoch instead of per-step, so v1.5 never actually reached its target learning rate) !!
- Blank tokens between input tokens
- Replaced the deterministic duration predictor with a stochastic one (from VITS), so timing isn't identical every time
- bf16, plus a Triton kernel for Monotonic Alignment Search (Super-MAS)
- HiFi-GAN vocoder trained longer and from scratch on ground-truth mels
Encoder, decoder, and the dataset are the same as v1.5.
29.6M parameters (acoustic model) + HiFi-GAN v1 vocoder (14M).
Samples
1. Introductory text
Text: "This is the first sample generated by SupraTTS zero point one Beta... Have fun trying it out."
Output:
2. Long Wikipedia text
Text: "Artificial intelligence (AI)" to "at least as well as a human." from https://en.wikipedia.org/wiki/Artificial_intelligence; adapted
Output:
3. Long numbers and complex words
Text: "On November 12th, 1998, quintessential astrophysicists meticulously calculated approximately 3,456,789 light-years, uncovering extraordinary phenomena through synchronization, authorization, and unprecedented logarithmic measurements."
Output:
Benchmarks
Evaluated on 10 out-of-domain sentences + 10 held-out LJSpeech sentences, 5 seeds each. UTMOS is an automatic MOS estimate (1-5). F0-Std is pitch variation inside a sentence, in semitones. WER is Whisper's word error rate, used here as a mispronunciation check.
| System | UTMOS (higher is better) | F0-Std [st] | WER % (lower is better) | Tempo vs. Ground Truth |
|---|---|---|---|---|
| Ground Truth (real recordings) | 4.34 ± 0.08 | 3.51 | 2.58 | 1.000 |
| Flare-TTS-v1.5 | 3.05 ± 0.08 | 1.87 | 3.72 | 1.087 |
| SupraTTS-0.1-Beta (this model) | 3.59 ± 0.07 | 1.01 | 3.64 | 1.004 |
UTMOS is up 0.53 vs v1.5. The old model dragged (1.087x the original speech rate); this one sits at 1.004x. WER is basically the same.
We also tried fine-tuning HiFi-GAN on the acoustic model's own mels (GTA). It made things worse: UTMOS 3.34, noisier pitch. LJSpeech is probably too small and too uniform for that, and the GT vocoder was already trained long enough. This release uses the ground-truth-mel vocoder only.
F0-Std is still well below the recordings. Pitch is stable, just flatter than a real speaker. That's the next thing to fix.
Config
- Architecture: Glow-TTS (12 flow blocks, 192 hidden channels, rel-pos-transformer encoder, 6 layers)
- Stochastic Duration Predictor + HiFi-GAN v1 vocoder
- Sample rate: 22050 Hz, 80 mel bands
- Dataset: LJSpeech-1.1 (single English female speaker, ~24h)
- Full configs:
config.json(acoustic model),vocoder_config.json
Training
| Acoustic model | Vocoder | |
|---|---|---|
| Steps | ~39,000 (145 epochs) | ~103,000 (GT mels) |
| Batch size | 48 | 16 |
| Precision | bf16 | fp16 |
| Hardware | 1x RTX 5060 Ti, 16GB VRAM | 1x RTX 5060 Ti, 16GB VRAM |
| Wall-clock time | ~25 hours | ~35 hours |
How to Run Locally
Make sure to have installed: Python 3.10+, PyTorch with CUDA, Coqui-TTS (pip install coqui-tts).
mkdir supratts-0.1-beta && cd supratts-0.1-beta
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/infer_v2.py
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/flare_glowtts.py
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/config.json
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/vocoder_config.json
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/model.pth
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/vocoder.pth
python3 infer_v2.py
Change TEXT at the top of infer_v2.py. Wav lands in output_v2.wav.
Thanks
- Keith Ito, LJSpeech
- Coqui-TTS (Idiap maintains it now)
- Kim et al., Glow-TTS
- Kim et al., VITS, stochastic duration predictor
- Kong et al., HiFi-GAN
- Supertone, Super-Monotonic-Align
What's Next
For Supra-TTS-1, if we get there:
- Stochastic pitch, better vocoder
- More than English
- More than one speaker
- Inline style tags (
[whisper],[scream],[laughing], etc...) - Zero-shot cloning
- Downloads last month
- 69
