Text-to-Speech
PyTorch
English
tts
glow-tts
hifi-gan
coqui-tts
supratts

SupraTTS-0.1-Beta

Hero-Design

About

SupraTTS-0.1-Beta is an English single-speaker TTS model. Glow-TTS, trained from scratch on LJSpeech in Coqui-TTS.

Successor to Flare-TTS-v1.5. Same dataset, almost the same size. The only goal was to sound better.

  • Fixed a learning-rate scheduler bug (the LR scheduler was stepping per-epoch instead of per-step, so v1.5 never actually reached its target learning rate) !!
  • Blank tokens between input tokens
  • Replaced the deterministic duration predictor with a stochastic one (from VITS), so timing isn't identical every time
  • bf16, plus a Triton kernel for Monotonic Alignment Search (Super-MAS)
  • HiFi-GAN vocoder trained longer and from scratch on ground-truth mels

Encoder, decoder, and the dataset are the same as v1.5.

29.6M parameters (acoustic model) + HiFi-GAN v1 vocoder (14M).

Samples

1. Introductory text

Text: "This is the first sample generated by SupraTTS zero point one Beta... Have fun trying it out."
Output:

2. Long Wikipedia text

Text: "Artificial intelligence (AI)" to "at least as well as a human." from https://en.wikipedia.org/wiki/Artificial_intelligence; adapted
Output:

3. Long numbers and complex words

Text: "On November 12th, 1998, quintessential astrophysicists meticulously calculated approximately 3,456,789 light-years, uncovering extraordinary phenomena through synchronization, authorization, and unprecedented logarithmic measurements."
Output:

Benchmarks

Evaluated on 10 out-of-domain sentences + 10 held-out LJSpeech sentences, 5 seeds each. UTMOS is an automatic MOS estimate (1-5). F0-Std is pitch variation inside a sentence, in semitones. WER is Whisper's word error rate, used here as a mispronunciation check.

System UTMOS (higher is better) F0-Std [st] WER % (lower is better) Tempo vs. Ground Truth
Ground Truth (real recordings) 4.34 ± 0.08 3.51 2.58 1.000
Flare-TTS-v1.5 3.05 ± 0.08 1.87 3.72 1.087
SupraTTS-0.1-Beta (this model) 3.59 ± 0.07 1.01 3.64 1.004

UTMOS is up 0.53 vs v1.5. The old model dragged (1.087x the original speech rate); this one sits at 1.004x. WER is basically the same.

We also tried fine-tuning HiFi-GAN on the acoustic model's own mels (GTA). It made things worse: UTMOS 3.34, noisier pitch. LJSpeech is probably too small and too uniform for that, and the GT vocoder was already trained long enough. This release uses the ground-truth-mel vocoder only.

F0-Std is still well below the recordings. Pitch is stable, just flatter than a real speaker. That's the next thing to fix.

Config

  • Architecture: Glow-TTS (12 flow blocks, 192 hidden channels, rel-pos-transformer encoder, 6 layers)
  • Stochastic Duration Predictor + HiFi-GAN v1 vocoder
  • Sample rate: 22050 Hz, 80 mel bands
  • Dataset: LJSpeech-1.1 (single English female speaker, ~24h)
  • Full configs: config.json (acoustic model), vocoder_config.json

Training

Acoustic model Vocoder
Steps ~39,000 (145 epochs) ~103,000 (GT mels)
Batch size 48 16
Precision bf16 fp16
Hardware 1x RTX 5060 Ti, 16GB VRAM 1x RTX 5060 Ti, 16GB VRAM
Wall-clock time ~25 hours ~35 hours

How to Run Locally

Make sure to have installed: Python 3.10+, PyTorch with CUDA, Coqui-TTS (pip install coqui-tts).

mkdir supratts-0.1-beta && cd supratts-0.1-beta

wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/infer_v2.py
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/flare_glowtts.py
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/config.json
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/vocoder_config.json
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/model.pth
wget https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta/resolve/main/vocoder.pth

python3 infer_v2.py

Change TEXT at the top of infer_v2.py. Wav lands in output_v2.wav.

Thanks

What's Next

For Supra-TTS-1, if we get there:

  • Stochastic pitch, better vocoder
  • More than English
  • More than one speaker
  • Inline style tags ([whisper], [scream], [laughing], etc...)
  • Zero-shot cloning
Downloads last month
69
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train SupraLabs/SupraTTS-0.1-Beta

Space using SupraLabs/SupraTTS-0.1-Beta 1

Collection including SupraLabs/SupraTTS-0.1-Beta

Papers for SupraLabs/SupraTTS-0.1-Beta