SimpleTTS

Decoder-only transformer TTS, using WavTokenizer as the audio codec.

Samples

Text Audio
Sample 1
Sample 2
Sample 3
Sample 4
Sample 5

Setup

git clone https://github.com/FirstPotatoCoder/SimpleTTS.git
cd SimpleTTS
bash setup.sh
python scripts/download_weights.py

Run

python examples/run_inference.py "Hello, this is a quick test-run. I'm checking whether the model can handle longer input text without any issues. This sentence is roughly five times the length of the original short test phrase."

Or from Python:

from tts.inference import TTSPipeline

pipe = TTSPipeline(
tts_weights="weights/tts.pt",
wavtokenizer_weights="weights/wavtokenizer.ckpt",
wavtokenizer_config="configs/wavtokenizer_config.yaml",
)
pipe.generate("Some text to speak.", out_path="out.wav")

Limitations

  • The model sometimes hallucinates, generating babbling or garbled output on rare or unseen words โ€” likely due to the limited amount of training data.
  • Works best when generating ~3s to 15s of audio, matching the length distribution of its training data.
  • No chunking support yet โ€” the current repo only supports clip-by-clip generation, one sample at a time.

Credits

Audio tokenizer/codec: WavTokenizer (vendored under wavtokenizer/, trimmed to inference-only code).

Data synthesis: Kokoro, used to generate ~100 hours of synthetic audio for training this TTS model.

Text phonemization: espeak-ng, used to convert input text into phonemes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support