Twi / Ghanaian English TTS (IPA-driven)

A Piper VITS voice for Twi that also reads the English words Ghanaian text is full of. It is driven by IPA phonemes rather than orthography, which is what lets one model serve both languages from a shared inventory.

Warm-started from en_US/libritts_r/medium and finetuned on 327 h (164 h Twi + 164 h Ghanaian English), 22.05 kHz, 1,555 speaker embeddings.

Use it

pip install stable-twi-tts && pip install "stable-twi-tts[twi]" && apt install espeak-ng
stable-twi-tts --model . --voice twi-6 --text "Akwaaba, wo ho te sen?" --out hello.wav
stable-twi-tts --model . --language mixed --voice twi-1 \
    --text "Mepe se mesua [computer science] wo [University of Ghana]." --out mixed.wav

Code: https://github.com/GhanaNLP/stable-twi-tts

Voices are ranked by measurement, not by training hours

12 voices are exposed of 1,555 speaker embeddings. Each synthesised the same held-out text, which was re-recognised and scored for phoneme error -- Twi with ghana-speech-phoneme-asr, the English spans with KoelLabs/xlsr-english-01.

pick for best
tiers.codeswitch text mixing English into Twi twi-1 (59.8%)
tiers.twi_only pure Twi twi-6 (26.8%; floor 25.9%)

The two rankings disagree sharply. twi-1 is the best code-switch voice yet 21st of 30 on pure Twi. Two of the three best Twi voices have under 3.3 h of audio each and would never have survived an hours-based shortlist. voices.json records every measurement.

Quality, honestly

Round-trip phoneme error against a real-audio floor, on held-out text:

floor model gap
Twi 25.9% 33.5% +7.6
English 32.2% 59.5% +27.3

Twi is good; English is markedly weaker. English had a third of the utterances (43k vs 151k) at 3.5x the length, and VITS learns alignment per utterance, so English got far less practice on harder material. English audio was also 16 kHz upsampled to 22.05 kHz, so English voices are band-limited to 8 kHz and sound duller. Twi was 24 kHz and is unaffected.

Training

Warm-started from en_US/libritts_r/medium and finetuned for 10 epochs / 52,039 steps, about 21.5 hours on a single H200. Batch 32, bf16-mixed, single device, 178,833 training utterances across 1,555 speakers. The shipped weights are epoch 7.

Per-validation metrics

Validation ran twice per epoch. val_mos is UTMOS, a learned naturalness predictor β€” higher is better. Everything else is a loss β€” lower is better.

step epoch val_loss val_mel val_dur val_kl val_mos
1,999 0 38.08 0.4792 1.6759 2.9333 2.563
3,999 0 39.06 0.4859 1.6750 2.9237 2.647
7,559 1 38.62 0.4851 1.6687 2.9323 2.658
9,559 1 40.12 0.4813 1.6738 2.8489 2.697
13,119 2 38.60 0.4738 1.6664 2.9387 2.746
15,119 2 38.05 0.4743 1.6708 2.9791 2.993
18,679 3 42.69 0.4821 1.6709 2.9007 2.695
20,679 3 37.51 0.4734 1.6643 2.9741 2.853
24,239 4 38.05 0.4774 1.6642 2.9261 2.809
26,239 4 38.12 0.4734 1.6645 2.9265 2.873
29,799 5 38.26 0.4815 1.6677 2.8775 2.977
31,799 5 38.50 0.4727 1.6695 2.9064 2.772
35,359 6 38.34 0.4743 1.6652 2.9190 2.795
37,359 6 38.85 0.4688 1.6620 2.9368 2.945
40,919 7 38.17 0.4741 1.6644 2.9444 3.018
42,919 7 38.04 0.4795 1.6656 2.9047 2.896
46,479 8 38.22 0.4686 1.6646 2.9374 2.761
48,479 8 37.96 0.4776 1.6653 2.8999 2.749
52,039 9 38.04 0.4732 1.6628 2.8492 2.879

Training losses

start best end
loss_g (generator, total) 41.05 33.50 37.86
loss_d (discriminator) 2.00 1.24 2.15
train_mel 0.5211 0.4047 0.4741

Reading these numbers

val_loss is not a quality signal here, and treating it as one would have been a mistake. It starts at 38.08 and ends at 38.04 β€” flat across the entire run, with a 42.69 spike at epoch 3 that means nothing. That is expected for a GAN: the generator's total loss is measured against a discriminator that is improving alongside it, so the number tracks the balance between the two, not how good the audio is. A run that looked converged at step 2,000 by val_loss in fact kept improving for another 40,000 steps.

The two signals that did carry information disagreed with each other:

  • val_mel (mel-spectrogram L1) bottomed at 0.4686 at epoch 8, having drifted down from 0.4792. A 2.2% total improvement β€” real, but small.
  • val_mos peaked at 3.018 at epoch 7 and never recovered: 2.90, 2.76, 2.75, 2.88 after. It is also noisy β€” epoch 3 swings 2.695 β†’ 2.853 within one epoch β€” so no single validation is trustworthy on its own.

Epoch 7 was shipped on val_mos plus listening, not on val_mel. Epoch 8's marginally better mel did not sound better. Where a spectral distance and a naturalness predictor disagree by this little, the honest reading is that the model plateaued somewhere around epoch 6–7 and the rest is noise.

The decisive check was neither: a round-trip evaluation β€” synthesise held-out text, re-recognise it with a phoneme ASR, score the phoneme error against the error the ASR makes on real audio of the same text. That is the "Quality, honestly" table above, and it is the number to trust, because it measures whether the audio says the right thing rather than whether it resembles a reference spectrogram.

checkpoint Twi UER English UER
epoch 0 32.1% 64.9%
epoch 1 31.3% 63.7%
epoch 7 33.5% 59.5%
real-audio floor 25.9% 32.2%

Twi was essentially converged by epoch 1; the remaining 8 epochs bought English (βˆ’4.2 points), which is the harder half. The Twi number moving up 2.2 points while listening quality clearly improved is a reminder that a 2-point UER difference is within this instrument's noise.

Training was stopped mid-epoch 9, at step 3,459 of 5,560. Nothing was still improving: val_mos had been below its epoch-7 peak for four consecutive validations, and val_mel's best was a checkpoint that sounded no better. Further epochs would have risked overfitting 1,555 pseudo-speakers without a metric able to tell us it was happening.

Finetuning

finetune/ has the Lightning checkpoint, the phoneme id map, the training command and dataset stats. The phoneme id map is required -- the weights encode "id 26 means /n/", so loading them against a different map yields nonsense. See finetune/FINETUNING.md, which also documents two traps: resize the speaker table before loading, and clear the phoneme cache or stale targets are silently reused.

Limits

  • Speakers are derived, not real identities -- pseudo-speakers from clustering unlabelled broadcast audio. One person may appear as two voices, and no voice is a consented identity.
  • Trained on broadcast speech, so it suits news-register text better than conversation.
  • Training targets came from an ASR, not human annotation, at roughly 17% phoneme error.
  • cc-by-nc-4.0: non-commercial, inherited from the Twi source audio.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support