Fatiman-TTS-v1

Try it in your browser: The Trio Serie Space runs Fatiman with the other two models of the series: the Fatiman tab reads your text with the six voices (from the ONNX export, Fatiman-TTS-v1-onnx), and the Klara tab speaks Makandal's answers. More on the Fatiman page of thetrio.space.

A Haitian Creole (Kreyòl) text-to-speech model: Kyutai's Pocket TTS (French 24-layer model) fine-tuned on Kreyòl speech. It runs on a CPU and streams: the first audio arrives in about 0.12 s on a laptop. It is the voice of Klara, a Kreyòl voice assistant that runs offline.

Results

80 held-out Kreyòl sentences (none in training), read aloud and transcribed back by two speech recognizers; character error rate, lower is better.

Voice CER (whisper / Qwen3-ASR Kreyòl) First audio
Fatiman-TTS-v1, voice man3 2.3 / 2.0 ~0.12 s, streaming
Fatiman-TTS-v1, six voices pooled 2.9 / 2.4 ~0.12 s, streaming
Kokoro (Kreyòl fine-tune) 4.4 / 3.6 after the whole phrase
Qwen3-TTS 0.6B 9.1 / 6.5 after the whole phrase
French Pocket TTS, no fine-tune ~20

The French model already reads Kreyòl at about 20% CER (its errors are French spelling habits, silent final letters above all); the fine-tune keeps its tokenizer and text embedding and teaches it Kreyòl spelling.

Voices

Six Kreyòl voices come with the model, in voices/. Each was made from about 10 s of a paid speaker's recording, collected with their signed consent to this use. Character error rate on 20 sample sentences (whisper / Qwen3-ASR):

Voice Speaker CER
female1 woman 2.4 / 1.5
man3 man 2.7 / 1.7
man1 man 2.8 / 2.0
female3 woman 3.1 / 2.2
man2 man 3.6 / 2.1
female2 woman 5.2 / 5.1

female2's score comes from one sentence cut short by an early stop that did not recur in five reruns.

How it was made

  • Base: kyutai/pocket-tts, languages/french_24l, its tokenizer kept.
  • Data: about 95 h of Kreyòl Bible readings (jsbeaudry/bible-kreyol-aligned, clips with a round-trip ASR CER of 15% or less) and about 27 h of 24 kHz Kreyòl speech, human and synthetic, CER-checked. Sentences of the evaluation set were removed from training.
  • Text: numbers and symbols are spelled out before training (17 → disèt), as Klara does before speaking; do the same, or digits will be read poorly.
  • Training: 12k steps, lr 1e-4, Kyutai's fine-tuning code.

Use

pip install pocket-tts==3.3.0 soundfile huggingface_hub
import os, tempfile
import numpy as np, soundfile as sf
from huggingface_hub import snapshot_download
from pocket_tts import TTSModel

folder = snapshot_download("jsbeaudry/Fatiman-TTS-v1")
config = os.path.join(tempfile.gettempdir(), "fatiman.yaml")
with open(config, "w") as f:                       # the config ships with @MODEL_DIR@ to fill in
    f.write(open(os.path.join(folder, "config.template.yaml")).read().replace("@MODEL_DIR@", folder))
model = TTSModel.load_model(config=config)

voice = model.get_state_for_audio_prompt(os.path.join(folder, "voices", "man3.safetensors"))
# or your own: model.get_state_for_audio_prompt("my_voice.wav"), ~10 s of clean Kreyòl speech
audio = np.concatenate([c.detach().cpu().numpy().reshape(-1)
                        for c in model.generate_audio_stream(voice, "Bonjou! Kijan ou ye jodi a?")])
sf.write("bonjou.wav", audio, model.sample_rate)

A new voice comes from a short reference clip: about 10 s of clean speech, starting and ending between words. 5 s clips did no better, and denoising an already clean clip made it worse.

Consent. This model can imitate a voice from a short recording. The six voices above are published with their speakers' signed consent; for anyone else's voice, get their permission first, and never use a voice to deceive (Kyutai's terms for Pocket TTS apply).

License

CC-BY-4.0, as the base model. Based on Pocket TTS by Kyutai.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jsbeaudry/Fatiman-TTS-v1

Finetuned
(28)
this model

Space using jsbeaudry/Fatiman-TTS-v1 1