Instructions to use jsbeaudry/Fatiman-TTS-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use jsbeaudry/Fatiman-TTS-v1 with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("jsbeaudry/Fatiman-TTS-v1") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Fatiman-TTS-v1
Try it in your browser: The Trio Serie Space runs Fatiman with the other two models of the series: the Fatiman tab reads your text with the six voices (from the ONNX export,
Fatiman-TTS-v1-onnx), and the Klara tab speaks Makandal's answers. More on the Fatiman page of thetrio.space.
A Haitian Creole (Kreyòl) text-to-speech model: Kyutai's Pocket TTS (French 24-layer model) fine-tuned on Kreyòl speech. It runs on a CPU and streams: the first audio arrives in about 0.12 s on a laptop. It is the voice of Klara, a Kreyòl voice assistant that runs offline.
Results
80 held-out Kreyòl sentences (none in training), read aloud and transcribed back by two speech recognizers; character error rate, lower is better.
| Voice | CER (whisper / Qwen3-ASR Kreyòl) | First audio |
|---|---|---|
Fatiman-TTS-v1, voice man3 |
2.3 / 2.0 | ~0.12 s, streaming |
| Fatiman-TTS-v1, six voices pooled | 2.9 / 2.4 | ~0.12 s, streaming |
| Kokoro (Kreyòl fine-tune) | 4.4 / 3.6 | after the whole phrase |
| Qwen3-TTS 0.6B | 9.1 / 6.5 | after the whole phrase |
| French Pocket TTS, no fine-tune | ~20 |
The French model already reads Kreyòl at about 20% CER (its errors are French spelling habits, silent final letters above all); the fine-tune keeps its tokenizer and text embedding and teaches it Kreyòl spelling.
Voices
Six Kreyòl voices come with the model, in voices/. Each was made from about 10 s of a paid speaker's recording,
collected with their signed consent to this use. Character error rate on 20 sample sentences (whisper / Qwen3-ASR):
| Voice | Speaker | CER |
|---|---|---|
female1 |
woman | 2.4 / 1.5 |
man3 |
man | 2.7 / 1.7 |
man1 |
man | 2.8 / 2.0 |
female3 |
woman | 3.1 / 2.2 |
man2 |
man | 3.6 / 2.1 |
female2 |
woman | 5.2 / 5.1 |
female2's score comes from one sentence cut short by an early stop that did not recur in five reruns.
How it was made
- Base:
kyutai/pocket-tts,languages/french_24l, its tokenizer kept. - Data: about 95 h of Kreyòl Bible readings (
jsbeaudry/bible-kreyol-aligned, clips with a round-trip ASR CER of 15% or less) and about 27 h of 24 kHz Kreyòl speech, human and synthetic, CER-checked. Sentences of the evaluation set were removed from training. - Text: numbers and symbols are spelled out before training (
17→disèt), as Klara does before speaking; do the same, or digits will be read poorly. - Training: 12k steps, lr 1e-4, Kyutai's fine-tuning code.
Use
pip install pocket-tts==3.3.0 soundfile huggingface_hub
import os, tempfile
import numpy as np, soundfile as sf
from huggingface_hub import snapshot_download
from pocket_tts import TTSModel
folder = snapshot_download("jsbeaudry/Fatiman-TTS-v1")
config = os.path.join(tempfile.gettempdir(), "fatiman.yaml")
with open(config, "w") as f: # the config ships with @MODEL_DIR@ to fill in
f.write(open(os.path.join(folder, "config.template.yaml")).read().replace("@MODEL_DIR@", folder))
model = TTSModel.load_model(config=config)
voice = model.get_state_for_audio_prompt(os.path.join(folder, "voices", "man3.safetensors"))
# or your own: model.get_state_for_audio_prompt("my_voice.wav"), ~10 s of clean Kreyòl speech
audio = np.concatenate([c.detach().cpu().numpy().reshape(-1)
for c in model.generate_audio_stream(voice, "Bonjou! Kijan ou ye jodi a?")])
sf.write("bonjou.wav", audio, model.sample_rate)
A new voice comes from a short reference clip: about 10 s of clean speech, starting and ending between words. 5 s clips did no better, and denoising an already clean clip made it worse.
Consent. This model can imitate a voice from a short recording. The six voices above are published with their speakers' signed consent; for anyone else's voice, get their permission first, and never use a voice to deceive (Kyutai's terms for Pocket TTS apply).
License
CC-BY-4.0, as the base model. Based on Pocket TTS by Kyutai.
- Downloads last month
- -
Model tree for jsbeaudry/Fatiman-TTS-v1
Base model
kyutai/pocket-tts