YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Igbo TTS — Isolated
An Igbo Text-to-Speech (TTS) model powered by YarnGPT2 and WavTokenizer.
The model generates natural-sounding Igbo speech directly from text and is packaged as a self-contained Hugging Face Transformers model with its required inference components included in the repository.
Model
- Hugging Face: nexusbert/igbo-tts-isolated
- Language: Igbo
- Default speaker:
igbo_male2 - Audio sampling rate: 24,000 Hz
Architecture
Igbo Text
│
▼
YarnGPT2
│
▼
Discrete Audio Codes
│
▼
WavTokenizer
│
▼
24 kHz Audio Waveform
The model combines YarnGPT2 for text-to-audio-token generation with WavTokenizer for decoding the generated audio tokens into a waveform.
Installation
Install the required packages:
pip install transformers torch huggingface_hub outetts uroman
Usage
Load the model directly from Hugging Face:
from transformers import AutoModelForTextToWaveform
model = AutoModelForTextToWaveform.from_pretrained(
"nexusbert/igbo-tts-isolated",
trust_remote_code=True
)
Generate Igbo speech:
text = "Ọ bụ oge ha si Enugwu steeti eme njem aga Anambra."
audio = model.generate_speech(text)
The returned value is a PyTorch audio tensor.
print(audio.shape)
print(audio.dtype)
The audio sampling rate is:
sampling_rate = model.config.sample_rate
print(sampling_rate)
Expected sampling rate: 24000 Hz.
Save Audio
You can save the generated waveform as a WAV file using torchaudio:
import torchaudio
audio_to_save = audio.detach().cpu()
if audio_to_save.dim() == 1:
audio_to_save = audio_to_save.unsqueeze(0)
torchaudio.save(
"igbo_output.wav",
audio_to_save,
model.config.sample_rate
)
Play Audio in Jupyter or Google Colab
import IPython.display as ipd
audio_cpu = audio.detach().cpu()
if audio_cpu.dim() == 2:
audio_cpu = audio_cpu.squeeze(0)
ipd.display(
ipd.Audio(
audio_cpu.numpy(),
rate=model.config.sample_rate
)
)
Speaker
The default speaker is:
igbo_male2
The model uses the configured default speaker when no speaker is explicitly provided.
The default can be inspected with:
print(model.config.default_speaker)
Generation Parameters
The model is configured with the following default generation settings:
| Parameter | Default |
|---|---|
| Temperature | 0.1 |
| Repetition penalty | 1.1 |
| Maximum length | 1000 |
| Sampling rate | 24000 Hz |
| Language | igbo |
| Speaker | igbo_male2 |
These parameters can be overridden during generation:
audio = model.generate_speech(
text="Ndewo, kedu ka ị mere?",
temperature=0.1,
repetition_penalty=1.1,
max_length=1000
)
Batch Generation
Multiple texts can also be supplied:
texts = [
"Ndewo, kedu ka ị mere?",
"Ọ bụ ezigbo ụbọchị taa.",
]
audio = model.generate_speech(texts)
The returned tensor contains the generated waveforms.
Hardware
The model can run on CUDA-enabled GPUs.
For GPU inference:
model = model.cuda()
For example, an NVIDIA A100 provides substantial acceleration for inference.
The model automatically keeps the YarnGPT2 and WavTokenizer components synchronized on the same device during synthesis.
Model Components
The repository contains the components required for inference:
config.json
configuration_igbo_tts.py
modeling_igbo_tts.py
audiotokenizer.py
model.safetensors
yarn_model/
wavtokenizer_large_speech_320_24k.ckpt
wavtokenizer_mediumdata_frame75_3s_nq1_code4096_dim512_kmeans200_attn.yaml
default_speakers_local/
This allows the model to load from the Hugging Face repository without requiring a separate YarnGPT2 repository checkout.
Technical Details
YarnGPT2
YarnGPT2 is responsible for converting the input text and speaker information into discrete audio tokens.
WavTokenizer
WavTokenizer converts the generated discrete audio codes into the final waveform.
Inference Flow
Input Igbo Text
│
▼
Text + Language + Speaker
│
▼
YarnGPT2
│
▼
Generated Audio Codes
│
▼
WavTokenizer
│
▼
Waveform
│
▼
24 kHz Audio
Example
Complete example:
from transformers import AutoModelForTextToWaveform
import torchaudio
model = AutoModelForTextToWaveform.from_pretrained(
"nexusbert/igbo-tts-isolated",
trust_remote_code=True
)
text = "Ọ bụ oge ha si Enugwu steeti eme njem aga Anambra."
audio = model.generate_speech(text)
audio = audio.detach().cpu()
if audio.dim() == 1:
audio = audio.unsqueeze(0)
torchaudio.save(
"igbo_output.wav",
audio,
model.config.sample_rate
)
print("Saved igbo_output.wav")
print("Sampling rate:", model.config.sample_rate)
Requirements
Recommended environment:
- Python 3.10+
- PyTorch
- Transformers
- Hugging Face Hub
- Outetts
- uroman
- CUDA-capable GPU for faster inference
Limitations
The quality of generated speech depends on:
- Text normalization
- Input sentence length
- Training data quality
- Speaker characteristics
- YarnGPT2 generation behavior
Very long inputs may require splitting the text into shorter sentences.
Citation
If you use this model in a research project, application, or publication, please reference the model repository:
nexusbert/igbo-tts-isolated
License
Please refer to the repository and the licenses of the underlying YarnGPT2 and WavTokenizer components before using the model for commercial or redistributed applications.
Acknowledgements
This model builds on the work of the developers of:
- YarnGPT2
- WavTokenizer
- Hugging Face Transformers
- Outetts
Special thanks to the open-source community for making these speech synthesis technologies available.
- Downloads last month
- 86