AhmedEladl/saudi-dialect-speech-female
Viewer โข Updated โข 4.96k โข 144 โข 1
How to use AhmedEladl/Magpie-TTS-Saudi-Arabic with NeMo:
# tag did not correspond to a valid NeMo domain.
This model is a fine-tuned version of the NVIDIA Magpie TTS Multilingual (357M) model, specifically adapted for Saudi Arabic text-to-speech generation.
Below are raw, unedited audio samples demonstrating the dialect adaptation.
| Text Prompt | Base Modelmagpie_tts_multilingual_357m |
Fine-Tuned Model |
|---|---|---|
| ุชุฑู ุงูุทุฑูู ุฌูุฉ ุงูุฑูุงุถ ุงูููู ู ุฑุฉ ุฒุญู ุฉุ ูุฅุฐุง ู ุณุชุนุฌููู ุฃุญุณู ูุทูุน ู ู ุงูุญูู ุนุดุงู ููุตู ูุจู ุงูู ูุนุฏ ูู ุง ูุชุฃุฎุฑ. | ||
| ุญูุงูู ุงููู ูู ุฃู ููุชุ ูุฅุฐุง ูุตูุชูุง ุนูู ูููุ ูุจุฌูุฒ ููู ุงููููุฉ ูุงูุญูุง ููุฌูุณ ูููุง ู ุน ุจุนุถ ููุณููู ููู ููุช ู ุชุฃุฎุฑ. | ||
| ู ู ูุซุฑ ู ุง ูุชุบูุฑ ุงูุฌู ูุงูุฃูุงู ุ ุงููุงุญุฏ ู ุง ูุฏุฑู ูุด ููุจุณุ ุงูุตุจุญ ูููู ุงูุฌู ุจุงุฑุฏ ุดูู ูุงูุนุตุฑ ูุฑุฌุน ุงูุญุฑ ู ุฑุฉ ุซุงููุฉ. | ||
| ุฃูุง ุฃุดูู ูุฎูู ุงูู ุดูุงุฑ ุจุนุฏ ุงูู ุบุฑุจุ ูุฃู ุงูุฌู ุจูููู ุฃุญุณู ูุงููุงุณ ุชุฎู ู ู ุงูุดูุงุฑุนุ ูููุฏุฑ ูุงุฎุฐ ุฑุงุญุชูุง ูู ุงูุทูุนุฉ. |
nvidia/magpie_tts_multilingual_357mar-SAnvidia/nemo-nano-codec-22khz-1.89kbps-21.5fpsThe model was fine-tuned using the AhmedEladl/saudi-voice-dataset dataset.
To run inference with this model, ensure you have the correct versions of the NeMo toolkit and audio handling libraries installed.
pip install torch>=2.1.0
pip install soundfile>=0.12.1
pip install librosa>=0.10.1
pip install huggingface_hub>=0.23.0
pip install nemo_toolkit[all]==2.8.0rc0
import soundfile as sf
from huggingface_hub import hf_hub_download
from nemo.collections.tts.modules.magpietts_inference.utils import ModelLoadConfig, load_magpie_model
# 1. Download Model & Codec from Hugging Face Hub
print("Downloading models...")
model_path = hf_hub_download(repo_id="AhmedEladl/Magpie-TTS-Saudi-Arabic", filename="Magpie-TTS-Saudi-Female.nemo")
codec_path = hf_hub_download(repo_id="nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps", filename="nemo-nano-codec-22khz-1.89kbps-21.5fps.nemo")
# 2. Load the fine-tuned model and codec
config = ModelLoadConfig(nemo_file=model_path, codecmodel_path=codec_path)
model, _ = load_magpie_model(config)
model.eval().cuda()
# 3. Generate Audio
prompt = "ุงูุฐูุงุก ุงูุงุตุทูุงุนู ุตุงุฑ ุฌุฒุก ุฃุณุงุณู ู
ู ุญูุงุชูุง ุงูููู
ูุฉุ ูุชุทููุฑ ูู
ุงุฐุฌ ุชุฏุนู
ููุฌุชูุง ุฎุทูุฉ ุฌุฏุงู ู
ูู
ุฉ."
print("Generating audio...")
res = model.do_tts(transcript=prompt, language="ar-SA", apply_TN=False)
# 4. Save the output
audio = res[0].cpu().numpy()
if len(audio.shape) == 2:
audio = audio[0]
sf.write("output.wav", audio, 22050)
print("โ
Audio saved successfully to output.wav")
If you use this model in your research or projects, please cite it as follows:
@misc{magpie_tts_sa-ar_arabic,
author = {Ahmed Eladl},
title = {Magpie TTS - Saudi Arabic (Fine-Tuned)},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/AhmedEladl/Magpie-TTS-Saudi-Arabic}
}
For any questions, issues, or inquiries regarding this model, please reach out:
Base model
nvidia/magpie_tts_multilingual_357m