🎙️ SUPERnova TeraiTTS Voice Model
SUPERnova TeraiTTS is a customized conversational voice cloning setup designed to synthesize highly natural, casual, and conversational speech in both English and Nepali.
It leverages the powerful foundation of Coqui XTTS-v2 alongside custom single-speaker vocal signatures optimized for realistic pitch variation and natural breathing rhythms.
📂 Repository Contents
ref_en.wav: Standardized high-quality reference audio representing the English vocal signature.ref_ne.wav: Standardized high-quality reference audio representing the Nepali vocal signature.
🚀 How to Use Directly from Hugging Face
You can run voice synthesis directly in Python on any machine (such as a local computer, server, or Google Colab) by importing the model's reference voices straight from this repository:
import os
import torch
from TTS.api import TTS
from huggingface_hub import hf_hub_download
# 1. Download reference audios directly from your Hugging Face Repo
repo_id = "Supernova11c/Supernova-TeraiTTS"
ref_en_path = hf_hub_download(repo_id=repo_id, filename="ref_en.wav")
ref_ne_path = hf_hub_download(repo_id=repo_id, filename="ref_ne.wav")
# 2. Load the XTTS-v2 Engine
device = "cuda" if torch.cuda.is_available() else "cpu"
tts = TTS(model_name="tts_models/multilingual/multi-dataset/xtts_v2").to(device)
# 3. Generate Speech with Natural Human Rhythms
# English
tts.tts_to_file(
text="This is a natural conversational voice generated directly from Hugging Face!",
speaker_wav=ref_en_path,
language="en",
file_path="output_en.wav",
temperature=0.85,
speed=0.95,
repetition_penalty=2.5,
top_k=50,
top_p=0.85
)
# Nepali (using Hindi phonetic engine mapping for clean output)
tts.tts_to_file(
text="नयाँ प्रविधिले हाम्रो जीवनलाई झन् सहज र सरल बनाउँछ।",
speaker_wav=ref_ne_path,
language="hi",
file_path="output_ne.wav",
temperature=0.85,
speed=0.95,
repetition_penalty=2.5,
top_k=50,
top_p=0.85
)
print("🎉 Audio generation completed successfully!")
Model tree for Supernova11c/Supernova-TeraiTTS
Base model
coqui/XTTS-v2