Instructions to use hyperneuronAILabs/quipus-0.6-speechv2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hyperneuronAILabs/quipus-0.6-speechv2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="hyperneuronAILabs/quipus-0.6-speechv2")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("hyperneuronAILabs/quipus-0.6-speechv2") model = AutoModelForCausalLM.from_pretrained("hyperneuronAILabs/quipus-0.6-speechv2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use hyperneuronAILabs/quipus-0.6-speechv2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hyperneuronAILabs/quipus-0.6-speechv2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hyperneuronAILabs/quipus-0.6-speechv2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for hyperneuronAILabs/quipus-0.6-speechv2 to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="hyperneuronAILabs/quipus-0.6-speechv2", max_seq_length=2048, )
HyperneuronAI Text-to-Speech Model
Multilingual Open-Source Text-to-Speech for Indian Languages
Model Details
Model Description
This is an open-source Text-to-Speech model developed by HyperneuronAI. The model is designed to generate natural speech from text and now supports all 22 officially recognized languages of India: Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, and Urdu.
What's new: this release expands the original 4-language model (Hindi, Assamese, Punjabi, Kannada) to full coverage of all 22 scheduled Indian languages, using a replay-based fine-tuning strategy so quality on the original 4 languages is preserved while adding 18 new ones. See Training Details below.
Quipus is 2x Faster than realtime and lower latency TTFB ~120ms.โ
โกQuipus Performance
| Metric | Value |
|---|---|
| First Audio | ~120 msโ |
| Generation Speed | 2ร Realtimeโ |
| Languages | 22 |
โ Inherited from the prior 4-language release โ architecture and generation loop are unchanged, so these should still hold, but they have not been re-benchmarked specifically against the 22-language checkpoint yet. Treat as provisional until re-verified.
Sneak Peak
In model Inference performance

๐บ๏ธ Roadmap
- ๐ Emotion-aware Speech
- ๐ Arabic & English support as part of worldwide contribution
- โก Streaming Backend Optimizations
- Training and Inference Code
The model uses a Qwen3 backbone and is intended for research, experimentation, and building voice AI applications. Users are free to fine-tune the model for custom voices and additional languages.
Voice cloning capabilities are not provided with this release to encourage responsible AI usage.
- Developed by: HyperneuronAI
- Funded by: HyperneuronAI
- Shared by: HyperneuronAI
- Model type: Text-to-Speech
- Backbone: Qwen3
- Languages: Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, Urdu
- License: MIT
Audio Samples
| Language | Sample |
|---|---|
| ๐ฎ๐ณ Kannada | |
| ๐ฎ๐ณ Hindi | |
| ๐ฎ๐ณ Punjabi | |
| ๐ฎ๐ณ Assamese | |
| ๐ฎ๐ณ Tamil | |
| ๐ฎ๐ณ Telugu | |
| ๐ฎ๐ณ Bengali | |
| ๐ฎ๐ณ Bodo | |
| ๐ฎ๐ณ Dogri | |
| ๐ฎ๐ณ Gujarati | |
| ๐ฎ๐ณ Kashmiri | |
| ๐ฎ๐ณ Konkani | |
| ๐ฎ๐ณ Maithili | |
| ๐ฎ๐ณ Malayalam | |
| ๐ฎ๐ณ Manipuri | |
| ๐ฎ๐ณ Marathi | |
| ๐ณ๐ต Nepali | |
| ๐ฎ๐ณ Odia | |
| ๐ฎ๐ณ Sanskrit | |
| ๐ฎ๐ณ Santali | |
| ๐ฎ๐ณ Sindhi | |
| ๐ฎ๐ณ Urdu |
Samples for the 18 newly added languages are not yet uploaded to this repo โ the table above only has the original 4. Upload new
.wavfiles toartifacts/and add a row per language (or generate them fresh with the code in "How to Get Started" below) before this table reflects full 22-language coverage.
๐๏ธ Available Voices
| Language | Speakers |
|---|---|
| ๐ฎ๐ณ Assamese | Dipankar(Male) โข Kavita(Female) |
| ๐ฎ๐ณ Bengali | Arindam(Male) โข Rupa(Female) |
| ๐ฎ๐ณ Bodo | Daimalu(Male) โข Mainao(Female) |
| ๐ฎ๐ณ Dogri | Balwant(Male) โข Shakti(Female) |
| ๐ฎ๐ณ Gujarati | Nirav(Male) โข Hetal(Female) |
| ๐ฎ๐ณ Hindi | Raman(Male) โข Anvita(Female) |
| ๐ฎ๐ณ Kannada | Manjunath(Male) โข Sindhura(Female) |
| ๐ฎ๐ณ Kashmiri | Farooq(Male) โข Habba(Female) |
| ๐ฎ๐ณ Konkani | Ramnath(Male) โข Shweta(Female) |
| ๐ฎ๐ณ Maithili | Vidyapati(Male) โข Mythili(Female) |
| ๐ฎ๐ณ Malayalam | Unnikrishnan(Male) โข Lakshmi(Female) |
| ๐ฎ๐ณ Manipuri | Ibomcha(Male) โข Chaoba(Female) |
| ๐ฎ๐ณ Marathi | Sachin(Male) โข Smita(Female) |
| ๐ฎ๐ณ Nepali | Prakash(Male) โข Sunita(Female) |
| ๐ฎ๐ณ Odia | Jagannath(Male) โข Sunanda(Female) |
| ๐ฎ๐ณ Punjabi | Amanjit(Male) โข Supreet(Female) |
| ๐ฎ๐ณ Sanskrit | Aryan(Male) โข Devavani(Female) |
| ๐ฎ๐ณ Santali | Sido(Male) โข Phulmani(Female) |
| ๐ฎ๐ณ Sindhi | Hari(Male) โข Kajal(Female) |
| ๐ฎ๐ณ Tamil | Murugan(Male) โข Kamala(Female) |
| ๐ฎ๐ณ Telugu | Ravi(Male) โข Padma(Female) |
| ๐ฎ๐ณ Urdu | Salman(Male) โข Zoya(Female) |
44 speaker voices total (2 per language ร 22 languages).
Model Sources
- Repository: Realtime optimized streaming inference code is planned for a future release.
- Demo: Coming soon.
Uses
Direct Use
This model can be used for Text-to-Speech generation in voice AI applications, including:
| โ |
|---|
| Conversational AI |
| AI Calling |
| Voice Assistants |
| Accessibility |
| Customer Support |
| IVR |
| Edge Devices |
| Research |
Fine-Tuning
Users may fine-tune this model for:
- New speakers
- Domain-specific speech styles
- Additional languages beyond the current 22
- Custom application-specific voices
Out-of-Scope Use
This model should not be used for unethical, harmful, deceptive, or illegal purposes, including but not limited to:
- Impersonation without consent
- Fraudulent voice generation
- Misinformation or manipulation
- Harassment or abuse
- Any use that violates applicable laws or platform policies
HyperneuronAI is not responsible for misuse of this model by third parties.
How to Get Started
vLLM code examples.
#Install libraries if not installed
# Download vlllm based on cuda version
#!pip install vllm==0.19.0
#!pip install -q soundfile
# ==========================================
# vLLM SERVING - Colab-safe audio saving
# ==========================================
"""
You can use vllm for inferencing as it provides optimised control over cuda profiling and inturn better response time.
"""
import time
import re
import torch
import numpy as np
import soundfile as sf
from IPython.display import Audio, display
from vllm import LLM, SamplingParams
from snac import SNAC
VLLM_MODEL_PATH = "hyperneuronAILabs/quipus-0.6-speechv1"
DEFAULT_SPEAKER = "Anvita" # any of the 44 voices in "Available Voices" above
FRAME_LAYER_PATTERN = [0, 1, 2, 2, 1, 2, 2]
SAMPLE_RATE = 24000
llm = LLM(
model=VLLM_MODEL_PATH,
dtype="bfloat16",
max_model_len=2048,
gpu_memory_utilization=0.20,
enforce_eager=False,
)
tok = llm.get_tokenizer()
AUDIO_END_ID = tok.convert_tokens_to_ids("<audio_end>")
snac_decoder = (
SNAC.from_pretrained("hubertsiuzdak/snac_24khz")
.eval()
.cuda()
)
def build_prompt_prefix(speaker, text):
return f"{speaker}: {text} <audio_start> "
def quipus_tts(
text,
speaker=DEFAULT_SPEAKER,
out="vllm_out.wav",
temperature=0.7,
top_p=0.9,
):
prompt = build_prompt_prefix(speaker, text)
sampling_params = SamplingParams(
temperature=temperature,
top_p=top_p,
repetition_penalty=1.1,
max_tokens=1024,
stop_token_ids=[AUDIO_END_ID],
)
t0 = time.time()
result = llm.generate([prompt], sampling_params)
gen_time = time.time() - t0
out_ids = list(result[0].outputs[0].token_ids)
toks = tok.convert_ids_to_tokens(out_ids)
parsed = []
for token in toks:
match = re.fullmatch(r"<snac_l(\d+)_c(\d+)>", token)
if match:
parsed.append((int(match.group(1)), int(match.group(2))))
l0, l1, l2 = [], [], []
i = 0
while i + 7 <= len(parsed):
window = parsed[i:i + 7]
if [layer for layer, _ in window] == FRAME_LAYER_PATTERN:
codes = [code for _, code in window]
l0.append(codes[0])
l1.extend([codes[1], codes[4]])
l2.extend([codes[2], codes[3], codes[5], codes[6]])
i += 7
else:
i += 1
if not l0:
print("No valid SNAC frames from vLLM output.")
return None
with torch.inference_mode():
wav = snac_decoder.decode(
[
torch.tensor([l0], dtype=torch.long, device="cuda"),
torch.tensor([l1], dtype=torch.long, device="cuda"),
torch.tensor([l2], dtype=torch.long, device="cuda"),
]
)
audio_sec = wav.shape[-1] / SAMPLE_RATE
pcm = wav.detach().squeeze().float().cpu().numpy()
pcm = np.clip(pcm, -1.0, 1.0)
sf.write(out, pcm, SAMPLE_RATE)
n = len(out_ids)
display(Audio(out, rate=SAMPLE_RATE))
return out
quipus_tts(
"เคฏเคน เคฎเฅเคกเคฒ เคธเฅเคชเฅเค-เคเฅ-เคเฅเคเฅเคธเฅเค เคฎเฅเคกเคฒ เคนเฅ , เคเคฟเคธเฅ เคจเคฟเคเคฟเคฒ เคจเฅ เคตเคฟเคเคธเคฟเคค เคเคฟเคฏเคพ เคนเฅเฅค เฆเฆ เฆฎเฆกเงเฆฒเฆเง เฆนเงเฆเง เฆธเงเฆชเฆฟเฆ เฆเง เฆเงเฆเงเฆธเฆ เฆฎเฆกเงเฆฒ, เฆจเฆฟเฆเฆฟเฆฒเง เฆฌเฆฟเฆเฆถเฆฟเฆค เฆเงฐเฆฟเฆเง",
speaker=DEFAULT_SPEAKER,
)
Training Details
Training Data
The model was fine-tuned on the AI4Bharat Rasa dataset, covering all 22 officially recognized languages of India:
Training Procedure
- Base checkpoint: this repo's prior release (4-language: Hindi, Assamese, Punjabi, Kannada).
- Method: full parameter fine-tuning โ all 608,362,496 parameters trained.
- Audio representation: SNAC neural codec at 24kHz, hierarchical 3-layer codebook, flattened into a 7-token-per-frame sequence (Orpheus-style interleave) so all three resolution layers stay time-aligned.
- Prompt format:
"{speaker}: {text} <audio_start> {snac_tokens} <audio_end>"โ loss is computed only on the audio-token portion; the text prompt is masked out. - Tokenizer: unchanged from the base checkpoint โ the SNAC control tokens were already present in its vocabulary, so no vocabulary resize was needed this round.
Training Hyperparameters
Training hyperparameters will be added in a future update.
Evaluation
Testing Data
Informal qualitative listening checks were run across all 22 languages using fixed test phrases and each language's paired speakers. A formal held-out evaluation set has not yet been built for this release.
Metrics
Formal evaluation metrics such as MOS, speaker similarity, intelligibility, word error rate, and latency benchmarks are not yet published for the 22-language checkpoint.
Results
Evaluation results will be added after broader testing and benchmarking.
Technical Specifications
Model Architecture and Objective
This is a Text-to-Speech model with a Qwen3 backbone. The model is optimized to generate speech from input text in supported Indian languages.
Further architecture details will be added in future documentation.
Compute Infrastructure
Compute details will be added in a future update.
Software
Software and inference dependencies will be added with the official inference code.
Limitations
- Output quality may vary depending on language, text normalization, punctuation, and input style.
- The model may struggle with code-mixed text, rare words, abbreviations, numerals, and domain-specific terminology.
- Voice cloning is not included in this release.
- Realtime streaming inference code is planned but not included yet.
Ethical Considerations
This model is released to support open-source development of Indian language voice AI. Users should ensure responsible deployment, obtain consent where required, and avoid deceptive or harmful applications.
Citation
Citation details will be added in a future release.
Authors
| Name |
|---|
| Nikhil Yadav |
| Ramanjit Singh |
| Pradeep Yadav |
| HyperneuronAI Research |
๐ค Contact
For questions, collaborations, or contributions, please contact HyperneuronAI
- Downloads last month
- 3
Model tree for hyperneuronAILabs/quipus-0.6-speechv2
Base model
hyperneuronAILabs/quipus-0.6-speechv1