Ghalib
Ghalib generates mono 48 kHz speech in the Ghalib voice. This private release bundles the Ghalib model weights, tokenizer, vocoder, speaker encoder, style references, inference runtime, and pronunciation processor. Firm is the default tone.
The pronunciation dictionary is part of the model package and is applied automatically by GhalibTTS and the HTTP server. There is no separate caller-side script to run. This is a packaged text processor, not a change to the neural weights. Existing callers using the low-level GhalibRuntime should switch to GhalibTTS to enable automatic processing and references.
Installation
Tested with Linux, Python 3.12, NVIDIA L40S, PyTorch 2.11.0 CUDA 13.0, and BF16. Use a CUDA-compatible driver. Install the package directly from an authenticated snapshot; it is not a package published on PyPI.
python3.12 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip huggingface_hub
hf auth login
pip install torch==2.11.0 torchaudio==2.11.0 --index-url https://download.pytorch.org/whl/cu130
python -c 'from huggingface_hub import snapshot_download; snapshot_download("theusamaaslam/ghalib", local_dir="ghalib-model")'
pip install ./ghalib-model
Access to this private repository is required. Never embed tokens in code. Download a fixed commit with revision="<commit>" for reproducible deployments. Pinned top-level runtime dependencies are in requirements.txt; file_inventory.json records model and code checksums.
Python
from ghalib import GhalibTTS
tts = GhalibTTS.from_pretrained("./ghalib-model")
tts.save_wav("آپ فکر نہ کریں، میں آپ کا internet check کر لیتا ہوں۔", "speech.wav")
Or load directly from the private Hub after installation:
tts = GhalibTTS.from_pretrained("theusamaaslam/ghalib")
result = tts.generate("میں آپ کی درخواست دیکھ رہا ہوں۔", style="firm")
# result["audio"] is a tensor; result["sample_rate"] is 48000.
Available styles: firm (default), neutral, warm, empathetic. Bundled references are resolved relative to the model directory and are cached by the runtime. You do not need to upload reference audio.
Native Audio Streaming
for audio_chunk in tts.generate_stream("میں آپ کا account check کر لیتا ہوں۔"):
pcm = (audio_chunk.detach().float().cpu().reshape(-1).clamp(-1, 1)
.numpy() * 32767).astype("<i2").tobytes()
# Send pcm to the audio output or LiveKit audio source immediately.
The complete input text is supplied once and audio is emitted incrementally from one acoustic session. This is not token-in double streaming, and text is not split into separately synthesized phrases. Defaults preserve BF16, optimized inference, Euler sampling, 10 solver steps, guidance 1.2, speaker scale 1.5, maximum 500 audio patches, and vocoder merge steps 4. First initialization includes compilation/warmup.
API
ghalib-serve --model ./ghalib-model --host 0.0.0.0 --port 8003 --workers 1
curl http://127.0.0.1:8003/health
curl -X POST http://127.0.0.1:8003/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"text":"آپ فکر نہ کریں، میں آپ کا internet check کر لیتا ہوں۔","stream":false,"response_format":"wav"}' \
--output speech.wav
Streaming example:
curl --no-buffer http://127.0.0.1:8003/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"text":"میں آپ کی درخواست دیکھ رہا ہوں۔","voice":"ghalib","style":"firm","stream":true,"native_stream":true,"response_format":"pcm"}' \
--output speech.pcm
PCM is headerless mono signed 16-bit little-endian, 48 kHz, not a WAV file. Streaming WAV requests are rejected; use stream:false for playable WAV. The legacy native_stream field remains accepted; audio streaming always uses the native path. For LiveKit, reframe PCM into 20 ms frames (960 samples, 1,920 bytes), carry incomplete frames between chunks, and resample only when the receiving audio source requires it.
Parallel Requests
--workers 2 starts two isolated model replicas, not two HTTP servers. Each replica owns its acoustic state. Requests queue with bounded admission; overload returns HTTP 503. A disconnected stream is drained internally before that worker is reused. More workers consume more VRAM and are not guaranteed to improve throughput on a saturated GPU. Use one HTTP process; do not combine this option with multiple Uvicorn workers.
CUDA memory depends on sequence length and concurrent workers. Measure peak use on the target GPU and leave headroom for other workloads. Low-precision quantization and reduced solver steps are not enabled. Warmed first-audio latency is not an end-to-end call latency guarantee; queued requests include waiting time.
Pronunciation Dictionary
pronunciation_lexicon.json preserves ordinary Urdu in Urdu script. Only explicitly listed English/technical loanwords become Latin, for example بینڈ وڈتھ to bandwidth and راؤٹر to router. The two approved ordinary-Urdu exceptions are شکریہ to shukria and امید to umeed. force_urdu is empty, and the frontend does not transliterate Latin input into Urdu. Number and punctuation normalization remain active. Whole-word boundaries and longest-phrase matching avoid replacing fragments inside unrelated words. Words containing ڑ are protected from Urdu-to-Roman conversion. URLs, email addresses, and path-like identifiers are excluded from rewriting.
print(tts.processor("main aap کا شکریہ ادا کرتا ہوں۔"))
Edit the bundled JSON atomically to add or adjust mappings; the processor notices file changes. Invalid or conflicting entries fail validation rather than being silently ignored. Listening is still required: dictionary coverage does not prove that every spelling sounds correct, and context-dependent homographs cannot always be fixed with one word-level mapping. Reference transcripts are never rewritten.
Limits And Safety
This release does not retrain weights, alter the tokenizer, or claim new pronunciation evaluation scores. Dictionary additions have structural tests but still need listening validation. Rare words, contextual homographs, code-switching, and pauses may remain imperfect. Synthesis can vary across requests.
For OOM, reduce worker count and prompt length; do not launch overlapping servers. For missing references, download the complete snapshot including styles/. For silence, verify the PCM format and inspect RMS, not just file size. For first-request delays, wait for startup warmup and /health readiness. Expose the service only behind authentication and network access controls; the example API has no built-in user authentication.
Use only with consent and disclose synthesized speech. Weights and references remain governed by LICENSE; redistribution is not granted by the code license. Portions of the runtime are licensed under Apache 2.0; see LICENSE-APACHE-2.0. No training datasets or credentials are bundled.
- Downloads last month
- 8