Instructions to use KhiredNetworks/PakUrdu-Conversational-ASR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use KhiredNetworks/PakUrdu-Conversational-ASR with PEFT:
from peft import PeftModel from transformers import AutoModelForSpeechSeq2Seq base_model = AutoModelForSpeechSeq2Seq.from_pretrained("openai/whisper-large-v3-turbo") model = PeftModel.from_pretrained(base_model, "KhiredNetworks/PakUrdu-Conversational-ASR") - Notebooks
- Google Colab
- Kaggle
PakUrdu-Conversational-ASR — Whisper Large-v3 Turbo Urdu Call Fine-tune (LoRA)
LoRA adapter (r=16, α=32, attention projections) on top of
openai/whisper-large-v3-turbo (immutable revision
41f01f3fe87f28c78e2fbf8b568835947dd65ed9), fine-tuned for
telephone-domain conversational Urdu on real customer-call audio from
SimpleRishta (سمپل رشتہ) — PII-masked, speaker-safe splits. No
SimpleRishta audio or transcripts are published; only these adapter weights.
Live demo: PakUrdu-ASR-Demo — runs fully in your browser (WebGPU); your audio never leaves your device.
Intended / out-of-scope use
- Intended: transcribing Urdu telephone-call audio (customer-call
traffic, conversational telephony speech), decoded with
language=ur,task=transcribe, greedy. - Out of scope: general-purpose studio/broadcast Urdu (our internal measurements place it behind general-purpose fine-tunes there), languages other than Urdu, speaker identification, or any decision-making about individuals.
Results
133 held-out SimpleRishta calls, zero caller overlap with training, corpus C0 WER (normalized, PII-masked), identical decode settings for both models:
| Model | WER on SR calls |
|---|---|
openai/whisper-large-v3 (off-the-shelf) |
15.6% |
| this adapter | 8.9% |
Reference transcription: an independent high-accuracy Whisper-family model
(eval-only, never trained on). Both models scored identically against it.
Head-to-head on the same calls: this adapter wins 121 / loses 9 / ties 3.
Zero degenerate generations across all stability checks (repetition-rate and
length-ratio guards, clean greedy decoding). Aggregate protocol and numbers:
eval/aggregate-results.json.
Independent reproduction (live run): the 8.9% figure is re-derived end-to-end on a clean GPU in this public Kaggle notebook — model loaded, 133 calls decoded, fresh WER printed against the certified value (observed drift 0.0000): r12-sr-eval-wer-proof.
Usage
import torch, soundfile as sf
from peft import PeftModel
from transformers import WhisperForConditionalGeneration, WhisperProcessor
base = "openai/whisper-large-v3-turbo"
processor = WhisperProcessor.from_pretrained(base)
model = WhisperForConditionalGeneration.from_pretrained(base)
model = PeftModel.from_pretrained(model, "KhiredNetworks/PakUrdu-Conversational-ASR").merge_and_unload()
speech, sr = sf.read("call.wav", dtype="float32")
inputs = processor(speech, sampling_rate=sr, return_tensors="pt")
ids = model.generate(inputs.input_features, language="ur",
task="transcribe", do_sample=False)
print(processor.batch_decode(ids, skip_special_tokens=True)[0])
A runnable version ships as inference_example.py.
Training data & governance
- Fine-tuned on real SimpleRishta customer-call audio (16 kHz telephone Urdu; transcripts PII-masked before training; brand name preserved; caller-group-safe train/dev split — no caller appears in both).
- Evaluation set frozen and never used in training.
- Training method: LoRA (PEFT), single epoch, fresh from the pinned base revision — no continuation from any other checkpoint.
- No SimpleRishta audio, transcripts, or caller-identifying data are published. Published results are aggregate-only.
- Full training/eval audit trail retained internally (sealed compute records; certified numbers independently reproduced end-to-end on a clean GPU run, drift 0.0000).
Transcription convention & inference settings
Urdu script output; evaluation text normalized under a corpus C0 protocol
(Unicode normalization, digit/punctuation folding, brand canonicalization,
PII masking) before scoring. Decode: greedy (do_sample=False),
language=ur, task=transcribe, 16 kHz mono input. FP16 on GPU / FP32 on
CPU both supported; results certified on a single T4.
Limitations
- Scores are vs a model reference, not human transcription.
- Evaluated on accepted (clean) call audio; performance on heavily noisy or overlapped speech is not characterized here.
- Urdu only (
language=urfixed in the intended decode config).
License, attribution & contact
- Code/adapter: Apache-2.0 (see
LICENSE). Base model:openai/whisper-large-v3-turbo— Whisper license and attribution chain apply downstream. - Citation: see
CITATION.cff(Khired Networks). - Issues: use the repo's Issues tab.
- Downloads last month
- 23