PakUrdu-Conversational-ASR — Whisper Large-v3 Turbo Urdu Call Fine-tune (LoRA)

LoRA adapter (r=16, α=32, attention projections) on top of openai/whisper-large-v3-turbo (immutable revision 41f01f3fe87f28c78e2fbf8b568835947dd65ed9), fine-tuned for telephone-domain conversational Urdu on real customer-call audio from SimpleRishta (سمپل رشتہ) — PII-masked, speaker-safe splits. No SimpleRishta audio or transcripts are published; only these adapter weights.

Live demo: PakUrdu-ASR-Demo — runs fully in your browser (WebGPU); your audio never leaves your device.

Intended / out-of-scope use

  • Intended: transcribing Urdu telephone-call audio (customer-call traffic, conversational telephony speech), decoded with language=ur, task=transcribe, greedy.
  • Out of scope: general-purpose studio/broadcast Urdu (our internal measurements place it behind general-purpose fine-tunes there), languages other than Urdu, speaker identification, or any decision-making about individuals.

Results

133 held-out SimpleRishta calls, zero caller overlap with training, corpus C0 WER (normalized, PII-masked), identical decode settings for both models:

Model WER on SR calls
openai/whisper-large-v3 (off-the-shelf) 15.6%
this adapter 8.9%

Reference transcription: an independent high-accuracy Whisper-family model (eval-only, never trained on). Both models scored identically against it. Head-to-head on the same calls: this adapter wins 121 / loses 9 / ties 3. Zero degenerate generations across all stability checks (repetition-rate and length-ratio guards, clean greedy decoding). Aggregate protocol and numbers: eval/aggregate-results.json.

Independent reproduction (live run): the 8.9% figure is re-derived end-to-end on a clean GPU in this public Kaggle notebook — model loaded, 133 calls decoded, fresh WER printed against the certified value (observed drift 0.0000): r12-sr-eval-wer-proof.

Usage

import torch, soundfile as sf
from peft import PeftModel
from transformers import WhisperForConditionalGeneration, WhisperProcessor

base = "openai/whisper-large-v3-turbo"
processor = WhisperProcessor.from_pretrained(base)
model = WhisperForConditionalGeneration.from_pretrained(base)
model = PeftModel.from_pretrained(model, "KhiredNetworks/PakUrdu-Conversational-ASR").merge_and_unload()

speech, sr = sf.read("call.wav", dtype="float32")
inputs = processor(speech, sampling_rate=sr, return_tensors="pt")
ids = model.generate(inputs.input_features, language="ur",
                     task="transcribe", do_sample=False)
print(processor.batch_decode(ids, skip_special_tokens=True)[0])

A runnable version ships as inference_example.py.

Training data & governance

  • Fine-tuned on real SimpleRishta customer-call audio (16 kHz telephone Urdu; transcripts PII-masked before training; brand name preserved; caller-group-safe train/dev split — no caller appears in both).
  • Evaluation set frozen and never used in training.
  • Training method: LoRA (PEFT), single epoch, fresh from the pinned base revision — no continuation from any other checkpoint.
  • No SimpleRishta audio, transcripts, or caller-identifying data are published. Published results are aggregate-only.
  • Full training/eval audit trail retained internally (sealed compute records; certified numbers independently reproduced end-to-end on a clean GPU run, drift 0.0000).

Transcription convention & inference settings

Urdu script output; evaluation text normalized under a corpus C0 protocol (Unicode normalization, digit/punctuation folding, brand canonicalization, PII masking) before scoring. Decode: greedy (do_sample=False), language=ur, task=transcribe, 16 kHz mono input. FP16 on GPU / FP32 on CPU both supported; results certified on a single T4.

Limitations

  • Scores are vs a model reference, not human transcription.
  • Evaluated on accepted (clean) call audio; performance on heavily noisy or overlapped speech is not characterized here.
  • Urdu only (language=ur fixed in the intended decode config).

License, attribution & contact

  • Code/adapter: Apache-2.0 (see LICENSE). Base model: openai/whisper-large-v3-turbo — Whisper license and attribution chain apply downstream.
  • Citation: see CITATION.cff (Khired Networks).
  • Issues: use the repo's Issues tab.
Downloads last month
23
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KhiredNetworks/PakUrdu-Conversational-ASR

Adapter
(156)
this model

Space using KhiredNetworks/PakUrdu-Conversational-ASR 1