Automatic Speech Recognition
NeMo
PyTorch
Persian
speech
audio
conformer
fastconformer
rnnt
Eval Results (legacy)

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

🎙️ SedaNevis_v1.0 (صدانویس)

Robust Multi-Domain Persian ASR Model with FastConformer-RNNT

SedaNevis_v1.0 is an enterprise-grade, state-of-the-art Automatic Speech Recognition (ASR) model specifically tailored for the Persian (Farsi) language. Built on top of NVIDIA's robust FastConformer-RNNT architecture, it is fine-tuned across diverse and challenging acoustic conditions, including elderly speech, conversational podcasts, studio reading, and emotional speech.

توضیح فارسی (Overview):
مدل SedaNevis_v1.0 (صدانویس) یک مدل پیشرفته و بهینه‌سازی‌شده برای بازشناسی خودکار گفتار زبان فارسی (ASR) است. این مدل با هدف عملکرد پایدار در شرایط چالش‌برانگیز آکوستیک (لهجه‌ها، گفتار سالمندان، مکالمات محاوره‌ای پادکست‌ها، و صدای همراه با احساسات) بر پایه معماری FastConformer-RNNT توسعه یافته است.


🏆 Benchmark & Latency Comparison

Evaluation Setup: Evaluated on 1,000 diverse Persian audio files under identical hardware conditions (NVIDIA GPU).

Rank Model Name Architecture Size Time (1k files) Speed & Throughput ⚡ CER (%) ↓ WER (%) ↓ Accuracy (%) ↑
🥇 1st SedaNevis_v1.0 (Ours) FastConformer-RNNT ~500 MB 62s (1 min) 16.13 audio/s (37.5× 🔥) 3.01% 9.78% 90.22%
🥈 2nd whisper-base-persian Seq2Seq Transformer ~290 MB 291s (4.8 min) 3.43 audio/s (8.0×) 6.91% 18.52% 81.48%
🥉 3rd Whisper-Base-Persian-Full Seq2Seq Transformer ~290 MB 290s (4.8 min) 3.44 audio/s (8.0×) 7.48% 19.69% 80.31%
4th Whisper-Large-v3-Turbo Seq2Seq Transformer ~1.6 GB 595s (10 min) 1.68 audio/s (3.9×) 9.22% 27.25% 72.75%
5th Whisper-Large-Persian-v4 Seq2Seq Transformer ~5.1 GB 2,325s (39 min) 0.43 audio/s (1.0× Baseline) 8.16% 20.97% 79.03%

⚡ Operational Advantage:
While achieving an unprecedented low error rate (9.78% WER), SedaNevis operates 5× to 38× faster than typical Whisper Large models, making it optimal for both real-time streaming pipelines and large-scale offline batch processing.


⚡ Unmatched Speed & Real-World Latency Comparison

The latency figures highlight a transformative efficiency gain:

  • 1 Minute vs. 39 Minutes: Transcribing the entire 1,000 benchmark clips takes only ~1 minute with SedaNevis_v1.0, whereas Whisper-Large-Persian-v4 requires nearly 39 minutes on the exact same workload!
  • 37.5× Higher Throughput: Delivering real-time streaming capability without the severe computational overhead of autoregressive decoder architectures.

توضیح فارسی (تفاوت شگفت‌انگیز در زمان و سرعت پردازش):
برای درک سرعت خارق‌العاده مدل به این مقایسه توجه کنید:
پیاده‌سازی ۱,۰۰۰ فایل صوتی تست، با مدل SedaNevis_v1.0 تنها حدود ۱ دقیقه طول کشید؛ در حالی که همین تعداد فایل با مدل Whisper-Large-Persian-v4 نزدیک به ۳۹ دقیقه زمان برد!
یعنی صدانویس بیش از ۳۷ برابر سریع‌تر از ویسپر لارج عمل می‌کند و پردازش هم‌زمان صدها مکالمه را به‌صورت بلادرنگ (Real-Time) ممکن می‌سازد.


💡 Exceptional Parameter Efficiency & Ultra-Low Footprint (~500 MB)

A standout characteristic of SedaNevis_v1.0 is its remarkable architectural efficiency:

  • Massive Size Reduction: Weighing in at only ~500 MB, SedaNevis achieves sub-10% WER, comfortably outperforming large-scale models (such as Whisper-Large variants that require 1.5GB – 3.1GB of storage and massive compute).
  • Edge & Production Ready: The compact footprint enables deployment in resource-constrained production settings, on-premise edge servers, CPU-only environments, and client-side applications with minimal VRAM utilization.
  • Cost-Efficient Serving: Delivers enterprise-grade transcription accuracy at a fraction of the cloud hosting and hardware inference costs compared to multi-gigabyte models.

توضیح فارسی (بهره‌وری فوق‌العاده و حجم کم):
یکی از مهم‌ترین نقاط قوت مدل صدانویس، حجم کم (حدود ۵۰۰ مگابایت) در کنار دقت بسیار بالا است. در حالی که مدل‌های بزرگ و سنگین چندگیگابایتی (مانند Whisper Large) به سخت‌افزارهای گران‌قیمت با VRAM بالا نیاز دارند، صدانویس با حجمی ۶ برابر کوچک‌تر خطای کمتری ثبت کرده و به‌سادگی بر روی سرورهای اقتصادی، سیستم‌های محلی (Edge) و حتی پردازنده‌های معمولی (CPU) قابل اجراست.


📊 Evaluation & Training Data Sources

The model's training and evaluation benchmarks leverage standard, peer-reviewed collections:

  1. Elderly Speech: AliAvd/persian-elderly-asr – Focuses on disfluent pronunciations, vocal cord fatigue, and slower speaking rates.
  2. Conversational Podcasts: farsi-asr/farsi-asr-dataset – Captures real-world colloquial Farsi, overlapping speech, and modern vernacular.
  3. Studio Read Speech: Thomcles/Persian-Farsi-Speech – Clean acoustic baseline with clear enunciation.
  4. Emotional Speech: ShEMO Database – Authentic recordings across anger, happiness, sadness, and surprise.
  5. Curated Broadcast & Web Corpora: Filtered speech segments (0.4s to 30.0s) ensuring linguistic variety.

🔬 Training Methodology

To achieve strong generalization while eliminating catastrophic forgetting, the training pipeline employed several advanced techniques:

1. Multi-Stage Adaptation & Terminal Layer Stabilization

  • Progressive Fine-Tuning Pipeline: The model was trained through a multi-stage adaptation pipeline. Following comprehensive end-to-end parameter training across diverse Persian corpora, late-stage targeted freezing was utilized to stabilize foundational representations.
  • Controlled Gradient Flow: In the terminal convergence phase, gradient updates were predominantly concentrated on high-level linguistic blocks and the RNN-T prediction network. This stabilized low-level acoustic representations while refining domain-specific speech nuances, dialects, and conversational patterns.

2. Speech Text Normalization & Persian ITN

  • Spoken-Form Expansion: Cardinal, ordinal, and fractional digits were systematically expanded into verbal Persian words.
  • Lexical Canonicalization: Normalized Persian solar calendar dates, clock times, metric units, and removed non-phonemic Arabic diacritics (Harakat) while preserving morphology.
  • Zero-Width Non-Joiner (ZWNJ) Handling: Standardized ZWNJ and punctuation characters to align with token boundaries in the SentencePiece vocabulary.

3. Hyperparameters & Optimization

  • Optimizer: AdamW with a base learning rate of $5.0 \times 10^{-6}$, combined with a Cosine Annealing learning rate schedule and warmup steps.
  • Precision: Mixed Precision (FP16) for enhanced GPU throughput and reduced memory footprint.
  • Gradient Accumulation: Step accumulation factor of 8 to simulate larger, stable effective batch sizes.

4. Cross-Corpus Data Harmonization & Balanced Ingestion

  • Acoustic & Duration Harmonization: Raw audio from disparate sources exhibited varying metadata formats, acoustic properties, and audio length distributions. The entire corpus was filtered and constrained strictly within a valid temporal boundary ($0.4\text{s} \le t \le 30.0\text{s}$) to eliminate corrupted clips and prevent gradient anomalies.
  • Stratified Domain Balancing: To prevent dominant domains (e.g., studio read speech) from overpowering underrepresented, challenging acoustic domains (e.g., elderly speech, spontaneous podcasts, or emotional expressions), a stratified sampling quota was enforced, yielding an evenly balanced acoustic representation across all training steps.

💻 Inference Guide

Installation

pip install torch torchaudio nemo_toolkit[asr] soundfile
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train jafarahmadi/SedaNevis_v1.0

Space using jafarahmadi/SedaNevis_v1.0 1

Evaluation results

  • Test WER on Balanced Persian Multi-Domain Benchmark (1,000 Utterances)
    self-reported
    9.780
  • Test CER on Balanced Persian Multi-Domain Benchmark (1,000 Utterances)
    self-reported
    3.010