Instructions to use jafarahmadi/SedaNevis_v1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use jafarahmadi/SedaNevis_v1.0 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("jafarahmadi/SedaNevis_v1.0") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
🎙️ SedaNevis_v1.0 (صدانویس)
Robust Multi-Domain Persian ASR Model with FastConformer-RNNT
SedaNevis_v1.0 is an enterprise-grade, state-of-the-art Automatic Speech Recognition (ASR) model specifically tailored for the Persian (Farsi) language. Built on top of NVIDIA's robust FastConformer-RNNT architecture, it is fine-tuned across diverse and challenging acoustic conditions, including elderly speech, conversational podcasts, studio reading, and emotional speech.
توضیح فارسی (Overview):
مدل SedaNevis_v1.0 (صدانویس) یک مدل پیشرفته و بهینهسازیشده برای بازشناسی خودکار گفتار زبان فارسی (ASR) است. این مدل با هدف عملکرد پایدار در شرایط چالشبرانگیز آکوستیک (لهجهها، گفتار سالمندان، مکالمات محاورهای پادکستها، و صدای همراه با احساسات) بر پایه معماری FastConformer-RNNT توسعه یافته است.
🏆 Benchmark & Latency Comparison
Evaluation Setup: Evaluated on 1,000 diverse Persian audio files under identical hardware conditions (NVIDIA GPU).
| Rank | Model Name | Architecture | Size | Time (1k files) | Speed & Throughput ⚡ | CER (%) ↓ | WER (%) ↓ | Accuracy (%) ↑ |
|---|---|---|---|---|---|---|---|---|
| 🥇 1st | SedaNevis_v1.0 (Ours) | FastConformer-RNNT | ~500 MB | 16.13 audio/s (37.5× 🔥) | 3.01% | 9.78% | 90.22% | |
| 🥈 2nd | whisper-base-persian | Seq2Seq Transformer | ~290 MB | 3.43 audio/s (8.0×) | 6.91% | 18.52% | 81.48% | |
| 🥉 3rd | Whisper-Base-Persian-Full | Seq2Seq Transformer | ~290 MB | 3.44 audio/s (8.0×) | 7.48% | 19.69% | 80.31% | |
| 4th | Whisper-Large-v3-Turbo | Seq2Seq Transformer | ~1.6 GB | 1.68 audio/s (3.9×) | 9.22% | 27.25% | 72.75% | |
| 5th | Whisper-Large-Persian-v4 | Seq2Seq Transformer | ~5.1 GB | 0.43 audio/s (1.0× Baseline) | 8.16% | 20.97% | 79.03% |
⚡ Operational Advantage:
While achieving an unprecedented low error rate (9.78% WER), SedaNevis operates 5× to 38× faster than typical Whisper Large models, making it optimal for both real-time streaming pipelines and large-scale offline batch processing.
⚡ Unmatched Speed & Real-World Latency Comparison
The latency figures highlight a transformative efficiency gain:
- 1 Minute vs. 39 Minutes: Transcribing the entire 1,000 benchmark clips takes only ~1 minute with SedaNevis_v1.0, whereas Whisper-Large-Persian-v4 requires nearly 39 minutes on the exact same workload!
- 37.5× Higher Throughput: Delivering real-time streaming capability without the severe computational overhead of autoregressive decoder architectures.
توضیح فارسی (تفاوت شگفتانگیز در زمان و سرعت پردازش):
برای درک سرعت خارقالعاده مدل به این مقایسه توجه کنید:
پیادهسازی ۱,۰۰۰ فایل صوتی تست، با مدل SedaNevis_v1.0 تنها حدود ۱ دقیقه طول کشید؛ در حالی که همین تعداد فایل با مدل Whisper-Large-Persian-v4 نزدیک به ۳۹ دقیقه زمان برد!
یعنی صدانویس بیش از ۳۷ برابر سریعتر از ویسپر لارج عمل میکند و پردازش همزمان صدها مکالمه را بهصورت بلادرنگ (Real-Time) ممکن میسازد.
💡 Exceptional Parameter Efficiency & Ultra-Low Footprint (~500 MB)
A standout characteristic of SedaNevis_v1.0 is its remarkable architectural efficiency:
- Massive Size Reduction: Weighing in at only ~500 MB, SedaNevis achieves sub-10% WER, comfortably outperforming large-scale models (such as Whisper-Large variants that require 1.5GB – 3.1GB of storage and massive compute).
- Edge & Production Ready: The compact footprint enables deployment in resource-constrained production settings, on-premise edge servers, CPU-only environments, and client-side applications with minimal VRAM utilization.
- Cost-Efficient Serving: Delivers enterprise-grade transcription accuracy at a fraction of the cloud hosting and hardware inference costs compared to multi-gigabyte models.
توضیح فارسی (بهرهوری فوقالعاده و حجم کم):
یکی از مهمترین نقاط قوت مدل صدانویس، حجم کم (حدود ۵۰۰ مگابایت) در کنار دقت بسیار بالا است. در حالی که مدلهای بزرگ و سنگین چندگیگابایتی (مانند Whisper Large) به سختافزارهای گرانقیمت با VRAM بالا نیاز دارند، صدانویس با حجمی ۶ برابر کوچکتر خطای کمتری ثبت کرده و بهسادگی بر روی سرورهای اقتصادی، سیستمهای محلی (Edge) و حتی پردازندههای معمولی (CPU) قابل اجراست.
📊 Evaluation & Training Data Sources
The model's training and evaluation benchmarks leverage standard, peer-reviewed collections:
- Elderly Speech: AliAvd/persian-elderly-asr – Focuses on disfluent pronunciations, vocal cord fatigue, and slower speaking rates.
- Conversational Podcasts: farsi-asr/farsi-asr-dataset – Captures real-world colloquial Farsi, overlapping speech, and modern vernacular.
- Studio Read Speech: Thomcles/Persian-Farsi-Speech – Clean acoustic baseline with clear enunciation.
- Emotional Speech: ShEMO Database – Authentic recordings across anger, happiness, sadness, and surprise.
- Curated Broadcast & Web Corpora: Filtered speech segments (0.4s to 30.0s) ensuring linguistic variety.
🔬 Training Methodology
To achieve strong generalization while eliminating catastrophic forgetting, the training pipeline employed several advanced techniques:
1. Multi-Stage Adaptation & Terminal Layer Stabilization
- Progressive Fine-Tuning Pipeline: The model was trained through a multi-stage adaptation pipeline. Following comprehensive end-to-end parameter training across diverse Persian corpora, late-stage targeted freezing was utilized to stabilize foundational representations.
- Controlled Gradient Flow: In the terminal convergence phase, gradient updates were predominantly concentrated on high-level linguistic blocks and the RNN-T prediction network. This stabilized low-level acoustic representations while refining domain-specific speech nuances, dialects, and conversational patterns.
2. Speech Text Normalization & Persian ITN
- Spoken-Form Expansion: Cardinal, ordinal, and fractional digits were systematically expanded into verbal Persian words.
- Lexical Canonicalization: Normalized Persian solar calendar dates, clock times, metric units, and removed non-phonemic Arabic diacritics (Harakat) while preserving morphology.
- Zero-Width Non-Joiner (ZWNJ) Handling: Standardized ZWNJ and punctuation characters to align with token boundaries in the SentencePiece vocabulary.
3. Hyperparameters & Optimization
- Optimizer: AdamW with a base learning rate of $5.0 \times 10^{-6}$, combined with a Cosine Annealing learning rate schedule and warmup steps.
- Precision: Mixed Precision (
FP16) for enhanced GPU throughput and reduced memory footprint. - Gradient Accumulation: Step accumulation factor of 8 to simulate larger, stable effective batch sizes.
4. Cross-Corpus Data Harmonization & Balanced Ingestion
- Acoustic & Duration Harmonization: Raw audio from disparate sources exhibited varying metadata formats, acoustic properties, and audio length distributions. The entire corpus was filtered and constrained strictly within a valid temporal boundary ($0.4\text{s} \le t \le 30.0\text{s}$) to eliminate corrupted clips and prevent gradient anomalies.
- Stratified Domain Balancing: To prevent dominant domains (e.g., studio read speech) from overpowering underrepresented, challenging acoustic domains (e.g., elderly speech, spontaneous podcasts, or emotional expressions), a stratified sampling quota was enforced, yielding an evenly balanced acoustic representation across all training steps.
💻 Inference Guide
Installation
pip install torch torchaudio nemo_toolkit[asr] soundfile
- Downloads last month
- -
Datasets used to train jafarahmadi/SedaNevis_v1.0
farsi-asr/farsi-asr-dataset
AliAvd/persian-elderly-asr
Space using jafarahmadi/SedaNevis_v1.0 1
Evaluation results
- Test WER on Balanced Persian Multi-Domain Benchmark (1,000 Utterances)self-reported9.780
- Test CER on Balanced Persian Multi-Domain Benchmark (1,000 Utterances)self-reported3.010