Streaming Speech Recognition with Whisper Encoder-Decoder

Browser-deployable automatic speech recognition system built around the Whisper encoder-decoder architecture, with LoRA/PEFT experimentation, robustness benchmarking, ONNX optimization, and static client-side deployment.

Production artifact in this repository: Whisper Tiny English FP32 ONNX browser model selected from measured evaluation evidence.

Live Resources

Evaluation Overview

Whisper Model Comparison

Evaluation Dashboard

Whisper Evaluation Dashboard

Model Selection

The project evaluated Whisper Tiny, Small, and Medium in both pretrained and LoRA fine-tuned configurations.

The final browser model was selected based on accuracy, latency, memory footprint, real-time factor, and ONNX parity, rather than assuming that fine-tuning would always improve performance.

Test-Set Results

Model WER CER Avg. Latency
Tiny pretrained 6.09% 2.64% 0.118 s
Tiny LoRA 16.76% 6.64% 0.262 s
Small pretrained 3.46% 1.44% 0.232 s
Small LoRA 6.26% 2.17% 0.619 s
Medium pretrained 4.68% 2.50% 0.405 s
Medium LoRA 4.23% 1.37% 1.188 s

All values above come from the project's held-out 1,452-example test evaluation.

Why Whisper Tiny for Browser Deployment?

Whisper Small pretrained achieved the lowest overall WER, but Whisper Tiny pretrained was selected as the browser champion because it provided a substantially lighter runtime:

  • WER: 6.09%
  • CER: 2.64%
  • Average latency: 0.118 s
  • Real-Time Factor: 0.017
  • Peak measured GPU memory: ~151 MB

This provides a stronger quality/performance trade-off for browser deployment.

Fine-Tuning Findings

LoRA/PEFT fine-tuning was performed for Tiny, Small, and Medium Whisper models.

Fine-tuning did not improve every model:

  • Tiny LoRA regressed relative to Tiny pretrained.
  • Small LoRA regressed relative to Small pretrained.
  • Medium LoRA produced a measurable improvement.

Successful Medium LoRA Result

Metric Medium pretrained Medium LoRA
WER 4.68% 4.23%
CER 2.50% 1.37%

The Medium LoRA experiment is retained as the strongest fine-tuning result.

Important: the Medium LoRA checkpoint is experimental evidence from the broader project. This Hugging Face repository contains the selected Tiny browser deployment artifact.

ONNX Optimization

The selected Tiny browser model was exported to ONNX and evaluated in both FP32 and Q8 configurations.

Runtime WER Relative WER Change Prediction Match
PyTorch 6.12% Baseline β€”
FP32 ONNX 5.90% -3.57% 98%
Q8 ONNX 6.45% +5.36% 94%

The project release criterion allowed no more than 2% relative WER regression.

Therefore:

  • βœ… FP32 ONNX is the production/default browser artifact
  • ⚠️ Q8 is retained only as an experimental optimization artifact

Robustness Evaluation

The selected Tiny pretrained model was evaluated on 1,800 robustness examples derived from 300 source recordings.

Condition WER
Clean 6.02%
Clipping 5.80%
Low volume 5.76%
Mild Gaussian noise 6.83%
Medium Gaussian noise 11.15%
Heavy Gaussian noise 21.14%

Overall robustness WER: 9.45%

The largest degradation occurs under heavy additive Gaussian noise.

Browser Architecture

Microphone / Audio Upload
        ↓
Browser Audio Preprocessing
        ↓
Whisper Feature Extraction
        ↓
FP32 ONNX Encoder
        ↓
Autoregressive Decoder
        ↓
Chunk Aggregation
        ↓
Transcript + Timestamps

The application performs inference client-side without requiring a Python inference server.

Browser Features

The deployed application includes:

  • Microphone recording
  • Audio file upload
  • Bundled sample audio
  • Audio preview
  • Chunked transcription
  • Transcription progress
  • Final transcript
  • Timestamps
  • Confidence/result presentation
  • Evaluation dashboard
  • WebGPU-capable execution
  • WASM fallback

Experimental Environment

Training and evaluation were executed locally using:

  • GPU: NVIDIA RTX 5090
  • VRAM: 32 GB
  • CUDA: 13.0
  • Python: 3.12
  • PyTorch: 2.13
  • Transformers: 4.57.6
  • PEFT: 0.20.0
  • Mixed Precision: BF16

Evaluation Pipeline

The evaluation workflow measures:

  • Word Error Rate (WER)
  • Character Error Rate (CER)
  • Average latency
  • P95 latency
  • Real-Time Factor
  • GPU memory
  • Robustness slices
  • Noise degradation
  • ONNX parity
  • Quantization regression
  • Browser deployment suitability

Intended Use

This model and project are suitable for:

  • Speech recognition experimentation
  • Browser-side ASR demonstrations
  • Whisper encoder-decoder research
  • ONNX inference experimentation
  • WebGPU/WASM deployment experiments
  • Machine learning portfolio demonstrations
  • Educational use

Limitations

  • The deployment model is English-only.
  • The main benchmark is based on LibriSpeech-style English speech.
  • Accent coverage in the current benchmark is limited.
  • Performance degrades under severe additive environmental noise.
  • Browser latency depends on hardware and browser runtime.
  • Q8 quantization exceeded the project's quality-regression threshold.
  • Speech recognition output may contain transcription errors.

Responsible Use

This project is intended for experimentation, education, research, portfolio demonstration, and general transcription.

Speech recognition output should not be treated as infallible.

Human review should be used before relying on transcripts in:

  • Medical applications
  • Legal applications
  • Emergency situations
  • Safety-critical systems
  • Security-sensitive workflows
  • Other high-stakes environments

Reproducibility

The complete source code includes:

  • Dataset preparation
  • Dataset auditing
  • LoRA hyperparameter sweeps
  • Tiny/Small/Medium Whisper training
  • Pretrained vs fine-tuned evaluation
  • Robustness benchmarking
  • Failure analysis
  • ONNX export
  • Q8 quantization
  • ONNX parity validation
  • Browser model selection
  • Static web deployment
  • Automated Python tests
  • JavaScript/browser testing

The full implementation is available here:

https://github.com/unit-mole/encoder-decoder-projects/tree/main/03-streaming-speech-recognition-whisper-encoder-decoder

License

MIT

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for anmol-unitmole/streaming-speech-recognition-whisper-encoder-decoder-model

Quantized
(12)
this model