seda-v0.1: streaming Turkish speech recognition

seda-v0.1 is a small, low-latency speech recognition model for Turkish. It is a streaming Zipformer2 transducer (66M parameters) trained with icefall, exported to int8 ONNX, and run with sherpa-onnx.

  • Streaming. Words appear while the speaker is still talking. The final result arrives about 40 ms after the audio ends.
  • Lightweight. On a single CPU thread it runs about 30× faster than real time and needs about 170 MB of RAM.
  • Accurate. It reaches 11.5% WER on FLEURS-TR. That is on par with Whisper small (11.9%), which is 3.7× larger and offline-only.

Model details

Architecture Zipformer2 encoder + stateless decoder + joiner (RNN-T / transducer), causal
Parameters 66.1M
Encoder 6 stacks, layers 2,2,3,4,3,2, dims 192,256,384,512,384,256, downsampling 1,2,4,8,4,2
Decoder / joiner Context size 2, dim 512
Input 80-dim log-mel fbank, 16 kHz mono
Output 500 BPE tokens. Text is lowercase spoken-form Turkish without punctuation; numbers are written as words.
Streaming Chunk 32 frames (0.64 s), left context 128 frames
Quantization int8 (dynamic) encoder and joiner, fp32 decoder
Language model LSTM RNN LM (2×512, 4.7 MB int8) for shallow fusion, plus a token bigram for LODR

Files

File Description
encoder.int8.onnx Encoder (70 MB)
decoder.onnx Decoder (2 MB)
joiner.int8.onnx Joiner (0.3 MB)
tokens.txt BPE-500 vocabulary
lm/rnnlm.int8.onnx, lm/2gram.fst RNN LM and LODR bigram for beam search

Usage

pip install sherpa-onnx huggingface_hub soundfile
import numpy as np
import sherpa_onnx
import soundfile as sf
from huggingface_hub import snapshot_download

d = snapshot_download("atasoglu/seda-v0.1")
recognizer = sherpa_onnx.OnlineRecognizer.from_transducer(
    tokens=f"{d}/tokens.txt",
    encoder=f"{d}/encoder.int8.onnx",
    decoder=f"{d}/decoder.onnx",
    joiner=f"{d}/joiner.int8.onnx",
    num_threads=1,
    sample_rate=16000,
    feature_dim=80,
    decoding_method="modified_beam_search",
    max_active_paths=4,
    lm=f"{d}/lm/rnnlm.int8.onnx",  # optional: drop these four lines
    lm_scale=0.6,                  # and use decoding_method="greedy_search"
    lodr_fst=f"{d}/lm/2gram.fst",  # for the fastest setup
    lodr_scale=-0.5,
)

audio, sr = sf.read("speech.wav", dtype="float32", always_2d=True)
audio = audio.mean(axis=1)  # mono; sherpa-onnx resamples to 16 kHz if needed

stream = recognizer.create_stream()
stream.accept_waveform(sr, audio)
stream.accept_waveform(sr, np.zeros(sr, dtype="float32"))  # 1 s tail so the last chunk is decoded
stream.input_finished()
while recognizer.is_ready(stream):
    recognizer.decode_stream(stream)
print(recognizer.get_result(stream))

For live input, call accept_waveform with small pieces (for example 100 ms from a microphone) and run the decode_stream loop after each piece. get_result returns the partial transcript at any time. To split speech into utterances, pass enable_endpoint_detection=True and check recognizer.is_endpoint(stream). See the sherpa-onnx examples.

Training

  • Data: about 415 h in total.
    • Common Voice 27.0 Turkish train: about 69 h of unique audio, about 206 h with 0.9/1.0/1.1 speed perturbation.
    • ISSAI Turkish Speech Corpus train: 209 h after cleanup.
    • Training sentences that also occur in the CV dev/test sets were removed.
  • Recipe:
    1. A CV-only model was trained from scratch for 40 epochs.
    2. Training continued from its weights for 12 epochs on the CV + ISSAI mix (LR 0.01).
    3. The release model averages epochs 5–12 (avg 8).
    • The released export uses a 0.64 s chunk with 128 frames of left context.
  • Language model:
    • Trained on Turkish Wikipedia (20231101 dump) plus the ASR training transcripts: 8.1M sentences, 109M words.
    • Sentences overlapping FLEURS, CV or ISSAI dev/test were filtered out.
    • The LM scales (0.6 / −0.5) were tuned on FLEURS dev only.

Evaluation

WER (%), streaming, CPU int8, chunk 32. References and hypotheses were normalized the same way: lowercase, no punctuation, numbers spelled out.

Test set greedy beam 4 beam 4 + LM
FLEURS-TR test (743 utt., zero-shot) 15.14 14.24 11.52
Common Voice 27.0 TR test (11,874 utt.) 16.32 15.59 13.61
ISSAI TSC test (3,484 utt.) 16.00 15.72 15.23

Speed and latency measured on 1 CPU thread with the int8 model:

greedy beam 4 beam 4 + LM
RTF 0.027 0.029 0.034
Peak RAM 169 MB 169 MB 175 MB
  • The RAM increase per extra concurrent stream is about 1.4 MB.
  • With beam 4 + LM, the median first-word latency is 0.70 s and the final result arrives about 40 ms after the audio ends.

On FLEURS-TR, Whisper small (244M, offline, beam 5) scores 11.90% WER under the same normalization. The 0.4-point difference is within the 95% bootstrap confidence interval, so the two models are on par.

Limitations

  • Output is spoken form. The output has no casing, punctuation or digits. A rule-based inverse text normalizer, which turns spoken numbers into digits, is not part of this release.
  • Foreign names and abbreviations are the main error source. Examples include USOC and Super-G.
  • Read speech dominates the training data. Accuracy on spontaneous, noisy, far-field or telephone speech has not been measured.
  • ISSAI test is in-domain. It comes from the same corpus as part of the training data.
  • The LM favours encyclopedic text. It gains −3.6 WER on FLEURS but only −0.8 on ISSAI. Some ISSAI audio comes from YouTube.

License

  • Acoustic model (encoder-*, decoder-*, joiner-*, tokens.txt): Apache-2.0.
    • It was trained on Common Voice (CC0) and the ISSAI Turkish Speech Corpus.
    • The ISSAI corpus is published under MIT on Hugging Face and under CC BY 4.0 by ISSAI; attribution is given below.
  • Language model files (lm/): CC BY-SA 4.0, because they were trained on Turkish Wikipedia text.
    • They are optional and can be left out if ShareAlike terms do not suit your use.

Acknowledgements and citation

The model was built with k2, icefall, lhotse and sherpa-onnx. It was trained on data from Mozilla Common Voice, the ISSAI Turkish Speech Corpus and Wikipedia.

@article{mussakhojayeva2023turkic,
  title   = {Multilingual Speech Recognition for Turkic Languages},
  author  = {Mussakhojayeva, Saida and Dauletbek, Kaisar and Yeshpanov, Rustem and Varol, Huseyin Atakan},
  journal = {Information},
  volume  = {14},
  number  = {2},
  pages   = {74},
  year    = {2023}
}
@inproceedings{ardila2020common,
  title     = {Common Voice: A Massively-Multilingual Speech Corpus},
  author    = {Ardila, Rosana and Branson, Megan and Davis, Kelly and Henretty, Michael and Kohler, Michael and Meyer, Josh and Morais, Reuben and Saunders, Lindsay and Tyers, Francis M. and Weber, Gregor},
  booktitle = {Proceedings of LREC},
  year      = {2020}
}
@inproceedings{yao2024zipformer,
  title     = {Zipformer: A Faster and Better Encoder for Automatic Speech Recognition},
  author    = {Yao, Zengwei and Guo, Liyong and Yang, Xiaoyu and Kang, Wei and Kuang, Fangjun and Yang, Yifan and Jin, Zengrui and Lin, Long and Povey, Daniel},
  booktitle = {ICLR},
  year      = {2024}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train atasoglu/seda-v0.1

Evaluation results