Instructions to use atasoglu/seda-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- K2
How to use atasoglu/seda-v0.1 with K2:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
seda-v0.1: streaming Turkish speech recognition
seda-v0.1 is a small, low-latency speech recognition model for Turkish. It is a streaming Zipformer2 transducer (66M parameters) trained with icefall, exported to int8 ONNX, and run with sherpa-onnx.
- Streaming. Words appear while the speaker is still talking. The final result arrives about 40 ms after the audio ends.
- Lightweight. On a single CPU thread it runs about 30× faster than real time and needs about 170 MB of RAM.
- Accurate. It reaches 11.5% WER on FLEURS-TR. That is on par with Whisper small (11.9%), which is 3.7× larger and offline-only.
Model details
| Architecture | Zipformer2 encoder + stateless decoder + joiner (RNN-T / transducer), causal |
| Parameters | 66.1M |
| Encoder | 6 stacks, layers 2,2,3,4,3,2, dims 192,256,384,512,384,256, downsampling 1,2,4,8,4,2 |
| Decoder / joiner | Context size 2, dim 512 |
| Input | 80-dim log-mel fbank, 16 kHz mono |
| Output | 500 BPE tokens. Text is lowercase spoken-form Turkish without punctuation; numbers are written as words. |
| Streaming | Chunk 32 frames (0.64 s), left context 128 frames |
| Quantization | int8 (dynamic) encoder and joiner, fp32 decoder |
| Language model | LSTM RNN LM (2×512, 4.7 MB int8) for shallow fusion, plus a token bigram for LODR |
Files
| File | Description |
|---|---|
encoder.int8.onnx |
Encoder (70 MB) |
decoder.onnx |
Decoder (2 MB) |
joiner.int8.onnx |
Joiner (0.3 MB) |
tokens.txt |
BPE-500 vocabulary |
lm/rnnlm.int8.onnx, lm/2gram.fst |
RNN LM and LODR bigram for beam search |
Usage
pip install sherpa-onnx huggingface_hub soundfile
import numpy as np
import sherpa_onnx
import soundfile as sf
from huggingface_hub import snapshot_download
d = snapshot_download("atasoglu/seda-v0.1")
recognizer = sherpa_onnx.OnlineRecognizer.from_transducer(
tokens=f"{d}/tokens.txt",
encoder=f"{d}/encoder.int8.onnx",
decoder=f"{d}/decoder.onnx",
joiner=f"{d}/joiner.int8.onnx",
num_threads=1,
sample_rate=16000,
feature_dim=80,
decoding_method="modified_beam_search",
max_active_paths=4,
lm=f"{d}/lm/rnnlm.int8.onnx", # optional: drop these four lines
lm_scale=0.6, # and use decoding_method="greedy_search"
lodr_fst=f"{d}/lm/2gram.fst", # for the fastest setup
lodr_scale=-0.5,
)
audio, sr = sf.read("speech.wav", dtype="float32", always_2d=True)
audio = audio.mean(axis=1) # mono; sherpa-onnx resamples to 16 kHz if needed
stream = recognizer.create_stream()
stream.accept_waveform(sr, audio)
stream.accept_waveform(sr, np.zeros(sr, dtype="float32")) # 1 s tail so the last chunk is decoded
stream.input_finished()
while recognizer.is_ready(stream):
recognizer.decode_stream(stream)
print(recognizer.get_result(stream))
For live input, call accept_waveform with small pieces (for example 100 ms from a microphone) and run the decode_stream loop after each piece. get_result returns the partial transcript at any time. To split speech into utterances, pass enable_endpoint_detection=True and check recognizer.is_endpoint(stream). See the sherpa-onnx examples.
Training
- Data: about 415 h in total.
- Common Voice 27.0 Turkish train: about 69 h of unique audio, about 206 h with 0.9/1.0/1.1 speed perturbation.
- ISSAI Turkish Speech Corpus train: 209 h after cleanup.
- Training sentences that also occur in the CV dev/test sets were removed.
- Recipe:
- A CV-only model was trained from scratch for 40 epochs.
- Training continued from its weights for 12 epochs on the CV + ISSAI mix (LR 0.01).
- The release model averages epochs 5–12 (avg 8).
- The released export uses a 0.64 s chunk with 128 frames of left context.
- Language model:
- Trained on Turkish Wikipedia (20231101 dump) plus the ASR training transcripts: 8.1M sentences, 109M words.
- Sentences overlapping FLEURS, CV or ISSAI dev/test were filtered out.
- The LM scales (0.6 / −0.5) were tuned on FLEURS dev only.
Evaluation
WER (%), streaming, CPU int8, chunk 32. References and hypotheses were normalized the same way: lowercase, no punctuation, numbers spelled out.
| Test set | greedy | beam 4 | beam 4 + LM |
|---|---|---|---|
| FLEURS-TR test (743 utt., zero-shot) | 15.14 | 14.24 | 11.52 |
| Common Voice 27.0 TR test (11,874 utt.) | 16.32 | 15.59 | 13.61 |
| ISSAI TSC test (3,484 utt.) | 16.00 | 15.72 | 15.23 |
Speed and latency measured on 1 CPU thread with the int8 model:
| greedy | beam 4 | beam 4 + LM | |
|---|---|---|---|
| RTF | 0.027 | 0.029 | 0.034 |
| Peak RAM | 169 MB | 169 MB | 175 MB |
- The RAM increase per extra concurrent stream is about 1.4 MB.
- With beam 4 + LM, the median first-word latency is 0.70 s and the final result arrives about 40 ms after the audio ends.
On FLEURS-TR, Whisper small (244M, offline, beam 5) scores 11.90% WER under the same normalization. The 0.4-point difference is within the 95% bootstrap confidence interval, so the two models are on par.
Limitations
- Output is spoken form. The output has no casing, punctuation or digits. A rule-based inverse text normalizer, which turns spoken numbers into digits, is not part of this release.
- Foreign names and abbreviations are the main error source. Examples include USOC and Super-G.
- Read speech dominates the training data. Accuracy on spontaneous, noisy, far-field or telephone speech has not been measured.
- ISSAI test is in-domain. It comes from the same corpus as part of the training data.
- The LM favours encyclopedic text. It gains −3.6 WER on FLEURS but only −0.8 on ISSAI. Some ISSAI audio comes from YouTube.
License
- Acoustic model (
encoder-*,decoder-*,joiner-*,tokens.txt): Apache-2.0.- It was trained on Common Voice (CC0) and the ISSAI Turkish Speech Corpus.
- The ISSAI corpus is published under MIT on Hugging Face and under CC BY 4.0 by ISSAI; attribution is given below.
- Language model files (
lm/): CC BY-SA 4.0, because they were trained on Turkish Wikipedia text.- They are optional and can be left out if ShareAlike terms do not suit your use.
Acknowledgements and citation
The model was built with k2, icefall, lhotse and sherpa-onnx. It was trained on data from Mozilla Common Voice, the ISSAI Turkish Speech Corpus and Wikipedia.
@article{mussakhojayeva2023turkic,
title = {Multilingual Speech Recognition for Turkic Languages},
author = {Mussakhojayeva, Saida and Dauletbek, Kaisar and Yeshpanov, Rustem and Varol, Huseyin Atakan},
journal = {Information},
volume = {14},
number = {2},
pages = {74},
year = {2023}
}
@inproceedings{ardila2020common,
title = {Common Voice: A Massively-Multilingual Speech Corpus},
author = {Ardila, Rosana and Branson, Megan and Davis, Kelly and Henretty, Michael and Kohler, Michael and Meyer, Josh and Morais, Reuben and Saunders, Lindsay and Tyers, Francis M. and Weber, Gregor},
booktitle = {Proceedings of LREC},
year = {2020}
}
@inproceedings{yao2024zipformer,
title = {Zipformer: A Faster and Better Encoder for Automatic Speech Recognition},
author = {Yao, Zengwei and Guo, Liyong and Yang, Xiaoyu and Kang, Wei and Kuang, Fangjun and Yang, Yifan and Jin, Zengrui and Lin, Long and Povey, Daniel},
booktitle = {ICLR},
year = {2024}
}
Datasets used to train atasoglu/seda-v0.1
issai/Turkish_Speech_Corpus
Evaluation results
- WER (beam 4 + LM, chunk 32) on FLEURS (tr_tr), testtest set self-reported11.520
- WER (beam 4 + LM, chunk 32) on Common Voice 27.0 (tr), testtest set self-reported13.610
- WER (beam 4 + LM, chunk 32) on ISSAI Turkish Speech Corpus, testtest set self-reported15.230