Neurlang Whipstr STT (ASR)
A deep learning automatic speech recognition (ASR) system for transcribing speech audio into IPA text using transformer-based sequence-to-sequence models.
- Language: Universal (IPA), 74+ languages
- Model Github: neurlang/whipstr https://github.com/neurlang/whipstr
- Model Dataset: Common Voice 21
- Model-Native Sample Rates: 8000 Hz, 16000 Hz, 24000 Hz, 32000 Hz, 48000 Hz
- Degraded-Performance Sample Rates: 11025 Hz, 22050 Hz, 44100 Hz
- License: GPL v2
- Release: 2026-07-07
- Size: 186 MB
- Total parameters:
- Encoder: 7 220 576
- Transformer: 7 537 184
- Total: 14 757 760
- CER: 48.68% (51.32% success rate)
- Note: Averaged across all supported languages, works better on higher resource languages
- WER: 92.58% (7.42% success rate)
- Note: Averaged across all supported languages, works better on higher resource languages
- Training Details:
- Hardware: Nvidia Spark
- Batch size: 1
- Samples: 960000
- Duration: 1:11:10:00
- Runs: Jul 5 14:58 - Jul 5 17:45, Jul 5 22:41 - Jul 6 04:38, Jul 6 04:46 - Jul 7 07:12 (shut down at Jul 7 07:36)
Inference code
git clone https://github.com/neurlang/whipstr.git
cd whipstr/
uv run --with torch --with transformers --with phase-spectrogram stt_infer_hf.py --audio /home/m/Downloads/LJ001-0001.wav --model neurlang/ipa-whipstr-base-48khz-cv-21
Output:
Loading weights: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 139/139 [00:00<00:00, 13371.75it/s]
Transcription: ˈɪt ˈæt juː pˈæst ðə kˈɔldz tˈɑɹk fˈeɪs vˈæst ðə ɹˈɛpɪtʃən lˈupənəl ˈi ˈæz bɪhˈeɪviɚ wˈi sˈɑloʊks lˈɛkəmˌɑɹəl hˈæn jˈæt lˈɚnd bˈeɪsɪks sˈɛktɪs ˈækmɪkmɪkəɹɪkənɪkəksɪksɪksɪksɪksɪksɪksɪksɪkstəkstəksɪksɪk
Explaination:
| IPA | Ground truth | Comment |
|---|---|---|
| ˈɪt | It | Exact match. |
| ˈæt juː | gets you | /ɡɛts/ was apparently lost, leaving something like "at you." |
| pˈæst | past | Good match. |
| ðə | the | Exact. |
| kˈɔldz | cold- | Reasonable; final /z/ is likely a transcription artifact. |
| tˈɑɹk | start | The /s/ was dropped and /st/ became /k/, a common recognition error in noisy speech. |
| fˈeɪs | phase | Very close. |
| vˈæst | fast | /f/→/v/ substitution. |
| ðə | The | Exact. |
| ɹˈɛpɪtʃən | repetition | Clearly recognizable. |
| lˈupənəl | loop / no-EOS | This portion collapsed badly; "loop no-EOS" became something like "lupənəl." |
| ˈi ˈæz | behavior (beginning omitted) | The start of "behavior" appears to have shifted. |
| bɪhˈeɪviɚ | behavior | Excellent match. |
| wˈi | we | Exact. |
| sˈɑloʊks | saw looks | This is actually close to "saw looks," with the pause removed. |
| lˈɛkəmˌɑɹəl | like the model | Quite distorted, but you can hear the rough rhythm. |
| hˈæn | hadn't | The final /t/ and /d/ were lost. |
| jˈæt | yet | Very close. |
| lˈɚnd | learned | Exact. |
| bˈeɪsɪks | basic | Essentially correct. |
| sˈɛktɪs | seq2seq | This was the biggest failure—only the initial /sɛk/ of "seq2seq" survived, while |
| loops | mechanics | "mechanics" disappeared entirely. |
End of training data:
Step 960000 is chosen as this release
| Step | WER | CER |
|---|---|---|
| 943000 | 94.21% | 49.58% |
| 944000 | 93.16% | 49.96% |
| 945000 | 94.09% | 50.22% |
| 946000 | 95.25% | 50.32% |
| 947000 | 96.18% | 50.30% |
| 948000 | 93.97% | 50.78% |
| 949000 | 92.93% | 49.63% |
| 950000 | 94.55% | 50.84% |
| 951000 | 94.90% | 50.11% |
| 952000 | 92.35% | 49.85% |
| 953000 | 93.74% | 49.08% |
| 954000 | 93.74% | 49.80% |
| 955000 | 92.70% | 49.87% |
| 956000 | 94.32% | 50.37% |
| 957000 | 94.21% | 48.90% |
| 958000 | 93.40% | 50.08% |
| 959000 | 92.82% | 50.20% |
| 960000 | 92.58% | 48.68% |
| 961000 | 95.13% | 50.32% |
| 962000 | 92.93% | 49.56% |
| 963000 | 94.21% | 50.84% |
| 964000 | 95.48% | 50.44% |
| 965000 | 94.32% | 50.39% |
| 966000 | 93.05% | 50.72% |
| 967000 | 93.97% | 51.04% |
| 968000 | 91.19% | 49.94% |
| 969000 | 95.60% | 50.82% |
| 970000 | 96.52% | 50.16% |
| 971000 | 92.82% | 51.22% |
- Downloads last month
- 94