Neurlang Whipstr STT (ASR)

A deep learning automatic speech recognition (ASR) system for transcribing speech audio into IPA text using transformer-based sequence-to-sequence models.

  • Language: Universal (IPA), 74+ languages
  • Model Github: neurlang/whipstr https://github.com/neurlang/whipstr
  • Model Dataset: Common Voice 21
  • Model-Native Sample Rates: 8000 Hz, 16000 Hz, 24000 Hz, 32000 Hz, 48000 Hz
  • Degraded-Performance Sample Rates: 11025 Hz, 22050 Hz, 44100 Hz
  • License: GPL v2
  • Release: 2026-07-07
  • Size: 186 MB
  • Total parameters:
    • Encoder: 7 220 576
    • Transformer: 7 537 184
    • Total: 14 757 760
  • CER: 48.68% (51.32% success rate)
    • Note: Averaged across all supported languages, works better on higher resource languages
  • WER: 92.58% (7.42% success rate)
    • Note: Averaged across all supported languages, works better on higher resource languages
  • Training Details:
    • Hardware: Nvidia Spark
    • Batch size: 1
    • Samples: 960000
    • Duration: 1:11:10:00
    • Runs: Jul 5 14:58 - Jul 5 17:45, Jul 5 22:41 - Jul 6 04:38, Jul 6 04:46 - Jul 7 07:12 (shut down at Jul 7 07:36)

Inference code

git clone https://github.com/neurlang/whipstr.git
cd whipstr/
uv run --with torch --with transformers --with phase-spectrogram stt_infer_hf.py --audio /home/m/Downloads/LJ001-0001.wav --model neurlang/ipa-whipstr-base-48khz-cv-21

Output:

Loading weights: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 139/139 [00:00<00:00, 13371.75it/s]
Transcription: ˈɪt ˈæt juː pˈæst ðə kˈɔldz tˈɑɹk fˈeɪs vˈæst ðə ɹˈɛpɪtʃən lˈupənəl ˈi ˈæz bɪhˈeɪviɚ wˈi sˈɑloʊks lˈɛkəmˌɑɹəl hˈæn jˈæt lˈɚnd bˈeɪsɪks sˈɛktɪs ˈækmɪkmɪkəɹɪkənɪkəksɪksɪksɪksɪksɪksɪksɪksɪkstəkstəksɪksɪk

Explaination:

IPA Ground truth Comment
ˈɪt It Exact match.
ˈæt juː gets you /ɡɛts/ was apparently lost, leaving something like "at you."
pˈæst past Good match.
ðə the Exact.
kˈɔldz cold- Reasonable; final /z/ is likely a transcription artifact.
tˈɑɹk start The /s/ was dropped and /st/ became /k/, a common recognition error in noisy speech.
fˈeɪs phase Very close.
vˈæst fast /f/→/v/ substitution.
ðə The Exact.
ɹˈɛpɪtʃən repetition Clearly recognizable.
lˈupənəl loop / no-EOS This portion collapsed badly; "loop no-EOS" became something like "lupənəl."
ˈi ˈæz behavior (beginning omitted) The start of "behavior" appears to have shifted.
bɪhˈeɪviɚ behavior Excellent match.
wˈi we Exact.
sˈɑloʊks saw looks This is actually close to "saw looks," with the pause removed.
lˈɛkəmˌɑɹəl like the model Quite distorted, but you can hear the rough rhythm.
hˈæn hadn't The final /t/ and /d/ were lost.
jˈæt yet Very close.
lˈɚnd learned Exact.
bˈeɪsɪks basic Essentially correct.
sˈɛktɪs seq2seq This was the biggest failure—only the initial /sɛk/ of "seq2seq" survived, while
loops mechanics "mechanics" disappeared entirely.

End of training data:

Step 960000 is chosen as this release

Step WER CER
943000 94.21% 49.58%
944000 93.16% 49.96%
945000 94.09% 50.22%
946000 95.25% 50.32%
947000 96.18% 50.30%
948000 93.97% 50.78%
949000 92.93% 49.63%
950000 94.55% 50.84%
951000 94.90% 50.11%
952000 92.35% 49.85%
953000 93.74% 49.08%
954000 93.74% 49.80%
955000 92.70% 49.87%
956000 94.32% 50.37%
957000 94.21% 48.90%
958000 93.40% 50.08%
959000 92.82% 50.20%
960000 92.58% 48.68%
961000 95.13% 50.32%
962000 92.93% 49.56%
963000 94.21% 50.84%
964000 95.48% 50.44%
965000 94.32% 50.39%
966000 93.05% 50.72%
967000 93.97% 51.04%
968000 91.19% 49.94%
969000 95.60% 50.82%
970000 96.52% 50.16%
971000 92.82% 51.22%
Downloads last month
94
Safetensors
Model size
17.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support