acoustic-abstention
Acoustic Abstention Encoding: Teaching Speech Recognition Encoders to Stay Silent on Non-Speech Audio
Aditya Pujari and Ajita Rattani, University of North Texas
Released checkpoints for acoustic-abstention. Acoustic abstention encoding (AAE) trains low-rank adapters in the upper audio encoder and the audio projector of a recognizer so that its frozen decoder reads non-speech as an empty transcript, while a per-frame loss keeps the features of speech, mixed with that same non-speech clip, close to their original values. The adapters are merged into the weights: a checkpoint has the original architecture and parameter count, and decodes with the same procedure as the original model. Every number below is measured on complete evaluation sets.
Hallucination rate (%, share of clips that receive any text) on the complete non-speech sets, original model β with AAE:
| model | UrbanSound8K (8,732) | WHAM! noise (8,000) | filtered FSD50K (6,414) |
|---|---|---|---|
| Whisper large-v3 | 100.00 β 0.34 | 100.00 β 0.00 | 100.00 β 0.02 |
| Qwen3-ASR-1.7B | 2.68 β 0.11 | 9.20 β 0.03 | 0.59 β 0.12 |
| ARK-ASR-3B | 99.86 β 0.11 | 100.00 β 0.00 | 99.81 β 0.03 |
| Hojo-ASR-V1 | 99.99 β 0.18 | 100.00 β 0.00 | 99.98 β 0.14 |
Speech error rate on the complete test sets (WER; CER for the Mandarin sets), original model β with AAE:
| model | LibriSpeech clean | other | tail | 5 dB | FLEURS Mandarin | AISHELL-1 |
|---|---|---|---|---|---|---|
| Whisper large-v3 | 2.15 β 2.15 | 4.02 β 4.07 | 3.44 β 3.57 | 5.17 β 5.62 | 7.60 β 7.72 | 8.10 β 8.10 |
| Qwen3-ASR-1.7B | 1.71 β 1.71 | 3.47 β 3.53 | 2.61 β 2.63 | 4.02 β 3.86 | 7.19 β 7.44 | 1.53 β 1.50 |
| ARK-ASR-3B | 1.43 β 1.43 | 2.84 β 2.84 | 2.01 β 2.07 | 3.96 β 3.98 | 8.76 β 8.75 | 1.83 β 1.83 |
| Hojo-ASR-V1 | 1.79 β 1.80 | 4.01 β 4.07 | 3.04 β 3.06 | 5.71 β 5.78 | 15.22 β 15.44 | 1.71 β 1.71 |
tail: a LibriSpeech test utterance, 1 s of silence, then 3 s of noise; 5 dB: noise added at 5 dB SNR.
Ablation study on the complete sets: mean hallucination rate over the three non-speech sets (%), and in parentheses the number of the six speech sets whose error rate exceeds the original model's by more than the tolerance fixed in advance (+0.3 WER test-clean, +0.5 test-other, +1.0 on each noisy set, +1.0 CER on each Mandarin set); bold: every speech set within its tolerance.
| variant | Whisper large-v3 | Qwen3-ASR-1.7B | ARK-ASR-3B | Hojo-ASR-V1 |
|---|---|---|---|---|
| AAE (full recipe) | 0.35 (0) | 0.09 (0) | 0.05 (0) | 0.06 (0) |
| no preservation, Ξ» = 0 | 0.00 (6) | 0.00 (6) | 0.00 (6) | 0.00 (6) |
| transcript CE on m | 0.18 (2) | 0.04 (3) | 0.12 (2) | 0.02 (2) |
| independent noise in m | 0.03 (0) | 0.08 (0) | 0.39 (0) | 0.04 (0) |
| adapters in the decoder | 1.09 (0) | 0.35 (2) | 0.05 (0) | 0.06 (3) |
Checkpoints
Twenty-two checkpoints: the released model for each recognizer, the four variants of the ablation study for each,
and the two full-recipe rows of the ablation that differ from the released model: the first Hojo-ASR-V1 run
(aae-run1) and Whisper large-v3 at 8,500 steps (aae-8500). Together they cover every evaluation behind the paper
and the repository's results.
| folder | what it is | adapter parameters |
|---|---|---|
whisper-large-v3/aae/ |
the full recipe | 5.90M |
whisper-large-v3/aae-8500/ |
Whisper large-v3 only: the aae run at 8,500 steps, the length of every variant (the full-recipe row of the ablation study); aae continues the same run to 9,250 steps | 5.90M |
whisper-large-v3/no-preservation/ |
preservation loss removed (lambda = 0) | 5.90M |
whisper-large-v3/transcript-ce/ |
preservation loss replaced by cross-entropy on m toward the original model's transcript of s | 5.90M |
whisper-large-v3/independent-noise/ |
the noise in m drawn independently of n | 5.90M |
whisper-large-v3/decoder-adapter/ |
rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m | 2.95M |
qwen3-asr-1.7b/aae/ |
the full recipe | 3.62M |
qwen3-asr-1.7b/no-preservation/ |
preservation loss removed (lambda = 0) | 3.62M |
qwen3-asr-1.7b/transcript-ce/ |
preservation loss replaced by cross-entropy on m toward the original model's transcript of s | 3.62M |
qwen3-asr-1.7b/independent-noise/ |
the noise in m drawn independently of n | 3.62M |
qwen3-asr-1.7b/decoder-adapter/ |
rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m | 3.61M |
ark-asr-3b/aae/ |
the full recipe | 6.14M |
ark-asr-3b/no-preservation/ |
preservation loss removed (lambda = 0) | 6.14M |
ark-asr-3b/transcript-ce/ |
preservation loss replaced by cross-entropy on m toward the original model's transcript of s | 6.14M |
ark-asr-3b/independent-noise/ |
the noise in m drawn independently of n | 6.14M |
ark-asr-3b/decoder-adapter/ |
rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m | 4.15M |
hojo-asr-v1/aae/ |
the full recipe | 7.29M |
hojo-asr-v1/aae-run1/ |
Hojo-ASR-V1 only: the first run of the full recipe (the full-recipe row of the ablation study) | 7.29M |
hojo-asr-v1/no-preservation/ |
preservation loss removed (lambda = 0) | 7.29M |
hojo-asr-v1/transcript-ce/ |
preservation loss replaced by cross-entropy on m toward the original model's transcript of s | 7.29M |
hojo-asr-v1/independent-noise/ |
the noise in m drawn independently of n | 7.29M |
hojo-asr-v1/decoder-adapter/ |
rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m | 6.64M |
Each folder holds model.safetensors, the merged float16 weights of every tensor the adapter changed (the exact
tensors that were evaluated), and training.json, the training arguments, loss curve and held-out curve.
hojo-asr-v1/aae is the second of two runs of the recipe; the first, aae-run1, is the full-recipe row of the
ablation study because its sampling matches the variants. whisper-large-v3/aae trained for 9,250 steps;
whisper-large-v3/aae-8500 is the same run at 8,500 steps, the length of every variant, and is that model's
full-recipe row.
Use
git clone https://github.com/pujariaditya/acoustic-abstention && cd acoustic-abstention
./setup.sh # two environments + the four original recognizers, pinned
envs/main/bin/python scripts/download_weights.py # these checkpoints, sha256-verified
Applying a checkpoint copies its tensors into the original model:
from aae import models, weights
ad = models.load("whisper-large-v3") # original model, pinned revision
weights.apply(ad.model, weights.resolve("whisper-large-v3/aae"))
print(ad.transcribe([waveform_16k])) # [''] for non-speech
The Whisper checkpoints use the parameter names of openai-whisper's large-v3.pt (sha256 e5b1a55b...), not the
transformers layout. The other base models are pinned to Qwen/Qwen3-ASR-1.7B-hf @ bcd2b5b7, Edge0/ARK-ASR-3B @
1e28271b and HojoAI/Hojo-ASR-V1 @ a22c3818. ./scripts/reproduce.sh re-runs every evaluation behind the paper and
checks each number against the paper's value.
Intended use and limitations
Research artifact for reproducing the reported results. The recipe preserves each original model's speech behavior rather than improving it. When every utterance of LibriSpeech test-other and FLEURS Mandarin test is mixed with training-pool noise at -5 dB SNR, 2.4% to 6.0% receive an empty output with AAE, against at most 1.0% for the original models; at 5 and 15 dB, at most 0.15% do. The noisy speech test sets use noise from the training pools. One training run per model and variant (two runs of the full recipe for Hojo-ASR-V1, both released).
Licenses
Each checkpoint is a derivative of its base model and is distributed under that model's license: Whisper large-v3
(MIT), Qwen3-ASR-1.7B, ARK-ASR-3B and Hojo-ASR-V1 (Apache-2.0); LICENSE.md has the folder-by-folder terms and the
full texts. The checkpoints were trained on audio from WHAM! noise, FSD50K, MUSAN, ESC-50, DEMAND, LibriSpeech, FLEURS
and AISHELL-1, each under its own terms; some are non-commercial (ESC-50 is CC BY-NC 3.0, and FSD50K clips carry
per-clip Creative Commons licenses, some of them CC BY-NC). The code is MIT.
Model tree for RootAccess4Life/acoustic-abstention
Base model
Edge0/ARK-ASR-3B