acoustic-abstention

Acoustic Abstention Encoding: Teaching Speech Recognition Encoders to Stay Silent on Non-Speech Audio

Aditya Pujari and Ajita Rattani, University of North Texas

Released checkpoints for acoustic-abstention. Acoustic abstention encoding (AAE) trains low-rank adapters in the upper audio encoder and the audio projector of a recognizer so that its frozen decoder reads non-speech as an empty transcript, while a per-frame loss keeps the features of speech, mixed with that same non-speech clip, close to their original values. The adapters are merged into the weights: a checkpoint has the original architecture and parameter count, and decodes with the same procedure as the original model. Every number below is measured on complete evaluation sets.

Hallucination rate (%, share of clips that receive any text) on the complete non-speech sets, original model β†’ with AAE:

model UrbanSound8K (8,732) WHAM! noise (8,000) filtered FSD50K (6,414)
Whisper large-v3 100.00 β†’ 0.34 100.00 β†’ 0.00 100.00 β†’ 0.02
Qwen3-ASR-1.7B 2.68 β†’ 0.11 9.20 β†’ 0.03 0.59 β†’ 0.12
ARK-ASR-3B 99.86 β†’ 0.11 100.00 β†’ 0.00 99.81 β†’ 0.03
Hojo-ASR-V1 99.99 β†’ 0.18 100.00 β†’ 0.00 99.98 β†’ 0.14

Speech error rate on the complete test sets (WER; CER for the Mandarin sets), original model β†’ with AAE:

model LibriSpeech clean other tail 5 dB FLEURS Mandarin AISHELL-1
Whisper large-v3 2.15 β†’ 2.15 4.02 β†’ 4.07 3.44 β†’ 3.57 5.17 β†’ 5.62 7.60 β†’ 7.72 8.10 β†’ 8.10
Qwen3-ASR-1.7B 1.71 β†’ 1.71 3.47 β†’ 3.53 2.61 β†’ 2.63 4.02 β†’ 3.86 7.19 β†’ 7.44 1.53 β†’ 1.50
ARK-ASR-3B 1.43 β†’ 1.43 2.84 β†’ 2.84 2.01 β†’ 2.07 3.96 β†’ 3.98 8.76 β†’ 8.75 1.83 β†’ 1.83
Hojo-ASR-V1 1.79 β†’ 1.80 4.01 β†’ 4.07 3.04 β†’ 3.06 5.71 β†’ 5.78 15.22 β†’ 15.44 1.71 β†’ 1.71

tail: a LibriSpeech test utterance, 1 s of silence, then 3 s of noise; 5 dB: noise added at 5 dB SNR.

Ablation study on the complete sets: mean hallucination rate over the three non-speech sets (%), and in parentheses the number of the six speech sets whose error rate exceeds the original model's by more than the tolerance fixed in advance (+0.3 WER test-clean, +0.5 test-other, +1.0 on each noisy set, +1.0 CER on each Mandarin set); bold: every speech set within its tolerance.

variant Whisper large-v3 Qwen3-ASR-1.7B ARK-ASR-3B Hojo-ASR-V1
AAE (full recipe) 0.35 (0) 0.09 (0) 0.05 (0) 0.06 (0)
no preservation, Ξ» = 0 0.00 (6) 0.00 (6) 0.00 (6) 0.00 (6)
transcript CE on m 0.18 (2) 0.04 (3) 0.12 (2) 0.02 (2)
independent noise in m 0.03 (0) 0.08 (0) 0.39 (0) 0.04 (0)
adapters in the decoder 1.09 (0) 0.35 (2) 0.05 (0) 0.06 (3)

Checkpoints

Twenty-two checkpoints: the released model for each recognizer, the four variants of the ablation study for each, and the two full-recipe rows of the ablation that differ from the released model: the first Hojo-ASR-V1 run (aae-run1) and Whisper large-v3 at 8,500 steps (aae-8500). Together they cover every evaluation behind the paper and the repository's results.

folder what it is adapter parameters
whisper-large-v3/aae/ the full recipe 5.90M
whisper-large-v3/aae-8500/ Whisper large-v3 only: the aae run at 8,500 steps, the length of every variant (the full-recipe row of the ablation study); aae continues the same run to 9,250 steps 5.90M
whisper-large-v3/no-preservation/ preservation loss removed (lambda = 0) 5.90M
whisper-large-v3/transcript-ce/ preservation loss replaced by cross-entropy on m toward the original model's transcript of s 5.90M
whisper-large-v3/independent-noise/ the noise in m drawn independently of n 5.90M
whisper-large-v3/decoder-adapter/ rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m 2.95M
qwen3-asr-1.7b/aae/ the full recipe 3.62M
qwen3-asr-1.7b/no-preservation/ preservation loss removed (lambda = 0) 3.62M
qwen3-asr-1.7b/transcript-ce/ preservation loss replaced by cross-entropy on m toward the original model's transcript of s 3.62M
qwen3-asr-1.7b/independent-noise/ the noise in m drawn independently of n 3.62M
qwen3-asr-1.7b/decoder-adapter/ rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m 3.61M
ark-asr-3b/aae/ the full recipe 6.14M
ark-asr-3b/no-preservation/ preservation loss removed (lambda = 0) 6.14M
ark-asr-3b/transcript-ce/ preservation loss replaced by cross-entropy on m toward the original model's transcript of s 6.14M
ark-asr-3b/independent-noise/ the noise in m drawn independently of n 6.14M
ark-asr-3b/decoder-adapter/ rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m 4.15M
hojo-asr-v1/aae/ the full recipe 7.29M
hojo-asr-v1/aae-run1/ Hojo-ASR-V1 only: the first run of the full recipe (the full-recipe row of the ablation study) 7.29M
hojo-asr-v1/no-preservation/ preservation loss removed (lambda = 0) 7.29M
hojo-asr-v1/transcript-ce/ preservation loss replaced by cross-entropy on m toward the original model's transcript of s 7.29M
hojo-asr-v1/independent-noise/ the noise in m drawn independently of n 7.29M
hojo-asr-v1/decoder-adapter/ rank-9 LoRA on the decoder self-attention instead of the encoder, with transcript CE on m 6.64M

Each folder holds model.safetensors, the merged float16 weights of every tensor the adapter changed (the exact tensors that were evaluated), and training.json, the training arguments, loss curve and held-out curve. hojo-asr-v1/aae is the second of two runs of the recipe; the first, aae-run1, is the full-recipe row of the ablation study because its sampling matches the variants. whisper-large-v3/aae trained for 9,250 steps; whisper-large-v3/aae-8500 is the same run at 8,500 steps, the length of every variant, and is that model's full-recipe row.

Use

git clone https://github.com/pujariaditya/acoustic-abstention && cd acoustic-abstention
./setup.sh                                   # two environments + the four original recognizers, pinned
envs/main/bin/python scripts/download_weights.py     # these checkpoints, sha256-verified

Applying a checkpoint copies its tensors into the original model:

from aae import models, weights
ad = models.load("whisper-large-v3")                           # original model, pinned revision
weights.apply(ad.model, weights.resolve("whisper-large-v3/aae"))
print(ad.transcribe([waveform_16k]))                           # [''] for non-speech

The Whisper checkpoints use the parameter names of openai-whisper's large-v3.pt (sha256 e5b1a55b...), not the transformers layout. The other base models are pinned to Qwen/Qwen3-ASR-1.7B-hf @ bcd2b5b7, Edge0/ARK-ASR-3B @ 1e28271b and HojoAI/Hojo-ASR-V1 @ a22c3818. ./scripts/reproduce.sh re-runs every evaluation behind the paper and checks each number against the paper's value.

Intended use and limitations

Research artifact for reproducing the reported results. The recipe preserves each original model's speech behavior rather than improving it. When every utterance of LibriSpeech test-other and FLEURS Mandarin test is mixed with training-pool noise at -5 dB SNR, 2.4% to 6.0% receive an empty output with AAE, against at most 1.0% for the original models; at 5 and 15 dB, at most 0.15% do. The noisy speech test sets use noise from the training pools. One training run per model and variant (two runs of the full recipe for Hojo-ASR-V1, both released).

Licenses

Each checkpoint is a derivative of its base model and is distributed under that model's license: Whisper large-v3 (MIT), Qwen3-ASR-1.7B, ARK-ASR-3B and Hojo-ASR-V1 (Apache-2.0); LICENSE.md has the folder-by-folder terms and the full texts. The checkpoints were trained on audio from WHAM! noise, FSD50K, MUSAN, ESC-50, DEMAND, LibriSpeech, FLEURS and AISHELL-1, each under its own terms; some are non-commercial (ESC-50 is CC BY-NC 3.0, and FSD50K clips carry per-clip Creative Commons licenses, some of them CC BY-NC). The code is MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for RootAccess4Life/acoustic-abstention

Base model

Edge0/ARK-ASR-3B
Adapter
(1)
this model