Surround Sound: speech event detector
Surround Sound - Speech Event Detection
Detect non-verbal vocalizations and environmental sounds with timestamped predictions for richer speech transcription.
Demo: Surround Sound Space: a 59-minute documentary with the detected sounds, per-tag probabilities and an aligned transcript, synced to the video.
A small detector (5.8M parameters) that finds non-verbal vocal sounds (breath, cough, laugh, sigh, sniff, throat clearing, …) and background sounds (music, horn, vehicle, dog bark, siren, rain, applause, …) in recordings, with start and end times. It covers 54 classes. The included code also merges detected events into a Whisper transcript. An illustrative example:
so I was saying [cough] that the <horn> road was blocked </horn> [breath] anyway
[tag] marks a sound made by the speaker, placed where it starts. <tag> … </tag> wraps the words a background sound overlaps.
- Input: audio of any length (any format ffmpeg reads), resampled to 16 kHz mono.
- Output: a list of
{label, start, end, prob}events, one probability per class every 40 ms, or a transcript with tags. - Speed: about 39× real time on 8 CPU threads and 380× on an NVIDIA T4 (3 minutes of audio).
Quick start
pip install torch torchaudio numpy scipy safetensors huggingface_hub # plus the ffmpeg binary
pip install openai-whisper # optional, for transcripts
hf download theboringai-work/surround-sound --local-dir surround-sound
cd surround-sound
python inference.py my_audio.mp3 # transcript with tags + event list
python inference.py my_audio.mp3 --asr none # events only
python inference.py recordings/ --asr none --json events.jsonl
From Python, run from the downloaded folder:
from sed import Detector, rich_transcript, clean
det = Detector.from_pretrained("theboringai-work/surround-sound", profile="sensitive") # or a local folder path
events = det.detect("my_audio.mp3")
# [{'label': 'cough', 'start': 0.68, 'end': 2.0, 'prob': 0.92}, {'label': 'breath', 'start': 7.72, 'end': 8.72, 'prob': 0.94}, ...]
times, probs = det.frame_probs(audio_16k) # raw per-class probabilities on the 40 ms grid
# with any ASR that gives word timestamps: words = [(word, start_s, end_s), ...]
print(rich_transcript(words, clean(words, events)))
Useful options of inference.py:
| option | meaning |
|---|---|
--profile sensitive|balanced |
threshold set (below) |
--hide breath,lip_smack |
leave tags out |
--thr cough=0.5,machine=1.01 |
change thresholds for this run (above 1 switches a class off) |
--asr none|openai-whisper|transformers |
transcription backend |
--language, --asr-model |
Whisper language and size |
--times |
include times inside the tags |
--json out.jsonl |
machine-readable output |
Files
| file | content |
|---|---|
model.safetensors, config.json |
weights, class list and settings |
thresholds_sensitive.json |
per-class thresholds tuned for F2 (default) |
thresholds.json |
per-class thresholds tuned for F1 |
detect_config.json |
post-processing setting (hysteresis ratio) |
sed/, inference.py |
inference code (PyTorch) |
Model
audio 16 kHz ─► log-mel (64 bands, 25 ms window, 10 ms hop)
─► 5 conv blocks (3×3, BatchNorm, ReLU; frequency /16, time /4)
─► linear projection + sinusoidal positions
─► Transformer encoder (6 layers, d=256, 8 heads, pre-norm, GELU)
─► linear + sigmoid: 54 probabilities every 40 ms
The model sees 8 s windows. Long audio is processed with a 2 s hop, and overlapping predictions are averaged. Each class's probability track is then median-filtered over 200 ms and thresholded with hysteresis (an event starts above the class threshold and continues while above 0.6 × threshold). Events closer than 0.3 s (speaker sounds) or 1.0 s (background) are joined, and events shorter than 0.15 s are dropped.
Threshold profiles
Every class has its own threshold, tuned on real held-out recordings by 2-fold cross-validation.
| profile | tuned for | 43 reliably measured classes (frame level, cross-validated) |
|---|---|---|
sensitive (default) |
F2: a missed sound costs twice a false tag | recall 0.52, precision 0.27 |
balanced |
F1 | recall 0.41, precision 0.35 |
Classes and accuracy
frame AP: frame-level average precision on real held-out windows. CV F1: cross-validated frame F1 at the balanced threshold. found: share of held-out real events detected (overlapping prediction of the same class) with the sensitive profile. weak: too little held-out data (< 1 s of positive frames) or CV F1 < 0.1; treat these tags as hints.
Speaker sounds, [tag]
| tag | frame AP | CV F1 (balanced) | held-out events found (sensitive) | status |
|---|---|---|---|---|
sigh |
0.72 | 0.65 | 84% of 389 | reliable |
snore |
0.55 | 0.61 | 68% of 28 | reliable |
laugh |
0.52 | 0.47 | 79% of 587 | reliable |
breath |
0.50 | 0.50 | 80% of 1,116 | reliable |
sniff |
0.44 | 0.42 | 65% of 272 | reliable |
sneeze |
0.43 | 0.21 | 0% of 8 | reliable |
scream |
0.42 | 0.13 | 16% of 32 | reliable |
throat_clear |
0.34 | 0.38 | 62% of 190 | reliable |
cough |
0.31 | 0.40 | 56% of 329 | reliable |
whisper |
0.30 | 0.37 | 36% of 36 | reliable |
burp |
0.22 | 0.28 | 57% of 7 | reliable |
shout |
0.15 | 0.26 | 45% of 97 | reliable |
cry |
0.07 | 0.05 | 43% of 61 | weak |
gasp |
0.07 | 0.13 | 37% of 41 | reliable |
lip_smack |
0.05 | 0.13 | 22% of 9 | reliable |
chew |
0.04 | 0.06 | 10% of 29 | weak |
groan |
0.01 | 0.03 | 19% of 36 | weak |
yawn |
0.01 | 0.01 | 7% of 14 | weak |
hiccup |
0.00 | 0.01 | 21% of 14 | weak |
nose_blow |
0.00 | 0.00 | 0% of 1 | weak |
Background sounds, <tag> … </tag>
| tag | frame AP | CV F1 (balanced) | held-out events found (sensitive) | status |
|---|---|---|---|---|
music |
0.88 | 0.81 | 84% of 1,361 | reliable |
horn |
0.59 | 0.56 | 76% of 258 | reliable |
applause |
0.58 | 0.48 | 34% of 199 | reliable |
whistle |
0.52 | 0.49 | 28% of 83 | reliable |
dog_bark |
0.51 | 0.47 | 66% of 128 | reliable |
vehicle |
0.50 | 0.49 | 87% of 549 | reliable |
siren |
0.49 | 0.47 | 58% of 114 | reliable |
bird |
0.40 | 0.43 | 64% of 791 | reliable |
water |
0.39 | 0.38 | 75% of 192 | reliable |
insect |
0.38 | 0.43 | 42% of 87 | reliable |
cheer |
0.32 | 0.33 | 26% of 73 | reliable |
crowd |
0.32 | 0.24 | 61% of 41 | reliable |
wind |
0.32 | 0.30 | 68% of 206 | reliable |
gunshot |
0.31 | 0.32 | 81% of 75 | reliable |
bell |
0.26 | 0.35 | 80% of 87 | reliable |
explosion |
0.22 | 0.46 | 63% of 54 | reliable |
baby_cry |
0.22 | 0.40 | 39% of 31 | reliable |
construction |
0.21 | 0.33 | 67% of 220 | reliable |
glass_break |
0.20 | 0.21 | 12% of 58 | reliable |
dishes |
0.20 | 0.20 | 42% of 50 | reliable |
phone_vibrate |
0.17 | 0.40 | 25% of 4 | reliable |
rain |
0.16 | 0.28 | 12% of 43 | reliable |
animal |
0.15 | 0.14 | 52% of 538 | reliable |
child_voice |
0.15 | 0.19 | 48% of 179 | reliable |
machine |
0.14 | 0.28 | 51% of 137 | reliable |
thunder |
0.10 | 0.19 | 54% of 13 | reliable |
buzzer |
0.09 | 0.12 | 20% of 25 | reliable |
beep |
0.09 | 0.18 | 56% of 165 | reliable |
phone_ring |
0.09 | 0.20 | 70% of 64 | reliable |
door |
0.08 | 0.03 | 5% of 21 | weak |
tv_speaker |
0.04 | 0.07 | 44% of 23 | weak |
footsteps |
0.04 | 0.08 | 50% of 218 | weak |
fan |
0.04 | 0.04 | 15% of 13 | weak |
knock |
0.00 | 0.00 | 0% of 5 | weak |
Evaluation
All evaluation uses real recordings held out from training:
- NonVerbalSpeech-38K: a 10% hash split.
- Vaani Noise Event Timestamps: verified clips from 14 held-out districts.
- NonverbalTTS: dev + test.
- AudioSet-Strong: the evaluation split.
Each class is scored only on the datasets that label it. Whole clips are run through the full detector; a true event is found if a same-class prediction overlaps it, and a prediction is correct if it overlaps a same-class true event. Macro averages over classes with ≥ 10 held-out events:
| classes (≥10 held-out events each) | classes | events | recall (sensitive) | recall (balanced) | precision (sensitive) | precision (balanced) | F1 (sensitive) | F1 (balanced) |
|---|---|---|---|---|---|---|---|---|
| speaker sounds | 16 | 3,271 | 0.46 | 0.33 | 0.38 | 0.40 | 0.38 | 0.34 |
| background sounds | 32 | 6,096 | 0.51 | 0.39 | 0.26 | 0.34 | 0.29 | 0.33 |
| all | 48 | 9,367 | 0.49 | 0.37 | 0.30 | 0.36 | 0.32 | 0.33 |
Thresholds were tuned on held-out data that overlaps the evaluation clips, so absolute numbers are somewhat optimistic. No standard benchmark (e.g. DCASE) has been run.
Training data
Real annotated recordings, event crops pasted into speech at known times (synthetic mixtures), and augmentation with real background noise.
| dataset | role | license (as stated by the source) |
|---|---|---|
| NonVerbalSpeech-38K | Mandarin/English recordings with one timestamped vocal event each | CC BY-NC 4.0, non-commercial research |
| Vaani Noise Event Timestamps | Indic speech recorded on phones (58 Indian languages); 319 noise tags with timestamps (blog post) | CC BY 4.0 |
| NonverbalTTS | English (VoxCeleb, Expresso); sound positions in the transcript, timestamps obtained by forced alignment | annotations CC BY-NC-SA 4.0; audio under the source corpora's terms |
| AudioSet-Strong | YouTube clips with temporally strong labels | not stated on this mirror |
| OpenSLR 99 (Deeply Nonverbal Vocalization) | isolated vocalizations recorded on phones | CC BY-NC-ND 4.0 |
| DNC | home background recordings (TV, appliances, cars, quiet rooms) | MIT |
| DEMAND | multi-channel environmental noise | CC BY 4.0 (the record's text also states CC BY-SA 3.0) |
Training examples (171,000 windows of 8 s, including augmented copies) were drawn from about 168 h of real annotated training recordings, plus event crops, speech beds and background noise; held-out recordings (about 17 h) were used only for model selection, thresholds and evaluation. Training ran for 10 epochs on one NVIDIA T4 (about 18 minutes per epoch), warm-started from a preliminary 31-class checkpoint of the same architecture. The checkpoint with the best mean macro average precision on the real held-out sets is kept (epoch 6).
Limitations and intended use
- Intended for research and non-commercial use: rich transcription, corpus analysis, data filtering, and annotation assistance with human review.
- Domains: training audio is mostly Mandarin drama, Indic phone speech, English interviews and acted speech, and YouTube. Accuracy on other domains (e.g. call centres, meetings, far-field microphones) is unmeasured. Check it on your own labelled sample.
- Weak classes (
tv_speaker,fan,footsteps,cry,groan,chew,door,yawn,hiccup,nose_blow,knock) fire rarely and unreliably. - Steady hum from some microphones can trigger
machineorfan. Use--hide machine,fanif needed. - Timing: tag positions in transcripts depend on the ASR's word timestamps.
- Not suitable for health or diagnostic use (e.g. inferring illness from coughs), surveillance, or any decision about a person.
License
The weights are released under CC BY-NC-SA 4.0 to respect the non-commercial and share-alike terms of the training data above. One source (OpenSLR 99) is CC BY-NC-ND 4.0. Users are responsible for complying with the terms of the training data in their jurisdiction. Commercial use is not permitted.
Citation
@misc{surround_sound_2026,
title = {Surround Sound: speech event detector},
author = {Bar, Ayush Kumar},
year = {2026},
note = {Model on the Hugging Face Hub},
url = {https://huggingface.co/theboringai-work/surround-sound}
}
Please also cite the training datasets (see their pages).
- Downloads last month
- 38