Surround Sound: speech event detector

Surround Sound - Speech Event Detection

Surround Sound - Speech Event Detection

Detect non-verbal vocalizations and environmental sounds with timestamped predictions for richer speech transcription.

🤗 Try the Live Demo

Demo: Surround Sound Space: a 59-minute documentary with the detected sounds, per-tag probabilities and an aligned transcript, synced to the video.

A small detector (5.8M parameters) that finds non-verbal vocal sounds (breath, cough, laugh, sigh, sniff, throat clearing, …) and background sounds (music, horn, vehicle, dog bark, siren, rain, applause, …) in recordings, with start and end times. It covers 54 classes. The included code also merges detected events into a Whisper transcript. An illustrative example:

so I was saying [cough] that the <horn> road was blocked </horn> [breath] anyway

[tag] marks a sound made by the speaker, placed where it starts. <tag> … </tag> wraps the words a background sound overlaps.

  • Input: audio of any length (any format ffmpeg reads), resampled to 16 kHz mono.
  • Output: a list of {label, start, end, prob} events, one probability per class every 40 ms, or a transcript with tags.
  • Speed: about 39× real time on 8 CPU threads and 380× on an NVIDIA T4 (3 minutes of audio).

Quick start

pip install torch torchaudio numpy scipy safetensors huggingface_hub    # plus the ffmpeg binary
pip install openai-whisper                                               # optional, for transcripts
hf download theboringai-work/surround-sound --local-dir surround-sound
cd surround-sound
python inference.py my_audio.mp3                    # transcript with tags + event list
python inference.py my_audio.mp3 --asr none         # events only
python inference.py recordings/ --asr none --json events.jsonl

From Python, run from the downloaded folder:

from sed import Detector, rich_transcript, clean

det = Detector.from_pretrained("theboringai-work/surround-sound", profile="sensitive")   # or a local folder path
events = det.detect("my_audio.mp3")
# [{'label': 'cough', 'start': 0.68, 'end': 2.0, 'prob': 0.92}, {'label': 'breath', 'start': 7.72, 'end': 8.72, 'prob': 0.94}, ...]

times, probs = det.frame_probs(audio_16k)          # raw per-class probabilities on the 40 ms grid

# with any ASR that gives word timestamps: words = [(word, start_s, end_s), ...]
print(rich_transcript(words, clean(words, events)))

Useful options of inference.py:

option meaning
--profile sensitive|balanced threshold set (below)
--hide breath,lip_smack leave tags out
--thr cough=0.5,machine=1.01 change thresholds for this run (above 1 switches a class off)
--asr none|openai-whisper|transformers transcription backend
--language, --asr-model Whisper language and size
--times include times inside the tags
--json out.jsonl machine-readable output

Files

file content
model.safetensors, config.json weights, class list and settings
thresholds_sensitive.json per-class thresholds tuned for F2 (default)
thresholds.json per-class thresholds tuned for F1
detect_config.json post-processing setting (hysteresis ratio)
sed/, inference.py inference code (PyTorch)

Model

audio 16 kHz ─► log-mel (64 bands, 25 ms window, 10 ms hop)
            ─► 5 conv blocks (3×3, BatchNorm, ReLU; frequency /16, time /4)
            ─► linear projection + sinusoidal positions
            ─► Transformer encoder (6 layers, d=256, 8 heads, pre-norm, GELU)
            ─► linear + sigmoid: 54 probabilities every 40 ms

The model sees 8 s windows. Long audio is processed with a 2 s hop, and overlapping predictions are averaged. Each class's probability track is then median-filtered over 200 ms and thresholded with hysteresis (an event starts above the class threshold and continues while above 0.6 × threshold). Events closer than 0.3 s (speaker sounds) or 1.0 s (background) are joined, and events shorter than 0.15 s are dropped.

Threshold profiles

Every class has its own threshold, tuned on real held-out recordings by 2-fold cross-validation.

profile tuned for 43 reliably measured classes (frame level, cross-validated)
sensitive (default) F2: a missed sound costs twice a false tag recall 0.52, precision 0.27
balanced F1 recall 0.41, precision 0.35

Classes and accuracy

frame AP: frame-level average precision on real held-out windows. CV F1: cross-validated frame F1 at the balanced threshold. found: share of held-out real events detected (overlapping prediction of the same class) with the sensitive profile. weak: too little held-out data (< 1 s of positive frames) or CV F1 < 0.1; treat these tags as hints.

Speaker sounds, [tag]

tag frame AP CV F1 (balanced) held-out events found (sensitive) status
sigh 0.72 0.65 84% of 389 reliable
snore 0.55 0.61 68% of 28 reliable
laugh 0.52 0.47 79% of 587 reliable
breath 0.50 0.50 80% of 1,116 reliable
sniff 0.44 0.42 65% of 272 reliable
sneeze 0.43 0.21 0% of 8 reliable
scream 0.42 0.13 16% of 32 reliable
throat_clear 0.34 0.38 62% of 190 reliable
cough 0.31 0.40 56% of 329 reliable
whisper 0.30 0.37 36% of 36 reliable
burp 0.22 0.28 57% of 7 reliable
shout 0.15 0.26 45% of 97 reliable
cry 0.07 0.05 43% of 61 weak
gasp 0.07 0.13 37% of 41 reliable
lip_smack 0.05 0.13 22% of 9 reliable
chew 0.04 0.06 10% of 29 weak
groan 0.01 0.03 19% of 36 weak
yawn 0.01 0.01 7% of 14 weak
hiccup 0.00 0.01 21% of 14 weak
nose_blow 0.00 0.00 0% of 1 weak

Background sounds, <tag> … </tag>

tag frame AP CV F1 (balanced) held-out events found (sensitive) status
music 0.88 0.81 84% of 1,361 reliable
horn 0.59 0.56 76% of 258 reliable
applause 0.58 0.48 34% of 199 reliable
whistle 0.52 0.49 28% of 83 reliable
dog_bark 0.51 0.47 66% of 128 reliable
vehicle 0.50 0.49 87% of 549 reliable
siren 0.49 0.47 58% of 114 reliable
bird 0.40 0.43 64% of 791 reliable
water 0.39 0.38 75% of 192 reliable
insect 0.38 0.43 42% of 87 reliable
cheer 0.32 0.33 26% of 73 reliable
crowd 0.32 0.24 61% of 41 reliable
wind 0.32 0.30 68% of 206 reliable
gunshot 0.31 0.32 81% of 75 reliable
bell 0.26 0.35 80% of 87 reliable
explosion 0.22 0.46 63% of 54 reliable
baby_cry 0.22 0.40 39% of 31 reliable
construction 0.21 0.33 67% of 220 reliable
glass_break 0.20 0.21 12% of 58 reliable
dishes 0.20 0.20 42% of 50 reliable
phone_vibrate 0.17 0.40 25% of 4 reliable
rain 0.16 0.28 12% of 43 reliable
animal 0.15 0.14 52% of 538 reliable
child_voice 0.15 0.19 48% of 179 reliable
machine 0.14 0.28 51% of 137 reliable
thunder 0.10 0.19 54% of 13 reliable
buzzer 0.09 0.12 20% of 25 reliable
beep 0.09 0.18 56% of 165 reliable
phone_ring 0.09 0.20 70% of 64 reliable
door 0.08 0.03 5% of 21 weak
tv_speaker 0.04 0.07 44% of 23 weak
footsteps 0.04 0.08 50% of 218 weak
fan 0.04 0.04 15% of 13 weak
knock 0.00 0.00 0% of 5 weak

Evaluation

All evaluation uses real recordings held out from training:

Each class is scored only on the datasets that label it. Whole clips are run through the full detector; a true event is found if a same-class prediction overlaps it, and a prediction is correct if it overlaps a same-class true event. Macro averages over classes with ≥ 10 held-out events:

classes (≥10 held-out events each) classes events recall (sensitive) recall (balanced) precision (sensitive) precision (balanced) F1 (sensitive) F1 (balanced)
speaker sounds 16 3,271 0.46 0.33 0.38 0.40 0.38 0.34
background sounds 32 6,096 0.51 0.39 0.26 0.34 0.29 0.33
all 48 9,367 0.49 0.37 0.30 0.36 0.32 0.33

Thresholds were tuned on held-out data that overlaps the evaluation clips, so absolute numbers are somewhat optimistic. No standard benchmark (e.g. DCASE) has been run.

Training data

Real annotated recordings, event crops pasted into speech at known times (synthetic mixtures), and augmentation with real background noise.

dataset role license (as stated by the source)
NonVerbalSpeech-38K Mandarin/English recordings with one timestamped vocal event each CC BY-NC 4.0, non-commercial research
Vaani Noise Event Timestamps Indic speech recorded on phones (58 Indian languages); 319 noise tags with timestamps (blog post) CC BY 4.0
NonverbalTTS English (VoxCeleb, Expresso); sound positions in the transcript, timestamps obtained by forced alignment annotations CC BY-NC-SA 4.0; audio under the source corpora's terms
AudioSet-Strong YouTube clips with temporally strong labels not stated on this mirror
OpenSLR 99 (Deeply Nonverbal Vocalization) isolated vocalizations recorded on phones CC BY-NC-ND 4.0
DNC home background recordings (TV, appliances, cars, quiet rooms) MIT
DEMAND multi-channel environmental noise CC BY 4.0 (the record's text also states CC BY-SA 3.0)

Training examples (171,000 windows of 8 s, including augmented copies) were drawn from about 168 h of real annotated training recordings, plus event crops, speech beds and background noise; held-out recordings (about 17 h) were used only for model selection, thresholds and evaluation. Training ran for 10 epochs on one NVIDIA T4 (about 18 minutes per epoch), warm-started from a preliminary 31-class checkpoint of the same architecture. The checkpoint with the best mean macro average precision on the real held-out sets is kept (epoch 6).

Limitations and intended use

  • Intended for research and non-commercial use: rich transcription, corpus analysis, data filtering, and annotation assistance with human review.
  • Domains: training audio is mostly Mandarin drama, Indic phone speech, English interviews and acted speech, and YouTube. Accuracy on other domains (e.g. call centres, meetings, far-field microphones) is unmeasured. Check it on your own labelled sample.
  • Weak classes (tv_speaker, fan, footsteps, cry, groan, chew, door, yawn, hiccup, nose_blow, knock) fire rarely and unreliably.
  • Steady hum from some microphones can trigger machine or fan. Use --hide machine,fan if needed.
  • Timing: tag positions in transcripts depend on the ASR's word timestamps.
  • Not suitable for health or diagnostic use (e.g. inferring illness from coughs), surveillance, or any decision about a person.

License

The weights are released under CC BY-NC-SA 4.0 to respect the non-commercial and share-alike terms of the training data above. One source (OpenSLR 99) is CC BY-NC-ND 4.0. Users are responsible for complying with the terms of the training data in their jurisdiction. Commercial use is not permitted.

Citation

@misc{surround_sound_2026,
  title  = {Surround Sound: speech event detector},
  author = {Bar, Ayush Kumar},
  year   = {2026},
  note   = {Model on the Hugging Face Hub},
  url    = {https://huggingface.co/theboringai-work/surround-sound}
}

Please also cite the training datasets (see their pages).

Downloads last month
38
Safetensors
Model size
6.83M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train theboringai-work/surround-sound

Space using theboringai-work/surround-sound 1