Zencastr Laughter Detection

Detect laughter events in close-microphone conversational speech. Built at Zencastr for podcast audio and open-sourced as a tool for creators.

Usage

Whole audio in, events out. Input can use any format ffmpeg can decode (path, URL, bytes, array) at any length or sample rate.

from transformers import pipeline

pipe = pipeline(
    "laughter-detection",
    model="zencastr/laughter-detection",
    trust_remote_code=True,
    device=0,
)

pipe("episode.mp3")
# {
#   "duration_s": 3591.4,
#   "events": [
#     {
#       "start": 812.16,
#       "end": 814.4,
#       "score": 0.997,
#     },
#     ...
#   ],
#   "operating_point": {
#     "name": "high_precision",
#     "threshold": 0.99,
#     "min_duration_s": 0.2,
#     "merge_gap_s": 1.0
#   },
# }

From the model repository, run the command-line scripts with uv:

uv run events.py episode.mp3 --operating-point max_f1 --pretty  # PyTorch
uv run events_onnx.py episode.mp3                               # ONNX

For clip-level classification of audio up to 4 s:

clf = pipeline(
    "audio-classification",
    model="zencastr/laughter-detection",
    trust_remote_code=True,
)
clf("clip.wav")
# [
#   {"label": "laughter", "score": 0.93},
#   {"label": "no_laughter", "score": 0.07}
# ]

Inference

Inference architecture

Audio x is converted to a 128-bin log-mel spectrogram X at 16 kHz using 25 ms Hann windows and a 10 ms hop. The encoder f and laughter heads g then produce frame and clip probabilities. φ denotes the LoRA adapters. The model produces frame probabilities every 160 ms from 4 s windows evaluated with a 2 s hop. Scores from overlapping windows are combined by taking their maximum, then converted into events with approximate timestamps using a threshold, minimum duration, and merge gap. View the frame probabilities and tune the event parameters as follows:

pipe("episode.mp3", threshold=0.9, return_frame_probs=True)

The audio-classification pipeline instead returns one attention-pooled probability for the entire clip.

Operating points

Operating points are named presets for decoding frame probabilities into events.

pipe("episode.mp3", operating_point="max_f1")

The presets were selected on ~17 h of internal podcast audio from 21 recordings using a 0.2 s matching tolerance. high_precision maximizes recall at precision ≥ 0.88 and is recommended when false alarms are costly; high_recall reaches recall 0.88 and is recommended when missed laughter is costly; max_f1 maximizes F1 and provides a balance between precision and recall.

Preset Threshold Min duration Merge gap Precision Recall F1 Onset error
high_precision (default) 0.99 0.2 s 1.0 s 0.88 0.70 0.78 1.6 s
max_f1 0.97 0.16 s 1.0 s 0.80 0.77 0.78 1.6 s
high_recall 0.69 0.2 s 0.16 s 0.59 0.88 0.70 1.4 s

Evaluation

Our model achieves the highest mean clip-level average precision (AP) on three conversational benchmarks while maintaining a reasonable false positive rate (FPR) on other non-speech vocalizations.

Model Podcast holdout AP ↑ PodcastFillers AP ↑ AMI AP ↑ VocalSound AP ↑ FSD50K FPR ↓
Ours 0.973 ± 0.002 0.833 ± 0.011 0.948 ± 0.003 0.955 ± 0.005 15 ± 1%
EAT 0.948 ± 0.018 0.784 ± 0.049 0.923 ± 0.024 0.957 ± 0.013 8 ± 3%
SenseVoice-S 0.967 ± 0.012 0.771 ± 0.044 0.898 ± 0.030 0.935 ± 0.017 17 ± 4%
CED-base 0.950 ± 0.019 0.811 ± 0.060 0.854 ± 0.039 0.970 ± 0.010 4 ± 2%
CED-tiny 0.933 ± 0.024 0.798 ± 0.066 0.856 ± 0.032 0.956 ± 0.012 7 ± 2%
BEATs 0.908 ± 0.023 0.748 ± 0.064 0.842 ± 0.030 0.927 ± 0.017 15 ± 4%
Gillick et al. 0.951 ± 0.019 0.819 ± 0.042 0.853 ± 0.031 0.717 ± 0.031 48 ± 5%

Our row reports mean ± SD across multiple seeds while baselines rows report the performance of a single checkpoint ± the 95% bootstrap half-width. FSD50K FPR measures false positives on 1,136 clips of non-laughter vocalizations (coughing, sneezing, breathing, gasping, sighing, crying, and burping) at model-specific thresholds yielding 90% recall. CED, EAT, and BEATs use their AudioSet heads; SenseVoice uses its utterance-level laughter tag.

Inference speed

Real-time factors measured on an NVIDIA L4 using the same 1,800 s of audio, batch 64, and the median of nine passes. All models use fp16 and PyTorch except ours, which uses ONNX opset 23 with fused attention. Timings include windowing and host-to-device transfer, and exclude model loading, audio decoding, and event decoding.

Model Parameters Window / hop × realtime ↑
CED-tiny 5.5 M 4 s / 2 s 4,540×
CED-base 85.7 M 4 s / 2 s 2,145×
Ours (ONNX) 90.7 M 4 s / 2 s 1,355×
EAT (PyTorch) 90.4 M 4 s / 2 s 1,032×
BEATs 90.7 M 2 s / 1 s 410×
SenseVoice-Small 234.0 M 4 s / 2 s 328×
Gillick et al. 0.8 M 1 s / 23 ms 46×

Files

File Purpose
model.safetensors PyTorch weights with LoRA merged and the AudioSet teacher head removed. Loads through AutoModelForAudioClassification / AutoModelForAudioFrameClassification.
model.onnx Float32, opset 17. Takes 16 kHz waveforms, including the mel front end in the graph. Runs with onnxruntime + numpy on CPU or GPU.
model_fp16.onnx Float16 encoder, opset 17, for GPU inference.
model_fp16_opset23.onnx Float16 with native ONNX Attention; onnxruntime ≥ 1.22, GPU. Used for the throughput result above.
export.json ONNX input/output contracts and parity checks against PyTorch.
laughter_core.py Shared numpy windowing, stitching, and event decoding.

Provenance and licenses

Model and code: MIT (LICENSE). Base encoder: worstchan/EAT-base_epoch30_finetune_AS2M, revision 60d61e8b, weights SHA-256 5fa59385…69fff1b; MIT, © 2024 Wenxi Chen (LICENSE-EAT). The base encoder was pretrained and fine-tuned on AudioSet.

Adaptation used Zencastr platform audio collected under terms permitting model training. Public datasets were used only for evaluation and monitoring during adaptation. Dataset licenses: AudioSet annotations (CC BY 4.0), AMI (CC BY 4.0), VocalSound (CC BY-SA 4.0), PodcastFillers (CC BY-NC), and FSD50K (per-clip CC licenses).

Citation

@software{zencastr2026laughter,
  title  = {Zencastr Laughter Detection},
  author = {{Zencastr}},
  year   = {2026},
  url    = {https://huggingface.co/zencastr/laughter-detection}
}

@article{chen2024eat,
  title   = {EAT: Self-Supervised Pre-Training with Efficient Audio Transformer},
  author  = {Chen, Wenxi and Liang, Yuzhe and Ma, Ziyang and Zheng, Zhisheng and Chen, Xie},
  journal = {arXiv preprint arXiv:2401.03497},
  year    = {2024}
}
Downloads last month
38
Safetensors
Model size
90M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zencastr/laughter-detection

Paper for zencastr/laughter-detection