Instructions to use zencastr/laughter-detection with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zencastr/laughter-detection with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="zencastr/laughter-detection", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForAudioFrameClassification model = AutoModelForAudioFrameClassification.from_pretrained("zencastr/laughter-detection", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Zencastr Laughter Detection
Detect laughter events in close-microphone conversational speech. Built at Zencastr for podcast audio and open-sourced as a tool for creators.
Usage
Whole audio in, events out. Input can use any format ffmpeg can decode (path, URL, bytes, array) at any length or sample rate.
from transformers import pipeline
pipe = pipeline(
"laughter-detection",
model="zencastr/laughter-detection",
trust_remote_code=True,
device=0,
)
pipe("episode.mp3")
# {
# "duration_s": 3591.4,
# "events": [
# {
# "start": 812.16,
# "end": 814.4,
# "score": 0.997,
# },
# ...
# ],
# "operating_point": {
# "name": "high_precision",
# "threshold": 0.99,
# "min_duration_s": 0.2,
# "merge_gap_s": 1.0
# },
# }
From the model repository, run the command-line scripts with uv:
uv run events.py episode.mp3 --operating-point max_f1 --pretty # PyTorch
uv run events_onnx.py episode.mp3 # ONNX
For clip-level classification of audio up to 4 s:
clf = pipeline(
"audio-classification",
model="zencastr/laughter-detection",
trust_remote_code=True,
)
clf("clip.wav")
# [
# {"label": "laughter", "score": 0.93},
# {"label": "no_laughter", "score": 0.07}
# ]
Inference
Audio x is converted to a 128-bin log-mel spectrogram X at 16 kHz using 25 ms Hann windows and a 10 ms hop. The encoder f and laughter heads g then produce frame and clip probabilities. φ denotes the LoRA adapters. The model produces frame probabilities every 160 ms from 4 s windows evaluated with a 2 s hop. Scores from overlapping windows are combined by taking their maximum, then converted into events with approximate timestamps using a threshold, minimum duration, and merge gap. View the frame probabilities and tune the event parameters as follows:
pipe("episode.mp3", threshold=0.9, return_frame_probs=True)
The audio-classification pipeline instead returns one attention-pooled probability for the entire clip.
Operating points
Operating points are named presets for decoding frame probabilities into events.
pipe("episode.mp3", operating_point="max_f1")
The presets were selected on ~17 h of internal podcast audio from 21 recordings using a 0.2 s matching tolerance. high_precision maximizes recall at precision ≥ 0.88 and is recommended when false alarms are costly; high_recall reaches recall 0.88 and is recommended when missed laughter is costly; max_f1 maximizes F1 and provides a balance between precision and recall.
| Preset | Threshold | Min duration | Merge gap | Precision | Recall | F1 | Onset error |
|---|---|---|---|---|---|---|---|
high_precision (default) |
0.99 | 0.2 s | 1.0 s | 0.88 | 0.70 | 0.78 | 1.6 s |
max_f1 |
0.97 | 0.16 s | 1.0 s | 0.80 | 0.77 | 0.78 | 1.6 s |
high_recall |
0.69 | 0.2 s | 0.16 s | 0.59 | 0.88 | 0.70 | 1.4 s |
Evaluation
Our model achieves the highest mean clip-level average precision (AP) on three conversational benchmarks while maintaining a reasonable false positive rate (FPR) on other non-speech vocalizations.
| Model | Podcast holdout AP ↑ | PodcastFillers AP ↑ | AMI AP ↑ | VocalSound AP ↑ | FSD50K FPR ↓ |
|---|---|---|---|---|---|
| Ours | 0.973 ± 0.002 | 0.833 ± 0.011 | 0.948 ± 0.003 | 0.955 ± 0.005 | 15 ± 1% |
| EAT | 0.948 ± 0.018 | 0.784 ± 0.049 | 0.923 ± 0.024 | 0.957 ± 0.013 | 8 ± 3% |
| SenseVoice-S | 0.967 ± 0.012 | 0.771 ± 0.044 | 0.898 ± 0.030 | 0.935 ± 0.017 | 17 ± 4% |
| CED-base | 0.950 ± 0.019 | 0.811 ± 0.060 | 0.854 ± 0.039 | 0.970 ± 0.010 | 4 ± 2% |
| CED-tiny | 0.933 ± 0.024 | 0.798 ± 0.066 | 0.856 ± 0.032 | 0.956 ± 0.012 | 7 ± 2% |
| BEATs | 0.908 ± 0.023 | 0.748 ± 0.064 | 0.842 ± 0.030 | 0.927 ± 0.017 | 15 ± 4% |
| Gillick et al. | 0.951 ± 0.019 | 0.819 ± 0.042 | 0.853 ± 0.031 | 0.717 ± 0.031 | 48 ± 5% |
Our row reports mean ± SD across multiple seeds while baselines rows report the performance of a single checkpoint ± the 95% bootstrap half-width. FSD50K FPR measures false positives on 1,136 clips of non-laughter vocalizations (coughing, sneezing, breathing, gasping, sighing, crying, and burping) at model-specific thresholds yielding 90% recall. CED, EAT, and BEATs use their AudioSet heads; SenseVoice uses its utterance-level laughter tag.
Inference speed
Real-time factors measured on an NVIDIA L4 using the same 1,800 s of audio, batch 64, and the median of nine passes. All models use fp16 and PyTorch except ours, which uses ONNX opset 23 with fused attention. Timings include windowing and host-to-device transfer, and exclude model loading, audio decoding, and event decoding.
| Model | Parameters | Window / hop | × realtime ↑ |
|---|---|---|---|
| CED-tiny | 5.5 M | 4 s / 2 s | 4,540× |
| CED-base | 85.7 M | 4 s / 2 s | 2,145× |
| Ours (ONNX) | 90.7 M | 4 s / 2 s | 1,355× |
| EAT (PyTorch) | 90.4 M | 4 s / 2 s | 1,032× |
| BEATs | 90.7 M | 2 s / 1 s | 410× |
| SenseVoice-Small | 234.0 M | 4 s / 2 s | 328× |
| Gillick et al. | 0.8 M | 1 s / 23 ms | 46× |
Files
| File | Purpose |
|---|---|
model.safetensors |
PyTorch weights with LoRA merged and the AudioSet teacher head removed. Loads through AutoModelForAudioClassification / AutoModelForAudioFrameClassification. |
model.onnx |
Float32, opset 17. Takes 16 kHz waveforms, including the mel front end in the graph. Runs with onnxruntime + numpy on CPU or GPU. |
model_fp16.onnx |
Float16 encoder, opset 17, for GPU inference. |
model_fp16_opset23.onnx |
Float16 with native ONNX Attention; onnxruntime ≥ 1.22, GPU. Used for the throughput result above. |
export.json |
ONNX input/output contracts and parity checks against PyTorch. |
laughter_core.py |
Shared numpy windowing, stitching, and event decoding. |
Provenance and licenses
Model and code: MIT (LICENSE). Base encoder: worstchan/EAT-base_epoch30_finetune_AS2M, revision 60d61e8b, weights SHA-256 5fa59385…69fff1b; MIT, © 2024 Wenxi Chen (LICENSE-EAT). The base encoder was pretrained and fine-tuned on AudioSet.
Adaptation used Zencastr platform audio collected under terms permitting model training. Public datasets were used only for evaluation and monitoring during adaptation. Dataset licenses: AudioSet annotations (CC BY 4.0), AMI (CC BY 4.0), VocalSound (CC BY-SA 4.0), PodcastFillers (CC BY-NC), and FSD50K (per-clip CC licenses).
Citation
@software{zencastr2026laughter,
title = {Zencastr Laughter Detection},
author = {{Zencastr}},
year = {2026},
url = {https://huggingface.co/zencastr/laughter-detection}
}
@article{chen2024eat,
title = {EAT: Self-Supervised Pre-Training with Efficient Audio Transformer},
author = {Chen, Wenxi and Liang, Yuzhe and Ma, Ziyang and Zheng, Zhisheng and Chen, Xie},
journal = {arXiv preprint arXiv:2401.03497},
year = {2024}
}
- Downloads last month
- 38
Model tree for zencastr/laughter-detection
Base model
worstchan/EAT-base_epoch30_pretrain