Vocal Burst Detector β€” commercial (small)

A 17-class vocal-burst classifier head on top of laion/voiceclap-commercial, the 110 M dual-tower VoiceCLAP encoder trained only on commercially usable data. The whole chain is CC BY 4.0, which is the reason this model exists: the stronger laion/vocal-burst-detector-x2 rests on a 7 B encoder and is a much larger dependency.

What it does

Given a ~16 kHz cut containing a vocal burst, it names the burst as one of 16 burst classes plus no_burst. The 17 are not a curated shortlist β€” they are exactly the labels with at least 100 annotated segments in both halves of a two-source corpus (real recordings and synthetic voice-acting audio). Chance is 5.9 %.

Affirmative Grunt, Breathy Giggle, Chuckle, Deep Breath, Exasperated Sigh, Exhausted Groan, Frustrated Groan, Heavy Breathing, Humming, Panting, Relief Sigh, Scream, Sharp Inhale, Soft Hum, Wistful Sigh, Yawn, no_burst

How it compares

Same data, same grouped split, same head, same five initialisations β€” only the encoder differs. Held-out, balanced at 25 clips per class:

encoder licence params dim real, 17-way DramaBox, 17-way real, 23 groups DramaBox, 23 groups
voiceclap-large-v2 CC BY 4.0 7 B + LoRA 3584 0.466 0.574 0.605 0.697
voiceclap-commercial (this model) CC BY 4.0 110 M 768 0.393 0.536 0.515 0.628
voiceclap-small-v2 CC BY-NC 4.0 110 M 768 0.402 0.551 0.513 0.675

The 7 B encoder is 0.073 better on the real half at 17-way β€” for roughly 64Γ— the parameters. On the synthetic half the gap is smaller still.

The non-commercial small encoder buys nothing. voiceclap-small-v2 (CC BY-NC 4.0) scores 0.402 against this model's 0.393 at 17-way and 0.513 against 0.515 on groups β€” a difference smaller than the spread across split seeds. If you need a small encoder, there is no accuracy argument for taking the non-commercial licence.

Class groups

vocal_burst_groups.json and GROUPS.md carry a 23-group scheme over the burst label space, grouping names that denote the same or a very similar sound. Scoring the same predictions at group level lifts this model from 0.393 to 0.515 on the real half β€” but a random grouping of identical group sizes already reaches 0.425, so +0.090 of that is the grouping being right and the rest is arithmetic. Read GROUPS.md before using the group column for anything.

Note that chance rises with grouping too: 17-way chance is 5.9 %, the 23-group scheme on this test set is 0.103. Ratios to chance, not raw accuracies, are what compare across schemes.

Per-class recall

per_class_recall.json / .csv give recall per class and per source, strict and family-relaxed, with Wilson intervals and a reliable flag at n β‰₯ 30. This matters for anyone scoring generated audio: a class this model recalls at 20 % cannot show a measured hit rate meaningfully above 20 %, however good the generator is. Treat a low row as an instrument reading first and a generator verdict second.

Files

file what
vocal_burst_mlp_prod_s{0..4}.pt the shipped 5-member ensemble (one grouped split, five initialisations β€” never an ensemble over splits, which would be scored on items its own members trained on)
prod_report.json config, split, confusion matrices per source
per_class_recall.json / .csv per class, per source, strict + family
train_report_*.json the cross-source 2Γ—2 arm
bs_vcsmall.py the encoder wrapper, tested on real shard audio
vocal_burst_groups.json, GROUPS.md the 23-group scheme

Usage

import numpy as np, soundfile as sf, torch
from bs_vcsmall import VoiceCLAPSmall      # this repo

enc = VoiceCLAPSmall('laion/voiceclap-commercial')       # or a local snapshot
x, sr = sf.read('burst.wav', dtype='float32')
if x.ndim > 1: x = x.mean(1)
# resample to 16 kHz yourself; the encoder rejects anything else
emb = enc.encode([x])                      # (1, 768), L2-normalised

import json
classes = json.load(open('prod_report.json'))['classes']
logits = sum(torch.load(f'vocal_burst_mlp_prod_s{i}.pt')(torch.tensor(emb))
             for i in range(5)) / 5
print(classes[int(logits.argmax())])

Limitations, stated rather than discovered

  • No drop-in scorer. laion/vocal-burst-detector-x2 ships a ProductionBurstScorer that wires a burst locator, encoder and head together. That wrapper targets the 3584-d encoder and has not been ported to this one. An untested port would load, run and silently produce plausible-looking numbers, which is worse than shipping none.
  • The encoder pads to 30 s and mean-pools. A 0.4 s burst therefore occupies about 1.3 % of the frames being averaged. Measured on real segments the embeddings stay clearly distinct (pairwise cosine 0.29–0.95), and the per-group recall does not fall the way that story predicts β€” scream, short and loud, is the one group this model recalls better than the 7 B encoder. So the dilution is real but is not established as the cause of the gap.
  • Nine of 17 classes carry reliable: false β€” the held-out set is too small for a confident per-class number there.
  • No human has listened. Every figure rests on annotator labels.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for laion/vocal-burst-detector-commercial

Finetuned
(2)
this model