TTS contamination detector
Your TTS is learning from other TTS. Web-scraped speech corpora are built from YouTube and similar sites, and part of that audio is itself machine-generated: TTS-narrated explainers, list videos, article read-outs, product promos. A TTS model trained on it learns another system's voice and artefacts. This model is a filter for that: it scores clips for synthetic speech so that whole videos or voices can be dropped from a training set.
On Emilia-YODAS English, it flags the videos behind 20.05% of the 21,206 hours we scored (see The Emilia-YODAS finding). In a blind listening check, every flagged clip the listener could judge was TTS (25 of 25; 95% interval 86.7–100%): about 1 in 5 of the Emilia-YODAS English hours we scored is synthetic speech.
- What it is: frozen HuBERT-Large features, mean- and std-pooled over 8 hidden layers (16,384 numbers per clip), and one linear logistic model. 197 kB of head weights.
- What it is for: filtering training data. Decisions are made per video or per voice, from many clips, not per clip.
- What it is not: a deepfake, fraud or forensics detector. It was not built or tested for single-clip verdicts, adversarial audio, or deciding whether a particular person's recording is genuine. Do not use it for that.
Quick start
pip install torch transformers numpy scipy soundfile safetensors huggingface_hub
hf download cloud0day3/tts-contamination-detector --local-dir tts-contamination-detector
# Clips of one video (or one voice): per-clip p and the group decision.
python tts-contamination-detector/detector.py video_clip01.flac video_clip02.flac video_clip03.flac
Output on five clips of one flagged Emilia-YODAS video (file names shortened):
p 0.988 11.75 s clip1.flac
p 0.999 4.86 s clip2.flac
p 0.994 12.74 s clip3.flac
p 1.000 9.17 s clip4.flac
p 0.997 11.13 s clip5.flac
group: mean p 0.995 over 5 clips (threshold 0.32): SYNTHETIC
From Python:
import sys
sys.path.insert(0, "tts-contamination-detector")
from detector import TTSDetector, video_decision
det = TTSDetector.from_pretrained("tts-contamination-detector") # a local folder or "cloud0day3/tts-contamination-detector"
logits = det.score_files(paths) # one logit per clip (NaN if a clip cannot be read or is < 0.25 s)
p = det.probability(logits)
decision = video_decision(durations, logits) # threshold=0.32, min_clips=3
print(decision.synthetic, decision.prob, decision.clips)
detector.py is a standalone port of the pipeline that produced the scores. It depends on numpy, scipy (the
resampler: scipy.signal.resample_poly with the pipeline's Kaiser FIR, so features match exactly; torchaudio or soxr
would not), soundfile, torch and transformers. It does not need the Spero code base.
- Front end: mono (channel mean); leading and trailing 20 ms frames quieter than the loudest frame minus 40 dB are trimmed, keeping 0.1 s before and 0.2 s after the speech; the middle 10 s are kept; then 16 kHz.
- Encoder:
facebook/hubert-large-ll60kat revisionff022d09, float32 (TF32 on CUDA), per-utterance normalised input, batched with an attention mask. Hidden states 1, 3, 6, 9, 12, 15, 18 and 24; per layer, the frame mean and standard deviation. - Head:
logit = clip(w · x + b, −12, 12)on the pooled features (cast to float16 and back, as the pipeline stored them).config.jsonholds every constant;model.safetensorsholdsweight,bias,meanandscale(the last two are 0 and 1: the head was folded onto raw features). - Video decision: the duration-weighted mean of the clip probabilities (durations floored at 0.1 s) over at least 3 scored clips; flagged at 0.32 or above. Use the clips' own durations, not the 10 s crop.
Equivalence with the pipeline (equivalence.json): on 334 Emilia-YODAS clips (24 whole videos from two shards),
detector.py reproduces the pipeline's own code bit for bit (front end, all pooled features, logits). Against the
logits the pipeline stored during the Emilia-YODAS run, the largest difference is 0.018 logits (mean 0.0011), from
batch composition and GPU type under TF32; every clip and video decision at 0.32 and 0.54 agrees. On a CPU the
logits move by about 0.01.
How it was trained
Only permissive audio, with a lineage row for every clip (254,685 rows). No commercial-provider audio (ElevenLabs, Google, OpenAI or any other API) and no private recordings were used.
| Role | Audio | Clips | Licence |
|---|---|---|---|
| Synthetic | Renders of 8 open TTS systems: Kokoro, Chatterbox, VibeVoice 1.5B, Qwen3-TTS, Maya1, Zonos v0.1, OpenVoice v2, MetaVoice (built-in voices or LibriTTS-R cloning references) | 16,800 | Apache-2.0 or MIT models; references CC BY 4.0 |
| Synthetic | YODAS2 clips of the 12 YouTube voices that read out Wikipedia section markup ("== History ==") | 3,052 | CC BY 3.0 |
| Human | YODAS2 clips from videos that pass both of our earlier heuristics | 10,074 | CC BY 3.0 |
| Human | People's Speech | 6,742 | CC BY 4.0 |
| Human | LibriTTS-R train | 16,362 | CC BY 4.0 |
| Human | VCTK 0.92 (mic1, 110 speakers) | 2,640 | CC BY 4.0 |
| Test only | Magpie 357M and StyleTTS 2 renders (licences not Apache-2.0 or MIT, so never fitted) | 4,800 | |
| Test only | LibriVox read speech (MLS, LibriHeavy, HiFiTTS-2) | 20,370 | CC BY 4.0 or public domain |
- Every training clip is used twice: clean, and through a seeded mix of reverb, noise, band limiting and an MP3 or Opus round trip. Both classes get the same mix, so a codec cannot become a cue for "human".
- L2 logistic regression on standardised features, 4 folds that hold out whole TTS systems, whole Wikipedia voices, groups of videos or documents that share a voice, and LibriVox readers. Five layer and regularisation settings were compared on out-of-fold AUC; all eight layers won.
- The released head is the average of the four fold models of that setting, folded into one weight vector. Those folds held out single videos; a later review re-folded the same plan so that no voice spans two folds, refit, and chose the same setting. The numbers below come from that leak-free refit.
- The "human" YouTube negatives are presumed human: they passed two heuristics, nobody listened to them.
Results (out of fold)
From the training report (reports/data/tts-detector.md in the Spero TTS repository; section numbers in brackets).
Each number comes from a fold model that never saw the clip's TTS system, Wikipedia voice, video group, document
group or reader.
Clips (AUC of each held-out synthetic set against all held-out human clips, §3):
| Held-out set | Clean | Through Opus 64 kbit/s |
|---|---|---|
| Chatterbox, Qwen3-TTS | 0.999 | 0.999 |
| OpenVoice v2 / Maya1 / Zonos v0.1 | 0.998 / 0.997 / 0.997 | 0.998 / 0.997 / 0.997 |
| MetaVoice | 0.992 | 0.993 |
| Kokoro | 0.985 | 0.984 |
| VibeVoice 1.5B | 0.971 | 0.970 |
| Magpie (never trained on) | 0.997 | 0.997 |
| StyleTTS 2 (never trained on) | 0.996 | 0.997 |
| Wikipedia-reader voices on YouTube | 0.997 |
Videos and documents at 0.32 (§3, §4; YODAS2 pilot of 2,838 videos, 428.7 h):
- Wikipedia-voice videos found: 92% (225 of 245). One voice accounts for 140 of them and is always found; without it, 81% (85 of 105). Two of the 12 voices are mostly missed (2 of 13 and 0 of 8 videos).
- Certainly-human sets: 0 of 527 YouTube videos with filled pauses ("um", "uh"), 0 of 5,018 People's Speech documents, no VCTK speaker, LibriTTS-R chapter or MLS book. LibriVox read speech is the hardest case: 3 of 337 HiFiTTS-2 chapters (0.9%) and 4 of 199 LibriHeavy books (2.0%).
- Videos that passed both earlier heuristics: 42 of 2,566 flagged (1.6%); 24 of them share a voice with a Wikipedia-reader or uniform-voice video, and most transcripts read like TTS-narrated genres.
- Counting every flag outside the Wikipedia set as wrong, precision is at least 80.1% (a loose bound).
Held-out TTS systems, render pseudo-videos of 20 prompts flagged at 0.32 (§3):
| System | Flagged | Where it misses |
|---|---|---|
| Chatterbox, Maya1, OpenVoice v2, Qwen3-TTS | 100% | |
| Zonos v0.1 | 99% | one cloned voice at 97% |
| Magpie (never trained on) | 98% | one voice at 90% |
| StyleTTS 2 (never trained on) | 94% | one cloned voice at 77% |
| Kokoro | 75% | the am_michael voice: 0% |
| MetaVoice | 62% | its clones score 13-97% |
| VibeVoice 1.5B | 52% | the Alice voice: 3% |
The released head was trained on all eight training systems, these voices included; the table shows how a system it has never heard can slip through, often one clean voice at a time.
The threshold (§3). 0.32 is the smallest threshold at which no certainly-human set lost more than 1% of its recordings in the first evaluation. With leak-free folds the same rule gives 0.54, set by one LibriHeavy book. 0.32 is kept as the default because the detector is meant for YouTube-derived speech, where the filled-pause videos and People's Speech still lose nothing at 0.32, while at 0.54 held-out systems lose much recall (Magpie 86%, StyleTTS 2 72%, MetaVoice 15%). For data that is mostly audiobook-style read speech, 0.54 is what the 1% rule gives.
| Threshold | Wikipedia videos found | Pass-heuristics videos flagged | Filler videos | People's Speech | HiFiTTS-2 | LibriHeavy | MLS, LibriTTS-R, VCTK |
|---|---|---|---|---|---|---|---|
| 0.10 | 96.7% | 3.5% | 0.4% | 0.1% | 3.0% | 7.0% | 0 |
| 0.20 | 92.7% | 2.3% | 0 | 0.02% | 1.5% | 4.5% | 0 |
| 0.32 | 91.8% | 1.6% | 0 | 0 | 0.9% | 2.0% | 0 |
| 0.50 | 91.0% | 1.2% | 0 | 0 | 0.3% | 1.0% | 0 |
| 0.70 | 89.8% | 1.1% | 0 | 0 | 0 | 0 | 0 |
The Emilia-YODAS finding
Emilia is a widely used open TTS training set; its Emilia-YODAS part is cut from YouTube. We scored 8,091,785 English clips (21,205.8 h, 448,394 videos; at most 40 clips per video in each shard part) and applied the video rule above:
| Threshold | Videos flagged | Clips | Hours | Share of scored hours |
|---|---|---|---|---|
| 0.32 | 62,538 | 1,588,547 | 4,251.2 | 20.05% |
| 0.54 | 61,219 | 1,571,555 | 4,210.8 | 19.86% |
The two thresholds barely differ: only 1,319 of the videos flagged at 0.32 (2.1%) score below 0.54.
A flag is the detector's claim, not ground truth. A blind listening check (one listener; 44 clips: the first 44, in a score-independent page order, of 90 drawn in proportion to hours from flagged and unflagged videos; protocol fixed before listening) measured:
- precision of the flagged hours: 100%, 25 heard as TTS, 0 as human, 1 can't tell (86.7–100%, Wilson 95%);
- synthetic speech among unflagged hours: 0%, 0 TTS, 13 human, 5 can't tell (0–22.8%); a loose bound on what the detector misses;
- corrected share of synthetic hours: 20.05%.
About 1 in 5 hours of Emilia-YODAS English that we scored is synthetic speech. The scorer samples at most 40 clips per video in each shard part; counting every clip of the scored videos (23,263 h), the flagged videos hold 20.31%.
Per-clip and per-video scores are in the companion dataset
cloud0day3/emilia-yodas-tts-contamination,
keyed by Emilia item id and YouTube video id, so you can drop the flagged videos from your own copy of Emilia (or
choose your own threshold). Only the video rule is applied there; the voice-across-videos rule needs a speaker join
and is not reproduced.
Limitations
- Coverage of TTS systems. It learned from 8 open TTS systems and 12 YouTube Wikipedia-reader voices. Newer and
commercial voices were never in training. Held-out systems were usually
caught, but single voices slipped through whole (Kokoro
am_michael, VibeVoice Alice). Expect it to miss some modern TTS; treat the Emilia-YODAS share as an estimate with that bias. - Per video, not per clip. A single clip's probability is noisy. The thresholds and every evaluation number are for videos (or documents, books, voices) of at least 3 clips. Do not use it to label individual clips.
- Read speech. Careful audiobook reading is the human speech it confuses most (LibriHeavy 2.0% of books at 0.32). Calm, scripted human narration on YouTube can be flagged too.
- Not a deepfake detector. No adversarial testing, no voice-conversion or partial-splice tests, no calibration for forensic use, and no claim about any person. It is a data filter for training sets.
- English, YouTube-style audio. Trained and evaluated on English. Other languages are untested.
- Presumed-human negatives. Some "human" training videos may be TTS that both heuristics missed; that biases the detector towards calling such voices human.
- Numerics. Scores depend slightly on GPU, batch composition and library versions (about 0.01-0.02 logits); decisions near the threshold can flip.
Licence and attribution
- This model (head weights,
detector.py, config): Apache-2.0. - Encoder: facebook/hubert-large-ll60k (Apache-2.0), Hsu et al., "HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units" (2021). It is downloaded from its own repository at the pinned revision, not redistributed here.
- Training audio: YODAS2 (CC BY 3.0), People's Speech (CC BY 4.0), LibriTTS-R (CC BY 4.0), VCTK 0.92 (CC BY 4.0); synthetic renders from Kokoro, Chatterbox, VibeVoice 1.5B, Qwen3-TTS, Maya1, Zonos v0.1, OpenVoice v2 and MetaVoice (Apache-2.0 or MIT). Test-only: Magpie (NVIDIA Open Model License), StyleTTS 2, MLS, LibriHeavy, HiFiTTS-2.
- Emilia-YODAS is distributed by its authors under CC BY 4.0; the companion dataset holds scores and ids, not audio.
Citation
@misc{spero_tts_2026_tts_contamination,
title = {TTS contamination detector: filtering synthetic speech out of web-scraped TTS training data},
author = {{Spero TTS}},
year = {2026},
howpublished = {\url{https://huggingface.co/cloud0day3/tts-contamination-detector}}
}
- Downloads last month
- -
Model tree for cloud0day3/tts-contamination-detector
Base model
facebook/hubert-large-ll60kDatasets used to train cloud0day3/tts-contamination-detector
MLCommons/peoples_speech
mythicinfinity/libritts_r
Evaluation results
- Share of scored hours flagged at 0.32 (%) on Emilia-YODAS Englishself-reported20.050