TTS contamination detector

Your TTS is learning from other TTS. Web-scraped speech corpora are built from YouTube and similar sites, and part of that audio is itself machine-generated: TTS-narrated explainers, list videos, article read-outs, product promos. A TTS model trained on it learns another system's voice and artefacts. This model is a filter for that: it scores clips for synthetic speech so that whole videos or voices can be dropped from a training set.

On Emilia-YODAS English, it flags the videos behind 20.05% of the 21,206 hours we scored (see The Emilia-YODAS finding). In a blind listening check, every flagged clip the listener could judge was TTS (25 of 25; 95% interval 86.7–100%): about 1 in 5 of the Emilia-YODAS English hours we scored is synthetic speech.

  • What it is: frozen HuBERT-Large features, mean- and std-pooled over 8 hidden layers (16,384 numbers per clip), and one linear logistic model. 197 kB of head weights.
  • What it is for: filtering training data. Decisions are made per video or per voice, from many clips, not per clip.
  • What it is not: a deepfake, fraud or forensics detector. It was not built or tested for single-clip verdicts, adversarial audio, or deciding whether a particular person's recording is genuine. Do not use it for that.

Quick start

pip install torch transformers numpy scipy soundfile safetensors huggingface_hub
hf download cloud0day3/tts-contamination-detector --local-dir tts-contamination-detector

# Clips of one video (or one voice): per-clip p and the group decision.
python tts-contamination-detector/detector.py video_clip01.flac video_clip02.flac video_clip03.flac

Output on five clips of one flagged Emilia-YODAS video (file names shortened):

p  0.988    11.75 s  clip1.flac
p  0.999     4.86 s  clip2.flac
p  0.994    12.74 s  clip3.flac
p  1.000     9.17 s  clip4.flac
p  0.997    11.13 s  clip5.flac
group: mean p 0.995 over 5 clips (threshold 0.32): SYNTHETIC

From Python:

import sys
sys.path.insert(0, "tts-contamination-detector")
from detector import TTSDetector, video_decision

det = TTSDetector.from_pretrained("tts-contamination-detector")   # a local folder or "cloud0day3/tts-contamination-detector"
logits = det.score_files(paths)                 # one logit per clip (NaN if a clip cannot be read or is < 0.25 s)
p = det.probability(logits)
decision = video_decision(durations, logits)    # threshold=0.32, min_clips=3
print(decision.synthetic, decision.prob, decision.clips)

detector.py is a standalone port of the pipeline that produced the scores. It depends on numpy, scipy (the resampler: scipy.signal.resample_poly with the pipeline's Kaiser FIR, so features match exactly; torchaudio or soxr would not), soundfile, torch and transformers. It does not need the Spero code base.

  • Front end: mono (channel mean); leading and trailing 20 ms frames quieter than the loudest frame minus 40 dB are trimmed, keeping 0.1 s before and 0.2 s after the speech; the middle 10 s are kept; then 16 kHz.
  • Encoder: facebook/hubert-large-ll60k at revision ff022d09, float32 (TF32 on CUDA), per-utterance normalised input, batched with an attention mask. Hidden states 1, 3, 6, 9, 12, 15, 18 and 24; per layer, the frame mean and standard deviation.
  • Head: logit = clip(w · x + b, −12, 12) on the pooled features (cast to float16 and back, as the pipeline stored them). config.json holds every constant; model.safetensors holds weight, bias, mean and scale (the last two are 0 and 1: the head was folded onto raw features).
  • Video decision: the duration-weighted mean of the clip probabilities (durations floored at 0.1 s) over at least 3 scored clips; flagged at 0.32 or above. Use the clips' own durations, not the 10 s crop.

Equivalence with the pipeline (equivalence.json): on 334 Emilia-YODAS clips (24 whole videos from two shards), detector.py reproduces the pipeline's own code bit for bit (front end, all pooled features, logits). Against the logits the pipeline stored during the Emilia-YODAS run, the largest difference is 0.018 logits (mean 0.0011), from batch composition and GPU type under TF32; every clip and video decision at 0.32 and 0.54 agrees. On a CPU the logits move by about 0.01.

How it was trained

Only permissive audio, with a lineage row for every clip (254,685 rows). No commercial-provider audio (ElevenLabs, Google, OpenAI or any other API) and no private recordings were used.

Role Audio Clips Licence
Synthetic Renders of 8 open TTS systems: Kokoro, Chatterbox, VibeVoice 1.5B, Qwen3-TTS, Maya1, Zonos v0.1, OpenVoice v2, MetaVoice (built-in voices or LibriTTS-R cloning references) 16,800 Apache-2.0 or MIT models; references CC BY 4.0
Synthetic YODAS2 clips of the 12 YouTube voices that read out Wikipedia section markup ("== History ==") 3,052 CC BY 3.0
Human YODAS2 clips from videos that pass both of our earlier heuristics 10,074 CC BY 3.0
Human People's Speech 6,742 CC BY 4.0
Human LibriTTS-R train 16,362 CC BY 4.0
Human VCTK 0.92 (mic1, 110 speakers) 2,640 CC BY 4.0
Test only Magpie 357M and StyleTTS 2 renders (licences not Apache-2.0 or MIT, so never fitted) 4,800
Test only LibriVox read speech (MLS, LibriHeavy, HiFiTTS-2) 20,370 CC BY 4.0 or public domain
  • Every training clip is used twice: clean, and through a seeded mix of reverb, noise, band limiting and an MP3 or Opus round trip. Both classes get the same mix, so a codec cannot become a cue for "human".
  • L2 logistic regression on standardised features, 4 folds that hold out whole TTS systems, whole Wikipedia voices, groups of videos or documents that share a voice, and LibriVox readers. Five layer and regularisation settings were compared on out-of-fold AUC; all eight layers won.
  • The released head is the average of the four fold models of that setting, folded into one weight vector. Those folds held out single videos; a later review re-folded the same plan so that no voice spans two folds, refit, and chose the same setting. The numbers below come from that leak-free refit.
  • The "human" YouTube negatives are presumed human: they passed two heuristics, nobody listened to them.

Results (out of fold)

From the training report (reports/data/tts-detector.md in the Spero TTS repository; section numbers in brackets). Each number comes from a fold model that never saw the clip's TTS system, Wikipedia voice, video group, document group or reader.

Clips (AUC of each held-out synthetic set against all held-out human clips, §3):

Held-out set Clean Through Opus 64 kbit/s
Chatterbox, Qwen3-TTS 0.999 0.999
OpenVoice v2 / Maya1 / Zonos v0.1 0.998 / 0.997 / 0.997 0.998 / 0.997 / 0.997
MetaVoice 0.992 0.993
Kokoro 0.985 0.984
VibeVoice 1.5B 0.971 0.970
Magpie (never trained on) 0.997 0.997
StyleTTS 2 (never trained on) 0.996 0.997
Wikipedia-reader voices on YouTube 0.997

Videos and documents at 0.32 (§3, §4; YODAS2 pilot of 2,838 videos, 428.7 h):

  • Wikipedia-voice videos found: 92% (225 of 245). One voice accounts for 140 of them and is always found; without it, 81% (85 of 105). Two of the 12 voices are mostly missed (2 of 13 and 0 of 8 videos).
  • Certainly-human sets: 0 of 527 YouTube videos with filled pauses ("um", "uh"), 0 of 5,018 People's Speech documents, no VCTK speaker, LibriTTS-R chapter or MLS book. LibriVox read speech is the hardest case: 3 of 337 HiFiTTS-2 chapters (0.9%) and 4 of 199 LibriHeavy books (2.0%).
  • Videos that passed both earlier heuristics: 42 of 2,566 flagged (1.6%); 24 of them share a voice with a Wikipedia-reader or uniform-voice video, and most transcripts read like TTS-narrated genres.
  • Counting every flag outside the Wikipedia set as wrong, precision is at least 80.1% (a loose bound).

Held-out TTS systems, render pseudo-videos of 20 prompts flagged at 0.32 (§3):

System Flagged Where it misses
Chatterbox, Maya1, OpenVoice v2, Qwen3-TTS 100%
Zonos v0.1 99% one cloned voice at 97%
Magpie (never trained on) 98% one voice at 90%
StyleTTS 2 (never trained on) 94% one cloned voice at 77%
Kokoro 75% the am_michael voice: 0%
MetaVoice 62% its clones score 13-97%
VibeVoice 1.5B 52% the Alice voice: 3%

The released head was trained on all eight training systems, these voices included; the table shows how a system it has never heard can slip through, often one clean voice at a time.

The threshold (§3). 0.32 is the smallest threshold at which no certainly-human set lost more than 1% of its recordings in the first evaluation. With leak-free folds the same rule gives 0.54, set by one LibriHeavy book. 0.32 is kept as the default because the detector is meant for YouTube-derived speech, where the filled-pause videos and People's Speech still lose nothing at 0.32, while at 0.54 held-out systems lose much recall (Magpie 86%, StyleTTS 2 72%, MetaVoice 15%). For data that is mostly audiobook-style read speech, 0.54 is what the 1% rule gives.

Threshold Wikipedia videos found Pass-heuristics videos flagged Filler videos People's Speech HiFiTTS-2 LibriHeavy MLS, LibriTTS-R, VCTK
0.10 96.7% 3.5% 0.4% 0.1% 3.0% 7.0% 0
0.20 92.7% 2.3% 0 0.02% 1.5% 4.5% 0
0.32 91.8% 1.6% 0 0 0.9% 2.0% 0
0.50 91.0% 1.2% 0 0 0.3% 1.0% 0
0.70 89.8% 1.1% 0 0 0 0 0

The Emilia-YODAS finding

Emilia is a widely used open TTS training set; its Emilia-YODAS part is cut from YouTube. We scored 8,091,785 English clips (21,205.8 h, 448,394 videos; at most 40 clips per video in each shard part) and applied the video rule above:

Threshold Videos flagged Clips Hours Share of scored hours
0.32 62,538 1,588,547 4,251.2 20.05%
0.54 61,219 1,571,555 4,210.8 19.86%

The two thresholds barely differ: only 1,319 of the videos flagged at 0.32 (2.1%) score below 0.54.

A flag is the detector's claim, not ground truth. A blind listening check (one listener; 44 clips: the first 44, in a score-independent page order, of 90 drawn in proportion to hours from flagged and unflagged videos; protocol fixed before listening) measured:

  • precision of the flagged hours: 100%, 25 heard as TTS, 0 as human, 1 can't tell (86.7–100%, Wilson 95%);
  • synthetic speech among unflagged hours: 0%, 0 TTS, 13 human, 5 can't tell (0–22.8%); a loose bound on what the detector misses;
  • corrected share of synthetic hours: 20.05%.

About 1 in 5 hours of Emilia-YODAS English that we scored is synthetic speech. The scorer samples at most 40 clips per video in each shard part; counting every clip of the scored videos (23,263 h), the flagged videos hold 20.31%.

Per-clip and per-video scores are in the companion dataset cloud0day3/emilia-yodas-tts-contamination, keyed by Emilia item id and YouTube video id, so you can drop the flagged videos from your own copy of Emilia (or choose your own threshold). Only the video rule is applied there; the voice-across-videos rule needs a speaker join and is not reproduced.

Limitations

  • Coverage of TTS systems. It learned from 8 open TTS systems and 12 YouTube Wikipedia-reader voices. Newer and commercial voices were never in training. Held-out systems were usually caught, but single voices slipped through whole (Kokoro am_michael, VibeVoice Alice). Expect it to miss some modern TTS; treat the Emilia-YODAS share as an estimate with that bias.
  • Per video, not per clip. A single clip's probability is noisy. The thresholds and every evaluation number are for videos (or documents, books, voices) of at least 3 clips. Do not use it to label individual clips.
  • Read speech. Careful audiobook reading is the human speech it confuses most (LibriHeavy 2.0% of books at 0.32). Calm, scripted human narration on YouTube can be flagged too.
  • Not a deepfake detector. No adversarial testing, no voice-conversion or partial-splice tests, no calibration for forensic use, and no claim about any person. It is a data filter for training sets.
  • English, YouTube-style audio. Trained and evaluated on English. Other languages are untested.
  • Presumed-human negatives. Some "human" training videos may be TTS that both heuristics missed; that biases the detector towards calling such voices human.
  • Numerics. Scores depend slightly on GPU, batch composition and library versions (about 0.01-0.02 logits); decisions near the threshold can flip.

Licence and attribution

  • This model (head weights, detector.py, config): Apache-2.0.
  • Encoder: facebook/hubert-large-ll60k (Apache-2.0), Hsu et al., "HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units" (2021). It is downloaded from its own repository at the pinned revision, not redistributed here.
  • Training audio: YODAS2 (CC BY 3.0), People's Speech (CC BY 4.0), LibriTTS-R (CC BY 4.0), VCTK 0.92 (CC BY 4.0); synthetic renders from Kokoro, Chatterbox, VibeVoice 1.5B, Qwen3-TTS, Maya1, Zonos v0.1, OpenVoice v2 and MetaVoice (Apache-2.0 or MIT). Test-only: Magpie (NVIDIA Open Model License), StyleTTS 2, MLS, LibriHeavy, HiFiTTS-2.
  • Emilia-YODAS is distributed by its authors under CC BY 4.0; the companion dataset holds scores and ids, not audio.

Citation

@misc{spero_tts_2026_tts_contamination,
  title        = {TTS contamination detector: filtering synthetic speech out of web-scraped TTS training data},
  author       = {{Spero TTS}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/cloud0day3/tts-contamination-detector}}
}
Downloads last month
-
Safetensors
Model size
49.2k params
Tensor type
F64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cloud0day3/tts-contamination-detector

Finetuned
(24)
this model

Datasets used to train cloud0day3/tts-contamination-detector

Evaluation results

  • Share of scored hours flagged at 0.32 (%) on Emilia-YODAS English
    self-reported
    20.050