trashbin+ AI music detector
Detects AI-generated songs from a 30 second Spotify preview. Used by the trashbin+ Spicetify extension, where it runs locally in the browser through onnxruntime-web.
Model
detector-mn10-v5.onnx(17 MB): EfficientAT MobileNetV3mn10_as(AudioSet pretrained, 4.2M parameters) fine-tuned as a binary AI vs human classifier.- Input
waveform: float32[N, 320000], mono 32 kHz, 10 second windows. The log-mel frontend is inside the graph. - Output
logit:[N]. The extension scores the start, middle and end window of a preview and averages the logits, then applies a sigmoid.
Training data
About 22,000 clips from 50+ sources, all re-encoded to the Spotify preview format (96 kbps MP3, 44.1 kHz stereo, 30 s):
- AI: Suno v2 to v6, Udio, Lyria 3 and 3.5, ElevenLabs Music v1 and v2, Mureka, MiniMax, MusicGPT, Stable Audio 1 to 3, MusicGen, Riffusion/Producer, ACE-Step, YuE, DiffRhythm, SongGen, HeartMuLa, Mubert, Boomy, AIVA, Soundful, Soundraw, TemPolor, Brev, and AI artists on Spotify.
- Human: Spotify releases from before 2022, FMA, MTG-Jamendo, MusicCaps.
- Public datasets used include ArtifactBench, HAIM and SONICS (CC BY-NC 4.0), MUSIC8K (CC BY 4.0) and Echoes (CC BY-SA 4.0). Because part of the training data is non-commercial, the weights are released under CC BY-NC 4.0.
Results (v5, held-out eval set, threshold 0.8)
- AUC 0.994.
- Human songs flagged: Spotify 0 of 299, FMA 3%, MTG-Jamendo 0 to 10%, MusicCaps 0%.
- Caught: AI artists on Spotify 95%; neural generators (Suno, Lyria, ElevenLabs, Mureka, MiniMax, MusicGPT, Stable Audio, MusicGen, ACE-Step) 87 to 100%.
- Weak: AIVA 40%, Loudly releases 53%, Mubert 67%, Boomy 70%, current Udio 71%. These assemble recorded loops or render MIDI with sample libraries and leave few audio artifacts.
Scores are probabilities from one model, not proof. Training code: ai/.