DGE all-tags audio tagger

EfficientAT mn10_as (MIT, fschmid56/EfficientAT, AudioSet-pretrained) fine-tuned to tag instruments, sound types and effects for Document Graph Explorer. Input: 32 kHz mono, 128-band log-mel frames as in EfficientAT; output: one sigmoid per label in onnx/model.json, with per-label thresholds in thresholds.json.

Fine-tuning data: FSD50K (CC licences), NSynth (CC BY 4.0), MTG-Jamendo and OpenMIC (CC licences), Freesound clips (per-clip CC licences) and SoundCloud preview clips. No audio is redistributed here; only the trained weights.

eval.txt holds held-out scores (precision/recall per label). Research and app use; check the licences of the training sets before other uses.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support