AMTFlow

Model assets for AMTFlow, a modular instrument-conditioned automatic music transcription pipeline (instrument tagging → instrument-family routing → MuScriptor). AMTFlow downloads these files on first use; you do not need to fetch them by hand.

heads/openmic/ — instrument-tagging heads

One classification head per frozen audio encoder, trained on OpenMIC-2018 (20 instrument classes). AMTFlow's probe tagger runs the encoder, applies the head to one layer's time-averaged features, and marks an instrument family present when any of its classes reaches that class's threshold. Several taggers, including audio-language models, can be combined by union or majority vote.

head encoder layer OpenMIC test mAP macro-F1
openmic/af3-whisper-mlp nvidia/audio-flamingo-3-hf audio tower 31 0.861 ± 0.001 0.793
openmic/afnext-whisper-mlp nvidia/audio-flamingo-next-hf audio tower 29 0.860 ± 0.003 0.788
openmic/matpac-music-mlp auriankelen/matpac_music 10 0.860 ± 0.002 0.792
openmic/matpac-general-mlp auriankelen/matpac_general_audio 10 0.858 ± 0.002 0.787
openmic/pupum2d-mlp spellbrush/PupuM2D (large) 19 0.858 ± 0.001 0.790
openmic/muq-mlp OpenMuQ/MuQ-large-msd-iter 2 0.832 ± 0.001 0.764
openmic/qwen3-omni-aut-mlp Qwen/Qwen3-Omni-30B-A3B-Instruct AuT encoder 23 0.821 ± 0.001 0.753
openmic/mert-mlp m-a-p/MERT-v1-330M 9 0.818 ± 0.002 0.740

Test mAP is the mean over five probe seeds with a 95% t-interval over seeds (the file is the best-validation seed). For reference, the MATPAC++ authors' released OpenMIC probe scores 0.861 on the same test clips.

Protocol. Official split01 test partition (5,085 clips); the official train partition split 80/20 by track (seed 0) into train/validation. Labels: aggregated relevance ≥ 0.5 is positive; unobserved (clip, class) labels are masked out. Every encoder layer is averaged over the 10 s clip and z-scored on train; the head is an MLP (512 hidden, ReLU, dropout 0.2) trained with masked BCE; layer and learning rate are selected on validation mAP. Per-class thresholds are calibrated on validation by F-beta (β = 2, precision ≥ 0.25), i.e. recall-oriented, because the transcriber cannot produce an instrument outside its conditioning set.

Files. <encoder>-mlp.safetensors holds the head weights and the train-set feature mean/std; its metadata (amtflow key, JSON) holds the encoder, layer, class names, thresholds and calibration details. index.json lists every head with its test scores and SHA-256.

License

The heads are released under CC BY 4.0 (OpenMIC-2018 is CC BY 4.0). A head is only usable together with its encoder, and the encoder's own license applies: MATPAC++ (Apache-2.0), PupuM2D (MIT), Qwen3-Omni (Apache-2.0), MERT and MuQ (CC BY-NC 4.0), Audio Flamingo 3 / Next (NVIDIA non-commercial). MuScriptor, the transcriber AMTFlow uses, is CC BY-NC 4.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support