AMTFlow
Model assets for AMTFlow, a modular instrument-conditioned automatic music transcription pipeline (instrument tagging → instrument-family routing → MuScriptor). AMTFlow downloads these files on first use; you do not need to fetch them by hand.
heads/openmic/ — instrument-tagging heads
One classification head per frozen audio encoder, trained on
OpenMIC-2018 (20 instrument classes). AMTFlow's
probe tagger runs the encoder, applies the head to one layer's time-averaged
features, and marks an instrument family present when any of its classes reaches that
class's threshold. Several taggers, including audio-language models, can be combined
by union or majority vote.
| head | encoder | layer | OpenMIC test mAP | macro-F1 |
|---|---|---|---|---|
openmic/af3-whisper-mlp |
nvidia/audio-flamingo-3-hf audio tower | 31 | 0.861 ± 0.001 | 0.793 |
openmic/afnext-whisper-mlp |
nvidia/audio-flamingo-next-hf audio tower | 29 | 0.860 ± 0.003 | 0.788 |
openmic/matpac-music-mlp |
auriankelen/matpac_music | 10 | 0.860 ± 0.002 | 0.792 |
openmic/matpac-general-mlp |
auriankelen/matpac_general_audio | 10 | 0.858 ± 0.002 | 0.787 |
openmic/pupum2d-mlp |
spellbrush/PupuM2D (large) | 19 | 0.858 ± 0.001 | 0.790 |
openmic/muq-mlp |
OpenMuQ/MuQ-large-msd-iter | 2 | 0.832 ± 0.001 | 0.764 |
openmic/qwen3-omni-aut-mlp |
Qwen/Qwen3-Omni-30B-A3B-Instruct AuT encoder | 23 | 0.821 ± 0.001 | 0.753 |
openmic/mert-mlp |
m-a-p/MERT-v1-330M | 9 | 0.818 ± 0.002 | 0.740 |
Test mAP is the mean over five probe seeds with a 95% t-interval over seeds (the file is the best-validation seed). For reference, the MATPAC++ authors' released OpenMIC probe scores 0.861 on the same test clips.
Protocol. Official split01 test partition (5,085 clips); the official train
partition split 80/20 by track (seed 0) into train/validation. Labels: aggregated
relevance ≥ 0.5 is positive; unobserved (clip, class) labels are masked out. Every
encoder layer is averaged over the 10 s clip and z-scored on train; the head is an MLP
(512 hidden, ReLU, dropout 0.2) trained with masked BCE; layer and learning rate are
selected on validation mAP. Per-class thresholds are calibrated on validation by
F-beta (β = 2, precision ≥ 0.25), i.e. recall-oriented, because the transcriber
cannot produce an instrument outside its conditioning set.
Files. <encoder>-mlp.safetensors holds the head weights and the train-set
feature mean/std; its metadata (amtflow key, JSON) holds the encoder, layer, class
names, thresholds and calibration details. index.json lists every head with its
test scores and SHA-256.
License
The heads are released under CC BY 4.0 (OpenMIC-2018 is CC BY 4.0). A head is only usable together with its encoder, and the encoder's own license applies: MATPAC++ (Apache-2.0), PupuM2D (MIT), Qwen3-Omni (Apache-2.0), MERT and MuQ (CC BY-NC 4.0), Audio Flamingo 3 / Next (NVIDIA non-commercial). MuScriptor, the transcriber AMTFlow uses, is CC BY-NC 4.0.