BTC chord recognition β GGUF
GGUF conversions of BTC (A Bi-directional Transformer for Musical Chord Recognition, Park et al., ISMIR 2019) for use with CrispASR.
Audio in, a chord timeline out:
0.000 1.950 C
1.950 3.901 G
3.901 8.081 N
That is the .lab layout the standard chord datasets use, so the output drops
straight into mir_eval. N means "no chord".
β οΈ Licence β non-commercial weights
These weights are CC-BY-NC-SA 4.0 and may NOT be used commercially.
This matters because the surrounding software is not:
| Licence | |
|---|---|
| Upstream BTC code (jayg996/BTC-ISMIR19) | MIT |
| CrispASR | MIT |
| These weights | CC-BY-NC-SA 4.0 β non-commercial |
The checkpoints were trained on the Isophonics, Robbie Williams and UsPop2002 chord annotations, whose terms forbid commercial use. The MIT licence on the code does not extend to them.
CrispASR enforces this: its model registry refuses to download these files unless you explicitly accept the licence.
crispasr --chords -m auto --auto-download \
--accept-license cc-by-nc-sa-4.0 -f song.wav
If you are shipping a commercial product, you must train or supply your own
weights. The converter (models/convert-btc-to-gguf.py) works on any BTC
checkpoint.
Files
| File | Size | Classes | Notes |
|---|---|---|---|
btc-chords-large-f16.gguf |
5.6 MB | 170 | default |
btc-chords-large-q8_0.gguf |
4.5 MB | 170 | smallest with no measurable loss |
btc-chords-large-f32.gguf |
11.2 MB | 170 | reference precision |
btc-chords-f16.gguf |
5.6 MB | 25 | maj/min + N |
btc-chords-q8_0.gguf |
4.4 MB | 25 | |
btc-chords-f32.gguf |
11.2 MB | 25 | reference precision |
No q4_k is published, deliberately. It builds and runs, but measured against the torch reference on real music it costs real accuracy for very little space:
| size | root | mirex | tetrads | |
|---|---|---|---|---|
| f16 | 5.6 MB | 99.17 % | 98.81 % | 98.56 % |
| q8_0 | 4.5 MB | 99.17 % | 98.81 % | 98.52 % |
| q4_k | 3.8 MB | 98.42 % | 97.34 % | 95.46 % |
q8_0 is indistinguishable; q4_k gives up 3.1 points of tetrad accuracy to save 0.6 MB on an already-tiny model. Quantize from the f16, not the f32 β only 73 of 213 tensors are quantizable, so the rest inherit the source dtype and a q8_0 built from f32 comes out at 7.5 MB, larger than the f16.
Prefer the 170-class model. It collapses to maj/min on demand with
CRISPASR_BTC_MAJ_MIN=1, whereas a 25-class model can never be expanded. The
two upstream checkpoints differ only in classifier width.
f16 is the shipping default. It passes the per-stage parity harness identically to f32 (below), at half the size.
The 170-class vocabulary is root * 14 + quality for the 12 roots and 14
qualities (min, maj, dim, aug, min6, maj6, min7, minmaj7, maj7, 7, dim7, hdim7, sus2, sus4), plus X (unknown) and N (no chord). Quality maj
renders as the bare root, so index 1 is C, not C:maj.
Verification
Every file here was validated per-stage against the PyTorch reference, not just spot-checked end-to-end.
| Check | Result |
|---|---|
| numpy spec vs PyTorch (170-class) | cos 1.00000004, argmax agreement 1.0000 |
| numpy spec vs PyTorch (25-class) | cos 0.99999991, argmax agreement 1.0000 |
crispasr-diff btc β 170-class f32 |
13/13 stages, cos 1.000000 |
crispasr-diff btc β 170-class f16 |
13/13 stages, cos 1.000000 |
crispasr-diff btc β 25-class f32 |
13/13 stages, cos 1.000000 |
crispasr-diff btc β 25-class f16 |
13/13 stages, cos 1.000000 |
| Chord agreement vs the torch reference, 257 s of real music | 98.6 β 99.2 % (mir_eval) |
Stages compared: input projection + positional encoding, both attention directions of layer 0, all 8 layer outputs, the final norm, and the logits.
Reproduce with:
python tools/btc_torch_parity.py model.gguf btc_model_large_voca.pt BTC-ISMIR19/ ref.gguf
crispasr-diff btc model.gguf ref.gguf audio.wav
End-to-end on the upstream test clip (4 min 17 s of real music), scored with
mir_eval against the PyTorch implementation's own output:
| metric | agreement |
|---|---|
| root | 99.17 % |
| majmin | 98.86 % |
| mirex | 98.81 % |
| thirds / triads | 98.88 % / 98.78 % |
| sevenths / tetrads | 98.70 % / 98.56 % |
Known limitation: the above measures fidelity to the reference
implementation, which is what a port owes. It is not an accuracy claim
against human chord annotations β the published MIREX figures are the model's
property, and were not re-measured here. On synthetic test tones the 25-class
model returns N where the 170-class model resolves chords; that is
out-of-distribution behaviour, not a conversion defect. Evaluate on your own
material before relying on it.
Usage
# 170-class output (default)
crispasr --chords -m btc-chords-large-f16.gguf -f song.wav
# collapse to plain major/minor
CRISPASR_BTC_MAJ_MIN=1 crispasr --chords -m btc-chords-large-f16.gguf -f song.wav
# JSON, with per-span confidence
crispasr --chords --chords-format json -m btc-chords-large-f16.gguf -f song.wav
The backend is auto-detected from the GGUF (general.architecture = btc), and
audio is resampled internally to BTC's native 22.05 kHz mono.
Also available through the session C ABI (crispasr_session_chords*) and the
WASM binding (sessionChords).
Model details
- 8 layers, hidden size 128, 4 heads, FFN width 128
- Blocks of exactly 108 frames β a fixed block, not a sliding window
- Each layer runs two attention blocks over the same input: one causal, one with the mask transposed; outputs concatenated and projected 256 β 128
- Front end: constant-Q transform, 144 bins, 24 per octave, log magnitude, normalised by mean/std stored in the GGUF
- 213 tensors converted; 8 upstream tensors (
output_layer.lstm.*) are dead weight and are skipped
Conversion and runtime details, including the ten implementation details that
are not the obvious default, are documented in CrispASR under
docs/music-transcription/.
Citation
@inproceedings{park2019bi,
title={A Bi-Directional Transformer for Musical Chord Recognition},
author={Park, Jonggwon and Choi, Kyoyun and Jeon, Sungwook and Kim, Dokyun and Park, Jonghun},
booktitle={Proceedings of the 20th International Society for Music Information Retrieval Conference (ISMIR)},
year={2019}
}
- Downloads last month
- 150
8-bit
16-bit
32-bit