Beat This! — GGUF

GGUF conversions of Beat This! (Foscarin, Schlüter & Widmer, CPJKU — ISMIR 2024), a transformer beat and downbeat tracker, for use with CrispASR / ggml.

Checkpoint: final0 (trained on all datasets except GTZAN, seed 0).

Why this model

Nearly every published beat tracker post-processes its framewise output with madmom's Dynamic Bayesian Network, which is Böck-patented and licensed for non-commercial use only. Beat This! reaches state-of-the-art without one — its postprocessing is plain peak-picking, and its dependency list (numpy / torch / torchaudio / einops / rotary-embedding-torch / soxr) contains no part of madmom.

Both the upstream code and the published weights are MIT. That combination — SOTA accuracy, no DBN, MIT weights — is why this is usable in a commercial product where most beat trackers are not.

Files

file size notes
beat-this-f16.gguf 41 MB default. Recommended for all normal use.
beat-this-f32.gguf 81 MB Reference build for exact-parity debugging.

20.25 M parameters, 147 tensors. Both files bake in the [513, 128] mel filterbank the model was exported with, so the front end never re-derives it (slaney-vs-htk and the freq/mel layout are classic silent-drift sources).

Verification

Ported stage by stage against a PyTorch reference driven by the original checkpoint, comparing every sub-block rather than only the final output.

At f32, every stage is numerically exact:

stage cos max rel err
stem 1.00000000 3.0e-7
blk0_partial 1.00000000 3.0e-7
blk2 1.00000000 9.5e-7
linear 1.00000000 1.3e-6
transformer 1.00000000 9.3e-7
out_beat / out_downbeat 1.00000000 2.7e-6 / 1.4e-6

At f16 the same stages score cos ≥ 0.99999973 with max rel err ~5e-4, flat across all 12 attention/FF sub-blocks rather than compounding — i.e. the residual is weight quantisation, not drift. The log-mel front end matches torchaudio at cos = 1.00000000.

Windowing (1500-frame chunks, 6-frame border, keep_first overlap) and the peak-picking postprocessor reproduce upstream exactly: on a 45 s two-chunk fixture, running the reference's own logits through the ported peak-picker gives identical beat and downbeat times to 1e-6 s.

Usage

# one line per beat: time_sec <TAB> beat|downbeat
crispasr --beats -m beat-this-f16.gguf -f song.wav

# JSON, including a median-interval tempo estimate
crispasr --beats -m beat-this-f16.gguf --beats-format json -f song.wav

Input is decoded to the model's native 22.05 kHz mono automatically, and long files are chunked internally.

Every downbeat is also reported as a beat. The postprocessor snaps each downbeat onto its nearest detected beat, so downbeats are a strict subset and you never have to merge two lists to reconstruct the grid.

Licence and provenance

Code and weights are MIT, per upstream. Note that upstream's own README records that some of the training audio is copyrighted or under restrictive Creative Commons terms; the licence on the released weights is unambiguous, but the provenance of the training corpus is a separate question that upstream, not this conversion, is the authority on.

Citation

@inproceedings{foscarin2024beatthis,
    title     = {Beat this! Accurate beat tracking without DBN postprocessing},
    author    = {Foscarin, Francesco and Schl{\"u}ter, Jan and Widmer, Gerhard},
    booktitle = {Proceedings of the 25th International Society for
                 Music Information Retrieval Conference (ISMIR)},
    year      = {2024}
}
Downloads last month
84
GGUF
Model size
20.3M params
Architecture
beat-this
Hardware compatibility
Log In to add your hardware

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support