Piano transcription for gospel keyboards

A fine-tune of ByteDance's high-resolution piano transcription model (Kong et al., 2021) for the keyboard parts of worship and gospel band recordings, where organ, electric piano and pads play alongside the piano. It was made for PianoScript, which turns church musicians' recordings into piano scores.

The original model was trained on solo piano only. On organ and pads it finds new attacks in held, wavering chords (a held chord comes out as a run of repeated notes) and misses much of the harmony. This checkpoint keeps the original architecture and file format and was trained further on band arrangements with known notes.

Use

Same format as the original Note_pedal checkpoint: a dict whose model entry holds note_model and pedal_model. It loads with piano_transcription_inference 0.0.6:

from piano_transcription_inference import PianoTranscription, load_audio, sample_rate

transcriptor = PianoTranscription(device="cpu", checkpoint_path="piano_transcription_keyboards.pth")
audio, _ = load_audio("keyboards.wav", sr=sample_rate, mono=True)
transcriptor.transcribe(audio, "keyboards.mid")

It expects the keyboards with vocals and drums already removed. PianoScript feeds it the other + bass stems of Demucs htdemucs.

For solo piano, use the original checkpoint. On real solo piano this one stays close to it but finds fewer notes (see below).

Training

  • Start: the original checkpoint (Zenodo record 4034264, note F1 0.9677).
  • Band songs: 1,500 generated songs (about 15 hours) with every note known. They use gospel chord progressions with extended and passing chords, and the accompaniment is block chords, pushes, arpeggios or runs. Organ and pads hold the notes shared between chords.
    • Sounds come from the FluidR3_GM and MuseScore_General sound fonts, plus a procedural tonewheel organ (drawbars, percussion, key click, overdrive, Leslie), an FM electric piano, pads and synth bass.
    • Each song is mixed with bass, drums, lead and choir voices, given reverb, MP3-encoded half of the time, then separated with Demucs htdemucs, exactly as in production.
  • Real piano: 4.35 hours of solo piano recordings from Wikimedia Commons (mostly Musopen; public domain, CC0 and CC BY 3.0), with the original model's output as target (distillation), so the fine-tune does not drift away from real pianos.
  • Optimisation: 4,000 steps of 8 ten-second segments; Adam (AMSGrad), learning rate 1e-4 with 300 warm-up steps and a cosine decay to 1e-5; BatchNorm frozen; mixed precision. About 3 hours on one RTX 3060 Laptop GPU.

Evaluation

Band songs. 16 synthetic songs rendered with a sound font never used in training (GeneralUser GS), through PianoScript's whole pipeline (separation, notes, bass line). Notes match when their onset is within 50 ms.

Original This checkpoint
Precision 0.65 0.80
Recall 0.75 0.83
Keyboard recall 0.72 0.81
Repeated notes in excess 362 72
Harmonics in excess 1,005 681

On the organ songs alone, recall rises from 0.33โ€“0.51 to 0.63โ€“0.78.

Real solo piano. Six Chopin recordings kept out of training, with no score. Agreement with the original model (onset F1) is 0.85 to 0.98. The notes this checkpoint keeps are nearly all ones the original also found (precision 0.93 to 0.99), but it finds 1 to 17% fewer notes. The biggest loss is on a slow, pedalled nocturne.

Limitations

  • Trained on synthetic songs. On four real gospel recordings, which have no score, the share of notes struck again just as the same key is released fell on three of them, for example from 40% to 31%, and stayed flat on the fourth. That is much less than on the synthetic benchmark.
  • Organ recall stays well below piano recall.
  • Pedal detection was not trained further and is not evaluated here.

License and attribution

CC BY 4.0, like the original checkpoint it is derived from. These are fine-tuned weights: the architecture and the starting point are the authors' work.

Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Wan and Yuxuan Wang. High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 3707โ€“3717, 2021. Original checkpoint: https://zenodo.org/records/4034264

Real piano recordings: Wikimedia Commons, used under their own licences (public domain, CC0, CC BY 3.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support