Piano transcription for gospel keyboards
A fine-tune of ByteDance's high-resolution piano transcription model (Kong et al., 2021) for the keyboard parts of worship and gospel band recordings, where organ, electric piano and pads play alongside the piano. It was made for PianoScript, which turns church musicians' recordings into piano scores.
The original model was trained on solo piano only. On organ and pads it finds new attacks in held, wavering chords (a held chord comes out as a run of repeated notes) and misses much of the harmony. This checkpoint keeps the original architecture and file format and was trained further on band arrangements with known notes.
Use
Same format as the original Note_pedal checkpoint: a dict whose model
entry holds note_model and pedal_model. It loads with
piano_transcription_inference 0.0.6:
from piano_transcription_inference import PianoTranscription, load_audio, sample_rate
transcriptor = PianoTranscription(device="cpu", checkpoint_path="piano_transcription_keyboards.pth")
audio, _ = load_audio("keyboards.wav", sr=sample_rate, mono=True)
transcriptor.transcribe(audio, "keyboards.mid")
It expects the keyboards with vocals and drums already removed. PianoScript
feeds it the other + bass stems of Demucs htdemucs.
For solo piano, use the original checkpoint. On real solo piano this one stays close to it but finds fewer notes (see below).
Training
- Start: the original checkpoint (Zenodo record 4034264, note F1 0.9677).
- Band songs: 1,500 generated songs (about 15 hours) with every note
known. They use gospel chord progressions with extended and passing
chords, and the accompaniment is block chords, pushes, arpeggios or runs.
Organ and pads hold the notes shared between chords.
- Sounds come from the FluidR3_GM and MuseScore_General sound fonts, plus a procedural tonewheel organ (drawbars, percussion, key click, overdrive, Leslie), an FM electric piano, pads and synth bass.
- Each song is mixed with bass, drums, lead and choir voices, given
reverb, MP3-encoded half of the time, then separated with Demucs
htdemucs, exactly as in production.
- Real piano: 4.35 hours of solo piano recordings from Wikimedia Commons (mostly Musopen; public domain, CC0 and CC BY 3.0), with the original model's output as target (distillation), so the fine-tune does not drift away from real pianos.
- Optimisation: 4,000 steps of 8 ten-second segments; Adam (AMSGrad), learning rate 1e-4 with 300 warm-up steps and a cosine decay to 1e-5; BatchNorm frozen; mixed precision. About 3 hours on one RTX 3060 Laptop GPU.
Evaluation
Band songs. 16 synthetic songs rendered with a sound font never used in training (GeneralUser GS), through PianoScript's whole pipeline (separation, notes, bass line). Notes match when their onset is within 50 ms.
| Original | This checkpoint | |
|---|---|---|
| Precision | 0.65 | 0.80 |
| Recall | 0.75 | 0.83 |
| Keyboard recall | 0.72 | 0.81 |
| Repeated notes in excess | 362 | 72 |
| Harmonics in excess | 1,005 | 681 |
On the organ songs alone, recall rises from 0.33โ0.51 to 0.63โ0.78.
Real solo piano. Six Chopin recordings kept out of training, with no score. Agreement with the original model (onset F1) is 0.85 to 0.98. The notes this checkpoint keeps are nearly all ones the original also found (precision 0.93 to 0.99), but it finds 1 to 17% fewer notes. The biggest loss is on a slow, pedalled nocturne.
Limitations
- Trained on synthetic songs. On four real gospel recordings, which have no score, the share of notes struck again just as the same key is released fell on three of them, for example from 40% to 31%, and stayed flat on the fourth. That is much less than on the synthetic benchmark.
- Organ recall stays well below piano recall.
- Pedal detection was not trained further and is not evaluated here.
License and attribution
CC BY 4.0, like the original checkpoint it is derived from. These are fine-tuned weights: the architecture and the starting point are the authors' work.
Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Wan and Yuxuan Wang. High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 3707โ3717, 2021. Original checkpoint: https://zenodo.org/records/4034264
Real piano recordings: Wikimedia Commons, used under their own licences (public domain, CC0, CC BY 3.0).