RysUpAlign learned alignment model v6

Frame embeddings for audio-to-audio vocal alignment: aligning a double, gang or backing vocal to a guide vocal without lyrics. Embeddings of the two takes are compared with dynamic time warping (DTW); the model is trained so that the same sung moment in two takes gets the same embedding even across singers, timbre and pitch (unison, harmony, octave).

It is the timing model of the RysUpAlign plug-in by Rys Up Audio. Benchmark, code and evaluation: https://github.com/rysupaudio-lab/rysupalign-bench

Files

file size sha256
rysupalign_learned_v6.onnx 305 MB, fp32 db3bf5b88834bba926c147d8785e829845c7827248efc74e2966f0bec7ec8b32
rysupalign_learned_v6.int8.onnx 112 MB, dynamic int8 e48d692579af3d8ed710ff99e47a1c37d0c1e0205440d5a69d1521a69f505246

Both give the same accuracy on the benchmark (see below).

Input / output

  • input wav: float32 [1, S], mono 16 kHz audio, normalised to zero mean and unit variance over the file
  • output emb: float32 [1, T, 128], unit-norm embeddings at 100 fps; frame k is centred at 10k + 12.5 ms

For long files, run 30 s chunks with 2 s overlap and drop 1 s on each side of every join.

import numpy as np, librosa, onnxruntime as ort
sess = ort.InferenceSession("rysupalign_learned_v6.onnx", providers=["CPUExecutionProvider"])
y, _ = librosa.load("take.wav", sr=16000, mono=True)
y = ((y - y.mean()) / (y.std() + 1e-7)).astype(np.float32)
emb = sess.run(None, {"wav": y[None]})[0][0]          # [T, 128], 100 fps
# cost between two takes: 1 - emb_a @ emb_b.T, then DTW

The complete aligners used for the results (gating, log-mel block, DTW method M1) are runners/learned_onnx.py and runners/m1_dtw.py in the GitHub repository.

Architecture

  • Encoder: HuBERT-base (facebook/hubert-base-ls960, revision af46f65f540dc3ca7aa59f46c6c3d5dbb4374fa8, Apache-2.0), frozen, first 9 transformer layers.
  • Head (1.66 M parameters): LayerNorm and a learned softmax-weighted sum of layers 3โ€“9, Linear 768โ†’256, ร—2 transposed-conv upsampling to 100 fps, plus a log-mel branch (64 bands, 25 ms; 80 bands over 80โ€“2000 Hz, 64 ms) on the same grid, 3 residual dilated Conv1d blocks, Linear โ†’ 128, L2 normalisation.
  • The log-mel front end is part of the graph (STFT as fixed conv1d kernels), so the ONNX file takes raw audio.

Training

Symmetric frame-level InfoNCE (positive = the true fractional frame in the other take, negatives = all other frames within ยฑ1 s). Data: synthetic doubles with exact timing (WSOLA timing slop, pitch and formant shifts, EQ, drive, reverb, noise) made from Rys Up Audio's own studio vocal recordings and VocalSet 1.2 (CC BY 4.0), plus pseudo-labelled real double takes and VocalSet cross-singer pairs. No evaluation audio was used for training.

Results (RAB, mean / median / share within 20 ms of the timing error)

system R: real unison choir pairs (48) S: exact-truth cases (120), mean D: held-out Dagstuhl pairs (12), mean
no alignment 52.0 / 34.8 ms / 33.7% 39.5 ms 65.1 ms
HuBERT-base L6 + log-mel DTW 47.4 / 24.1 ms / 45.0% 6.1 ms 42.6 ms
MERT-v1-95M L4 + log-mel DTW 44.8 / 20.0 ms / 49.9% โ€“ 47.8 ms
MERT-v1-330M L10 + log-mel DTW โ€“ โ€“ 45.2 ms
this model + log-mel, DTW 43.8 / 20.3 ms / 49.5% 3.9 ms 42.0 ms
this model + log-mel, DTW method M1 39.8 / 14.8 ms / 58.4% 3.2 ms 43.0 ms (C++ engine)

On the held-out suite D the model is statistically tied with HuBERT-base features (12 cases); the gain on R comes mostly from the DTW method and did not carry over to D. Suite R was used for model selection. Commercial alignment tools were not evaluated. Full tables, caveats (including an invalid annotation in the Choral Singing Dataset that affects 2 of the 48 R cases) and the scripts to reproduce everything are in the GitHub repository.

Licence

  • Model weights: PolyForm Noncommercial License 1.0.0 (LICENSE). Noncommercial use (research, personal, educational, by noncommercial organisations) is permitted under its terms.
  • The HuBERT-base encoder weights contained in the files remain under the Apache License 2.0; see NOTICE and the appendix of LICENSE.
  • Commercial licensing of the weights is available from Rys Up Audio: https://rysupaudio.com/pages/contact-us

Required Notice: Copyright 2026 Rys Up Audio LLC (https://rysupaudio.com)

Citation

Please cite the repository (https://github.com/rysupaudio-lab/rysupalign-bench, CITATION.cff) and HuBERT (Hsu et al., IEEE/ACM TASLP 2021, arXiv:2106.07447) and VocalSet (Wilkins et al., ISMIR 2018).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rysupaudio/rysupalign-learned-v6

Quantized
(10)
this model

Paper for rysupaudio/rysupalign-learned-v6