AudioWeb models
ONNX exports of Meta's HTDemucs 6-source model (htdemucs_6s), used by
AudioWeb to split songs into vocals, drums, bass, guitar, piano and other,
entirely on the user's device (in the browser and in the iOS and Android apps).
| File | Window | Notes | SHA-256 |
|---|---|---|---|
demucs_htdemucs_6s_s78_fp16w_ca.onnx |
7.8 s | Default. fp16 weights, attention computed in query chunks for a lower memory peak (same output as the unchunked graph, max difference ~1e-6). | 05079e586874d966ceb1b8bca286f38f78d5602ed2c51d475a4092839891e0d5 |
demucs_htdemucs_6s_s39_fp16w_ca.onnx |
3.9 s | Phones and devices with 4 GB of memory or less. Same export, shorter window. | 761501fb6448709414eb4ce13d67dbd2866de165c94770c7d893d63b9ba0c2d1 |
Input is stereo audio at 44.1 kHz in fixed-length segments; the STFT and its inverse run outside the model (in JavaScript in AudioWeb). Outputs are the six sources in the order drums, bass, other, vocals, guitar, piano.
Source and license
The weights are Meta AI's Demucs (Alexandre Défossez, Simon Rouard and contributors), released under the MIT License: https://github.com/facebookresearch/demucs. These files are unmodified weights re-exported to ONNX (fp16 storage, chunked attention) and are distributed under the same MIT License.
If you use them, please cite the Demucs papers:
- Rouard, Massa, Défossez. Hybrid Transformers for Music Source Separation. ICASSP 2023.
- Défossez. Hybrid Spectrogram and Waveform Source Separation. ISMIR 2021 Workshop on Music Source Separation.
Music analysis models (analysis/)
Small models AudioWeb runs on the device after a split: beats and bars, chords, and notes for
MIDI export. analysis/manifest.json lists each file with its size and SHA-256, and the app
checks every download against it. analysis/THIRD_PARTY_NOTICES.txt carries the full license texts.
| File | Model | License | SHA-256 |
|---|---|---|---|
analysis/beat_this_small.onnx |
Beat This! small (CPJKU/beat_this) | MIT | 6d9e1b3ebdb953b0bacf6755e6fa421e04e74709cdbe4abcc6f0e4b432c89f09 |
analysis/btc_majmin.onnx |
BTC major/minor, 25 classes (jayg996/BTC-ISMIR19) | MIT | 9660ad90d38170d265bf97b6dc0c5f9f04cf7f7f21c507e0b27e18990c88b209 |
analysis/btc_large_voca.onnx |
BTC large vocabulary, 170 classes (jayg996/BTC-ISMIR19) | MIT | 3576d3a52628e130252af8b70378dbbd2f678efe27c9dd520930d671055bad2a |
analysis/btc_meta.json |
BTC normalization and label tables | MIT | cd56da51912598d0dacd7113dbd93bf03ebc4d96ecd546dd99ecbe6184254add |
analysis/basic_pitch_nmp.onnx |
Basic Pitch ICASSP 2022 (spotify/basic-pitch) | Apache-2.0 | 2c3c1d144bfa61ad236e92e169c13535c880469a12a047d4e73451f2c059a0ec |
Demo song (demo/)
demo/sunday-groove.zip is AudioWeb's demo song, already split and analysed. It was synthesized
from scratch for AudioWeb (no samples, no third-party audio) and is released under the MIT License.
SHA-256 891a9a0974d389d499a7bb92663b810bcbbb3804b8602c81d46093ae4c704a2b.
Download count
Hugging Face counts a repo's downloads by requests for config.json, not for its model files. So
when a device downloads one of the separation models for the first time, AudioWeb also requests
config.json (nothing else is sent). This repo's download count is therefore roughly the number
of devices that have split a song with AudioWeb.
- Downloads last month
- 39