muse sample types
Name any audio sample on a Mac, on-device. One Core AI model that listens to a sample once and answers two questions:
- What is it? One of 22 sample types, such as kick, snare, bass, pluck, pad, vocal or sweep, with the runner-up types and their probabilities.
- Does a musical key mean anything here? Tonal or atonal, so a kick drum is never labelled "C minor".
It also returns the sound's 512-number CLAP embedding, for "find sounds like this" search.
It is the model behind the class and tonal columns of the
muse command-line tool, which downloads it on first use.
Labels
22 types in five families (types.txt). A label says what the sound is; whether it is a
one-shot or a loop is a separate question, so a sub-bass hit and a rolling bass line are
both bass.
| Family | Types |
|---|---|
| drums | kick, snare, clap, hat_closed, hat_open, cymbal (crash, ride), tom, percussion (incl. shakers), drum_loop (a full loop, with a kick), perc_loop (hats, shakers, percussion, no kick) |
| bass | bass (hits, loops, sub, reese), acid (303-style lines) |
| synth | lead, pluck, pad, stab (incl. chords and synth hits), keys (incl. piano) |
| vocals | vocal (hits and loops) |
| fx | sweep (risers and downlifters), impact, atmo (incl. noise), fx |
Gate labels (gate.txt): atonal, tonal.
Accuracy
Measured with 5-fold cross-validation split by sample pack: every number is for packs the model never saw in training, counting only files whose label came from their folder name.
| v1 (Aug 2026) | v2 | |
|---|---|---|
| Sample type, right first time | 68.3% (39 types) | 85.4% (22 types) |
| Sample type, right in the top three | β | 97.5% |
| Tonal gate | 93.3% (Create ML) | 98.6% |
Per type, v2, right first time / in the top three:
| Type | Top-1 | Top-3 | Type | Top-1 | Top-3 | |
|---|---|---|---|---|---|---|
| kick | 97.5% | 99.7% | lead | 76.8% | 97.6% | |
| drum_loop | 96.7% | 99.3% | stab | 74.6% | 92.6% | |
| vocal | 93.1% | 98.7% | pluck | 74.3% | 97.6% | |
| perc_loop | 92.4% | 98.2% | percussion | 72.2% | 96.1% | |
| cymbal | 92.3% | 98.7% | acid | 71.7% | 93.0% | |
| hat_closed | 90.0% | 99.5% | keys | 69.3% | 93.7% | |
| impact | 89.7% | 96.4% | fx | 62.0% | 93.5% | |
| pad | 88.7% | 96.7% | atmo | 61.7% | 91.5% | |
| bass | 88.2% | 98.0% | sweep | 57.6% | 94.3% | |
| clap | 87.7% | 98.3% | ||||
| hat_open | 85.4% | 98.3% | ||||
| snare | 81.9% | 97.3% | ||||
| tom | 78.3% | 97.7% |
The most common confusions are between neighbours a listener would also hesitate over: lead β pluck, bass β stab, sweep and atmo β fx, snare β clap. Show the top three rather than one answer where it matters.
Use
With muse
muse downloads this model the first time it needs it. A real run on a snare one-shot:
muse analyze snare.wav
# key β (atonal β suppressed by gate)
# class snare (0.49) Β· drums β or percussion 0.38, tom 0.03
From Swift
With swift-music-analysis, pointing at a folder holding this repository's files:
let classifier = try await SampleClassifier(directory: modelsFolder)
let verdict = try await classifier.classify(url: sampleURL)
verdict.label // "kick"
verdict.ranking.prefix(3) // the three likeliest, best first
verdict.tonal // probability that a key means something
From Python (Core AI runtime)
import asyncio, numpy as np, coreai.runtime as rt
options = rt.SpecializationOptions.from_preferred_compute_unit_kind(rt.ComputeUnitKind.gpu())
model = asyncio.run(rt.AIModel.load("muse-sample-types.aimodel", options))
classify = model.load_function("classify")
out = asyncio.run(classify(inputs={"input_features": rt.NDArray(mel)})) # mel: [1, 1, 1001, 64] float32
types = out["types"].numpy()[0]; gate = out["gate"].numpy()[0]
mel is CLAP's feature extractor output (ClapFeatureExtractor from
laion/clap-htsat-unfused, padding="repeatpad") on the first ten seconds of the sound at
48 kHz mono. mel-filters.f32 is its 513Γ64 slaney filterbank for implementations that
compute the spectrogram themselves.
Files
| File | What it is |
|---|---|
muse-sample-types.aimodel/ |
The model: entry point classify; in input_features [1, 1, 1001, 64]; out types [1, 22], gate [1, 2] (softmax probabilities) and embedding [1, 512] (L2-normalised) |
types.txt, gate.txt |
The labels of the two outputs, in order |
mel-filters.f32 |
513Γ64 slaney mel filterbank, little-endian float32, row-major |
Batch size is fixed at 1. Requires macOS 27 and Apple silicon (Core AI).
How it was made
- Encoder. laion/clap-htsat-unfused (LAION, Apache-2.0), weights unchanged. It won a contest on the same 10,322 samples, packs held out: CLAP 70.7% top-1 with a linear probe, against larger_clap_general 68.2%, AST-AudioSet 64.6% and MERT-v1-95M 58.1%.
- Data. A licensed library of electronic-music sample packs: 80,102 one-shots, loops and synth preset previews from 159 packs (construction-kit stems, demos and duplicates left out). Every file was fingerprinted on its first ten seconds, the same window the model hears in use.
- Labels. From folder and file names, mapped onto the 22 types. Then checked: a head trained without each file's pack predicted it, and the 4,163 files (5.5%) where it was confident in another type and gave the folder's type almost nothing were reviewed. A listening check found the folder name wrong in those cases (a sub-bass drop filed as FX, vocal chops filed as stabs), so they took the model's label. The accuracy above counts only files whose label came from their name.
- Heads. Two small GELU networks on the normalised embedding, 512 β 512 β 22 and 512 β 128 β 2, trained with label smoothing and a cap of 3,000 files per type per epoch (there are 15,848 kicks). The tonal gate's classes follow v1: pitched instruments are tonal, drums and percussion atonal, FX left out of its training as ambiguous.
- Packaging. Encoder and both heads exported together to Core AI; on the GPU the outputs match PyTorch to within 5Γ10β»β·.
Limitations
- Trained on electronic music sample packs (house, techno, trance, psytrance and neighbours). Acoustic instruments, orchestral sounds and field recordings sit outside it; guitar, brass, strings and choir were too rare to learn and have no label of their own.
- It hears the first ten seconds only.
- The FX family is the weakest because its types overlap in use; the top three is the better answer there.
- Rising versus falling sweeps, and one-shot versus loop, are not this model's job: CLAP barely hears direction, and both are measured separately in muse.
History
- v2 (October 2026): one asset with both heads; 22 types (from 39), cleaned labels, synth preset previews added, and the tonal gate moved onto CLAP. Supersedes the separate muse-tonal-gate.
- v1 (October 2026): CLAP encoder plus a separate 39-type head; tonal gate as a Create ML model.
Licence
Apache-2.0, following CLAP. The training audio is not distributed and cannot be recovered from these files.
Model tree for arraypress/muse-sample-types
Base model
laion/clap-htsat-unfused