muse sample types

Name any audio sample on a Mac, on-device. One Core AI model that listens to a sample once and answers two questions:

  1. What is it? One of 22 sample types, such as kick, snare, bass, pluck, pad, vocal or sweep, with the runner-up types and their probabilities.
  2. Does a musical key mean anything here? Tonal or atonal, so a kick drum is never labelled "C minor".

It also returns the sound's 512-number CLAP embedding, for "find sounds like this" search. It is the model behind the class and tonal columns of the muse command-line tool, which downloads it on first use.

Labels

22 types in five families (types.txt). A label says what the sound is; whether it is a one-shot or a loop is a separate question, so a sub-bass hit and a rolling bass line are both bass.

Family Types
drums kick, snare, clap, hat_closed, hat_open, cymbal (crash, ride), tom, percussion (incl. shakers), drum_loop (a full loop, with a kick), perc_loop (hats, shakers, percussion, no kick)
bass bass (hits, loops, sub, reese), acid (303-style lines)
synth lead, pluck, pad, stab (incl. chords and synth hits), keys (incl. piano)
vocals vocal (hits and loops)
fx sweep (risers and downlifters), impact, atmo (incl. noise), fx

Gate labels (gate.txt): atonal, tonal.

Accuracy

Measured with 5-fold cross-validation split by sample pack: every number is for packs the model never saw in training, counting only files whose label came from their folder name.

v1 (Aug 2026) v2
Sample type, right first time 68.3% (39 types) 85.4% (22 types)
Sample type, right in the top three β€” 97.5%
Tonal gate 93.3% (Create ML) 98.6%

Per type, v2, right first time / in the top three:

Type Top-1 Top-3 Type Top-1 Top-3
kick 97.5% 99.7% lead 76.8% 97.6%
drum_loop 96.7% 99.3% stab 74.6% 92.6%
vocal 93.1% 98.7% pluck 74.3% 97.6%
perc_loop 92.4% 98.2% percussion 72.2% 96.1%
cymbal 92.3% 98.7% acid 71.7% 93.0%
hat_closed 90.0% 99.5% keys 69.3% 93.7%
impact 89.7% 96.4% fx 62.0% 93.5%
pad 88.7% 96.7% atmo 61.7% 91.5%
bass 88.2% 98.0% sweep 57.6% 94.3%
clap 87.7% 98.3%
hat_open 85.4% 98.3%
snare 81.9% 97.3%
tom 78.3% 97.7%

The most common confusions are between neighbours a listener would also hesitate over: lead ↔ pluck, bass ↔ stab, sweep and atmo ↔ fx, snare ↔ clap. Show the top three rather than one answer where it matters.

Use

With muse

muse downloads this model the first time it needs it. A real run on a snare one-shot:

muse analyze snare.wav
#   key     β€”  (atonal β€” suppressed by gate)
#   class   snare  (0.49)  Β· drums  β€” or percussion 0.38, tom 0.03

From Swift

With swift-music-analysis, pointing at a folder holding this repository's files:

let classifier = try await SampleClassifier(directory: modelsFolder)
let verdict = try await classifier.classify(url: sampleURL)
verdict.label            // "kick"
verdict.ranking.prefix(3) // the three likeliest, best first
verdict.tonal            // probability that a key means something

From Python (Core AI runtime)

import asyncio, numpy as np, coreai.runtime as rt

options = rt.SpecializationOptions.from_preferred_compute_unit_kind(rt.ComputeUnitKind.gpu())
model = asyncio.run(rt.AIModel.load("muse-sample-types.aimodel", options))
classify = model.load_function("classify")
out = asyncio.run(classify(inputs={"input_features": rt.NDArray(mel)}))   # mel: [1, 1, 1001, 64] float32
types = out["types"].numpy()[0]; gate = out["gate"].numpy()[0]

mel is CLAP's feature extractor output (ClapFeatureExtractor from laion/clap-htsat-unfused, padding="repeatpad") on the first ten seconds of the sound at 48 kHz mono. mel-filters.f32 is its 513Γ—64 slaney filterbank for implementations that compute the spectrogram themselves.

Files

File What it is
muse-sample-types.aimodel/ The model: entry point classify; in input_features [1, 1, 1001, 64]; out types [1, 22], gate [1, 2] (softmax probabilities) and embedding [1, 512] (L2-normalised)
types.txt, gate.txt The labels of the two outputs, in order
mel-filters.f32 513Γ—64 slaney mel filterbank, little-endian float32, row-major

Batch size is fixed at 1. Requires macOS 27 and Apple silicon (Core AI).

How it was made

  1. Encoder. laion/clap-htsat-unfused (LAION, Apache-2.0), weights unchanged. It won a contest on the same 10,322 samples, packs held out: CLAP 70.7% top-1 with a linear probe, against larger_clap_general 68.2%, AST-AudioSet 64.6% and MERT-v1-95M 58.1%.
  2. Data. A licensed library of electronic-music sample packs: 80,102 one-shots, loops and synth preset previews from 159 packs (construction-kit stems, demos and duplicates left out). Every file was fingerprinted on its first ten seconds, the same window the model hears in use.
  3. Labels. From folder and file names, mapped onto the 22 types. Then checked: a head trained without each file's pack predicted it, and the 4,163 files (5.5%) where it was confident in another type and gave the folder's type almost nothing were reviewed. A listening check found the folder name wrong in those cases (a sub-bass drop filed as FX, vocal chops filed as stabs), so they took the model's label. The accuracy above counts only files whose label came from their name.
  4. Heads. Two small GELU networks on the normalised embedding, 512 β†’ 512 β†’ 22 and 512 β†’ 128 β†’ 2, trained with label smoothing and a cap of 3,000 files per type per epoch (there are 15,848 kicks). The tonal gate's classes follow v1: pitched instruments are tonal, drums and percussion atonal, FX left out of its training as ambiguous.
  5. Packaging. Encoder and both heads exported together to Core AI; on the GPU the outputs match PyTorch to within 5Γ—10⁻⁷.

Limitations

  • Trained on electronic music sample packs (house, techno, trance, psytrance and neighbours). Acoustic instruments, orchestral sounds and field recordings sit outside it; guitar, brass, strings and choir were too rare to learn and have no label of their own.
  • It hears the first ten seconds only.
  • The FX family is the weakest because its types overlap in use; the top three is the better answer there.
  • Rising versus falling sweeps, and one-shot versus loop, are not this model's job: CLAP barely hears direction, and both are measured separately in muse.

History

  • v2 (October 2026): one asset with both heads; 22 types (from 39), cleaned labels, synth preset previews added, and the tonal gate moved onto CLAP. Supersedes the separate muse-tonal-gate.
  • v1 (October 2026): CLAP encoder plus a separate 39-type head; tonal gate as a Create ML model.

Licence

Apache-2.0, following CLAP. The training audio is not distributed and cannot be recovered from these files.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for arraypress/muse-sample-types

Finetuned
(2)
this model