NovelSep
NovelSep: Bridging Optimization-Based Separation and Deep Neural Refinement for Novelty Detection with Listenable Explanations
Koki Shoda, Jun Younes Louhi Kasahara, Qi An, and Atsushi Yamashita — The University of Tokyo.
Code and installation · Project page and audio examples
NovelSep separates an observed recording into normal and novel sound components. A normal dictionary first produces an optimization-based separation. NovelSep uses the encoded components to initialize a learned flow that refines both estimates. The energy of the estimated novel component provides the novelty score. The accompanying Normal Region Exclusion geometry filters training sound candidates that are too close to the normal reference region; this filtering stage is separate from inference.
This repository contains the learned adapters and supporting artifacts used in the accompanying research manuscript.
Released environments
| Subdirectory | Normal environment | Recording duration | Normal prompt |
|---|---|---|---|
airport |
TAU airport | 5.0 s | Background noise |
metro_station |
TAU metro station | 5.0 s | Background noise |
public_square |
TAU public square | 5.0 s | Background noise |
sonyc_ust_alert_signal |
SONYC-UST urban audio without alert signals | 10 s | Urban background noise without alert signals |
Each environment contains:
adapter.safetensors: NovelSep LoRA parameters, retaining the tensor values and names from the selected training checkpoint.nne_model.npzandnne_model.json: the matching Nonnegative Novelty Extraction (NNE) dictionary and portable inference metadata.normal_region_exclusion.npz: fitted Principal Component Analysis (PCA), projected normal reference embeddings, calibration distances, and the exclusion threshold. Recording identifiers and local filesystem paths are omitted.config.json: the prompt, audio settings, adapter settings, numerical precision requirements, domain-specific detection threshold, and artifact hashes.
The base SAM-Audio weights are downloaded separately from Meta's model repository. Access to that repository and agreement to its terms may be required. These adapters are intended for the pinned base revision recorded in config.json.
Usage
Install NovelSep and SAM-Audio following the code repository instructions. To retrieve the portable artifacts without executing any model code:
from huggingface_hub import snapshot_download
model_directory = snapshot_download(
repo_id="kokieto/NovelSep",
allow_patterns=["airport/*", "LICENSE", "README.md", "SHA256SUMS"],
)
Use the NovelSep class to load the adapter, its matching dictionary, and the pinned SAM-Audio base model:
import soundfile as sf
from novelsep import NovelSep
waveform, sample_rate = sf.read("input.wav", dtype="float32", always_2d=True)
separator = NovelSep.from_pretrained(
"kokieto/NovelSep", environment="airport", device="cuda"
)
result = separator.separate(waveform.T, sample_rate=int(sample_rate))
sf.write("normal.wav", result.normal, result.sample_rate, subtype="FLOAT")
sf.write("novel.wav", result.novel, result.sample_rate, subtype="FLOAT")
print(result.novelty_score, result.is_novel)
The NormalRegionExclusion class loads normal_region_exclusion.npz independently with NormalRegionExclusion.load(path). No training audio or full SAM-Audio checkpoint is included here.
Training and numerical settings
Each environment has a separate normal dictionary and adapter. Normal training recordings come from TAU Urban Acoustic Scenes 2019 or SONYC-UST. Novel training candidates come from the locally prepared FSD50K development audio and are filtered by Normal Region Exclusion. The selected checkpoint minimizes the fixed validation flow-matching loss. Detection thresholds maximize accuracy on synthetic tuning mixtures and are not chosen from final evaluation labels.
The adapter rank is 16 and its scaling parameter is 32. Audio preprocessing uses mono recordings at 16 kHz and peak normalization. SAM-Audio encoding and decoding operate at 48 kHz. Inference uses midpoint integration; the exact integration step is preserved in config.json. CUDA inference uses bfloat16 autocast for SAM-Audio, float32 without autocast for NNE, and enabled TensorFloat-32 (TF32). Changing this numerical path can change the separation output. The public implementation contains the required settings.
The Normal Region Exclusion geometry uses L2-normalized embeddings from Meta's PE_AV small model, PCA fitted on normal reference embeddings and development sound candidates, and mean distance to the 32 closest normal references. Its calibrated threshold is Q95 + lambda × (Q95 − Q50), with lambda = 1.0. Q95 and Q50 are the normal calibration distance quantiles. The stored geometry must be used with the same embedding model and preprocessing.
Intended use and limitations
The release supports experimental, noncommercial research in environmental sound separation and novelty detection. Normality is specific to each training environment. Detection thresholds are tied to the released preprocessing and tuning distributions; they should be validated again before use in a different environment. A separated waveform is a model estimate and can contain missing events, leakage, or artifacts. The novelty score does not identify an event class or establish the cause of a sound.
The SONYC-UST environment treats alert signals as novel and uses real urban recordings for evaluation. The TAU environments use synthetic mixtures for evaluation. Performance in other environments, recording devices, durations, and sound mixtures is not established by this release. See the manuscript and project page for the evaluation protocol; selected audio examples are illustrative and are not an unbiased performance estimate.
GPU inference depends on the upstream SAM-Audio implementation and compatible PyTorch, torchaudio, and CUDA libraries. The adapter files use the NovelSep loading code and are not generic PEFT checkpoints. A CPU execution path may be useful for inspection but is not expected to reproduce the CUDA numerical path exactly.
Licenses and acknowledgments
The adapters are derived from SAM-Audio by Meta and are distributed subject to the SAM License, reproduced without modification. This release acknowledges Meta's SAM-Audio models and software. The base model is not redistributed here. The presence of a model on the Hub does not replace or remove upstream model or dataset terms.
Training data have separate licenses and attribution requirements:
- TAU Urban Acoustic Scenes 2019, by Toni Heittola, Annamaria Mesaros, Tuomas Virtanen, and the Audio Research Group at Tampere University, permits experimental and noncommercial use under its custom license. See the official dataset and documentation. Commercial permission is not granted by this model card.
- SONYC-UST, by Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez, Yu Wang, Ho-Hsiang Wu, Vincent Lostanlen, Magdalena Fuentes, Graham Dove, Charlie Mydlarz, Justin Salamon, Oded Nov, and Juan Pablo Bello, is provided under CC BY 4.0, as stated in its official README.
- FSD50K, by Eduardo Fonseca and collaborators, has a dataset-level CC BY license and per-recording Creative Commons licenses, including noncommercial terms for some recordings. Consult the official license information and recording metadata for the applicable terms.
- PE_AV by Meta provides the audio embeddings used for Normal Region Exclusion. The PE_AV model itself is not included in this release; consult its model repository for access and license terms.
Integrity and provenance
SHA256SUMS covers every release file except the checksum list itself. Run sha256sum -c SHA256SUMS after downloading the complete repository. Adapter tensors are verified byte-for-byte against the selected source checkpoints during export. NNE dictionary archives are copied unchanged.
In config.json, nne_sha256 is the verified original training identity: a hash of the dictionary archive followed by the canonical original metadata. The original metadata contains machine-specific provenance and is not distributed. nne_file_sha256 and nne_metadata_sha256 independently verify the distributed dictionary archive and sanitized metadata. source_checkpoint_sha256 identifies the original training checkpoint, and adapter_sha256 identifies the exported safe tensor file. Geometry contains only a numeric allowlist and portable settings. The reproducible export utility is scripts/export_models.py.
Model tree for kokieto/NovelSep
Base model
facebook/sam-audio-small