SAGE: Semantic Audio Generative Encoder
SAGE is a compact variational autoencoder for stereo music at 44.1 kHz. A Swin Transformer V2 encoder and decoder operate on the STFT, and the latent is shaped by distilling the embeddings of LAION-CLAP. The model has 104.6M parameters and compresses a stereo waveform ×64 into a 16-channel latent at 86 frames per second (one latent frame per 512 samples).
- Paper: arXiv 2609.32755
- Code: github.com/francescobrigante/SAGE
- Project page: sage-music.pages.dev
Checkpoint
| File | Size | SHA-256 |
|---|---|---|
SAGE_FTe992.ckpt (EMA weights, epoch 992) |
402 MiB | dd87d01eaee88ca92f96c80ffe0e504a1d7cbc02e4271d8c48033df591d6ac99 |
The checkpoint holds the model configuration and the EMA weights of the decoder fine-tuning (phase 2, epoch 992), the model evaluated in the paper.
Usage
Install the code (Python 3.11, uv):
git clone https://github.com/francescobrigante/SAGE.git && cd SAGE
uv sync
import torch
from huggingface_hub import hf_hub_download
from sage import SAGE
ckpt = hf_hub_download("francescobrigante/SAGE", "SAGE_FTe992.ckpt")
codec = SAGE.from_checkpoint(ckpt) # CUDA if available
wav = torch.randn(1, 2, 10 * codec.sample_rate) # (B, 2, N) at 44.1 kHz
rec = codec.reconstruct(wav) # encode + decode, same length as the input
padded, n = codec.pad(wav) # right-pad to the model's frame grid
z = codec.encode(padded, deterministic=True) # posterior mean, (B, 16, frames)
y = codec.decode(z, target_length=padded.shape[-1])[..., :n] # back to audio, input length
encode samples the latent unless deterministic=True. Training, evaluation and the commands
that reproduce every result of the paper are in the
repository.
Training data
About 10.5K hours of 44.1 kHz stereo music from three public corpora: the FMA-full training split, the MTG-Jamendo training partition (without the tracks of the Song Describer Dataset) and M4Singer. Two phases on 16 A100 GPUs: pretraining from scratch for 500 epochs, then decoder fine-tuning with the encoder frozen for 992 epochs (1,536 and 3,043 GPU-hours).
Intended use and limitations
SAGE is meant for research on music: a latent space for generative models (e.g. latent diffusion) and semantic analysis, and audio reconstruction. It is trained on music and solo singing; speech and general audio are out of domain. The weights are released under CC BY-NC-SA 4.0, following the non-commercial terms of part of the training data, and may not be used commercially.
Citation
@misc{brigante2026sage,
title = {{SAGE}: Semantic Audio Generative Encoder},
author = {Brigante, Francesco and Cerovaz, Luca and Marincione, Davide and Strano, Giorgio and
Zhou, Luca and Rodol{\`a}, Emanuele and Mancusi, Michele},
year = {2026},
eprint = {2609.32755},
archivePrefix = {arXiv},
primaryClass = {cs.SD},
url = {https://arxiv.org/abs/2609.32755}
}