SAGE: Semantic Audio Generative Encoder

SAGE is a compact variational autoencoder for stereo music at 44.1 kHz. A Swin Transformer V2 encoder and decoder operate on the STFT, and the latent is shaped by distilling the embeddings of LAION-CLAP. The model has 104.6M parameters and compresses a stereo waveform ×64 into a 16-channel latent at 86 frames per second (one latent frame per 512 samples).

Checkpoint

File Size SHA-256
SAGE_FTe992.ckpt (EMA weights, epoch 992) 402 MiB dd87d01eaee88ca92f96c80ffe0e504a1d7cbc02e4271d8c48033df591d6ac99

The checkpoint holds the model configuration and the EMA weights of the decoder fine-tuning (phase 2, epoch 992), the model evaluated in the paper.

Usage

Install the code (Python 3.11, uv):

git clone https://github.com/francescobrigante/SAGE.git && cd SAGE
uv sync
import torch
from huggingface_hub import hf_hub_download
from sage import SAGE

ckpt = hf_hub_download("francescobrigante/SAGE", "SAGE_FTe992.ckpt")
codec = SAGE.from_checkpoint(ckpt)                               # CUDA if available
wav = torch.randn(1, 2, 10 * codec.sample_rate)                  # (B, 2, N) at 44.1 kHz

rec = codec.reconstruct(wav)                                     # encode + decode, same length as the input

padded, n = codec.pad(wav)                                       # right-pad to the model's frame grid
z = codec.encode(padded, deterministic=True)                     # posterior mean, (B, 16, frames)
y = codec.decode(z, target_length=padded.shape[-1])[..., :n]     # back to audio, input length

encode samples the latent unless deterministic=True. Training, evaluation and the commands that reproduce every result of the paper are in the repository.

Training data

About 10.5K hours of 44.1 kHz stereo music from three public corpora: the FMA-full training split, the MTG-Jamendo training partition (without the tracks of the Song Describer Dataset) and M4Singer. Two phases on 16 A100 GPUs: pretraining from scratch for 500 epochs, then decoder fine-tuning with the encoder frozen for 992 epochs (1,536 and 3,043 GPU-hours).

Intended use and limitations

SAGE is meant for research on music: a latent space for generative models (e.g. latent diffusion) and semantic analysis, and audio reconstruction. It is trained on music and solo singing; speech and general audio are out of domain. The weights are released under CC BY-NC-SA 4.0, following the non-commercial terms of part of the training data, and may not be used commercially.

Citation

@misc{brigante2026sage,
  title         = {{SAGE}: Semantic Audio Generative Encoder},
  author        = {Brigante, Francesco and Cerovaz, Luca and Marincione, Davide and Strano, Giorgio and
                   Zhou, Luca and Rodol{\`a}, Emanuele and Mancusi, Michele},
  year          = {2026},
  eprint        = {2609.32755},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SD},
  url           = {https://arxiv.org/abs/2609.32755}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for francescobrigante/SAGE