hybridcodec_scratch_25hz

HybridCodec is a speech codec that represents audio with a discrete semantic stream and a continuous residual stream. This repository contains the Scratch quantizer configuration, operating at a semantic token rate of 25 Hz. The codec uses the pretrained FocalCodec 50 Hz encoder and vocoder with the matching HybridCodec quantizer weights.

Intended use

Encode mono speech audio into HybridCodec semantic and residual streams, and decode those streams back into audio. The model configuration expects audio at 16 kHz. Use the same repository for encoding and decoding.

Files

  • hyperparams.yaml: model architecture and inference configuration.
  • weights/: quantizer parameters in safetensors format.

Install

Install HybridCodec from GitHub:

pip install "git+https://github.com/hybridcodec/hybridcodec.git"

Encode and decode audio

Load this model from the Hugging Face Hub, encode a WAV file, and write the reconstruction:

import soundfile as sf
import torch
from hybridcodec.inference import HybridCodec

codec = HybridCodec.from_hparams(source="hybridcodec/hybridcodec_scratch_25hz")
waveform = codec.load_audio("sample.wav").squeeze(-1).to(codec.device)
lengths = torch.ones(1, device=codec.device)
semantic, residual = codec.encode(waveform.unsqueeze(0), lengths)
reconstruction = codec.decode(semantic, residual)
sf.write("reconstructed.wav", reconstruction[0].cpu().numpy(), 16000)

Paper

See HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for hybridcodec/hybridcodec_scratch_25hz