Higgs Audio V2 Tokenizer

Hybrid semantic + acoustic speech tokenizer. Encodes audio to discrete codes and decodes discrete codes to a 24 kHz waveform.

Model Details

Property Value
Model type higgs_audio_v2_tokenizer
Architecture HiggsAudioV2TokenizerModel
Developed by Higgs audio team
License Apache-2.0
Language English

Architecture

Component Description
Semantic branch 7-layer strided convolutional feature extractor followed by a 12-layer transformer encoder (hidden size 768, 12 attention heads, intermediate size 3072); 16 kHz input, downsample factor 320, 50 frames/s
Acoustic branch DAC-style convolutional encoder and decoder (hidden size 256, decoder hidden size 1024); downsampling ratios 8, 5, 4, 2, 3; hop length 960 at 24 kHz, 25 frames/s
Quantiser Residual vector quantiser, 8 codebooks, codebook size 1024, codebook dimension 64, no quantiser dropout
Fusion Linear projections fc, fc1, fc2 combine semantic and acoustic features before quantisation
Target bitrates 0.5, 1, 1.5, 2 kbps

Uses

Direct Use

  • Encoding speech into discrete codes for analysis, compression, or speech language model input.
  • Decoding discrete codes to 24 kHz audio.

Out-of-Scope Use

  • Speaker identification.
  • Voice cloning without consent.
  • Music or general sound effects.

Bias, Risks, and Limitations

  • English-focused model.
  • Reconstruction quality varies with accent and channel conditions.

Recommendations

  • Measure reconstruction on target accents and channel conditions before using the codes.
  • Evaluate decoded audio with a speech recognizer when intelligibility matters.

Model Files and Format

Property Value
File model.safetensors
Format safetensors, single file, no sharding
File size 805,665,628 bytes (805.7 MB, 768.3 MiB)
Tensors 527
Tensor dtype float32
Parameters 201,400,553
Header size 63,408 bytes
Metadata {"format": "pt"}
SHA-256 51c9b982f4e0d26907eb419ce2346a97b24583f7d30057c585f0847a894b8b6b

Parameter Split

Module Parameters
semantic_model.* 94,370,944
acoustic_encoder.* 51,314,496
acoustic_decoder.* 20,238,113
encoder_semantic.* 14,747,136
decoder_semantic.* 16,516,608
quantizer.* 2,114,056
fc.* 1,049,600
fc1.* 787,200
fc2.* 262,400

Companion Files

File Description
config.json Model configuration
preprocessor_config.json Feature extractor configuration
LICENSE Apache-2.0 license text
WEIGHTS.md Weight file size and checksum

Training Details

Training Data

Property Value
Total audio 14,300 hours
Language English
License openly licensed
Sources read audiobooks, broadcast news, conversational podcasts, telephone-style recordings
Telephone-style audio ~20% of total
Clip length 2 to 15 seconds
Filtering overlapping speakers, clipping

Training Procedure

Hyperparameter Value
Weight precision float32
Mixed precision bfloat16
Optimiser AdamW
Peak learning rate 2e-4
Learning rate schedule cosine decay
Warm-up steps 4,000
Batch size 64 clips (~6.4 minutes of audio)
Training steps 310,000
Reconstruction loss reconstruction + multi-scale spectral loss
Codebook loss weight 1.0
Commitment loss weight 0.25
Semantic distillation loss weight 0.5
Best validation step 287,000

Checkpoint Size and Throughput

Property Value
Checkpoint size 805.7 MB
Training throughput ~38 audio-hours per wall-clock hour
Selected checkpoint best validation step, 287,000

Evaluation

Testing Data

Property Value
Utterances 1,850
Audio ~4.1 hours
Speakers 46, disjoint from training
Clean studio clips 925
Telephone-style clips 925, simulated narrowband channel effects

Metrics

Metric Definition Direction
Mel distance L1 distance between log-mel spectrograms lower is better
WER word error rate of a fixed reference recognizer on decoded audio lower is better
Speaker similarity cosine similarity of speaker embeddings higher is better

Results

Condition Mel distance WER (decoded) Speaker similarity
Clean studio 0.42 4.8% 0.91
Telephone-style 0.67 11.3% 0.84

Reference Deltas

Condition WER increase vs. uncoded audio
Clean studio 1.1 percentage points
Telephone-style 3.9 percentage points

Codebook Usage

Property Value
Codebook entries used ~93% per codebook on the test set
First two codebooks most speech content
Later codebooks timbre and channel noise refinement

Environmental Impact

Property Value
Hardware 4 x 80 GB accelerators
Hours used 212
Cloud provider regional cloud provider (unnamed)
Compute region European region
Carbon emitted ~46 kg CO2eq

Carbon estimates can be reproduced with the Machine Learning Impact calculator from Lacoste et al. (2019).

Technical Specifications

Property Value
Model type higgs_audio_v2_tokenizer
Architecture HiggsAudioV2TokenizerModel
Acoustic model type dac
Semantic model type hubert
Semantic feature extractor layers 7
Semantic transformer layers 12
Codebook loss weight 1.0
Commitment loss weight 0.25
Codebooks 8
Codebook size 1024
Codebook dimension 64
Quantiser dropout 0
Sample rate 24000
Semantic sample rate 16000
Serialisation safetensors
Framework Transformers 5.3.0.dev0, PyTorch

Model Card Contact

Open an issue on the model repository.

Downloads last month
10
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for BobJames/open-higgs-tokenizer