Higgs Audio V2 Tokenizer
Hybrid semantic + acoustic speech tokenizer. Encodes audio to discrete codes and decodes discrete codes to a 24 kHz waveform.
Model Details
| Property |
Value |
| Model type |
higgs_audio_v2_tokenizer |
| Architecture |
HiggsAudioV2TokenizerModel |
| Developed by |
Higgs audio team |
| License |
Apache-2.0 |
| Language |
English |
Architecture
| Component |
Description |
| Semantic branch |
7-layer strided convolutional feature extractor followed by a 12-layer transformer encoder (hidden size 768, 12 attention heads, intermediate size 3072); 16 kHz input, downsample factor 320, 50 frames/s |
| Acoustic branch |
DAC-style convolutional encoder and decoder (hidden size 256, decoder hidden size 1024); downsampling ratios 8, 5, 4, 2, 3; hop length 960 at 24 kHz, 25 frames/s |
| Quantiser |
Residual vector quantiser, 8 codebooks, codebook size 1024, codebook dimension 64, no quantiser dropout |
| Fusion |
Linear projections fc, fc1, fc2 combine semantic and acoustic features before quantisation |
| Target bitrates |
0.5, 1, 1.5, 2 kbps |
Uses
Direct Use
- Encoding speech into discrete codes for analysis, compression, or speech language model input.
- Decoding discrete codes to 24 kHz audio.
Out-of-Scope Use
- Speaker identification.
- Voice cloning without consent.
- Music or general sound effects.
Bias, Risks, and Limitations
- English-focused model.
- Reconstruction quality varies with accent and channel conditions.
Recommendations
- Measure reconstruction on target accents and channel conditions before using the codes.
- Evaluate decoded audio with a speech recognizer when intelligibility matters.
Model Files and Format
| Property |
Value |
| File |
model.safetensors |
| Format |
safetensors, single file, no sharding |
| File size |
805,665,628 bytes (805.7 MB, 768.3 MiB) |
| Tensors |
527 |
| Tensor dtype |
float32 |
| Parameters |
201,400,553 |
| Header size |
63,408 bytes |
| Metadata |
{"format": "pt"} |
| SHA-256 |
51c9b982f4e0d26907eb419ce2346a97b24583f7d30057c585f0847a894b8b6b |
Parameter Split
| Module |
Parameters |
semantic_model.* |
94,370,944 |
acoustic_encoder.* |
51,314,496 |
acoustic_decoder.* |
20,238,113 |
encoder_semantic.* |
14,747,136 |
decoder_semantic.* |
16,516,608 |
quantizer.* |
2,114,056 |
fc.* |
1,049,600 |
fc1.* |
787,200 |
fc2.* |
262,400 |
Companion Files
| File |
Description |
config.json |
Model configuration |
preprocessor_config.json |
Feature extractor configuration |
LICENSE |
Apache-2.0 license text |
WEIGHTS.md |
Weight file size and checksum |
Training Details
Training Data
| Property |
Value |
| Total audio |
14,300 hours |
| Language |
English |
| License |
openly licensed |
| Sources |
read audiobooks, broadcast news, conversational podcasts, telephone-style recordings |
| Telephone-style audio |
~20% of total |
| Clip length |
2 to 15 seconds |
| Filtering |
overlapping speakers, clipping |
Training Procedure
| Hyperparameter |
Value |
| Weight precision |
float32 |
| Mixed precision |
bfloat16 |
| Optimiser |
AdamW |
| Peak learning rate |
2e-4 |
| Learning rate schedule |
cosine decay |
| Warm-up steps |
4,000 |
| Batch size |
64 clips (~6.4 minutes of audio) |
| Training steps |
310,000 |
| Reconstruction loss |
reconstruction + multi-scale spectral loss |
| Codebook loss weight |
1.0 |
| Commitment loss weight |
0.25 |
| Semantic distillation loss weight |
0.5 |
| Best validation step |
287,000 |
Checkpoint Size and Throughput
| Property |
Value |
| Checkpoint size |
805.7 MB |
| Training throughput |
~38 audio-hours per wall-clock hour |
| Selected checkpoint |
best validation step, 287,000 |
Evaluation
Testing Data
| Property |
Value |
| Utterances |
1,850 |
| Audio |
~4.1 hours |
| Speakers |
46, disjoint from training |
| Clean studio clips |
925 |
| Telephone-style clips |
925, simulated narrowband channel effects |
Metrics
| Metric |
Definition |
Direction |
| Mel distance |
L1 distance between log-mel spectrograms |
lower is better |
| WER |
word error rate of a fixed reference recognizer on decoded audio |
lower is better |
| Speaker similarity |
cosine similarity of speaker embeddings |
higher is better |
Results
| Condition |
Mel distance |
WER (decoded) |
Speaker similarity |
| Clean studio |
0.42 |
4.8% |
0.91 |
| Telephone-style |
0.67 |
11.3% |
0.84 |
Reference Deltas
| Condition |
WER increase vs. uncoded audio |
| Clean studio |
1.1 percentage points |
| Telephone-style |
3.9 percentage points |
Codebook Usage
| Property |
Value |
| Codebook entries used |
~93% per codebook on the test set |
| First two codebooks |
most speech content |
| Later codebooks |
timbre and channel noise refinement |
Environmental Impact
| Property |
Value |
| Hardware |
4 x 80 GB accelerators |
| Hours used |
212 |
| Cloud provider |
regional cloud provider (unnamed) |
| Compute region |
European region |
| Carbon emitted |
~46 kg CO2eq |
Carbon estimates can be reproduced with the Machine Learning Impact calculator from Lacoste et al. (2019).
Technical Specifications
| Property |
Value |
| Model type |
higgs_audio_v2_tokenizer |
| Architecture |
HiggsAudioV2TokenizerModel |
| Acoustic model type |
dac |
| Semantic model type |
hubert |
| Semantic feature extractor layers |
7 |
| Semantic transformer layers |
12 |
| Codebook loss weight |
1.0 |
| Commitment loss weight |
0.25 |
| Codebooks |
8 |
| Codebook size |
1024 |
| Codebook dimension |
64 |
| Quantiser dropout |
0 |
| Sample rate |
24000 |
| Semantic sample rate |
16000 |
| Serialisation |
safetensors |
| Framework |
Transformers 5.3.0.dev0, PyTorch |
Model Card Contact
Open an issue on the model repository.