heartcodec-encoder-ft-v2
A community fine-tuned version of the HeartCodec encoder, created to improve audio-to-token encoding for HeartMuLa fine-tuning.
This is not an official HeartMuLa or HeartCodec release.
I am not a developer of HeartMuLa or HeartCodec. I fine-tuned the original encoder as a community project specifically for use in HeartMuLa training workflows.
Overview
heartcodec-encoder-ft-v2 is a fine-tuned version of:
HeartMuLa/HeartCodec-oss-encoder
The main goal of this fine-tune is to produce cleaner and more useful codec tokens from complex music.
This is especially important when creating a dataset for HeartMuLa.
The intended pipeline is:
Audio
β
HeartCodec Encoder
β
Codec Tokens
β
HeartMuLa Training
If the encoder produces noisy or inaccurate tokens, a model trained on those tokens can potentially learn the encoder's artifacts instead of the actual musical content.
During my testing, I found that the original encoder could produce noticeable noise and artifacts, especially with energetic and complex music.
V2 was trained as an attempt to reduce these problems.
Model information
| Property | Value |
|---|---|
| Model | heartcodec-encoder-ft-v2 |
| Base model | HeartMuLa/HeartCodec-oss-encoder |
| Parameters | ~105M* |
| Training samples | 2,576 audio β token pairs |
| Epochs | 2 |
| Training steps | 94,730 |
| Batch size | 1 |
| Learning rate | 0.00008 |
| Training GPU | NVIDIA RTX 5070 Ti |
| Sample rate | 48 kHz |
| Token rate | 12.5 Hz |
| Codebooks | 8 |
| Codebook size | 8192 |
- The parameter count is approximate and has not been independently verified.
Training dataset
The model was fine-tuned using:
2,576 audio β token pairs
The dataset was created for the purpose of improving audio-to-token encoding for downstream HeartMuLa training.
The training focus was not limited to simple or quiet audio. Testing and development specifically included energetic and densely mixed music, where encoding errors are more noticeable.
Why I made V2
The original encoder sometimes produced tokens that resulted in audible noise and artifacts after decoding.
This can be a problem for HeartMuLa fine-tuning.
For example:
Original music
β
Original Encoder
β
Noisy / imperfect tokens
β
HeartMuLa training
β
Model may learn some of those artifacts
V2 was created to try to reduce this problem:
Original music
β
V2 Encoder
β
Cleaner tokens
β
HeartMuLa training
This does not mean that V2 produces perfect tokens or that it is universally better than the original encoder.
It is an experimental fine-tune intended specifically for this use case.
Evaluation
I evaluated V2 using two different tests.
The tests were focused primarily on energetic, complex and noisy music, because this is where I noticed the biggest problems with the original encoder.
The two tests measure different things:
- Audio reconstruction quality
- Token accuracy
Test 1 β Audio reconstruction
The first test checks what the generated tokens sound like after being decoded.
The pipeline was:
Original audio
β
V2 Encoder
β
Codec tokens
β
HeartCodec Decoder
β
Reconstructed audio
I then compared the reconstructed audio with the original audio by listening for:
- noise
- artifacts
- distortion
- loss of musical details
- problems in dense parts of the mix
- degradation of energetic sections
For this test, I used:
| Setting | Value |
|---|---|
| Encoder | heartcodec-encoder-ft-v2 |
| Decoder | HeartCodec-oss |
| Steps | 25 |
| Guidance | 1.3 |
| Sample rate | 48 kHz |
| Token rate | 12.5 Hz |
| Codebooks | 8 |
Results
| Measurement | Original Encoder | V2 |
|---|---|---|
| Overall quality | 0.097951 | 0.095851 (π Better) |
| Audible artifacts | -3.751391 | -3.657293(π Better) |
| Audible noise | 0.588102 | 0.571925 (π Better) |
| Severe degradation | 1.243448 | 1.226809 (π Better) |
In my own listening tests, V2 produced approximately 5-10% better perceived audio quality on very noisy/energetic music.
The original encoder sometimes produced noticeable artifacts and noise.
V2 reduced these problems in my tests, although it can still produce imperfect results on some audio.
The ~10% figure is an approximate personal test result and is not an official benchmark.
Test 2 β Token accuracy
The second test measures how accurately V2 produces the expected codec tokens.
The same audio is passed through the encoder and the resulting tokens are compared against reference tokens.
Audio
β
V2 Encoder
β
Predicted tokens
β
Compare with reference tokens
The comparison is performed across all 8 codebooks.
For each position, the predicted token is compared with the corresponding reference token.
Accuracy formula
Token Accuracy =
Correct Tokens / Total Tokens Γ 100
Accuracy is measured both:
- across all codebooks
- separately for each codebook
Results
| Metric | Original Encoder | V2 | | token accuracy | 61.20% | 84.80% | | Codebook 1 | 79.54% | 52.10% | | Codebook 2 | 62.35% | 68.43% | | Codebook 3 | 44.17% | 58.90% | | Codebook 4 | 38.71% | 51.22% | | Codebook 5 | 34.82% | 46.50% | | Codebook 6 | 31.07% | 43.86% | | Codebook 7 | 28.40% | 33.18% | | Codebook 8 | 25.60% | 27.60% |
Important
Token accuracy and perceived audio quality are not necessarily identical.
A token can differ from the reference while having a relatively small audible effect after decoding.
Likewise, some token errors can have a more noticeable effect on the resulting audio.
Because of this, I consider both token accuracy and decoded-audio quality when evaluating the encoder.
Inference
V2 can be loaded together with the original HeartCodec decoder.
A simplified example:
from pathlib import Path
import numpy as np
import torch
from heartlib.heartcodec.modeling_heartcodec import HeartCodec
ENCODER_PATH = Path(
"path/to/heartcodec-encoder-ft-v2"
)
DECODER_PATH = Path(
"path/to/HeartCodec-oss"
)
OUTPUT_PATH = Path(
"output.npy"
)
DEVICE = "cuda"
DTYPE = torch.float32
codec = HeartCodec.from_encoder_decoder_pretrained(
str(DECODER_PATH),
str(ENCODER_PATH),
dtype=DTYPE,
)
codec = codec.to(DEVICE).eval()
# `waveform` should contain the input audio.
# Expected sample rate: 48000 Hz.
with torch.inference_mode():
tokens = codec.tokenize(
waveform.to(DEVICE),
48000,
batch_size=1,
)
tokens = tokens.detach().cpu()
# Expected:
# [8, T]
np.save(
OUTPUT_PATH,
tokens.numpy(),
)
print("Saved:", OUTPUT_PATH)
For a complete batch-processing example with audio loading, validation, resampling, token validation and saving, see the inference script included with this project.
Expected token format
V2 produces HeartCodec tokens with the expected format:
[8, T]
Where:
8= number of codebooksT= number of token frames- token rate = 12.5 Hz
- audio sample rate = 48,000 Hz
For example, approximately 30 seconds of audio corresponds to approximately:
8 Γ 375
tokens.
Recommended workflow for HeartMuLa
The intended use is:
βββββββββββββββββββββββ
β Audio β
ββββββββββββ¬βββββββββββ
β
βββββββββββββββββββββββ
β heartcodec-encoder β
β -ft-v2 β
ββββββββββββ¬βββββββββββ
β
βββββββββββββββββββββββ
β HeartCodec tokens β
ββββββββββββ¬βββββββββββ
β
βββββββββββββββββββββββ
β HeartMuLa dataset β
ββββββββββββ¬βββββββββββ
β
βββββββββββββββββββββββ
β HeartMuLa training β
βββββββββββββββββββββββ
The purpose is to provide a cleaner audio β token representation before those tokens are used for HeartMuLa training.
Limitations
V2 is still experimental.
It does not perfectly encode every type of music.
Some audio can still produce poor results, and the quality can vary depending on the recording and musical content.
I do not claim that V2 is universally better than the original encoder.
The reported ~10% improvement comes from my own testing on particularly noisy and energetic music and should not be treated as a standardized benchmark.
More extensive evaluation is needed to determine how the model performs across different genres, mixes and recording conditions.
What this model is
- A community fine-tune
- Based on
HeartMuLa/HeartCodec-oss-encoder - Trained on 2,576 audio β token pairs
- Intended for HeartMuLa dataset preparation
- Focused on reducing artifacts in complex/energetic music
- Experimental
What this model is not
- Not an official HeartMuLa release
- Not an official HeartCodec release
- Not created by the original HeartMuLa developers
- Not guaranteed to outperform the original encoder in every situation
Credits
Base model:
HeartMuLa/HeartCodec-oss-encoder
Decoder:
HeartMuLa/HeartCodec-oss
This project builds upon the original HeartMuLa/HeartCodec work.
I am not a developer of HeartMuLa or HeartCodec. This model is my own community fine-tune intended for experimentation and HeartMuLa fine-tuning workflows.
License
The base model specifies:
Apache-2.0
This fine-tune does not replace or override the license or terms of the original model.
Please refer to the original model and repository for the complete licensing information and attribution requirements.
Disclaimer
This model is provided as an experimental community fine-tune.
Approximate results from my local testing; not a standardized benchmark.
The results reported here are based on my own testing.
The model may perform differently on other music, datasets or hardware.
No claim is made that V2 is universally superior to the original encoder.
Model summary
Model: heartcodec-encoder-ft-v2
Base: HeartMuLa/HeartCodec-oss-encoder
Purpose: Improved audio β codec tokens for HeartMuLa fine-tuning
Training data: 2,576 audio β token pairs
Parameters: ~105M
Training: 2 epochs / 134,730 steps / batch size 1 / LR 0.00008
GPU: NVIDIA RTX 5070 Ti
Codebooks: 8
Token rate: 12.5 Hz
Model tree for valdj/heartcodec-encoder-ft-v2
Base model
HeartMuLa/HeartCodec-oss-encoder