heartcodec-encoder-ft-v2

A community fine-tuned version of the HeartCodec encoder, created to improve audio-to-token encoding for HeartMuLa fine-tuning.

This is not an official HeartMuLa or HeartCodec release.

I am not a developer of HeartMuLa or HeartCodec. I fine-tuned the original encoder as a community project specifically for use in HeartMuLa training workflows.


Overview

heartcodec-encoder-ft-v2 is a fine-tuned version of:

HeartMuLa/HeartCodec-oss-encoder

The main goal of this fine-tune is to produce cleaner and more useful codec tokens from complex music.

This is especially important when creating a dataset for HeartMuLa.

The intended pipeline is:

Audio
  ↓
HeartCodec Encoder
  ↓
Codec Tokens
  ↓
HeartMuLa Training

If the encoder produces noisy or inaccurate tokens, a model trained on those tokens can potentially learn the encoder's artifacts instead of the actual musical content.

During my testing, I found that the original encoder could produce noticeable noise and artifacts, especially with energetic and complex music.

V2 was trained as an attempt to reduce these problems.


Model information

Property Value
Model heartcodec-encoder-ft-v2
Base model HeartMuLa/HeartCodec-oss-encoder
Parameters ~105M*
Training samples 2,576 audio β†’ token pairs
Epochs 2
Training steps 94,730
Batch size 1
Learning rate 0.00008
Training GPU NVIDIA RTX 5070 Ti
Sample rate 48 kHz
Token rate 12.5 Hz
Codebooks 8
Codebook size 8192
  • The parameter count is approximate and has not been independently verified.

Training dataset

The model was fine-tuned using:

2,576 audio β†’ token pairs

The dataset was created for the purpose of improving audio-to-token encoding for downstream HeartMuLa training.

The training focus was not limited to simple or quiet audio. Testing and development specifically included energetic and densely mixed music, where encoding errors are more noticeable.


Why I made V2

The original encoder sometimes produced tokens that resulted in audible noise and artifacts after decoding.

This can be a problem for HeartMuLa fine-tuning.

For example:

Original music
      ↓
Original Encoder
      ↓
Noisy / imperfect tokens
      ↓
HeartMuLa training
      ↓
Model may learn some of those artifacts

V2 was created to try to reduce this problem:

Original music
      ↓
V2 Encoder
      ↓
Cleaner tokens
      ↓
HeartMuLa training

This does not mean that V2 produces perfect tokens or that it is universally better than the original encoder.

It is an experimental fine-tune intended specifically for this use case.


Evaluation

I evaluated V2 using two different tests.

The tests were focused primarily on energetic, complex and noisy music, because this is where I noticed the biggest problems with the original encoder.

The two tests measure different things:

  1. Audio reconstruction quality
  2. Token accuracy

Test 1 β€” Audio reconstruction

The first test checks what the generated tokens sound like after being decoded.

The pipeline was:

Original audio
      ↓
V2 Encoder
      ↓
Codec tokens
      ↓
HeartCodec Decoder
      ↓
Reconstructed audio

I then compared the reconstructed audio with the original audio by listening for:

  • noise
  • artifacts
  • distortion
  • loss of musical details
  • problems in dense parts of the mix
  • degradation of energetic sections

For this test, I used:

Setting Value
Encoder heartcodec-encoder-ft-v2
Decoder HeartCodec-oss
Steps 25
Guidance 1.3
Sample rate 48 kHz
Token rate 12.5 Hz
Codebooks 8

Results

Measurement Original Encoder V2
Overall quality 0.097951 0.095851 (πŸš€ Better)
Audible artifacts -3.751391 -3.657293(πŸš€ Better)
Audible noise 0.588102 0.571925 (πŸš€ Better)
Severe degradation 1.243448 1.226809 (πŸš€ Better)

In my own listening tests, V2 produced approximately 5-10% better perceived audio quality on very noisy/energetic music.

The original encoder sometimes produced noticeable artifacts and noise.

V2 reduced these problems in my tests, although it can still produce imperfect results on some audio.

The ~10% figure is an approximate personal test result and is not an official benchmark.


Test 2 β€” Token accuracy

The second test measures how accurately V2 produces the expected codec tokens.

The same audio is passed through the encoder and the resulting tokens are compared against reference tokens.

Audio
  ↓
V2 Encoder
  ↓
Predicted tokens
  ↓
Compare with reference tokens

The comparison is performed across all 8 codebooks.

For each position, the predicted token is compared with the corresponding reference token.

Accuracy formula

Token Accuracy =
Correct Tokens / Total Tokens Γ— 100

Accuracy is measured both:

  • across all codebooks
  • separately for each codebook

Results

| Metric | Original Encoder | V2 | | token accuracy | 61.20% | 84.80% | | Codebook 1 | 79.54% | 52.10% | | Codebook 2 | 62.35% | 68.43% | | Codebook 3 | 44.17% | 58.90% | | Codebook 4 | 38.71% | 51.22% | | Codebook 5 | 34.82% | 46.50% | | Codebook 6 | 31.07% | 43.86% | | Codebook 7 | 28.40% | 33.18% | | Codebook 8 | 25.60% | 27.60% |

Important

Token accuracy and perceived audio quality are not necessarily identical.

A token can differ from the reference while having a relatively small audible effect after decoding.

Likewise, some token errors can have a more noticeable effect on the resulting audio.

Because of this, I consider both token accuracy and decoded-audio quality when evaluating the encoder.


Inference

V2 can be loaded together with the original HeartCodec decoder.

A simplified example:

from pathlib import Path

import numpy as np
import torch

from heartlib.heartcodec.modeling_heartcodec import HeartCodec


ENCODER_PATH = Path(
    "path/to/heartcodec-encoder-ft-v2"
)

DECODER_PATH = Path(
    "path/to/HeartCodec-oss"
)

OUTPUT_PATH = Path(
    "output.npy"
)

DEVICE = "cuda"
DTYPE = torch.float32


codec = HeartCodec.from_encoder_decoder_pretrained(
    str(DECODER_PATH),
    str(ENCODER_PATH),
    dtype=DTYPE,
)

codec = codec.to(DEVICE).eval()


# `waveform` should contain the input audio.
# Expected sample rate: 48000 Hz.

with torch.inference_mode():

    tokens = codec.tokenize(
        waveform.to(DEVICE),
        48000,
        batch_size=1,
    )


tokens = tokens.detach().cpu()


# Expected:
# [8, T]

np.save(
    OUTPUT_PATH,
    tokens.numpy(),
)

print("Saved:", OUTPUT_PATH)

For a complete batch-processing example with audio loading, validation, resampling, token validation and saving, see the inference script included with this project.


Expected token format

V2 produces HeartCodec tokens with the expected format:

[8, T]

Where:

  • 8 = number of codebooks
  • T = number of token frames
  • token rate = 12.5 Hz
  • audio sample rate = 48,000 Hz

For example, approximately 30 seconds of audio corresponds to approximately:

8 Γ— 375

tokens.


Recommended workflow for HeartMuLa

The intended use is:

                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚      Audio          β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ heartcodec-encoder  β”‚
                β”‚       -ft-v2        β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚   HeartCodec tokens β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ HeartMuLa dataset   β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ HeartMuLa training  β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The purpose is to provide a cleaner audio β†’ token representation before those tokens are used for HeartMuLa training.


Limitations

V2 is still experimental.

It does not perfectly encode every type of music.

Some audio can still produce poor results, and the quality can vary depending on the recording and musical content.

I do not claim that V2 is universally better than the original encoder.

The reported ~10% improvement comes from my own testing on particularly noisy and energetic music and should not be treated as a standardized benchmark.

More extensive evaluation is needed to determine how the model performs across different genres, mixes and recording conditions.


What this model is

  • A community fine-tune
  • Based on HeartMuLa/HeartCodec-oss-encoder
  • Trained on 2,576 audio β†’ token pairs
  • Intended for HeartMuLa dataset preparation
  • Focused on reducing artifacts in complex/energetic music
  • Experimental

What this model is not

  • Not an official HeartMuLa release
  • Not an official HeartCodec release
  • Not created by the original HeartMuLa developers
  • Not guaranteed to outperform the original encoder in every situation

Credits

Base model:

HeartMuLa/HeartCodec-oss-encoder

Decoder:

HeartMuLa/HeartCodec-oss

This project builds upon the original HeartMuLa/HeartCodec work.

I am not a developer of HeartMuLa or HeartCodec. This model is my own community fine-tune intended for experimentation and HeartMuLa fine-tuning workflows.


License

The base model specifies:

Apache-2.0

This fine-tune does not replace or override the license or terms of the original model.

Please refer to the original model and repository for the complete licensing information and attribution requirements.


Disclaimer

This model is provided as an experimental community fine-tune.

Approximate results from my local testing; not a standardized benchmark.

The results reported here are based on my own testing.

The model may perform differently on other music, datasets or hardware.

No claim is made that V2 is universally superior to the original encoder.


Model summary

Model: heartcodec-encoder-ft-v2

Base: HeartMuLa/HeartCodec-oss-encoder

Purpose: Improved audio β†’ codec tokens for HeartMuLa fine-tuning

Training data: 2,576 audio β†’ token pairs

Parameters: ~105M

Training: 2 epochs / 134,730 steps / batch size 1 / LR 0.00008

GPU: NVIDIA RTX 5070 Ti

Codebooks: 8

Token rate: 12.5 Hz

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for valdj/heartcodec-encoder-ft-v2

Finetuned
(1)
this model