Baseline for the Second Cadenza Lyric Intelligibility Prediction (CLIP2) challenge

The baseline model for the CLIP2 Cadenza challenge consisting of a regression model that predicts a scalar in [0, 1] representing how intelligible a song's lyrics are. Built on a partially fine-tuned Whisper encoder as front-end and a small MLP regression head.

The lyric_intelligibility_model.py is self-contained: it has the model (WhisperIntelligibilityModel, with save_pretrained()/from_pretrained() mirroring the standard HF API) and a ready-to-use inference wrapper (IntelligibilityPredictor) in one file.

Files in this repo

File Purpose
config.json Model hyperparameters (backbone, hidden_dim, layer selection)
model.safetensors Trained weights (state_dict)
lyric_intelligibility_model.py WhisperIntelligibilityModel + IntelligibilityPredictor โ€” everything needed to load and run the model
inference.py Optional CLI wrapper around IntelligibilityPredictor that writes results to a CSV file

Loading and using the model

from lyric_intelligibility_model import IntelligibilityPredictor

predictor = IntelligibilityPredictor.from_pretrained(
    "cadenzachallenge/CLIP2-BaselineMono"
)
result = predictor.predict("path/to/song.wav")
print(result["score"])          # e.g. 0.83
print(result["channel"])        # "mono", or "left"/"right" for stereo (better-ear)

predict() also accepts a directory: it scores every audio file at the top level (not recursive) and returns {filename: result} instead of a single result dict.

results = predictor.predict("path/to/songs_dir")
for filename, r in results.items():
    print(filename, r["score"])

Stereo audio is automatically scored per-channel and combined via the better-ear strategy (max(score_left, score_right)) by default โ€” pass better_ear=False to downmix to mono instead. Audio is resampled to 16 kHz internally using soxr (SoX HQ).

If you just want the raw model (e.g. for further fine-tuning) without the audio-loading wrapper:

from lyric_intelligibility_model import WhisperIntelligibilityModel
from transformers import WhisperProcessor

model = WhisperIntelligibilityModel.from_pretrained(
    "cadenzachallenge/CLIP2-BaselineMono"
)
processor = WhisperProcessor.from_pretrained(
    model.SUPPORTED_MODELS['whisper-large-v3']
)

# unfreeze some encoder layers for fine-tuning
model.unfreeze_backbone(num_top_layers=4, context_layers=2)

# preprocess audio -> log-mel features
inputs = processor(waveform, sampling_rate=16000, return_tensors="pt")
scores = model(inputs.input_features)  # forward pass, then compute your loss

Command-line usage

inference.py is a thin CLI wrapper around IntelligibilityPredictor that writes a results.csv (columns: filename stem, score โ€” no header) instead of returning a dict:

pip install torch transformers huggingface_hub safetensors soundfile soxr

python inference.py \
    --repo_id cadenzachallenge/CLIP2-BaselineMono \
    --audio path/to/song.wav          # or a directory

# options:
#   --no-better-ear   downmix stereo to mono instead of left/right max
#   --max-100         scale scores to [0, 100] instead of [0, 1]
#   --output FILE     custom CSV path (default: results.csv)

Model details

  • Backbone: Whisper encoder (size set by config.json, default whisper-large-v3)

  • Layer fusion: mean-pooled encoder layer(s) โ†’ learned per-layer weights โ†’ optional cross-layer attention (when multiple layers are selected)

  • Head: 2-layer MLP + sigmoid โ†’ score in (0, 1)

  • Output: single scalar intelligibility score per input

  • Data: The model was trained in CLIP2 train dataset.

More Information

For more information, please see the Cadenza Challenge website.

Downloads last month
50
Safetensors
Model size
0.6B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support