Instructions to use alefiury/CALMOS-MMS-300m-BRSpeech with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alefiury/CALMOS-MMS-300m-BRSpeech with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="alefiury/CALMOS-MMS-300m-BRSpeech", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("alefiury/CALMOS-MMS-300m-BRSpeech", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
CAL-MOS · MMS 300m · BRSpeech
CAL-MOS is a non-intrusive Mean Opinion Score (MOS) predictor built on speech foundation models. This checkpoint uses the Adapters + Mean (A+M) configuration with MMS 300m, trained on BRSpeech.
Model
| Backbone | facebook/mms-300m |
| Dataset | BRSpeech |
| Strategy | Adapters + Mean (A+M) |
| Backbone | Frozen |
| Pooling | Mean |
| Input sampling rate | 16,000 Hz |
CAL-MOS collects representations across the frozen encoder depth, calibrates each layer with a lightweight adapter, and combines the adapted representations before MOS regression.
Usage
Load the model directly from the Hugging Face Hub with Transformers. No clone or manual snapshot download is required.
from transformers import AutoModel
model = AutoModel.from_pretrained(
"alefiury/CALMOS-MMS-300m-BRSpeech",
trust_remote_code=True,
).to("cuda")
mos = model.predict("audio.wav")
print(mos)
Batch inference is also supported:
scores = model.predict(
[
"audio_1.wav",
"audio_2.wav",
"audio_3.wav",
],
batch_size=8,
)
print(scores)
Audio is converted to mono and resampled to 16,000 Hz automatically. Predictions are clipped to the [1, 5] MOS range by default.
trust_remote_code=True is required because the lightweight CAL-MOS
architecture is shipped with this model repository. The frozen backbone is
downloaded automatically from its original Hugging Face repository.
Citation
@inproceedings{ferreira26_interspeech,
title = {{CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models}},
author = {Alef Iury Ferreira and Pedro Botelho and Fernanda Silva and Daniel Casanova and Rafael Faustino and Frederico Oliveira and Arlindo Galvão Filho and Anderson da Silva Soares},
year = {2026},
booktitle = {{Interspeech 2026}},
pages = {174--179},
doi = {10.21437/Interspeech.2026-2960},
issn = {2958-1796}
}
License
MIT. See the CAL-MOS repository for the project license.
The MMS-300m backbone remains subject to its own license.
- Downloads last month
- 21