Instructions to use Harmonium/brage-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Harmonium/brage-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Harmonium/brage-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("Harmonium/brage-v1") model = AutoModelForSpeechSeq2Seq.from_pretrained("Harmonium/brage-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Brage carries the use restrictions of CoRal, its training data. Please read LICENSE before downloading.
Log in or Sign Up to review the conditions and access this model content.
Brage v1: Danish speech recognition
Brage, named after the Norse god of poetry and eloquence, is
openai/whisper-large-v3 (1.55 B parameters, 32-layer decoder) fully
fine-tuned on 2,660 hours of transcribed Danish speech. On the
open Danish ASR leaderboard harness it scores a mean WER of
8.34 over the five test sets, at about 18x real time on one RTX 4090 with the decoding below.
| test set | WER | CER |
|---|---|---|
| CoRal-v3 conversation | 16.20 | 9.40 |
| CoRal-v3 read-aloud | 8.30 | 3.31 |
| Common Voice Danish (leaderboard set, 2,756 clips) | 5.56 | 1.90 |
| FLEURS da_dk | 6.01 | 2.45 |
| FTSpeech | 5.64 | 3.14 |
| mean | 8.34 | 4.04 |
Usage
The repository is gated: accept the terms on this page once, and log in with hf auth login.
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Harmonium/brage-v1")
sys.path.insert(0, path)
from brage_decode import BrageASR
asr = BrageASR.from_pretrained(path, device="cuda")
print(asr.transcribe("clip.wav"))
brage_decode.py is the decoder behind the leaderboard scores. It needs torch, transformers 5.17 or later (below 6),
kenlm, soundfile and the leaderboard's text normaliser
(pip install git+https://github.com/Rye-A1/danish-asr-leaderboard). On first use it downloads the Danish language model
danish-foundation-models/munin-7b-alpha (14.5 GB,
Apache 2.0). Whisper and Munin take turns on the GPU, which needs about 15 GB free; neural_lm=None skips Munin
(8.58 mean WER). Once the leaderboard merges the brage backend, the harness runs it with
danish-asr-eval --model Harmonium/brage-v1 --backend brage. Decoding:
- beam search with 5 beams returns its 5 best transcripts, rescored as acoustic log-probability + 0.15 x Munin
log-probability + 0.025 x word 4-gram (
lm/brage-4gram.bin, KenLM) + 0.5 per word; - a candidate whose text compresses better than 2.4:1 under zlib is a repetition loop and is dropped; if all five are, the clip is decoded greedily with Whisper's temperature fallback;
- clips over 30 s are cut at quiet points into pieces of at most 28 s, and the pieces' texts are joined.
The 4-gram was built from the training transcripts minus every sentence that also occurs in a dev or test set. Mean dev WER: beam search with the loop filter 8.83, with the 4-gram 8.70, with Munin as well 8.37.
Training data
Training used six sets, gold transcripts only: 1.66 M clips and 2,660 hours after filtering.
| corpus (train split) | clips | hours | share of audio |
|---|---|---|---|
| FTSpeech (parliament) | 980,558 | 1,683 | 43 % |
| CoRal-v3 read-aloud | 299,253 | 521 | 24 % |
| NST Danish (Språkbanken, main and supplement) | 227,471 | 303 | 16 % |
| CoRal-v3 conversation | 146,991 | 144 | 12 % |
| FLEURS da_dk | 2,463 | 7.5 | 3 % |
| Common Voice 17 Danish | 3,176 | 4 | 2 % |
Transcripts are used as published (FTSpeech without casing or punctuation); the leaderboard's normaliser removes both before scoring. Clips by the 119 Common Voice test speakers were removed.
CoRal reuses its prompt sentences across speakers. We kept the 24,616 training clips whose sentence also occurs in a CoRal dev or test set; on the read-aloud test set, the 6,435 clips with such a sentence score 7.67 WER and the other 2,687 score 9.78.
Training recipe
A full fine-tune (encoder and decoder) following the published recipe of Edda v0.1.
| audio seen | about 12,400 h in 38,363 updates of about 256 clips, 3 days on one RTX 4090 |
| batching | 30 s windows packed first-fit decreasing for the first 90 % of the audio, single clips for the last 10 % (the learning-rate decay) |
| optimizer | 8-bit AdamW (torchao) on bf16 weights with stochastic rounding |
| schedule | warmup, then 1e-5 flat, then linear decay |
| loss | cross-entropy with label smoothing 0.05 |
| augmentation | speed 0.9 to 1.1, coloured noise at 0 to 20 dB SNR (p 0.6), light SpecAugment, 2.5 % noise-only windows with an empty transcript |
| weights | EMA 0.9998; the release averages four EMA snapshots (updates 28,000 to 38,363), chosen on dev |
Limitations
- On parliamentary speech the model tends to leave out restarts and repetitions, as the edited FTSpeech record does.
- On silence or noise alone it can produce a short phrase that was never said.
- On the Alvenir evaluation sets it scores 2.52 WER on
ossand 4.78 onwiki(Edda v0.1: 2.94 and 5.60). - Munin pulls transcripts toward written Danish: it gains 0.72 WER on read speech and loses 0.27 on conversation. It is a public model and predicts the read-aloud and FLEURS references much better than the others, so it may have seen some of their source texts.
- Children, strong non-native accents, telephone-band audio and medical dictation are untested.
License and attribution
Trained by eriksp and cpius (Harmonium).
The weights and the 4-gram are released under the Apache License 2.0 with use restrictions added (LICENSE);
brage_decode.py is Apache 2.0 alone. CoRal's license (LICENSE-CORAL) makes its use restrictions apply
to any model built from it: among them, no imitating a specific person's voice, no inferring traits such as age, gender
or health from speech, and machine-generated text must be disclosed as such when shared. Redistributors must pass them
on.
The base model openai/whisper-large-v3 is Apache 2.0. NOTICE credits every training corpus. FT Speech
contains data from Folketinget, which has not endorsed this model.
- Downloads last month
- 2
Model tree for Harmonium/brage-v1
Base model
openai/whisper-large-v3Datasets used to train Harmonium/brage-v1
CoRal-project/coral-v3
alexandrainst/ftspeech
Evaluation results
- WER on CoRal-v3 conversation (test)test set self-reported16.200
- CER on CoRal-v3 conversation (test)test set self-reported9.400
- WER on CoRal-v3 read-aloud (test)test set self-reported8.300
- CER on CoRal-v3 read-aloud (test)test set self-reported3.310
- WER on Common Voice Danish (leaderboard settest set self-reported5.560
- CER on Common Voice Danish (leaderboard settest set self-reported1.900
- WER on FLEURS da_dk (test)test set self-reported6.010
- CER on FLEURS da_dk (test)test set self-reported2.450