You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Brage carries the use restrictions of CoRal, its training data. Please read LICENSE before downloading.

Log in or Sign Up to review the conditions and access this model content.

Brage v1: Danish speech recognition

Brage, named after the Norse god of poetry and eloquence, is openai/whisper-large-v3 (1.55 B parameters, 32-layer decoder) fully fine-tuned on 2,660 hours of transcribed Danish speech. On the open Danish ASR leaderboard harness it scores a mean WER of 8.34 over the five test sets, at about 18x real time on one RTX 4090 with the decoding below.

test set WER CER
CoRal-v3 conversation 16.20 9.40
CoRal-v3 read-aloud 8.30 3.31
Common Voice Danish (leaderboard set, 2,756 clips) 5.56 1.90
FLEURS da_dk 6.01 2.45
FTSpeech 5.64 3.14
mean 8.34 4.04

Usage

The repository is gated: accept the terms on this page once, and log in with hf auth login.

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("Harmonium/brage-v1")
sys.path.insert(0, path)
from brage_decode import BrageASR

asr = BrageASR.from_pretrained(path, device="cuda")
print(asr.transcribe("clip.wav"))

brage_decode.py is the decoder behind the leaderboard scores. It needs torch, transformers 5.17 or later (below 6), kenlm, soundfile and the leaderboard's text normaliser (pip install git+https://github.com/Rye-A1/danish-asr-leaderboard). On first use it downloads the Danish language model danish-foundation-models/munin-7b-alpha (14.5 GB, Apache 2.0). Whisper and Munin take turns on the GPU, which needs about 15 GB free; neural_lm=None skips Munin (8.58 mean WER). Once the leaderboard merges the brage backend, the harness runs it with danish-asr-eval --model Harmonium/brage-v1 --backend brage. Decoding:

  • beam search with 5 beams returns its 5 best transcripts, rescored as acoustic log-probability + 0.15 x Munin log-probability + 0.025 x word 4-gram (lm/brage-4gram.bin, KenLM) + 0.5 per word;
  • a candidate whose text compresses better than 2.4:1 under zlib is a repetition loop and is dropped; if all five are, the clip is decoded greedily with Whisper's temperature fallback;
  • clips over 30 s are cut at quiet points into pieces of at most 28 s, and the pieces' texts are joined.

The 4-gram was built from the training transcripts minus every sentence that also occurs in a dev or test set. Mean dev WER: beam search with the loop filter 8.83, with the 4-gram 8.70, with Munin as well 8.37.

Training data

Training used six sets, gold transcripts only: 1.66 M clips and 2,660 hours after filtering.

corpus (train split) clips hours share of audio
FTSpeech (parliament) 980,558 1,683 43 %
CoRal-v3 read-aloud 299,253 521 24 %
NST Danish (Språkbanken, main and supplement) 227,471 303 16 %
CoRal-v3 conversation 146,991 144 12 %
FLEURS da_dk 2,463 7.5 3 %
Common Voice 17 Danish 3,176 4 2 %

Transcripts are used as published (FTSpeech without casing or punctuation); the leaderboard's normaliser removes both before scoring. Clips by the 119 Common Voice test speakers were removed.

CoRal reuses its prompt sentences across speakers. We kept the 24,616 training clips whose sentence also occurs in a CoRal dev or test set; on the read-aloud test set, the 6,435 clips with such a sentence score 7.67 WER and the other 2,687 score 9.78.

Training recipe

A full fine-tune (encoder and decoder) following the published recipe of Edda v0.1.

audio seen about 12,400 h in 38,363 updates of about 256 clips, 3 days on one RTX 4090
batching 30 s windows packed first-fit decreasing for the first 90 % of the audio, single clips for the last 10 % (the learning-rate decay)
optimizer 8-bit AdamW (torchao) on bf16 weights with stochastic rounding
schedule warmup, then 1e-5 flat, then linear decay
loss cross-entropy with label smoothing 0.05
augmentation speed 0.9 to 1.1, coloured noise at 0 to 20 dB SNR (p 0.6), light SpecAugment, 2.5 % noise-only windows with an empty transcript
weights EMA 0.9998; the release averages four EMA snapshots (updates 28,000 to 38,363), chosen on dev

Limitations

  • On parliamentary speech the model tends to leave out restarts and repetitions, as the edited FTSpeech record does.
  • On silence or noise alone it can produce a short phrase that was never said.
  • On the Alvenir evaluation sets it scores 2.52 WER on oss and 4.78 on wiki (Edda v0.1: 2.94 and 5.60).
  • Munin pulls transcripts toward written Danish: it gains 0.72 WER on read speech and loses 0.27 on conversation. It is a public model and predicts the read-aloud and FLEURS references much better than the others, so it may have seen some of their source texts.
  • Children, strong non-native accents, telephone-band audio and medical dictation are untested.

License and attribution

Trained by eriksp and cpius (Harmonium).

The weights and the 4-gram are released under the Apache License 2.0 with use restrictions added (LICENSE); brage_decode.py is Apache 2.0 alone. CoRal's license (LICENSE-CORAL) makes its use restrictions apply to any model built from it: among them, no imitating a specific person's voice, no inferring traits such as age, gender or health from speech, and machine-generated text must be disclosed as such when shared. Redistributors must pass them on.

The base model openai/whisper-large-v3 is Apache 2.0. NOTICE credits every training corpus. FT Speech contains data from Folketinget, which has not endorsed this model.

Downloads last month
2
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Harmonium/brage-v1

Finetuned
(1076)
this model

Datasets used to train Harmonium/brage-v1

Evaluation results