Edda v0.2 duo — Danish speech recognition

Edda v0.2 duo is two fine-tuned Whisper models decoded as one: Edda v0.2 large (Whisper large-v3, 1.54 B parameters) and Edda v0.2 (Whisper large-v3-turbo, 0.81 B parameters). They share a single beam search: at every decoding step both models score the same partial transcripts, and their log-probabilities are averaged with equal weight before the beams are pruned. The models make different mistakes, so the pair is more accurate than either model alone.

On the open Danish ASR leaderboard harness it scores a mean WER of 7.94 over the five test sets.

test set Edda v0.2 duo Edda v0.2 large Edda v0.2
CoRal-v3 conversation 14.50 15.97 15.54
CoRal-v3 read-aloud 8.11 8.65 9.59
Common Voice Danish (leaderboard set, 2,756 clips) 5.14 5.69 5.71
FLEURS da_dk 6.35 7.13 7.34
FTSpeech 5.58 5.80 5.77
mean WER 7.94 8.65 8.79

All three were scored with the leaderboard's own harness (Rye-A1/danish-asr-leaderboard), using 5 beams.

Parameters

component base model parameters
large — Edda v0.2 large openai/whisper-large-v3 (32-layer encoder, 32-layer decoder) 1,543,490,560
turbo — Edda v0.2 openai/whisper-large-v3-turbo (32-layer encoder, 4-layer decoder) 808,878,080
Edda v0.2 duo 2,352,368,640

Both models are in this repository's model.safetensors (fp16, 4.7 GB), under the large. and turbo. prefixes. Running the pair needs about 5 GB of GPU memory for the weights.

How decoding works

  1. Encoding. Each model encodes the audio with its own encoder. The weights differ, so the encoders cannot be shared.
  2. One beam search, 5 beams, run by the large model. At each step the large model computes log-probabilities for the next token of every beam.
  3. Turbo scores the same beams. The turbo model computes its own log-probabilities for the same 5 prefixes. It keeps a separate cache that follows the beams as they are reordered.
  4. Averaging. The two distributions are averaged, 0.5 · log p_large + 0.5 · log p_turbo, and the beam search keeps the 5 best continuations under the averaged score.
  5. Finished transcripts are ranked by the averaged, length-normalised score.

This is a log-linear ensemble (a product of experts), not a mixture of experts: both models take part in every token and nothing is routed. The pair costs about 1.15× the large model alone, because turbo's 4-layer decoder is small next to the large model's 32 layers.

Why it helps. On clips where the two models disagree, the pair beats both models on 4.5–9.3 % of clips, depending on the test set. On 55–66 % it matches the better of the two, and on only 1.5–2.3 % is it worse than both. On 5–21 % of clips it produces a transcript that neither model produced on its own: it combines the words each model is confident about.

Usage

import torch
from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2-duo",
               trust_remote_code=True, dtype=torch.float16, device="cuda")
print(asr("clip.wav")["text"])                       # any format ffmpeg reads; long recordings in 30 s windows
print([r["text"] for r in asr(["a.wav", "b.wav"], batch_size=8)])

trust_remote_code=True is required: the pair's decoding code ships in this repository (modeling_edda_duo.py and pipeline_edda_duo.py). The pipeline accepts a file path, a 16 kHz float array, a dict with "raw" (or "array") and "sampling_rate", or a list of these, and returns {"text": ...} for each. Recordings longer than 30 s are cut into consecutive 30 s windows, each transcribed on its own, and the pieces joined, which is also how the scores above were measured. Decoding defaults to Danish transcription with 5 beams.

Without the pipeline:

import soundfile as sf
import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo = "danish-foundation-models/edda-v0.2-duo"
model = AutoModelForSpeechSeq2Seq.from_pretrained(repo, trust_remote_code=True, dtype=torch.float16).to("cuda")
processor = AutoProcessor.from_pretrained(repo)

audio, sr = sf.read("clip.wav", dtype="float32")    # 16 kHz mono; resample first if needed
print(model.transcribe([audio], processor)[0])

model.transcribe takes a list of 16 kHz mono arrays and returns one transcript per array. model.generate(input_features) accepts the usual WhisperForConditionalGeneration.generate arguments. ensemble_weight= changes the weight on turbo's log-probabilities for one call.

Leaderboard harness. danish-asr-eval --model danish-foundation-models/edda-v0.2-duo --backend transformers-remote. That backend loads with trust_remote_code=True in the dtype the config declares (float16), then calls model.transcribe(processor=, language="da", audio_arrays=, sample_rates=). The parameter count, 2.35 B, is read from the safetensors.

The code in this repository (modeling_edda_duo.py, about 100 lines) runs with trust_remote_code=True. It works with both transformers 4 and 5: we tested 4.57 (the last 4.x release) and 5.10, which gave identical transcripts in our tests.

Training

The two components are trained separately with the same recipe and data. See the model cards of Edda v0.2 large and Edda v0.2. In brief:

  • Data: six public Danish corpora (CoRal-v3, FTSpeech, NST, FLEURS, Common Voice 17), training splits only, 2,605 h.
  • Fine-tuning: a full fine-tune at a constant learning rate of 1e-5. Exits are taken every 20,000 steps, each cooled down for 10,000 steps, and the four cooled-down models are averaged with equal weights.
  • Combining the pair: no further training. The 0.5 weight was not tuned.

Limitations

  • Speed: slower than either component alone: decoding costs about 1.15× Edda v0.2 large, which is itself about twice as slow as Edda v0.2.
  • Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, while CoRal references are verbatim.
  • Long audio is handled in fixed 30 s windows, so a word cut at a window boundary can be lost or garbled.
  • Untested domains: domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.

License and attribution

The model weights and code are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor of the CoRal-v3 dataset. The base models openai/whisper-large-v3 and openai/whisper-large-v3-turbo are MIT-licensed. The training corpora carry their own licenses; see the respective dataset cards. Trained by the Alexandra Institute within the CoRal project and released by Danish Foundation Models.

Downloads last month
25
Safetensors
Model size
2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for danish-foundation-models/edda-v0.2-duo

Finetuned
(1085)
this model

Datasets used to train danish-foundation-models/edda-v0.2-duo

Evaluation results