Instructions to use danish-foundation-models/edda-v0.2-duo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use danish-foundation-models/edda-v0.2-duo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2-duo", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSpeechSeq2Seq model = AutoModelForSpeechSeq2Seq.from_pretrained("danish-foundation-models/edda-v0.2-duo", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Edda v0.2 duo — Danish speech recognition
Edda v0.2 duo is two fine-tuned Whisper models decoded as one: Edda v0.2 large (Whisper large-v3, 1.54 B parameters) and Edda v0.2 (Whisper large-v3-turbo, 0.81 B parameters). They share a single beam search: at every decoding step both models score the same partial transcripts, and their log-probabilities are averaged with equal weight before the beams are pruned. The models make different mistakes, so the pair is more accurate than either model alone.
On the open Danish ASR leaderboard harness it scores a mean WER of 7.94 over the five test sets.
| test set | Edda v0.2 duo | Edda v0.2 large | Edda v0.2 |
|---|---|---|---|
| CoRal-v3 conversation | 14.50 | 15.97 | 15.54 |
| CoRal-v3 read-aloud | 8.11 | 8.65 | 9.59 |
| Common Voice Danish (leaderboard set, 2,756 clips) | 5.14 | 5.69 | 5.71 |
| FLEURS da_dk | 6.35 | 7.13 | 7.34 |
| FTSpeech | 5.58 | 5.80 | 5.77 |
| mean WER | 7.94 | 8.65 | 8.79 |
All three were scored with the leaderboard's own harness (Rye-A1/danish-asr-leaderboard), using 5 beams.
Parameters
| component | base model | parameters |
|---|---|---|
large — Edda v0.2 large |
openai/whisper-large-v3 (32-layer encoder, 32-layer decoder) |
1,543,490,560 |
turbo — Edda v0.2 |
openai/whisper-large-v3-turbo (32-layer encoder, 4-layer decoder) |
808,878,080 |
| Edda v0.2 duo | 2,352,368,640 |
Both models are in this repository's model.safetensors (fp16, 4.7 GB), under the large. and turbo. prefixes.
Running the pair needs about 5 GB of GPU memory for the weights.
How decoding works
- Encoding. Each model encodes the audio with its own encoder. The weights differ, so the encoders cannot be shared.
- One beam search, 5 beams, run by the large model. At each step the large model computes log-probabilities for the next token of every beam.
- Turbo scores the same beams. The turbo model computes its own log-probabilities for the same 5 prefixes. It keeps a separate cache that follows the beams as they are reordered.
- Averaging. The two distributions are averaged,
0.5 · log p_large + 0.5 · log p_turbo, and the beam search keeps the 5 best continuations under the averaged score. - Finished transcripts are ranked by the averaged, length-normalised score.
This is a log-linear ensemble (a product of experts), not a mixture of experts: both models take part in every token and nothing is routed. The pair costs about 1.15× the large model alone, because turbo's 4-layer decoder is small next to the large model's 32 layers.
Why it helps. On clips where the two models disagree, the pair beats both models on 4.5–9.3 % of clips, depending on the test set. On 55–66 % it matches the better of the two, and on only 1.5–2.3 % is it worse than both. On 5–21 % of clips it produces a transcript that neither model produced on its own: it combines the words each model is confident about.
Usage
import torch
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="danish-foundation-models/edda-v0.2-duo",
trust_remote_code=True, dtype=torch.float16, device="cuda")
print(asr("clip.wav")["text"]) # any format ffmpeg reads; long recordings in 30 s windows
print([r["text"] for r in asr(["a.wav", "b.wav"], batch_size=8)])
trust_remote_code=True is required: the pair's decoding code ships in this repository (modeling_edda_duo.py and
pipeline_edda_duo.py). The pipeline accepts a file path, a 16 kHz float array, a dict with "raw" (or "array") and
"sampling_rate", or a list of these, and returns {"text": ...} for each. Recordings longer than 30 s are cut into
consecutive 30 s windows, each transcribed on its own, and the pieces joined, which is also how the scores above were
measured. Decoding defaults to Danish transcription with 5 beams.
Without the pipeline:
import soundfile as sf
import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
repo = "danish-foundation-models/edda-v0.2-duo"
model = AutoModelForSpeechSeq2Seq.from_pretrained(repo, trust_remote_code=True, dtype=torch.float16).to("cuda")
processor = AutoProcessor.from_pretrained(repo)
audio, sr = sf.read("clip.wav", dtype="float32") # 16 kHz mono; resample first if needed
print(model.transcribe([audio], processor)[0])
model.transcribe takes a list of 16 kHz mono arrays and returns one transcript per array. model.generate(input_features)
accepts the usual WhisperForConditionalGeneration.generate arguments. ensemble_weight= changes the weight on turbo's
log-probabilities for one call.
Leaderboard harness. danish-asr-eval --model danish-foundation-models/edda-v0.2-duo --backend transformers-remote. That
backend loads with trust_remote_code=True in the dtype the config declares (float16), then calls
model.transcribe(processor=, language="da", audio_arrays=, sample_rates=). The parameter count, 2.35 B, is read from the
safetensors.
The code in this repository (modeling_edda_duo.py, about 100 lines) runs with trust_remote_code=True. It works with both
transformers 4 and 5: we tested 4.57 (the last 4.x release) and 5.10, which gave identical transcripts in our tests.
Training
The two components are trained separately with the same recipe and data. See the model cards of Edda v0.2 large and Edda v0.2. In brief:
- Data: six public Danish corpora (CoRal-v3, FTSpeech, NST, FLEURS, Common Voice 17), training splits only, 2,605 h.
- Fine-tuning: a full fine-tune at a constant learning rate of 1e-5. Exits are taken every 20,000 steps, each cooled down for 10,000 steps, and the four cooled-down models are averaged with equal weights.
- Combining the pair: no further training. The 0.5 weight was not tuned.
Limitations
- Speed: slower than either component alone: decoding costs about 1.15× Edda v0.2 large, which is itself about twice as slow as Edda v0.2.
- Transcription convention follows the corpora. FTSpeech references are the edited parliamentary record, while CoRal references are verbatim.
- Long audio is handled in fixed 30 s windows, so a word cut at a window boundary can be lost or garbled.
- Untested domains: domains not represented in training (children, strong non-native accents, telephone-band audio, singing) are untested.
License and attribution
The model weights and code are released under the Apache License 2.0 by the Alexandra Institute, which is also the licensor
of the CoRal-v3 dataset. The base models openai/whisper-large-v3
and openai/whisper-large-v3-turbo are MIT-licensed. The training corpora carry their own licenses; see the respective
dataset cards. Trained by the Alexandra Institute within the CoRal project and released by Danish Foundation Models.
- Downloads last month
- 25
Model tree for danish-foundation-models/edda-v0.2-duo
Base model
openai/whisper-large-v3Datasets used to train danish-foundation-models/edda-v0.2-duo
CoRal-project/coral-v3
mozilla-foundation/common_voice_17_0
Evaluation results
- WER on CoRal-v3 conversation (test)test set self-reported14.500
- CER on CoRal-v3 conversation (test)test set self-reported8.280
- WER on CoRal-v3 read-aloud (test)test set self-reported8.110
- CER on CoRal-v3 read-aloud (test)test set self-reported3.140
- WER on Common Voice Danish (leaderboard settest set self-reported5.140
- CER on Common Voice Danish (leaderboard settest set self-reported1.660
- WER on FLEURS da_dk (test)test set self-reported6.350
- CER on FLEURS da_dk (test)test set self-reported2.510