Mimir

DFM Mimir v1.5

Danish Foundation Models SDU UCloud

Trained on SDU UCloud

DFM Mimir v1.5 is the final epoch-10 EMA checkpoint of the XL HRM-Text training lineage, at 2,877,261 optimizer steps. It continues the Danish and English Mimir model family on subsequent dataset mixtures, ending with DFM11. It is a continuation checkpoint, not a new model trained from scratch for ten identical epochs. It is not an XXL model.

The architecture has approximately 0.981B non-embedding/head parameters, or 1.787B total parameters including its separate 262,144-token input embedding and output head. The family is historically described as a 1B model. This repository contains BF16 inference weights, not optimizer state or a resumable FSDP training checkpoint.

Model details

Property Value
Architecture HRM-Text XL, PrefixLM
Total parameters 1,786,775,040
Parameters excluding input embedding and output head 981,468,672
Hidden dimension 1,536
Stored transformer layers 16, reused across recurrent cycles
Attention heads 12
H / L cycles 2 / 3; six L and two H executions per complete reasoning pass
Vocabulary 262,144
Context length 4,096
Weight dtype BF16
Exported weights EMA, epoch 10, step 2,877,261
Final epoch dataset DFM11, 103,214,604,702 sampled tokens
Final base learning rate 1e-5, with automatic H/L module scaling
License Apache 2.0

The final-epoch token count is not a claim about every preceding epoch or the total training-token budget. Earlier epochs used different mixtures.

Training data and procedure

Mimir began as a model trained from scratch on Danish and English post-training and instruction data. v1.5 continues that XL lineage through later mixtures; it is not a fresh ten-epoch run over one fixed dataset. The original v1 card reports 161 datasets and approximately 70.479B sampled tokens per epoch for its DFM8 mixture. Those figures describe the earlier mixture, not DFM11.

The final epoch resumes the completed DFM10 epoch-9 checkpoint at step 2,482,084, preserves optimizer and EMA state, and trains on DFM11 to step 2,877,261. DFM11 contains 103,214,604,702 sampled tokens per epoch and 235,520,711 sampled index rows; caps and repeats are included in that token count. The preceding DFM10 mixture contained 92,658,813,451 sampled tokens per epoch. These are mixture sizes, not unique-token counts or a total lifetime training budget. Multiplying the DFM11 size by ten would misstate this lineage.

DFM11 retains the DFM10 base and introduces Danish/English grounded FineInstructions conversations, Koolbardi data, mathematical/programmatic reasoning, repaired native tool-use conversations and repaired Danish parliamentary document error correction. Its union manifest records 15,733 retained base task files, 13 replaced base task files and 290 addition task files; these file counts are not counts of independent datasets.

The data policy combines openly licensed/public-domain sources and their synthetic derivatives with agreement-backed material. The full corpus is not reconstructible from public Hub downloads alone. Public DFM11 additions are listed with their exact revisions in training_data_manifest.json; the sampling policy records source caps and repeats. The eleven packages are:

For FineInstructions, the training selection uses 416,022 English chats (316,022 controlled chats plus 100,000 source-balanced legacy chats) and 503,740 Danish controlled chats. Published initial question/answer pairs are not separately tokenized again, because the conversations already contain those exchanges. The sampling policy excludes held-out validation/test rows; this is not a claim that an independent contamination audit has certified every benchmark in this card.

Training uses Gemma-native chat rendering with thinking disabled, a 4,096-token context and target-only loss under the PrefixLM objective. The final epoch uses a global token batch of 262,144, gradient accumulation of 2, BF16 compute with FP32 FSDP parameters/optimizer, H2/L3 recurrence and BP8. The learning-rate schedule ends at a base/embedding/head rate of 1e-5, with automatic module scaling to 5e-6 for H and approximately 1.6667e-6 for L. The released weights are the BF16 EMA export.

Source documentation: DFM8 inventory, DFM11 construction and sampling, and XL epoch-10 training continuation. The packaged manifest and sampling policy above preserve the local records used for this card even if upstream documentation changes.

Standalone benchmark comparisons

These are the completed HRM-test standalone evaluations, separate from the training-integrated evaluation below. Scores use a 0–100 scale. Category macro-averages give equal weight to 7 English, 3 Math/Code and 10 Danish tasks, with one primary metric per task. All listed category averages are complete; Munin models were evaluated for Danish only.

Checkpoint identity matters: ¹ the comparison tested the 2,850,000-step EMA checkpoint from the v1.5 training lineage, whereas this repository contains the later 2,877,261-step EMA checkpoint. Their weight hashes differ. Thus the 2,850K scores are nearby-checkpoint comparisons, not measurements of the exact released weights. ² The earlier Mimir comparison uses the local 1,650,000-step EMA baseline; the original published card separately lists 1,750,000 training steps. We do not infer weight identity from the family name.

Category macro-averages

Model / evaluated checkpoint English Math/Code Danish
Mimir v1.5 lineage, 2,850K EMA¹ 77.81 71.12 62.97
Mimir v1 baseline, 1,650K EMA² 68.97 64.14 57.04
Ouro 1.4B 60.68 56.63 24.80
Ouro 2.6B 70.23 60.80 25.39
Qwen3.5 0.8B 50.65 38.59 30.89
Qwen3.5 2B 58.67 58.97 32.73
Qwen3.5 4B 69.26 65.02 49.82
Qwen3.5 9B 77.05 85.84 53.05
Gemma 4 E2B 52.79 75.44 44.43
Gemma 4 E4B 54.89 82.85 53.58
Gemma 4 E4B (thinking) 72.26 82.17 54.53
Gemma 3 1B 37.44 43.19 34.00
OLMo 2 1B 42.82 31.35 25.72
SmolLM3 3B 63.08 67.93 28.82
EuroLLM 9B 60.61 17.30 45.74
Apertus 8B 61.63 44.11 46.52
Ministral 3 8B 71.34 76.24 42.24
Munin Apertus 8B — — 44.35
Munin-Mistral 8B — — 42.13
Munin Qwen 9B — — 44.80

Under these recorded protocols, the 2,850K Mimir checkpoint improves over the 1,650K baseline by 8.84 English, 6.97 Math/Code and 5.94 Danish percentage points. Qwen 3.5 9B and Gemma 4 E4B score higher on Math/Code; Mimir 2,850K has the highest English and Danish means in this comparison set. These are descriptive comparisons with the prompt/runtime differences below, not a claim of universal model rankings or statistical significance.

Download category averages, per-task scores and measurement provenance / checkpoint hashes.

English per-benchmark results

Model / checkpoint BoolQ Winogrande HellaSwag MMLU ARC-C DROP F1 GovReport R1 Mean
Mimir v1.5 lineage, 2,850K EMA¹ 91.62 81.77 83.88 67.41 90.36 88.77 40.85 77.81
Mimir v1 baseline, 1,650K EMA² 87.80 73.48 67.28 57.49 81.57 83.10 32.04 68.97
Ouro 1.4B 83.88 55.96 59.64 67.05 86.95 38.87 32.40 60.68
Ouro 2.6B 89.02 71.67 78.37 74.49 93.00 51.95 33.08 70.23
Qwen3.5 0.8B 69.82 50.04 36.97 51.53 68.43 45.23 32.51 50.65
Qwen3.5 2B 80.76 56.99 64.59 62.78 82.68 31.35 31.51 58.67
Qwen3.5 4B 87.03 70.01 83.15 75.79 92.92 48.03 27.89 69.26
Qwen3.5 9B 89.30 75.14 88.56 79.51 94.11 82.05 30.71 77.05
Gemma 4 E2B 64.11 54.62 45.99 44.07 69.82 57.26 33.63 52.79
Gemma 4 E4B 85.24 63.97 44.58 35.40 49.80 74.04 31.22 54.89
Gemma 4 E4B (thinking) 88.43 76.05 69.82 68.60 88.80 79.93 34.20 72.26
Gemma 3 1B 62.39 51.70 30.56 37.52 43.52 6.96 29.46 37.44
OLMo 2 1B 67.16 50.36 42.37 41.61 48.12 12.38 37.72 42.82
SmolLM3 3B 84.34 60.30 65.10 60.16 79.54 53.96 38.13 63.08
EuroLLM 9B 85.35 61.25 58.63 58.25 76.54 54.75 29.50 60.61
Apertus 8B 79.66 56.04 66.87 62.14 79.27 52.47 34.96 61.63
Ministral 3 8B 88.23 67.17 67.61 73.93 90.10 81.03 31.32 71.34

Math/Code per-benchmark results

Model / checkpoint GSM8K MATH HumanEval Mean
Mimir v1.5 lineage, 2,850K EMA¹ 91.74 48.44 73.17 71.12
Mimir v1 baseline, 1,650K EMA² 89.92 45.80 56.71 64.14
Ouro 1.4B 63.46 51.56 54.88 56.63
Ouro 2.6B 74.00 59.02 49.39 60.80
Qwen3.5 0.8B 49.13 36.16 30.49 38.59
Qwen3.5 2B 73.69 55.66 47.56 58.97
Qwen3.5 4B 60.50 56.52 78.05 65.02
Qwen3.5 9B 95.53 74.18 87.80 85.84
Gemma 4 E2B 88.32 64.22 73.78 75.44
Gemma 4 E4B 91.96 71.84 84.76 82.85
Gemma 4 E4B (thinking) 93.78 75.30 77.44 82.17
Gemma 3 1B 49.66 37.22 42.68 43.19
OLMo 2 1B 59.36 18.84 15.85 31.35
SmolLM3 3B 79.98 62.22 61.59 67.93
EuroLLM 9B 0.00 29.94 21.95 17.30
Apertus 8B 66.72 25.36 40.24 44.11
Ministral 3 8B 91.66 61.44 75.61 76.24

Danish per-benchmark results

Model / checkpoint AngryTweets F1 DaLA F1 GEC EM PIQA Daisy EM WikiQA EM WMT chrF N.News R1 IFEval strict HellaSwag-DA Mean
Mimir v1.5 lineage, 2,850K EMA¹ 63.31 96.83 92.48 73.15 12.33 66.94 54.57 29.39 66.36 74.37 62.97
Mimir v1 baseline, 1,650K EMA² 65.76 96.14 85.64 53.70 9.63 66.80 53.85 23.76 56.56 58.50 57.04
Ouro 1.4B 18.77 35.62 0.29 86.11 0.51 12.45 27.69 18.98 22.55 25.01 24.80
Ouro 2.6B 17.78 33.33 0.78 75.93 1.35 24.76 29.97 19.05 26.80 24.17 25.39
Qwen3.5 0.8B 43.61 51.03 0.68 56.48 0.68 41.55 37.84 19.37 30.31 27.31 30.89
Qwen3.5 2B 62.49 36.43 8.01 25.00 2.53 49.37 45.63 19.47 47.13 31.24 32.73
Qwen3.5 4B 68.84 50.08 42.58 70.37 4.73 57.13 52.14 23.65 67.10 61.62 49.82
Qwen3.5 9B 68.92 70.11 51.95 62.04 8.45 50.44 54.82 22.31 70.79 70.70 53.05
Gemma 4 E2B 64.71 56.66 36.91 46.30 5.57 44.09 55.21 20.53 69.50 44.84 44.43
Gemma 4 E4B 69.74 63.03 49.41 71.30 7.94 53.52 57.30 22.01 77.26 64.29 53.58
Gemma 4 E4B (thinking) 69.90 72.55 39.65 69.44 7.60 60.60 57.52 21.41 80.41 66.19 54.53
Gemma 3 1B 50.69 41.03 3.32 72.22 1.35 42.63 45.13 21.73 37.34 24.58 34.00
OLMo 2 1B 26.55 48.70 0.20 75.00 0.00 8.40 29.97 19.30 24.21 24.90 25.72
SmolLM3 3B 62.70 33.51 3.32 51.85 2.20 0.29 37.27 18.83 41.04 37.16 28.82
EuroLLM 9B 64.67 37.76 47.66 71.30 14.86 50.88 56.77 23.18 50.28 40.04 45.74
Apertus 8B 63.47 51.46 37.60 69.44 10.81 59.38 55.78 21.27 56.01 39.99 46.52
Ministral 3 8B 63.93 58.91 17.87 62.96 7.26 48.73 50.19 18.43 48.61 45.45 42.24
Munin Apertus 8B 61.42 46.08 42.09 81.48 12.50 49.90 55.85 12.66 43.81 37.74 44.35
Munin-Mistral 8B 59.31 48.85 26.37 76.85 8.45 48.39 51.81 16.09 59.15 26.06 42.13
Munin Qwen 9B 68.22 60.62 11.43 38.89 5.41 55.66 56.05 19.87 64.88 66.94 44.80

Protocol and interpretation

  • The Ouro 1.4B and 2.6B results added on 2026-09-25 cover all 20 benchmarks for ByteDance/Ouro-1.4B and ByteDance/Ouro-2.6B (non-thinking releases). Both use vLLM with the configured four recurrent steps, an 8,192-token context, native chat templates for generative tasks and the established raw-prompt MCQ protocol. They use the non-thinking comparator token budgets below. Exact model revisions and sample counts are recorded in the accompanying JSON.
  • These are the full configured benchmark splits, rather than the initial 500-example smoke-test settings. Nordjylland retains its established 500-example subset. MATH uses all 5,000 examples; GovReport uses all 973. Eight disjoint shards for each of these tasks are merged by sample count.
  • Decoding is greedy (temperature 0); sample shuffle seed is 4242. Mimir uses the original HF Transformers BF16/SDPA PrefixLM evaluation path, native Gemma chat template, thinking disabled and a 3,840-token prompt cap. Comparator serving uses vLLM, with model-specific chat/passthrough templates. Historical runs retain their recorded generation settings; this is not a uniform compute-budget comparison.
  • Mimir GSM8K is 0-shot, whereas the newer comparator runs use 10-shot. Mimir uses 1,024 generation tokens for GSM8K and 2,048 for MATH. New non-thinking comparators use 2,048 for GSM8K/MATH, 1,024 for HumanEval/DROP, and 512 for GovReport. Gemma E4B thinking uses 4,096 for MATH and 2,048 for HumanEval/DROP/GovReport. Context limits also differ: Mimir and EuroLLM use 4,096; the other new comparator servers use 8,192.
  • Danish HellaSwag includes “Svar kun med ét bogstav: a, b, c eller d.” before the answer cue. Revised non-thinking runs allow eight output tokens; E4B thinking has a larger thinking budget. The inherited scorer reads the first answer character and gives 0.25 credit to invalid answers; it is not strict whole-response letter accuracy. Older English MCQ pipelines also retain their recorded invalid-answer handling.
  • Primary metrics: accuracy except DROP F1; GovReport/Nordjylland ROUGE-1; AngryTweets/DaLA macro F1; GEC-DaLA/Daisy/MultiWikiQA exact match; WMT24++ chrF3++; and IFEval-DA prompt-strict accuracy. MATH in this standalone comparison uses the recorded mathematical answer-verification pipeline, rather than substituting the training-integrated judge score.
  • New Inspect DROP runs contain 9,535 examples; Mimir's custom DROP contains 9,536. Historical English prompt variants and backends differ across models. EuroLLM's recorded GSM8K score is 0.00; it is retained in the mean and should be read as performance under this generation/extraction configuration.
  • Gemma E4B thinking PIQA is complete (108 examples, 69.44%). Its recovery reused 26 completed samples and finished using four-way tensor parallelism; no missing-task estimate enters its Danish mean.
  • Model IDs for the new comparisons are Qwen/Qwen3.5-9B, google/gemma-4-E4B-it, utter-project/EuroLLM-9B-Instruct, swiss-ai/Apertus-8B-Instruct-2509 and mistralai/Ministral-3-8B-Instruct-2512-BF16. Apertus is the original 2509 model, not Apertus v1.5; EuroLLM is the original Instruct release. “Thinking” and non-thinking E4B use the same model weights.

Comparisons from the original DFM-Mimir card

The complete original English, Math/Code and Danish tables are retained, including HRM-Text 1B and Gemma 4 E2B with thinking. These two model variants have no complete revised-protocol table here. Their original published macro-averages are shown separately:

Historical protocol only English Math/Code Danish
HRM-Text 1B 66.1 46.9 21.7
Gemma 4 E2B (thinking) 66.6 70.5 49.9

Do not combine these historical averages with the updated table: the original Danish card uses AngryTweets accuracy, Nordjylland chrF and the old HellaSwag-DA prompt, and some English prompt variants also differ. The current tables above include the other original comparison models with their selected standalone results and revised HellaSwag-DA measurements.

Released-checkpoint evaluation (training-integrated)

This separate section preserves the released checkpoint’s existing production results; its metrics/protocols must not be mixed with the standalone tables. The following scores come from the completed production evaluation of this exact epoch-10 EMA export. Each row names its suite and metric; accuracy/F1/ ROUGE/normalized chrF values are expressed on a 0-100 scale. The plot averages the individual rows in each section, not suite averages or W&B headline averages.

Mean scores by subject area

English

Benchmark / metric Score (0-100)
BoolQ accuracy (standard) 91.16
Winogrande accuracy (standard) 82.72
HellaSwag accuracy (standard) 83.90
MMLU accuracy (standard) 69.23
ARC accuracy (standard) 89.42
DROP F1 (standard) 89.91
GovReport ROUGE-1 (DFM) 43.43
Unweighted mean of rows above 78.54

Math & Code

Benchmark / metric Score (0-100)
GSM8K accuracy (standard) 89.54
MATH accuracy (standard) 49.90
HumanEval sanitized accuracy (DFM) 73.17
Unweighted mean of rows above 70.87

Danish

Benchmark / metric Score (0-100)
AngryTweets macro F1 (EuroEval) 70.80
DaLA macro F1 (DFM) 97.07
GEC-DaLA exact match (DFM) 91.99
PIQA-DA accuracy (DFM) 74.07
Daisy / generative idioms judge accuracy (DFM) 0.12
MultiWikiQA exact match (DFM) 66.65
WMT24++ en-da chrF3++ (DFM) 54.58
NordjyllandNews chrF3++ (DFM) 29.78
IFEval-DA prompt strict accuracy (DFM) 65.99
HellaSwag-DA accuracy (EuroEval) 75.47
Unweighted mean of rows above 62.65

Evaluation caveats

  • These tables use the current production evaluation protocols. They are not a controlled re-evaluation of the comparison models in the v1 card. Do not compare similarly named tasks across suites as if prompts and scorers match.
  • In particular AngryTweets is macro F1 here, and HellaSwag-DA is EuroEval; the original v1 card used different reported metrics/protocols for some rows.
  • The very low generative-idiom/Daisy score is retained as measured, not omitted from the mean. It is a judge-based task and is not the EuroEval multiple-choice Danish idioms benchmark.
  • Standard Winogrande here is distinct from DFM's zero-shot Winogrande. The latter has known formatting/truncation and duplicate-shard issues in this evaluation pipeline and is not used in this table.
  • evaluation_results.json records the raw merged metric keys and values, table scaling, and source-artifact checksums for this release.

Usage instructions

Use a Transformers/vLLM build with native hrm_text / HrmTextForCausalLM support and the HRM PrefixLM attention implementation. A generic causal-only decoder or a build without this architecture is not an equivalent runtime. See the training and evaluation repository for the compatible export and serving path. No standalone remote model code is bundled here, matching the v1 package.

  • Use the supplied Gemma 4 tokenizer and chat_template.jinja; do not substitute a generic chat template. Keep enable_thinking=False when reproducing the non-thinking evaluation prompts.
  • Respect the 4,096-token context limit and PrefixLM prompt masking.
  • The uploaded files preserve the production export exactly, including fix_mistral_regex=true. Training tokenization used the unpatched tokenizer graph; a separate 2750K comparison established that changing this flag can change token IDs and scores. The training-integrated epoch-10 scores below are for the uploaded fix-enabled export, not a training-tokenizer/no-fix re-evaluation. Do not silently change the flag when comparing these published results.
  • Generated code is untrusted and should be executed only in a sandbox.

Technical report

The original model-family report is DFM Mimir v1. It describes the earlier release, not all additional data, training steps or results of v1.5. Training uses the HRM-Text fork.

Memorisation audit

The v1 report and v1 model card describe two memorisation audits of that release. Those numerical findings must not be treated as measurements of this later checkpoint. No new checkpoint-specific memorisation audit is supplied with v1.5.

Limitations

The model is primarily intended for Danish and English. It may hallucinate, produce incorrect code or reasoning, fail format constraints, and reproduce social biases. It is not guaranteed safe or reliable for high-stakes uses. The v1 safety and memorisation findings do not certify the v1.5 weights. Benchmark results depend on the tokenizer, template, generation limits, answer extraction and judge implementation.

License

This model is released under the Apache License 2.0. See LICENSE. The license and DFM logo are carried over from the original DFM-Mimir package.

Project partners & funding

The Mimir project was developed in collaboration between University of Southern Denmark, Aarhus University, University of Copenhagen and the Alexandra Institute, as part of Danish Foundation Models. The original project acknowledges funding from the Ministry of Science, Higher Education and Digital Affairs.

How to cite

For the model family, cite the v1 report and additionally identify this repository and its pinned revision when reporting v1.5 results:

@misc{schneiderkamp2026dfmmimirv1open,
  title={DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data},
  author={Peter Schneider-Kamp and Jacob Nielsen and Gianluca Barmina and Kenneth Enevoldsen and Lukas Galke Poech},
  year={2026},
  eprint={2608.13517},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2608.13517}
}
Downloads last month
456
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for danish-foundation-models/DFM-Mimir-v1.5

Quantizations
1 model

Paper for danish-foundation-models/DFM-Mimir-v1.5