DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
2026-08-13 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors developed Mimir v1, a language model with 1 billion parameters, that is trained using only permissible and ethically sourced data. It performs well in English and sets a new record for Danish language tasks, without relying on large, non-permissible datasets. They trained the model on 161 different datasets and found it competes with larger models on tests involving English, math, code, and Danish. Mimir v1 is openly available for other researchers to use and build on.
language modelparametersHierarchical Reasoning Modelpermissible dataopen-sourcebenchmarkEnglish NLPDanish NLPpost-training dataHugging Face Hub
Authors
Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech
Abstract
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir