DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Summary
Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained on ethically sourced data that achieves competitive English performance and sets state-of-the-art results for Danish, matching larger frontier models.
View Cached Full Text
Cached at: 08/17/26, 03:48 PM
Paper page - DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Source: https://huggingface.co/papers/2608.13517
Abstract
Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained solely on permissible data that achieves competitive English results and state-of-the-art Danish performance across multiple benchmarks.
Current largelanguage modeldevelopment relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameterlanguage modelbased on theHierarchical Reasoning Model(HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissiblepost-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20benchmarksfor English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2608\.13517
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### danish-foundation-models/DFM-Mimir Text Generation• 2B• Updated2 days ago • 1.38k • 9
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.13517 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.13517 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Mimir: Did the vikings train a 1.7B killer model?
A new 1.7B parameter model called Mimir claims to outperform Qwen 3.5 and Gemma 4 E2B on benchmarks, with strengths in math and coding, and is based on Sapient's HRM-text architecture.
sapientinc/HRM-Text-1B
Sapient Intelligence released HRM-Text-1B, a 1-billion-parameter language model with a novel dual-timescale recurrent architecture (Hierarchical Reasoning Model) that provides unbounded compute depth at bounded parameter count. The pre-alignment checkpoint is available on Hugging Face.
@Sapient_Int: Introducing HRM-Text. An ultra-lean 1B-parameter reasoning language model designed to deliver strong general performanc…
Sapient Intelligence introduces HRM-Text, a 1B-parameter reasoning language model trained on only 40B tokens with a budget of $1,000, achieving competitive performance while drastically reducing data and compute requirements.
HRM-Text: Trained on only 1k$ and 40b tokens with brain inspired hierarchical latent architecture
HRM-Text is a 1B parameter text generation model that uses a brain-inspired hierarchical recurrent architecture to achieve efficient pretraining with only 40B tokens and ~$1000, enabling accessible foundation model training with dramatically reduced compute and data requirements.
A 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task
The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.