Back to Feed
Training & Fine-Tuning / Benchmarks & Evals

Building Capable Models With Only Permissible Data

Original: DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 1 concepts

Key Takeaways

  • Mimir v1 achieves a 36.7% improvement in Math Code benchmarks compared to the baseline HRM-Text 1B model.
  • The model sets a new state of the art for Danish language benchmarks among the models tested.
  • Development focused on using only permissible, copyright-compliant training data to solve accessibility barriers in research.
  • Performance is achieved despite using a small 1B parameter footprint.

Summary & Methodology Analysis

The researchers developed Mimir v1 to address the dependency on non-permissible datasets that create high barriers for LLM research, particularly for low-resource languages like Danish. The training pipeline utilizes the HRM-Text framework and includes a curated mixture of 161 permissible datasets, totaling 70.5B tokens per epoch. To ensure data compliance, the team integrated synthetic transplant datasets, which are generated data points designed to replace non-compliant sources. The model was trained from scratch using the Gemma-4 tokenizer, chat templates, Fully Sharded Data Parallelism, and the AdamW optimizer to manage compute efficiency across the cluster. The training process also leveraged pre-norm layer normalization and Rotary Position Embedding, which is a technique for encoding the relative positions of tokens to maintain structural coherence in the sequence.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the main goal of this paper?

To build a capable, fully permissible foundation model that avoids copyright infringement and non-permissible data.

Q2. Does this model work for Danish?

Yes, it sets a new state of the art for Danish benchmarks among the models tested.

Q3. How does it compare to other models?

It outperforms the baseline HRM-Text 1B on English, Math, and Code tasks.

Q4. What architecture does Mimir v1 use?

It adopts the HRM-Text hierarchical reasoning architecture with a hidden size of 1,536, 12 attention heads, 2 H-cycles, and 3 L-cycles.

Q5. How much data was used during training?

The model used a mixture of 161 permissible datasets, which provided 70.5B tokens per epoch.

Q6. What specific tools were used during training?

The team used the Gemma-4 tokenizer, chat templates, Fully Sharded Data Parallelism, and the AdamW optimizer.

Q7. What are the limitations of this model?

Performance on Math and Code domains is still behind the 5B parameter Gemma 4 model, and its capabilities as an assistant are limited compared to the current state of the art.

Q8. Is the model faster than other versions?

The paper does not specify the inference latency or throughput compared to other models.

Q9. What were the exact scores on Math and Code?

Mimir v1 achieved a 64.1 score compared to 46.9 for the baseline, representing a 36.7% improvement.

Flag an issue

What is wrong with this summary?

What is wrong?