A Sovereign, Open-Source Foundation Model for German and English

arXiv cs.CL Models

Summary

The Soofi S 30B-A3B is a sovereign, open-source Mixture-of-Experts foundation model for German and English, pretrained on 27 trillion tokens. It matches dense 14-27B models on benchmarks while offering superior long-context throughput and outperforms all European open baselines.

arXiv:2607.09424v1 Announce Type: new Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.
Original Article
View Cached Full Text

Cached at: 07/13/26, 07:58 AM

# 1 Introduction
Source: [https://arxiv.org/html/2607.09424](https://arxiv.org/html/2607.09424)
![[Uncaptioned image]](https://arxiv.org/html/2607.09424v1/images/logo.png)

A Sovereign, Open\-Source Foundation Model for German and English

Soofi S Pretraining Report v1\.0

The Soofi\-Team\*

Core Team:Benedikt Droste10, David Fitzek3,9, Ruben Härle5, Lukas Helff2,5, Maximilian Idahl10, Alex Jude3,9, Abbas Goher Khan3, Maurice Kraus5, Timm Ruland3,9, Richard Rutmann3,9, Sebastian Sztwiertnia5

Contributors:Markus Frey3,9, Daniil Gurgurov2, Jan Pfister6, Tom Röhr7, Sebastian von Rohrscheidt7

Advisors:Jörg Bienert1, Nicolas Flores\-Herr3, Simon Gottschalk8, Andreas Hotho6, Kristian Kersting2,5,11, Joachim Köhler3, Alexander Löser7, Wolfgang Nejdl8, Simon Ostermann2, Jan Plogsties4, Patrick Putzky12

Technical Leads:Mehdi Ali3,9, Michael Fromm3,9, Max Lübbering3,9

Affiliations:1KI Bundesverband,2DFKI,3Fraunhofer IAIS,4Fraunhofer IIS,5Technische Universität Darmstadt,6Universität Würzburg,7Berliner Hochschule für Technik,8L3S Research Center,9Lamarr,10ellamind,11hessian\.AI,12Merantix Momentum

Coordination & Funding:Consortium coordinated by the KI Bundesverband\. Funded by the German Federal Ministry for Economic Affairs and Energy \(BMWE\)\.

∗\*Authors are listed alphabetically\. Detailed Contributions in Appendix[A](https://arxiv.org/html/2607.09424#A1)\.

1

Abstract

We present Soofi S 30B\-A3B, a sovereign, open\-source Mixture\-of\-Experts \(MoE\) hybrid Mamba Transformer foundation model for German and English\. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near\-constant as context grows, giving it a decisive throughput advantage over dense models for long\-context, high\-concurrency deployment\. Pretrained on roughly 27 trillion tokens with deliberately up\-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters\. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B\. Soofi S was built end\-to\-end on the German Industrial AI Cloud, a sovereign HPC\-scale AI infrastructure operated by Deutsche Telekom in Munich\. Soofi S will be released under highly permissive, open\-access terms: weights, selected intermediate checkpoints111[https://huggingface\.co/Soofi\-Project](https://huggingface.co/Soofi-Project), full per\-source data accounting, hyperparameters, and training and evaluation code\. Where source licenses permit, data\-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x1.png)\(a\)Capability vs\. measured aggregate decode TPS\.
![Refer to caption](https://arxiv.org/html/2607.09424v1/x2.png)\(b\)Aggregate decode TPS scaling with context\.

Figure 1:Long\-context serving efficiency\.Soofi S combines frontier\-level capability with the highest measured aggregate long\-context decode TPS, and unlike full\-attention dense baselines maintains high throughput as context grows\. Panel \([1\(a\)](https://arxiv.org/html/2607.09424#S1.F1.sf1)\) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32\. The Capability Index averages five benchmark groups, i\.e\., Code, GSM8K, GPQA\-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model\. Aggregate decode TPS/GPU is measured with a TP=1, one\-B200 vLLM latency\-subtraction protocol\. Panel \([1\(b\)](https://arxiv.org/html/2607.09424#S1.F1.sf2)\) shows measured aggregate decode TPS/GPU as a function of input context length under the same batch\-32 protocol\. For both panels higher is better\.Open language models have improved at remarkable speed, yet three gaps remain conspicuous for anyone deciding what to actually deploy\.

The first is openness in substance rather than name: despite a proliferation of capable models, the majority of releases remain weight\-only releases, documenting their training with little more than an aggregate token count and omitting the data, recipes, and decisions needed to reproduce or audit them\.

The second is language: general\-purpose multilingual models are either English\-centric or spread their capacity thinly across dozens of languages, leaving German underrepresented relative to its economic and scientific weight\. Dedicated European efforts to date\[[3](https://arxiv.org/html/2607.09424#bib.bib3),[26](https://arxiv.org/html/2607.09424#bib.bib26),[53](https://arxiv.org/html/2607.09424#bib.bib53),[5](https://arxiv.org/html/2607.09424#bib.bib5)\]have prioritized openness and language coverage over frontier capability\.

The third gap is the one that most directly governs the deployment cost, and where our chosen architecture is aimed\. At economic concurrency the price of generation is set not by how many parameters a model nominally contains, nor even by how many it activates per token, but by memory bandwidth: every decoded token must re\-read the model weights and, for a Transformer, the attention cache of every sequence in the batch\. As contexts grow into the tens or hundreds of thousands of tokens and many requests are served in parallel, this key–value \(KV\) cache comes to dominate, and full\-attention dense models slow down accordingly\. A model that keeps its per\-sequence state small and near\-constant in context length therefore enjoys a structural advantage that compounds in exactly the regimelong context, high concurrencythat matters most in production\.

Soofi S 30B\-A3B addresses all three at once: It is a Mixture\-of\-Experts \(MoE\) hybrid Mamba Transformer\[[61](https://arxiv.org/html/2607.09424#bib.bib61),[60](https://arxiv.org/html/2607.09424#bib.bib60)\]trained to excel in both German and English and will be released radically open: not weights alone, but the full set of artifacts required to audit every stage of training and, where source licenses permit, rebuild the data mixture, in the spirit of recent fully open efforts\[[5](https://arxiv.org/html/2607.09424#bib.bib5),[64](https://arxiv.org/html/2607.09424#bib.bib64)\]\. Architecturally it adopts the openly published Nemotron 3 Nano reference design\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\([Figure˜2](https://arxiv.org/html/2607.09424#S2.F2)\): Mamba\-2 layers\[[16](https://arxiv.org/html/2607.09424#bib.bib16)\]carry most of the sequence mixing with a fixed\-size recurrent state, only 6 of its 52 layers maintain a KV cache, and sparse MoE layers activate just 3\.2 of 31\.6 billion parameters per token—the capacity of a 30B network at roughly the inference cost of a 3B one\.

##### Contributions\.

We summarize our contributions below:

- •German–English champion\.Soofi S achieves best\-in\-class performance on aggregate English and German base model benchmarks, including strong performance on code, mathematics and German regional\-knowledge\. It is the strongest fully open model in our evaluation on English and German benchmarks, and matches or outperforms every European sovereign baseline in our comparison on every German benchmark in our suite, while matching dense 14–27B international models on English and German aggregate performance at a fraction of their active\-parameter cost \([Section˜4](https://arxiv.org/html/2607.09424#S4)\)\.
- •Full data transparency\.We release the complete pretraining corpus statistics \(LABEL:tab:phase1\-sources,LABEL:tab:phase2\-sources,[Table˜10](https://arxiv.org/html/2607.09424#A2.T10)\) and reproducible construction scripts222[https://github\.com/soofi\-project/Soofi\-Pretraining](https://github.com/soofi-project/Soofi-Pretraining)with*per\-source and per\-language token accounting*, distinguishing dataset\-card estimates from tokenizer\-exact consumed\-token counts\. We also provide the German : English : code mixing ratio, and the rationale behind it, in contrast to reports that disclose only an aggregate token count\.
- •Reproducible recipe\.We publish the full learning\-rate schedule \(Warmup–Stable–Decay\), optimizer, all hyperparameters, the per\-phase token budgets, and the phase boundaries, so a third party can rebuild the run\.
- •Long\-context serving efficiency\.The hybrid Mamba–MoE design keeps the per\-sequence cache near\-constant in context length, yielding measured aggregate decode TPS/GPU88–9×9\\timesthat of dense 14–24B models at 40K context and batch 32, and aggregate decode TPS that stays essentially flat from 4K to 256K where full\-attention models degrade \([Figure˜1](https://arxiv.org/html/2607.09424#S1.F1),[Section˜4\.3](https://arxiv.org/html/2607.09424#S4.SS3)\)\.
- •Documented design\.We report the data and design ablations behind each choice, exposing the*why*rather than only the final configuration\.
- •

The remainder of this report documents Soofi S 30B\-A3B Base end to end\.[Figure˜2](https://arxiv.org/html/2607.09424#S2.F2)presents the model and the training recipe: the Nemotron 3 Nano reference architecture we adopt and the rationale for adopting it, the Warmup–Stable–Decay optimization schedule, all hyperparameters and per\-phase token budgets, the training dynamics of the run, and the compute infrastructure\.[Section˜3](https://arxiv.org/html/2607.09424#S3)details the three\-phase, 26\.68T\-token German–English data curriculum, with full per\-source token accounting for every stage\.[Section˜4](https://arxiv.org/html/2607.09424#S4)then evaluates the resulting base model against 16 open models of comparable or larger active size along two axes: capability across parallel English and German benchmarks, and serving efficiency\. Section[5](https://arxiv.org/html/2607.09424#S5)situates the work relative to prior open and European efforts, and Section[6](https://arxiv.org/html/2607.09424#S6)concludes\.

## 2Model Architecture and Training

![Refer to caption](https://arxiv.org/html/2607.09424v1/x3.png)Figure 2:Training dynamics over the full∼27​T\{\\sim\}27\\text\{T\}\-token run\.The quantity is plotted against the number of consumed tokens \(in trillions\); the pretraining\-to\-annealing transition occurs at∼20​T\{\\sim\}20\\text\{T\}tokens\. The solid line is a rolling median over1,0001\{,\}000steps and the faint trace is the raw per\-step signal\.Soofi S 30B\-A3B Base adopts the hybrid Mamba–Transformer Mixture\-of\-Experts \(MoE\) reference architecture of Nemotron 3 Nano\[[61](https://arxiv.org/html/2607.09424#bib.bib61),[60](https://arxiv.org/html/2607.09424#bib.bib60)\]: a 52\-layer network interleaving 23 Mamba\-2 sequence\-mixing layers\[[16](https://arxiv.org/html/2607.09424#bib.bib16)\], 23 granular MoE layers with shared experts\[[77](https://arxiv.org/html/2607.09424#bib.bib77),[19](https://arxiv.org/html/2607.09424#bib.bib19),[15](https://arxiv.org/html/2607.09424#bib.bib15),[41](https://arxiv.org/html/2607.09424#bib.bib41)\], and 6 Grouped\-Query Attention \(GQA\) layers\[[1](https://arxiv.org/html/2607.09424#bib.bib1)\]distributed sparsely through the network depth\.

The model totals∼31\.6\{\\sim\}31\.6B parameters, of which only∼3\.2\{\\sim\}3\.2B are active per forward pass \(∼3\.6\{\\sim\}3\.6B including embeddings\), and only the 6 GQA layers maintain a KV cache\. Because we reuse the reference design without modification, we refer the reader to the Nemotron reports\[[60](https://arxiv.org/html/2607.09424#bib.bib60),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]for the design motivation and the ablations behind each architectural choice; Table[1](https://arxiv.org/html/2607.09424#S2.T1)records the exact configuration for reproducibility\.

##### Why a reference architecture\.

Reusing an established, openly specified architecture rather than designing a bespoke one was a deliberate decision, on three grounds\. First,*deployability*: the Nemotron 3 Nano architecture is already integrated into the major open inference and serving stacks, including the vLLM stack used for our own serving measurements \([Section˜4\.3](https://arxiv.org/html/2607.09424#S4.SS3)\), with mature, heavily optimized kernels for its Mamba\-2, GQA, and MoE components, so Soofi S can be hosted efficiently by existing software from the day of release, without bespoke integration work by downstream users\. Second,*serving efficiency*: the predominantly Mamba\-2 backbone makes the architecture exceptionally fast in exactly the regime we target, prefill and decode costs grow near\-linearly in sequence length and the per\-sequence cache stays near\-constant, which is what underpins the long\-context throughput results of[Section˜4\.3](https://arxiv.org/html/2607.09424#S4.SS3)and makes the 1M\-token context extension of[Section˜3\.4](https://arxiv.org/html/2607.09424#S3.SS4)practical\. Third,*scientific control*: sharing the backbone with Nemotron 3 Nano turns that model into an architecture\-identical baseline, so the effect of our German–English data recipe can be measured in isolation \([Section˜4\.2](https://arxiv.org/html/2607.09424#S4.SS2)\)\.

Table 1:Soofi S 30B\-A3B architecture, following the Nemotron 3 Nano reference configuration\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. The layer pattern interleaves Mamba\-2 and MoE layers, with 6 GQA layers distributed through the network depth\.
### 2\.1Optimization and Hyperparameters

We train Soofi S with the Megatron\-Bridge framework777For this training run, we adopted Megatron\-Bridge \([https://github\.com/NVIDIA/Megatron\-Bridge](https://github.com/NVIDIA/Megatron-Bridge)\) rather than integrating the Nemotron\-3 reference architecture into our open\-source stack\[[50](https://arxiv.org/html/2607.09424#bib.bib50)\], as Megatron\-Bridge already provided mature, reliable support for the newly released architecture\.using AdamW\[[49](https://arxiv.org/html/2607.09424#bib.bib49)\]under a Warmup–Stable–Decay \(WSD\)\[[35](https://arxiv.org/html/2607.09424#bib.bib35),[30](https://arxiv.org/html/2607.09424#bib.bib30)\]learning\-rate schedule whose decay segment follows aminus\_sqrtshape\. With the exception of the long\-context phase \(Section[2\.2](https://arxiv.org/html/2607.09424#S2.SS2)\), all stages share the same parallelism and batching configuration: tensor\-model\-parallel \(TP\) size11, expert\-parallel \(EP\) size 8, sequence parallelism disabled, micro\-batch size22, and global batch size30723072at a sequence length of81928192tokens, with the remaining GPUs used for data parallelism and a distributed \(optimizer\-state\-sharded\) optimizer\. This corresponds to25,165,82425\{,\}165\{,\}824tokens per optimizer step\. All stages are trained in bf16 mixed precision\. Since the granular MoE layers route each token to experts held on different expert\-parallel ranks, all\-to\-all communication sits on the critical path, which motivates the node topology and interconnect reported in Section[2\.4](https://arxiv.org/html/2607.09424#S2.SS4)\. Manual garbage collection is triggered every101101iterations\. The schedule is realised across one warmup\-plus\-stable phase and three successive annealing continuations, summarised in Table[2](https://arxiv.org/html/2607.09424#S2.T2); the data mixture consumed by each stage is documented in[Section˜3](https://arxiv.org/html/2607.09424#S3):

- •Base pretraining \(stable\)\.794,728794\{,\}728iterations \(19,999,984,975,87219\{,\}999\{,\}984\{,\}975\{,\}872tokens,∼20\{\\sim\}20T\) on the 20T EN/DE mixture \(see[Section˜3\.2](https://arxiv.org/html/2607.09424#S3.SS2)\)\. After a254254\-iteration warmup to the peak learning rate of1​e−31\\mathrm\{e\}\{\-\}3, training proceeds on the WSD stable plateau\.
- •Main annealing \(decay\)\.A continuation of198,682198\{,\}682iterations \(4,999,996,243,9684\{,\}999\{,\}996\{,\}243\{,\}968tokens,∼5\{\\sim\}5T\) on the 5T high\-quality EN/DE mixture \(see[Section˜3\.3](https://arxiv.org/html/2607.09424#S3.SS3)\), applying the WSDminus\_sqrtdecay from1​e−31\\mathrm\{e\}\{\-\}3to1​e−51\\mathrm\{e\}\{\-\}5\. The configured total after this continuation is993,410993\{,\}410iterations\.
- •Constant annealing\.Since the slope at the end of the decay segment remained relatively steep, we append a further62,59062\{,\}590iterations \(1,575,128,924,1601\{,\}575\{,\}128\{,\}924\{,\}160tokens \(see[Section˜3\.3](https://arxiv.org/html/2607.09424#S3.SS3)\),∼1\.58\{\\sim\}1\.58T\) at a*constant*learning rate of1​e−51\\mathrm\{e\}\{\-\}5on the same 5T mixture, allowing additional training at the tail of the annealing curve\.
- •Final annealing \(discarded\)\.A final WSDminus\_sqrtdecay of11,92011\{,\}920iterations \(∼0\.30\{\\sim\}0\.30T, see[Section˜3\.3](https://arxiv.org/html/2607.09424#S3.SS3)\) from1​e−51\\mathrm\{e\}\{\-\}5to0, beginning at iteration reference1,056,0001\{,\}056\{,\}000and targeting1,067,9201\{,\}067\{,\}920\. We report this stage for completeness but*do not*use its checkpoints, as they showed no clear additional benchmark improvement over the constant annealing stage\.

Table 2:Training stages and learning\-rate schedule\. “Iters” are the iterations added in each stage; all88K\-context stages use25,165,82425\{,\}165\{,\}824tokens per iteration\. The final annealing stage was run but its checkpoints were not used\. The long\-context stage uses a distinct configuration \(Section[2\.2](https://arxiv.org/html/2607.09424#S2.SS2)\)\.StageRoleItersTokensLR \(peak→\\tomin\)Base pretrainingStable \(warmup254254\)794,728∼20\{\\sim\}20T1​e−31\\mathrm\{e\}\{\-\}3\(plateau\)Main annealingWSD decay \(minus\_sqrt\)198,682∼5\{\\sim\}5T1​e−3→1​e−51\\mathrm\{e\}\{\-\}3\\to 1\\mathrm\{e\}\{\-\}5Constant annealingConstant LR62,590∼1\.58\{\\sim\}1\.58T1​e−51\\mathrm\{e\}\{\-\}5Final annealing†WSD decay \(minus\_sqrt\)11,920∼0\.30\{\\sim\}0\.30T1​e−5→01\\mathrm\{e\}\{\-\}5\\to 0Long contextConstant \(warmup100100\)2,000∼0\.10\{\\sim\}0\.10T1​e−5→1​e−71\\mathrm\{e\}\{\-\}5\\to 1\\mathrm\{e\}\{\-\}7†Run for completeness; checkpoints not used \(no clear benchmark gain\)\.
### 2\.2Long\-Context Extension

The long\-context phase \(Phase 3, Section[3\.4](https://arxiv.org/html/2607.09424#S3.SS4)\) is trained with a distinct parallelism configuration to accommodate the11M\-token sequences\. We use a context\-parallel size of1616, a micro\-batch size of11, and a global batch size of4848at a sequence length of1,048,5761\{,\}048\{,\}576tokens, giving50,331,64850\{,\}331\{,\}648tokens per optimizer step\. The stage runs for2,0002\{,\}000iterations, for an approximate total of100\.66100\.66B tokens\. Optimization again uses AdamW\. The learning rate follows a constant schedule with a100100\-iteration warmup, holding at1​e−51\\mathrm\{e\}\{\-\}5\(minimum1​e−71\\mathrm\{e\}\{\-\}7\)\. Two settings distinguish this phase from the earlier stages: HybridEP is enabled \(it was not used previously\), and we apply*no*intra\-document masking, matching NVIDIA’s long\-context recipe\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]rather than the alternative adopted by some other efforts\.

### 2\.3Training Dynamics

Figure[2](https://arxiv.org/html/2607.09424#S2.F2)summarizes the four central signals we logged over the full∼27\{\\sim\}27T\-token run, each plotted against the number of consumed tokens\. The pretraining\-to\-annealing transition at∼20\{\\sim\}20T tokens—the end of the stable plateau and the start of the main WSD decay \(Table[2](https://arxiv.org/html/2607.09424#S2.T2)\)—is the reference point for reading all four panels\. For the loss, gradient\-norm, and throughput traces, the solid line is a rolling median over1,0001\{,\}000steps and the faint trace is the raw per\-step signal; loss and gradient norm are clipped at2\.02\.0for readability\.

##### Training logs\.

In addition to the static training\-dynamics plot in Figure[2](https://arxiv.org/html/2607.09424#S2.F2), we provide the corresponding Weights & Biases dashboard for the full pretraining run888[https://api\.wandb\.ai/links/soofi\-exchange/j11vi7rg](https://api.wandb.ai/links/soofi-exchange/j11vi7rg)\. The dashboard contains the raw and smoothed traces used to inspect optimization stability, throughput, checkpointing effects, and the transitions between the base\-pretraining, annealing, and long\-context phases\.

The learning\-rate schedule \(Figure[2](https://arxiv.org/html/2607.09424#S2.F2)a\) follows the Warmup–Stable–Decay \(WSD\) shape: a short linear warmup to the peak of1×10−31\\times 10^\{\-3\}, a constant plateau held throughout base pretraining, and aminus\_sqrtdecay toward∼0\{\\sim\}0once annealing begins\. The language\-modeling loss \(Figure[2](https://arxiv.org/html/2607.09424#S2.F2)b\) declines steadily across the stable phase and then drops sharply at the onset of annealing due to the dataset switch, tracking the learning\-rate decay as the model is concentrated on the high\-quality mixture\. The gradient norm \(Figure[2](https://arxiv.org/html/2607.09424#S2.F2)c\) remains stable across the entire run, with only a mild increase during the annealing phase and no divergence or sustained spikes\. Training throughput \(Figure[2](https://arxiv.org/html/2607.09424#S2.F2)d\), measured in tokens per second, is steady for most of the run; the downward excursions correspond to checkpointing, evaluation, and restart iterations rather than to changes in the training dynamics themselves\.

### 2\.4Compute Infrastructure

Soofi S was trained on the Industrial AI Cloud\[[17](https://arxiv.org/html/2607.09424#bib.bib17),[62](https://arxiv.org/html/2607.09424#bib.bib62)\]operated by Deutsche Telekom in Munich built together with NVIDIA and brought into operation in February 2026\. Our run used up to 512 NVIDIA B200 GPUs, i\.e\. 64 DGX B200 nodes of 8 GPUs each\. Within a node the 8 B200s are fully connected by fifth\-generation NVLink/NVSwitch, while nodes are interconnected by an eight\-rail NVIDIA Quantum\-2 NDR InfiniBand fabric, one 400 Gb/s ConnectX\-7 port per GPU,3\.23\.2Tb/s of scale\-out bandwidth per node\. Keeping expert\-parallel groups within a node lets the MoE all\-to\-all \(Section[2\.1](https://arxiv.org/html/2607.09424#S2.SS1)\) run over intra\-node NVLink rather than the slower inter\-node fabric, which matters for MoE throughput at this scale\. The run took place from 24 March 2026 to 13 May 2026 and consumed approximately 253,000 B200 GPU\-hours across the stable, annealing, and long\-context stages\. Training on this infrastructure is part of the sovereignty of the model: it was executed on German soil under European operational and data\-protection requirements rather than on extra\-European hyperscale compute\. Soofi S was one of the first flagship workloads on the Industrial AI Cloud, whose infrastructure was procured for the sovereign open\-source foundation\-model effort under which this model was developed\[[17](https://arxiv.org/html/2607.09424#bib.bib17)\]\. The Munich facility is powered entirely by renewable energy, designed for high energy efficiency, cooled with water drawn from the nearby Eisbach canal, and integrated with a waste\-heat\-reuse concept that feeds the surrounding Tucherpark district\.

## 3Pretraining Data

![Refer to caption](https://arxiv.org/html/2607.09424v1/x4.png)Figure 3:Effective\-token mixture across the three training phases\. A single flow diagram tracing seven data categories \(English Web, Academic & Wiki, SFT, Reasoning, Code, Math, and German\) from left to right across the phases\.Phase 1\(diverse pretraining\),23,051\.1323\{,\}051\.13B effective tokens\.Phase 2\(high\-quality annealing\),6,303\.06\{,\}303\.0B effective tokens, showing increased density of skill\-oriented and German data relative to Phase 1\.Phase 3\(long\-context extension\),188188B effective tokens, where the SFT band branches into its General, Code, and Math SFT components\.Consistent with our commitment to full reproducibility, we document the pretraining corpus of Soofi S at the granularity of individual source datasets\. For every constituent, we report its public identifier, its raw token count, the number of epochs it was repeated, the resulting effective token count, and its share of the phase\. We deliberately also list sources that were enumerated but*excluded*from training \(epoch count of zero\), so that the mixture can be audited and rebuilt end to end\. This stands in contrast to the common practice of disclosing only an aggregate token count for the training data, and follows the openness ethos of recent fully open efforts\[[5](https://arxiv.org/html/2607.09424#bib.bib5),[64](https://arxiv.org/html/2607.09424#bib.bib64)\]\.

Soofi S is trained on a three\-phase curriculum, consistent with the Warmup–Stable–Decay \(WSD\) learning\-rate schedule \([Section˜2\.1](https://arxiv.org/html/2607.09424#S2.SS1)\)\. Phase 1 \(see[Section˜3\.2](https://arxiv.org/html/2607.09424#S3.SS2)\) maximizes diversity over a large, quality\-tiered mixture of web, synthetic, code, mathematics, and multilingual data\. Phase 2 \([Section˜3\.3](https://arxiv.org/html/2607.09424#S3.SS3)\) is an annealing \(decay\) phase that concentrates the highest\-quality web data together with skill\-focused code, mathematics, STEM, reasoning, and instruction data, while further up\-weighting German to15\.3%15\.3\\%of the constructed annealing pool\. Phase 3 \([Section˜3\.4](https://arxiv.org/html/2607.09424#S3.SS4)\) extends the usable context length up to 1M tokens via length\-bucketed up\-sampling\. Table[3](https://arxiv.org/html/2607.09424#S3.T3)summarizes the token budget of each phase\. Across all phases, the corpus comprises approximately 27 trillion tokens, of which a deliberately elevated fraction is German, reflecting the design goal of a German–English model rather than a broadly multilingual one\. The full per\-source composition of every phase is documented in Appendix[B](https://arxiv.org/html/2607.09424#A2); the figure in this section summarize those tables as a mixture flow diagram\.

Table 3:Three\-phase pretraining curriculum and token budget\. The*Pool*column is the effective\-token mixture constructed for each phase \(documented per source in Appendix[B](https://arxiv.org/html/2607.09424#A2)\); for Phases 1–2 these are dataset\-card counts and are approximate, whereas the Phase 3 pool is tokenized with our own tokenizer\. The*Consumed*column is the number of tokens actually trained on, counted exactly from the optimizer schedule \(iterations×\\timestokens/iteration; Table[2](https://arxiv.org/html/2607.09424#S2.T2)\)\. A phase may consume less than one epoch of its pool \(Phase 1\) or more than one \(Phase 2\); the headline∼27\{\\sim\}27T figure refers to consumed tokens\.### 3\.1Quality Tiers, Synthetic Data, and Epoching

Most of the web data is drawn from the openly released Nemotron\-CC datasets\[[80](https://arxiv.org/html/2607.09424#bib.bib80),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\(v1\.0, v2\.0, and v2\.1\), which provide documents pre\-sorted into quality tiers \(High, Medium\-High, Medium\) and include several synthetically rephrased variants \(\-Synthetic\), diverse question–answer reformulations \(Diverse QA/DQA\), and English translations of non\-English documents \(Translated\-To\-English\)\. We exploit these tiers directly: the highest\-quality and synthetic tiers are repeated for multiple epochs to increase their effective contribution, whereas the large Medium\-Quality pools are listed but set to zero epochs and thus excluded from the final mixture\. An epoch count greater than one, therefore, denotes deliberate up\-sampling of a high\-value source, and an epoch count of zero denotes a source we evaluated but chose not to train on\. All effective\-token figures and shares below are reported*after*applying these epoch multipliers\.

##### Token accounting and its precision\.

For Phases 1 and 2, the raw and effective token counts below are taken from the token statistics reported on each source’s public \(HuggingFace\) dataset card\. Since different datasets are tokenized with different tokenizers, these counts are not all expressed in our model’s tokens; they should be read as close approximations that document the*relative composition*of the mixture rather than an exact token ledger\. By contrast, all per\-iteration training\-token figures \(Section[2\.1](https://arxiv.org/html/2607.09424#S2.SS1), Table[2](https://arxiv.org/html/2607.09424#S2.T2)\) and the entire Phase 3 long\-context pool \(Section[3\.4](https://arxiv.org/html/2607.09424#S3.SS4), Tables[10](https://arxiv.org/html/2607.09424#A2.T10)–[11](https://arxiv.org/html/2607.09424#A2.T11)\) were obtained by tokenizing the data with the Nemotron\-3 tokenizer and are therefore exact\. This distinction also explains why the constructed\-pool totals reported below \(e\.g\.∼23\.05\{\\sim\}23\.05T effective for Phase 1\) differ from the exactly\-counted tokens actually consumed during training \(e\.g\.∼20\{\\sim\}20T; Table[3](https://arxiv.org/html/2607.09424#S3.T3)\): the former are dataset\-card estimates of the*available*pool, the latter are exact counts of what the optimizer*saw*, and a phase may consume less than one epoch of its pool \(Phase 1\) or slightly more than one \(Phase 2\)\.

### 3\.2Phase 1: Diverse Pretraining

Phase 1 provides the bulk of the training signal\. Its composition is given in full inLABEL:tab:phase1\-sources, grouped by source family with per\-family subtotals\. The phase totals∼\\sim16\.35T raw tokens, which after epoching yield∼\\sim23\.05T effective tokens in dataset\-card terms; the run consumes∼\\sim20T of this pool \(Section[2\.1](https://arxiv.org/html/2607.09424#S2.SS1)\)\. The mixture is anchored on quality\-filtered and synthetic web text from the three Nemotron\-CC releases \(CC\-v2\.1, CC\-v2\.0, and CC\-v1\.0 contribute∼\\sim2\.84T,∼\\sim5\.60T, and∼\\sim3\.15T effective tokens respectively\), supplemented by a large code component \(∼\\sim3\.38T across the Nemotron code datasets\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\), specialized STEM and scientific data \(Nemotron\-Pretraining\-Specialized\-v1\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\],∼\\sim1\.35T\), pretraining\-stage SFT mixtures \(∼\\sim1\.57T of Math, Code, and General SFT\), and dedicated mathematics corpora \(∼\\sim1\.06T from Nemotron\-CC\-Math v1/v2\[[51](https://arxiv.org/html/2607.09424#bib.bib51)\]and a further352352B from UltraData\-Math\[[91](https://arxiv.org/html/2607.09424#bib.bib91)\]\)\. PDF\-derived text from FinePDFs\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]and Dolma3\_POOL\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]adds high\-value long\-form document data\.[Figure˜3](https://arxiv.org/html/2607.09424#S3.F3)\(Left\) summarises the Phase 1 mixture by source family\.

For the German–English objective, German is intentionally over\-represented relative to the base Nemotron recipe: German sources contribute∼\\sim1\.65T effective tokens, or7\.2%7\.2\\%of Phase 1, against the5%5\\%multilingual share of the reference Nemotron 3 Nano mixture\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. The German component combines naturally occurring web and document text \(HPLT Monolingual Datasets 3\.0\[[63](https://arxiv.org/html/2607.09424#bib.bib63)\], German Commons\[[24](https://arxiv.org/html/2607.09424#bib.bib24)\], Genios\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\], the German subset of FinePDFs\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]and FineWiki\[[67](https://arxiv.org/html/2607.09424#bib.bib67)\]\) with machine\-translated \(MT\) and synthetic German \(MultiSynt/MT\[[37](https://arxiv.org/html/2607.09424#bib.bib37)\], MT\-Reasoning\[[25](https://arxiv.org/html/2607.09424#bib.bib25),[58](https://arxiv.org/html/2607.09424#bib.bib58)\], MT of Nemotron\-Multilingual\-Reasoning\[[29](https://arxiv.org/html/2607.09424#bib.bib29)\], PleIAs/Synth\[[69](https://arxiv.org/html/2607.09424#bib.bib69)\]\)\. Table[7](https://arxiv.org/html/2607.09424#A2.T7)compares our full Phase 1 mixture against the published Nemotron 3 Nano mixture; relative to it we raise German and academic share, lean more heavily on high\-quality web, and trim synthetic web and code\-SFT\.

### 3\.3Phase 2: High\-Quality Annealing

The annealing phase coincides with the learning\-rate decay and is composed exclusively of high\-value data\. The per\-source breakdown is given in TableLABEL:tab:phase2\-sources, organised by category with per\-category subtotals; the constructed pool totals∼6\.30\{\\sim\}6\.30T effective tokens \(dataset\-card counts\), of which the annealing schedule trains for∼6\.58\{\\sim\}6\.58T\. Relative to Phase 1, we drop the lower\-tier web pools entirely, retain only the High\-Quality and High\-Quality\-Synthetic web tiers, and substantially increase the density of skill\-oriented data: code \(∼\\sim1\.03T\), mathematics \(∼\\sim330B\), and a broad SFT mixture \(∼\\sim0\.93T\) spanning math, code, agentic, competitive\-programming, instruction\-following, science, finance, software\-engineering, safety, and multilingual subsets\. A dedicated reasoning bucket—Nemotron\-Pretraining\-Specialized\-v1 and related sources \(∼\\sim353B\)—strengthens chain\-of\-thought ability ahead of post\-training\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. English Web accounts for36\.4%36\.4\\%of the annealing mixture; including Academic & Wiki, English web/document data accounts for42\.8%42\.8\\%\. Figure[3](https://arxiv.org/html/2607.09424#S3.F3)\(Mid\) summarizes the Phase 2 mixture by category\.

German is up\-weighted again during the annealing phase\. The German category alone contributes965\.37965\.37B effective tokens, drawn from a pre\-release version of HPLT\-4999[https://hplt\-project\.org/datasets/v4\.0](https://hplt-project.org/datasets/v4.0), a German translation of ClimbMix\[[18](https://arxiv.org/html/2607.09424#bib.bib18)\]produced with the KletterMix pipeline\[[42](https://arxiv.org/html/2607.09424#bib.bib42)\], German FinePDFs\-Edu\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]and FineWiki\[[67](https://arxiv.org/html/2607.09424#bib.bib67)\], and synthetic/translated German reasoning sources; this brings the multilingual share to15\.32%15\.32\\%, more than triple the5%5\\%of the reference mixture\.[Table˜9](https://arxiv.org/html/2607.09424#A2.T9)reports the category\-level composition of the annealing phase against Nemotron 3 Nano\.

As part of the SFT mixture, we additionally include QA\-base \(∼0\.05%\{\\sim\}0\.05\\%of the pool; the1\.431\.43B English tokens are counted under the SFT category and the1\.871\.87B German tokens under the German category in TableLABEL:tab:phase2\-sources\), paraphrased training splits of 25 standard NLP benchmarks in English and German, analogous to the paraphrase\-augmented benchmark training data in Olmo 3’s mid\-training mix \(e\.g\. TinyMATH, Dolmino Flan\)\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]and the benchmark\-seeded synthetic data in Nemotron 3\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\.

### 3\.4Phase 3: Long\-Context Extension

To extend the usable context window up to 1M tokens, we assemble a long\-context data pool of approximately188\.5B tokensdrawn from∼\\sim21 million documents, partitioned into nine sequence\-length buckets \(4K, 8K, 16K, 32K, 64K, 128K, 256K, 512K, and 1M tokens\)\. The schema targets balanced exposure across context lengths by allocating comparable token mass to each bucket; in practice the realized mass per bucket varies with source availability, and several buckets at the extremes are sparsely populated or empty \(Table[11](https://arxiv.org/html/2607.09424#A2.T11)\)\. Seven domains contribute to the pool: general web \(28\.78B effective tokens\), code \(9\.51B\), mathematics \(6\.73B\), German \(6\.68B\), and three supervised\-fine\-tuning streams, general SFT \(31\.25B\), code SFT \(29\.24B\), and mathematics SFT \(76\.31B\), which together account for roughly73%73\\%of the pool\. Table[10](https://arxiv.org/html/2607.09424#A2.T10)lists the per\-domain token budgets, document counts, and source priorities; Table[11](https://arxiv.org/html/2607.09424#A2.T11)gives the per\-bucket document counts\. Unlike Phases 1–2, every Phase 3 token count was produced by tokenizing the data with the Nemotron\-3 tokenizer, so the figures in this section and in Tables[10](https://arxiv.org/html/2607.09424#A2.T10)–[11](https://arxiv.org/html/2607.09424#A2.T11)are exact rather than dataset\-card estimates\.

The long\-context stage trained for 2,000 optimizer steps at a 1M\-token sequence length, i\.e\. approximately 100\.66B tokens \(Section[2\.2](https://arxiv.org/html/2607.09424#S2.SS2)\)—about53%53\\%of the pool, or roughly half an epoch\. We stopped at this point because no further loss improvement was observed from additional long\-context training, mirroring the rationale for the discarded final\-annealing stage \(Section[2\.1](https://arxiv.org/html/2607.09424#S2.SS1)\)\.

When a bucket draws from several sources, we fill it according to a fixed priority order: for web, ClimbMix\[[18](https://arxiv.org/html/2607.09424#bib.bib18)\]is preferred over OlmoOCR\[[70](https://arxiv.org/html/2607.09424#bib.bib70)\], which is preferred over FinePDFs\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]; for code, SwallowCode\[[20](https://arxiv.org/html/2607.09424#bib.bib20)\]is preferred over Nemotron\-Pretraining\-Code\-v1/v2\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]; and German is composed of 40% HPLTv4 and 60% German translation of ClimbMix\[[18](https://arxiv.org/html/2607.09424#bib.bib18)\]with the pipeline of KletterMix\[[42](https://arxiv.org/html/2607.09424#bib.bib42)\]\.

We note the population irregularities in the interest of full disclosure\. Our mathematics sources do not provide documents beyond the 64K bucket, so the longer mathematics buckets are effectively unpopulated \(2 documents at 128K, none beyond\)\. The web and German domains have no populated 256K bucket, and the code domain’s 4K and 8K buckets are effectively empty \(whole\-file repository code is scarce at these lengths after packing\); these short sequence lengths are instead carried by the web, mathematics, and SFT streams, the latter being densest at 8K–64K\. The SFT mathematics buckets nominally extend to 256K but are negligible above 64K \(19 and 2 documents at 128K and 256K, respectively\)\.

### 3\.5Data Provenance and Release

The overwhelming majority of our pretraining data is openly available\. The web, code, mathematics, specialized, and SFT components build on NVIDIA’s publicly released Nemotron pretraining datasets\[[80](https://arxiv.org/html/2607.09424#bib.bib80),[51](https://arxiv.org/html/2607.09424#bib.bib51),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\(Nemotron\-CC v1\.0/v2\.0/v2\.1, Nemotron\-CC\-Math, Nemotron\-CC\-Code, Nemotron\-Pretraining\-Code/Specialized/SFT, and the Nemotron post\-training SFT collections\), complemented by open corpora including the Dolma 3 pools\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\], the Fine\* family\[[68](https://arxiv.org/html/2607.09424#bib.bib68),[44](https://arxiv.org/html/2607.09424#bib.bib44)\]\(FinePDFs, FinePDFs\-edu, FineWiki\), ClimbMix\[[18](https://arxiv.org/html/2607.09424#bib.bib18)\], SwallowCode\[[20](https://arxiv.org/html/2607.09424#bib.bib20)\], UltraData\-Math\[[91](https://arxiv.org/html/2607.09424#bib.bib91)\], and AceReason\[[48](https://arxiv.org/html/2607.09424#bib.bib48)\]\. German coverage is provided by open resources such as HPLT \(v3 and v4\)\[[11](https://arxiv.org/html/2607.09424#bib.bib11),[63](https://arxiv.org/html/2607.09424#bib.bib63)\], German\-Commons\[[24](https://arxiv.org/html/2607.09424#bib.bib24)\], KletterMix\[[42](https://arxiv.org/html/2607.09424#bib.bib42)\], and the Fine\* German subsets, augmented with machine\-translated and synthetic German data \(the MultiSynt/MT\-\* sources, Soofi\-Think\-SFT\[[28](https://arxiv.org/html/2607.09424#bib.bib28)\]\), plus the commercially licensed Genios\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]corpus\. Consistent with the openness goals of this work, we release the complete mixture specification; every source, its raw token count, epoch multiplier, and effective contribution for all three phases, so that the corpus can be independently reconstructed where source licenses permit; for commercially licensed Genios, we release aggregate statistics and exact mixture accounting rather than redistributing raw text\.

##### Openness classification\.

The definition of open\-source AI remains contested: the OSI’s Open Source AI Definition 1\.0\[[65](https://arxiv.org/html/2607.09424#bib.bib65)\]permits documented but unsharable training data, whereas stricter proposals for a European definition require every training token to be redistributable\[[46](https://arxiv.org/html/2607.09424#bib.bib46)\]\. Soofi S satisfies OSAID 1\.0: we will release weights, intermediate checkpoints, training and evaluation code, and exact per\-source data accounting under permissive licenses\. Under the stricter open\-data standard, Soofi S falls short in exactly one documented component: the commercially licensed Genios\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]corpus \(1\.3%1\.3\\%of Phase 1 effective tokens, reported in aggregate in Appendix[D](https://arxiv.org/html/2607.09424#A4)\)\. Every other source is publicly obtainable, so∼99%\{\\sim\}99\\%of the mixture can be independently reconstructed\. We state this boundary explicitly rather than claim a stronger openness status than the release supports\.

#### 3\.5\.1Code Web Data

Our code\-from\-web component is Nemotron\-CC\-Code\-v1\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\], a corpus of code and code\-adjacent documents recovered directly from Common Crawl rather than from repository hosting platforms\. Standard web\-extraction pipelines tend to mangle source code embedded in HTML, collapsing the whitespace and indentation that is syntactically meaningful in languages such as Python, so this corpus is built with a code\-aware extraction pipeline that preserves the formatting of code blocks, retains the surrounding natural\-language context \(tutorials, documentation, Q&A threads, blog posts\), and applies dedicated cleaning, language identification, and deduplication stages\. The result is code in its*instructional habitat*: implementations interleaved with the prose that explains them, which is a complementary signal to raw repository files\. We use theActualsubset at 3 epochs \(427\.9427\.9B raw,1,283\.71\{,\}283\.7B effective tokens; TableLABEL:tab:phase1\-sources\), making it the single largest code source in Phase 1\.

#### 3\.5\.2Curated Code Data

Repository\-sourced code is drawn from Nemotron\-Pretraining\-Code\-v1\[[59](https://arxiv.org/html/2607.09424#bib.bib59)\]and \-v2\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\], which curate permissively licensed source code from public repositories with quality filtering, per\-language balancing, and aggressive deduplication\. Both releases pair the curatedActualcode with synthetic derivatives generated from it, and we train on five such variants from v2:*code review*\(critiques and improvement suggestions for real code\),*question answering*\(Q&A pairs grounded in repository code\),*rewriting*\(refactorings and alternative implementations of the same functionality\),*student–teacher*\(dialogues that explain code step\-by\-step\), and*transpilation*\(translations of programs between programming languages\)\. These variants convert passive code exposure into bidirectional code–language supervision during pretraining itself\. In Phase 1, the curated\-code family contributes2,091\.142\{,\}091\.14B effective tokens \(theActualsubsets at 3 epochs, synthetic subsets at 2\), and the full family is retained at 1 epoch during annealing \(TableLABEL:tab:phase2\-sources\)\. During annealing, we additionally include Swallow\-Code\-v2\[[20](https://arxiv.org/html/2607.09424#bib.bib20)\]\(stage 5 subset for 2 epochs\) and the Dolma 3 Dolmino code pool\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]as high\-quality curated complements\.

#### 3\.5\.3German and English Web Data

General web text is the backbone of the corpus\. The English side builds on the three Nemotron\-CC releases\[[80](https://arxiv.org/html/2607.09424#bib.bib80),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\(v1\.0, v2\.0, v2\.1\), which classify Common Crawl documents into quality tiers using ensembles of model\-based quality classifiers, and augment the high\-value tiers with synthetic transformations: rephrased variants of high\-quality pages \(\-Synthetic\), diverse question–answer reformulations \(Diverse QA\), and English translations of high\-quality non\-English crawl data \(Translated\-To\-English\)\. As described in Section[3\.1](https://arxiv.org/html/2607.09424#S3.SS1), we up\-sample the High\-Quality and synthetic tiers and exclude the Medium tier entirely; the three releases together contribute11,601\.411\{,\}601\.4B effective tokens to Phase 1\. Long\-form English document data comes from PDF\-derived corpora \(FinePDFs\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]and the Dolma 3 PDF pool\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]\), which supply the book\-, report\-, and paper\-style text that is underrepresented in HTML crawls\.

The German web component is assembled from sources of complementary character\. Naturally occurring German is provided by quality\-filtered crawl data—the HPLT corpora\[[11](https://arxiv.org/html/2607.09424#bib.bib11),[63](https://arxiv.org/html/2607.09424#bib.bib63)\]restricted to the top decile of their educational\-quality score HPLT\-3\-Top10% in Phase 1 \(edu\-scores based on JQL\[[2](https://arxiv.org/html/2607.09424#bib.bib2)\], HPLT\-4\-Top10% in Phase 2 \(edu\-scores based on Propella\[[36](https://arxiv.org/html/2607.09424#bib.bib36)\]\)—together with German\-Commons\[[24](https://arxiv.org/html/2607.09424#bib.bib24)\]\(openly licensed German text\), Genios\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\], a commercially licensed corpus of 916 German newspaper and trade\-press archives comprising 193M articles \(57\.6B words, 2010–2025; see Appendix[D](https://arxiv.org/html/2607.09424#A4)\), the German FinePDFs/FinePDFs\-Edu subsets, German FineWiki, and the curated mixture used during annealing\.

Since the supply of high\-quality native German text is far smaller than for English, we extend it with translated and synthetic German: KletterMix\[[42](https://arxiv.org/html/2607.09424#bib.bib42)\]\(machine translations of ClimbMix\[[18](https://arxiv.org/html/2607.09424#bib.bib18)\]\), MultiSynt/MT\[[37](https://arxiv.org/html/2607.09424#bib.bib37)\]\(machine translations of high\-quality Nemotron\-CC\[[80](https://arxiv.org/html/2607.09424#bib.bib80)\]English documents into German\), MT\-Reasoning\[[58](https://arxiv.org/html/2607.09424#bib.bib58)\], and Nemotron\-Multilingual\-Reasoning\[[29](https://arxiv.org/html/2607.09424#bib.bib29)\]\(translated reasoning traces\), and the German subset of PleIAs/Synth\[[69](https://arxiv.org/html/2607.09424#bib.bib69)\]\.

Phase 1 German up\-sampling leans on the scarce high\-quality crawl \(the HPLT v3 top\-decile pool runs8\.48\.4epochs\), while annealing switches to the fresher HPLT v4 pool and a German translation of ClimbMix produced with the KletterMix pipeline\[[42](https://arxiv.org/html/2607.09424#bib.bib42)\]\. In total, German receives7\.2%7\.2\\%of Phase 1 and15\.3%15\.3\\%of Phase 2 effective tokens, well above the∼5%\{\\sim\}5\\%*total*multilingual share of the reference recipe, concentrated in a single language\.

#### 3\.5\.4Specialized Synthetic Data

Beyond web and code, we train on three families of specialized, largely synthetic data\. First, nvidia/Nemotron\-CC\-Math\[[51](https://arxiv.org/html/2607.09424#bib.bib51)\]\(v1 and v2\) recovers mathematical content from Common Crawl with a rendering\-based pipeline that preserves equations and converts them to a uniform LaTeX representation, sorted into quality bands; we train on all bands for 4 epochs in Phase 1 and keep the top v1 bands during annealing, complemented by UltraData\-Math\[[91](https://arxiv.org/html/2607.09424#bib.bib91)\]\. Second, Nemotron\-Pretraining\-Specialized\-v1\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]provides targeted synthetic corpora for STEM and reasoning, including synthetic math, Wikipedia\-style rewrites of encyclopedic content, and competition\-style problems, which we up\-sample aggressively \(5 epochs,1,353\.51\{,\}353\.5B effective tokens,5\.9%5\.9\\%of Phase 1\) and retain at 1 epoch during annealing alongside the Dolma 3Thinkingpool\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]\. Third, we include SFT\-formatted data already at the pretraining stage: Nemotron\-Pretraining\-SFT\-v1 \(Math, Code, and General splits\) in Phase 1, broadened during annealing by the Nemotron post\-training collections101010[https://huggingface\.co/collections/nvidia/nemotron\-post\-training\-v3](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)\(agentic, competitive programming, instruction following, math proofs, science, finance, software engineering, safety, and multilingual subsets;LABEL:tab:phase2\-sources\)\. Exposing the model to instruction\- and reasoning\-formatted text before post\-training shortens the distribution shift at the SFT stage and measurably strengthens chain\-of\-thought behaviour of the base model\.

### 3\.6Data Mixture and Ordering

We organise all sources into seven unified categories, English Web, Academic & Wiki, Code, Mathematics, SFT, Reasoning, and German, and steer the mixture at this category level \(the per\-source realisation is given in Appendix[B](https://arxiv.org/html/2607.09424#A2)\)\. The appendix tables retain the finer Nemotron\-style accounting categories; Figure[3](https://arxiv.org/html/2607.09424#S3.F3)folds them into the seven unified categories used in the main text\. In Phase 2, the main\-text Reasoning category corresponds to the appendix row “Reasoning / STEM\-SFT”, while the narrower longtable “Reasoning subtotal” reports only the non\-SFT reasoning sources\. The guiding principle of the ordering is a quality\- and skill\-based curriculum aligned with the WSD learning\-rate schedule: breadth while the learning rate is high, concentration while it decays\.

During the stable phase \(Phase 1\), the mixture is dominated by diverse web text \(50\.3%50\.3\\%English Web plus8\.0%8\.0\\%academic and wiki text\), with code at14\.6%14\.6\\%, reasoning at10\.4%10\.4\\%, German at7\.2%7\.2\\%, mathematics at6\.0%6\.0\\%, and SFT at3\.5%3\.5\\%\(Figure[3](https://arxiv.org/html/2607.09424#S3.F3)\)\. Up\-sampling via epoch multipliers, rather than the inclusion of lower\-quality pools, is the primary lever for hitting these targets: high\-quality and synthetic tiers run22–66epochs, the medium tiers run zero\. At the onset of learning\-rate decay \(Phase 2\), the mixture shifts decisively toward skill density and German depth: English Web drops to36\.4%36\.4\\%, Academic & Wiki accounts for6\.4%6\.4\\%, SFT contributes8\.5%8\.5\\%, Reasoning accounts for11\.8%11\.8\\%, Code accounts for16\.4%16\.4\\%, Mathematics accounts for5\.2%5\.2\\%, and German rises to15\.3%15\.3\\%\(Figure[3](https://arxiv.org/html/2607.09424#S3.F3)\)\.

This places the highest\-value tokens in the regime where the decaying learning rate consolidates them most effectively\. Sources seen for multiple epochs in Phase 1 are re\-weighted downward in Phase 2 \(typically to a single epoch, or fresh replacements are substituted, e\.g\. HPLT v3→\\tov4\) so that the annealing phase adds new signal rather than repeating saturated data\.

Phase 3 builds on top of the highest\-quality documents from Phase 2 and spans seven document\-level domains \(web, code, math, German, Math SFT, General\-SFT, and Code\-SFT\)\. It is organized by*length*, using the length\-bucketed scheme of Section[3\.4](https://arxiv.org/html/2607.09424#S3.SS4)\.

Relative to the reference Nemotron 3 Nano mixture, our category targets differ in one deliberate respect: German replaces the broad multilingual bucket and is raised from5%5\\%\(Phase 1 & Phase 2\) to7\.2%7\.2\\%\(our Phase 1\) and15\.32%15\.32\\%\(our Phase 2\), funded by reductions in synthetic web and code\-SFT share \(Tables[7](https://arxiv.org/html/2607.09424#A2.T7)and[9](https://arxiv.org/html/2607.09424#A2.T9)\)\. Within each phase, category proportions are held stationary: data of all categories is shuffled and interleaved uniformly at the batch level, so the curriculum acts between phases, not within them\.

## 4Evaluations

We evaluate Soofi S 30B\-A3B Base against 16 open base models with the same harness, prompts, and few\-shot configuration; the complete per\-task results for all models are released alongside this report\.[Figure˜1](https://arxiv.org/html/2607.09424#S1.F1)includes the full set of models in the throughput–capability comparison, while Tables[4](https://arxiv.org/html/2607.09424#S4.T4)and[5](https://arxiv.org/html/2607.09424#S4.T5)report detailed results for the main comparison subsets\.111111Baselines: Nemotron 3 Nano 30B\-A3B; Qwen3\.5 35B\-A3B/9B; Gemma 3 12B/27B; Ministral 3 3B/8B/14B; Olmo 3 32B; Alia 40B; Apertus 8B/70B; EuroLLM 9B/22B; Teuken 7B; Salamandra 7B\.In the main text, we report detailed results for the strongest baselines in two comparison settings\. The first compares against large open\-source base models \(Alia 40B\[[26](https://arxiv.org/html/2607.09424#bib.bib26)\], EuroLLM 22B\[[53](https://arxiv.org/html/2607.09424#bib.bib53),[52](https://arxiv.org/html/2607.09424#bib.bib52)\], Apertus 70B\[[5](https://arxiv.org/html/2607.09424#bib.bib5)\], and Olmo 3 32B\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]\); the second compares against its architectural reference, Nemotron 3 Nano 30B\-A3B\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\], and large open\-weight models of comparable or larger active size \(Qwen3\.5 35B\-A3B\[[88](https://arxiv.org/html/2607.09424#bib.bib88),[71](https://arxiv.org/html/2607.09424#bib.bib71)\], Ministral 3 14B\[[47](https://arxiv.org/html/2607.09424#bib.bib47)\], and Gemma 3 27B\[[83](https://arxiv.org/html/2607.09424#bib.bib83)\]\)\. All models are evaluated with the samelm\-evaluation\-harness\[[22](https://arxiv.org/html/2607.09424#bib.bib22)\]pipeline, using the same task configurations \(prompts, number of few\-shot samples, etc\), on English and German benchmark suites closely following the evaluation setup of Olmo 3\[[64](https://arxiv.org/html/2607.09424#bib.bib64)\]covering code, mathematics, knowledge, reasoning, science, reading comprehension, and German language proficiency\. This includes a held\-out set of benchmarks that were not used during any ablation experiments\. Soofi S is evaluated at the selected base checkpointiter\_1056000, the final checkpoint of the constant\-annealing stage \(Section[2\.1](https://arxiv.org/html/2607.09424#S2.SS1), Appendix[F](https://arxiv.org/html/2607.09424#A6)\); for other checkpointed models we report the highest available training step\. The complete per\-task results files are released alongside the model; headline numbers are collected in Tables[4](https://arxiv.org/html/2607.09424#S4.T4)and[5](https://arxiv.org/html/2607.09424#S4.T5)and Figures[4](https://arxiv.org/html/2607.09424#S4.F4)and[9](https://arxiv.org/html/2607.09424#S4.F9)\.

### 4\.1Soofi S vs Open\-Source Models

The open\-source comparison reports results against Alia 40B, EuroLLM 22B, Apertus 70B, and Olmo 3 32B\. We separate these models from the larger open\-weight baselines in Section[4\.2](https://arxiv.org/html/2607.09424#S4.SS2)so that each table compares models with similar release categories\. Within this set, Soofi S is the strongest model overall: it obtains the highest English aggregate \(\+2\.8 over Olmo 3 32B\), German aggregate \(\+6\.3 over Apertus 70B\), and the highest held\-out English \(\+8\.3\) and German \(\+5\.6\) scores \(Figure[4](https://arxiv.org/html/2607.09424#S4.F4)\)\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x5.png)Figure 4:Evaluation overview for the open\-source comparison\.Soofi Sis compared against largeopen\-source models\(Alia, EuroLLM, Apertus, and Olmo 3\)\. Aggregates are the harness\-level English and German suite means\. Code EN averages HumanEval and MBPP, Code DE averages HumanEval\-DE and MBPP\-DE, and LBPP is reported separately\.Table[4](https://arxiv.org/html/2607.09424#S4.T4)gives the per\-task results\. The largest margins over the next\-best open\-source baseline occur on German and technical benchmarks: German aggregate \(\+6\.3\+6\.3\), held\-out English \(\+8\.3\+8\.3\), held\-out German \(\+5\.6\+5\.6\), HumanEval \(\+10\.8\+10\.8\), MBPP\-DE \(\+13\.4\+13\.4\), GSM8K\-Platinum\-DE \(\+9\.7\+9\.7\), INCLUDE\-DE \(\+10\.1\+10\.1\), and GLP\-DE \(\+7\.6\+7\.6\)\. Olmo 3 32B is the strongest non\-Soofi model overall and leads this subset on LBPP, SocialIQA, SQuAD, DROP, and NaturalQuestions\. These results identify the main residual gaps for Soofi S in this comparison as contamination\-aware code evaluation, open\-domain factual recall, and extractive reading comprehension\.

Table 4:Base model evaluation results \(%\) against large*open\-source*models\. Best result per row inbold, second bestunderlined\. All models evaluated with identical harness, prompts, and few\-shot settings; “\-DE” denotes the German variant of a benchmark\. Aggregates are harness\-level suite means; Column shading:Soofi S\(ours\),open\-source\(Alia, EuroLLM, Apertus, Olmo 3\)\.##### Code performance\.

Soofi S ranks first on four of the five code benchmarks in the open\-source comparison \(Figure[5](https://arxiv.org/html/2607.09424#S4.F5)\)\. It exceeds the next\-best open\-source baseline by10\.810\.8points on HumanEval,7\.47\.4points on MBPP,3\.03\.0points on HumanEval\-DE, and13\.413\.4points on MBPP\-DE\. LBPP is the only code benchmark in this subset where Soofi S is not first; Olmo 3 32B scores32\.132\.1compared with31\.031\.0for Soofi S\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x6.png)Figure 5:Code generation results \(pass@1\) against large open\-source models on English and German benchmarks\. Soofi S leads this comparison on HumanEval, MBPP, HumanEval\-DE, and MBPP\-DE; Olmo 3 32B is strongest on LBPP\.
##### Mathematics, knowledge, and reasoning\.

Soofi S also ranks first within the open\-source subset on all mathematics benchmarks shown in Table[4](https://arxiv.org/html/2607.09424#S4.T4)and Figure[6](https://arxiv.org/html/2607.09424#S4.F6)\. The largest mathematics margins are on GSM8K\-Platinum\-DE \(\+9\.7\+9\.7over Olmo 3 32B\), Minerva\-500 \(\+24\.2\+24\.2over Olmo 3 32B\), Minerva Math\-EN \(\+27\.0\+27\.0over Olmo 3 32B\), and Minerva MATH\-DE \(\+7\.5\+7\.5over Olmo 3 32B\)\. In knowledge, reasoning, and science \(Figure[7](https://arxiv.org/html/2607.09424#S4.F7)\), Soofi S ranks first on MMLU\-STEM, MMLU\-Pro, MMLU\-Pro\-DE, INCLUDE\-DE, BBH, AGIEval, GPQA\-Diamond, GPQA\-Diamond\-DE, and ARC\-Challenge\. The largest margins in this group include GPQA\-Diamond \(\+10\.1\+10\.1over Olmo 3 32B\), INCLUDE\-DE \(\+10\.1\+10\.1over EuroLLM 22B\), and GPQA\-Diamond\-DE \(\+7\.8\+7\.8over Olmo 3 32B\)\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x7.png)Figure 6:Mathematics results against large open\-source models on English and German benchmarks\. Soofi S leads the open\-source comparison on GSM8K, GSM8K\-Platinum\-DE, Math\-EN, and Math\-DE\.![Refer to caption](https://arxiv.org/html/2607.09424v1/x8.png)Figure 7:Knowledge \(left\) and reasoning/science \(right\) benchmarks against large open\-source models\. Soofi S leads the open\-source comparison on MMLU\-STEM, MMLU\-Pro, INCLUDE\-DE, BBH, AGIEval, GPQA\-Diamond, GPQA\-Diamond\-DE, and ARC\-Challenge\.
##### German capabilities\.

Figure[8](https://arxiv.org/html/2607.09424#S4.F8)isolates the German benchmarks for the open\-source comparison\. Soofi S ranks first on every German row in this subset, with margins of\+6\.3\+6\.3on the German aggregate,\+7\.6\+7\.6on GLP\-DE,\+7\.0\+7\.0on ARC\-Challenge\-DE,\+10\.1\+10\.1on INCLUDE\-DE,\+9\.7\+9\.7on GSM8K\-Platinum\-DE,\+13\.4\+13\.4on MBPP\-DE, and\+5\.6\+5\.6on the held\-out German suite relative to the strongest open\-source baseline for each metric\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x9.png)Figure 8:German benchmark results against large open\-source models\. Soofi S ranks first on the German aggregate, GLP\-DE, ARC\-Challenge\-DE, INCLUDE\-DE, GSM8K\-Platinum\-DE, MBPP\-DE, and the held\-out German suite in this comparison\.

### 4\.2Soofi S vs Open\-Weight Models

The open\-weight comparison contains two types of baselines: the architecture\-identical Nemotron 3 Nano reference, and larger public\-weight models from Qwen, Ministral, and Gemma\. This split separates the effect of the German–English data recipe from comparisons against larger models with different architectures and training corpora\. In this subset, Soofi S is not the top model on aggregate scores, but it is competitive with larger dense baselines and improves over Nemotron 3 Nano on all aggregate and held\-out metrics\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x10.png)Figure 9:Base model evaluation overview forSoofi SagainstNemotron\(same architecture\) and largeopen\-weight models\(Qwen, Ministral, and Gemma\)\. Aggregates are the harness\-level English and German suite means\. Code EN averages HumanEval and MBPP, Code DE averages HumanEval\-DE and MBPP\-DE, and LBPP is reported separately\.At the aggregate level, Qwen3\.5 35B\-A3B has the highest English, German, and held\-out means in this subset \(Table[5](https://arxiv.org/html/2607.09424#S4.T5)\)\. Soofi S scores70\.170\.1on the English aggregate, compared with70\.370\.3for Gemma 3 27B and Ministral 3 14B\. On the German aggregate, Soofi S scores79\.179\.1, compared with78\.378\.3for Ministral 3 14B and78\.478\.4for Gemma 3 27B\. Relative to Nemotron 3 Nano, Soofi S improves the English aggregate by\+1\.8\+1\.8, the German aggregate by\+4\.2\+4\.2, held\-out English by\+6\.7\+6\.7, and held\-out German by\+2\.0\+2\.0\.

Table 5:Base model evaluation results \(%\) against Nemotron 3 Nano and large*open\-weight*models\. Best result per row inbold, second bestunderlined\. All models evaluated with identical harness, prompts, and few\-shot settings; “\-DE” denotes the German variant of a benchmark\. Aggregates are harness\-level suite means; Nemotron 3 Nano shares the same 30B\-A3B Mixture\-of\-Experts architecture as Soofi S\. Column shading:Soofi S\(ours\),Nemotron\(same architecture\),open\-weight\(Qwen, Ministral, Gemma\)\.##### Code performance\.

Figure[10](https://arxiv.org/html/2607.09424#S4.F10)reports the code\-generation results for the open\-weight comparison\. Soofi S sets the best score in this subset on HumanEval\[[12](https://arxiv.org/html/2607.09424#bib.bib12)\]\(73\.873\.8\), MBPP\[[6](https://arxiv.org/html/2607.09424#bib.bib6)\]\(70\.270\.2\), and MBPP\-DE \(84\.284\.2\), and is second only to its architectural reference on HumanEval\-DE \(65\.565\.5vs\.68\.868\.8\)\. The German code benchmarks \(problem statements and docstrings in German\) show the largest positive margins for Soofi S in this subset: on MBPP\-DE, Soofi S leads the best dense open\-weight model by8\.68\.6points\. The contamination\-aware LBPP benchmark\[[54](https://arxiv.org/html/2607.09424#bib.bib54)\]\(31\.031\.0\) is the one code benchmark where Soofi S trails, behind Nemotron \(38\.138\.1\) and Qwen3\.5 35B\-A3B \(32\.432\.4\), while staying ahead of Ministral and Gemma\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x11.png)Figure 10:Code generation results \(pass@1\) against Nemotron 3 Nano and large open\-weight models on English and German benchmarks\. Soofi S achieves the best HumanEval, MBPP, and MBPP\-DE scores in this comparison\. The code aggregates used in Figure[9](https://arxiv.org/html/2607.09424#S4.F9)average HumanEval with MBPP and HumanEval\-DE with MBPP\-DE; LBPP is reported separately\.
##### Mathematics\.

On grade\-school reasoning, Soofi S scores86\.186\.1on GSM8K\[[14](https://arxiv.org/html/2607.09424#bib.bib14)\]and87\.187\.1on GSM8K\-Platinum\-DE\[[86](https://arxiv.org/html/2607.09424#bib.bib86)\]—within0\.40\.4and4\.14\.1points of the best open\-weight result on the respective tasks \(Figure[11](https://arxiv.org/html/2607.09424#S4.F11)\)\. On competition\-style mathematics\[[32](https://arxiv.org/html/2607.09424#bib.bib32),[45](https://arxiv.org/html/2607.09424#bib.bib45)\], Soofi S reaches79\.479\.4on Minerva\-500 and81\.081\.0on Minerva Math\-EN—the second\-best results in this comparison, behind only Qwen3\.5 35B\-A3B—lifting its Math\-EN score to82\.882\.8, essentially tied with Qwen3\.5 35B\-A3B under this suite definition\. German competition mathematics is weaker: on Minerva MATH\-DE \(56\.056\.0\) Soofi S trails Qwen3\.5 35B\-A3B \(76\.576\.5\), Gemma 3 27B \(65\.665\.6\), and its architectural reference \(58\.158\.1\), leaving Math\-DE at71\.571\.5\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x12.png)Figure 11:Mathematics results against Nemotron 3 Nano and large open\-weight models on English and German benchmarks\. Soofi S is second\-best on GSM8K\[[14](https://arxiv.org/html/2607.09424#bib.bib14)\]; on the German side it is essentially level with its architectural reference on GSM8K\-Platinum\-DE, while competition\-style Math\-DE trails Qwen3\.5 35B\-A3B\.
##### Knowledge\.

On broad academic knowledge, Soofi S is competitive with dense models several times its active size: MMLU\-STEM\[[31](https://arxiv.org/html/2607.09424#bib.bib31)\]75\.975\.9and MMLU\-Pro\[[87](https://arxiv.org/html/2607.09424#bib.bib87)\]51\.451\.4sit between Gemma 3 27B and Ministral 3 14B, and the German MMLU\-Pro\-DE \(49\.449\.4\) shows the same pattern \(Figure[12](https://arxiv.org/html/2607.09424#S4.F12), left\)\. On INCLUDE\-DE\[[76](https://arxiv.org/html/2607.09424#bib.bib76)\]—regional, Germany\-specific knowledge spanning driving\-licence, social\-science, and STEM exams—Soofi S ties Qwen3\.5 35B\-A3B for the best score in the table \(61\.261\.2\), consistent with the up\-weighted native German data described in Sections[3\.2](https://arxiv.org/html/2607.09424#S3.SS2)and[3\.3](https://arxiv.org/html/2607.09424#S3.SS3)\. Open\-domain factual recall as measured by NaturalQuestions\[[43](https://arxiv.org/html/2607.09424#bib.bib43)\]\(79\.079\.0\) narrowly trails the largest dense models, consistent with storing world knowledge in33B active parameters; we return to this in the Limitations paragraph\.

##### Reasoning and science\.

Soofi S achieves strong general\-reasoning performance, with scores of78\.878\.8on BBH\[[81](https://arxiv.org/html/2607.09424#bib.bib81)\],66\.966\.9on AGIEval\[[90](https://arxiv.org/html/2607.09424#bib.bib90)\], and90\.690\.6on ARC\-Challenge\[[13](https://arxiv.org/html/2607.09424#bib.bib13)\]\. These results place it consistently within a point of Ministral 3 14B \(Figure[12](https://arxiv.org/html/2607.09424#S4.F12), right\)\. Its graduate\-level science performance is particularly strong: GPQA\-Diamond\[[75](https://arxiv.org/html/2607.09424#bib.bib75)\]increases to43\.443\.4, a\+9\.6\+9\.6\-point improvement over the reference Nemotron recipe, trailing Qwen3\.5 35B\-A3B among the open\-weight models shown here\. On the German GPQA\-Diamond\-DE benchmark, Soofi S reaches41\.941\.9, a\+4\.5\+4\.5\-point improvement over its reference and within0\.60\.6points of the best remaining dense model \(Ministral 3 14B,42\.542\.5\)\.

We attribute these gains primarily to the reasoning\-oriented annealing phase and the high density of STEM\-SFT data described in Section[3\.5\.4](https://arxiv.org/html/2607.09424#S3.SS5.SSS4)\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x13.png)Figure 12:Knowledge \(left\) and reasoning/science \(right\) benchmarks against Nemotron 3 Nano and large open\-weight models\. Soofi S ties Qwen3\.5 35B\-A3B on INCLUDE\-DE and improves GPQA\-Diamond by\+9\.6\+9\.6points over its architectural reference, while matching dense 14–27B models on MMLU\-STEM, MMLU\-Pro, BBH, and AGIEval\.
##### German capabilities\.

Figure[13](https://arxiv.org/html/2607.09424#S4.F13)reports the German benchmarks for the open\-weight comparison\. Soofi S ranks first on MBPP\-DE \(84\.284\.2\) and ties Qwen3\.5 35B\-A3B on INCLUDE\-DE \(61\.261\.2\)\. It is close to the larger dense models on the German aggregate \(79\.179\.1\), behind Qwen3\.5 35B\-A3B \(81\.681\.6\) and ahead of Gemma 3 27B \(78\.478\.4\), Ministral 3 14B \(78\.378\.3\), and Nemotron 3 Nano \(74\.974\.9\)\. Relative to Nemotron, Soofi S improves GLP\-DE by\+15\.1\+15\.1, INCLUDE\-DE by\+1\.5\+1\.5, and the held\-out German suite by\+2\.0\+2\.0, while remaining slightly below Nemotron on GSM8K\-Platinum\-DE and HumanEval\-DE\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x14.png)Figure 13:German benchmark results against Nemotron 3 Nano and large open\-weight models\. Soofi S ranks first on MBPP\-DE, ties Qwen3\.5 35B\-A3B on INCLUDE\-DE, and improves the German aggregate over the architecture\-identical Nemotron baseline by\+4\.2\+4\.2points\.
##### Effect of the German–English recipe\.

Since Soofi S shares its architecture with Nemotron 3 Nano 30B\-A3B, the pairwise comparison in Figure[14](https://arxiv.org/html/2607.09424#S4.F14)isolates the effect of our data recipe from architecture\. The German interventions deliver large, targeted gains—GLP\-DE\+15\.1\+15\.1, German aggregate\+4\.2\+4\.2, INCLUDE\-DE\+1\.5\+1\.5, German code\+0\.5\+0\.5—without the regression on English that monolingual specialisation usually incurs: the English aggregate*improves*by\+1\.8\+1\.8, code by\+2\.2\+2\.2, CommonsenseQA\[[82](https://arxiv.org/html/2607.09424#bib.bib82)\]by\+5\.1\+5\.1, and GPQA\-Diamond by\+9\.6\+9\.6\. The held\-out benchmark suite scores improve by\+6\.7\+6\.7\(EN\) and\+2\.0\+2\.0\(DE\)\. The only notable regressions are German competition mathematics \(Math\-DE aggregate−1\.4\-1\.4\), open\-domain factual recall \(NaturalQuestions−1\.3\-1\.3\), and a−0\.4\-0\.4on GSM8K, consistent with German tokens displacing a portion of English web knowledge\. Overall, the Nemotron comparison indicates that the German\-focused data recipe substantially improves German capability while preserving or improving the English aggregate, code, reasoning, and held\-out scores reported here\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x15.png)Figure 14:Per\-benchmark score difference between Soofi S and the architecture\-identical Nemotron 3 Nano 30B\-A3B, isolating the effect of the German–English data recipe\. German\-focused benchmarks improve by up to1515points while English capability is preserved or improved; the main cost is open\-domain English factual recall\.
##### Limitations\.

We report three caveats in the interest of full transparency\. First, the Minerva mathematics\[[32](https://arxiv.org/html/2607.09424#bib.bib32),[45](https://arxiv.org/html/2607.09424#bib.bib45)\]scores required a corrected evaluation protocol\. Our initial harness configuration used a\\n\\nstop sequence and a generation limit of1,0241\{,\}024tokens, which truncated long chain\-of\-thought solutions before a final answer was produced and collapsed the Minerva scores of the strongest models \(Soofi S measured9\.89\.8on Minerva\-500 under this configuration, and the architecture\-identical Nemotron 3 Nano5\.65\.6\)\. Removing the stop condition and raising the generation limit to4,0964\{,\}096tokens yields the results reported in Tables[4](https://arxiv.org/html/2607.09424#S4.T4)and[5](https://arxiv.org/html/2607.09424#S4.T5)and the figures \(79\.479\.4and81\.081\.0for Soofi S on Minerva\-500 and Minerva Math\-EN\); all models were re\-evaluated under the same corrected configuration\. Second, competition\-style mathematics in German remains the clearest capability gap to the frontier: on Minerva MATH\-DE \(56\.056\.0\) Soofi S trails Qwen3\.5 35B\-A3B \(76\.576\.5\) and Gemma 3 27B \(65\.665\.6\), even though its German grade\-school mathematics \(GSM8K\-Platinum\-DE,87\.187\.1\) is essentially level with the architectural reference\. Third, open\-domain factual recall remains capacity\-limited: NaturalQuestions \(79\.079\.0\) trails the largest dense baselines \(Gemma 3 27B83\.583\.5\), consistent with storing world knowledge in 3B active parameters; we expect retrieval\-augmented deployment to close this gap in practice\.

### 4\.3Serving efficiency\.

Active parameter count is only a proxy for inference cost\. In deployment, the relevant quantity is sustained throughput at realistic context lengths and request concurrency, where decoding is often limited by memory bandwidth: each generated token must stream the active model weights and read the attention cache for every active sequence\. This is the regime targeted by the hybrid Mamba–MoE architecture\. In Soofi S, only 6 of 52 layers maintain a KV cache, with 2 KV heads each, while the 23 Mamba\-2 layers carry a fixed\-size recurrent state\. Consequently, the incremental attention\-cache footprint is only about 6 KB per token per sequence, which is1111–53×53\{\\times\}lower than for the dense models in our comparison\. As context length grows, only this small attention component scales with sequence length; the Mamba recurrent state remains constant\-size\.

Figure[1](https://arxiv.org/html/2607.09424#S1.F1)quantifies this effect using measured aggregate decode tokens per second \(TPS\) per GPU\. We use a TP=1, single\-B200 vLLM latency\-subtraction protocol\. For each context length, we measure fixed\-batch latency as a function of output lengthOO\. Lett​\(O\)t\(O\)denote this latency\. We estimate aggregate decode TPS per GPU by subtracting theOshort=1O\_\{\\mathrm\{short\}\}=1run from theOlong=1024O\_\{\\mathrm\{long\}\}=1024run:

TPSdecodeagg/GPU=B​\(Olong−Oshort\)NGPU​\[t​\(Olong\)−t​\(Oshort\)\]\\mathrm\{TPS\}\_\{\\mathrm\{decode\}\}^\{\\mathrm\{agg/GPU\}\}=\\frac\{B\\left\(O\_\{\\mathrm\{long\}\}\-O\_\{\\mathrm\{short\}\}\\right\)\}\{N\_\{\\mathrm\{GPU\}\}\\left\[t\\left\(O\_\{\\mathrm\{long\}\}\\right\)\-t\\left\(O\_\{\\mathrm\{short\}\}\\right\)\\right\]\}\(1\)HereOshort=1O\_\{\\mathrm\{short\}\}=1,Olong=1024O\_\{\\mathrm\{long\}\}=1024,NGPU=1N\_\{\\mathrm\{GPU\}\}=1, andB=32B=32for all plotted measurements\. This subtraction removes most prompt\-prefill cost and isolates the per\-GPU cache\-bandwidth pressure that tensor parallelism can otherwise mask\. At a 40K context and batch size 32, Soofi S sustains a measured aggregate decode rate of 4\.82k TPS/GPU, which is9\.2×9\.2\{\\times\}higher than Ministral 3 14B, while fitting the weights and all 32 sequence states on a single GPU\. In contrast, Apertus 70B does not fit on a single GPU under the current fixed\-256K, TP=1 setup\.

The same measurements also expose the prefill/TTFT side of the hybrid architecture:t​\(Oshort\)=t​\(1\)t\(O\_\{\\mathrm\{short\}\}\)=t\(1\)is the measured batch\-32 latency to process the prompt and produce the first output token, i\.e\., the TTFT\-like measurement in this protocol\. At 40K context, Soofi S reachest​\(1\)=22\.7t\(1\)=22\.7s, compared with71\.971\.9s for Ministral 3 14B,92\.992\.9s for Gemma 3 27B,101\.7101\.7s for OLMo 3\.1 32B, and213\.4213\.4s for the Qwen3 32B dense control\. At 256K context, Soofi S remains the fastest complete sweep in our batch\-32 protocol \(372\.7372\.7s\), ahead of Qwen3\.5 35B\-A3B \(579\.2579\.2s\), Gemma 3 27B \(999\.4999\.4s\), OLMo 3\.1 32B \(1,487\.31\{,\}487\.3s\), Ministral 3 14B \(2,058\.92\{,\}058\.9s\), and Qwen3 32B dense \(6,428\.66\{,\}428\.6s\)\. This is the prefill analogue of the decode\-cache advantage: most sequence mixing in Soofi S is carried by Mamba layers rather than full attention, so long\-context TTFT grows much more favorably than in full\-attention dense baselines\.

Panel \([1\(b\)](https://arxiv.org/html/2607.09424#S1.F1.sf2)\) shows that the decode gap widens with context length: dense\-model throughput decreases as KV\-cache reads dominate decoding, whereas Soofi S remains nearly flat from 4K to 256K in the current measurement snapshot, with endpoint aggregate decode rate changing from 4\.29k to 4\.30k TPS/GPU and no point more than∼34%\{\\sim\}34\\%below the 4K value\. Among the comparison models, only Qwen3\.5, a Gated\-DeltaNet hybrid\[[71](https://arxiv.org/html/2607.09424#bib.bib71)\], exhibits similar scaling behavior\. In its published 9B configuration, 8 of 32 layers remain full\-attention layers, with 4 KV heads of dimension 256, corresponding to a per\-sequence cache of 32 KB per token, or5\.3×5\.3\{\\times\}that of Soofi S\. For the 35B\-A3B variant, we measure 2\.60k aggregate decode TPS/GPU at 40K under the same protocol, which is1\.9×1\.9\{\\times\}lower than Soofi S\. This qualitative separation is consistent with NVIDIA’s Nemotron\-H measurements, where a related hybrid engine achieved3\.3×3\.3\{\\times\}higher throughput than Qwen3\-30B\-A3B on production inference engines\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\.

## 5Related Work

We situate Soofi S with respect to five strands of prior work: open language\-model pretraining, data acquisition and filtering, training curricula, efficient sparse and hybrid architectures, and multilingual European foundation models\. We focus on base\-model pretraining; instruction tuning, preference optimization, and reasoning\-specific reinforcement learning are complementary and outside the scope of this report\.

##### Open language\-model pretraining\.

The modern causal language\-modeling paradigm was established by GPT\-style pretraining: GPT\-2 showed that a decoder\-only Transformer trained with next\-token prediction on large web text could perform a wide range of tasks in a zero\-shot setting, and GPT\-3 demonstrated the emergence of strong in\-context few\-shot behavior at a substantially larger scale\[[72](https://arxiv.org/html/2607.09424#bib.bib72),[10](https://arxiv.org/html/2607.09424#bib.bib10)\]\. Subsequent scaling\-law work made model size, data size, and compute budget explicit design variables, while the Chinchilla analysis shifted attention from parameter count alone toward data–compute balance and the importance of training sufficiently long on enough tokens\[[40](https://arxiv.org/html/2607.09424#bib.bib40),[33](https://arxiv.org/html/2607.09424#bib.bib33)\]\. These results motivate the central design principle of recent pretraining recipes: capability is a joint function of architecture, token budget, data quality, curriculum, and serving cost, not merely of total parameters\.

Open releases have followed a separate but related trajectory\. OPT released a family of decoder\-only models together with code and a detailed account of training infrastructure, GPT\-NeoX\-20B provided a permissively released large autoregressive model trained on the Pile, and Pythia made training dynamics inspectable by releasing many intermediate checkpoints and a fixed data order\[[89](https://arxiv.org/html/2607.09424#bib.bib89),[9](https://arxiv.org/html/2607.09424#bib.bib9),[7](https://arxiv.org/html/2607.09424#bib.bib7),[21](https://arxiv.org/html/2607.09424#bib.bib21)\]\. BLOOM extended this community\-scale approach to multilingual pretraining, combining a large open\-access model with the ROOTS multilingual corpus and a collaborative development process\[[8](https://arxiv.org/html/2607.09424#bib.bib8)\]\. These efforts established the scientific value of releasing more than a final checkpoint: intermediate states, training code, data documentation, and evaluation recipes make it possible to study memorization, bias, scaling behavior, and data effects\.

The broader open\-weight ecosystem has since produced increasingly strong general\-purpose baselines, including LLaMA and Llama 2, Falcon, Mistral 7B, Mixtral, Qwen, Gemma, and recent Mistral/Ministral releases\[[84](https://arxiv.org/html/2607.09424#bib.bib84),[85](https://arxiv.org/html/2607.09424#bib.bib85),[4](https://arxiv.org/html/2607.09424#bib.bib4),[38](https://arxiv.org/html/2607.09424#bib.bib38),[39](https://arxiv.org/html/2607.09424#bib.bib39),[88](https://arxiv.org/html/2607.09424#bib.bib88),[71](https://arxiv.org/html/2607.09424#bib.bib71),[83](https://arxiv.org/html/2607.09424#bib.bib83),[47](https://arxiv.org/html/2607.09424#bib.bib47),[55](https://arxiv.org/html/2607.09424#bib.bib55)\]\. However, many such models are open\-weight rather than fully reproducible: the weights and inference code are available, but the exact data mixture, filtering pipeline, training order, logs, intermediate checkpoints, and rejected data sources are often missing\. OLMo, OLMoE, OLMo 3, and Apertus mark a stronger notion of openness by releasing training data or data recipes, code, model artifacts, and broader documentation for scientific audit and reuse\[[27](https://arxiv.org/html/2607.09424#bib.bib27),[57](https://arxiv.org/html/2607.09424#bib.bib57),[64](https://arxiv.org/html/2607.09424#bib.bib64),[5](https://arxiv.org/html/2607.09424#bib.bib5)\]\. Soofi S follows this fully open line, but differs in scope and architecture: it is a German–English base model trained with exact per\-source token accounting on a sparse hybrid Mamba–Transformer MoE backbone, rather than a dense general\-purpose or broadly multilingual Transformer\.

##### Pretraining data acquisition and filtering\.

Pretraining data pipelines have evolved from relatively simple web\-crawl selection toward multi\-stage corpus construction\. Early influential corpora such as WebText, C4, and the Pile used web\-scale collection, heuristic filtering, deduplication, and source balancing to turn noisy internet data into training data suitable for language modeling\[[72](https://arxiv.org/html/2607.09424#bib.bib72),[73](https://arxiv.org/html/2607.09424#bib.bib73),[21](https://arxiv.org/html/2607.09424#bib.bib21)\]\. BLOOM’s ROOTS corpus added a multilingual community\-curated data effort, combining web data with manually selected sources across many languages\[[8](https://arxiv.org/html/2607.09424#bib.bib8)\]\. More recent corpora make the filtering pipeline itself a central contribution: FineWeb and Dolma document large\-scale web cleaning, deduplication, and mixture construction; Nemotron\-CC provides quality\-tiered Common Crawl data with synthetic and translated variants; and JQL shows that language\-model\-based quality judgments can be distilled into multilingual data filters that transfer across languages more robustly than English\-centric heuristics\[[68](https://arxiv.org/html/2607.09424#bib.bib68),[79](https://arxiv.org/html/2607.09424#bib.bib79),[80](https://arxiv.org/html/2607.09424#bib.bib80),[2](https://arxiv.org/html/2607.09424#bib.bib2)\]\.

Domain\-specific acquisition has become equally important\. PDF\- and OCR\-derived corpora such as FinePDFs and OlmoOCR add long\-form reports, books, papers, and educational material that are underrepresented in HTML\-only crawls\[[44](https://arxiv.org/html/2607.09424#bib.bib44),[70](https://arxiv.org/html/2607.09424#bib.bib70)\]\. Code corpora such as Nemotron\-Pretraining\-Code\-v1/v2 and SwallowCode complement web\-extracted code with repository\-derived or rewritten programming data\[[61](https://arxiv.org/html/2607.09424#bib.bib61),[59](https://arxiv.org/html/2607.09424#bib.bib59),[20](https://arxiv.org/html/2607.09424#bib.bib20)\]\. Mathematical data pipelines such as Nemotron\-CC\-Math and UltraData\-Math target equation\-rich documents and problem\-solving text, where generic web extraction often destroys structure\[[51](https://arxiv.org/html/2607.09424#bib.bib51),[91](https://arxiv.org/html/2607.09424#bib.bib91)\]\. For German, large multilingual crawls and national resources such as HPLT, German Commons, and KletterMix are especially relevant because high\-quality native German tokens are scarcer than English web tokens\[[11](https://arxiv.org/html/2607.09424#bib.bib11),[63](https://arxiv.org/html/2607.09424#bib.bib63),[24](https://arxiv.org/html/2607.09424#bib.bib24),[42](https://arxiv.org/html/2607.09424#bib.bib42)\]\. Soofi S builds on these trends but reports the corpus at finer granularity: for every phase, we disclose source identifiers, raw tokens, epoch multipliers, effective token counts, and sources considered but excluded\. This makes the data mixture auditable in a way that aggregate token counts cannot provide\.

##### Training curricula and optimization recipes\.

Large\-scale pretraining recipes increasingly separate the problem of collecting tokens from the problem of ordering and weighting them\. Compute\-optimal scaling results motivate training smaller or sparse models for more tokens when inference efficiency matters, while modern open reports show that curriculum phase boundaries, learning\-rate schedules, and annealing mixtures can have large effects on downstream behavior\[[33](https://arxiv.org/html/2607.09424#bib.bib33),[35](https://arxiv.org/html/2607.09424#bib.bib35),[30](https://arxiv.org/html/2607.09424#bib.bib30),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. A common pattern is to allocate early training to broad coverage and later training to higher\-quality or skill\-focused data, often under a Warmup–Stable–Decay schedule\. OLMo\-style releases emphasize logging, intermediate checkpoints, and reproducible recipes; Nemotron\-style recipes emphasize quality\-tiered web data, synthetic skill data, and high\-quality annealing\[[27](https://arxiv.org/html/2607.09424#bib.bib27),[64](https://arxiv.org/html/2607.09424#bib.bib64),[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. Soofi S follows the same high\-level philosophy but adapts it to a German–English target: Phase 1 maximizes diversity, Phase 2 concentrates high\-quality web, code, mathematics, reasoning, SFT\-formatted, and German data, and Phase 3 extends context length with document\-length buckets\. The difference is the explicit bilingual reallocation of token budget and the disclosure of the exact realized mixtures\.

##### Efficient sparse and hybrid architectures\.

Most open LLMs remain dense Transformers, but dense attention becomes costly at long context and high concurrency because decoding must repeatedly read both model weights and the per\-sequence KV cache\. Sparse Mixture\-of\-Experts models address the weight side of this problem by increasing total capacity while activating only a small expert subset per token\. Foundational MoE work, Switch Transformers, DeepSeekMoE, and fine\-grained MoE studies show that sparse routing can improve the capability–compute trade\-off when routing and load balancing are stable\[[77](https://arxiv.org/html/2607.09424#bib.bib77),[19](https://arxiv.org/html/2607.09424#bib.bib19),[15](https://arxiv.org/html/2607.09424#bib.bib15),[41](https://arxiv.org/html/2607.09424#bib.bib41)\]\. OLMoE demonstrates that this sparse route can also be made fully open at smaller active\-parameter scales\[[57](https://arxiv.org/html/2607.09424#bib.bib57)\]\.

A complementary line of work reduces the sequence\-state cost of attention\. Mamba\-2 and related state\-space models replace much of the quadratic attention machinery with recurrent state updates, yielding linear\-time sequence mixing and a near\-constant state during decoding\[[16](https://arxiv.org/html/2607.09424#bib.bib16)\]\. Nemotron\-H and Nemotron 3 combine Mamba\-style sequence mixing, sparse attention, and MoE layers to obtain strong long\-context serving efficiency, while Qwen3\.5 explores a different hybrid sequence\-modeling path with Gated DeltaNet layers\[[60](https://arxiv.org/html/2607.09424#bib.bib60),[61](https://arxiv.org/html/2607.09424#bib.bib61),[71](https://arxiv.org/html/2607.09424#bib.bib71)\]\. Soofi S adopts the Nemotron\-style 30B\-A3B hybrid Mamba–Transformer MoE design, but evaluates it under a distinct sovereign bilingual pretraining recipe\. This makes the architectural comparison unusually clean: relative to Nemotron 3 Nano, gains and trade\-offs can largely be attributed to data mixture, German up\-weighting, annealing, and long\-context continuation rather than to a different backbone\.

##### Multilingual and European language models\.

Multilingual pretraining aims to reduce the English bias of large language models, but it requires careful allocation of model capacity and high\-quality tokens across languages\. BLOOM was an early large\-scale open\-access demonstration of multilingual decoder\-only pretraining\[[8](https://arxiv.org/html/2607.09424#bib.bib8)\]\. More recent European efforts add requirements around sovereignty, transparency, language coverage, and regulatory compatibility\. Teuken\-7B targets the official EU languages with a European multilingual tokenizer and a large non\-English data share\[[3](https://arxiv.org/html/2607.09424#bib.bib3)\]\. EuroLLM develops a family of European multilingual models across several scales, with multilingual data filtering, tokenizer design, and evaluation as central components\[[53](https://arxiv.org/html/2607.09424#bib.bib53),[52](https://arxiv.org/html/2607.09424#bib.bib52),[74](https://arxiv.org/html/2607.09424#bib.bib74)\]\. Salamandra and the subsequent ALIA family focus on European and Iberian language modeling\[[26](https://arxiv.org/html/2607.09424#bib.bib26)\]\. Apertus emphasizes fully open and compliant multilingual foundation models, and OpenEuroLLM extends this direction as a coordinated European initiative\[[5](https://arxiv.org/html/2607.09424#bib.bib5),[66](https://arxiv.org/html/2607.09424#bib.bib66)\]\.

Soofi S is complementary to these broad\-coverage efforts\. Rather than optimizing for many languages at once, it studies a narrower German–English setting in which the data mixture, annealing phase, and evaluation suite are designed around bilingual depth\. Broad European models address coverage across languages, while Soofi S tests how much capability and efficiency can be gained when one bilingual deployment setting is given a dedicated, fully documented pretraining recipe\.

##### Positioning of Soofi S\.

The closest architectural reference for Soofi S is Nemotron 3 Nano, because both use a 30B\-A3B hybrid Mamba–Transformer MoE architecture\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. The closest openness references are OLMo, OLMoE, OLMo 3, and Apertus, because they emphasize releases that enable audit and reconstruction rather than merely inference\[[27](https://arxiv.org/html/2607.09424#bib.bib27),[57](https://arxiv.org/html/2607.09424#bib.bib57),[64](https://arxiv.org/html/2607.09424#bib.bib64),[5](https://arxiv.org/html/2607.09424#bib.bib5)\]\. The closest European language references are EuroLLM, Teuken, Salamandra, Apertus, and OpenEuroLLM\[[53](https://arxiv.org/html/2607.09424#bib.bib53),[52](https://arxiv.org/html/2607.09424#bib.bib52),[74](https://arxiv.org/html/2607.09424#bib.bib74),[3](https://arxiv.org/html/2607.09424#bib.bib3),[26](https://arxiv.org/html/2607.09424#bib.bib26),[5](https://arxiv.org/html/2607.09424#bib.bib5),[66](https://arxiv.org/html/2607.09424#bib.bib66)\]\. Soofi S combines these lines in a configuration not covered by prior work: a fully documented European pretraining run, a German–English data curriculum with exact per\-source accounting, and a sparse hybrid architecture designed for long\-context, high\-concurrency serving\. The resulting model fills a gap between broadly multilingual European sovereignty efforts and efficient international open\-weight baselines: it asks whether a sovereign model can be simultaneously open, bilingual\-depth\-oriented, and competitive in capability per active parameter\.

## 6Conclusion

We presented Soofi S 30B\-A3B, a sovereign, open\-source MoE hybrid Mamba–Transformer foundation model for German and English\. Built on the Nemotron 3 Nano architecture—52 layers combining Mamba\-2, Grouped\-Query Attention, and granular MoE layers that activate roughly 3B of∼30\{\\sim\}30B parameters per token—Soofi S was pretrained on approximately 27 trillion tokens under a three\-phase Warmup–Stable–Decay curriculum: 20T tokens of diverse, quality\-tiered pretraining,∼7\{\\sim\}7T tokens of high\-quality annealing, and a length\-bucketed long\-context phase extending the usable context to 1M tokens\. Throughout, German was deliberately up\-weighted—to7\.2%7\.2\\%of the stable phase and15\.32%15\.32\\%of the annealing mixture, more than triple the multilingual share of the reference recipe—realizing the design goal of a German–English champion rather than a thinly spread multilingual model\.

The result is a model that reaches the capability frontier at a fraction of the inference cost of dense alternatives\. On a unified evaluation of 17 open base models \([Section˜4](https://arxiv.org/html/2607.09424#S4)\), Soofi S achieves the best English and German code aggregates among the measured models \(HumanEval/MBPP averages, with LBPP reported separately\), a Math\-EN score essentially tied with the strongest model, second\-best scores on GSM8K\[[14](https://arxiv.org/html/2607.09424#bib.bib14)\], the Minerva mathematics benchmarks, and INCLUDE\-DE \(tied\-best within the reported open\-weight comparison\), and English and German aggregates that match dense 14–27B models—all while activating only 3B parameters per token\. It outperforms every European sovereign baseline in our comparison—including those an order of magnitude larger in active parameters—matching or outperforming them on every German benchmark in the suite, often by 10–30 points \([Figure˜1](https://arxiv.org/html/2607.09424#S1.F1),[Figure˜8](https://arxiv.org/html/2607.09424#S4.F8), and[Figure˜13](https://arxiv.org/html/2607.09424#S4.F13)\)\. To our knowledge, this makes Soofi S the first European sovereign model to sit on the same capability\-per\-active\-parameter frontier as the strongest international open\-weight releases, the strongest open German base model in its inference\-cost class, and the strongest fully open base model in our evaluation on both English and German aggregates\.

Equally central to this work is*how*the model is released\. Trained end\-to\-end on the German Industrial AI Cloud by a consortium of German research institutions, Soofi S ships not only weights but the complete set of artifacts needed to audit and rebuild it: the full per\-source token accounting of all three pretraining phases \(including sources we evaluated and excluded\), every hyperparameter and learning\-rate stage— including a discarded final annealing stage, reported for completeness—and the training and evaluation code under permissive licenses, with licensed data sources documented through aggregate statistics and exact mixture accounting rather than redistributed raw text\. We hope this level of transparency moves the open ecosystem from open\-weight toward genuinely open\-source, and provides a reproducible template for other language communities seeking capable, efficient, sovereign foundation models\.

Future work will extend Soofi S along three axes: open post\-training \(SFT and large\-scale RL\) toward instruct and reasoning variants \(including modular reasoning in the spirit of FlexOlmo and Bar\[[78](https://arxiv.org/html/2607.09424#bib.bib78),[56](https://arxiv.org/html/2607.09424#bib.bib56)\]\), broader and deeper German evaluation suites, and continued scaling of the high\-quality German data pipeline that this release identified as the principal bottleneck for further gains\.

## Acknowledgments

This work was supported by the German Federal Ministry for Economic Affairs and Energy \(BMWE\) in the context of IPCEI\-CIS and 8ra through “Soofi: Souveräne KI für Europa” \(grant number 13IPC040A\-J\)\. Parts of it have benefited from the hessian\.AI Service Center \(funded by the Federal Ministry of Research, Technology and Space, BMFTR, grant no\. 16IS22091\) and the hessian\.AI Innovation Lab \(funded by the Hessian Ministry for Digital Strategy and Innovation, grant no\. SDIW04/0013/003\)\. We are also grateful to all the many people who have supported and enabled this project, including the Telekom Industrial AI Cloud and NVIDIA teams\. In particular, we would like to thank Pramod Kumbhar and Oleg Sudakov, whose in\-depth expertise and tremendous dedication have been invaluable to our work\. We further thank Miroslav Shaltev and Oleh Astappiev from the L3S Research Center for operating the CPU cluster and its Slurm scheduling\. We would also like to thank Christian Kotulek, Marek Soha and all their colleagues for their hard work in ensuring that the cluster runs round the clock at full capacity\. Finally, we thank Lara Lawniczak, Nora Malke and Synje Jungbehr for their project coordination, organizing project meetings, and keeping the project running smoothly\.

## References

- \[1\]J\. Ainslie, J\. Lee\-Thorp, M\. De Jong, Y\. Zemlyanskiy, F\. Lebrón, and S\. Sanghai\.Gqa: Training generalized multi\-query transformer models from multi\-head checkpoints\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, 2023\.
- \[2\]M\. Ali, M\. Brack, M\. Lübbering, E\. Wendt, A\. G\. Khan, R\. Rutmann, A\. Jude, M\. Kraus, A\. A\. Weber, F\. Stollenwerk, D\. Kaczér, F\. Mai, L\. Flek, R\. Sifa, N\. Flores\-Herr, J\. Koehler, P\. Schramowski, M\. Fromm, and K\. Kersting\.Judging quality across languages: A multilingual approach to pretraining data filtering with language models\.In C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng, editors,Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 8859–8898, Suzhou, China, Nov\. 2025\. Association for Computational Linguistics\.
- \[3\]M\. Ali, M\. Fromm, K\. Thellmann, J\. Ebert, A\. A\. Weber, R\. Rutmann, C\. Jain, M\. Lübbering, D\. Steinigen, J\. Leveling, K\. Klug, J\. S\. Buschhoff, L\. Jurkschat, H\. Abdelwahab, B\. J\. Stein, K\.\-H\. Sylla, P\. Denisov, N\. Brandizzi, Q\. Saleem, A\. Bhowmick, L\. Helmer, C\. John, P\. O\. Suarez, M\. Ostendorff, A\. Jude, L\. Manjunath, S\. Weinbach, C\. Penke, O\. Filatov, F\. Barth, P\. Mirza, L\. Weber, I\. Wendler, R\. Sifa, F\. Küch, A\. Herten, R\. Jäkel, G\. Rehm, S\. Kesselheim, J\. Köhler, and N\. Flores\-Herr\.Teuken\-7b\-base & teuken\-7b\-instruct: Towards european llms\.arXiv preprint arXiv:2410\.03730, 2025\.
- \[4\]E\. Almazrouei, H\. Alobeidli, A\. Alshamsi, A\. Cappelli, R\. Cojocaru, M\. Debbah, É\. Goffinet, D\. Hesslow, J\. Launay, Q\. Malartic, D\. Mazzotta, B\. Noune, B\. Pannier, and G\. Penedo\.The falcon series of open language models\.arXiv preprint arXiv:2311\.16867, 2023\.
- \[5\]P\. Apertus, A\. Hernández\-Cano, A\. Hägele, A\. H\. Huang, A\. Romanou, A\.\-J\. Solergibert, B\. Pasztor, B\. Messmer, D\. Garbaya, E\. F\. Ďurech, I\. Hakimi, J\. G\. Giraldo, M\. Ismayilzada, N\. Foroutan, S\. Moalla, T\. Chen, V\. Sabolčec, Y\. Xu, M\. Aerni, B\. AlKhamissi, I\. A\. Mariñas, M\. H\. Amani, M\. Ansaripour, I\. Badanin, H\. Benoit, E\. Boros, N\. Browning, F\. Bösch, M\. Böther, N\. Canova, C\. Challier, C\. Charmillot, J\. Coles, J\. Deriu, A\. Devos, L\. Drescher, D\. Dzenhaliou, M\. Ehrmann, D\. Fan, S\. Fan, S\. Gao, M\. Gila, M\. Grandury, D\. Hashemi, A\. Hoyle, J\. Jiang, M\. Klein, A\. Kucharavy, A\. Kucherenko, F\. Lübeck, R\. Machacek, T\. Manitaras, A\. Marfurt, K\. Matoba, S\. Matrenok, H\. Mendonça, F\. R\. Mohamed, S\. Montariol, L\. Mouchel, S\. Najem\-Meyer, J\. Ni, G\. Oliva, M\. Pagliardini, E\. Palme, A\. Panferov, L\. Paoletti, M\. Passerini, I\. Pavlov, A\. Poiroux, K\. Ponkshe, N\. Ranchin, J\. Rando, M\. Sauser, J\. Saydaliev, M\. A\. Sayfiddinov, M\. Schneider, S\. Schuppli, M\. Scialanga, A\. Semenov, K\. Shridhar, R\. Singhal, A\. Sotnikova, A\. Sternfeld, A\. K\. Tarun, P\. Teiletche, J\. Vamvas, X\. Yao, H\. Zhao, A\. Ilic, A\. Klimovic, A\. Krause, C\. Gulcehre, D\. Rosenthal, E\. Ash, F\. Tramèr, J\. VandeVondele, L\. Veraldi, M\. Rajman, T\. Schulthess, T\. Hoefler, A\. Bosselut, M\. Jaggi, and I\. Schlag\.Apertus: Democratizing open and compliant llms for global language environments\.arXiv preprint arXiv:2509\.14233, 2025\.
- \[6\]J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, et al\.Program synthesis with large language models\.arXiv preprint arXiv:2108\.07732, 2021\.
- \[7\]S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. Van Der Wal\.Pythia: A suite for analyzing large language models across training and scaling\.In A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett, editors,Proceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 2397–2430\. PMLR, 23–29 Jul 2023\.
- \[8\]BigScience Workshop\.BLOOM: A 176b\-parameter open\-access multilingual language model\.arXiv preprint arXiv:2211\.05100, 2022\.
- \[9\]S\. Black, S\. Biderman, E\. Hallahan, Q\. Anthony, L\. Gao, L\. Golding, H\. He, C\. Leahy, K\. McDonell, J\. Phang, M\. Pieler, U\. S\. Prashanth, S\. Purohit, L\. Reynolds, J\. Tow, B\. Wang, and S\. Weinbach\.GPT\-NeoX\-20B: An open\-source autoregressive language model\.In A\. Fan, S\. Ilic, T\. Wolf, and M\. Gallé, editors,Proceedings of BigScience Episode \#5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pages 95–136, virtual\+Dublin, May 2022\. Association for Computational Linguistics\.
- \[10\]T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei\.Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems, volume 33, pages 1877–1901, 2020\.
- \[11\]L\. Burchell, O\. de Gibert, N\. Arefyev, M\. Aulamo, M\. Bañón, P\. Chen, M\. Fedorova, L\. Guillou, B\. Haddow, J\. Hajič, J\. Helcl, E\. Henriksson, M\. Klimaszewski, V\. Komulainen, A\. Kutuzov, J\. Kytöniemi, V\. Laippala, P\. Mæhlum, B\. Malik, F\. Mehryary, V\. Mikhailov, N\. Moghe, A\. Myntti, D\. O’Brien, S\. Oepen, P\. Pal, J\. Piha, S\. Pyysalo, G\. Ramírez\-Sánchez, D\. Samuel, P\. Stepachev, J\. Tiedemann, D\. Variš, T\. Vojtěchová, and J\. Zaragoza\-Bernabeu\.An expanded massive multilingual dataset for high\-performance language technologies \(HPLT\)\.In W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 17452–17485, Vienna, Austria, July 2025\. Association for Computational Linguistics\.
- \[12\]M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, et al\.Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374, 2021\.
- \[13\]P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord\.Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018\.
- \[14\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, et al\.Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168, 2021\.
- \[15\]D\. Dai, C\. Deng, C\. Zhao, R\. Xu, H\. Gao, D\. Chen, J\. Li, W\. Zeng, X\. Yu, Y\. Wu, et al\.Deepseekmoe: Towards ultimate expert specialization in mixture\-of\-experts language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 1280–1297, 2024\.
- \[16\]T\. Dao and A\. Gu\.Transformers are ssms: Generalized models and efficient algorithms through structured state space duality\.arXiv preprint arXiv:2405\.21060, 2024\.
- \[17\]Deutsche Telekom AG\.Germany’s first AI factory for industry officially goes into operation in Munich\.[https://www\.telekom\.com/en/media/media\-information/archive/germany\-s\-first\-ai\-factory\-for\-industry\-1101670](https://www.telekom.com/en/media/media-information/archive/germany-s-first-ai-factory-for-industry-1101670), Feb\. 2026\.Industrial AI Cloud, Munich; accessed 2026\-07\-08\.
- \[18\]S\. Diao, Y\. Yang, Y\. Fu, X\. Dong, D\. Su, M\. Kliegl, Z\. Chen, P\. Belcak, Y\. Suhara, H\. Yin, et al\.Nemotron\-climb: Clustering\-based iterative data mixture bootstrapping for language model pre\-training\.Advances in Neural Information Processing Systems, 38, 2026\.
- \[19\]W\. Fedus, B\. Zoph, and N\. Shazeer\.Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity\.Journal of Machine Learning Research, 23\(120\):1–39, 2022\.
- \[20\]K\. Fujii, Y\. Tajima, S\. Mizuki, M\. Kawamura, H\. Shimada, T\. Shiotani, K\. Saito, M\. Oi, T\. Nakamura, T\. Okamoto, S\. Ishida, K\. Hattori, Y\. Ma, H\. Takamura, R\. Yokota, J\. Sakuma, and N\. Okazaki\.Rewriting pre\-training data boosts llm performance in math and code\.arXiv preprint arXiv:2505\.02881, 2026\.
- \[21\]L\. Gao, S\. Biderman, S\. Black, L\. Golding, T\. Hoppe, C\. Foster, J\. Phang, H\. He, A\. Thite, N\. Nabeshima, S\. Presser, and C\. Leahy\.The pile: An 800gb dataset of diverse text for language modeling\.arXiv preprint arXiv:2101\.00027, 2020\.
- \[22\]L\. Gao, J\. Tow, B\. Abbasi, S\. Biderman, S\. Black, A\. DiPofi, C\. Foster, L\. Golding, J\. Hsu, A\. Le Noac’h, H\. Li, K\. McDonell, N\. Muennighoff, C\. Ociepa, J\. Phang, L\. Reynolds, H\. Schoelkopf, A\. Skowron, L\. Sutawika, E\. Tang, A\. Thite, B\. Wang, K\. Wang, and A\. Zou\.The language model evaluation harness, 07 2024\.
- \[23\]GBI\-Genios Deutsche Wirtschaftsdatenbank GmbH\.GENIOS – german business and press database\.[https://www\.genios\.de/browse/Alle](https://www.genios.de/browse/Alle), 2025\.Commercially licensed corpus of German newspaper and trade\-press archives; obtained under a license that does not permit redistribution\.
- \[24\]L\. Gienapp, C\. Schröder, S\. Schweter, C\. Akiki, F\. Schlatt, A\. Zimmermann, P\. Genêt, and M\. Potthast\.The german commons \- 154 billion tokens of openly licensed text for german language models\.arXiv preprint arXiv:2510\.13996, 2025\.
- \[25\]GlaiveAI\.Reasoning\-v1\-20m, 2025\.A synthetic reasoning dataset containing 22mil\+ general reasoning questions and responses generated using deepseek\-ai/DeepSeek\-R1\-Distill\-Llama\-70B\.
- \[26\]A\. Gonzalez\-Agirre, M\. Pàmies, J\. Llop, I\. Baucells, S\. D\. Dalt, D\. Tamayo, J\. J\. Saiz, F\. Espuña, J\. Prats, J\. Aula\-Blasco, M\. Mina, I\. Pikabea, A\. Rubio, A\. Shvets, A\. Sallés, I\. Lacunza, J\. Palomar, J\. Falcão, L\. Tormo, L\. Vasquez\-Reina, M\. Marimon, O\. Pareras, V\. Ruiz\-Fernández, and M\. Villegas\.Salamandra technical report\.arXiv preprint arXiv:2502\.08489, 2025\.
- \[27\]D\. Groeneveld, I\. Beltagy, E\. Walsh, A\. Bhagia, R\. Kinney, O\. Tafjord, A\. Jha, H\. Ivison, I\. Magnusson, Y\. Wang, S\. Arora, D\. Atkinson, R\. Authur, K\. Chandu, A\. Cohan, J\. Dumas, Y\. Elazar, Y\. Gu, J\. Hessel, T\. Khot, W\. Merrill, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, V\. Pyatkin, A\. Ravichander, D\. Schwenk, S\. Shah, W\. Smith, E\. Strubell, N\. Subramani, M\. Wortsman, P\. Dasigi, N\. Lambert, K\. Richardson, L\. Zettlemoyer, J\. Dodge, K\. Lo, L\. Soldaini, N\. Smith, and H\. Hajishirzi\.OLMo: Accelerating the science of language models\.In L\.\-W\. Ku, A\. Martins, and V\. Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 15789–15809, Bangkok, Thailand, Aug\. 2024\. Association for Computational Linguistics\.
- \[28\]D\. Gurgurov and T\. Röhr\.ReasonXL: A multilingual cross\-domain reasoning corpus\.[https://huggingface\.co/datasets/toroe/Soofi\-Think\-SFT\-10B\-multilingual](https://huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual), 2026\.
- \[29\]D\. Gurgurov, T\. Röhr, and S\. Ostermann\.Nemotron\-multilingual\-reasoning: A multilingual science reasoning dataset\.[https://huggingface\.co/datasets/DGurgurov/Nemotron\-Multilingual\-Reasoning](https://huggingface.co/datasets/DGurgurov/Nemotron-Multilingual-Reasoning), 2025\.Derived from nvidia/Llama\-Nemotron\-Post\-Training\-Dataset with machine\-translated reasoning traces\.
- \[30\]A\. Hägele, E\. Bakouch, A\. Kosson, L\. B\. Allal, L\. Von Werra, and M\. Jaggi\.Scaling laws and compute\-optimal training beyond fixed training durations\.Advances in Neural Information Processing Systems, 37:76232–76264, 2024\.
- \[31\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt\.Measuring massive multitask language understanding, 2021\.
- \[32\]D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt\.Measuring mathematical problem solving with the math dataset\.In J\. Vanschoren and S\. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021\.
- \[33\]J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, O\. Vinyals, J\. Rae, and L\. Sifre\.An empirical analysis of compute\-optimal large language model training\.InAdvances in Neural Information Processing Systems, volume 35, pages 30016–30030\. Curran Associates, Inc\., 2022\.
- \[34\]C\.\-P\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. Ginsburg\.Ruler: What’s the real context size of your long\-context language models?, 2024\.
- \[35\]S\. Hu, Y\. Tu, X\. Han, C\. He, G\. Cui, X\. Long, Z\. Zheng, Y\. Fang, Y\. Huang, W\. Zhao, et al\.Minicpm: Unveiling the potential of small language models with scalable training strategies\.arXiv preprint arXiv:2404\.06395, 2024\.
- \[36\]M\. Idahl, B\. Droste, B\. Plüster, and J\. P\. Harries\.propella\-1: Multi\-property document annotation for llm data curation at scale, 2026\.
- \[37\]M\. Idahl, J\. Tiedemann, S\. Pyysalo, D\. Salinas, T\. Galica, S\. Qian, T\. N\. Mateiu, Z\. Li, A\. Lokrantz, F\. Vitiugin, A\. F\. T\. Martins, J\. Kanerva, F\. Ginter, M\. Lindemann, T\. Isbister, B\. Moell, J\. Lindh, J\. Hajič, J\. Jitsev, A\. Kutuzov, S\. Oepen, and G\. Ramírez\-Sánchez\.Multisynt/mt: Trillion\-token multi\-parallel pre\-training data translated across 36 languages, 2026\.
- \[38\]A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\.\-A\. Lachaux, P\. Stock, T\. Le Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. El Sayed\.Mistral 7b\.arXiv preprint arXiv:2310\.06825, 2023\.
- \[39\]A\. Q\. Jiang, A\. Sablayrolles, A\. Roux, A\. Mensch, B\. Savary, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, E\. Bou Hanna, F\. Bressand, G\. Lengyel, G\. Bour, G\. Lample, L\. R\. Lavaud, L\. Saulnier, M\.\-A\. Lachaux, P\. Stock, S\. Subramanian, S\. Yang, S\. Antoniak, T\. Le Scao, T\. Gervet, T\. Lavril, T\. Wang, T\. Lacroix, and W\. El Sayed\.Mixtral of experts\.arXiv preprint arXiv:2401\.04088, 2024\.
- \[40\]J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei\.Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361, 2020\.
- \[41\]J\. Krajewski, J\. Ludziejewski, K\. Adamczewski, M\. Pióro, M\. Krutul, S\. Antoniak, K\. Ciebiera, K\. Król, T\. Odrzygóźdź, P\. Sankowski, et al\.Scaling laws for fine\-grained mixture of experts\.arXiv preprint arXiv:2402\.07871, 2024\.
- \[42\]M\. Kraus, R\. Härle, S\. Sztwiertnia, A\. G\. Khan, M\. Ali, M\. Fromm, and K\. Kersting\.Klettermix: Climbing toward high\-quality german pretraining data\.arXiv preprint arXiv:2606\.03773, 2026\.
- \[43\]T\. Kwiatkowski, J\. Palomaki, O\. Redfield, M\. Collins, A\. Parikh, C\. Alberti, D\. Epstein, I\. Polosukhin, J\. Devlin, K\. Lee, K\. Toutanova, L\. Jones, M\. Kelcey, M\.\-W\. Chang, A\. M\. Dai, J\. Uszkoreit, Q\. Le, and S\. Petrov\.Natural questions: A benchmark for question answering research\.Transactions of the Association for Computational Linguistics, 7:452–466, 2019\.
- \[44\]H\. Kydlíček, G\. Penedo, and L\. von Werra\.Finepdfs\.[https://huggingface\.co/datasets/HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs), 2025\.
- \[45\]A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, Y\. Wu, B\. Neyshabur, G\. Gur\-Ari, and V\. Misra\.Solving quantitative reasoning problems with language models\.In S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh, editors,Advances in Neural Information Processing Systems, volume 35, pages 3843–3857\. Curran Associates, Inc\., 2022\.
- \[46\]A\. Liesenfeld, M\. Dingemanse, D\. Blankvoort, N\. Kalra, and A\. R\. Golkhandan\.European open source ai definitions\.[https://osai\-index\.eu/osai\-definitions](https://osai-index.eu/osai-definitions), 2026\.Accessed: 2026\-07\-06\.
- \[47\]A\. H\. Liu, K\. Khandelwal, S\. Subramanian, V\. Jouault, A\. Rastogi, A\. Sadé, A\. Jeffares, A\. Jiang, A\. Cahill, A\. Gavaudan, A\. Sablayrolles, A\. Héliou, A\. You, A\. Ehrenberg, A\. Lo, A\. Eliseev, A\. Calvi, A\. Sooriyarachchi, B\. Bout, B\. Rozière, B\. D\. Monicault, C\. Lanfranchi, C\. Barreau, C\. Courtot, D\. Grattarola, D\. Dabert, D\. de las Casas, E\. Chane\-Sane, F\. Ahmed, G\. Berrada, G\. Ecrepont, G\. Guinet, G\. Novikov, G\. Kunsch, G\. Lample, G\. Martin, G\. Gupta, J\. Ludziejewski, J\. Rute, J\. Studnia, J\. Amar, J\. Delas, J\. S\. Roberts, K\. Yadav, K\. Chandu, K\. Jain, L\. Aitchison, L\. Fainsin, L\. Blier, L\. Zhao, L\. Martin, L\. Saulnier, L\. Gao, M\. Buyl, M\. Jennings, M\. Pellat, M\. Prins, M\. Poirée, M\. Guillaumin, M\. Dinot, M\. Futeral, M\. Darrin, M\. Augustin, M\. Chiquier, M\. Schimpf, N\. Grinsztajn, N\. Gupta, N\. Raghuraman, O\. Bousquet, O\. Duchenne, P\. Wang, P\. von Platen, P\. Jacob, P\. Wambergue, P\. Kurylowicz, P\. R\. Muddireddy, P\. Chagniot, P\. Stock, P\. Agrawal, Q\. Torroba, R\. Sauvestre, R\. Soletskyi, R\. Menneer, S\. Vaze, S\. Barry, S\. Gandhi, S\. Waghjale, S\. Gandhi, S\. Ghosh, S\. Mishra, S\. Aithal, S\. Antoniak, T\. L\. Scao, T\. Cachet, T\. S\. Sorg, T\. Lavril, T\. N\. Saada, T\. Chabal, T\. Foubert, T\. Robert, T\. Wang, T\. Lawson, T\. Bewley, T\. Bewley, T\. Edwards, U\. Jamil, U\. Tomasini, V\. Nemychnikova, V\. Phung, V\. Maladière, V\. Richard, W\. Bouaziz, W\.\-D\. Li, W\. Marshall, X\. Li, X\. Yang, Y\. E\. Ouahidi, Y\. Wang, Y\. Tang, and Z\. Ramzi\.Ministral 3\.arXiv preprint arXiv:2601\.08584, 2026\.
- \[48\]Z\. Liu, Z\. Yang, Y\. Chen, C\. Lee, M\. Shoeybi, B\. Catanzaro, and W\. Ping\.Acereason\-nemotron 1\.1: Advancing math and code reasoning through sft and rl synergy\.arXiv preprint arXiv:2506\.13284, 2025\.
- \[49\]I\. Loshchilov and F\. Hutter\.Decoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101, 2017\.
- \[50\]M\. Lübbering, T\. Ruland, R\. Rutmann, F\. Stollenwerk, D\. Fitzek, M\. Fromm, A\. Weber, R\. Sifa, N\. Flores\-Herr, J\. Köhler, et al\.Modalities, a pytorch\-native framework for large\-scale llm training and research\.arXiv preprint arXiv:2602\.08387, 2026\.
- \[51\]R\. K\. Mahabadi, S\. Satheesh, S\. Prabhumoye, M\. Patwary, M\. Shoeybi, and B\. Catanzaro\.Nemotron\-cc\-math: A 133 billion\-token\-scale high quality math pretraining dataset\.arXiv preprint arXiv:2508\.15096, 2025\.
- \[52\]P\. H\. Martins, J\. Alves, P\. Fernandes, N\. M\. Guerreiro, R\. Rei, A\. Farajian, M\. Klimaszewski, D\. M\. Alves, J\. Pombal, N\. Boizard, M\. Faysse, P\. Colombo, F\. Yvon, B\. Haddow, J\. G\. C\. de Souza, A\. Birch, and A\. F\. T\. Martins\.Eurollm\-9b: Technical report\.arXiv preprint arXiv:2506\.04079, 2025\.
- \[53\]P\. H\. Martins, P\. Fernandes, J\. Alves, N\. M\. Guerreiro, R\. Rei, D\. M\. Alves, J\. Pombal, A\. Farajian, M\. Faysse, M\. Klimaszewski, P\. Colombo, B\. Haddow, J\. G\. C\. de Souza, A\. Birch, and A\. F\. T\. Martins\.Eurollm: Multilingual language models for europe\.arXiv preprint arXiv:2409\.16235, 2024\.
- \[54\]A\. Matton, T\. Sherborne, D\. Aumiller, E\. Tommasone, M\. Alizadeh, J\. He, R\. Ma, M\. Voisin, E\. Gilsenan\-McMahon, and M\. Gallé\.On leakage of code generation evaluation datasets\.In Y\. Al\-Onaizan, M\. Bansal, and Y\.\-N\. Chen, editors,Findings of the Association for Computational Linguistics: EMNLP 2024, pages 13215–13223, Miami, Florida, USA, Nov\. 2024\. Association for Computational Linguistics\.
- \[55\]Mistral AI\.Mistral small 3\.1\.[https://mistral\.ai/news/mistral\-small\-3\-1/](https://mistral.ai/news/mistral-small-3-1/), 2025\.Model release, March 2025\.
- \[56\]J\. Morrison, S\. Adhikesaven, A\. Bhagia, M\. Zaharia, N\. A\. Smith, and S\. Min\.Train separately, merge together: Modular post\-training with mixture\-of\-experts, 2026\.
- \[57\]N\. Muennighoff, L\. Soldaini, D\. Groeneveld, K\. Lo, J\. Morrison, S\. Min, W\. Shi, E\. P\. Walsh, O\. Tafjord, N\. Lambert, Y\. Gu, S\. Arora, A\. Bhagia, D\. Schwenk, D\. Wadden, A\. Wettig, B\. Hui, T\. Dettmers, D\. Kiela, A\. Farhadi, N\. A\. Smith, P\. W\. Koh, A\. Singh, and H\. Hajishirzi\.OLMoE: Open mixture\-of\-experts language models\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24–28, 2025\. OpenReview\.net, 2025\.
- \[58\]MultiSynt\.Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025\.An automatic translations into 2 languages of Glaive AI reasoning dataset\.
- \[59\]NVIDIA, :, A\. Basant, A\. Khairnar, A\. Paithankar, A\. Khattar, A\. Renduchintala, A\. Malte, A\. Bercovich, A\. Hazare, A\. Rico, A\. Ficek, A\. Kondratenko, A\. Shaposhnikov, A\. Bukharin, A\. Taghibakhshi, A\. Barton, A\. S\. Mahabaleshwarkar, A\. Shen, A\. Tao, A\. Guan, A\. Shors, A\. Mandarwal, A\. Mehta, A\. Venkatesan, A\. Sharabiani, A\. Aithal, A\. Poojary, A\. Dattagupta, B\. Buddharaju, B\. Zhu, B\. Simkin, B\. Kartal, B\. D\. Rouhani, B\. Chen, B\. Ginsburg, B\. Norick, B\. Yu, B\. Catanzaro, C\. Wang, C\. Truong, C\. Mungekar, C\. Patel, C\. Alexiuk, C\. Munley, C\. Parisien, D\. Su, D\. Afrimi, D\. Korzekwa, D\. Rohrer, D\. Gitman, D\. Mosallanezhad, D\. Narayanan, D\. Rekesh, D\. Yared, D\. Pykhtar, D\. Ahn, D\. Riach, E\. Long, E\. Ning, E\. Chung, E\. Galinkin, E\. Bakhturina, G\. Prasad, G\. Shen, H\. Qian, H\. Elisha, H\. Sharma, H\. Ross, H\. Ngo, H\. Sahota, H\. Wang, H\. C\. Shin, H\. Huang, I\. Cunningham, I\. Gitman, I\. Moshkov, J\. Jung, J\. Kautz, J\. P\. Scowcroft, J\. Casper, J\. Zhang, J\. Zeng, J\. Zhang, J\. Xue, J\. Huang, J\. Conway, J\. Kamalu, J\. Cohen, J\. Jennings, J\. V\. Vialard, J\. Yi, J\. Parmar, K\. Briski, K\. Cheung, K\. Luna, K\. Wyss, K\. Santhanam, K\. Kong, K\. Pawelec, K\. Anik, K\. Li, K\. Ahmadian, L\. McAfee, L\. Sleiman, L\. Derczynski, L\. Vega, M\. R\. de Melo, M\. N\. Sreedhar, M\. Chochowski, M\. Cai, M\. Kliegl, M\. Stepniewska\-Dziubinska, M\. Novikov, M\. Samadi, M\. Price, M\. Boubdir, M\. Boone, M\. Evans, M\. Bien, M\. Zawalski, M\. Martinez, M\. Chrzanowski, M\. Shoeybi, M\. Patwary, N\. Dhameja, N\. Assaf, N\. Habibi, N\. Bhatia, N\. Pope, N\. Tajbakhsh, N\. K\. Juluru, O\. Rybakov, O\. Hrinchuk, O\. Kuchaiev, O\. Olabiyi, P\. Ribalta, P\. Subramanian, P\. Chadha, P\. Molchanov, P\. Dykas, P\. Jin, P\. Bialecki, P\. Januszewski, P\. Thalasta, P\. Gaikwad, P\. Varshney, P\. Gundecha, P\. Tredak, R\. K\. Mahabadi, R\. Patel, R\. El\-Yaniv, R\. Rajan, R\. Cheruvu, R\. Shahbazyan, R\. Borkar, R\. Gala, R\. Waleffe, R\. Zhang, R\. J\. Hewett, R\. Prenger, S\. Jain, S\. Kriman, S\. Satheesh, S\. Kaji, S\. Yurick, S\. Muralidharan, S\. Narenthiran, S\. Bak, S\. Sameni, S\. Han, S\. Ramasamy, S\. Ghosh, S\. T\. Sreenivas, S\. Thomas, S\. Diao, S\. Gopal, S\. Prabhumoye, S\. Toshniwal, S\. Ding, S\. Singh, S\. Jain, S\. Majumdar, S\. Singhal, S\. Alborghetti, S\. N\. Akter, T\. Kong, T\. Moon, T\. Hliwiak, T\. Asida, T\. Wang, T\. Konuk, T\. Vashishth, T\. Poon, U\. Karpas, V\. Noroozi, V\. Srinivasan, V\. Korthikanti, V\. Fugro, V\. Kalluru, V\. Kurin, V\. Lavrukhin, W\. U\. Ahmad, W\. Du, W\. Byeon, X\. Lu, X\. Dong, Y\. Karnati, Y\. Choi, Y\. Zhang, Y\. Lin, Y\. Fu, Y\. Suhara, Z\. Dong, Z\. Li, Z\. Zhu, and Z\. Chen\.Nvidia nemotron nano 2: An accurate and efficient hybrid mamba\-transformer reasoning model, 2025\.
- \[60\]NVIDIA, :, A\. Blakeman, A\. Basant, A\. Khattar, A\. Renduchintala, A\. Bercovich, A\. Ficek, A\. Bjorlin, A\. Taghibakhshi, A\. S\. Deshmukh, A\. S\. Mahabaleshwarkar, A\. Tao, A\. Shors, A\. Aithal, A\. Poojary, A\. Dattagupta, B\. Buddharaju, B\. Chen, B\. Ginsburg, B\. Wang, B\. Norick, B\. Butterfield, B\. Catanzaro, C\. del Mundo, C\. Dong, C\. Harvey, C\. Parisien, D\. Su, D\. Korzekwa, D\. Yin, D\. Gitman, D\. Mosallanezhad, D\. Narayanan, D\. Fridman, D\. Rekesh, D\. Ma, D\. Pykhtar, D\. Ahn, D\. Riach, D\. Stosic, E\. Long, E\. Segal, E\. Evans, E\. Chung, E\. Galinkin, E\. Bakhturina, E\. Dobrowolska, F\. Jia, F\. Liu, G\. Prasad, G\. Shen, G\. Liu, G\. Chen, H\. Qian, H\. Ngo, H\. Liu, H\. Li, I\. Gitman, I\. Karmanov, I\. Moshkov, I\. Golan, J\. Kautz, J\. P\. Scowcroft, J\. Casper, J\. Seppanen, J\. Lu, J\. Sewall, J\. Zeng, J\. You, J\. Zhang, J\. Zhang, J\. Huang, J\. Xue, J\. Huang, J\. Conway, J\. Kamalu, J\. Barker, J\. Cohen, J\. Jennings, J\. Parmar, K\. Sapra, K\. Briski, K\. Chumachenko, K\. Luna, K\. Santhanam, K\. Kong, K\. Sivamani, K\. Pawelec, K\. Anik, K\. Li, L\. McAfee, L\. Derczynski, L\. Pavao, L\. Vega, L\. Voegtle, M\. Bala, M\. R\. de Melo, M\. N\. Sreedhar, M\. Chochowski, M\. Kliegl, M\. Stepniewska\-Dziubinska, M\. Le, M\. Novikov, M\. Samadi, M\. Andersch, M\. Evans, M\. Martinez, M\. Chrzanowski, M\. Ranzinger, M\. Blaz, M\. Smelyanskiy, M\. Fawzy, M\. Shoeybi, M\. Patwary, N\. Lee, N\. Tajbakhsh, N\. Xu, O\. Rybakov, O\. Kuchaiev, O\. Delalleau, O\. Nitski, P\. Chadha, P\. Shamis, P\. Micikevicius, P\. Molchanov, P\. Dykas, P\. Fischer, P\.\-Y\. Aquilanti, P\. Bialecki, P\. Varshney, P\. Gundecha, P\. Tredak, R\. Karimi, R\. Kandu, R\. El\-Yaniv, R\. Joshi, R\. Waleffe, R\. Zhang, S\. Kavanaugh, S\. Jain, S\. Kriman, S\. Lym, S\. Satheesh, S\. Muralidharan, S\. Narenthiran, S\. Anandaraj, S\. Bak, S\. Kashirsky, S\. Han, S\. Acharya, S\. Ghosh, S\. T\. Sreenivas, S\. Clay, S\. Thomas, S\. Prabhumoye, S\. Pachori, S\. Toshniwal, S\. Prayaga, S\. Jain, S\. Das, S\. Kierat, S\. Majumdar, S\. Han, S\. Singhal, S\. Niverty, S\. Alborghetti, S\. Panguluri, S\. Bhendigeri, S\. N\. Akter, S\. Migacz, T\. Shiri, T\. Kong, T\. Roman, T\. Ronen, T\. Saar, T\. Konuk, T\. Rintamaki, T\. Poon, U\. De, V\. Noroozi, V\. Singh, V\. Korthikanti, V\. Kurin, W\. U\. Ahmad, W\. Du, W\. Ping, W\. Dai, W\. Byeon, X\. Ren, Y\. Xu, Y\. Choi, Y\. Zhang, Y\. Lin, Y\. Suhara, Z\. Yu, Z\. Li, Z\. Li, Z\. Zhu, Z\. Yang, and Z\. Chen\.Nemotron\-h: A family of accurate and efficient hybrid mamba\-transformer models\.arXiv preprint arXiv:2504\.03624, 2025\.
- \[61\]NVIDIA, :, A\. Blakeman, A\. Grattafiori, A\. Basant, A\. Gupta, A\. Khattar, A\. Renduchintala, A\. Vavre, A\. Shukla, A\. Bercovich, A\. Ficek, A\. Shaposhnikov, A\. Kondratenko, A\. Bukharin, A\. Milesi, A\. Taghibakhshi, A\. Liu, A\. Barton, A\. S\. Mahabaleshwarkar, A\. Klein, A\. Zuker, A\. Geifman, A\. Shen, A\. Bhiwandiwalla, A\. Tao, A\. Guan, A\. Mandarwal, A\. Mehta, A\. Aithal, A\. Poojary, A\. Ahamed, A\. K\. Thekkumpate, A\. Dattagupta, B\. Zhu, B\. Sadeghi, B\. Simkin, B\. Lanir, B\. Schifferer, B\. Nushi, B\. Kartal, B\. D\. Rouhani, B\. Ginsburg, B\. Norick, B\. Soubasis, B\. Kisacanin, B\. Yu, B\. Catanzaro, C\. del Mundo, C\. Hwang, C\. Wang, C\.\-P\. Hsieh, C\. Zhang, C\. Yu, C\. Mungekar, C\. Patel, C\. Alexiuk, C\. Parisien, C\. Neale, D\. Mosk\-Aoyama, D\. Su, D\. Corneil, D\. Afrimi, D\. Rohrer, D\. Serebrenik, D\. Gitman, D\. Levy, D\. Stosic, D\. Mosallanezhad, D\. Narayanan, D\. Nathawani, D\. Rekesh, D\. Yared, D\. Kakwani, D\. Ahn, D\. Riach, D\. Stosic, E\. Minasyan, E\. Lin, E\. Long, E\. P\. Long, E\. Lantz, E\. Evans, E\. Ning, E\. Chung, E\. Harper, E\. Tramel, E\. Galinkin, E\. Pounds, E\. Briones, E\. Bakhturina, F\. Ladhak, F\. Wang, F\. Jia, F\. Soares, F\. Chen, F\. Galko, F\. Siino, G\. H\. Agam, G\. Ajjanagadde, G\. Bhatt, G\. Prasad, G\. Armstrong, G\. Shen, G\. Batmaz, G\. Nalbandyan, H\. Qian, H\. Sharma, H\. Ross, H\. Ngo, H\. Sahota, H\. Wang, H\. Soni, H\. Upadhyay, H\. Mao, H\. C\. Nguyen, H\. Q\. Nguyen, I\. Cunningham, I\. Shahaf, I\. Gitman, I\. Loshchilov, I\. Moshkov, I\. Putterman, J\. Kautz, J\. P\. Scowcroft, J\. Casper, J\. Mitra, J\. Glick, J\. Chen, J\. Oliver, J\. Zhang, J\. Zeng, J\. Lou, J\. Zhang, J\. Huang, J\. Conway, J\. Guman, J\. Kamalu, J\. Greco, J\. Cohen, J\. Jennings, J\. Daw, J\. V\. Vialard, J\. Yi, J\. Parmar, K\. Xu, K\. Zhu, K\. Briski, K\. Cheung, K\. Luna, K\. Santhanam, K\. Shih, K\. Kong, K\. Bhardwaj, K\. C\. Puvvada, K\. Pawelec, K\. Anik, L\. McAfee, L\. Sleiman, L\. Derczynski, L\. Ding, L\. Liebenwein, L\. Vega, M\. Grover, M\. V\. Segbroeck, M\. R\. de Melo, M\. N\. Sreedhar, M\. Kilaru, M\. Ashkenazi, M\. Romeijn, M\. Cai, M\. Kliegl, M\. Moosaei, M\. Novikov, M\. Samadi, M\. Corpuz, M\. Wang, M\. Price, M\. Boone, M\. Evans, M\. Martinez, M\. Chrzanowski, M\. Shoeybi, M\. Patwary, N\. Mulepati, N\. Hereth, N\. Assaf, N\. Habibi, N\. Zmora, N\. Haber, N\. Sessions, N\. Bhatia, N\. Jukar, N\. Pope, N\. Ludwig, N\. Tajbakhsh, N\. Juluru, O\. Hrinchuk, O\. Kuchaiev, O\. Delalleau, O\. Olabiyi, O\. U\. Argov, O\. Xie, P\. Chadha, P\. Shamis, P\. Molchanov, P\. Morkisz, P\. Dykas, P\. Jin, P\. Xu, P\. Januszewski, P\. P\. Thombre, P\. Varshney, P\. Gundecha, Q\. Miao, R\. K\. Mahabadi, R\. El\-Yaniv, R\. Zilberstein, R\. Shafipour, R\. Harang, R\. Izzo, R\. Shahbazyan, R\. Garg, R\. Borkar, R\. Gala, R\. Islam, R\. Waleffe, R\. Watve, R\. Koren, R\. Zhang, R\. J\. Hewett, R\. Prenger, R\. Timbrook, S\. Mahdavi, S\. Modi, S\. Kriman, S\. Kariyappa, S\. Satheesh, S\. Kaji, S\. Pasumarthi, S\. Narentharen, S\. Narenthiran, S\. Bak, S\. Kashirsky, S\. Poulos, S\. Mor, S\. Ramasamy, S\. Acharya, S\. Ghosh, S\. T\. Sreenivas, S\. Thomas, S\. Fan, S\. Gopal, S\. Prabhumoye, S\. Pachori, S\. Toshniwal, S\. Ding, S\. Singh, S\. Sun, S\. Ithape, S\. Majumdar, S\. Singhal, S\. Alborghetti, S\. Ge, S\. D\. Devare, S\. K\. Barua, S\. Panguluri, S\. Gupta, S\. Priyadarshi, S\. N\. Akter, T\. Bui, T\.\-D\. Ene, T\. Kong, T\. Do, T\. Blankevoort, T\. Balough, T\. Asida, T\. B\. Natan, T\. Konuk, T\. Vashishth, U\. Karpas, U\. De, V\. Noorozi, V\. Noroozi, V\. Srinivasan, V\. Elango, V\. Korthikanti, V\. Kurin, V\. Lavrukhin, W\. Jiang, W\. U\. Ahmad, W\. Du, W\. Ping, W\. Zhou, W\. Jennings, W\. Zhang, W\. Prazuch, X\. Ren, Y\. Karnati, Y\. Choi, Y\. Meyer, Y\.\-F\. Wu, Y\. Zhang, Y\. Lin, Y\. Geifman, Y\. Fu, Y\. Subara, Y\. Suhara, Y\. Gao, Z\. Moshe, Z\. Dong, Z\. Liu, Z\. Chen, and Z\. Yan\.Nemotron 3 nano: Open, efficient mixture\-of\-experts hybrid mamba\-transformer model for agentic reasoning\.arXiv preprint arXiv:2512\.20848, 2025\.
- \[62\]NVIDIA\.Deutsche telekom and NVIDIA launch industrial AI cloud\.[https://blogs\.nvidia\.com/blog/germany\-industrial\-ai\-cloud\-launch/](https://blogs.nvidia.com/blog/germany-industrial-ai-cloud-launch/), Nov\. 2025\.Accessed 2026\-07\-08\.
- \[63\]S\. Oepen, N\. Arefev, M\. Aulamo, M\. Bañón, M\. Buljan, L\. Burchell, L\. Charpentier, P\. Chen, M\. Fedorova, O\. de Gibert, B\. Haddow, J\. Hajič, J\. Helcl, A\. Kutuzov, V\. Laippala, Z\. Li, R\. Luukkonen, B\. Malik, V\. Mikhailov, A\. Myntti, D\. O’Brien, L\. Poláková, S\. Pyysalo, G\. R\. Sánchez, J\. Siewert, P\. Stepachev, J\. Tiedemann, T\. Vahtola, D\. Variš, F\. Vitiugin, T\. Vojtěchová, and J\. Zaragoza\.Hplt 3\.0: Very large\-scale multilingual resources for llms and mt\. mono\- and bi\-lingual data, multilingual evaluation, and pre\-trained models\.arXiv preprint arXiv:2511\.01066, 2026\.
- \[64\]T\. Olmo, :, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison, J\. Morrison, J\. Poznanski, K\. Lo, L\. Soldaini, M\. Jordan, M\. Chen, M\. Noukhovitch, N\. Lambert, P\. Walsh, P\. Dasigi, R\. Berry, S\. Malik, S\. Shah, S\. Geng, S\. Arora, S\. Gupta, T\. Anderson, T\. Xiao, T\. Murray, T\. Romero, V\. Graf, A\. Asai, A\. Bhagia, A\. Wettig, A\. Liu, A\. Rangapur, C\. Anastasiades, C\. Huang, D\. Schwenk, H\. Trivedi, I\. Magnusson, J\. Lochner, J\. Liu, L\. J\. V\. Miranda, M\. Sap, M\. Morgan, M\. Schmitz, M\. Guerquin, M\. Wilson, R\. Huff, R\. L\. Bras, R\. Xin, R\. Shao, S\. Skjonsberg, S\. Z\. Shen, S\. S\. Li, T\. Wilde, V\. Pyatkin, W\. Merrill, Y\. Chang, Y\. Gu, Z\. Zeng, A\. Sabharwal, L\. Zettlemoyer, P\. W\. Koh, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi\.Olmo 3\.arXiv preprint arXiv:2512\.13961, 2026\.
- \[65\]Open Source Initiative\.The open source ai definition, version 1\.0\.[https://opensource\.org/ai/open\-source\-ai\-definition](https://opensource.org/ai/open-source-ai-definition), 2024\.Accessed: 2026\-07\-06\.
- \[66\]OpenEuroLLM Consortium\.OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025\.Project website, accessed 2026\-06\-25\.
- \[67\]G\. Penedo\.Finewiki, 2025\.Source: Wikimedia Enterprise Snapshot API \(https://api\.enterprise\.wikimedia\.com/v2/snapshots\)\. Text licensed under CC BY\-SA 4\.0 with attribution to Wikipedia contributors\.
- \[68\]G\. Penedo, H\. Kydlíček, A\. Lozhkov, M\. Mitchell, C\. Raffel, L\. Von Werra, T\. Wolf, et al\.The fineweb datasets: Decanting the web for the finest text data at scale\.Advances in Neural Information Processing Systems, 37:30811–30849, 2024\.
- \[69\]Pleias and AI Alliance\.SYNTH: An open generalist synthetic dataset for training small reasoning models\.[https://huggingface\.co/datasets/PleIAs/SYNTH](https://huggingface.co/datasets/PleIAs/SYNTH), 2025\.Dataset comprising 79,648,272 text samples \(over 41 billion words\) derived from the synthetic amplification of 58,698 Wikipedia and Wikibooks articles\. Licensed under CC\-BY\-SA 4\.0\.
- \[70\]J\. Poznanski, A\. Rangapur, J\. Borchardt, J\. Dunkelberger, R\. Huff, D\. Lin, C\. Wilhelm, K\. Lo, and L\. Soldaini\.olmocr: Unlocking trillions of tokens in pdfs with vision language models\.arXiv preprint arXiv:2502\.18443, 2025\.
- \[71\]Qwen Team\.Qwen3\.5: Towards native multimodal agents, February 2026\.
- \[72\]A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever\.Language models are unsupervised multitask learners\.Technical report, OpenAI, 2019\.
- \[73\]C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu\.Exploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of Machine Learning Research, 21\(140\):1–67, 2020\.
- \[74\]M\. M\. Ramos, D\. M\. Alves, H\. Gisserot\-Boukhlef, J\. Alves, P\. H\. Martins, P\. Fernandes, J\. Pombal, N\. M\. Guerreiro, R\. Rei, N\. Boizard, A\. Farajian, M\. Klimaszewski, J\. G\. C\. de Souza, B\. Haddow, F\. Yvon, P\. Colombo, A\. Birch, and A\. F\. T\. Martins\.Eurollm\-22b: Technical report\.arXiv preprint arXiv:2602\.05879, 2026\.
- \[75\]D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman\.Gpqa: A graduate\-level google\-proof q&a benchmark\.arXiv preprint arXiv:2311\.12022, 2023\.
- \[76\]A\. Romanou, N\. Foroutan, A\. Sotnikova, S\. H\. Nelaturu, S\. Singh, R\. Maheshwary, M\. Altomare, Z\. Chen, M\. Haggag, S\. A, A\. Amayuelas, A\. H\. Amirudin, D\. Boiko, M\. Chang, J\. Chim, G\. Cohen, A\. K\. Dalmia, A\. Diress, S\. Duwal, D\. Dzenhaliou, D\. Florez, F\. Farestam, J\. M\. Imperial, S\. Islam, P\. Isotalo, M\. Jabbarishiviari, B\. F\. Karlsson, E\. Khalilov, C\. Klamm, F\. Koto, D\. Krzemiński, G\. de Melo, S\. Montariol, Y\. Nan, J\. Niklaus, J\. Novikova, J\. S\. Obando Ceron, D\. Paul, E\. Ploeger, J\. Purbey, S\. Rajwal, S\. S\. Ravi, S\. Rydell, R\. Santhosh, D\. Sharma, M\. Prifti Skenduli, A\. Soltani Moakhar, B\. moakhar, A\. Tarun, A\. T\. Wasi, T\. Weerasinghe, S\. Yilmaz, M\. Zhang, I\. Schlag, M\. Fadaee, S\. Hooker, and A\. Bosselut\.Include: Evaluating multilingual language understanding with regional knowledge\.In Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu, editors,International Conference on Learning Representations, volume 2025, pages 83291–83322, 2025\.
- \[77\]N\. Shazeer, A\. Mirhoseini, K\. Maziarz, A\. Davis, Q\. Le, G\. Hinton, and J\. Dean\.Outrageously large neural networks: The sparsely\-gated mixture\-of\-experts layer\.arXiv preprint arXiv:1701\.06538, 2017\.
- \[78\]W\. Shi, A\. Bhagia, K\. Farhat, N\. Muennighoff, J\. Morrison, E\. Walsh, D\. Schwenk, S\. Longpre, J\. Poznanski, A\. Ettinger, et al\.Flexolmo: Open language models for flexible data use\.Advances in Neural Information Processing Systems, 38:165943–165974, 2026\.
- \[79\]L\. Soldaini, R\. Kinney, A\. Bhagia, D\. Schwenk, D\. Atkinson, R\. Authur, B\. Bogin, K\. Chandu, J\. Dumas, Y\. Elazar, V\. Hofmann, A\. Jha, S\. Kumar, L\. Lucy, X\. Lyu, N\. Lambert, I\. Magnusson, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, A\. Ravichander, K\. Richardson, Z\. Shen, E\. Strubell, N\. Subramani, O\. Tafjord, E\. Walsh, L\. Zettlemoyer, N\. Smith, H\. Hajishirzi, I\. Beltagy, D\. Groeneveld, J\. Dodge, and K\. Lo\.Dolma: An open corpus of three trillion tokens for language model pretraining research\.In L\.\-W\. Ku, A\. Martins, and V\. Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 15725–15788, Bangkok, Thailand, Aug\. 2024\. Association for Computational Linguistics\.
- \[80\]D\. Su, K\. Kong, Y\. Lin, J\. Jennings, B\. Norick, M\. Kliegl, M\. Patwary, M\. Shoeybi, and B\. Catanzaro\.Nemotron\-CC: Transforming Common Crawl into a refined long\-horizon pretraining dataset\.In W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 2459–2475, Vienna, Austria, July 2025\. Association for Computational Linguistics\.
- \[81\]M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. Le, E\. Chi, D\. Zhou, and J\. Wei\.Challenging BIG\-bench tasks and whether chain\-of\-thought can solve them\.In A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki, editors,Findings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051, Toronto, Canada, July 2023\. Association for Computational Linguistics\.
- \[82\]A\. Talmor, J\. Herzig, N\. Lourie, and J\. Berant\.CommonsenseQA: A question answering challenge targeting commonsense knowledge\.In J\. Burstein, C\. Doran, and T\. Solorio, editors,Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\), pages 4149–4158, Minneapolis, Minnesota, June 2019\. Association for Computational Linguistics\.
- \[83\]G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. bastien Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\.\-T\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. yeong Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\.\-B\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot\.Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786, 2025\.
- \[84\]H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\.\-A\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar, A\. Rodriguez, A\. Joulin, E\. Grave, and G\. Lample\.LLaMA: Open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971, 2023\.
- \[85\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale, D\. Bikel, L\. Blecher, C\. Canton Ferrer, M\. Chen, G\. Cucurull, D\. Esiobu, J\. Fernandes, J\. Fu, W\. Fu, B\. Fuller, C\. Gao, V\. Goswami, N\. Goyal, A\. Hartshorn, S\. Hosseini, R\. Hou, H\. Inan, M\. Kardas, V\. Kerkez, M\. Khabsa, I\. Kloumann, A\. Korenev, P\. S\. Koura, M\.\-A\. Lachaux, T\. Lavril, J\. Lee, D\. Liskovich, Y\. Lu, Y\. Mao, X\. Martinet, T\. Mihaylov, P\. Mishra, I\. Molybog, Y\. Nie, A\. Poulton, J\. Reizenstein, R\. Rungta, K\. Saladi, A\. Schelten, R\. Silva, E\. M\. Smith, R\. Subramanian, X\. E\. Tan, B\. Tang, R\. Taylor, A\. Williams, J\. X\. Kuan, P\. Xu, Z\. Yan, I\. Zarov, Y\. Zhang, A\. Fan, M\. Kambadur, S\. Narang, A\. Rodriguez, R\. Stojnic, S\. Edunov, and T\. Scialom\.Llama 2: Open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288, 2023\.
- \[86\]J\. Vendrow, E\. Vendrow, S\. Beery, and A\. Madry\.Do large language model benchmarks test reliability?arXiv preprint arXiv:2502\.03461, 2025\.
- \[87\]Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, et al\.Mmlu\-pro: A more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems, 37:95266–95290, 2024\.
- \[88\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388, 2025\.
- \[89\]S\. Zhang, S\. Roller, N\. Goyal, M\. Artetxe, M\. Chen, S\. Chen, C\. Dewan, M\. Diab, X\. Li, X\. V\. Lin, T\. Mihaylov, M\. Ott, S\. Shleifer, K\. Shuster, D\. Simig, P\. S\. Koura, A\. Sridhar, T\. Wang, and L\. Zettlemoyer\.OPT: Open pre\-trained transformer language models\.arXiv preprint arXiv:2205\.01068, 2022\.
- \[90\]W\. Zhong, R\. Cui, Y\. Guo, Y\. Liang, S\. Lu, Y\. Wang, A\. Saied, W\. Chen, and N\. Duan\.AGIEval: A human\-centric benchmark for evaluating foundation models\.In K\. Duh, H\. Gomez, and S\. Bethard, editors,Findings of the Association for Computational Linguistics: NAACL 2024, pages 2299–2314, Mexico City, Mexico, June 2024\. Association for Computational Linguistics\.
- \[91\]C\. Zhou, H\. Lyu, X\. Lin, H\. Zhao, J\. Guo, X\. Zhang, S\. Xue, Q\. Ma, J\. Zhou, Y\. Wang, and Z\. Liu\.Ultradata\-math, 2026\.

## Appendix AAuthor Contributions

### A\.1Training

1. 1\.Pretraining stack development and evaluation: Max Lübbering, Richard Rutmann, Timm Ruland, David Fitzek, Mehdi Ali
2. 2\.Model architecture, training methodology and framework\-correctness validation: Timm Ruland, David Fitzek, Max Lübbering, Richard Rutmann
3. 3\.Compute infrastructure, cluster benchmarking and interconnect tuning: David Fitzek, Timm Ruland, Richard Rutmann, Max Lübbering
4. 4\.Distributed\-training scaling and memory/throughput optimization: David Fitzek, Timm Ruland, Max Lübbering, Richard Rutmann
5. 5\.Training execution, stability analysis, emergency debugging, framework bug fixes and experiment tracking: Timm Ruland, David Fitzek, Max Lübbering, Richard Rutmann

### A\.2Data

1. 1\.Base model data acquisition: Michael Fromm, Alex Jude, Abbas Khan, Ruben Härle, Maurice Kraus, Jan Pfister, Daniil Gurgurov
2. 2\.Pretraining data mixture: Michael Fromm
3. 3\.Data curation infrastructure and experimentation: Michael Fromm, Alex Jude, Abbas Khan, Ruben Härle, Maurice Kraus, Richard Rutmann, Mehdi Ali, Max Lübbering, Maximilian Idahl
4. 4\.Data preprocessing / tokenization pipeline: Richard Rutmann, Max Lübbering, Alex Jude
5. 5\.Mid\- and long\-context data curation and experimentation: Michael Fromm, Alex Jude, Abbas Khan, Ruben Härle, Maurice Kraus, Sebastian Sztwiertnia, Tom Röhr, Sebastian von Rohrscheidt

### A\.3Evaluation

1. 1\.Evaluation methodology and infrastructure: Maximilian Idahl, Benedikt Droste, Alex Jude, Abbas Khan

### A\.4Other

1. 1\.Mentorship, advising, program management, and broader strategy: Nicolas Flores\-Herr, Simon Gottschalk, Jörg Bienert, Kristian Kersting, Andreas Hotho, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky
2. 2\.Technical leadership and cross\-workstream contributions: Mehdi Ali, Michael Fromm, Max Lübbering, Sebastian Sztwiertnia, Tom Röhr

## Appendix BDetailed Pretraining Data Composition

This appendix gives the full per\-source token accounting underlying the mixture flow diagram in Section[3](https://arxiv.org/html/2607.09424#S3)\([Figure˜3](https://arxiv.org/html/2607.09424#S3.F3)\)\. TablesLABEL:tab:phase1\-sourcesand[7](https://arxiv.org/html/2607.09424#A2.T7)cover Phase 1, TablesLABEL:tab:phase2\-sourcesand[9](https://arxiv.org/html/2607.09424#A2.T9)cover Phase 2 \(annealing\), and Tables[10](https://arxiv.org/html/2607.09424#A2.T10)and[11](https://arxiv.org/html/2607.09424#A2.T11)cover the Phase 3 long\-context extension\. Rows with zero epochs are sources we enumerated but excluded from training\.

Table 6:Phase 1 \(diverse pretraining\) data composition\. “Raw” is the source token count in billions; “Ep\.” is the number of epochs; “Eff\.” is the effective token count after epoching; “Share” is the percentage of Phase 1 effective tokens\. Rows with zero epochs are enumerated but excluded from training\. Raw and effective counts are taken from the source dataset cards \(HuggingFace\) and may reflect different tokenizers; they are approximate\. Exact tokenizer counts of consumed tokens appear in Table[2](https://arxiv.org/html/2607.09424#S2.T2)\.SourceSubset / QualityRawEp\.Eff\.Share[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality26\.0378\.00\.3%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)Medium\-High\-Quality16\.9116\.90\.1%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)Medium\-Quality53\.500\.00\.0%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality\-Synthetic93\.52187\.00\.8%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)Medium\-High\-Quality\-Synthetic2122\.812122\.89\.2%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)HQ\-Translated\-To\-English39\.6279\.20\.3%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)MHQ\-Translated\-To\-English26\.8126\.80\.1%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)HQ\-Translated\-To\-English\-Synthetic157\.82315\.61\.4%[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality\-DQA8\.0216\.00\.1%Nemotron\-CC\-v2\.1 subtotal2842\.312\.3%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)High\-Quality613\.731841\.18\.0%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Medium\-High\-Quality545\.61545\.62\.4%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Medium\-Quality2200\.800\.00\.0%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)High\-Quality\-Synthetic1257\.022514\.010\.9%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Diverse QA692\.41692\.43\.0%[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Translated\-Diverse QA \(DE\)2\.024\.00\.0%Nemotron\-CC\-v2\.0 subtotal5597\.124\.3%[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)High\-Quality553\.031659\.07\.2%[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)Medium\-High\-Quality504\.01504\.02\.2%[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)Medium\-Quality2023\.000\.00\.0%[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)High\-Synthetic\-Diverse QA Pairs499\.52999\.04\.3%Nemotron\-CC\-v1\.0 subtotal3162\.013\.7%[nvidia/Nemotron\-CC\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Code-v1)Actual427\.931283\.75\.6%[nvidia/Nemotron\-Pretraining\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1)Synthetic174\.92349\.81\.5%[nvidia/Nemotron\-Pretraining\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1)Actual125\.03375\.01\.6%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-code\-review71\.382142\.760\.6%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-question\-answering212\.522425\.041\.8%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-rewriting76\.842153\.680\.7%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-student\-teacher28\.34256\.680\.2%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-transpilation24\.09248\.180\.2%[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)Actual180\.03540\.02\.3%Code subtotal3374\.8414\.6%[nvidia/Nemotron\-Pretraining\-Specialized\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1)—270\.751353\.55\.9%[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Math SFT190\.65953\.04\.1%[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Code SFT58\.56351\.01\.5%[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)General SFT87\.53262\.51\.1%Specialized \+ SFT subtotal2920\.012\.7%[nvidia/Nemotron\-CC\-Math\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1)3plus133\.04532\.02\.3%[nvidia/Nemotron\-CC\-Math\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1)4plus52\.04208\.00\.9%[nvidia/Nemotron\-CC\-Math\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1)v173\.04292\.01\.3%[nvidia/Nemotron\-Math\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Math-v2)high2\.7410\.80\.0%[nvidia/Nemotron\-Math\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Math-v2)medium2\.048\.00\.0%[nvidia/Nemotron\-Math\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Math-v2)low1\.244\.80\.0%[openbmb/UltraData\-Math](https://huggingface.co/datasets/openbmb/UltraData-Math)en88\.04352\.01\.5%Mathematics subtotal1407\.66\.1%[MultiSynt/MT\-Reasoning](https://huggingface.co/datasets/MultiSynt/MT-Reasoning)en35\.8271\.60\.3%[nvidia/AceReason\-1\.1\-SFT](https://huggingface.co/datasets/nvidia/AceReason-1.1-SFT)en30\.0260\.00\.3%[HuggingFaceFW/finewiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki)en10\.0550\.00\.2%[allenai/dolma3\_pool](https://huggingface.co/datasets/allenai/dolma3_pool)pdfs600\.01600\.02\.6%[PleIAs/Synth](https://huggingface.co/datasets/PleIAs/Synth)en60\.02120\.00\.5%[HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs)en1190\.6511190\.655\.2%English \(other\) subtotal2092\.259\.1%[HuggingFaceFW/finepdfs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs)de177\.562355\.121\.5%[HPLT\-3\-Top10%](https://hplt-project.org/datasets/v3.0)de60\.98\.4511\.562\.2%[coral\-nlp/german\-commons](https://huggingface.co/datasets/coral-nlp/german-commons)de154\.561154\.560\.7%[MultiSynt/MT\-Nemotron\-CC](https://huggingface.co/datasets/MultiSynt/MT-Nemotron-CC)de117\.02234\.01\.0%[Genios](https://www.genios.de/browse/Alle)de150\.02300\.01\.3%[HuggingFaceFW/finewiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki)de3\.527\.00\.0%[MultiSynt/MT\-Reasoning](https://huggingface.co/datasets/MultiSynt/MT-Reasoning)de42\.0284\.00\.4%[DGurgurov/Nemotron\-Multilingual\-Reasoning](https://huggingface.co/datasets/DGurgurov/Nemotron-Multilingual-Reasoning)de2\.024\.00\.0%[PleIAs/Synth](https://huggingface.co/datasets/PleIAs/Synth)de2\.424\.80\.0%German subtotal1655\.047\.2%Total16,350\.0423,051\.13100\.0%Table 7:Phase 1 composition by category, with Nemotron 3 Nano’s reported shares for comparison\. These categories cover the full Phase 1 mixture of23,051\.1323\{,\}051\.13B effective tokens \(100%100\\%\)\.Table 8:Phase 2 \(high\-quality annealing\) data composition\. Columns as in TableLABEL:tab:phase1\-sources\. Rows with zero epochs are enumerated but excluded from training\. Raw and effective counts are taken from the source dataset cards \(HuggingFace\) and may reflect different tokenizers; they are approximate\. Exact tokenizer counts of consumed tokens appear in Table[2](https://arxiv.org/html/2607.09424#S2.T2)\.SourceSubsetRawEp\.Eff\.Web[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality26\.0126\.0[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality\-Synthetic93\.5193\.5[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)HQ\-Translated\-To\-English39\.6139\.6[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)HQ\-Translated\-To\-English\-Synthetic157\.800\.0[nvidia/Nemotron\-CC\-v2\.1](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1)High\-Quality\-DQA8\.018\.0[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)High\-Quality613\.700\.0[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)High\-Quality\-Synthetic1257\.00\.5628\.5[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Diverse QA692\.41692\.4[nvidia/Nemotron\-CC\-v2\.0](https://huggingface.co/datasets/nvidia/Nemotron-CC-v2)Translated\-Diverse QA \(DE\)2\.000\.0[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)High\-Quality553\.000\.0[nvidia/Nemotron\-CC\-v1\.0](https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html)High\-Synthetic\-Diverse QA Pairs499\.500\.0[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)Web5\.2115\.21[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)pdfs240\.01240\.0[karpathy/climbmix\-400b\-shuffle](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle)en400\.02800\.0[HuggingFaceFW/finepdfs\-edu](https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu)en142\.01142\.0Web subtotal2675\.21Code[nvidia/Nemotron\-Pretraining\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1)Synthetic174\.91174\.9[nvidia/Nemotron\-Pretraining\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1)Actual125\.01125\.0[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-code\-review71\.38171\.38[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-question\-answering212\.521212\.52[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-rewriting76\.84176\.84[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-student\-teacher28\.34128\.34[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)synthetic\-transpilation24\.09124\.09[nvidia/Nemotron\-Pretraining\-Code\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)Actual180\.01180\.0[tokyotech\-llm/swallow\-code\-v2](https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2)stage549\.8299\.6[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)Code40\.0140\.0Code subtotal1032\.67Mathematics[nvidia/Nemotron\-CC\-Math\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1)3plus133\.01133\.0[nvidia/Nemotron\-CC\-Math\-v1](https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1)4plus52\.0152\.0[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)Math21\.34121\.34[openbmb/UltraData\-Math](https://huggingface.co/datasets/openbmb/UltraData-Math)en88\.01\.4123\.2Mathematics subtotal329\.54SFT[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Math SFT190\.62381\.2[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Code SFT58\.52117\.0[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)General SFT87\.5187\.5[nvidia/Nemotron\-Agentic\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Agentic-v1)agentic1\.022\.0[nvidia/Nemotron\-Competitive\-Programming\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Competitive-Programming-v1)code\-sft50\.02100\.0[nvidia/Nemotron\-Instruction\-Following\-Chat\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Instruction-Following-Chat-v1)General SFT1\.523\.0[nvidia/Nemotron\-Math\-Proofs\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Math-Proofs-v1)stem\-sft6\.0212\.0[nvidia/Nemotron\-Math\-v2](https://huggingface.co/datasets/nvidia/Nemotron-Math-v2)stem\-sft30\.0260\.0[nvidia/Nemotron\-RLHF\-GenRM\-v1](https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1)—0\.521\.0[nvidia/Nemotron\-Science\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Science-v1)stem\-sft0\.521\.0[nvidia/Nemotron\-SFT\-Agentic\-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2)agentic1\.523\.0[nvidia/Nemotron\-SFT\-Competitive\-Programming\-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2)code\-sft20\.0240\.0[nvidia/Nemotron\-SFT\-Instruction\-Following\-Chat\-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v2)General SFT3\.026\.0[nvidia/Nemotron\-SFT\-Math\-v3](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v3)stem\-sft1\.022\.0[nvidia/Nemotron\-SFT\-Multilingual\-v1](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v1)multilingual\-sft3\.527\.0[nvidia/Nemotron\-SFT\-OpenCode\-v1](https://huggingface.co/datasets/nvidia/Nemotron-SFT-OpenCode-v1)code\-sft7\.0214\.0[nvidia/Nemotron\-SFT\-Safety\-v1](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1)safety\-sft0\.0320\.06[nvidia/Nemotron\-SpecializedDomains\-Finance\-v1](https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1)stem\-sft4\.028\.0[nvidia/Nemotron\-SWE\-v1](https://huggingface.co/datasets/nvidia/Nemotron-SWE-v1)code\-sft0\.721\.4[nvidia/Nemotron\-SFT\-SWE\-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-SWE-v2)code\-sft2\.525\.0[nvidia/AceReason\-1\.1\-SFT](https://huggingface.co/datasets/nvidia/AceReason-1.1-SFT)en30\.0130\.0[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)QA25\.8125\.8[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)Instruction\-Data18\.41118\.41[AIML\-TUDA/QA\-base](https://huggingface.co/datasets/AIML-TUDA/QA-base)en0\.143101\.43SFT subtotal926\.80Reasoning[nvidia/Nemotron\-Pretraining\-Specialized\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1)—270\.71270\.7[allenai/dolma3\_dolmino\_pool](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)Thinking37\.6137\.6[nvidia/Nemotron\-Pretraining\-Specialized\-v1\.1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1)en9\.319\.3[MultiSynt/MT\-Reasoning](https://huggingface.co/datasets/MultiSynt/MT-Reasoning)en35\.8135\.8Reasoning subtotal353\.40Wiki[HuggingFaceFW/finewiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki)en10\.0220\.0Wiki subtotal20\.00German[HuggingFaceFW/finepdfs\-edu](https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu)de20\.0240\.0[AIML\-TUDA/QA\-base](https://huggingface.co/datasets/AIML-TUDA/QA-base)de0\.187101\.87[HPLT\-4\-Top10%](https://hplt-project.org/datasets/v4.0)de291\.01291\.0[coral\-nlp/german\-commons](https://huggingface.co/datasets/coral-nlp/german-commons)de154\.5600\.0[German Translation of ClimbMix](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle)de571\.01571\.0[Genios](https://www.genios.de/browse/Alle)de150\.000\.0[HuggingFaceFW/finewiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki)de3\.513\.5[MultiSynt/MT\-Reasoning](https://huggingface.co/datasets/MultiSynt/MT-Reasoning)de42\.0142\.0[DGurgurov/Nemotron\-Multilingual\-Reasoning](https://huggingface.co/datasets/DGurgurov/Nemotron-Multilingual-Reasoning)de2\.012\.0[toroe/Soofi\-Think\-SFT\-10B\-multilingual](https://huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual)de7\.0214\.0German subtotal965\.37Total6,303\.0Table 9:Phase 2 \(annealing\) composition by category, with Nemotron 3 Nano’s reported shares for comparison\. Percentages are shares of the6,3036\{,\}303B annealing pool\. For the seven\-category main\-text view in Figure[3](https://arxiv.org/html/2607.09424#S3.F3), the Reasoning/STEM\-SFT bucket is folded into Reasoning\.Table 10:Long\-context phase: per\-domain token budgets, document counts, and source priorities for the released data pool\. The run consumed∼\\sim100\.66B of the 188\.5B pool \([Section˜3\.4](https://arxiv.org/html/2607.09424#S3.SS4)\)\.DomainTokens \(B\)DocumentsSources \(priority order\)Web28\.783,051,984[ClimbMix](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle)\>\>[OlmoOCR](https://huggingface.co/datasets/allenai/dolma3_dolmino_pool)\>\>[FinePDFs](https://huggingface.co/datasets/HuggingFaceFW/finepdfs)Code9\.51344,625[Swallow\-Code\-v2](https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2)\>\>[Nemotron\-Pretraining\-Code\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1),[v2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2)Mathematics6\.732,526,637[openbmb/UltraData\-Math\-L3](https://huggingface.co/datasets/openbmb/UltraData-Math); no data\>\>64KGerman6\.68714,98940%[HPLT\-4\-Top10%](https://hplt-project.org/datasets/v4.0)\+\+60%[German Translation of ClimbMix](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle)General SFT31\.254,386,847[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Code SFT29\.242,541,896[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Math SFT76\.317,990,170[nvidia/Nemotron\-Pretraining\-SFT\-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1)Total188\.4921,557,148Table 11:Long\-context phase: number of documents per sequence\-length bucket and domain in the released pool\. “–” denotes a bucket left unpopulated\. Unlike the idealized symmetric schema, the realized counts do not halve cleanly across buckets\.
## Appendix CProxy Data\-Mixture Ablations

Before fixing the data recipe used for Soofi S, we ran controlled small\-scale data\-mixture ablations\. These runs were designed as a proxy study: their role was to select a robust German–English data family under a much cheaper setup, not to predict the absolute performance of the final 30B\-A3B hybrid\-Mamba MoE model\. To make the comparison realistic for the final setting, we held the English backbone fixed and changed only the German or multilingual part of the mixture\. English therefore remained the dominant language in every candidate, contributing roughly 90–93% of the 100B\-token proxy budget, while the remaining budget tested different allocations among German web, PDF, synthetic, commons, news, and lightly multilingual sources\. This isolates the main design question for Soofi S: how to spend the scarce non\-English budget without confounding the result with changes in English data\.

##### Ablated mixtures\.

The English backbone follows a fixed Nemotron\-style composition of web, code, math, SFT, and PDF sources\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. The ablation identifiersD01,D03–D13correspond to the non\-English mixtures summarized in[Table˜12](https://arxiv.org/html/2607.09424#A3.T12)\.

The candidates cover several qualitatively different hypotheses\.D01combines all German source families with a strong HPLT\-3\[[63](https://arxiv.org/html/2607.09424#bib.bib63)\]web anchor and small amounts of German\-Commons\[[24](https://arxiv.org/html/2607.09424#bib.bib24)\]and Genios\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]\.D12uses the same high\-level shares but swaps the German HPLT quality classifier for the Propella\-filtered variant\[[36](https://arxiv.org/html/2607.09424#bib.bib36)\]\.D04is the only explicitly multilingual candidate, adding Spanish, Italian, and French HPLT\-3 top\-decile data to the otherwise German tail\.D08tests an almost pure German FinePDFs\[[44](https://arxiv.org/html/2607.09424#bib.bib44)\]tail,D09andD10emphasize MultiSynt\[[37](https://arxiv.org/html/2607.09424#bib.bib37)\], andD11tests a broader HPLT\-3 German filter by moving from the top decile to the top two deciles\.

Table 12:Non\-English mixture shares for the evaluated proxy ablations\. Entries are percentages of the 100B\-token proxy budget recorded in the ablation sheet\. “0\.1q” denotes the top 10% HPLT\-3 filter, and “0\.2q” denotes the top 20% filter\. The “Other HPLT” column is non\-German HPLT\-3 data \(Spanish, Italian, and French\) and is nonzero only for the explicitly multilingual ablationD04\.
##### Proxy training setup\.

All ablations used the same training configuration and differed only in the chosen tokenized dataset blend\. The proxy model was based on the Qwen3\-1\.7B pretraining recipe\[[88](https://arxiv.org/html/2607.09424#bib.bib88)\]\. We trained with bf16 mixed precision, global batch size 256, and micro\-batch size 1\. The runs were launched on four Leonardo nodes with four GPUs per node\. Checkpoints were written during training and evaluated offline with thelm\-evaluation\-harness\[[22](https://arxiv.org/html/2607.09424#bib.bib22)\]pipeline\. We used a training batch size of 2,097,152 tokens per optimizer step, and the final comparable checkpoint at step 47,683 corresponds to∼100​B\{\\sim\}100Btrained tokens\. Checkpoints were evaluated at approximately 10B\-token intervals, giving a learning curve rather than a single endpoint for each candidate mixture\.

##### Evaluation protocol\.

For mixture selection we used the same English–German harness family as in the main report\. The primary scalar criterion was the average rank over four suite\-level signals: English bits\-per\-byte, German bits\-per\-byte, English normalized rank\-choice accuracy, and German normalized rank\-choice accuracy\. Bits\-per\-byte is lower\-better; normalized rank\-choice accuracy is higher\-better\. We preferred ranks over raw\-score averaging because the four metrics have different scales\. All terminal averages in Table[13](https://arxiv.org/html/2607.09424#A3.T13)are computed using only the step\-47,683 checkpoint\. We explicitly ignore the near\-duplicate step\-47,680 endpoint so that the final evaluation is not double\-counted\.

Table 13:Proxy data\-mixture ablations at the terminal 100B\-token checkpoint, computed using only step 47,683\. “EN/DE bpb” are suite\-level bits\-per\-byte scores \(lower is better\); “EN/DE acc\.” are normalized rank\-choice accuracies in percent \(higher is better\)\.R100​BR\_\{100\\mathrm\{B\}\}is the average rank over these four terminal suite metrics\. Best values are bolded and second\-best values are underlined\.
##### Results\.

The rank\-based selection criterion gives a clear winner\.D01has the best terminal average rank \(R100​B=1\.75R\_\{100\\mathrm\{B\}\}=1\.75\) and the best average rank across the full training trace \(Rall=3\.55R\_\{\\mathrm\{all\}\}=3\.55\)\. At the final checkpoint it is best on English bits\-per\-byte and German normalized accuracy, second on English normalized accuracy, and close to the strongest group on German bits\-per\-byte\. This balance matters because several alternatives win a single metric but are less stable overall\.D06, for example, achieves the lowest German bits\-per\-byte, but its average rank is only sixth at the final checkpoint and seventh over the full trace\.D12is the strongest runner\-up and is slightly better on English normalized accuracy, but it gives back performance on both bits\-per\-byte suites and on German normalized accuracy relative toD01\.

The source\-level patterns are also informative\. A mixture consisting almost entirely of German FinePDFs \(D08\) performs competitively on English but collapses on German bits\-per\-byte, suggesting that document\-style PDF text alone is too narrow for the German tail\. Heavy MultiSynt mixtures improve some German likelihood tasks but do not produce the best bilingual aggregate \(D09,D10\)\. The explicitly multilingual mixture \(D04\) is a strong candidate and obtains the second\-best German bits\-per\-byte score, but it does not matchD01on the combined English–German selection criterion\. The comparison betweenD01andD12further suggests that the original HPLT\-DE component is preferable to the Propella\-filtered substitute under this proxy setup, even when the high\-level source shares are otherwise unchanged\.

Task\-group aggregates support the same conclusion\.D01is best on the English likelihood suite, best on the English math and English QA bits\-per\-byte aggregates, and best on the German code bits\-per\-byte and German grammar\-fluency aggregates\. Its weaker German QA and German math likelihood scores are offset by stronger German rank\-choice and fluency performance, which better matches the intended downstream profile of a German–English base model\.

##### Selection for the full run\.

We therefore used the mixture represented byD01as the basis for the full Soofi S data recipe\. The proxy result was not copied mechanically into the final 26\.68T\-token curriculum: the final training plan still separates broad pretraining from high\-quality annealing, up\-weights German in both phases, and adds the long\-context extension described in Section[2\.2](https://arxiv.org/html/2607.09424#S2.SS2)\. Nevertheless, the ablations provide the empirical justification for selecting a balanced German data family—HPLT\-DE, German FinePDFs, MultiSynt, German\-Commons, and Genios—rather than optimizing a single benchmark, a single German source, or a broad multilingual tail in isolation\.

## Appendix DFurther Dataset Information

##### Genios\.

The Genios corpus\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]is a commercially licensed collection of German\-language newspaper and trade\-press archives obtained from GBI\-Genios\. It comprises 916 distinct publications with 193\.1M articles \(∼\\sim57\.6B words\) spanning 2010–2025, with per\-year volumes between 9\.7M and 14\.4M documents \(Figure[15](https://arxiv.org/html/2607.09424#A4.F15)\)\. The collection is dominated by regional daily newspapers \(e\.g\.*Rheinische Post*,*Rhein\-Zeitung*,*Neue Westfälische*\), complemented by a long tail of national outlets and specialist trade and academic periodicals \(Table[14](https://arxiv.org/html/2607.09424#A4.T14)\)\. Articles average∼\\sim298 words\. Because the data was purchased under a commercial license, it cannot be redistributed; we therefore report aggregate corpus statistics only\.

Table 14:Composition of the Genios corpus\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]: the ten largest publications by document count, out of 916 German newspaper and trade\-press sources totalling 193\.1M articles and∼\\sim57\.6B words\.2,0102\{,\}0102,0112\{,\}0112,0122\{,\}0122,0132\{,\}0132,0142\{,\}0142,0152\{,\}0152,0162\{,\}0162,0172\{,\}0172,0182\{,\}0182,0192\{,\}0192,0202\{,\}0202,0212\{,\}0212,0222\{,\}0222,0232\{,\}0232,0242\{,\}0242,0252\{,\}02505510101515Documents \(millions\)Figure 15:Temporal distribution of the 193\.1M Genios articles\[[23](https://arxiv.org/html/2607.09424#bib.bib23)\]\. Coverage is roughly uniform across 2010–2025 \(9\.7–14\.4M documents per year\)\.

## Appendix EFurther Base Model Evaluations

##### Long\-context\.

![Refer to caption](https://arxiv.org/html/2607.09424v1/x16.png)Figure 16:RULER accuracy averaged over the selected subtasks, by input context length \(4K–1M\)\. \(a\) all 13 subtasks; \(b\) all subtasks except CWE; \(c\) CWE \(common\-word extraction\) only, where Soofi S degrades sharply past 32K\. Shaded band = accuracy gap between the two models\.We evaluate long\-context behaviour with RULER\[[34](https://arxiv.org/html/2607.09424#bib.bib34)\]on the checkpoint after the long\-context stage \(Section[2\.2](https://arxiv.org/html/2607.09424#S2.SS2)\), comparing against Nemotron 3 Nano 30B\-A3B\. Because the two models share an identical backbone, any difference in long\-context accuracy isolates the effect of the long\-context*data*and continuation recipe rather than the architecture\. Figure[16](https://arxiv.org/html/2607.09424#A5.F16)reports accuracy averaged over RULER subtasks as a function of input length from 4K to 1M tokens\.

On the full 13\-subtask suite Soofi S trails the reference by a mean of6\.86\.8points across lengths, reaching5050versus6060at 1M \(Figure[16](https://arxiv.org/html/2607.09424#A5.F16)a\)\. The gap is almost entirely attributable to a single subtask, common\-word extraction \(CWE\), which requires aggregating and reproducing the most frequently occurring words across the entire input rather than retrieving a localized span\. Excluding CWE, the two models track each other to within a mean of4\.44\.4points across all lengths and to within∼6\{\\sim\}6points at 1M \(5454versus6060\), i\.e\. effectively on par out to the full context window \(Figure[16](https://arxiv.org/html/2607.09424#A5.F16)b\)\. On CWE itself Soofi S matches the reference up to 32K but degrades sharply at longer inputs \(Figure[16](https://arxiv.org/html/2607.09424#A5.F16)c\), falling to∼3%\{\\sim\}3\\%at 256K–1M while Nemotron retains6060–64%64\\%\.

We attribute this regression to the long\-context data mixture rather than the backbone, for two reasons\. First, the two models are architecturally identical, so the fixed\-size Mamba\-2 recurrent state and the66\-of\-5252attention layers cannot by themselves explain a gap against the same architecture trained on a different long\-context blend\. Second, the reference recipe deliberately targets this exact class of task: the Nemotron long\-context phase devotes its blend to long\-context document\-QA data \(scaled3×3\\timesover the previous generation\) together with a dedicated slice of synthetic*retrieval\-focused*data at up to 256K tokens, added specifically to improve RULER\-style subtasks, with the remainder being down\-weighted high\-quality pretraining data\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\. Our long\-context pool \(Section[3\.4](https://arxiv.org/html/2607.09424#S3.SS4)\), by contrast, is dominated by length\-bucketed SFT and general document text sampled uniformly, and contains no dedicated retrieval\- or aggregation\-oriented long\-context data\. The two runs consume a comparable long\-context token budget \(∼100\.66\{\\sim\}100\.66B here versus∼121\{\\sim\}121B for the reference\[[61](https://arxiv.org/html/2607.09424#bib.bib61)\]\), so the difference reflects mixture composition rather than long\-context training volume\. We flag CWE beyond 32K as a known limitation and a concrete target for the next iteration of the long\-context data pipeline, specifically, adding retrieval\- and aggregation\-style synthetic data in the 32K–1M range, and note that, this single subtask aside, Soofi S retains near\-parity with its architecture\-matched reference across the full 1M\-token range\.

## Appendix FCheckpoint Merging Ablations

In addition to selecting a single late\-stage checkpoint, we evaluated a series of post\-hoc checkpoint merges over the annealing trajectory\. These experiments were intended to test whether weight\-space averaging could reduce checkpoint noise at the end of training and improve robustness without adding any new training tokens\. In all cases, checkpoints came from the same training run or from direct continuations of it, so the models had identical architecture, tokenizer, tensor layout, and optimizer history\. The merges therefore average model weights only; optimizer states were not included\.

The relevant baseline for this comparison isiter\_1056000, the checkpoint selected for the Section[4](https://arxiv.org/html/2607.09424#S4)base\-model evaluation and used to initialize the long\-context extension\. This checkpoint is the model we would have released without any post\-hoc merge\. Later final\-annealing checkpoints and their merges are therefore treated as ablations against this selected checkpoint, not as replacements by default\.

##### Merge implementation\.

All merges were performed in the original distributed\-checkpoint layout with a shared reference checkpoint\. Floating\-point tensors were accumulated infloat32, while non\-floating tensors were copied from the reference checkpoint\. We considered three families of merge weights\. First, uniform averaging assigns equal weight to each checkpoint in a window\. Second, exponential averaging assigns larger weight to later checkpoints, with decay valuesα∈\{0\.2,0\.8\}\\alpha\\in\\\{0\.2,0\.8\\\}; largerα\\alphakeeps the merge closer to the end of the trajectory\. Third, two\-checkpoint manual blends mix the penultimate or earlier endpoint with the final checkpoint using either0\.2/0\.80\.2/0\.8or0\.1/0\.90\.1/0\.9weights\. All manual weights were normalized and constrained to be non\-negative\.

Table 15:Checkpoint\-merge strategies evaluated after annealing\. “Reference” is the checkpoint whose metadata and non\-floating tensors were used for the merged export\.
##### Evaluation\.

We evaluated 22 merged checkpoints with the same benchmark harness used for late\-stage checkpoint selection and compared them directly to the selectediter\_1056000checkpoint\. For the main comparison we used four aggregate suite metrics: English bits\-per\-byte, German bits\-per\-byte, English normalized rank\-choice accuracy, and German normalized rank\-choice accuracy\. Bits\-per\-byte metrics are lower\-better, while normalized accuracies are higher\-better\. As in the data\-mixture ablations, we summarize the trade\-off using an average rank over the four aggregate metrics rather than averaging raw scores with different scales\. One two\-checkpoint final merge had incomplete bits\-per\-byte coverage in the filtered evaluation file and was therefore excluded from the rank table\.

Table 16:Checkpoint merges compared against the selected base checkpoint used for Section[4](https://arxiv.org/html/2607.09424#S4)and long\-context initialization\. Accuracies are reported as percentages\.RRis the average rank over English bpb, German bpb, English normalized accuracy, and German normalized accuracy, computed over the selected checkpoint and all complete merge evaluations\.
##### Outcome\.

The merge ablation did not reveal a uniformly better model than the selectediter\_1056000checkpoint\. The best merged variants were concentrated in the final annealing window: the uniform average over all 13 final\-window checkpoints had the best aggregate rank, and the late\-biased exponential average withα=0\.8\\alpha=0\.8was nearly tied\. These merges slightly improved the German suite metrics relative toiter\_1056000: German bits\-per\-byte improved from0\.36560\.3656to0\.36370\.3637, and German normalized accuracy improved from80\.1080\.10to80\.3480\.34\. However, the selected checkpoint remained better on the English suite metrics, with English bits\-per\-byte0\.43900\.4390and English normalized accuracy77\.3977\.39, compared with0\.43960\.4396and77\.2877\.28for the best uniform merge\. This trade\-off explains why we did not treat checkpoint merging as a decisive source of additional capability\. The final\-window merges provide useful evidence that the end of annealing lies in a stable basin and that small German\-side gains are possible through weight averaging\. At the same time, those gains come with small regressions on the English aggregate metrics and do not clearly dominate the checkpoint already used for the long\-context continuation and the main Section[4](https://arxiv.org/html/2607.09424#S4)evaluation\. We therefore report the merge results for transparency, but keepiter\_1056000as the primary selected base checkpoint\.

Similar Articles

Soofi – Sovereign Open Source Foundation Models

Hacker News Top

Soofi introduces Soofi S, a 30B parameter Mixture-of-Experts open source foundation model trained on 27 trillion tokens, targeting industrial AI applications in German and English. The model is part of a European sovereign AI initiative.

Apertus – Open Foundation Model for Sovereign AI

Hacker News Top

Apertus is a fully open foundation model for sovereign AI, developed by the Swiss AI Initiative. It is open weights, open data, open science, compliant with EU AI Act, and competitive with top open models at 8B and 70B parameters, supporting over 1000 languages.

Sumi: Open Uniform Diffusion Language Model from Scratch

Hugging Face Daily Papers

Sumi is a 7B uniform diffusion language model pretrained from scratch on 1.5T tokens, achieving competitive performance on knowledge and reasoning tasks while being fully open-source with released weights and training recipe.