Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
摘要
This paper shows that four seemingly minor architectural choices—normalization, GQA, pretraining context length, and sliding window attention—have a compoundingly negative effect on long context extensibility, dropping performance by up to 47% when combined. The authors release OlmPool, a set of 26 comparable 7B models, after 170,000 GPU hours of controlled ablations.
查看缓存全文
缓存时间: 2026/08/12 08:34
# Seemingly Minor Architectural Choices Impact Long Context Extension
Source: [https://arxiv.org/html/2608.10296](https://arxiv.org/html/2608.10296)
## Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
Amanda Bertsch1,2Luca Soldaini1Matthew R\. Gormley2Graham Neubig2 Hannaneh Hajishirzi1,3Kyle Lo1,3Dirk Groeneveld1
1
Ai22Carnegie Mellon University3University of Washington###### Abstract
One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy\. However, we demonstrate that this is not the case in the long context setting\. Specifically, we show that a set of four minor architectural decisions — all made by at least one of the Olmo, Llama, and Qwen dense model families — have a compoundingly negative effect on long context extensibility\. Any one of these choices alone has a minor impact on long context performance, but combining three or more can drop the performance downstream by up to 47%\. Furthermore, these differences are not detectable from short\-context loss or validation datasets\. We show that much of the variation in long context ability across model families is driven by these architectural features and detectable from applying context extension early in pretraining\. We demonstrate this with controlled ablations that hold data, tokenizer, and extension recipe fixed while varying normalization, GQA, pretraining context length, and sliding window attention\. After over 170,000 GPU hours of training, we release the resulting set of models as OlmPool, a set of 26 comparable 7B models with checkpoints before and after long\-context extension\. This pool includes several architectures that outperform the Llama 3 architecture on long context extensibility\. In an analysis of our ablation models, we identify patterns in attention sink behavior and attention distributions across context that are attributable to specific architectural differences\.
## 1Introduction
Pretraining large language models is an expensive and time\-consuming process\. From architectural choices to data selection, many design choices carry the potential to shift the behavior of the resulting downstream model — and these factors can interact in complex ways\. Because running experiments at full scale is prohibitively expensive, a critical question is how to validate design decisions early in training or at smaller scale\. This challenge is especially acute for capabilities that are elicited later in the development cycle — such as reasoning\(Wanget al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib52)\)or agentic behavior\(Qinet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib51)\)— since architectural choices must be made long before these capabilities can be directly observed\.
Long\-context processing is an important instance of this problem\. Context length is typically extended by modifying positional embeddings and continuing to pretrain at longer context lengths during a midtraining phase at the end of pretraining\(Xionget al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib50)\)\. Because this phase comes late in the development cycle, practitioners must commit to architectural decisions before they can observe how those decisions affect long\-context behavior\. Compounding this, most long\-context extension recipes are developed on a small set of base models: the majority of works focus on extending Llama family models \(Fuet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib43)\); Gaoet al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib48)\); Luet al\.\([2024b](https://arxiv.org/html/2608.10296#bib.bib49)\); Chenet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib47)\); Penget al\.\([2026](https://arxiv.org/html/2608.10296#bib.bib8)\),inter alia\), with relatively few considering other architectures \(e\.g\.Dinget al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib45)\); Zhaoet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib46)\); Huet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib44)\)\)\. As a result, it is unclear how broadly existing recipes transfer, or whether the base architecture itself is a decisive factor in downstream long\-context performance, even when comparing only transformer models\.111In parallel, recent work on architectures outside of the standard transformer has been motivated by improving long context performance \(e\.gGu and Dao \([2024](https://arxiv.org/html/2608.10296#bib.bib40)\); Yanget al\.\([2025c](https://arxiv.org/html/2608.10296#bib.bib38)\); Penget al\.\([2023](https://arxiv.org/html/2608.10296#bib.bib39)\)\); we focus here on variationswithinthe transformer family\.
0202040406060OLQΔ\\Delta26\.5Models sorted by HELMET 32KHELMET 32K0 features1 feature2 features3 features4 featuresFigure 1:HELMET 32K scores across all OlmPool models with identical data and context extension strategy, sorted worst to best\. The colors indicate the count of \(individually minor\) choices made that downweight long context performance; the combination of these features can dramatically degrade performance, up to 26\.5 points on HELMET\. \(Q\), \(O\), and \(L\) indicate the Qwen 3, Olmo 3, and Llama 3 architectures; none of these is optimal\.In this work, we demonstrate that a small set of cross\-model\-family architectural variations account for substantial variation in downstream long context performance by performing a set of data\- and optimization\-controlled pretraining experiments\. We first show that short context performance is not sufficient to predict long context performance by training a sweep of models with the same short context behavior but dramatically variable long context behavior; then, we study how normalization decisions, the use of GQA, sliding window attention, and the pretraining context length shift long context performance\. We select only values for these four factors that have been used in Llama 2, Llama 3, Qwen 3, or Olmo 3, and pretrain a sweep of 26 models, OlmPool, which represent varying choices for each factor\. We show that even interpolating in this narrow design space can result in dramatically divergent long context performance \(e\.g\. in Figure[1](https://arxiv.org/html/2608.10296#S1.F1)\)\. This presents a tradeoff: each of these design decisions has desirable properties for training efficiency, inference cost, or stability, while also negatively impacting long context extensibility\. We identify minimal pairs of architectural changes that substantially damage downstream long context performance, characterize the compounding impact of applying multiple of these architectural changes at once, and analyze the attention patterns of these model pairs\.
## 2Setting
We perform a set of controlled pretraining experiments to construct OlmPool,222Anolmis a type of cave\-dwelling aquatic salamander\.a set of 26 comparable models in the 7\-8B parameter range\. Appendix[A](https://arxiv.org/html/2608.10296#A1)provides full descriptions\.
### 2\.1Architectural choices ablated
We consider four primary architectural design decisions: normalization strategy, grouped\-query attention, the use of sliding windows, and pretraining context length\. These features were selected because they differ across several major recent model releases and have explicit connections to the attention mechanism or context length\.
#### Normalization
The two normalization factors we consider are layernorm ordering and the presence of QK norm\. The observation that QK norm can limit long context performance was first made byYanget al\.\([2025b](https://arxiv.org/html/2608.10296#bib.bib54)\); we replicate this finding in our setting and further explore how specific variants of QK norm and norm order play a role\.
QK norm is generally added to improve training stability \(e\.g\. inChameleon Team \([2025](https://arxiv.org/html/2608.10296#bib.bib11)\); Marin Community \([2025](https://arxiv.org/html/2608.10296#bib.bib10)\)\) and is often implemented layerwise, as a normalization step applied before the concatenated query matrix is split for specific attention heads\. This may be implemented as an RMS norm:
Q^=QRMS\(Q\)γQ,K^=KRMS\(K\)γK\\hat\{Q\}=\\frac\{Q\}\{\\text\{RMS\}\(Q\)\}\\,\\gamma^\{Q\},\\qquad\\hat\{K\}=\\frac\{K\}\{\\text\{RMS\}\(K\)\}\\,\\gamma^\{K\}\(1\)
WhereγQ,γK\\gamma^\{Q\},\\gamma^\{K\}are normalization parameters learned for each layer\. Thislayerwiseimplementation of QK norm is used in Olmo 2 and 3\(OLMoet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib36); Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\)\. We also consider the headwise variant of QK norm, used by Qwen 3\(Yanget al\.,[2025a](https://arxiv.org/html/2608.10296#bib.bib24)\), Gemma 3\(Gemma Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib23)\), and Marin 32B\(Marin Community,[2025](https://arxiv.org/html/2608.10296#bib.bib10)\)\. Headwise QK norm applies normalization separately to each attention head’s queries and keys, learning per\-head valuesγhQ,γhK\\gamma\_\{h\}^\{Q\},\\gamma\_\{h\}^\{K\}forh=1,…,Hh=1,\\dots,H:
Q^h=QhRMS\(Qh\)γhQ,K^h=KhRMS\(Kh\)γhK\\hat\{Q\}\_\{h\}=\\frac\{Q\_\{h\}\}\{\\mathrm\{RMS\}\(Q\_\{h\}\)\}\\,\\gamma^\{Q\}\_\{h\},\\qquad\\hat\{K\}\_\{h\}=\\frac\{K\_\{h\}\}\{\\mathrm\{RMS\}\(K\_\{h\}\)\}\\,\\gamma^\{K\}\_\{h\}\(2\)
A related question is whether to place the layernorm before or after the sublayer \(i\.e\. prenorm or post\-sublayer\-norm333Closely related to perinorm, which normalizes a sublayer’s inputsandoutputs\(Kimet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib21)\)\.\)\. Applying post\-sublayer\-norm without QK norm can lead to training divergence because models that apply normalization after the sublayer are more sensitive to gradient instability caused by large or high\-variance attention logits\.444We confirm this experimentally as well; in runs with post\-sublayer\-norm and no QK norm, training consistently diverged before 140B tokens\.
#### Grouped query attention \(GQA\)\.
GQA\(Ainslieet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib25)\)increases inference efficiency by reusing the same key\-value matrices for multiple query heads in the same layer, reducing the size of the key\-value cache\. Typical GQA models share 8 key\-value heads for 32 query heads\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib68); Yanget al\.,[2025a](https://arxiv.org/html/2608.10296#bib.bib24)\)\. Note that because GQA shares someWKW\_\{K\}andWVW\_\{V\}parameters, it reduces the total capacity of the network; when this occurs, we adjust the intermediate size slightly to result in the same total parameter count\. This should benefit models with GQA in our comparisons\.
#### Sliding window attention \(SWA\)\.
Modern sliding window attention implementations generally intersperse layers of local window attention with layers of full attention in a many\-to\-one pattern\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67); Gemma Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib23)\), which reduces the KV cache size at inference time\. We use the Olmo 3 configuration, which has 3 local attention layers of 4096 context for every 1 global attention layer\.
#### Pretrained context length\.
Zhaoet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib46)\)observe that long context ability is impacted by the pretraining context length, with models trained at a longer context length better able to adapt to context extension\. Conversely, shorter context lengths generally allow for faster training throughput\. We pretrain Olmo and Llama variants at 4096 context length, matching the context length of the prior generation for each model\. This allows us to compare models trained at 4K directly with models that use a 4K sliding window on 8K context\.
### 2\.2Training
We pretrain each model for 140B tokens \(the Chinchilla\-optimal amount of data\(Hoffmannet al\.,[2022b](https://arxiv.org/html/2608.10296#bib.bib65)\)\)\. Then, we adjust the RoPE theta for context extension\(Xionget al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib66)\)and continue pretraining for 10B tokens on 64K context data from the Longmino mix\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\), annealing the learning rate to 0\. We hold the data selection and ordering, the learning rate and schedule, and the tokenizer constant across all models in OlmPool\. We reuse the same optimization learning rate and schedule as Olmo 3\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\)and further discuss the rationale for standardizing optimization in Appendix[C](https://arxiv.org/html/2608.10296#A3)\.
Ideally, we would also standardize initialization\. However, because not all models have the same parameters \(e\.g\. if GQA or QK norm are added\), we cannot exactly synchronize the initializations\. Where possible, we reuse the same initalization across runs; if there is a minor difference in parameterization, we use the remainder of the same initialization and only re\-initialize the new parameters\. Appendix[A](https://arxiv.org/html/2608.10296#A1)identifies the initialization for each model, and we further discuss the impact of model initialization in Appendix[B](https://arxiv.org/html/2608.10296#A2)\.
### 2\.3Evaluation
We evaluate all models downstream on 3 popular measures of long context performance: RULER\(Hsiehet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib73)\), which represents a sweep of synthetic Needle\-in\-a\-Haystack \(NIAH\) style tasks of increasing complexity; HELMET\(Yenet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib71)\), which additionally considers in\-context learning, reranking, and question\-answering tasks; and LongPPL\(Fanget al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib76)\), a variant of perplexity that only considers tokens that require long\-range dependencies to predict\. Because we are working with weak models early in pretraining, we consider only the subtasks from HELMET that do not require generating long spans of text to evaluate with an LM judge\. We observe that all three measures correlate closely; for readability, we primarily report HELMET at 32K in the main text, unless trends differ across the three measures\. To compare model architectures, we construct multiple paired comparisons, holding constant all other factors, wherever possible\. We also fit linear regressions from architectural features, short\-context metrics, and observed attention distributions to measure the predictive power of each feature across OlmPool\.
## 3Short context metrics are not predictive of long context performance
0\.670\.670\.670\.670\.680\.680\.680\.680\.690\.690\.690\.690\.70\.70\.70\.7303035354040454550505555Short context evals \(BPB\) \(↓\\downarrow\)HELMET 32K101020203030404050503030404050506060Pre\-extension HELMET 8K \(↑\\uparrow\)HELMET 32K \(↑\\uparrow\)
Figure 2:Benchmark scores pre\-extension largely fail to predict long\-context benchmark scores post\-extension, even when evaluating on the shorter\-context version of the same benchmark\. Each point is a model from OlmPool\.The models in OlmPool demonstrate that seemingly small changes in the model recipe have a dramatic downstream effect on long context extensibility\. The performance of these models ranges from 29\.9 to 56\.4 on HELMET at 32K, or 44\.7 to 67\.7 on RULER at 32K\.
Can this downstream effect be predicted from pretraining metrics? We observe that short context metrics are surprisingly poor predictors of long\-context behavior\. We consider a sweep of measures to try to predict long context performance downstream, demonstrating that standard pretraining metrics are insufficient to predict long context performance\.
#### Intrinsic metrics\.
We measure the loss for each training run at the end of pretraining and at the end of long context extension and compute the correlation between these values and the resulting HELMET score\. Training loss correlates only weakly with downstream long context performance \(R2=0\.29\)R^\{2\}=0\.29\); surprisingly, the loss during pretraining ismore predictiveof downstream score than the loss during context extension \(R2=0\.06R^\{2\}=0\.06\)\. We measure perplexity of the pre\-context\-extension model over 11 held\-out text samples from differing domains\(Magnussonet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib75)\); these scores range from completely uncorrelated to weakly correlated with downstream long context scores\. The best\-correlated splits do not align with common understanding of the types of data that require long\-range dependencies: a sample of data from WikiText\(Merityet al\.,[2016](https://arxiv.org/html/2608.10296#bib.bib9)\)is the most correlated \(R2=0\.39R^\{2\}=0\.39\), with perplexity of samples from academic texts and code showing little to no correlation with long context behavior downstream \(0\.01≤R2≤0\.200\.01\\leq R^\{2\}\\leq 0\.20\)\.555These trends hold even if we measure correlation between perplexity and LongPPL, which is a more intrinsic metric of long context quality\. See a full per\-dataset breakdown in Appendix[D\.3](https://arxiv.org/html/2608.10296#A4.SS3)\.
#### Downstream evaluations\.
During pretraining, models are often periodically evaluated on a set of development benchmarks\. We evaluate on a set of 16 in\-loop benchmarks, detailed in Appendix[D\.3](https://arxiv.org/html/2608.10296#A4.SS3); because early pretraining checkpoints are often inconsistent at answering in multiple\-choice format\(Bunnet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib42)\), we score by the bits\-per\-byte on the correct answer for each benchmark question instead of accuracy\(Heinemanet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib83)\)\. Figure[2](https://arxiv.org/html/2608.10296#S3.F2)\(left\) shows the average of these evaluations graphed against downstream HELMET score; there is a slight correlation between the two \(R2=0\.17R^\{2\}=0\.17\), but this fails to exceed the predictiveness of the best perplexity measures\. Note that all scores are very close together: in our setting, where models are trained with the same data and very similar architectures, little difference is observable between models pre\-context\-extension\. Clearly, standard in\-loop evaluations are not sufficient to provide signal for downstream long context extensibility\.
Is this merely an issue of choosing the wrong benchmarks? We consider a benchmark more directly related to our evaluations for performance downstream: the shortest context split of HELMET, which, at 8K, is possible for most of our models to process without context extension\. Recall that these models are being evaluated pre\-anneal and early in training, so we expect scores to be quite poor\. Figure[2](https://arxiv.org/html/2608.10296#S3.F2)\(right\) graphs these pre\-extension scores against the post\-extension performance; while HELMET scores pre\-extension are both more variable and slightly more predictive \(R2=0\.32R^\{2\}=0\.32\) than other short\-context evaluations, they still fail to predict double\-digit swings in HELMET performance downstream\.
Clearly, the differing performance of these models is not solely attributable to how well each model has fit the pretraining corpus\. We consider the individual effect of each of the four factors we have identified in turn and show that their behavior is best modeled as an additive impairment to long context: the individual features selected do not matter nearly as much as the number of long\-context\-inhibiting features present\. In Appendix[B](https://arxiv.org/html/2608.10296#A2), we also discuss the impact of additional features that we considered: pretraining with linear layers quantized to float8 and changing the pretraining run’s random initialization\.
#### All four features reduce long context capability downstream\.
In paired comparison runs, each architectural feature results in some degradation of long context performance downstream\. Pretraining at 4096 instead of 8192 context length and training with sliding window layers result in modest average degradations of 1\-2 points on HELMET at 32K\. Figure[3](https://arxiv.org/html/2608.10296#S4.F3)shows that pretraining with GQA configurations that use increasingly aggressive degrees of query sharing \(i\.e\. decreasing the number of KV heads\) degrades performance from the Llama 3 configuration, and pretraining withmoreKV heads than the Llama 3 architecture improves performance\. But the largest individual performance impact by far arises from normalization choice\. On the Olmo architecture, changing Olmo 3’s choice of QK norm and post\-sublayer\-norm to prenorm results in a 6 point gain on HELMET; conversely, adding these features to Llama 3’s architecture results in a 3\.8 point drop\.666We find that the choice of headwise versus layerwise QK norm and the ordering of prenorm vs post\-sublayer\-norm have some effect as well, although the decision to apply QK norm at all dominates; for more details, see Appendix[B](https://arxiv.org/html/2608.10296#A2)\.
481632484850505252545456565858\# KV HeadsHELMET 32K \(↑\\uparrow\)48163260606262646466666868\# KV HeadsRULER 32K \(↑\\uparrow\)4816323\.43\.43\.53\.53\.63\.6\# KV HeadsLongPPL \(↓\\downarrow\)Figure 3:GQA is harmful to long context performance in our setting, even when adjusting for comparable parameter count\. Variants of Llama that increase the number of KV heads improve performance on long context benchmarks\. 8 KV heads is the Llama GQA configuration, and 32 KV heads indicates no GQA\. Note that the model with 16 KV heads is slightly larger than the other models in this comparison \(8\.1B vs 7\.7\-7\.8B\)\.
#### Minor effects of architecture compound\.
With the exception of QK norm, most individual architectural features have relatively minor effects when they are the only potentially detrimental feature tested\. However, when paired together, these features can have a much more significant impact\. For instance, adding a sliding window has a minor negative effect on a model without GQA \(\-1\.1 points from the full attention configuration\)\. However, in multiple comparison runs, adding sliding window layers to a model with GQA results in a dramatic drop in performance: \-9 points on average\. The worst\-scoring runs in OlmPool combine two or more features that limit the expressivity of attention — e\.g\. the single worst configuration combines GQA, sliding windows, and headwise QK norm, for a total effect far worse than the sum of these individual impacts\.
We find that downstream long context performance can be well\-estimated by simply counting how many of these four architectural choices are present in a OlmPool model — this single numerical feature is the single most predictive feature for downstream long context performance \(in\-sampleR2=0\.67R^\{2\}=0\.67, LOOR2=0\.61R^\{2\}=0\.61\), outperforming even a linear regression over the four individual axes of architectural variation\. It is not any single feature which results in catastrophically poor long context extensibility, but the combination of several features that each reduce the expressivity of attention\.
Figure[1](https://arxiv.org/html/2608.10296#S1.F1)also shows that combining all four features is not necessarily worse than combining three features in our setting\. We hypothesize that this may be because at least some of these features impact long context abilities in overlapping ways: for instance, sliding window attention restricts the model to only 4K context at some layers during pretraining, and pretraining at 4K restricts the model to only 4K context atalllayers during pretraining\. However, the predictor that counts the presence of any of the four features remains more predictive than any predictor that collapses two of the features into the same category\.
#### Llama 3 is a particularly good architecture for long context\.
While it was previously unclear whether the ease of extensibility for Llama 3 was due to architectural or data factors \(since the pretraining data for Llama 3 was not publicly disclosed\), we find evidence that this is primarily an architectural phenomenon\. The Llama 3 architecture model is one of the best models in the design space\. This suggests that context extension recipes developed with this model may require additional effort to apply to other architectures — for instance, our results show that same context extension recipe applied to the Llama 3, Qwen 3, and Olmo 3 architectures is far more effective for the Llama\-like model, even when the other models were pretrained for the same duration and data as the Llama\-like model\. This also helps explain our empirical observation that Olmo 3 Base is more challenging to context extend than Llama 3 Base\.
## 5Analysis
### 5\.1Is this measuring token efficiency or a fundamental capability gap?
Our main experiments compare models after the same 10B token context extension\. However, it’s possible that extending with a much longer context extension phase would wash out these differences\. To test this, we choose three representative models at different quality points in OlmPool: the models with the architectures of Llama 3, Olmo 3, and the worst\-scoring architecture on HELMET downstream\. On this trio of models, we conduct 1B and 50B context extensions and compare performance with increasing amounts of data\.777Note that because we perform each context extension run as an anneal to a learning rate of 0, the 1B, 10B, and 50B extensions are separate runs, not points sampled from one long extension\.Figure[5](https://arxiv.org/html/2608.10296#S5.F5)shows the performance of these models with increasing amounts of data\. As expected, all models see better long context performance with 50B token extensions; however, this effect is not enough to compensate for architecture\-driven differences in base quality\. Even after 50B tokens, the worst architecture does not reach the same performance as the Llama architecture achieves after 1B tokens, and the differences between the three architectures remains relatively stable\. While it’s possible that context extensions at a more extreme scale could overpower these architectural effects, we see no evidence of this occurring up to the 50B token scale, where the context extension phrase represents 26% of the total tokens seen by the model\.
1B10B50B0202040406060Context extension tokensHELMET scoreWorstOlmoLlamaFigure 4:Performance on HELMET after 1B, 10B, or 50B token extension for three representative runs \(the worst architecture, the Olmo 3 architecture, and the Llama 3 architecture\)\. Longer context extension fails to wash out architectural differences\.70B140B280B2T0202040406060Pretraining tokensHELMET scoreOlmo\-likeWorstFigure 5:HELMET score at extensions performed after progressively more pretraining tokens for two training runs\. The difference in long context behavior is consistent, especially from 140B onwards\.
### 5\.2Do these differences persist in longer pretraining runs?
To be able to measure many architectural configurations, we extend models after only 140B tokens of pretraining\. However, most modern models are trained far longer, with long context extension often applied after the model has seen trillions of tokens\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\)\. Would these effects wash out with a longer pretraining run?
We measure the difference between two configurations of interest on much longer pretraining runs\. At 70B, 140B, 280B, and 2T tokens into a longer pretraining run, we adjust the RoPE theta and perform a 10B anneal on long context \(64K\) data\. Figure[5](https://arxiv.org/html/2608.10296#S5.F5)shows the relative behavior of these models on downstream long context evaluations at each point\. We show that nontrivial long context performance can be recovered from extensions at least as early as 70B, and the relative performance of the two architectures remains fairly consistent over the course of pretraining\. Because we observe that the two models are closer in performance at 70B than at any point 140B or later, we train all other OlmPool models to 140B tokens before extending to ensure we are not missing differences across models\.
### 5\.3Do these differences persist across context extension strategies?
8K16K32K64K0202040406060HELMET scoreOlmoLlamaBestFigure 6:Solid color indicates the YaRN 2\-stage extension performance; shaded is the NTK 1\-stage extension performance\.To understand if the differences we observe are sensitive to context extension strategy, we consider an alternate long context extension recipe\. We perform a two\-stage context extension using YaRN\(Penget al\.,[2026](https://arxiv.org/html/2608.10296#bib.bib8)\)\.We perform continued pretraining at 32K for 5B tokens and then at 64K for 5B tokens, over the course of a combined linear anneal of the learning rate to 0\. We warm up the learning rate for 200 steps at the start of the first phase, and use the same YaRN factor across phases\.
We perform this modified extension on 3 models: the Olmo\-like model, the Llama\-like model, and the best model in OlmPool \(measured by HELMET at 32K after our standard extension\)\. Figure[6](https://arxiv.org/html/2608.10296#S5.F6)compares these extensions with the extension to 64K used in the primary OlmPool analysis\. This strategy uniformly underperforms our other context extensions, likely because of the reduced number of tokens at 64K context\. However, the same relative ranking of architectures is recovered by both methods, and the gap between architectures is larger with thebettercontext extension method\.
### 5\.4Attention behaviors in OlmPool
Armed with a set of models that differ in downstream performance, we now seek to understand the finegrained differences between these models\. We compute a number of statistics of the attention distribution across all models, measured over a set of 100 long documents selected at random from an even split of Project Gutenberg books, government reports\(Huanget al\.,[2021](https://arxiv.org/html/2608.10296#bib.bib27)\), FineWeb PDFs\(Penedoet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib28)\), and legal texts\(Henderson\*et al\.,[2022](https://arxiv.org/html/2608.10296#bib.bib26)\)\. We measure the entropy of the attention distribution at each attention head at positions 1K, 4K, 16K, and 32K\. We also measure the percentage of attention mass in the attention sink \(defined as the first 100 tokens of visible context\) and the local context \(defined as the 100 tokens immediately preceding the current token\)\.
55667788055101015152020Attention Entropy @ 32K% Sink AttentionFull\-Attention Layersno QKNormQKNormheadwise QKNorm4\.54\.5555\.55\.5666\.56\.5055101015152020Attention Entropy @ 32KSWA LayersFigure 7:Attention entropy and the presence of a a strong attention sink both cluster by presence of QK norm\. OlmPool models with sliding window attention have less attention sink behavior on full attention layers; these represent the lower cluster of blue dots\.#### High entropy and attention sinks are positive indicators — and QK norm reduces both\.
Yanget al\.\([2025b](https://arxiv.org/html/2608.10296#bib.bib54)\)observe that models with QK\-norm have higher entropy in the attention distribution \(i\.e\. less peaky attention\), which they suggest makes it more challenging for the model to attend over longer contexts\. Attention entropy is also influenced by the presence of attention sinks\(Xiaoet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib29)\): positions early in the context window that consistently receive substantial attention mass\. The presence of these sinks is theorized to be a result of models attempting to “discard” excess attention weight\(Bondarenkoet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib33); Qiuet al\.,[2026](https://arxiv.org/html/2608.10296#bib.bib32); Guet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib7)\)and avoid representation collapse\(Barberoet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib6)\); it is also believed to make models more difficult to quantize\(Yeet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib31)\)\.
Figure[7](https://arxiv.org/html/2608.10296#S5.F7)visualizes both sink attention and attention entropy\. While the presence of attention sinks is generally considered negative \(e\.g\.Qiuet al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib30)\)names reduction of attention sinks as a core benefit of their approach\), the presence of attention sinks here correlates with improved long context performance \(R2=0\.38R^\{2\}=0\.38\)\. In the absence of another mechanism such as differential attention\(Yeet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib31)\)or gating\(Qiuet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib30)\), attention sinks appear to be the default strategy learned by QK\-norm\-less transformers to compensate for excess attention\. Thus, attention sink behavior corresponds with long context abilities in OlmPool\.
#### Retrieval heads
Wuet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib41)\)propose the existence ofretrieval heads, specialized attention heads that are primarily responsible for retrieving information from prior context in both short and long context processing\. We measure this using the same set of documents with a needle\-in\-a\-haystack \(NIAH\) task injected\. We compute, at the end of prefill, the percentage of attention that is placed on the needle tokens\. Then, we generate a completion for each example and look at examples where the model successfully generates the needle text\. We compute per\-head retrieval scores by measuring how much each individual attention head attended to the needle tokens in the input over the course of generation\.
At prefill time, we observe no more than a slight correlation between the ability to place more attention on the needle tokens and performance on downstream evaluations\. Models with QK norm \(headwiseorlayerwise\) uniformly place less attention on the needle tokens than models without QK norm\. However, during generation we observe little difference in retrieval head behavior across models in OlmPool, with very low retrieval head scores across models\. These models may be too weak to reliably identify retrieval heads\. Alternatively, some other mechanism may govern long context abilities in models early in training\.
## 6Related work
#### Architectural choices and long\-context extensibility\.
Prior works in this area primarily focus on positional embeddings or modifications to the extension method itself\. The closest work to our setting,Yanget al\.\([2025b](https://arxiv.org/html/2608.10296#bib.bib54)\), compares RoPE, NoPE, and QK\-normalized RoPE models and shows that these three variants can differ substantially on long\-context evaluation despite similar standard performance\. More broadly, work on positional design and extrapolation shows that even a single architectural axis can have a large effect on longer\-range behavior:Kazemnejadet al\.\([2023](https://arxiv.org/html/2608.10296#bib.bib55)\),Presset al\.\([2022](https://arxiv.org/html/2608.10296#bib.bib56)\), andWanget al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib60)\)find substantial differences across positional embedding variants, andLuet al\.\([2024a](https://arxiv.org/html/2608.10296#bib.bib64)\)find that sparse attention generally lags full attention for long context\. Other work studies how to improve extrapolation by changing attention or positional embedding settings during context extension\(Sunet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib58); Chenet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib57)\)\. A separate line of long\-context work changes the architecture, for example by modifying the attention mechanism\(Zimerman and Wolf,[2023](https://arxiv.org/html/2608.10296#bib.bib59)\), through recurrence and memory\(Daiet al\.,[2019](https://arxiv.org/html/2608.10296#bib.bib61)\), or newer architectures designed for effectively unbounded context\(Maet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib62)\); seeHuanget al\.\([2023](https://arxiv.org/html/2608.10296#bib.bib63)\)for a survey of this space\. In contrast, we focus on choices made before pretraining that mayunintentionallydetermine downstream long\-context extension outcomes\.
#### Performance prediction from smaller scales\.
A separate methodological literature asks how to make expensive model\-development decisions using smaller or earlier experiments\. Controlled pretraining suites such as Pythia\(Bidermanet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib69)\), open\-science efforts such as OLMo and LLM360\(Groeneveldet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib78); Liuet al\.,[2023](https://arxiv.org/html/2608.10296#bib.bib79)\), and small\-experiment prediction work such as DataDecide\(Magnussonet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib70)\)show that many choices can be studied scientifically before full\-scale deployment; scaling\-law work shows that even expensive allocation decisions can be forecast from smaller runs\(Hoffmannet al\.,[2022a](https://arxiv.org/html/2608.10296#bib.bib82)\)\. Some recent work has also looked at integrating features of the data distribution or architecture into scaling law predictions for short\-context tasks\(Liuet al\.,[2026](https://arxiv.org/html/2608.10296#bib.bib22)\)\.Magnussonet al\.\([2024](https://arxiv.org/html/2608.10296#bib.bib75)\); Fanget al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib76)\)argue that easy scalar proxies such as perplexity can miss important behavior during training runs, while recent work on evaluation reliability shows that the signal\-to\-noise properties of benchmarks themselves can strongly affect how useful small experiments are for model\-development decisions\(Heinemanet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib83)\)\. More general work on capability prediction emphasizes that downstream behavior is often harder to forecast from pretraining signals alone\(Schaefferet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib77)\)\. Our setting is far cheaper than full\-scale long context training, but more expensive than measuring pretraining signals alone\. We also hope that OlmPool can serve as a set of models on which to evaluate future performance prediction metrics\.
## 7Conclusion
We demonstrate that a series of small, individually reasonable architectural perturbations, well grounded in the literature, can result in dramatically reduced long context capabilities\. We show that this degradation is difficult to detect in short context metrics but detectable from context extension runs very early into pretraining\. This suggests two interesting directions for future research: evaluating the minimal size or token budget at which these effects can be reliably measured, to further reduce the cost of architectural experimentation; and devising better proxy metrics for short context models to estimate long context performance without performing a context extension\. We also believe there is more to understand mechanistically about the differences between OlmPool models\. Finally, OlmPool’s parallel runs, traversing a similar optimization problem with slightly different architectures, may be useful for research into other phenomena in early pretraining\. To this end, we release 38 checkpoints for each model, representing the full pretraining and long context extension\.
Each feature we ablate has some clear benefit— stability for normalization, pretraining efficiency for context length, and inference efficiency for sliding window and GQA— and individually may be justified because of these other factors\. Yet the combination of these features results in unacceptable long context extensibility\. By exposing the interplay between these factors in a controlled setting, we hope to enable model developers to make more informed choices about their architecture design and to spur future research into alternatives that better navigate these tradeoffs\.
## Ethics Statement
Pretraining ablations carry a heavy computational cost\. We estimate that the cost of training the 26 models for this paper was approximately 170,000 H100 hours, in addition to the costs of initial experimentation, evaluation, and additional ablation runs\. Using the calculations fromMorrisonet al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib37)\), this is equivalent to approximately 42\.63 metric tons of CO2 emitted, if our hardware was of equivalent power usage\. In releasing all artifacts including intermediate checkpoints, we hope that this cost can be amortized by the reuse of OlmPool\.
## Acknowledgements
We thank Sewon Min, Adithya Pratapa, Prasann Singhal, Jacob Morrison, and the anonymous reviewers for helpful feedback on this work, and Taira Anderson, Kyle Wiggers, and Bailey Kuehl for release assistance\.
This material is based upon work supported by the National Science Foundation under Award No\. 2413244\. AB was supported by a grant from the National Science Foundation Graduate Research Fellowship Program under Grant No\. DGE2140739\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author\(s\) and do not necessarily reflect the views of the sponsors\.
## References
- GQA: training generalized multi\-query transformer models from multi\-head checkpoints\.External Links:2305\.13245,[Link](https://arxiv.org/abs/2305.13245)Cited by:[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px2.p1.2)\.
- J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. Sutton \(2021\)Program synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- F\. Barbero, Á\. Arroyo, X\. Gu, C\. Perivolaropoulos, M\. Bronstein, P\. Veličković, and R\. Pascanu \(2025\)Why do llms attend to the first token?\.External Links:2504\.02732,[Link](https://arxiv.org/abs/2504.02732)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1)\.
- S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. van der Wal \(2023\)Pythia: a suite for analyzing large language models across training and scaling\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 2397–2430\.External Links:[Link](https://proceedings.mlr.press/v202/biderman23a.html)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- Y\. Bondarenko, M\. Nagel, and T\. Blankevoort \(2023\)Quantizable transformers: removing outliers by helping attention heads do nothing\.External Links:2306\.12929,[Link](https://arxiv.org/abs/2306.12929)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1)\.
- A\. Bunn, S\. Wiegreffe, and B\. Bogin \(2025\)Fine\-tune on the format: first improving multiple\-choice evaluation for intermediate LLM checkpoints\.InProceedings of the Fourth Workshop on Generation, Evaluation and Metrics \(GEM²\),O\. Arviv, M\. Clinciu, K\. Dhole, R\. Dror, S\. Gehrmann, E\. Habba, I\. Itzhak, S\. Mille, Y\. Perlitz, E\. Santus, J\. Sedoc, M\. Shmueli Scheuer, G\. Stanovsky, and O\. Tafjord \(Eds\.\),Vienna, Austria and virtual meeting,pp\. 511–521\.External Links:[Link](https://aclanthology.org/2025.gem-1.46/),ISBN 979\-8\-89176\-261\-9Cited by:[§3](https://arxiv.org/html/2608.10296#S3.SS0.SSS0.Px2.p1.1)\.
- Chameleon Team \(2025\)Chameleon: mixed\-modal early\-fusion foundation models\.External Links:2405\.09818,[Link](https://arxiv.org/abs/2405.09818)Cited by:[Appendix B](https://arxiv.org/html/2608.10296#A2.SS0.SSS0.Px1.p1.1),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p2.1)\.
- M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. Zaremba \(2021\)Evaluating large language models trained on code\.External Links:2107\.03374,[Link](https://arxiv.org/abs/2107.03374)Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- S\. Chen, S\. Wong, L\. Chen, and Y\. Tian \(2023\)Extending context window of large language models via positional interpolation\.External Links:2306\.15595,[Link](https://arxiv.org/abs/2306.15595)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- Y\. Chen, S\. Qian, H\. Tang, X\. Lai, Z\. Liu, S\. Han, and J\. Jia \(2024\)LongLoRA: efficient fine\-tuning of long\-context large language models\.External Links:2309\.12307,[Link](https://arxiv.org/abs/2309.12307)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord \(2018\)Think you have solved question answering? Try ARC, the AI2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- Z\. Dai, Z\. Yang, Y\. Yang, J\. Carbonell, Q\. V\. Le, and R\. Salakhutdinov \(2019\)Transformer\-XL: attentive language models beyond a fixed\-length context\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 2978–2988\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1285),[Link](https://aclanthology.org/P19-1285/)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- Y\. Ding, L\. L\. Zhang, C\. Zhang, Y\. Xu, N\. Shang, J\. Xu, F\. Yang, and M\. Yang \(2024\)LongRoPE: extending llm context window beyond 2 million tokens\.External Links:2402\.13753,[Link](https://arxiv.org/abs/2402.13753)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- L\. Fang, Y\. Wang, Z\. Liu, C\. Zhang, S\. Jegelka, J\. Gao, B\. Ding, and Y\. Wang \(2025\)What is wrong with perplexity for long\-context language modeling?\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=fL4qWkSmtM)Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p2.1),[§2\.3](https://arxiv.org/html/2608.10296#S2.SS3.p1.1),[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- Y\. Fu, R\. Panda, X\. Niu, X\. Yue, H\. Hajishirzi, Y\. Kim, and H\. Peng \(2024\)Data engineering for scaling language models to 128k context\.External Links:2402\.10171,[Link](https://arxiv.org/abs/2402.10171)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- L\. Gao, S\. Biderman, S\. Black, L\. Golding, T\. Hoppe, C\. Foster, J\. Phang, H\. He, A\. Thite, N\. Nabeshima, S\. Presser, and C\. Leahy \(2020\)The Pile: an 800GB dataset of diverse text for language modeling\.arXiv preprint arXiv:2101\.00027\.Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1)\.
- T\. Gao, A\. Wettig, H\. Yen, and D\. Chen \(2025\)How to train long\-context language models \(effectively\)\.External Links:2410\.02660,[Link](https://arxiv.org/abs/2410.02660)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- Gemma Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p4.3),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px3.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px2.p1.2)\.
- S\. Greenbaum and G\. Nelson \(1996\)The International Corpus of English \(ICE\) project\.World Englishes15\(1\),pp\. 3–15\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-971X.1996.tb00088.x)Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1)\.
- D\. Groeneveld, I\. Beltagy, E\. Walsh, A\. Bhagia, R\. Kinney, O\. Tafjord, A\. Jha, H\. Ivison, I\. Magnusson, Y\. Wang, S\. Arora, D\. Atkinson, R\. Authur, K\. Chandu, A\. Cohan, J\. Dumas, Y\. Elazar, Y\. Gu, J\. Hessel, T\. Khot, W\. Merrill, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, V\. Pyatkin, A\. Ravichander, D\. Schwenk, S\. Shah, W\. Smith, E\. Strubell, N\. Subramani, M\. Wortsman, P\. Dasigi, N\. Lambert, K\. Richardson, L\. Zettlemoyer, J\. Dodge, K\. Lo, L\. Soldaini, N\. A\. Smith, and H\. Hajishirzi \(2024\)OLMo: accelerating the science of language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15789–15809\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.841),[Link](https://aclanthology.org/2024.acl-long.841/)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- A\. Gu and T\. Dao \(2024\)Mamba: linear\-time sequence modeling with selective state spaces\.External Links:2312\.00752,[Link](https://arxiv.org/abs/2312.00752)Cited by:[footnote 1](https://arxiv.org/html/2608.10296#footnote1)\.
- X\. Gu, T\. Pang, C\. Du, Q\. Liu, F\. Zhang, C\. Du, Y\. Wang, and M\. Lin \(2025\)When attention sink emerges in language models: an empirical view\.External Links:2410\.10781,[Link](https://arxiv.org/abs/2410.10781)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1)\.
- D\. Heineman, V\. Hofmann, I\. Magnusson, Y\. Gu, N\. A\. Smith, H\. Hajishirzi, K\. Lo, and J\. Dodge \(2025\)Signal and noise: a framework for reducing uncertainty in language model evaluation\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=sAFottNlra)Cited by:[§3](https://arxiv.org/html/2608.10296#S3.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- P\. Henderson\*, M\. S\. Krass\*, L\. Zheng, N\. Guha, C\. D\. Manning, D\. Jurafsky, and D\. E\. Ho \(2022\)Pile of law: learning responsible data filtering from the law and a 256gb open\-source legal dataset\.arXiv\.External Links:[Link](https://arxiv.org/abs/2207.00220)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.Proceedings of the International Conference on Learning Representations \(ICLR\)\.Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. Sifre \(2022a\)Training compute\-optimal large language models\.External Links:2203\.15556,[Link](https://arxiv.org/abs/2203.15556)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. Sifre \(2022b\)Training compute\-optimal large language models\.External Links:2203\.15556,[Link](https://arxiv.org/abs/2203.15556)Cited by:[§2\.2](https://arxiv.org/html/2608.10296#S2.SS2.p1.1)\.
- C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, and B\. Ginsburg \(2024\)RULER: what’s the real context size of your long\-context language models?\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=kIoBbc76Sy)Cited by:[§2\.3](https://arxiv.org/html/2608.10296#S2.SS3.p1.1)\.
- Z\. Hu, Y\. Liu, J\. Zhao, S\. Wang, Y\. Wang, W\. Shen, Q\. Gu, A\. T\. Luu, S\. Ng, Z\. Jiang, and B\. Hooi \(2024\)LongRecipe: recipe for efficient long context generalization in large language models\.External Links:2409\.00509,[Link](https://arxiv.org/abs/2409.00509)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- L\. Huang, S\. Cao, N\. Parulian, H\. Ji, and L\. Wang \(2021\)Efficient attentions for long document summarization\.External Links:2104\.02112,[Link](https://arxiv.org/abs/2104.02112)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.p1.1)\.
- Y\. Huang, J\. Xu, Z\. Jiang, J\. Lai, Z\. Li, Y\. Yao, T\. Chen, L\. Yang, Zhou Xin, and X\. Ma \(2023\)Advancing transformer architecture in long\-context large language models: a comprehensive survey\.External Links:2311\.12351,[Link](https://arxiv.org/abs/2311.12351)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- A\. Kazemnejad, I\. Padhi, K\. N\. Ramamurthy, P\. Das, and S\. Reddy \(2023\)The impact of positional encoding on length generalization in transformers\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://openreview.net/forum?id=Drrl2gcjzl)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- J\. Kim, B\. Lee, C\. Park, Y\. Oh, B\. Kim, T\. Yoo, S\. Shin, D\. Han, J\. Shin, and K\. M\. Yoo \(2025\)Peri\-ln: revisiting normalization layer in the transformer architecture\.External Links:2502\.02732,[Link](https://arxiv.org/abs/2502.02732)Cited by:[footnote 3](https://arxiv.org/html/2608.10296#footnote3)\.
- A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo, Y\. Wu, B\. Neyshabur, G\. Gur\-Ari, and V\. Misra \(2022\)Solving quantitative reasoning problems with language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 3843–3857\.Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- E\. Liu, A\. Bertsch, L\. Sutawika, L\. Tjuatja, P\. Fernandes, L\. Marinov, M\. Chen, S\. Singhal, C\. Lawrence, A\. Raghunathan, K\. Gashteovski, and G\. Neubig \(2026\)Not\-just\-scaling laws: towards a better understanding of the downstream impact of language model design decisions\.External Links:2503\.03862,[Link](https://arxiv.org/abs/2503.03862)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- Z\. Liu, A\. Qiao, W\. Neiswanger, H\. Wang, B\. Tan, T\. Tao, J\. Li, Y\. Wang, S\. Sun, O\. Pangarkar, R\. Fan, Y\. Gu, V\. Miller, Y\. Zhuang, G\. He, H\. Li, F\. Koto, L\. Tang, N\. Ranjan, Z\. Shen, X\. Ren, R\. Iriondo, M\. Cun, Z\. Hu, M\. Schulze, P\. Nakov, T\. Baldwin, and E\. Xing \(2023\)LLM360: towards fully transparent open\-source LLMs\.External Links:2312\.06550,[Link](https://arxiv.org/abs/2312.06550)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- Y\. Lu, J\. N\. Yan, S\. Yang, J\. T\. Chiu, S\. Ren, F\. Yuan, W\. Zhao, Z\. Wu, and A\. M\. Rush \(2024a\)A controlled study on long context extension and generalization in LLMs\.External Links:2409\.12181,[Link](https://arxiv.org/abs/2409.12181)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- Y\. Lu, J\. N\. Yan, S\. Yang, J\. T\. Chiu, S\. Ren, F\. Yuan, W\. Zhao, Z\. Wu, and A\. M\. Rush \(2024b\)A controlled study on long context extension and generalization in llms\.External Links:2409\.12181,[Link](https://arxiv.org/abs/2409.12181)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- X\. Ma, X\. Yang, W\. Xiong, B\. Chen, L\. Yu, H\. Zhang, J\. May, L\. Zettlemoyer, O\. Levy, and C\. Zhou \(2024\)Megalodon: efficient LLM pretraining and inference with unlimited context length\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/840abfadd04c967feaa2a49aba94a32d-Abstract-Conference.html)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- I\. Magnusson, A\. Bhagia, V\. Hofmann, L\. Soldaini, A\. H\. Jha, O\. Tafjord, D\. Schwenk, E\. P\. Walsh, Y\. Elazar, K\. Lo, D\. Groeneveld, I\. Beltagy, H\. Hajishirzi, N\. A\. Smith, K\. Richardson, and J\. Dodge \(2024\)Paloma: a benchmark for evaluating language model fit\.InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/760b2d94398aa61468aa3bc11506d9ea-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1),[§3](https://arxiv.org/html/2608.10296#S3.SS0.SSS0.Px1.p1.4),[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- I\. Magnusson, N\. Tai, B\. Bogin, D\. Heineman, J\. D\. Hwang, L\. Soldaini, A\. Bhagia, J\. Liu, D\. Groeneveld, O\. Tafjord, N\. A\. Smith, P\. W\. Koh, and J\. Dodge \(2025\)DataDecide: how to predict best pretraining data with small experiments\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.External Links:[Link](https://proceedings.mlr.press/v267/magnusson25a.html)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- Marin Community \(2025\)Marin 32b retrospective\.Technical reportMarin\.External Links:[Link](https://marin.readthedocs.io/en/latest/reports/marin-32b-retro/)Cited by:[Appendix B](https://arxiv.org/html/2608.10296#A2.SS0.SSS0.Px1.p1.1),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p2.1),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p4.3)\.
- S\. McCandlish, J\. Kaplan, D\. Amodei, and O\. D\. Team \(2018\)An empirical model of large\-batch training\.External Links:1812\.06162,[Link](https://arxiv.org/abs/1812.06162)Cited by:[Appendix C](https://arxiv.org/html/2608.10296#A3.SS0.SSS0.Px1.p2.1)\.
- S\. Merity, C\. Xiong, J\. Bradbury, and R\. Socher \(2016\)Pointer sentinel mixture models\.External Links:1609\.07843Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1),[§3](https://arxiv.org/html/2608.10296#S3.SS0.SSS0.Px1.p1.4)\.
- W\. Merrill, S\. Arora, D\. Groeneveld, and H\. Hajishirzi \(2025\)Critical batch size revisited: a simple empirical approach to large\-batch language model training\.External Links:2505\.23971,[Link](https://arxiv.org/abs/2505.23971)Cited by:[Appendix C](https://arxiv.org/html/2608.10296#A3.SS0.SSS0.Px1.p2.1)\.
- J\. D\. Morrison, C\. Na, J\. Fernandez, T\. Dettmers, E\. Strubell, and J\. Dodge \(2025\)Holistically evaluating the environmental impact of creating language models\.ArXivabs/2503\.05804\.External Links:[Link](https://api.semanticscholar.org/CorpusID:276902612)Cited by:[Ethics Statement](https://arxiv.org/html/2608.10296#Sx1.p1.1)\.
- Olmo Team, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison, J\. Morrison, J\. Poznanski, K\. Lo, L\. Soldaini, M\. Jordan, M\. Chen, M\. Noukhovitch, N\. Lambert, P\. Walsh, P\. Dasigi, R\. Berry, S\. Malik, S\. Shah, S\. Geng, S\. Arora, S\. Gupta, T\. Anderson, T\. Xiao, T\. Murray, T\. Romero, V\. Graf, A\. Asai, A\. Bhagia, A\. Wettig, A\. Liu, A\. Rangapur, C\. Anastasiades, C\. Huang, D\. Schwenk, H\. Trivedi, I\. Magnusson, J\. Lochner, J\. Liu, L\. Miranda, M\. Sap, M\. Morgan, M\. Schmitz, M\. Guerquin, M\. Wilson, R\. Huff, R\. L\. Bras, R\. Xin, R\. Shao, S\. Skjonsberg, S\. Z\. Shen, S\. S\. Li, T\. Wilde, V\. Pyatkin, W\. Merrill, Y\. Chang, Y\. Gu, Z\. Zeng, A\. Sabharwal, L\. Zettlemoyer, P\. W\. Koh, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi \(2025\)Olmo 3: charting a path through the model flow to lead open\-source ai\.Technical reportAllen Institute for AI\.Note:[https://allenai\.org/papers/olmo3](https://allenai.org/papers/olmo3)Technical reportCited by:[Appendix C](https://arxiv.org/html/2608.10296#A3.p1.1),[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p4.3),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2608.10296#S2.SS2.p1.1),[§5\.2](https://arxiv.org/html/2608.10296#S5.SS2.p1.1)\.
- T\. OLMo, P\. Walsh, L\. Soldaini, D\. Groeneveld, K\. Lo, S\. Arora, A\. Bhagia, Y\. Gu, S\. Huang, M\. Jordan, N\. Lambert, D\. Schwenk, O\. Tafjord, T\. Anderson, D\. Atkinson, F\. Brahman, C\. Clark, P\. Dasigi, N\. Dziri, A\. Ettinger, M\. Guerquin, D\. Heineman, H\. Ivison, P\. W\. Koh, J\. Liu, S\. Malik, W\. Merrill, L\. J\. V\. Miranda, J\. Morrison, T\. Murray, C\. Nam, J\. Poznanski, V\. Pyatkin, A\. Rangapur, M\. Schmitz, S\. Skjonsberg, D\. Wadden, C\. Wilhelm, M\. Wilson, L\. Zettlemoyer, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi \(2025\)2 olmo 2 furious\.External Links:2501\.00656,[Link](https://arxiv.org/abs/2501.00656)Cited by:[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p4.3)\.
- G\. Penedo, H\. Kydlíček, L\. B\. allal, A\. Lozhkov, M\. Mitchell, C\. Raffel, L\. V\. Werra, and T\. Wolf \(2024\)The fineweb datasets: decanting the web for the finest text data at scale\.External Links:2406\.17557,[Link](https://arxiv.org/abs/2406.17557)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.p1.1)\.
- B\. Peng, E\. Alcaide, Q\. Anthony, A\. Albalak, S\. Arcadinho, S\. Biderman, H\. Cao, X\. Cheng, M\. Chung, M\. Grella, K\. K\. GV, X\. He, H\. Hou, J\. Lin, P\. Kazienko, J\. Kocon, J\. Kong, B\. Koptyra, H\. Lau, K\. S\. I\. Mantri, F\. Mom, A\. Saito, G\. Song, X\. Tang, B\. Wang, J\. S\. Wind, S\. Wozniak, R\. Zhang, Z\. Zhang, Q\. Zhao, P\. Zhou, Q\. Zhou, J\. Zhu, and R\. Zhu \(2023\)RWKV: reinventing rnns for the transformer era\.External Links:2305\.13048,[Link](https://arxiv.org/abs/2305.13048)Cited by:[footnote 1](https://arxiv.org/html/2608.10296#footnote1)\.
- B\. Peng, J\. Quesnelle, H\. Fan, and E\. Shippole \(2026\)YaRN: efficient context window extension of large language models\.External Links:2309\.00071,[Link](https://arxiv.org/abs/2309.00071)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1),[§5\.3](https://arxiv.org/html/2608.10296#S5.SS3.p1.1)\.
- O\. Press, N\. A\. Smith, and M\. Lewis \(2022\)Train short, test long: attention with linear biases enables input length extrapolation\.InThe Tenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=R8sQPpGCv0)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- J\. Qin, Y\. Xi, J\. Huang, R\. Rui, D\. Yin, W\. Liu, Y\. Yu, W\. Zhang, and X\. Sun \(2025\)APTBench: benchmarking agentic potential of base llms during pre\-training\.External Links:2510\.24397,[Link](https://arxiv.org/abs/2510.24397)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p1.1)\.
- Z\. Qiu, Z\. Huang, K\. Wen, P\. Jin, B\. Zheng, Y\. Zhou, H\. Huang, Z\. Wang, X\. Li, H\. Zhang, Y\. Xu, H\. Lian, S\. Zhang, R\. Men, J\. Zhang, I\. Titov, D\. Liu, J\. Zhou, and J\. Lin \(2026\)A unified view of attention and residual sinks: outlier\-driven rescaling is essential for transformer training\.External Links:2601\.22966,[Link](https://arxiv.org/abs/2601.22966)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1)\.
- Z\. Qiu, Z\. Wang, B\. Zheng, Z\. Huang, K\. Wen, S\. Yang, R\. Men, L\. Yu, F\. Huang, S\. Huang, D\. Liu, J\. Zhou, and J\. Lin \(2025\)Gated attention for large language models: non\-linearity, sparsity, and attention\-sink\-free\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=1b7whO4SfY)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p2.1)\.
- C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. Liu \(2020\)Exploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of Machine Learning Research21\(140\),pp\. 1–67\.External Links:[Link](http://jmlr.org/papers/v21/20-074.html)Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1)\.
- M\. Reid, V\. Zhong, S\. Gururangan, and L\. Zettlemoyer \(2022\)M2D2: a massively multi\-domain language modeling dataset\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 964–975\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.63)Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1)\.
- R\. Schaeffer, H\. Schoelkopf, B\. Miranda, G\. Mukobi, V\. Madan, A\. Ibrahim, H\. Bradley, S\. Biderman, and S\. Koyejo \(2025\)Why has predicting downstream capabilities of frontier AI models with scale remained elusive?\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.External Links:[Link](https://proceedings.mlr.press/v267/schaeffer25b.html)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px2.p1.1)\.
- L\. Soldaini, R\. Kinney, A\. Bhagia, D\. Schwenk, D\. Atkinson, R\. Authur, B\. Bogin, K\. Chandu, J\. Dumas, Y\. Elazar, V\. Hofmann, A\. Jha, S\. Kumar, L\. Lucy, X\. Lyu, N\. Lambert, I\. Magnusson, J\. Morrison, N\. Muennighoff, A\. Naik, C\. Nam, M\. Peters, A\. Ravichander, K\. Richardson, Z\. Shen, E\. Strubell, N\. Subramani, O\. Tafjord, E\. Walsh, L\. Zettlemoyer, N\. Smith, H\. Hajishirzi, I\. Beltagy, D\. Groeneveld, J\. Dodge, and K\. Lo \(2024\)Dolma: an open corpus of three trillion tokens for language model pretraining research\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15725–15788\.External Links:[Link](https://aclanthology.org/2024.acl-long.840/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.840)Cited by:[§D\.2](https://arxiv.org/html/2608.10296#A4.SS2.p1.1)\.
- Y\. Sun, L\. Dong, B\. Patra, S\. Ma, S\. Huang, A\. Benhaim, V\. Chaudhary, X\. Song, and F\. Wei \(2023\)A length\-extrapolatable transformer\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Toronto, Canada,pp\. 14590–14604\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.816),[Link](https://aclanthology.org/2023.acl-long.816/)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- J\. Wang, T\. Ji, Y\. Wu, H\. Yan, T\. Gui, Q\. Zhang, X\. Huang, and X\. Wang \(2024\)Length generalization of causal transformers without position encoding\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 14024–14040\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.834),[Link](https://aclanthology.org/2024.findings-acl.834/)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- Z\. Wang, F\. Zhou, X\. Li, and P\. Liu \(2025\)OctoThinker: mid\-training incentivizes reinforcement learning scaling\.External Links:2506\.20512,[Link](https://arxiv.org/abs/2506.20512)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p1.1)\.
- W\. Wu, Y\. Wang, G\. Xiao, H\. Peng, and Y\. Fu \(2024\)Retrieval head mechanistically explains long\-context factuality\.External Links:2404\.15574,[Link](https://arxiv.org/abs/2404.15574)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px2.p1.1)\.
- G\. Xiao, Y\. Tian, B\. Chen, S\. Han, and M\. Lewis \(2024\)Efficient streaming language models with attention sinks\.External Links:2309\.17453,[Link](https://arxiv.org/abs/2309.17453)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1)\.
- W\. Xiong, J\. Liu, I\. Molybog, H\. Zhang, P\. Bhargava, R\. Hou, L\. Martin, R\. Rungta, K\. A\. Sankararaman, B\. Oguz, M\. Khabsa, H\. Fang, Y\. Mehdad, S\. Narang, K\. Malik, A\. Fan, S\. Bhosale, S\. Edunov, M\. Lewis, S\. Wang, and H\. Ma \(2023\)Effective long\-context scaling of foundation models\.External Links:2309\.16039,[Link](https://arxiv.org/abs/2309.16039)Cited by:[§2\.2](https://arxiv.org/html/2608.10296#S2.SS2.p1.1)\.
- W\. Xiong, J\. Liu, I\. Molybog, H\. Zhang, P\. Bhargava, R\. Hou, L\. Martin, R\. Rungta, K\. A\. Sankararaman, B\. Oguz, M\. Khabsa, H\. Fang, Y\. Mehdad, S\. Narang, K\. Malik, A\. Fan, S\. Bhosale, S\. Edunov, M\. Lewis, S\. Wang, and H\. Ma \(2024\)Effective long\-context scaling of foundation models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 4643–4663\.External Links:[Link](https://aclanthology.org/2024.naacl-long.260/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.260)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025a\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p4.3),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px2.p1.2)\.
- B\. Yang, B\. Venkitesh, D\. Talupuru, H\. Lin, D\. Cairuz, P\. Blunsom, and A\. Locatelli \(2025b\)Rope to nope and back again: a new hybrid attention strategy\.External Links:2501\.18795,[Link](https://arxiv.org/abs/2501.18795)Cited by:[Figure 8](https://arxiv.org/html/2608.10296#A2.F8),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px1.p1.1),[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
- G\. Yang, E\. Hu, I\. Babuschkin, S\. Sidor, X\. Liu, D\. Farhi, N\. Ryder, J\. Pachocki, W\. Chen, and J\. Gao \(2021\)Tuning large neural networks via zero\-shot hyperparameter transfer\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 17084–17097\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2021/file/8df7c2e3c3c3be098ef7b382bd2c37ba-Paper.pdf)Cited by:[Appendix C](https://arxiv.org/html/2608.10296#A3.SS0.SSS0.Px1.p1.1)\.
- S\. Yang, J\. Kautz, and A\. Hatamizadeh \(2025c\)Gated delta networks: improving mamba2 with delta rule\.External Links:2412\.06464,[Link](https://arxiv.org/abs/2412.06464)Cited by:[footnote 1](https://arxiv.org/html/2608.10296#footnote1)\.
- T\. Ye, L\. Dong, Y\. Xia, Y\. Sun, Y\. Zhu, G\. Huang, and F\. Wei \(2025\)Differential transformer\.External Links:2410\.05258,[Link](https://arxiv.org/abs/2410.05258)Cited by:[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p1.1),[§5\.4](https://arxiv.org/html/2608.10296#S5.SS4.SSS0.Px1.p2.1)\.
- H\. Yen, T\. Gao, M\. Hou, K\. Ding, D\. Fleischer, P\. Izsak, M\. Wasserblat, and D\. Chen \(2025\)HELMET: how to evaluate long\-context models effectively and thoroughly\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=293V3bJbmE)Cited by:[§2\.3](https://arxiv.org/html/2608.10296#S2.SS3.p1.1)\.
- R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. Choi \(2019\)HellaSwag: can a machine really finish your sentence?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 4791–4800\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1472)Cited by:[§D\.3](https://arxiv.org/html/2608.10296#A4.SS3.p1.1)\.
- H\. Zhang, D\. Morwani, N\. Vyas, J\. Wu, D\. Zou, U\. Ghai, D\. Foster, and S\. Kakade \(2025\)How does critical batch size scale in pre\-training?\.External Links:2410\.21676,[Link](https://arxiv.org/abs/2410.21676)Cited by:[Appendix C](https://arxiv.org/html/2608.10296#A3.SS0.SSS0.Px1.p2.1)\.
- L\. Zhao, T\. Wei, L\. Zeng, C\. Cheng, L\. Yang, P\. Cheng, L\. Wang, C\. Li, X\. Wu, B\. Zhu, Y\. Gan, R\. Hu, S\. Yan, H\. Fang, and Y\. Zhou \(2024\)LongSkywork: a training recipe for efficiently extending context length in large language models\.External Links:2406\.00605,[Link](https://arxiv.org/abs/2406.00605)Cited by:[§1](https://arxiv.org/html/2608.10296#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.10296#S2.SS1.SSS0.Px4.p1.1)\.
- I\. Zimerman and L\. Wolf \(2023\)On the long range abilities of transformers\.External Links:2311\.16620,[Link](https://arxiv.org/abs/2311.16620)Cited by:[§6](https://arxiv.org/html/2608.10296#S6.SS0.SSS0.Px1.p1.1)\.
## Appendix AOlmPool
Tables[1](https://arxiv.org/html/2608.10296#A1.T1)and[2](https://arxiv.org/html/2608.10296#A1.T2)describe all 26 models in OlmPool\. The architectural columns of the table are identical, and replicated in the second table solely for ease of reading; Table[1](https://arxiv.org/html/2608.10296#A1.T1)reports HELMET and LongPPL and Table[2](https://arxiv.org/html/2608.10296#A1.T2)reports RULER scores for all runs\. Both tables are sorted in order of HELMET score at 32K, with the best score bolded and the worst score italicized in each column\.
The column labeled SWA indicates training with sliding window layers: three out of every four layers use a 4096 local window in lieu of full attention\. The QKNorm column indicates the presence of layerwise \(✓\\checkmark\) or headwise \(hw\) QK norm\. The initializations for each run is labeled with a letter code\. The number of KV heads indicates the presence and degree of GQA used: 32 KV heads indicates full MHA\. The feedforward dimensiondffd\_\{ff\}was adjusted for some runs with GQA to keep the parameter counts as close as possible across training runs\.
Table 1:All runs sorted by HELMET 32K \(worst to best\): HELMET scores and LongPPL\.Table 2:All runs sorted by HELMET 32K \(worst to best\): RULER scores\.
## Appendix BAdditional features studied
#### Norm ordering and type of QK norm have small effects\.
We ablate the three combinations of QK norm and norm order that are stable in our training regime: preorder without QK norm, preorder with QK norm, and post\-sublayer order with QK norm in Figure[8](https://arxiv.org/html/2608.10296#A2.F8)\. We find that norm order alone has an inconsistent effect, with QK norm accounting for the vast majority of the performance difference between runs\. Headwise QK norm causes an additional slight degradation\. QK norm is most often added to improve training stability \(e\.g\.Chameleon Team \([2025](https://arxiv.org/html/2608.10296#bib.bib11)\); Marin Community \([2025](https://arxiv.org/html/2608.10296#bib.bib10)\)\); we also confirm that it improves stability in our setting in Appendix[D\.4](https://arxiv.org/html/2608.10296#A4.SS4)\.
pre/no QKNpre/QKNpost/QKNpost/hw QKN0202040406060HELMET 32KOlmo\(SWA, 32 KV heads\)pre/no QKNpre/QKNpost/QKN0202040406060Llama\(no SWA, 8 KV heads\)Figure 8:QK norm is harmful for long context \(as first observed byYanget al\.\([2025b](https://arxiv.org/html/2608.10296#bib.bib54)\)\); we further note that the headwise variant is an additional slight detriment to downstream performance, and the less harmful norm order with QK norm is model\-family\-dependent\.
#### Float8 pretraining does not appear to have a consistent effect\.
We consider training linear layers in float8 instead of bfloat16 precision\. In two controlled comparisons, adding float8 training results in a slight degradation in one comparison and an improvement in the other\. We find no evidence that float8 pretraining degrades long context performance— instead, it appears that float8’s effect in OlmPool models may be purely noise, like initialization, because float8 optimization results in a slightly different optimization path\. We note that we do not consider a setting where attention weights are quantized during pretraining; it is possible that this would result in degradation, as long context abilities inherently require the model to be able to attend precisely over a long context window\.
#### Initialization causes more variation in short context than long context\.
We are not able to completely standardize initialization across runs because some models differ in parameterization\. In OlmPool, we construct four pairs of runs that are identical except for initialization and measure the swing in performance for both short context and long context metrics\. To standardize across metrics with differing scales, we compute the swing due to initialization as a percentage of the maximum variation between runs in OlmPool\.
Initialization has minimal impact on train loss and perplexity \(shifting runs by less than 3% of the cross\-run range on average\) and causes the largest swings in the short context benchmarks and pre\-context\-extension HELMET scores, where runs with the same architecture but different initialization may vary by up to 58% of the total range of variance\. This effect is reduced after long context extension; on the long context benchmarks, initialization can account for smaller but still substantial swings of up to 17% \(and on average 7\.7%\) of the observed range of values\. In the results, we discuss only differences across architectural variations that result, on average, in long context score shifts substantially larger than the mean shift from initialization\. Note that the sole dimension where we cannot construct an initalization\-controlled trial is in measuring the impact of varying GQA degree; thus, the effect size for the difference there may be slightly larger or smaller than the one we observe\.
## Appendix CAdditional Training Details
For additional details on training and the pretraining data used, we refer the reader to the Olmo 3 technical report\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\); we follow the configuration for Olmo 3 7B Base pretraining, although we use substantially less than 1024 concurrent GPUs for these much shorter runs\. The[project github repository](https://github.com/allenai/olmpool/)provides configurations for each pretraining and context extension run, and is the best reference point for specific training details for each run\.
#### On optimization hyperparameters\.
We use the same learning rate, batch size, and optimization schedule for all models in OlmPool\. While it would be infeasible to sweep all hyperparameters for all models in OlmPool, we believe that our selections for these hyperparameters are reasonable for all models in the pool for several reasons\. First, because we only pretrain for 140B tokens in a 5T learning rate schedule, changing the learning rate schedule causes only marginal differences\. For instance, we trained one development run at both 5T and 7T learning rate schedules, and observed near\-identical performance on both loss and downstream metrics\. We do not change the depth of networks888With one exception: the Qwenlike model is 36 instead of 32 layers to match Qwenand hold parameter count as close to constant as possible, standardizing two major factors often identified as driving changes in optimal learning rate\(Yanget al\.,[2021](https://arxiv.org/html/2608.10296#bib.bib5)\)\.
Additionally, several works observe that batch size has minimal impact below a critical batch size at which learning slows\(McCandlishet al\.,[2018](https://arxiv.org/html/2608.10296#bib.bib2)\)\. We are below the estimated critical batch size for models of this configuration\(Merrillet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib3)\)for the vast majority of training, and we have no reason to believe that any of these changes would impact the critical batch size; in fact,Zhanget al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib4)\)show empirically that context length does not impact critical batch size\.
#### Skip\-step optimizer
A final important detail that we highlight here: like all recent Olmo models, OlmPool models are trained with a skip\-step optimizer, which skips steps that have an abnormally high gradient norm to increase run stability\. This triggers very rarely, but when it does occur, it causes a small amount of variance in pretraining data across OlmPool \(i\.e\., for a small percentage of data, not every model performs a gradient update\)\. For more details, see the[documentation](https://olmo-core.readthedocs.io/en/latest/optim.html#olmo_core.optim.SkipStepOptimizer)\.
## Appendix DFurther evaluation details
This appendix briefly describes each short\-context evaluation metric and validation perplexity set and its correlation with long context downstream\.
### D\.1Loss
In Figure[9](https://arxiv.org/html/2608.10296#A4.F9), we graph the \(slight\) correlation between loss and long context ability\. Both pretraining and long context loss fail to correlate with long context capabilities\.
2\.122\.122\.142\.142\.162\.162\.182\.182\.22\.23030404050506060Pretraining CE lossHELMET 32K1\.41\.41\.51\.51\.61\.61\.71\.73030404050506060LC extension CE loss
Figure 9:Training loss does not correlate well with long context ability, in either the pretraining or long context extension runs\.
### D\.2Perplexity on held\-out text
We measure perplexity across a diverse set of held\-out validation splits, drawn mostly from Paloma\(Magnussonet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib75)\)\.C4\(Raffelet al\.,[2020](https://arxiv.org/html/2608.10296#bib.bib16)\)is a filtered and deduplicated web corpus derived from Common Crawl; we use the English validation split\.Dolma\(Soldainiet al\.,[2024](https://arxiv.org/html/2608.10296#bib.bib20)\)is Ai2’s open pretraining corpus; we evaluate on six domain\-specific splits: books, Common Crawl, peS2o \(scientific papers\), Reddit, Stack Exchange, and Wikipedia\.ICE\(Greenbaum and Nelson,[1996](https://arxiv.org/html/2608.10296#bib.bib17)\)\(International Corpus of English\) provides text spanning multiple varieties of dialectal English\.M2D2\(Reidet al\.,[2022](https://arxiv.org/html/2608.10296#bib.bib18)\)is a massively multi\-domain dataset; we use the S2ORC split covering scientific text\.The Pile\(Gaoet al\.,[2020](https://arxiv.org/html/2608.10296#bib.bib19)\)is an 800GB diverse text corpus from EleutherAI, evaluated on its held\-out validation set\.WikiText\-103\(Merityet al\.,[2016](https://arxiv.org/html/2608.10296#bib.bib9)\)is a standard language modeling benchmark of Wikipedia articles\.
Figure[10](https://arxiv.org/html/2608.10296#A4.F10)shows the per\-split correlation with the downstream long context metrics\.
00\.10\.10\.20\.20\.30\.30\.40\.40\.50\.50\.60\.60\.70\.70\.80\.80\.90\.911TrainWikiTextC4WikiBooksCommonCrawlRedditICEPileStackPes2oS2ORCR2R^\{2\}LongPPLRULER\_32KHELMET\_32KFigure 10:R2R^\{2\}between each standard validation perplexity metric and downstream long\-context metrics \(HELMET 32K, RULER 32K, LongPPL\)\.
### D\.3BPB on downstream benchmarks
We evaluate short\-context capabilities across several standard benchmarks\.ARC\(Easy and Challenge splits\)\(Clarket al\.,[2018](https://arxiv.org/html/2608.10296#bib.bib13)\)is grade\-school science question answering, with the Challenge split requiring more complex reasoning\.HellaSwag\(Zellerset al\.,[2019](https://arxiv.org/html/2608.10296#bib.bib14)\)measures commonsense natural language inference via sentence completion\.MMLU\(Hendryckset al\.,[2021](https://arxiv.org/html/2608.10296#bib.bib15)\)measures graduate\-level knowledge across 57 subjects; we use the general split into humanities, social sciences, STEM, and other domains \.HumanEval\(Chenet al\.,[2021](https://arxiv.org/html/2608.10296#bib.bib12)\)requires Python code generation conditioned on docstrings\.MBPP\(Austinet al\.,[2021](https://arxiv.org/html/2608.10296#bib.bib34)\)similarly evaluates code generation on entry\-level Python programming problems\.MINERVA\(Lewkowyczet al\.,[2022](https://arxiv.org/html/2608.10296#bib.bib35)\)is a set of 500 math problems requiring multi\-step numerical and symbolic computation\. Finally,Basic Skills\(Olmo Teamet al\.,[2025](https://arxiv.org/html/2608.10296#bib.bib67)\)is a datatset covering six fundamental competencies: arithmetic, coding, common knowledge, logical reasoning, pattern recognition, and string operations\. All downstream metrics are reported as bits\-per\-byte \(BPB\) on the gold answer string\.
Figure[11](https://arxiv.org/html/2608.10296#A4.F11)shows the correlation between each individual short context benchmark BPB and the downstream long context scores\. No benchmark correlates strongly, and we do not observe any consistent trend in which types of benchmarks correlate most\. Note one of the three coding\-related benchmarks correlates strongly, while the others do not\. Since LongPPL correlates very poorly with validation perplexity for these datasets, we believe this is likely because these datasets contain a supermajority of tokens which can be predicted using only short\-context dependencies; for more details on this problem, seeFanget al\.\([2025](https://arxiv.org/html/2608.10296#bib.bib76)\)\.
00\.10\.10\.20\.20\.30\.30\.40\.40\.50\.50\.60\.60\.70\.70\.80\.80\.90\.911SC\_avgARC\-CARC\-EHellaSwagMMLU\-HumMMLU\-OtherMMLU\-SocSciMMLU\-STEMHumanEvalMBPPMATHBasicSkillsBS\-ArithBS\-CodeBS\-LogicBS\-PatternBS\-StringR2R^\{2\}LongPPLRULER\_32KHELMET\_32KFigure 11:R2R^\{2\}between each short\-context benchmark metric and downstream long\-context metrics \(HELMET 32K, RULER 32K, LongPPL\)\.
### D\.4Training stability
Another important feature in pretraining is the stability of the training process\. We measure a score for stability by computing the percentage of gradient norms that are more than 6 standard deviations from the mean during the pretraining run\. We observe that runs with QK norm have, on average, a lower spike score, which aligns with prior observations that QK norm is beneficial for stability\.
We then calculate correlation between this score and long context performance downstream and visualize this in Figure[12](https://arxiv.org/html/2608.10296#A4.F12)\. In OlmPool, there is a slight negative correlation \(R2=0\.22\)R^\{2\}=0\.22\)between pretraining stability and long context performance, mostly because QK norm both improves stability and damages long context performance\.
223344⋅10−3\\cdot 10^\{\-3\}303040405050SpikeScore \(↓\\downarrow\)HELMET 32K \(↑\\uparrow\)Figure 12:Spike score vs\. HELMET 32K across OlmPool\. Stability weakly correlates with worse LC performance downstream, mostly due to QK\-norm’s influence on both factors\.相似文章
@NVIDIAAI: A long-context model's serving speed is largely decided before training starts. Attention used to be a small part of a …
NVIDIA explains how attention architecture choices (group size, head dimension, KV-cache size, parallelism) set the ceiling for long-context inference performance, with guidelines for co-designing models for faster serving.
Jet-Long: 具有动态双焦RoPE的高效长上下文扩展
Jet-Long提出了一种无需微调的零样本方法,通过动态调整RoPE缩放来扩展LLM上下文长度,在高达128K上下文的基准测试中取得了强劲性能,且推理开销极小。
@VukRosic99: 长上下文Transformer面临两大瓶颈:二次注意力计算和KV缓存(在1M tokens时可达数百GB)…
MiniCPM-SALA是一款9B参数的混合注意力模型,通过在稀疏注意力和线性注意力之间交替插入(每3个线性层插入1个稀疏层)来克服长上下文Transformer的二次计算和KV缓存瓶颈。在256K tokens下,其推理速度比Qwen3-8B快3.5倍,并能在消费级GPU上支持高达1M tokens。该模型采用经济高效的持续训练方法,训练成本降低约75%。
@GergelyOrosz: 试图弄清楚,在上下文窗口中使用更多上下文(我称之为上下文深度),在更长的运行中,错误会如何累积/代理会如何漂移……
Gergely Orosz 指出,在AI代理的上下文窗口中使用更多上下文进行更长的运行会增加错误和漂移,建议使用更短的运行和更少的上下文以提高可靠性。
@akshay_pachaar: 扩展上下文窗口不仅仅是关于更大的矩阵。在传统的Transformer中,将token数量扩大8倍会…
解释了由于注意力的二次复杂度,扩展Transformer上下文窗口所带来的内存挑战,并暗示了解决方案。