Two AI Metrics Diverged: Will it Make All the Difference?
Summary
This paper analyzes how different AI performance metrics (bounded vs unbounded) determine whether frontier AI capabilities remain concentrated among wealthy actors or diffuse to smaller models, with implications for regulation.
View Cached Full Text
Cached at: 07/02/26, 05:41 AM
# Two AI Metrics Diverged: Will it Make All the Difference?
Source: [https://arxiv.org/html/2607.00913](https://arxiv.org/html/2607.00913)
###### Abstract
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with “meek models inheriting the earth”? Building onGundlachet al\.\([2025b](https://arxiv.org/html/2607.00913#bib.bib7)\), we show that the answer depends on how we value and measure AI capabilities\. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on other metrics frontier models grow their lead forever\. Classifying performance metrics by their functional forms in relation to training \(and inference\) compute, we provide tight mathematical conditions for determining which metrics favor meek models, and show that bounded performance metrics always do\. But careful interpretation of performance metrics is essential: we show that many common bounded metrics have closely\-related counterpart metrics that are unbounded \(and vice versa\)\. Determining the apt metric in a domain is a prerequisite for policy, since bounded and unbounded metrics may suggest opposing policy responses\. If a particular capability — like software engineering, synthetic biology, or rhetorical persuasiveness — is unbounded when measured in the terms we care about, frontier\-level capability will likely be concentrated in the hands of a few wealthy actors\. Conversely, if that capability is instead bounded, frontier\-level capabilities proliferate through meek models into the hands of the many\.
Machine Learning, Scaling Laws, Compute, Benchmarks, Time Horizon
## 1Introduction
### 1\.1Background
When AI models are trained with more compute they gain increased capabilities, as measured by \(for example\) validation loss\(Kaplanet al\.,[2020](https://arxiv.org/html/2607.00913#bib.bib34); Rosenfeld,[2021](https://arxiv.org/html/2607.00913#bib.bib32); Hoffmannet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib35); Bahriet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib33)\)\. Similarly, when AI models are allowed to reason about problems for longer — by using more compute at inference time — capabilities improve, as measured by \(for example\) success on benchmark tasks\(Jones,[2021](https://arxiv.org/html/2607.00913#bib.bib2); Villalobos and Atkinson,[2023a](https://arxiv.org/html/2607.00913#bib.bib55)\)\. As a result of these regularities, companies have invested exponentially increasing amounts in compute for pre\-training and inference\(Sevillaet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib36); Juniewicz,[2026](https://arxiv.org/html/2607.00913#bib.bib72)\)\.
These large expenditures have enabled a small handful of companies to offer more powerful AI capabilities than available elsewhere\.
But is this oligopolistic equilibrium guaranteed? Providers of smaller, cheaper open\-weight models have been able to replicate the capabilities of expensive proprietary models with only a short delay — often less than a year\(Cottieret al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib54); Emberson,[2025](https://arxiv.org/html/2607.00913#bib.bib21)\)\. If the capabilities gap doesn’t widen, but instead shrinks, we might expect frontier\-level capabilities to diffuse very widely\. In such a world, regulating AI capabilities may require compliance from many actors, or require enforcing regulations at the hardware level — which would require substantial new technological means and political will\(O’Garaet al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib27)\)\.
Given the regulatory relevance, it would be valuable to know whether the capabilities gap between expensive proprietary models and cheaper open models will shrink or widen\. However the empirical literature is sparse, ambivalent, and — as we’ll see — sensitive to one’s individual utility function over a range of capabilities\(Emberson,[2025](https://arxiv.org/html/2607.00913#bib.bib21); AI Index Steering Committee,[2025](https://arxiv.org/html/2607.00913#bib.bib22); Ihle,[2025](https://arxiv.org/html/2607.00913#bib.bib20)\)\.
Against this backdrop, a recent paperGundlachet al\.\([2025b](https://arxiv.org/html/2607.00913#bib.bib7)\)provides a theoretical argument that the gap could shrink\. The authors \(also authors of this paper\) compare the performance of models with exponentially growing investment in training compute to “meek” models which are trained at a constant level of compute investment\. Both models benefit from exponential improvements in hardware and algorithm efficiency \(meek models’effectivecompute grows, even if theircompute investmentdoesn’t\), but meek models ride a slower exponential growth rate\. The authors show that, on a few common performance metrics — validation loss, sigmoidal benchmarks — the performance of meek models and frontier models converge in the long run\. They analyze several specific capabilities, but the long\-run properties do not depend on the specific domain, only the functional form of how compute translates into performance\. The underlying intuition is that both models are traversing a fundamentally bounded performance metric at different paces as they increase their effective compute: while the frontier model gets near the bound sooner, eventually both models are in the region where the metric doesn’t change much \(or in some cases, at all\) over time\.
### 1\.2Capabilities Versus Metrics
As we will show, this result is sensitive to the specific mathematical properties of the performance metric analyzed\. This gives rise to a problem: in machine learning, two metrics used to measure ostensibly similar capabilities can often have very different functional forms with respect to compute\. Metrics are often selected for accurately capturing the ordinal ranking of models in some domain, and there is typically no requirement that the cardinal performance level of models capture something especially meaningful\. This is an issue if we aim to draw conclusions about a performance gap in units that are meaningful — if we do, that would require us to use a performance metric that reflects how capabilities improvements arevaluedin some domain\. Utility as a function of model capability may be highly non\-linear, and the behavior of the selected performance metric needs to take into account this non\-linearity if it aims to make meaningful claims about performance gaps or other cardinal properties of the metric\.
In this paper, we show that some performance metrics are “meek” metrics and some are “mighty” metrics\. “Meek” metrics are ones where “meek models inherit the earth” — that is, they eventually see models converge in performance under exponentially diverging compute expenditure\. In contrast, “mighty” metrics do not converge\. We show that if a metric is bounded, then it is meek\.Critically, we also show that subtle differences in utility function111A note on terms: by “utility function”, we are referring to any function that evaluates how much a particular gain in capabilities matters for some real\-world aim\. Utility functions may differ between actors, tasks, and domains\.over a model’s capabilities can produce metrics which change from unbounded to bounded, from mighty to meek, from sensitive to exponential investment gaps to \(in the long run\) completely indifferent to them\.
Consider the case of two software engineers using coding agents\. The first is required to briefly validate the code produced at regular task intervals \(e\.g\. every few human work hours\) and aims to improve the accuracy within that interval; the second aims to maximize the length of the task that can be achieved within some reasonable error tolerance\. Both engineers prefer models with improving capabilities, but their utility functions produce dramatically different real\-world preferences\. The first engineer’s utility function isbounded\(since there is a maximum 100% success rate\)\. Therefore, as time progresses, the first engineer will find less and less difference between models trained with exponentially more compute and models which only leverage shared improvements in algorithms and hardware\. By contrast, the second engineer will see continual improvements on their metric from larger and larger models\.
### 1\.3Contributions
Upon seeing empirical claims about converging or diverging performance in some domain, it is important to understand that the result may be imposed by the choice of performance metric\. The task of an AI analyst, economist, or policymaker is then to determine whether the metric is appropriate\.
On the metrics we care about, it is useful to understand the long\-term behavior of model capabilities\. Our expectation, informed by our discussion of metrics below, is that we will see meek models inherit some parts of the world, and the mighty dominate in other domains\. Concentration in some areas; diffusion in others\.
The remainder of this paper is as follows\. We summarize the contribution ofGundlachet al\.\([2025b](https://arxiv.org/html/2607.00913#bib.bib7)\); define “meek metrics” and describe the boundary conditions that confer meek metric status \(Section[2](https://arxiv.org/html/2607.00913#S2)\); survey common performance metrics and discuss their sensitivity to interpretation \(Section[3](https://arxiv.org/html/2607.00913#S3)\); discuss how the framework can be generalized to model gaps in inference compute \(Section[4](https://arxiv.org/html/2607.00913#S4)\); and conclude by discussing limitations, complexities, and implications for society \(Section[5](https://arxiv.org/html/2607.00913#S5)\)\.
## 2Core Mathematical Definitions and Results
### 2\.1Meek Models Argument
Our goal here is to generalize the argument made inGundlachet al\.\([2025b](https://arxiv.org/html/2607.00913#bib.bib7)\), which analyzed theoretical differences invalidation lossbetween so\-calledmeek modelsandfrontier models\. Since validation loss is a power law in effective training compute, one can analytically observe the closed form of the difference as:
ΔL=A\(Cmeek\)−α−A\(Cfrontier\)−α\\Delta L=A\(C\_\{meek\}\)^\{\-\\alpha\}\-A\(C\_\{frontier\}\)^\{\-\\alpha\}
In this context,0<α<10<\\alpha<1andAAare parameters governing the power law, andCCiseffective training compute: raw training compute scaling multiplied by shared exponential growth factors fromalgorithmic / data progressandhardware efficiencyprogress\. While both frontier and meek models see exponentially growing effective training compute, frontier models havefasterexponential growth, since they exponentially increase investment in raw training compute over time\.
More concretely, letghg\_\{h\}be the shared annual growth factor of hardware efficiency \(in FLOPs per dollar\);gag\_\{a\}be the shared annual growth factor of algorithmic and data progress;gig\_\{i\}be the annual growth factor of training compute scaling for the frontier model only; andC0C\_\{0\}be some initial effective compute\. Then the loss difference yields a difference of two decaying exponentials, which quickly approach zero after a one\-time peak\.
ΔL=A\[\(gagh\)tC0\]−α−A\[\(gigagh\)tC0\]−α\\Delta L=A\[\(g\_\{a\}g\_\{h\}\)^\{t\}C\_\{0\}\]^\{\-\\alpha\}\-A\[\(g\_\{i\}g\_\{a\}g\_\{h\}\)^\{t\}C\_\{0\}\]^\{\-\\alpha\}
However, as we discuss in Section[3\.1\.1](https://arxiv.org/html/2607.00913#S3.SS1.SSS1), loss can be an opaque metric for measuring capabilities; it’s not obvious how to translate loss into utility\. Thus we ask, when does this argument generalize to other performance metrics?
### 2\.2Generalized Notion of “Meek Metrics”
We briefly present a collection of definitions and results which allow for straightforward categorization of performance metrics\. Proofs can be found in Appendix[A](https://arxiv.org/html/2607.00913#A1)\.
###### Definition 2\.1\(Normal Performance Metric\)\.
If a function mapping training compute to performance,P:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}, is both differentiable and weakly monotonically increasing, we call it anormal performance metric\.
These criteria are natural: \(1\) Performance growth is typically smooth in compute, and even when performance exhibits sharp jumps, such transitions are still smooth\(Poweret al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib8)\), \(2\) Although adding compute can diminish performance \(e\.g\. overfitting\), since we only require weak monotonicity, any model which degrades in performance can simply be discarded in favor of the prior, better model\.
Notice that here and throughout when we talk about “metrics” we are talking not just about a particular way of evaluating performance \(e\.g\. a benchmark\) but the curve that describes howrealized or forecastedperformance on that evaluation changes with compute\.
###### Definition 2\.2\(Meek Metric\)\.
LetP:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}\. We sayPPis a meek metric under exponential scaling if it is a normal performance metric and for all pairs of exponentsb\>a\>1b\>a\>1, and for all initial values of training computeC0\>0C\_\{0\}\>0, the following limit holds:
limt→∞P\(btC0\)−P\(atC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)=0
Intuitively,PPis invariant to exponential differences in effective training compute, whereaacomes from hardware/algorithmic efficiency, andbbadds exponential increases in investment\. Given some initial training compute investmentC0C\_\{0\}, two actors who see exponential differences in the growth of that effective training compute will see no durable difference in performance in the long run\.
###### Definition 2\.3\(Mighty Metric\)\.
A metric is mighty if and only if it is a normal performance metric and it is not meek\. That is, ifPPis a normal performance metric butlimt→∞P\(btC0\)−P\(atC0\)\>0\\lim\_\{t\\rightarrow\\infty\}P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)\>0, thenPPis mighty\.
Note that a mighty metric does notnecessarilyexhibitdivergencebetween frontier and meek model performance\. There could be a constant capabilities gap, for example \(see Theorem[2\.5](https://arxiv.org/html/2607.00913#S2.Thmtheorem5)\)\.
### 2\.3Results
We now present two results which classify the space of normal performance metrics\. We start with a common and convenient special case of bounded performance metrics\.
###### Theorem 2\.4\.
IfPPis a normal performance metric which is bounded above, thenPPis a meek metric\.
Although this result follows immediately from the Monotone Convergence Theorem222The Monotone Convergence Theorem asserts that a monotone real\-valued function with an upper bound must converge to its least upper bound\., it is practically useful given the diversity of performance metrics which are indeed bounded\. For example, the original result fromGundlachet al\.\([2025b](https://arxiv.org/html/2607.00913#bib.bib7)\)follows as an immediate corollary\.
For the broader class of unbounded normal performance metrics, we have the following characterization\.
###### Theorem 2\.5\.
LetPPbe a normal performance metric which is unbounded above\. Moreover, suppose the derivative ofPPwith respect tologlogC\\log\\log Cis eventually monotone\. ThenPPis a meek metric if and only if
limC→∞P\(C\)loglogC=0\\lim\_\{C\\rightarrow\\infty\}\\frac\{P\(C\)\}\{\\log\\log C\}=0
Intuitively, this bound holds for the following reason: a single application of the logarithm to inputsbtC0b^\{t\}C\_\{0\}andatC0a^\{t\}C\_\{0\}still results in two terms which grow distinctly: \(tlogb\+logC0t\\log b\+\\log C\_\{0\}\) and \(tloga\+logC0t\\log a\+\\log C\_\{0\}\)\. Yet their ratio clearly approacheslogb/loga\\log b/\\log a, such that after applying another logarithm, the limit approaches a constantlog\(logbloga\)=loglogb−logloga\\log\(\\frac\{\\log b\}\{\\log a\}\)=\\log\\log b\-\\log\\log a\. Thus, in the limit, the functionP\(C\)=loglog\(C\)P\(C\)=\\log\\log\(C\)results in a constant difference and acts as a sort of boundary point for this convergence\.
Together, these provide a tight criterion on the space of normal performance metrics: \(1\) Boundedness suffices to show meekness; and \(2\) for the set of unbounded normal metrics which have well\-behaved growth, functions are meek if and only if they grow slower thanloglogC\\log\\log C\.
Finally we show that being a meek metric under exponential compute scaling is equivalent to being a meek metric under power\-law compute scaling\. This shows meekness holds under more general assumptions about the future of progress in computing\.
###### Theorem 2\.6\(Equivalence Under Power Law Scaling\)\.
LetP:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}\. Then the following are equivalent:
1. 1\.For all exponentsb\>a\>1b\>a\>1and computeC0\>0C\_\{0\}\>0,limt→∞P\(btC0\)−P\(atC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)=0
2. 2\.For all powersb\>a\>0b\>a\>0and computeC0\>0C\_\{0\}\>0,limt→∞P\(tbC0\)−P\(taC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(t^\{b\}C\_\{0\}\)\-P\(t^\{a\}C\_\{0\}\)=0
IsP\(C\)P\(C\)bounded above?MEEKgap closesDoesP\(C\)loglogC→0\\frac\{P\(C\)\}\{\\log\\log C\}\\\!\\to\\\!0asC→∞C\\rightarrow\\infty?MEEKgap closesMIGHTYgap persistsYesNoYesNoe\.g\.benchmark %, win\-rate, pass@kk,−\-val\. losse\.g\.loglogC\\sqrt\{\\log\\log C\} \(theoretical; no known real\-world instance\)e\.g\.ELO, task horizon,−logϵ\-\\\!\\log\\epsilon, log\-compute index
Figure 1:Metric classification decision tree\.Any normal performance metric ismeek\(gap converges\) ormighty\(gap persists\) and can be classified with two questions\. Bounded metrics are always meek \(Theorem[2\.4](https://arxiv.org/html/2607.00913#S2.Thmtheorem4)\); for unbounded metrics meekness holds iff the metric grows slower thanloglogC\\log\\log C, subject to the weak growth conditions given in Theorem[2\.5](https://arxiv.org/html/2607.00913#S2.Thmtheorem5)\.
## 3The Subtleties of Performance Metrics in the Wild
The meek metric criterion is a strong one: if a performance metric is meek, it will eventually be indifferent to exponential differences in compute investment\. Naturally, one might wonder which metrics, if any, are meek metrics\. This question turns out to be both subtle and consequential\. At first glance, the literature consists mostly of two types of metrics: bounded metrics which are meek \(e\.g\. benchmarks, negative validation loss, inference scaling\) and non\-meek power law metrics \(e\.g\. game performance, time horizon\)\. Yet upon further examination, many meek \(or non\-meek\) metrics have closely\-related alternative metrics which are non\-meek \(or meek\)\. This has important implications for those tracking and governing AI progress: two slightly different interpretations of the same capability imply profoundly different relationships to compute expenditure\.
We now present examples from the literature of common performance metrics to highlight these subtleties\. In addition, we hope to supply the reader with a concise overview of many common metrics and their relationship to compute\.
### 3\.1Meek Metrics
#### 3\.1\.1Validation Loss
The most notable example of a meek performance metric is, of course, validation loss predicted by neural scaling laws\(Kaplanet al\.,[2020](https://arxiv.org/html/2607.00913#bib.bib34); Rosenfeld,[2021](https://arxiv.org/html/2607.00913#bib.bib32); Hoffmannet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib35)\)\. However, validation loss is a difficult metric to use on its own since \(1\) its interpretation is information\-theoretic, rather than directly describing performance on a real\-world task and \(2\) itdiminishesto zero\. To resolve the latter trouble, one can apply transformations to loss to make it a Normal Performance Metric, one that is unbounded and increasing, but this choice has inherent freedom\. As we discuss in Section[3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2), negative log\-loss is a common and useful transformation in predicting other metrics \(e\.g\. benchmarks\)\.
#### 3\.1\.2Benchmarks
In the era of large language models \(LLMs\), bounded metrics are by far the most common meek performance metric due to the prevalence of benchmarks scored from0\-100100\. Even pre\-dating LLMs, benchmarks like ImageNet\(Denget al\.,[2009](https://arxiv.org/html/2607.00913#bib.bib41)\), CIFAR\(Krizhevsky,[2009](https://arxiv.org/html/2607.00913#bib.bib42)\), and MNIST\(LeCunet al\.,[1998](https://arxiv.org/html/2607.00913#bib.bib43)\)were foundational for measuring progress across computer vision\.
Benchmarks which are bounded between0and100100are meek precisely due to their boundedness\. As analyzed originally in\(Gundlachet al\.,[2025b](https://arxiv.org/html/2607.00913#bib.bib7)\), we should expect benchmark scores to converge across exponentially diverging compute investment as progress in hardware and algorithms / data brings even fixed\-compute models closer to 100%\.
A body of existing work relating benchmarks, loss, and training compute demonstrates that benchmarks appear to be sigmoidal in negative log\-loss \(or log\-compute333Under compute optimality, neural scaling laws derive a power law between compute and loss\. Therefore, one expects log\-loss and log\-compute to be affine\.\)\. That is:BenchmarkAccuracy=σ\(−logL\)=σ\(logC\)\\text\{BenchmarkAccuracy\}=\\sigma\(\-\\log L\)=\\sigma\(\\log C\)\.Owen \([2024](https://arxiv.org/html/2607.00913#bib.bib15)\)fits a logistic sigmoid to Big\-Bench\(Srivastavaet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib39)\)and MMLU\(Hendryckset al\.,[2020](https://arxiv.org/html/2607.00913#bib.bib40)\)scores as a function of compute\. More generally,Ruanet al\.\([2024](https://arxiv.org/html/2607.00913#bib.bib17)\)computes the principal components of \(logit\-transformed\) benchmarks and finds correlations between the primary component and log\-compute\. Both papers find this sigmoidal relationship between performance and log\-compute\. Though one might naturally expect sigmoidal shapes for metrics which are bounded on both sides, why is the right transformation of compute \(or loss\)logarithmic?
ExtendingSchaefferet al\.\([2023](https://arxiv.org/html/2607.00913#bib.bib19)\), we posit an explanation which not only justifies the presence of the logarithm, but also argues that benchmarks can be naturally thought of as a transformation oflog\(1/ϵ\)\\log\(1/\\epsilon\)whereϵ\\epsilonis a per\-token, amortized, task level error rate, andlog\(1/ϵ\)\\log\(1/\\epsilon\)can be thought of as error rate orders of magnitude or “the number of nines of reliability\.”444Technically, it’s proportional to the number of nines, since we use natural log throughout\.Importantly, unlike the benchmarks it explains,log\(1/ϵ\)\\log\(1/\\epsilon\)is a mighty metric \(the so\-called “march of nines”\)\.
Schaeffer first notes that validation loss is roughly the \(negative\) log\-probability of correctly producing a particular token,L≈−logpiL\\approx\-\\log p\_\{i\}\(when irreducible loss is negligible\)\. In this case,LLis the reducible loss — total cross\-entropy minus the irreducible floor\. For some number of tokensTTrequired for a single question, idealizing these draws as independent, one arrives at the following expression for accuracy as a function of compute, for some scaling exponent0<α<10<\\alpha<1:
Accuracy\(C\)=\(elogpi\)⏟success rateT=\(e−L\)T=e\(−AC−αT\)\\text\{Accuracy\}\(C\)=\{\\underbrace\{\(e^\{\\log p\_\{i\}\}\)\}\_\{\\text\{success rate\}\}\}^\{T\}=\(e^\{\-L\}\)^\{T\}=e^\{\(\-AC^\{\-\\alpha\}T\)\}
This function is a very gradual sigmoid inCC, but as a function oflogC\\log C, it becomes the Gompertz function, a more abrupt sigmoid \(explaining why benchmarks are empirically well modeled as sigmoidal in log compute\)\. Writing the error rate using the approximatione−x≈1−xe^\{\-x\}\\approx 1\-xfor smallxx, we get
ϵ\(C\)=1−Accuracy\(C\)=1−e−LT≈LT\\epsilon\(C\)=1\-\\text\{Accuracy\}\(C\)=1\-e^\{\-LT\}\\approx LT
From this one straightforwardly sees that halving the error rate requires halving the loss, which itself requires a21/α×2^\{1/\\alpha\}\\timesincrease in compute \(for Chinchilla scaling roughly a 90× increase in training compute would be required to halve one’s error\(Hoffmannet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib35)\)\)\. Equivalently,log\(1/ϵ\)\\log\(1/\\epsilon\)is linear in log compute:
log\(1/ϵ\)≈−log\(LT\)≈αlogC−log\(AT\)\\log\(1/\\epsilon\)\\approx\-\\log\(LT\)\\approx\\alpha\\log\{C\}\-\\log\(AT\)
This connection to benchmarks is consistent with other work fromHoet al\.\([2025](https://arxiv.org/html/2607.00913#bib.bib18)\), which finds the fundamental capability scale is linear in log\-compute\. Using a suite of benchmark scores, and assuming a sigmoidal form relating some latent model capability to downstream benchmark score, Ho et al\. use an item\-response theory \(IRT\) model to back out a universal model capability index\. That index correlates strongly with log\-compute \(and thus likely log\-loss\), suggesting that the right scale for measuring capabilities is logarithmic in compute \(or loss\)\.
Both of these expositions of benchmarks suggest benchmark performance may be best explained by a fundamental scale which is linear in log\-compute – and therefore amighty metric– despite the benchmark itself being ameek metric\. Which scale is correct depends on one’s relationship to the underlying measure: if one’s utility is in error rate orders of magnitude, meek and frontier models will diverge\. If near perfect accuracy on a fixed benchmark suffices, meek models will indeed inherit the earth\.
#### 3\.1\.3Misalignment
This phenomenon can also arise when the functional form of benchmarks does not reflect the true utility of the underlying capability\. For example, suppose one is testing a model’s misalignment through a fixed benchmark\(Zhanget al\.,[2023](https://arxiv.org/html/2607.00913#bib.bib48); Mazeikaet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib49); Gaboret al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib50)\)\. Wecouldmeasure misalignment on a suite of benchmarks with scores between 0 and 100%\. But suppose that the harm caused by an instance of misaligned behavior increases with timett— since ever\-more capable models can cause bigger harms\(Anthropic,[2026](https://arxiv.org/html/2607.00913#bib.bib56)\)— and so the size of a harm scales by \(for example\)rtr^\{t\}\. Utility \(or rather, disutility\) might be better described by theexpected harm, rather than the frequency of harm, and thusH\(t,ϵ\)=ϵrtH\(t,\\epsilon\)=\\epsilon r^\{t\}may be a more suitable measure of disutility thanϵ\\epsilonalone \(whereϵ\\epsilonis the proportion of the time the model displays misaligned behavior on the benchmark tasks\)\. This metric grows unboundedly for sufficiently largerrrelative to the rate of effective compute growth\. This example again illustrates a broader pattern: benchmark accuracy is bounded and therefore meek, but benchmark accuracy often has closely related transformations with natural interpretations that are unbounded, and therefore possibly mighty\.
### 3\.2Non\-Meek Metrics
We first investigate two examples of power law performance metrics, where multiplicative increases in compute result in \(usually smaller\) multiplicative increases in performance\. In contrast to the power laws in neural scaling laws, these metrics are monotonicallyincreasingin compute\. We also discuss inference time scaling — which is linear in log\-compute — requiring multiplicative changes in compute for linear changes in performance\.
#### 3\.2\.1Reinforcement Learning in Games
Game environments have been fundamental in the development of reinforcement learning and deep learning more generally\(Bellemareet al\.,[2013](https://arxiv.org/html/2607.00913#bib.bib45); Mnihet al\.,[2015](https://arxiv.org/html/2607.00913#bib.bib44)\)\. A robust literature on compute scaling gives ample examples across domains, though typically using an ELO score\. Notably, asNeumann and Gros \([2022](https://arxiv.org/html/2607.00913#bib.bib12)\)points out, ELO is merely the Bradley\-Terry strength on a logarithmic scale, where two Bradley\-Terry scoresγi\\gamma\_\{i\}andγj\\gamma\_\{j\}correspond to a win\-rate for playeriiofγi/\(γi\+γj\)\\gamma\_\{i\}/\(\\gamma\_\{i\}\+\\gamma\_\{j\}\)\.555Bradley\-Terry has a natural interpretation: each player getsγi\\gamma\_\{i\}lottery tickets \(equal to their Bradley\-Terry strength\), and a single random draw determines the winner of the lottery\. Thanks to Toby Ord for mentioning this interpretation and related discussions\.Thus any relationship which relates ELO as log\-linear in compute, in fact, yields a power law in Bradley\-Terry strength\.
Indeed this is exactly what is found in the literature\. In Hex, Pentago, and Connect Four,Jones \([2021](https://arxiv.org/html/2607.00913#bib.bib2)\)andNeumann and Gros \([2022](https://arxiv.org/html/2607.00913#bib.bib12)\)show exponential fits between ELO scores and training compute\. This relationship is also observed byThompsonet al\.\([2022](https://arxiv.org/html/2607.00913#bib.bib31)\)for ELO in historical Chess and Go systems for a mixture of machine learning and classical AI systems\.
Performance in deterministic games with unknown optimal strategies may be unbounded, depending on the game and the possibility of draws\. In a game without draws, for example, the Bradley\-Terry strengthγi\\gamma\_\{i\}could be increased arbitrarily against some fixed reference opponent with strengthγj\\gamma\_\{j\}\. This would make Bradley\-Terry an unbounded metric, even though the win\-rate \(γi/\(γi\+γj\)\\gamma\_\{i\}/\(\\gamma\_\{i\}\+\\gamma\_\{j\}\)\) exists on a separate \(bounded\) scale\. Bradley\-Terry scores are mighty, while win\-rates are meek\.
This fact is quite intuitive\. Consider chess engines: although chess engines may be able to improve indefinitely with exponential compute \(or at least for many, many orders of magnitude\), engines beyond a certain strength already have what they need to best any human\. Though it was once surprising to see a computer beat grandmaster Garry Kasparov, improvements in hardware and algorithms have rendered any budget smartphone equally capable of defeating world\-champions\. With respect to chess win\-rates, meek models have inherited the earth, a fact that now seems unsurprising\.
#### 3\.2\.2Task Horizon Length
Notably deployed by\(Kwaet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib9)\), the task horizon length of a model measures the difficulty of reference tasks by how long completion takes humans, on average\. Across a range of tasks of varying duration, one first fits a curve to predict the model’s probability of success\. After deciding on some threshold success probability \(typically 50%\), one can derive the expected task length that the model will complete\. Over time, the expected duration at a 50% success rate appears to be increasing exponentially\.
Analyses of the relationship between task horizon length and training compute show power\-law relationships \(Whitfillet al\.\([2025](https://arxiv.org/html/2607.00913#bib.bib10)\), although some results are imputed from benchmark scores\)\. Related theoretical work gives a simple mechanism for this pattern\(Sinhaet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib67)\): if a task of lengthnnrequiresnnindependent steps each to succeed with probabilityp=1−ϵp=1\-\\epsilon, then achieving fixed task success probabilityqqrequires
n=log\(q\)log\(1−ϵ\)≈−log\(q\)ϵ\.n=\\frac\{\\log\(q\)\}\{\\log\(1\-\\epsilon\)\}\\approx\\frac\{\-\\log\(q\)\}\{\\epsilon\}\.Hence fixed\-threshold task horizon grows approximately in proportion to the inverse single\-step error rate\. As discussed above, error typically falls as a power law in compute, meaning it falls exponentially in time under exponential compute growth\. So task horizon rises exponentially in time\.
Here again we find that an unbounded mighty metric — task horizon length — is a transformation of a bounded meek metric — the per task error rate\. Depending on circumstances, either metric may be the utility\-relevant one\.
## 4Extending to Inference Time Scaling
We’ve defined meekness with respect to gaps in effectivetrainingcompute\. However, the Meek Models Framework extends naturally to gaps ininference\-time computationbudgets\. Suppose two models have equal effective training compute, but one has an inference token budget growing at a quick exponential rate, as a result of an exponentially growing dollar budget on top of inference efficiency improvements and hardware efficiency improvements\. Meanwhile, the other has a token budget growing at a slower exponential rate, getting the benefit only of inference and hardware efficiency improvements\. Will the performance of the latter model catch up over time?
The existing literature on inference scaling laws mostly uses benchmark accuracy as a metric, which \(as we’ve seen\) is inherently bounded and therefore meek, regardless of how fast benchmark accuracy rises\. In practice, the particular functional forms for benchmark accuracy with respect to inference compute vary, depending on the inference\-scaling technique used\(Villalobos and Atkinson,[2023b](https://arxiv.org/html/2607.00913#bib.bib66)\)\. But in general, a common empirical pattern is that performance improves predictably with additional inference compute over a substantial range before eventually exhibiting diminishing returns\(Brownet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib3); Ellis\-Mohret al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib4)\)\.
For example, take the particular inference scaling technique of repeated sampling with verification\. Benchmark performance using this technique is often well fit by an exponentiated power law,
log\(passi@k\)≈akb,\\log\\left\(\\mathrm\{pass\}\_\{i\}@k\\right\)\\approx ak^\{b\},wherekkis the number of attempts \(i\.e\. the amount of inference compute\)\(Brownet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib3)\)anda,b<0a,b<0are fitted constants\. This is the same functional form as we derived above for benchmark performance as a function of log training compute, just this time in terms of log inference compute\.666Interestingly, one paper argues that the reason we see this functional form in particular depends not just on the fundamental dynamics of test\-time compute scaling, but also on the distribution of question difficulty across multi\-question benchmarks\(Schaefferet al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib6)\)\. \(If all questions were equally difficult, performance would still scale to a bound, but you’d expectexponentiallydecaying performance\.\) This is another reason to be careful when interpreting benchmarks’ scaling trajectories, whether with respect to inference or training compute — especially since the difficulty of questions across benchmarks is not something that benchmark designers typically design deliberately\.Since the pass rate is bounded above by11, it is a meek metric under our definition: even exponential differences in inference expenditure do not generate a permanent gap in this metric\. \(As we saw with training compute, though, the corresponding “march of nines” metric is mighty\.\)
Jones \([2021](https://arxiv.org/html/2607.00913#bib.bib2)\)examines ELO vs perfect play in various games, where ELO eventually plateaus in inference compute\. Note that ELO isn’t constructed in such a way that makes unbounded performance theoretically impossible: inference compute just doesn’t empirically yield continually increasing ELO performance, and so ELO\-with\-respect\-to\-inference\-compute turns out to be a meek metric\. Finally, we see similar behavior so far for performance on METR’s time horizon, where we see sharply diminishing returns with inference scaling\(Ord,[2025](https://arxiv.org/html/2607.00913#bib.bib1)\), even though the metric is theoretically unbounded\.
The preceding discussion in this section has considered two models with equal effective training compute, but different inference compute budgetson a single task\.Two related but different questions are also worth considering:
1. 1\.How do performance gaps on a task change when one model has a growing training compute budgetanda growing inference compute budget?
2. 2\.How does the gap ineconomic returnschange when one actor has an exponentially growing inference compute budget, and uses it to do more economic tasks while holding the amount of inference per task fixed?
We leave both of these questions as avenues for future work: the first because more work is needed to define a joint training\-inference scaling law, the second because this is a fundamentally economic question, beyond our scope here, and the answer will likely vary a great deal by domain\.
## 5Discussion
### 5\.1Limitations
#### 5\.1\.1Positional metrics
In some domains, rewards are positional, where ordinal rankings of performance matter but cardinal positions don’t\. An extreme case of this is winner\-takes\-all competition, such as: a government contract that goes to the model that performs best on some metric\. Here, a real\-world reward is distributed on the basis of the ordinal position of models on an underlying metric\. Positional metrics can’t be modeled with the “meek” framework\. This is because these metrics are functions that take inmultipleplayers’ capabilities\. They are also often non\-differentiable in compute: small compute increases can cause a player to leapfrog a competitor\.
However, the transformation of raw capabilities into positional metrics is in key respects similar to the transformations of bounded metrics\. Just as many of the meek metrics we describe have related metrics that are non\-meek that we might care about \(and vice versa\), many capabilities that have meek metrics have positional metrics where rewards always go to frontier models\. Consider: even when an underlying performance metric is asymptotically bounded, the positional reward might always go to a frontier model with more compute\. This is because, though all models may converge to equivalent performance in the limit, at any particulartt, the frontier model has some \(possibly infinitesimal\) advantage over the meek model\.
That said, this case can be subtle, since often what matters is theexpectedpositional reward ex ante when there is some stochasticity in how performance translates to position\. This can be modeled with win\-rates or ELO, described above\.
#### 5\.1\.2Alternative forms of compute scaling
In our discussion of benchmarks, we’ve assumed that frontier model builders scale investment in compute exponentially\. A reader may wonder if this assumption is reasonable, and if it is consequential for our conclusions\.
##### Should we expect an exponentially growing compute gap?
Historically, frontier model compute scalinghasbeen exponential\(Epoch AI,[2026](https://arxiv.org/html/2607.00913#bib.bib46)\)\. However, the historical pattern may not persist\. Persistence will require revenue to grow sufficiently quickly with compute scale to justify the expenditure\. So far, it has\(Somala,[2025](https://arxiv.org/html/2607.00913#bib.bib47)\)\. But if, at some point, the revenue companies obtain from scaling is bounded, or does not grow sufficiently quickly, thegap in computebetween frontier models and meek models may itself shrink over time\. As a result,revenue generated by AI systemsis a singularly important metric of AI capabilities\. We welcome more efforts to understand how revenue scales with compute investment\.
##### If compute gaps grow sub\-exponentially, would this change our findings?
Theorem[2\.6](https://arxiv.org/html/2607.00913#S2.Thmtheorem6)says that power\-law scaling of compute would not overturn our assessment of meekness in any case\. That said, convergence would be delayed\. And substantially slower\-than\-polynomial scaling could have different results, especially for unbounded metrics\.
Figure 2:Gap between frontier model performance improvements and the imputed meek model performance \(following the same trend more slowly, coinciding with the frontier at GPT\-4\) on two related metrics: the METR time horizon\(Kwaet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib9)\)\(Top\) and the probability of task success at 4\-minute, 1\-hour, and 4\-hour tasks \(Bottom\)\. Both metrics are generated from the same underlying trend, but some lead to a meek outcome and some do not\. These results are stylized estimates, not forecasts\.
#### 5\.1\.3Alternative forms of algorithmic and data progress
##### Scale\-dependent algorithmic and data progress\.
For simplicity, we’ve modeled algorithmic and data progress with the single coefficientgag\_\{a\}\. This functions like a constant multiplier on a model’s pretraining compute\. Recent work\(Gundlachet al\.,[2025a](https://arxiv.org/html/2607.00913#bib.bib52)\)finds that much empirical algorithmic progress is scale\-dependent, increasing the scaling lawexponent, and thereby implying an increasing multiplier on models with more total compute\. Incorporating exponent\-shift algorithmic progress of this sort in our framework would not change our results, since this is equivalent to multiplyinggig\_\{i\}by some factor\.
##### Proprietary algorithmic and data progress\.
We model hardware and algorithmic/data efficiency increases as benefiting all AI developers; this is a simplification\. A frontier AI company may be able to generate hardware, algorithmic or data innovations that they prevent from diffusing to the rest of the industry\(Mertenset al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib37)\)due to proprietary data, algorithmic innovations that are kept secret, or innovations that are specific to the firm’s combinations of models, hardware, and agentic scaffolding\(Gundlachet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib74)\)\.
For this reason, rates of hardware and algorithmic progress for meek models could conceivably slow substantially, attenuating the core mechanism by which meek models catch up to the frontier\. However, as long as the rate of shared, field\-wide algorithmic progress doesn’t goallthe way to zero, the meek framework still applies in the long run\.
#### 5\.1\.4Near\-term Predictions\.
This paper addresses whether metrics will show convergent capabilities in the long term\. But even if capabilities converge in the limit, frontier models may pull ahead in the short\-term\. All meek metrics \(which are non\-zero\) have gaps that first rise and then fall \(some unusual ones may have multiple peaks\)\. Convergence may take a long time, yielding a substantial window where the regulatory landscape looks more like a “mighty” world\. See Fig\.[2](https://arxiv.org/html/2607.00913#S5.F2)for an example of what this can look like in practice, using multiple metrics drawn from METR analysis\(Kwaet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib9)\)\(details in[B](https://arxiv.org/html/2607.00913#A2)\)\.
### 5\.2Governance and Implications
##### Incentives to invest\.
Whether a metric is meek has direct bearing on the private return to compute scaling\. For meek metrics, exponential investment yields only a transient advantage: a frontier developer who outspends competitors by orders of magnitude can expect that lead to erode as hardware and algorithmic progress lift the meek competitor’s compute budget along the same curve\. For mighty metrics, by contrast, the return to scaling is durable, and investment is a moat\. The commercial case for continued exponential compute growth therefore depends on which metrics frontier firms can monetize and what the returns on improved performance are — a question that is itself contested and may differ across domains\.
##### Concentration of power\.
When unbounded performance is valuable, owners of compute capital can entrench their advantage: firms or nations who can sustain exponential compute expenditure can consistently control capabilities that smaller players can’t access\. Those actors may have the ability to dictate terms of use, restrict harmful or privately disadvantageous applications of frontier capabilities, and leverage their capabilities to capture private rents\. When the performance metrics we care about are bounded, we see the opposite dynamic: a broadened set of actors who can operate at the frontier in the long run\.
The overall future is likely mixed: some sectors concentrated, others proliferated\.
##### Compute controls and national advantages\.
Export controls on advanced compute\(O’Garaet al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib27)\)presume that restricting a rival’s effective compute will durably restrict its capabilities\. If the capabilities in question are ones where unbounded performance matters, compute controls can indeed sustain a durable gap\. But for capabilities where performance is bounded when measured in the terms we care about — that is, cases where meek metrics are apt — the presumption that compute restrictions lead to persistent capabilities gaps fails in the long run\. Provided the restricted actor still enjoys some exponential rate of hardware and algorithmic progress \(even if it’s a slower exponential rate than the frontier actor enjoys\), its capabilities on meek metrics will eventually catch up\.Compute controls then function as a delay rather than a permanent ceiling on an adversary’s capabilities\.
##### Dangerous capabilities\.
For some dangerous capabilities, the relevant social harm may be largely realized once a model crosses a fixed capability threshold: for example, reliably helping a user reproduce a software vulnerability, automate a phishing campaign, synthesize a dangerous protocol from dispersed biological information, or guide a novice through key steps in a CBRN\-relevant workflow\(Frontier Model Forum,[2025](https://arxiv.org/html/2607.00913#bib.bib59); Moutonet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib60); OpenAI,[2024](https://arxiv.org/html/2607.00913#bib.bib61); Peppinet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib62); CyberGym authors,[2026](https://arxiv.org/html/2607.00913#bib.bib63); Anthropic,[2025](https://arxiv.org/html/2607.00913#bib.bib64)\)\.In these cases, marginal improvements above the threshold may matter much less than whether the threshold is crossed at all\.If hardware and algorithmic progress continue to raise the effective compute available to fixed\-budget developers, then the meek\-models result has a direct governance implication: even if only frontier developers can access the dangerous capability today, many meek model builders may eventually access substantially the same threshold capability\. Ifanysuch capabilities impose sufficiently high existential risks, this may be unacceptable\(Jones,[2024](https://arxiv.org/html/2607.00913#bib.bib65)\)\.
This does not mean that every dangerous capability is best understood as meek\. Some misuse\-relevant quantities may be unbounded or positional: the number of targets that can be attacked, the speed of exploitation, the ability to adapt to active defenders\. In particular, societal safety often depends on an offense\-defense balance\.777Offense and defense need not be measured on the same performance scale\. A model that marginally improves vulnerability discovery may help attackers, while a different model that improves patch generation, intrusion detection, or incident response may help defenders; the ultimate outcome depends on how these twodifferenttypes of performance interact\. The meek framework can help ask whether a particular underlying capability — like cyber knowledge — will diffuse to meek models, but it does not by itself determine whether diffusion favors attackers or defenders\.
##### Alignment\.
If meek models inherit the earth, the future is shaped not by one or a few aligned frontier systems but by a population of many roughly frontier\-capable models\. Which statistic of that population matters—average alignment, maximum alignment, or minimum alignment among any accessible system—depends on the threat model\. On pessimistic views in which a single misaligned frontier system suffices for catastrophe, proliferation of frontier\-level capabilities is alarming\(Hammondet al\.,[2025](https://arxiv.org/html/2607.00913#bib.bib71); Bostrom,[2019](https://arxiv.org/html/2607.00913#bib.bib70)\)\. On more optimistic views in which defense aggregates across aligned systems, and average alignment is what matters, proliferation may be protective\.
#### 5\.2\.1A call for utility\-aware analysis\.
Empirical claims about convergence or divergence in AI capabilities cannot be naively read off a single metric\. The meekness of a metric is a property of the construction of the metric, and closely\-related metrics can differ in meekness despite measuring ostensibly similar phenomena\. For example, on one metric, we might see the gap in capabilities between the US and China growing, and on another similar metric see it shrinking; which metric is “right” is a question about how utility changes with respect to a measured increase in capabilities, but metrics are often designed without this question in mind\. Analysts, economists, and policymakers should therefore use caution when using existing metrics to draw conclusions about how capabilities gaps are changing\. And they should treat the construction and selection of a performance metric as a substantive decision, one that necessarily implies a point of view on how AI capabilities matter\. Finally: a better understanding of which metrics most faithfully capture capabilities of economic and social value — and how models scale on those metrics — is essential for long\-term AI policy\.
## Impact Statement
The authors believe AI capabilities are of fundamental importance to effective governance, geopolitical strategy, scientific progress, economic prosperity, and global welfare\. The current understanding of these capabilities relies heavily on the metrics discussed in this paper\. Therefore, we see careful analysis of these metrics — and their relation to compute and capital — as essential for informing public discourse, improving national security, and guiding policy decisions\.
## LLM Usage Statement
The authors of this paper used LLMs for generating graphics, styling the text, validating the mathematical results, guidance during mathematical derivations, and literature review\.
## References
- AI Index Steering Committee \(2025\)Chapter 2: technical performance\.InArtificial Intelligence Index Report 2025,Stanford University\.External Links:[Link](https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter2_final.pdf)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p4.1)\.
- Anthropic \(2025\)Why do we take LLMs seriously as a potential source of biorisk?\.Note:Published September 5, 2025External Links:[Link](https://red.anthropic.com/2025/biorisk/)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- Anthropic \(2026\)Claude Mythos preview system card\.System CardAnthropic\.Note:Accessed: 2026\-04\-24External Links:[Link](https://www.anthropic.com/claude-mythos-preview-system-card)Cited by:[§3\.1\.3](https://arxiv.org/html/2607.00913#S3.SS1.SSS3.p1.6)\.
- Y\. Bahri, E\. Dyer, J\. Kaplan, J\. Lee, and U\. Sharma \(2024\)Explaining neural scaling laws\.Proceedings of the National Academy of Sciences121\(27\)\.External Links:ISSN 1091\-6490,[Link](http://dx.doi.org/10.1073/pnas.2311878121),[Document](https://dx.doi.org/10.1073/pnas.2311878121)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1)\.
- M\. G\. Bellemare, Y\. Naddaf, J\. Veness, and M\. Bowling \(2013\)The arcade learning environment: An evaluation platform for general agents\.Journal of Artificial Intelligence Research47,pp\. 253–279\.Cited by:[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p1.4)\.
- N\. Bostrom \(2019\)The vulnerable world hypothesis\.Global Policy10\(4\),pp\. 455–476\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/1758-5899.12718),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/1758-5899.12718),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/1758\-5899\.12718Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px5.p1.1)\.
- B\. Brown, J\. Juravsky, R\. Ehrlich, R\. Clark, Q\. V\. Le, C\. Ré, and A\. Mirhoseini \(2024\)Large language monkeys: scaling inference compute with repeated sampling\.arXiv preprint arXiv:2407\.21787\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.21787)Cited by:[§4](https://arxiv.org/html/2607.00913#S4.p2.1),[§4](https://arxiv.org/html/2607.00913#S4.p3.3)\.
- B\. Cottier, J\. You, N\. Martemianova, and D\. Owen \(2024\)How far behind are open models?\.Note:Accessed: 2026\-04\-24External Links:[Link](https://epoch.ai/blog/open-models-report)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p3.1)\.
- CyberGym authors \(2026\)CyberGym: Evaluating AI agents’ real\-world cybersecurity capabilities at scale\.InInternational Conference on Learning Representations,Note:Author list to be verified from the final ICLR proceedings entryExternal Links:2506\.02548,[Link](https://arxiv.org/abs/2506.02548)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- E\. Del Sozzo, M\. Fleming, K\. Flamm, and N\. Thompson \(2026\)How much progress has there been in NVIDIA datacenter GPUs?\.arXiv preprint arXiv:2601\.20115\.Cited by:[Appendix B](https://arxiv.org/html/2607.00913#A2.p2.9)\.
- J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-Fei \(2009\)ImageNet: A large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,pp\. 248–255\.Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p1.2)\.
- A\. R\. Ellis\-Mohr, A\. K\. Nayak, and L\. R\. Varshney \(2025\)A theory of inference compute scaling: reasoning through directed stochastic skill search\.arXiv preprint arXiv:2507\.00004\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.00004)Cited by:[§4](https://arxiv.org/html/2607.00913#S4.p2.1)\.
- L\. Emberson \(2025\)Epoch AI\.External Links:[Link](https://epoch.ai/data-insights/open-weights-vs-closed-weights-models)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p3.1),[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p4.1)\.
- Epoch AI \(2026\)Data on AI models\.Note:Accessed: 23 Apr 2026External Links:[Link](https://epoch.ai/data/ai-models)Cited by:[§5\.1\.2](https://arxiv.org/html/2607.00913#S5.SS1.SSS2.Px1.p1.1)\.
- Frontier Model Forum \(2025\)Frontier capability assessments: Technical report\.Technical reportFrontier Model Forum\.Note:Implementing Frontier AI Safety Frameworks report seriesExternal Links:[Link](https://www.frontiermodelforum.org/uploads/2025/04/FMF-PDF-Frontier-Capability-Assessments_-Technical-Report.pdf)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- J\. Gabor, J\. Lynch, and J\. Rosenfeld \(2025\)EvilGenie: A reward hacking benchmark\.External Links:2511\.21654,[Link](https://arxiv.org/abs/2511.21654)Cited by:[§3\.1\.3](https://arxiv.org/html/2607.00913#S3.SS1.SSS3.p1.6)\.
- H\. Gundlach, Z\. A\. Brown, J\. Lynch, and N\. Thompson \(2026\)Just a wrapper? How much do scaffolds matter?\.Note:[https://mitfuturetech\.substack\.com/p/just\-a\-wrapper\-how\-much\-do\-scaffolds](https://mitfuturetech.substack.com/p/just-a-wrapper-how-much-do-scaffolds)Mixture of Experts \(MIT FutureTech\), Substack, accessed July 1, 2026Cited by:[§5\.1\.3](https://arxiv.org/html/2607.00913#S5.SS1.SSS3.Px2.p1.1)\.
- H\. Gundlach, A\. Fogelson, J\. Lynch, A\. Trisovic, J\. Rosenfeld, A\. Sandhu, and N\. Thompson \(2025a\)On the origin of algorithmic progress in AI\.External Links:2511\.21622,[Document](https://dx.doi.org/10.48550/arXiv.2511.21622),[Link](https://arxiv.org/abs/2511.21622)Cited by:[§5\.1\.3](https://arxiv.org/html/2607.00913#S5.SS1.SSS3.Px1.p1.2)\.
- H\. Gundlach, J\. Lynch, and N\. Thompson \(2025b\)Meek models shall inherit the earth\.arXiv:2507\.07931\.External Links:2507\.07931,[Link](https://arxiv.org/abs/2507.07931)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p5.1),[§1\.3](https://arxiv.org/html/2607.00913#S1.SS3.p3.1),[§2\.1](https://arxiv.org/html/2607.00913#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2607.00913#S2.SS3.p2.1),[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p2.2)\.
- L\. Hammond, A\. Chan, J\. Clifton, J\. Hoelscher\-Obermaier, A\. Khan, E\. McLean, C\. Smith, W\. Barfuss, J\. Foerster, T\. Gavenčiak, T\. A\. Han, E\. Hughes, V\. Kovařík, J\. Kulveit, J\. Z\. Leibo, C\. Oesterheld, C\. S\. de Witt, N\. Shah, M\. Wellman, P\. Bova, T\. Cimpeanu, C\. Ezell, Q\. Feuillade\-Montixi, M\. Franklin, E\. Kran, I\. Krawczuk, M\. Lamparth, N\. Lauffer, A\. Meinke, S\. Motwani, A\. Reuel, V\. Conitzer, M\. Dennis, I\. Gabriel, A\. Gleave, G\. Hadfield, N\. Haghtalab, A\. Kasirzadeh, S\. Krier, K\. Larson, J\. Lehman, D\. C\. Parkes, G\. Piliouras, and I\. Rahwan \(2025\)Multi\-agent risks from advanced AI\.External Links:2502\.14143,[Link](https://arxiv.org/abs/2502.14143)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px5.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2020\)Measuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.External Links:2009\.03300Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p3.1)\.
- A\. Ho, T\. Besiroglu, E\. Erdil, D\. Owen, R\. Rahman, Z\. C\. Guo, D\. Atkinson, N\. Thompson, and J\. Sevilla \(2024\)Algorithmic progress in language models\.External Links:2403\.05812,[Link](https://arxiv.org/abs/2403.05812)Cited by:[Appendix B](https://arxiv.org/html/2607.00913#A2.p2.9)\.
- A\. Ho, J\. Denain, D\. Atanasov, S\. Albanie, and R\. Shah \(2025\)A Rosetta Stone for AI Benchmarks\.arXiv preprint arXiv:2512\.00193\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.00193),[Link](https://arxiv.org/abs/2512.00193)Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p9.1)\.
- J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, J\. W\. Rae, O\. Vinyals, and L\. Sifre \(2022\)Training compute\-optimal large language models\.External Links:2203\.15556,[Link](https://arxiv.org/abs/2203.15556)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1),[§3\.1\.1](https://arxiv.org/html/2607.00913#S3.SS1.SSS1.p1.1),[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p8.2)\.
- H\. T\. Ihle \(2025\)LessWrong\.External Links:[Link](https://www.lesswrong.com/posts/NLnGRDRXATW2pqXuE/is-the-gap-between-open-and-closed-models-growing-evidence)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p4.1)\.
- A\. L\. Jones \(2021\)Scaling scaling laws with board games\.arXiv preprint arXiv:2104\.03113\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2104.03113)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1),[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p2.1),[§4](https://arxiv.org/html/2607.00913#S4.p4.1)\.
- C\. I\. Jones \(2024\)The AI dilemma: Growth versus existential risk\.American Economic Review: Insights6\(4\),pp\. 575–90\.External Links:[Document](https://dx.doi.org/10.1257/aeri.20230570),[Link](https://www.aeaweb.org/articles?id=10.1257/aeri.20230570)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- I\. Juniewicz \(2026\)Hyperscaler capex has quadrupled since GPT\-4’s release\.Note:Accessed: 2026\-06\-10External Links:[Link](https://epoch.ai/data-insights/hyperscaler-capex-trend)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.External Links:2001\.08361,[Link](https://arxiv.org/abs/2001.08361)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1),[§3\.1\.1](https://arxiv.org/html/2607.00913#S3.SS1.SSS1.p1.1)\.
- A\. Krizhevsky \(2009\)Learning multiple layers of features from tiny images\.Technical reportUniversity of Toronto\.Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p1.2)\.
- T\. Kwa, B\. West, J\. Becker, A\. Deng, K\. Garcia, M\. Hasin, S\. Jawhar, M\. Kinniment, N\. Rush, S\. Von Arx, R\. Bloom, T\. Broadley, H\. Du, B\. Goodrich, N\. Jurkovic, L\. H\. Miles, S\. Nix, T\. Lin, C\. Painter, N\. Parikh, D\. Rein, L\. J\. K\. Sato, H\. Wijk, D\. M\. Ziegler, E\. Barnes, and L\. Chan \(2026\)Measuring AI ability to complete long software tasks\.arXiv preprint arXiv:2503\.14499\.Cited by:[Appendix B](https://arxiv.org/html/2607.00913#A2.p1.5),[§3\.2\.2](https://arxiv.org/html/2607.00913#S3.SS2.SSS2.p1.1),[Figure 2](https://arxiv.org/html/2607.00913#S5.F2),[Figure 2](https://arxiv.org/html/2607.00913#S5.F2.3.2),[§5\.1\.4](https://arxiv.org/html/2607.00913#S5.SS1.SSS4.p1.1)\.
- Y\. LeCun, L\. Bottou, Y\. Bengio, and P\. Haffner \(1998\)Gradient\-based learning applied to document recognition\.Proceedings of the IEEE86\(11\),pp\. 2278–2324\.Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p1.2)\.
- M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang, N\. Mu, E\. Sakhaee, N\. Li, S\. Basart, B\. Li, D\. Forsyth, and D\. Hendrycks \(2024\)HarmBench: A standardized evaluation framework for automated red teaming and robust refusal\.InProceedings of the 41st International Conference on Machine Learning,External Links:[Link](https://arxiv.org/abs/2402.04249)Cited by:[§3\.1\.3](https://arxiv.org/html/2607.00913#S3.SS1.SSS3.p1.6)\.
- M\. Mertens, N\. Fischl\-Lanzoni, and N\. Thompson \(2026\)Is there “secret sauce” in large language model development?\.arXiv preprint arXiv:2602\.07238\.External Links:[Link](https://arxiv.org/abs/2602.07238)Cited by:[§5\.1\.3](https://arxiv.org/html/2607.00913#S5.SS1.SSS3.Px2.p1.1)\.
- V\. Mnih, K\. Kavukcuoglu, D\. Silver, A\. A\. Rusu, J\. Veness, M\. G\. Bellemare, A\. Graves, M\. Riedmiller, A\. K\. Fidjeland, G\. Ostrovski, S\. Petersen, C\. Beattie, A\. Sadik, I\. Antonoglou, H\. King, D\. Kumaran, D\. Wierstra, S\. Legg, and D\. Hassabis \(2015\)Human\-level control through deep reinforcement learning\.Nature518\(7540\),pp\. 529–533\.Cited by:[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p1.4)\.
- C\. A\. Mouton, C\. Lucas, and E\. Guest \(2024\)The operational risks of AI in large\-scale biological attacks: results of a red\-team study\.Technical reportTechnical ReportRR\-A2977\-2,RAND Corporation\.External Links:[Document](https://dx.doi.org/10.7249/RRA2977-2),[Link](https://www.rand.org/pubs/research_reports/RRA2977-2.html)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- O\. Neumann and C\. Gros \(2022\)Scaling laws for a multi\-agent reinforcement learning model\.arXiv preprint arXiv:2210\.00849\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2210.00849)Cited by:[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p1.4),[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p2.1)\.
- A\. O’Gara, G\. Kulp, W\. Hodgkins, J\. Petrie, V\. Immler, A\. Aysu, K\. Basu, S\. Bhasin, S\. Picek, and A\. Srivastava \(2025\)Hardware\-enabled mechanisms for verifying responsible AI development\.External Links:2505\.03742,[Link](https://arxiv.org/abs/2505.03742)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p3.1),[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px3.p1.1)\.
- OpenAI \(2024\)Building an early warning system for LLM\-aided biological threat creation\.Note:Published January 31, 2024External Links:[Link](https://openai.com/index/building-an-early-warning-system-for-llm-aided-biological-threat-creation/)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- T\. Ord \(2025\)External Links:[Link](https://www.tobyord.com/writing/hourly-costs-for-ai-agents)Cited by:[§4](https://arxiv.org/html/2607.00913#S4.p4.1)\.
- D\. Owen \(2024\)How predictable is language model benchmark performance?\.arXiv preprint arXiv:2401\.04757\.External Links:[Link](https://arxiv.org/abs/2401.04757)Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p3.1)\.
- A\. Peppin, A\. Reuel, S\. Casper, E\. Jones, A\. Strait, U\. Anwar, A\. Agrawal, S\. Kapoor, S\. Koyejo, M\. Pellat, R\. Bommasani, N\. Frosst, and S\. Hooker \(2024\)The reality of AI and biorisk\.External Links:2412\.01946,[Document](https://dx.doi.org/10.48550/arXiv.2412.01946),[Link](https://arxiv.org/abs/2412.01946)Cited by:[§5\.2](https://arxiv.org/html/2607.00913#S5.SS2.SSS0.Px4.p1.1)\.
- A\. Power, Y\. Burda, H\. Edwards, I\. Babuschkin, and V\. Misra \(2022\)Grokking: generalization beyond overfitting on small algorithmic datasets\.arXiv preprint arXiv:2201\.02177\.Cited by:[Appendix A](https://arxiv.org/html/2607.00913#A1.p1.1),[§2\.2](https://arxiv.org/html/2607.00913#S2.SS2.p2.1)\.
- J\. S\. Rosenfeld \(2021\)Scaling laws for deep learning\.arXiv preprint arXiv:2108\.07686\.Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1),[§3\.1\.1](https://arxiv.org/html/2607.00913#S3.SS1.SSS1.p1.1)\.
- Y\. Ruan, C\. J\. Maddison, and T\. Hashimoto \(2024\)Observational scaling laws and the predictability of language model performance\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/1cded4f97cf5f01a284c574110b7e3b9-Paper-Conference.pdf)Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p3.1)\.
- R\. Schaeffer, J\. Kazdan, J\. Hughes, J\. Juravsky, S\. Price, A\. Lynch, E\. Jones, R\. Kirk, A\. Mirhoseini, and S\. Koyejo \(2025\)How do large language monkeys get their power \(laws\)?\.arXiv preprint arXiv:2502\.17578\.Cited by:[footnote 6](https://arxiv.org/html/2607.00913#footnote6)\.
- R\. Schaeffer, B\. Miranda, and S\. Koyejo \(2023\)Are emergent abilities of large language models a mirage?\.arXiv preprint arXiv:2304\.15004\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2304.15004),[Link](https://arxiv.org/abs/2304.15004)Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p4.4)\.
- J\. Sevilla, L\. Heim, A\. Ho, T\. Besiroglu, M\. Hobbhahn, and P\. Villalobos \(2022\)Compute trends across three eras of machine learning\.In2022 International Joint Conference on Neural Networks \(IJCNN\),External Links:[Link](http://dx.doi.org/10.1109/IJCNN55064.2022.9891914),[Document](https://dx.doi.org/10.1109/ijcnn55064.2022.9891914)Cited by:[Appendix B](https://arxiv.org/html/2607.00913#A2.p2.9),[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1)\.
- A\. Sinha, A\. Arun, S\. Goel, S\. Staab, and J\. Geiping \(2026\)The illusion of diminishing returns: Measuring long horizon execution in LLMs\.InThe Fourteenth International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=3lm8lWYxiq)Cited by:[§3\.2\.2](https://arxiv.org/html/2607.00913#S3.SS2.SSS2.p2.4)\.
- V\. Somala \(2025\)OpenAI’s revenue has been growing 3x a year since 2024\.Note:Accessed: 2026\-04\-23External Links:[Link](https://epoch.ai/data-insights/openai-revenue)Cited by:[§5\.1\.2](https://arxiv.org/html/2607.00913#S5.SS1.SSS2.Px1.p1.1)\.
- A\. Srivastava, A\. Rastogi, A\. Rao, A\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso, A\. Kluska, A\. Zhou,et al\.\(2022\)Beyond the imitation game: Quantifying and extrapolating the capabilities of language models\.arXiv preprint arXiv:2206\.04615\.External Links:2206\.04615Cited by:[§3\.1\.2](https://arxiv.org/html/2607.00913#S3.SS1.SSS2.p3.1)\.
- N\. C\. Thompson, S\. Ge, and G\. F\. Manso \(2022\)The importance of \(exponentially more\) computing power\.arXiv preprint arXiv:2206\.14007\.Cited by:[§3\.2\.1](https://arxiv.org/html/2607.00913#S3.SS2.SSS1.p2.1)\.
- P\. Villalobos and D\. Atkinson \(2023a\)Trading off compute in training and inference\.Note:Accessed: 2024\-07\-24External Links:[Link](https://epochai.org/blog/trading-off-compute-in-training-and-inference)Cited by:[§1\.1](https://arxiv.org/html/2607.00913#S1.SS1.p1.1)\.
- P\. Villalobos and D\. Atkinson \(2023b\)Trading off compute in training and inference\.Note:Accessed: 2026\-04\-24External Links:[Link](https://epoch.ai/blog/trading-off-compute-in-training-and-inference)Cited by:[§4](https://arxiv.org/html/2607.00913#S4.p2.1)\.
- P\. Whitfill, B\. Snodin, and J\. Becker \(2025\)Forecasting AI time horizon under compute slowdowns\.arXiv preprint arXiv:2511\.19492\.Cited by:[§3\.2\.2](https://arxiv.org/html/2607.00913#S3.SS2.SSS2.p2.4)\.
- Z\. Zhang, L\. Lei, L\. Wu, R\. Sun, Y\. Huang, C\. Long, X\. Liu, X\. Lei, J\. Tang, and M\. Huang \(2023\)SafetyBench: Evaluating the safety of large language models\.External Links:2309\.07045,[Link](https://arxiv.org/abs/2309.07045)Cited by:[§3\.1\.3](https://arxiv.org/html/2607.00913#S3.SS1.SSS3.p1.6)\.
## Appendix AFormal Definitions and Proofs
###### Definition A\.1\(Normal Performance Metric\)\.
IfP:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}is both differentiable and weakly monotonically increasing, we call it anormal performance metric\.
These criteria are natural: \(1\) Performance growth is typically smooth, and even when performance exhibits sharp jumps, such transitions are still smooth\(Poweret al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib8)\), \(2\) although adding compute can diminish performance \(e\.g\. overfitting\), since we only require weak monotonicity, any model which degrades in performance can simply be discarded\.
###### Definition A\.2\(Meek Metric\)\.
LetP:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}\. We sayPPis a meek metric if for all pairs of exponentsb\>a\>1b\>a\>1, and for all initial values of computeC0\>0C\_\{0\}\>0, the following limit holds:
limt→∞P\(btC0\)−P\(atC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)=0
Intuitively,PPis invariant to exponential differences in effective compute, whereaacomes from hardware/algorithmic efficiency, andbbadds exponential increases in investment\. Given some initial compute investmentC0C\_\{0\}, two actors who see exponential differences in the growth of that effective compute will see no durable difference in performance over time\.
###### Theorem A\.3\(Equivalence Under Power Law Scaling\)\.
LetP:ℝ\+→ℝP:\\mathbb\{R\}^\{\+\}\\rightarrow\\mathbb\{R\}\. Then the following are equivalent:
1. 1\.For all exponentsb\>a\>1b\>a\>1and computeC0\>0C\_\{0\}\>0,limt→∞P\(btC0\)−P\(atC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)=0
2. 2\.For all powersb\>a\>0b\>a\>0and computeC0\>0C\_\{0\}\>0,limt→∞P\(tbC0\)−P\(taC0\)=0\\lim\_\{t\\rightarrow\\infty\}P\(t^\{b\}C\_\{0\}\)\-P\(t^\{a\}C\_\{0\}\)=0
###### Proof\.
The proof is an exercise in reparameterization\.
To show\(1\)⟹\(2\)\(1\)\\implies\(2\), lets=logts=\\log t\. \(Wherelog\\logis the natural log throughout this appendix\.\) For some powersb\>a\>0b\>a\>0, we have that
P\(tbC0\)−P\(taC0\)\\displaystyle P\(t^\{b\}C\_\{0\}\)\-P\(t^\{a\}C\_\{0\}\)=P\(\(eb\)sC0\)−P\(\(ea\)sC0\)\\displaystyle=P\(\(e^\{b\}\)^\{s\}C\_\{0\}\)\-P\(\(e^\{a\}\)^\{s\}C\_\{0\}\)Sinceeb\>ea\>1e^\{b\}\>e^\{a\}\>1, the limit ast→∞t\\rightarrow\\infty\(and equivalentlys→∞s\\rightarrow\\infty\) is zero by\(1\)\(1\)\. To show\(2\)⟹\(1\)\(2\)\\implies\(1\), for some exponentsb\>a\>1b\>a\>1, instead lets=ets=e^\{t\}, so we have
P\(btC0\)−P\(atC0\)\\displaystyle P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)=P\(\(elogb\)tC0\)−P\(\(eloga\)tC0\)\\displaystyle=P\(\(e^\{\\log b\}\)^\{t\}C\_\{0\}\)\-P\(\(e^\{\\log a\}\)^\{t\}C\_\{0\}\)=P\(slogbC0\)−P\(slogaC0\)\\displaystyle=P\(s^\{\\log b\}C\_\{0\}\)\-P\(s^\{\\log a\}C\_\{0\}\)Sincelogb\>loga\>0\\log b\>\\log a\>0, the limit ass→∞s\\rightarrow\\infty\(and equivalentlyt→∞t\\rightarrow\\infty\) is zero by\(2\)\(2\)\.
∎
###### Theorem A\.4\.
IfPPis a normal performance metric which is bounded above, thenPPis a meek metric\.
###### Proof\.
LetPPbe a normal performance metric with least upper boundMM\. Letfx\(t\)=P\(xtC0\)f\_\{x\}\(t\)=P\(x^\{t\}C\_\{0\}\)\. SincebtC0b^\{t\}C\_\{0\}andatC0a^\{t\}C\_\{0\}are monotonically increasing intt, as isPP, the functionsfa\(t\)f\_\{a\}\(t\)andfb\(t\)f\_\{b\}\(t\)are monotonically increasing and bounded above byMM, and therefore both converge toMMast→∞t\\rightarrow\\inftyby the Monotone Convergence Theorem\. ∎
###### Theorem A\.5\.
LetP\(C\)P\(C\)be a monotonically increasing, differentiable, unbounded function\. Moreover, suppose the derivative ofPPwith respect tologlogC\\log\\log\{C\}is eventually monotone\. ThenPPis a meek metric if and only if
limC→∞P\(C\)loglogC=0\\lim\_\{C\\rightarrow\\infty\}\\frac\{P\(C\)\}\{\\log\\log\{C\}\}=0
###### Proof\.
We’ll show two results, which together give the result:
1. 1\.LetD\(C\)D\(C\)be the derivative ofP\(C\)P\(C\)with respect tologlogC\\log\\log C\. Then P\(C\)loglogC→0⇔D\(C\)→0\\frac\{P\(C\)\}\{\\log\\log C\}\\rightarrow 0\\iff D\(C\)\\rightarrow 0
2. 2\.For some arbitrary exponentsb\>a\>1b\>a\>1and some initial computeC0C\_\{0\}, letΔ\(x\)=P\(btC0\)−P\(atC0\)\\Delta\(x\)=P\(b^\{t\}C\_\{0\}\)\-P\(a^\{t\}C\_\{0\}\)— the relevant difference in the meek metric criteria\. IfD\(x\)D\(x\)is eventuallyincreasing, then for some functionIIsuch thatlimx→∞I\(x\)=logλ\\lim\_\{x\\rightarrow\\infty\}I\(x\)=\\log\\lambda, we have the following bounds\. D\(x\)I\(x\)≤Δ\(x\)≤D\(kxλ\)I\(x\)D\(x\)I\(x\)\\leq\\Delta\(x\)\\leq D\(kx^\{\\lambda\}\)I\(x\)whereλ\\lambdaandkkare non\-zero constants which depend onbb,aa, andC0C\_\{0\}\. IfD\(x\)D\(x\)is eventuallydecreasing, the inequalities are flipped\.
From \(2\), we can see that
Δ\(x\)\\Delta\(x\)goes to zero if and only if
D\(x\)→0D\(x\)\\rightarrow 0since
logλ\\log\\lambdais non\-zero, and therefore if and only if
P\(C\)/loglogC→0P\(C\)/\\log\\log\{C\}\\rightarrow 0by \(1\)\.
Proof of Result 1:First note thatP\(C\)→∞P\(C\)\\rightarrow\\infty\(by assumption\) andloglogC→∞\\log\\log C\\rightarrow\\inftyasC→∞C\\rightarrow\\infty\. SinceD\(C\)D\(C\)is monotonic, the extended limit certainly exists\. If the limit is finite, we can invoke L’Hopital’s rule to show thatP\(C\)/loglogCP\(C\)/\\log\\log CandD\(C\)D\(C\)share a limit:
=\\displaystyle=limC→∞P\(C\)loglogC\\displaystyle\\lim\_\{C\\rightarrow\\infty\}\\frac\{P\(C\)\}\{\\log\\log C\}=LH\\displaystyle\\overset\{\\text\{LH\}\}\{=\}limC→∞dP/dCd\(loglogC\)/dC\\displaystyle\\lim\_\{C\\rightarrow\\infty\}\\frac\{dP/dC\}\{d\(\\log\\log C\)/dC\}=\\displaystyle=limC→∞dPdCdCd\(loglogC\)\\displaystyle\\lim\_\{C\\rightarrow\\infty\}\\frac\{dP\}\{dC\}\\frac\{dC\}\{d\(\\log\\log C\)\}=\\displaystyle=limC→∞dPd\(loglogC\)\\displaystyle\\lim\_\{C\\rightarrow\\infty\}\\frac\{dP\}\{d\(\\log\\log C\)\}=\\displaystyle=limC→∞D\(C\)\\displaystyle\\lim\_\{C\\rightarrow\\infty\}D\(C\)If the limit ofD\(C\)D\(C\)is infinite, we instead showP\(C\)/loglogCP\(C\)/\\log\\log Cgoes to infinity\. First, for allM\>0M\>0, we knowD\(C\)≥MD\(C\)\\geq Mfor large enoughCC, and thus\(dP/dC\)≥M/\(ClogC\)\(dP/dC\)\\geq M/\(C\\log C\)\. Choosing some small fixedc0c\_\{0\}, we can integratedP/dCdP/dCto get:
P\(C\)−P\(c0\)=∫c0CdPdu𝑑u≥∫c0CMulogu𝑑u=M\(loglogC−loglogc0\)P\(C\)\-P\(c\_\{0\}\)=\\int\_\{c\_\{0\}\}^\{C\}\\frac\{dP\}\{du\}du\\geq\\int\_\{c\_\{0\}\}^\{C\}\\frac\{M\}\{u\\log u\}du=M\(\\log\\log C\-\\log\\log c\_\{0\}\)Rearranging, we get
P\(C\)loglogC≥M\+𝒪\(1\)loglogC\\frac\{P\(C\)\}\{\\log\\log C\}\\geq M\+\\frac\{\\mathcal\{O\}\(1\)\}\{\\log\\log C\}So clearlylim infC→∞P\(C\)/loglogC≥M\\liminf\_\{C\\rightarrow\\infty\}P\(C\)/\\log\\log C\\geq M, and sinceMMis arbitrary, it is unbounded likeD\(C\)D\(C\)\.
Proof of Result 2:We first re\-parameterize the second argument in the meek metric criteria,atC0a^\{t\}C\_\{0\}, asxx\. Then we define two constantsλ=logab\>1\\lambda=\\log\_\{a\}b\>1andk=C01−λk=C\_\{0\}^\{1\-\\lambda\}, such that we expressbtC0b^\{t\}C\_\{0\}askxλkx^\{\\lambda\}\. We re\-write the meek metric difference,Δ\(x\)\\Delta\(x\), using these parameters:
Δ\(x\)=P\(kxλ\)−P\(x\)\\Delta\(x\)=P\(kx^\{\\lambda\}\)\-P\(x\)
To boundΔ\(x\)\\Delta\(x\), we write it in its integral form, using the relationD\(C\)=ClogCdP\(C\)dCD\(C\)=C\\log C\\frac\{dP\(C\)\}\{dC\}:
limx→∞Δ\(x\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\Delta\(x\)=\\displaystyle=limx→∞P\(kxλ\)−P\(x\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}P\(kx^\{\\lambda\}\)\-P\(x\)=∗\\displaystyle\\overset\{\*\}\{=\}limx→∞∫xkxλdPdu𝑑u\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{dP\}\{du\}du=\\displaystyle=limx→∞∫xkxλ\(ulogu\)dPdu\(ulogu\)𝑑u\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{\(u\\log\{u\}\)\\frac\{dP\}\{du\}\}\{\(u\\log\{u\}\)\}du=\\displaystyle=limx→∞∫xkxλD\(u\)\(ulogu\)𝑑u\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{D\(u\)\}\{\(u\\log\{u\}\)\}duThe use of the fundamental theorem of calculus is justified sinceDDis monotone and bounded on the interval for sufficiently largexx, thus Riemann integrable, and since1/\(ulogu\)1/\(u\\log u\)is continuous on the interval and thus Riemann integrable,P′P^\{\\prime\}is too as the product of two Riemann integrable functions\. Without loss of generality, assumeD\(u\)D\(u\)is eventually monotonicallyincreasing\. We can bound the integral as:
D\(x\)∫xkxλ1\(ulogu\)𝑑u⏟I\(x\)≤∫xkxλD\(u\)\(ulogu\)𝑑u≤D\(kxλ\)∫xkxλ1\(ulogu\)𝑑u⏟I\(x\)D\(x\)\\underbrace\{\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{1\}\{\(u\\log u\)\}du\}\_\{I\(x\)\}\\leq\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{D\(u\)\}\{\(u\\log\{u\}\)\}du\\leq D\(kx^\{\\lambda\}\)\\underbrace\{\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{1\}\{\(u\\log u\)\}du\}\_\{I\(x\)\}
Finally, we can show thatI\(x\)→logλI\(x\)\\rightarrow\\log\\lambdaby computation:
limx→∞∫xkxλ1\(ulogu\)𝑑u\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\int\_\{x\}^\{kx^\{\\lambda\}\}\\frac\{1\}\{\(u\\log u\)\}du=\\displaystyle=limx→∞loglog\(kxλ\)−loglogx\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\log\\log\\left\(kx^\{\\lambda\}\\right\)\-\\log\\log x=\\displaystyle=limx→∞log\(logkxλlogx\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\log\\left\(\\frac\{\\log kx^\{\\lambda\}\}\{\\log x\}\\right\)=\\displaystyle=limx→∞log\(\(logk\+λlogx\)logx\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\log\\left\(\\frac\{\\left\(\\log k\+\\lambda\\log x\\right\)\}\{\\log x\}\\right\)=\\displaystyle=limx→∞log\(1logx\(logk\+λlogx\)\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\log\\left\(\\frac\{1\}\{\\log x\}\(\\log k\+\\lambda\\log x\)\\right\)=\\displaystyle=limx→∞log\(logklogx\+λ\)\\displaystyle\\lim\_\{x\\rightarrow\\infty\}\\log\\left\(\\frac\{\\log k\}\{\\log x\}\+\\lambda\\right\)=\\displaystyle=logλ=log\(logab\)\>0\\displaystyle\\log\\lambda=\\log\(\\log\_\{a\}b\)\>0
∎
###### Corollary A\.6\(Original “Meek Models”\)\.
LetL\(C\)L\(C\)be the validation loss of a model at compute optimality, given byL\(C\)=AC−α\+EL\(C\)=AC^\{\-\\alpha\}\+E, whereA,EA,Eandα\\alphaare positive constants\. Moreover letga,gh,g\_\{a\},g\_\{h\},andgig\_\{i\}be annual rates of shared algorithmic progress, shared hardware progress, and compute scaling by a single actor, each greater than11\. Then the loss difference between the scaling actor and a static actor, given byL\(\(ghga\)tC0\)−L\(\(ghgagi\)tC0\)L\(\(g\_\{h\}g\_\{a\}\)^\{t\}C\_\{0\}\)\-L\(\(g\_\{h\}g\_\{a\}g\_\{i\}\)^\{t\}C\_\{0\}\), goes to zero ast→∞t\\rightarrow\\infty\.
This corollary can be seen immediately by lettingb=gighga\>ghga=ab=g\_\{i\}g\_\{h\}g\_\{a\}\>g\_\{h\}g\_\{a\}=a, andP\(C\)=−L\(C\)P\(C\)=\-L\(C\)\. SinceLLis bounded below,PPis bounded above and is therefore a meek metric\.
## Appendix BPlot details
We construct smooth trends from constants reported by METR\(Kwaet al\.,[2026](https://arxiv.org/html/2607.00913#bib.bib9)\), starting from the release date of GPT\-4–0314\. We use their exponential doubling time for the unbounded 50% time horizon, reported as∼\\sim7\-month\. We also use the logistic\-in\-task\-duration relationship for the probability of task success, which shifts rightward with time according to 50% time horizon, and estimate probabilities at three fixed human\-completion times \(4 min, 1 hr, and 4 hr\)\. We anchor the horizon trend to Claude 3\.7 Sonnet’s reported∼\\sim1\-hour 50% horizon \(released1\.951\.95years after GPT\-4 0314\), and place the frontier and meek models at the same effective compute att=0t=0\(GPT\-4 0314\)\. The success logistic is explicitly reported to have a5×5\\timesratio between the 50% and 80% horizons, which locks in the slope of that curve, while the exponential locks in the x\-axis shift at a particular point in time\.
For the meek versus mighty estimation, the meek model follows an identical time trend more slowly, such thatmeek\(t\)=frontier\(ρt\)\\text\{meek\}\(t\)=\\text\{frontier\}\(\\rho t\)\. Intuitively, forb\>a\>1b\>a\>1, we wantbρtC0=atC0b^\{\\rho t\}C\_\{0\}=a^\{t\}C\_\{0\}, and thereforeρ=log\(a\)/log\(b\)=log\(ghga\)/log\(ghgagi\)≈0\.55\\rho=\\log\(a\)/\\log\(b\)=\\log\(g\_\{h\}g\_\{a\}\)/\\log\(g\_\{h\}g\_\{a\}g\_\{i\}\)\\approx 0\.55\. Here,gh≈1\.36g\_\{h\}\\approx 1\.36/yr is the shared hardware\-efficiency growth, taken as the FP16 performance\-per\-dollar trend for top\-performing datacenter GPUs without sparsity \(a2\.252\.25\-year doubling\) fromDel Sozzoet al\.\([2026](https://arxiv.org/html/2607.00913#bib.bib73)\); we use FP16 as the precision on which modern mixed\-precision frontier training runs, and note that the more precision\-neutral FP32 trend gives a similar, slightly more conservative rate\. The shared algorithmic effective\-compute rate isga=2\.8g\_\{a\}=2\.8/yr\(Hoet al\.,[2024](https://arxiv.org/html/2607.00913#bib.bib58)\)\. Frontier training compute has grown at roughlyghgi≈4\.1g\_\{h\}g\_\{i\}\\approx 4\.1/yr\(Sevillaet al\.,[2022](https://arxiv.org/html/2607.00913#bib.bib36)\), implying a frontier\-only investment scale\-upgi≈3\.0g\_\{i\}\\approx 3\.0/yr after dividing out hardware\. All curves plot the frontier\-minus\-meek gap of the titular metric, projected to 2040\.
1"""Frontier\-vs\-meekperformancegaponMETRtime\-horizonmetrics\(constantsonly\)\."""
2
3importmath
4importnumpyasnp
5
6
7
8
9HW\_DOUBLING\_YR=2\.25
10G\_H=2\.0\*\*\(1\.0/HW\_DOUBLING\_YR\)
11G\_A,G\_I=2\.80,3\.00
12
13
14DOUBLING\_MONTHS=7\.0
15RATIO\_50\_80=5\.0
16P\_HI=0\.80
17
18ANNUAL\_GROWTH=12\.0/DOUBLING\_MONTHS
19H\_ANCHOR\_MIN,T\_ANCHOR\_YR=60\.0,1\.95
20
21
22REF\_YEAR,END\_YEAR=2023\.20,2040\.0
23TASK\_MIN=\{"4minutes":4\.0,"1hour":60\.0,"4hours":240\.0\}
24
25
26
27deffrontier\_50\_percent\_horizon\(t\):
28"""Frontier50%timehorizon\(minutes\)attyearssincethereferencedate\."""
29return2\.0\*\*\(ANNUAL\_GROWTH\*\(t\-T\_ANCHOR\_YR\)\+math\.log2\(H\_ANCHOR\_MIN\)\)
30
31deffrontier\_p\_success\(t,task\_min\):
32"""FrontierP\(success\)onataskoflengthtask\_min,atyear\-offsett\."""
33logit=lambdax:math\.log\(x/\(1\.0\-x\)\)
34
35log\_space\_distance\_at\_t=np\.log2\(frontier\_50\_percent\_horizon\(t\)\)\-np\.log2\(task\_min\)
36logit\_space\_increase\_per\_log\_space=logit\(P\_HI\)/math\.log2\(RATIO\_50\_80\)
37return1\.0/\(1\.0\+np\.exp\(\-logit\_space\_increase\_per\_log\_space\*log\_space\_distance\_at\_t\)\)
38
39defgap\(curve,t\):
40"""Frontierminusmeek,wherethemeekpathreachesthefrontieratrateRHO\."""
41
42\#whereb=G\_H\*G\_A\*G\_I,anda=G\_H\*G\_A
43RHO=math\.log\(G\_H\*G\_A\)/math\.log\(G\_H\*G\_A\*G\_I\)
44returncurve\(t\)\-curve\(RHO\*t\)
45
46years=np\.linspace\(0,END\_YEAR\-REF\_YEAR,400\)
47calendar=REF\_YEAR\+years
48gap\_horizon=gap\(frontier\_50\_percent\_horizon,years\)
49gap\_p=\{lbl:gap\(lambdat,T=T:frontier\_p\_success\(t,T\),years\)
50forlbl,TinTASK\_MIN\.items\(\)\}
51
52
53
54importmatplotlib\.pyplotasplt
55frommatplotlib\.tickerimportFuncFormatter,FixedLocator,MultipleLocator
56
57LABEL\_FS,TICK\_FS,TITLE\_FS,LEGEND\_FS=16,13,18,16
58
59TIME\_TICKS=\[\(1,"1min"\),\(60,"1hr"\),\(1440,"1day"\),\(43200,"1mo"\),
60\(525600,"1yr"\),\(5256000,"10yr"\),\(52560000,"100yr"\),
61\(525600000,"1000yr"\)\]
62year\_fmt=FuncFormatter\(lambdax,\_:f"\{int\(round\(x\)\)\}"\)
63pct\_fmt=FuncFormatter\(lambdav,\_:f"\{v:\.0f\}%"\)
64
65fig,\(axL,axR\)=plt\.subplots\(2,1,figsize=\(7\.5,9\.5\)\)
66
67axL\.plot\(calendar,gap\_horizon,color="\#1f77b4",lw=2\.5\)
68axL\.set\_yscale\("log"\)
69axL\.yaxis\.set\_major\_locator\(FixedLocator\(\[mform,\_inTIME\_TICKS\]\)\)
70axL\.yaxis\.set\_major\_formatter\(FuncFormatter\(lambdam,\_:dict\(TIME\_TICKS\)\.get\(m,""\)\)\)
71axL\.set\_ylim\(min\(gap\_horizon\[gap\_horizon\>0\]\),gap\_horizon\.max\(\)\*1\.5\)
72axL\.set\_xlabel\("Year",fontsize=LABEL\_FS\)
73axL\.set\_ylabel\("FrontierModelvs\.MeekModelGap",fontsize=LABEL\_FS\)
74axL\.set\_title\("METR50%TimeHorizon",fontsize=TITLE\_FS,fontweight="bold"\)
75axL\.grid\(True,which="both",alpha=0\.25\)
76
77for\(lbl,g\),colorinzip\(gap\_p\.items\(\),\("\#2ca02c","\#d62728","\#9467bd"\)\):
78axR\.plot\(calendar,100\.0\*g,color=color,lw=2\.5,label=lbl\)
79axR\.yaxis\.set\_major\_formatter\(pct\_fmt\)
80axR\.set\_xlabel\("Year",fontsize=LABEL\_FS\)
81axR\.set\_ylabel\("FrontierModelvs\.MeekModelGap",fontsize=LABEL\_FS\)
82axR\.set\_title\("METRProjectedSuccessatFixedTimeHorizon",fontsize=TITLE\_FS,fontweight="bold"\)
83axR\.set\_ylim\(0,None\)
84axR\.legend\(frameon=False,fontsize=LEGEND\_FS\)
85axR\.grid\(True,alpha=0\.25\)
86
87foraxin\(axL,axR\):
88ax\.xaxis\.set\_major\_locator\(MultipleLocator\(4\)\)
89ax\.xaxis\.set\_major\_formatter\(year\_fmt\)
90ax\.tick\_params\(axis="both",labelsize=TICK\_FS\)
91
92fig\.tight\_layout\(\)
93fig\.align\_ylabels\(\(axL,axR\)\)
94fig\.savefig\("meek\_gap\_metr\.png",dpi=180,bbox\_inches="tight"\)Similar Articles
One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation
This study tests whether economic benchmarks in AI leaderboards measure distinct capabilities or just a general trend, finding they add incremental information but are largely time-driven.
Are we slowly moving toward two different kinds of AI?
An observation about the growing divergence between heavily restricted mainstream AI models and more open, less restricted local models, and a question about whether this divide will persist or one side will dominate.
Today's Frontier AI companies will never exceed the AI capability frontier again (18 minute read)
Article argues that networks of smaller AI models are now surpassing frontier AI systems in speed, accuracy, and cost, predicting a shift to decentralized 'network-source AI'.
The Economics of the Intelligence Frontier (20 minute read)
AI tasks become commodities once models exceed maximum necessary intelligence, shifting competition to cost and infrastructure, but frontier labs can thrive by creating valuable new markets before commoditization.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.