One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

arXiv cs.LG Papers

Summary

This study tests whether economic benchmarks in AI leaderboards measure distinct capabilities or just a general trend, finding they add incremental information but are largely time-driven.

arXiv:2608.29420v1 Announce Type: new Abstract: Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent-variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R^2 = 0.505), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed. Removing the date trend lowers that share by 14.9 points, and by 24.1 with one row per base model. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave-one-benchmark-out test with factors re-estimated inside every fold shows that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]). Economic benchmarks therefore add incremental predictive information to a largely date-driven general factor, and the evidence does not support treating them as a distinct latent capability. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date-adjusted before being read as a capability difference.
Original Article
View Cached Full Text

Cached at: 09/01/26, 01:15 PM

# One Capability or Many?Testing the Economic Validity of Frontier AI Evaluation
Source: [https://arxiv.org/html/2608.29420](https://arxiv.org/html/2608.29420)
Louis Yiven Zhu††thanks:ORCID 0009\-0001\-5579\-0340\.

###### Abstract

Frontier\-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change\. Whether such benchmarks measure a capability distinct from general test\-taking or re\-express the one axis along which every benchmark rises as models improve is a question of construct validity that has not yet been studied\. We test it on a hash\-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent\-variable model with four hypotheses and their thresholds fixed before analysis\. A single factor explains 74\.5% of common variance and tracks model release date \(R2=0\.505R^\{2\}=0\.505\), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed\. Removing the date trend lowers that share by 14\.9 percentage points and by 24\.1 when the analysis is repeated with one row per base model\. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave\-one\-benchmark\-out test with factors re\-estimated inside every fold shows that a multi\-factor representation predicts held\-out economic scores better than a single general\-capability index \(pooledΔ​MSE\\Delta\\mathrm\{MSE\}0\.037, 95% bootstrap interval\[0\.019,0\.055\]\[0\.019,0\.055\]\)\. Economic benchmarks therefore add incremental predictive information to a largely date\-driven general factor, and the evidence does not support treating them as a distinct latent capability\. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date\-adjusted before it is read as a capability difference; for benchmark builders, we give a two\-test protocol for showing that a new suite measures more\.

## 1Introduction

Public leaderboards have become the dominant instrument for comparing and procuring frontier language models, and they increasingly inform policy, so a recent shift in what they are based on raises a measurement question with direct economic stakes\. Alongside academic knowledge tests such as GPQA \(Graduate\-Level Google\-Proof Q&A\)\([Rein et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib37)\)and Humanity’s Last Exam \(HLE\)\([Phan et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib34)\), leaderboards now carry*economic*benchmarks such as GDPval\([Patwardhan et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib32)\), theτ\\tau\-Bench family\([Yao et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib52);[Barres et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib3)\), and terminal or agentic coding suites\([Merrill et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib29)\), all built to proxy paid professional work and multi\-step tool use, one of them \(τ3\\tau^\{3\}\-Banking\) in banking customer support, a policy\-compliance setting that evaluation suites cover thinly\. Whether these benchmarks deserve a separate column depends on what they measure\. If they capture a capability distinct from general test\-taking, they carry information that a single “intelligence” score cannot; if they merely re\-express one dominant axis of model progress, a separate column overstates what the leaderboard knows\.This is a question of construct validity, the degree to which a measurement captures the construct it claims to\([Cronbach and Meehl, 1955](https://arxiv.org/html/2608.29420#bib.bib12);[Messick, 1998](https://arxiv.org/html/2608.29420#bib.bib30)\)\. Although machine learning has been urged to hold its benchmarks to that psychometric standard\([Raji et al\., 2021](https://arxiv.org/html/2608.29420#bib.bib36);[Jacobs and Wallach, 2021](https://arxiv.org/html/2608.29420#bib.bib21);[Bowman and Dahl, 2021](https://arxiv.org/html/2608.29420#bib.bib7)\), recent audits find construct validity systematically under\-examined in AI evaluation\([Bean et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib5);[Kearns, 2026](https://arxiv.org/html/2608.29420#bib.bib25)\)\.

We frame a snapshot of frontier models scored across twelve benchmarks as a psychometric problem, with benchmarks as items, models as respondents, and scores as the responses from which latent structure is inferred, and answer it under a pre\-specified design \(Figure[1](https://arxiv.org/html/2608.29420#S1.F1)\)\. Four research questions organise the analysis, each mapping to one pre\-specified hypothesis \(Section[4\.3](https://arxiv.org/html/2608.29420#S4.SS3)\); the first three concern structure \(Task 1, unsupervised\) and the fourth prediction \(Task 2, supervised\)\.

1. RQ1\(dominance\)How many latent capabilities does the battery measure, and how dominant is the leading one?
2. RQ2\(temporal confound\)Is the leading axis a release\-date artefact, and how much survives date adjustment?
3. RQ3\(structural distinctiveness\)After date adjustment, do the economic benchmarks load on a factor of their own?
4. RQ4\(predictive validity\)Does a multi\-factor representation predict held\-out economic scores better than a single general index?

Figure 1:Study design\.\(a\) Benchmarks as items \(columns, coloured by capability block, economic in red\) and model configurations as respondents \(rows, ordered by release date\); a coverage rule fixed in advance reduces 548 configurations to a421×12421\\times 12matrix\. \(b\) The two candidate structures, a single date\-driven general factorgg, orggtogether with a distinct economic factoree\. \(c\) The two pre\-specified tests, date adjustment followed by re\-fitting \(H1–H3\) and leave\-one\-benchmark\-out prediction \(H4\); grids are stylised\.#### Contributions\.

The paper makes four contributions\. First, it subjects the economic column of a frontier leaderboard to a construct\-validity test for the first time\. Second, it isolates model release date as a first\-order confound and removes it before reading any factor structure, whereupon the dominant capability axis proves substantially a temporal artefact, and where prior work controls for scale, the calendar carries what scale carries across generations\. Third, it introduces a leave\-one\-benchmark\-out predictive\-validity test, since a benchmark that adds nothing to held\-out prediction adds nothing a practitioner can use\. Fourth, it delivers the analysis under a protocol fixed in advance with hash\-pinned inputs, documenting every deviation and reporting the two failed tests as failed\.

## 2Related work

Two strands of the construct\-validity programme bear on this paper\. The first treats LLM scores as psychometric data, following the case for measuring artificial systems with the instruments built for measuring minds\([Hernández\-Orallo, 2017](https://arxiv.org/html/2608.29420#bib.bib17)\)and the century\-old observation that cognitive tests yield a positive manifold and a general factor\([Spearman, 1904](https://arxiv.org/html/2608.29420#bib.bib45)\)\.[Ilić and Gignac \(2024\)](https://arxiv.org/html/2608.29420#bib.bib20)recover that manifold across 591 models on twelve tests, with a mean inter\-test correlation of 0\.73 and a general factor correlating with parameter count at about 0\.6;[Burnell et al\. \(2023\)](https://arxiv.org/html/2608.29420#bib.bib9), applying factor analysis to 29 LLMs across 27 cognitive tasks, resolve capability instead into three factors\.[Ruan et al\. \(2024\)](https://arxiv.org/html/2608.29420#bib.bib41)then put the low\-dimensional structure to work, using principal components of a score matrix as a capability space in which performance becomes predictable, including agentic performance forecast from non\-agentic benchmarks\([Zhou et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib54), see also\)\. Closest to our design,[Kearns \(2026\)](https://arxiv.org/html/2608.29420#bib.bib25)makes the validity question quantitative on the Open LLM Leaderboard, where a plain factor analysis fits a single factor explaining 72% of communal variance that tracks parameter count, and residualising every score on a fitted scaling law lowers the mean inter\-item correlation from 0\.64 to 0\.48; that residualisation is the template for our date adjustment\. Across all of these studies scale is the variable to control; on a frontier snapshot we find that release date takes that role and compute adds little once date is removed, which suggests the confound has moved with the field\.

The second strand audits benchmarks as artefacts\. Calls to repair benchmarking predate the frontier era\([Bowman and Dahl, 2021](https://arxiv.org/html/2608.29420#bib.bib7);[Kiela et al\., 2021](https://arxiv.org/html/2608.29420#bib.bib26)\), and holistic evaluation frameworks\([Liang et al\., 2023](https://arxiv.org/html/2608.29420#bib.bib27)\), reproducibility guidance\([Biderman et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib6)\), and proposals for a full evaluation science\([Weidinger et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib51)\)have since sharpened them\. Reviewing 445 benchmarks,[Bean et al\. \(2025\)](https://arxiv.org/html/2608.29420#bib.bib5)find that 53\.4% present any evidence for the construct validity of their benchmark and 16\.0% report uncertainty estimates or statistical tests when comparing models, a gap[Miller \(2024\)](https://arxiv.org/html/2608.29420#bib.bib31)addresses with error bars and[Reuel et al\. \(2024\)](https://arxiv.org/html/2608.29420#bib.bib38)with a quality checklist\. On the temporal side,[Akhtar et al\. \(2026\)](https://arxiv.org/html/2608.29420#bib.bib1)find that nearly half of sixty benchmarks are saturated and that saturation rises with benchmark age, which makes release date a first\-order confound whenever capabilities are compared across a moving frontier, and[Reuel et al\. \(2025\)](https://arxiv.org/html/2608.29420#bib.bib39)map an uneven division of evaluation labour between developers and third parties\. None of this work isolates the economic column, and none pairs a structural test with an out\-of\-sample predictive one\.

## 3Data and design

#### Source and provenance\.

The primary data set is a snapshot of the Artificial Analysis model leaderboard\([Artificial Analysis, 2026](https://arxiv.org/html/2608.29420#bib.bib2)\), captured on 6 July 2026 directly from the public leaderboard page and pinned by its SHA\-256 hash \(Appendix[N](https://arxiv.org/html/2608.29420#A14)\)\. It contains 548 model configurations with per\-benchmark scores, release dates, and provider metadata\. A secondary source, the Epoch AI Notable AI Models data set\([Epoch AI, 2026](https://arxiv.org/html/2608.29420#bib.bib14)\), supplies training\-compute figures for a subset of models and serves only a scale robustness check\. Both snapshots are pinned, so every number here is reproducible from fixed inputs\.

#### Inclusion rule and analysis grids\.

To avoid selecting benchmarks or models on the very outcomes under study, we fixed the coverage rule in advance\. A benchmark is retained when at least sixty models carry a score, and a model when it carries at least eight of the thirteen sufficiently covered benchmarks\. Both floors were set as indicative values in the analysis plan and confirmed unchanged at the coverage audit; that confirmation used coverage counts alone, never a score, a correlation, or a factor solution\. Appendix[G](https://arxiv.org/html/2608.29420#A7)shows what each protects\. Under this rule thirteen of fourteen candidates pass; we drop APEX\-Agents\([Vidgen et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib50)\)for sparsity and demote MMMU\-Pro\([Yue et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib53)\)to a sensitivity\-only role, leaving twelve primary benchmarks and 421 model configurations, four of them economic benchmarks proxying paid professional or agentic work and eight spanning academic knowledge, scientific coding\([Tian et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib47);[Zhu et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib55)\), long\-context retrieval, and instruction following\([Pyatkin et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib35)\)\(Appendix[F](https://arxiv.org/html/2608.29420#A6)\)\. The 421 rows are configurations, since one base model can appear under several reasoning or effort settings; the plan deduplicates to one row per base model as robustness check R1, and we keep the configuration level as primary because it is the more conservative choice\.

Because the economic benchmarks are scored on far fewer models than the academic core, we use two grids\. The complete\-case grid of 96 models scored on all twelve benchmarks is the only one on which the economic block is complete, and it carries every hypothesis\. The dense grid of 409 models scored on the nine near\-universal benchmarks serves the corroborating clustering \(Appendix[K](https://arxiv.org/html/2608.29420#A11)\), where sample size matters more than economic completeness\. This trade\-off is the central data\-design decision, and the 96\-model grid is the set of frontier configurations on which the economic benchmarks have actually been run, so it is representative of that population by construction\.

#### Preprocessing and scope\.

Since the benchmarks are reported on incommensurable scales, mostly accuracy fractions alongside one hallucination\-penalised score and GDPval’s Elo rating \(Appendix[F](https://arxiv.org/html/2608.29420#A6)\), we standardise each benchmark to zero mean and unit variance before any analysis, using training\-fold statistics only in the supervised task\. Missingness is structural, so we use complete\-case analysis and impute nothing\. On the complete\-case grid the Kaiser–Meyer–Olkin \(KMO\) measure of sampling adequacy is 0\.933, well above the 0\.8 threshold\([Kaiser, 1974](https://arxiv.org/html/2608.29420#bib.bib24)\), and Bartlett’s test of sphericity is decisively rejected\([Bartlett, 1950](https://arxiv.org/html/2608.29420#bib.bib4)\), so factor analysis is appropriate\. The battery is also highly collinear before any modelling, with a mean off\-diagonal Spearman correlation ofρ=0\.79\\rho=0\.79, every benchmark pair positive, and the economic benchmarks correlating with the academic core almost as strongly as with one another \(Appendix[H](https://arxiv.org/html/2608.29420#A8)\); the manifold is thus at least as strong here as the 0\.73[Ilić and Gignac \(2024\)](https://arxiv.org/html/2608.29420#bib.bib20)report on a broader population\. We fixed the scope of the work to internal validity, meaning the correlational and predictive structure of aggregate scores on one cross\-section; item\-level responses, external economic impact such as adoption or revenue, and saturation dynamics lie outside it\.

## 4Methods

### 4\.1Task 1: latent structure

Each structural question has its own quantity, and Appendix[D](https://arxiv.org/html/2608.29420#A4)states each method formally with the reasoning behind it\. The first factor’s share of common variance measures dominance \(RQ1\), the change in that share after date adjustment measures the temporal confound \(RQ2\), and the date\-adjusted loading pattern tests distinctiveness \(RQ3\)\. To fix dimensionality, we use Horn’s parallel analysis\([Horn, 1965](https://arxiv.org/html/2608.29420#bib.bib19)\), retaining factors whose eigenvalue exceeds the 95th percentile of eigenvalues from random data of the same shape, with the scree criterion\([Cattell, 1966](https://arxiv.org/html/2608.29420#bib.bib11)\)as a cross\-check, since on collinear data it is more conservative than the Kaiser rule\([Kaiser, 1960](https://arxiv.org/html/2608.29420#bib.bib23)\)and than information criteria \(Appendix[D](https://arxiv.org/html/2608.29420#A4)\)\. We prefer factor analysis to the principal\-component description of[Ruan et al\. \(2024\)](https://arxiv.org/html/2608.29420#bib.bib41)because components mix common and unique variance, whereas the common\-factor model isolates shared capability, the distinction a construct\-validity question turns on\. The model is accordingly a maximum\-likelihood exploratory factor analysis \(EFA\) with an oblique oblimin rotation\([Jennrich and Sampson, 1966](https://arxiv.org/html/2608.29420#bib.bib22);[Thurstone, 1947](https://arxiv.org/html/2608.29420#bib.bib46)\), chosen because the date\-adjusted factors are strongly correlated \(Section[5\.3](https://arxiv.org/html/2608.29420#S5.SS3)\) and an orthogonal rotation would spread shared variance across artificially independent axes\. Every factor is oriented so that higher means more capable\.

To separate a general capability from a shared time trend, we regress each standardised benchmark on release date, and on a subsample also on training compute, and re\-fit the factor analysis on the residuals; we call this date adjustment\. It removes from the benchmark covariance the rank\-one component aligned with date \(Appendix[D](https://arxiv.org/html/2608.29420#A4)\), so the resulting change in the first factor’s variance share measures how much of the apparent general factor is a date artefact, the analogue of the scaling\-law residualisation[Kearns \(2026\)](https://arxiv.org/html/2608.29420#bib.bib25)reports\. Clustering corroborates the structure without assuming a factor model \(Appendix[K](https://arxiv.org/html/2608.29420#A11)\)\.

### 4\.2Task 2: predictive validity

Task 2 recasts the distinctiveness question as a forecast, asking whether a model’s score on a held\-out economic benchmark follows from its behaviour on the others, a sharper test than the in\-sample capability space of[Ruan et al\. \(2024\)](https://arxiv.org/html/2608.29420#bib.bib41)because the target column is withheld from estimation entirely\. To that end, the leave\-one\-benchmark\-out \(LOBO\) protocol takes each economic benchmark as the target in turn and predicts it from representations of the remaining eleven on then=96n=96complete\-case grid, so that success demands transfer across benchmarks\. A five\-rung predictor ladder makes the comparison concrete\. Rung \(i\) uses release timing and scale only, rung \(ii\) a single mean\-score general index, rung \(iii\) the first factor only, rung \(iv\) thekkfactors \(the three\-factor solution of Section[4\.3](https://arxiv.org/html/2608.29420#S4.SS3)\), and rung \(v\) thekkfactors plus model covariates\. The contrast of interest sets rung \(iv\) against the single\-index rung \(ii\)\.

So that the ladder comparison is tied to no single model class, we fit each rung with four learners spanning the linear\-regularised and tree families\([Hastie et al\., 2009](https://arxiv.org/html/2608.29420#bib.bib16)\), ridge\([Hoerl and Kennard, 1970](https://arxiv.org/html/2608.29420#bib.bib18)\), elastic net\([Zou and Hastie, 2005](https://arxiv.org/html/2608.29420#bib.bib56)\), random forest\([Breiman, 2001](https://arxiv.org/html/2608.29420#bib.bib8)\), and gradient boosting\([Friedman, 2001](https://arxiv.org/html/2608.29420#bib.bib15)\)\. Hyperparameters are tuned by grid search in an inner five\-fold cross\-validation nested inside an outer five\-fold split that supplies the reported out\-of\-fold error, the nesting that removes selection bias from the estimate\([Varma and Simon, 2006](https://arxiv.org/html/2608.29420#bib.bib49)\)\(grid in Appendix[I](https://arxiv.org/html/2608.29420#A9)\)\. Because the factors are estimated quantities, we re\-estimate them inside the loop, fitting the factor analysis on each training fold alone, projecting the held\-out fold through the training\-fold Thurstone factor\-score weights, and re\-orienting by the training\-fold mean index; a naive fit\-on\-all\-data pipeline would silently open that leakage path \(full procedure in Appendix[B](https://arxiv.org/html/2608.29420#A2)\)\.

Two error metrics then serve two purposes\. Root mean squared error \(RMSE\) on the standardised target is the descriptive metric since it stays on the target’s own scale\. The test statistic isΔ​MSE=MSE⁡\(single index\)−MSE⁡\(k​factors\)\\Delta\\mathrm\{MSE\}=\\mathrm\{MSE\}\(\\text\{single index\}\)\-\\mathrm\{MSE\}\(k\\text\{ factors\}\)per economic target, positive when the multi\-factor representation wins, with a 95% interval from a paired bootstrap of the pooled out\-of\-fold errors\([Efron and Tibshirani, 1993](https://arxiv.org/html/2608.29420#bib.bib13)\)withB=2000B=2000\.

### 4\.3Pre\-specified hypotheses

Before analysis, the design fixed four hypotheses with quantitative thresholds set with reference to the figures[Kearns \(2026\)](https://arxiv.org/html/2608.29420#bib.bib25)reports for the Open LLM Leaderboard\.H1\(dominance, RQ1\) holds that the first factor carries more than 50% of common variance\.H2\(date artefact, RQ2\) has two parts, \(i\) that the dominant factor’s release\-dateR2R^\{2\}is at least 0\.30 and the strongest of the factors, and \(ii\) that date adjustment lowers the first factor’s share by at least fifteen percentage points\.H3\(structural distinctiveness, RQ3\) holds that after date adjustment parallel analysis retains a factor on which the four economic benchmarks load at least 0\.40 and exceed their cross\-loadings, which is the discriminant criterion of[Campbell and Fiske \(1959\)](https://arxiv.org/html/2608.29420#bib.bib10)applied to a benchmark battery\.H4\(predictive validity, RQ4\) holds that the multi\-factor representation’sΔ​MSE\\Delta\\mathrm\{MSE\}interval excludes zero and a majority of the four economic targets individually favour it\. We report every outcome, including the two that fail\.

Two factor counts are used in what follows\. The primary count isk=1k=1, returned by parallel analysis, the Kaiser rule, and the scree elbow, and it governs H1 to H3\. The second isk=3k=3, an exploratory over\-extraction that no criterion selects, which supplies the loading pattern of Section[5\.3](https://arxiv.org/html/2608.29420#S5.SS3)and thekk\-factor predictor of Task 2 and is a documented deviation from the plan \(Appendix[J](https://arxiv.org/html/2608.29420#A10)\); the predictive conclusion does not hinge on it, since sweeping rung \(iv\) overk∈\{2,3,4,5\}k\\in\\\{2,3,4,5\\\}beats the single index at every count \(Appendix[L](https://arxiv.org/html/2608.29420#A12)\)\.

## 5Results

### 5\.1Dominance of a single factor \(H1\)

Every dimensionality diagnostic points to a dominant capability axis\. The first principal component alone accounts for 79\.4% of total score variance, and parallel analysis retains a single factor, since only the first observed eigenvalue of 9\.53 exceeds its random\-data 95th percentile of 1\.79 \(Appendix Figure[5](https://arxiv.org/html/2608.29420#A4.F5)\)\. In the maximum\-likelihood factor analysis the first factor holds 74\.5% of common variance, above the 50% dominance threshold and close to the 72%[Kearns \(2026\)](https://arxiv.org/html/2608.29420#bib.bib25)reports on a different leaderboard, and more concentrated than the three\-factor structure[Burnell et al\. \(2023\)](https://arxiv.org/html/2608.29420#bib.bib9)recover\. H1 is therefore supported, which answers RQ1 in full, and because rank\-based and logit\-transformed variants only raise the share \(Section[5\.5](https://arxiv.org/html/2608.29420#S5.SS5)\), the Pearson\-based headline is the conservative one\.

### 5\.2A date\-driven artefact \(H2\)

Having established a dominant factor, we ask whether it reflects stable capability or the passage of time\. The factor’s scores track model release date with a logistic\-fitR2R^\{2\}of 0\.505 \(0\.477 under an ordinary\-least\-squares fit, Appendix[L](https://arxiv.org/html/2608.29420#A12)\), the strongest of the three factors, so H2\(i\) is supported and the leading axis of “capability” is substantially a time trend\. Adjusting every benchmark for release date and re\-fitting then lowers the first factor’s common\-variance share from 74\.5% to 59\.6% \(Figure[2](https://arxiv.org/html/2608.29420#S5.F2)c\), a drop of 14\.9 percentage points, which means date alone accounts for about a fifth of the general factor\. That point estimate falls just short of the pre\-specified fifteen\-point threshold, with a bootstrap 95% interval of\[−5\.3,\+32\.7\]\[\-5\.3,\+32\.7\], so on the primary grid the drop is indistinguishable both from zero and from the threshold\. We report H2\(ii\) as not met on the point estimate without moving the line after seeing the result\.

![Refer to caption](https://arxiv.org/html/2608.29420v1/fig_loadings.png)Figure 2:Factor loadings from the three\-factor oblimin EFA\.\(a\) Raw scores; the economic benchmarks \(red\) already concentrate on F1\. \(b\) After date adjustment they still concentrate on F1, with a largest cross\-loading of 0\.38\. \(c\) The first factor’s common\-variance share falls from 74\.5% to 59\.6%, a 14\.9\-point drop\.Whether H2\(ii\) passes depends on the unit of analysis\. Under the plan’s base\-model unit the drop is 24\.1 points and clears the threshold; under the configuration\-level primary it is 14\.9 points and fails; on the 58\-model compute\-known subsample date alone gives 16\.5 points and clears it\. Two of the three specifications therefore pass\. The alternative, promoting the deduplicated grid to primary, would pass the threshold but would turn the plan’s robustness unit into the headline on fewer models \(89 against 96\), so we keep the configuration\-level specification as primary and the confirmation of a temporal confound rests on the subsample and the deduplicated grid\. RQ2 is therefore answered in part, since the leading axis is substantially a release\-date artefact while the size of the correction is established on two of the three specifications\. Adding training compute does not deepen the correction, since on the compute\-known subsample the joint date\-plus\-compute drop is smaller, at 9\.3 points, plausibly because later models are also larger and the two predictors split their shared variance\. This inverts the usual control, since prior work treats scale as the variable to model or remove\([Kearns, 2026](https://arxiv.org/html/2608.29420#bib.bib25);[Ilić and Gignac, 2024](https://arxiv.org/html/2608.29420#bib.bib20);[Ruan et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib41)\), whereas on a frontier snapshot the calendar carries what scale carries across generations\.

### 5\.3Economic distinctiveness \(H3\)

With the temporal confound removed, we test whether the economic benchmarks form a distinct factor, a pre\-specified test with two conditions, that parallel analysis retain a factor and that its loadings concentrate discriminantly on the economic block\. On the date\-adjusted data parallel analysis retains a single factor, which leaves no block to separate; the separation appears only once three factors are extracted, which the dimensionality rule does not endorse\. Holding H3 to the plan as H2\(ii\) was held to its 15\-point line, we report H3 as not supported as specified, which answers RQ3 negatively under the plan\.

Under the three\-factor over\-extraction the economic block does separate, and we report the pattern as exploratory\. All four economic benchmarks load at least 0\.40 on a single factor, F1, with GDPval at 0\.84, Terminal\-Bench v2\.1 at 1\.01,τ3\\tau^\{3\}\-Banking at 0\.54, andτ2\\tau^\{2\}\-Bench at 0\.50, against a largest economic cross\-loading of 0\.38 \(Figure[2](https://arxiv.org/html/2608.29420#S5.F2)b, tabulated in Appendix[E](https://arxiv.org/html/2608.29420#A5)\); the mean absolute F1 loading is 0\.72 for the economic benchmarks and 0\.32 for the rest\. Because the three factors are strongly correlated \(0\.67 between F1 and F3 and 0\.72 between F2 and F3\) the separation reads as a secondary refinement of one dominant axis\. Bootstrap intervals on the economic loadings are wide \(0\.10 to 0\.78 forτ2\\tau^\{2\}\-Bench and 0\.44 to 0\.99 for GDPval\), excluding zero but overlapping one another, so the pattern is suggestive and loosely estimated\. As discussed in Section[5\.4](https://arxiv.org/html/2608.29420#S5.SS4), out\-of\-sample transfer is a stronger dimensionality argument than an in\-sample retention rule, and we let the predictive test carry the claim the structural test cannot\.

### 5\.4Predictive validity \(H4\)

The final test asks whether the economic factor improves prediction\. Across the predictor ladder a single mean\-score index is already a strong baseline, at economic\-block test RMSE 0\.474 andR2R^\{2\}0\.771; thekk\-factor representation improves on it to RMSE 0\.433 andR2R^\{2\}0\.808, and adding covariates gives no further gain \(Table[1](https://arxiv.org/html/2608.29420#S5.T1), Figure[3](https://arxiv.org/html/2608.29420#S5.F3)a\)\. The first factor alone predicts poorly at RMSE 0\.950, which locates the economic signal in the secondary factors\. BootstrappingΔ​MSE\\Delta\\mathrm\{MSE\}over the pooled out\-of\-fold errors gives a pooled improvement of\+0\.037\+0\.037with a 95% interval of\[\+0\.019,\+0\.055\]\[\+0\.019,\+0\.055\], and because that resample includes near\-duplicate settings of the same base model, we re\-run it on the deduplicated grid of 89 distinct models, where the improvement is\+0\.038\+0\.038with\[\+0\.020,\+0\.056\]\[\+0\.020,\+0\.056\]\(Appendix[L](https://arxiv.org/html/2608.29420#A12)\)\. Three of the four economic benchmarks improve individually, GDPval by\+0\.026\+0\.026, Terminal\-Bench v2\.1 by\+0\.028\+0\.028, andτ3\\tau^\{3\}\-Banking by\+0\.056\+0\.056, whileτ2\\tau^\{2\}\-Bench’s interval includes zero \(Table[1](https://arxiv.org/html/2608.29420#S5.T1), Figure[3](https://arxiv.org/html/2608.29420#S5.F3)b\)\. H4 is therefore supported on the rule fixed in advance \(Section[4\.3](https://arxiv.org/html/2608.29420#S4.SS3)\), which answers RQ4 affirmatively\. In direction this agrees with[Ruan et al\. \(2024\)](https://arxiv.org/html/2608.29420#bib.bib41), and in size it qualifies them\.

Table 1:Predictor\-ladder cross\-validated errors \(economic block\) and the H4 bootstrap test\.Left, train and test RMSE on standardised targets, out\-of\-foldR2R^\{2\}, and the best learner per rung \(RF random forest, EN elastic net\)\. Right,Δ​MSE=MSE⁡\(mean index\)−MSE⁡\(k​factors\)\\Delta\\mathrm\{MSE\}=\\mathrm\{MSE\}\(\\text\{mean index\}\)\-\\mathrm\{MSE\}\(k\\ \\text\{factors\}\),B=2000B=2000; every interval exceptτ2\\tau^\{2\}\-Bench’s excludes zero\.
Figure 3:Predictive validity under the LOBO protocol\.\(a\) Test RMSE by rung, economic block \(red\) and all twelve targets \(grey\)\. \(b\) BootstrapΔ​MSE\\Delta\\mathrm\{MSE\}\(single index minuskkfactors\) with 95% intervals; positive favours the multi\-factor representation\. \(c\) SHAP \(SHapley Additive exPlanations\) importances for the rung\-\(iv\) gradient\-boosting model of GDPval\.An improvement of 0\.037 sits on top of a single index that already explains 77% of economic\-block variance, so the gain is incremental, yet it survives a validation that rewards transfer and penalises memorisation\. The one exception,τ2\\tau^\{2\}\-Bench, carries the widest interval, and its null reflects limited coverage as much as absent signal\. The 96 configurations mix reasoning and effort variants of base models across providers; two design features address their comparability, since one operator scores every model under a single harness\([Artificial Analysis, 2026](https://arxiv.org/html/2608.29420#bib.bib2)\)and the deduplicated re\-run reproduces the gain\. On GDPval, SHAP attribution\([Lundberg and Lee, 2017](https://arxiv.org/html/2608.29420#bib.bib28)\)on the rung\-\(iv\) boosting learner ranks the agentic factor F1 first \(Figure[3](https://arxiv.org/html/2608.29420#S5.F3)c, Appendix[M](https://arxiv.org/html/2608.29420#A13)\)\.

### 5\.5Robustness

The headline structure is stable under perturbations to the unit of analysis, the correlation measure, the adjustment covariates, and the factor count \(Appendices[J](https://arxiv.org/html/2608.29420#A10)and[L](https://arxiv.org/html/2608.29420#A12)\)\. Collapsing the 421 configurations to one row per base model raises the first\-factor share to 90\.2% and the date\-adjustment drop to 24\.1 points, rank\-based and logit\-transformed correlations raise the share to 94\.5% and 91\.8%, so H1 is no metric artefact\([Schaeffer et al\., 2023](https://arxiv.org/html/2608.29420#bib.bib42)\), the predictor ladder’s ordering holds on all twelve targets, and the H4 verdict depends on neither the regularisation strength nor the factor count\.

## 6Discussion

Taken together, the two tasks answer the four research questions\. On RQ1 the battery is governed by one dominant axis holding 74\.5% of common variance; on RQ2 that axis is substantially a release\-date artefact, as removing the date trend takes away about a fifth of it \(74\.5% to 59\.6%\)\. On RQ3 they form no separate capability in the pre\-specified structure and separate only under an over\-extraction, yet on RQ4 the secondary economic factor still predicts held\-out economic scores better than a single general index on three of four benchmarks\. Economic validity is therefore limited, in the precise sense that a date\-driven progress factor accounts for most of the economic signal\.

For leaderboard use, a single intelligence index captures most of what distinguishes models, yet the residual economic signal is the part most relevant to claims about professional value, so reporting economic benchmarks separately is justified\. Small gaps between contemporaneous models still call for caution\. A leaderboard ranking models released months apart is ranking them largely on release date, an external factor set by the pace at which new models and capabilities arrive\. A date\-adjusted score or a comparison restricted to one release window would show how much of a gap is capability and how much is calendar\. In both respects the findings reinforce the call for psychometric scrutiny of AI benchmarks\([Bean et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib5);[Kearns, 2026](https://arxiv.org/html/2608.29420#bib.bib25)\)and the evidence that benchmarks lose discriminating power as they age\([Akhtar et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib1)\)\.

#### Contributing a two\-test protocol for benchmark builders\.

Beyond how leaderboards are read, the result points to how new economic benchmarks should be validated, since a benchmark’s value for distinguishing contemporaneous models depends on the variance it adds once the shared date trend is removed\. A builder who wants to show that a new suite measures something distinct can apply the two tests directly with what a leaderboard already publishes, per\-model scores on the new benchmark and the existing battery and their release dates\. Step one, the structural test, regresses every benchmark on release date, fits an oblique factor model to the residuals, and checks whether the new benchmark’s loading separates from the general factor; step two, the predictive test, holds the new benchmark out, predicts it from akk\-factor representation of the others with factors re\-estimated in every fold, and checks whether that beats a single mean\-score index with a bootstrap interval excluding zero\. A benchmark that passes both adds a construct, and one that passes neither is re\-measuring the progress trend\. An operator can run both tests on every benchmark it adds and publish the two statistics beside the score, as the code accompanying this paper does from a single score matrix \([https://github\.com/louisyzhu/frontier\-ai\-economic\-validity](https://github.com/louisyzhu/frontier-ai-economic-validity)\)\.

#### Limitations\.

Four limitations bound the claims\. First, the economic factor comes from an exploratory over\-extraction, so it stands as a hypothesis for a confirmatory factor model on an independent snapshot\. Second, the analysis is observational, so it establishes how strongly release date correlates with the general factor and not what drives that association\. Third, coverage is uneven, so the economic results carry wider uncertainty; fourth, the results describe one leaderboard on one date and will shift as new models arrive, a risk the hash\-pinned snapshot mitigates\.

#### Conclusion\.

Economic benchmarks measure neither one capability nor many\. Most of what they record is a single general factor that is largely a release\-date trend, and what remains is a small but real economic component that predicts held\-out economic scores\. Read with the date in view they add information to a general index without yet constituting a distinct capability\. Every test and threshold here was fixed before analysis, every verdict is reported as it fell, and every number is reproducible from pinned inputs, which is the standard we propose for evaluation claims of this kind\.

## Acknowledgements

I thank Marcos Barreto for comments and guidance on this work\.

## References

- Akhtar et al\. \[2026\]Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, et al\.When AI benchmarks plateau: A systematic study of benchmark saturation\.In*Proceedings of the 43rd International Conference on Machine Learning \(ICML\)*, 2026\.arXiv:2602\.16763\.
- Artificial Analysis \[2026\]Artificial Analysis\.Artificial Analysis LLM Leaderboard and Intelligence Index\.[https://artificialanalysis\.ai](https://artificialanalysis.ai/), 2026\.Snapshot accessed 6 July 2026\.
- Barres et al\. \[2026\]Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan\.τ2\\tau^\{2\}\-Bench: Evaluating conversational agents in a dual\-control environment\.In*Proceedings of the 43rd International Conference on Machine Learning \(ICML\)*, 2026\.arXiv:2506\.07982\.
- Bartlett \[1950\]Maurice S\. Bartlett\.Tests of significance in factor analysis\.*British Journal of Psychology \(Statistical Section\)*, 3:77–85, 1950\.
- Bean et al\. \[2025\]Andrew M\. Bean, Ryan O\. Kearns, Angelika Romanou, Franziska S\. Hafner, Harry Mayne, et al\.Measuring what matters: Construct validity in large language model benchmarks\.In*Advances in Neural Information Processing Systems 38 \(NeurIPS 2025\), Datasets and Benchmarks Track*, 2025\.arXiv:2511\.04703\.
- Biderman et al\. \[2024\]Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, et al\.Lessons from the trenches on reproducible evaluation of language models, 2024\.arXiv:2405\.14782\.
- Bowman and Dahl \[2021\]Samuel R\. Bowman and George E\. Dahl\.What will it take to fix benchmarking in natural language understanding?In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics \(NAACL\)*, pages 4843–4855, 2021\.
- Breiman \[2001\]Leo Breiman\.Random forests\.*Machine Learning*, 45\(1\):5–32, 2001\.
- Burnell et al\. \[2023\]Ryan Burnell, Han Hao, Andrew R\. A\. Conway, and José Hernández\-Orallo\.Revealing the structure of language model capabilities, 2023\.arXiv:2306\.10062\.
- Campbell and Fiske \[1959\]Donald T\. Campbell and Donald W\. Fiske\.Convergent and discriminant validation by the multitrait\-multimethod matrix\.*Psychological Bulletin*, 56\(2\):81–105, 1959\.
- Cattell \[1966\]Raymond B\. Cattell\.The scree test for the number of factors\.*Multivariate Behavioral Research*, 1\(2\):245–276, 1966\.
- Cronbach and Meehl \[1955\]Lee J\. Cronbach and Paul E\. Meehl\.Construct validity in psychological tests\.*Psychological Bulletin*, 52\(4\):281–302, 1955\.
- Efron and Tibshirani \[1993\]Bradley Efron and Robert J\. Tibshirani\.*An Introduction to the Bootstrap*\.Chapman & Hall, 1993\.
- Epoch AI \[2026\]Epoch AI\.Notable AI Models \(data set, CC\-BY 4\.0\)\.[https://epoch\.ai/data/notable\-ai\-models](https://epoch.ai/data/notable-ai-models), 2026\.
- Friedman \[2001\]Jerome H\. Friedman\.Greedy function approximation: a gradient boosting machine\.*Annals of Statistics*, 29\(5\):1189–1232, 2001\.
- Hastie et al\. \[2009\]Trevor Hastie, Robert Tibshirani, and Jerome Friedman\.*The Elements of Statistical Learning: Data Mining, Inference, and Prediction*\.Springer, 2nd edition, 2009\.
- Hernández\-Orallo \[2017\]José Hernández\-Orallo\.*The Measure of All Minds: Evaluating Natural and Artificial Intelligence*\.Cambridge University Press, 2017\.
- Hoerl and Kennard \[1970\]Arthur E\. Hoerl and Robert W\. Kennard\.Ridge regression: biased estimation for nonorthogonal problems\.*Technometrics*, 12\(1\):55–67, 1970\.
- Horn \[1965\]John L\. Horn\.A rationale and test for the number of factors in factor analysis\.*Psychometrika*, 30\(2\):179–185, 1965\.
- Ilić and Gignac \[2024\]David Ilić and Gilles E\. Gignac\.Evidence of interrelated cognitive\-like capabilities in large language models: indications of artificial general intelligence or achievement?*Intelligence*, 106:101858, 2024\.
- Jacobs and Wallach \[2021\]Abigail Z\. Jacobs and Hanna Wallach\.Measurement and fairness\.In*Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency*, pages 375–385, 2021\.
- Jennrich and Sampson \[1966\]Robert I\. Jennrich and Paul F\. Sampson\.Rotation for simple loadings\.*Psychometrika*, 31\(3\):313–323, 1966\.
- Kaiser \[1960\]Henry F\. Kaiser\.The application of electronic computers to factor analysis\.*Educational and Psychological Measurement*, 20\(1\):141–151, 1960\.
- Kaiser \[1974\]Henry F\. Kaiser\.An index of factorial simplicity\.*Psychometrika*, 39\(1\):31–36, 1974\.
- Kearns \[2026\]Ryan O\. Kearns\.Quantifying construct validity in large language model evaluations, 2026\.MSc thesis, Oxford Internet Institute, University of Oxford, Trinity Term 2025\. arXiv:2602\.15532\.
- Kiela et al\. \[2021\]Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, et al\.Dynabench: Rethinking benchmarking in NLP\.In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics \(NAACL\)*, pages 4110–4124, 2021\.
- Liang et al\. \[2023\]Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, et al\.Holistic evaluation of language models\.*Transactions on Machine Learning Research*, 2023\.arXiv:2211\.09110\.
- Lundberg and Lee \[2017\]Scott M\. Lundberg and Su\-In Lee\.A unified approach to interpreting model predictions\.In*Advances in Neural Information Processing Systems 30 \(NeurIPS 2017\)*, pages 4765–4774, 2017\.
- Merrill et al\. \[2026\]Mike A\. Merrill, Alexander G\. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, et al\.Terminal\-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026\.arXiv:2601\.11868\.
- Messick \[1998\]Samuel Messick\.Test validity: a matter of consequence\.*Social Indicators Research*, 45:35–44, 1998\.
- Miller \[2024\]Evan Miller\.Adding error bars to evals: A statistical approach to language model evaluations, 2024\.arXiv:2411\.00640\.
- Patwardhan et al\. \[2026\]Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, et al\.GDPval: Evaluating AI model performance on real\-world economically valuable tasks\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.arXiv:2510\.04374\.
- Pedregosa et al\. \[2011\]Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, et al\.Scikit\-learn: Machine learning in Python\.*Journal of Machine Learning Research*, 12:2825–2830, 2011\.
- Phan et al\. \[2025\]Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, et al\.Humanity’s Last Exam, 2025\.arXiv:2501\.14249\.
- Pyatkin et al\. \[2025\]Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, and Hannaneh Hajishirzi\.Generalizing verifiable instruction following, 2025\.arXiv:2507\.02833\.
- Raji et al\. \[2021\]Inioluwa Deborah Raji, Emily M\. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna\.AI and the everything in the whole wide world benchmark\.In*Advances in Neural Information Processing Systems 34 \(NeurIPS 2021\), Datasets and Benchmarks Track*, 2021\.
- Rein et al\. \[2024\]David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R\. Bowman\.GPQA: A graduate\-level Google\-proof Q&A benchmark\.In*First Conference on Language Modeling \(COLM\)*, 2024\.arXiv:2311\.12022\.
- Reuel et al\. \[2024\]Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J\. Kochenderfer\.BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices\.In*Advances in Neural Information Processing Systems 37 \(NeurIPS 2024\), Datasets and Benchmarks Track*, 2024\.
- Reuel et al\. \[2025\]Anka Reuel, Avijit Ghosh, Jenny Chim, Andrew Tran, et al\.Who evaluates AI’s social impacts? Mapping coverage and gaps in first and third party evaluations, 2025\.Preprint, Evaluating Evaluations \(EvalEval\) Coalition\. arXiv:2511\.05613\.
- Rousseeuw \[1987\]Peter J\. Rousseeuw\.Silhouettes: a graphical aid to the interpretation and validation of cluster analysis\.*Journal of Computational and Applied Mathematics*, 20:53–65, 1987\.
- Ruan et al\. \[2024\]Yangjun Ruan, Chris J\. Maddison, and Tatsunori Hashimoto\.Observational scaling laws and the predictability of language model performance\.In*Advances in Neural Information Processing Systems 37 \(NeurIPS 2024\)*, 2024\.
- Schaeffer et al\. \[2023\]Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo\.Are emergent abilities of large language models a mirage?In*Advances in Neural Information Processing Systems 36 \(NeurIPS 2023\)*, 2023\.
- Schwarz \[1978\]Gideon Schwarz\.Estimating the dimension of a model\.*The Annals of Statistics*, 6\(2\):461–464, 1978\.
- Shi et al\. \[2026\]Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan, and Victor Barres\.τ\\tau\-Knowledge: Evaluating conversational agents over unstructured knowledge, 2026\.arXiv:2603\.04370\.
- Spearman \[1904\]Charles Spearman\.“General Intelligence,” Objectively Determined and Measured\.*The American Journal of Psychology*, 15\(2\):201–292, 1904\.
- Thurstone \[1947\]Louis L\. Thurstone\.*Multiple\-Factor Analysis*\.University of Chicago Press, 1947\.
- Tian et al\. \[2024\]Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, et al\.SciCode: A research coding benchmark curated by scientists\.In*Advances in Neural Information Processing Systems 37 \(NeurIPS 2024\), Datasets and Benchmarks Track*, 2024\.arXiv:2407\.13168\.
- Tibshirani et al\. \[2001\]Robert Tibshirani, Guenther Walther, and Trevor Hastie\.Estimating the number of clusters in a data set via the gap statistic\.*Journal of the Royal Statistical Society: Series B*, 63\(2\):411–423, 2001\.
- Varma and Simon \[2006\]Sudhir Varma and Richard Simon\.Bias in error estimation when using cross\-validation for model selection\.*BMC Bioinformatics*, 7:91, 2006\.
- Vidgen et al\. \[2026\]Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, et al\.APEX\-Agents, 2026\.arXiv:2601\.14242\.
- Weidinger et al\. \[2025\]Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, et al\.Toward an evaluation science for generative AI systems, 2025\.arXiv:2503\.05336\.
- Yao et al\. \[2025\]Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan\.τ\\tau\-bench: A benchmark for tool\-agent\-user interaction in real\-world domains\.In*International Conference on Learning Representations \(ICLR\)*, 2025\.arXiv:2406\.12045\.
- Yue et al\. \[2025\]Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, et al\.MMMU\-Pro: A more robust multi\-discipline multimodal understanding benchmark\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 15134–15186, 2025\.arXiv:2409\.02813\.
- Zhou et al\. \[2025\]Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez\-Plumed, et al\.General scales unlock AI evaluation with explanatory and predictive power, 2025\.arXiv:2503\.06378\.
- Zhu et al\. \[2025\]Minhui Zhu, Minyang Tian, Xiaocheng Yang, Tianci Zhou, Lifan Yuan, Penghao Zhu, Eli Chertkov, Shengyan Liu, Yufeng Du, Ziming Ji, et al\.Probing the critical point \(CritPt\) of AI reasoning: A frontier physics research benchmark, 2025\.arXiv:2509\.26574\.
- Zou and Hastie \[2005\]Hui Zou and Trevor Hastie\.Regularization and variable selection via the elastic net\.*Journal of the Royal Statistical Society: Series B*, 67\(2\):301–320, 2005\.

## Appendix

The appendix collects the material that supports the main text without interrupting its argument\. Each section opens with the main\-text section it supports, and every figure and table is placed with the section that discusses it\.

## Appendix AAbbreviations and notation

Table 2:Abbreviations and symbols, in order of first appearance in the main text\.
## Appendix BLeave\-one\-benchmark\-out procedure

*Supports Section[4\.2](https://arxiv.org/html/2608.29420#S4.SS2), which describes the protocol in words\.*Figure[4](https://arxiv.org/html/2608.29420#A2.F4)states the complete procedure, including the two safeguards on which the predictive test depends\. First, the factor model is re\-fitted inside every outer training fold, so the held\-out models never influence the loadings used to score them\. Second, hyperparameters are tuned in an inner loop nested within the outer split, so the reported out\-of\-fold error is never used for selection\[[Varma and Simon, 2006](https://arxiv.org/html/2608.29420#bib.bib49)\]\.

ProcedureLeave\-one\-benchmark\-out predictive validity with nested cross\-validationInput:standardised score matrixZZ\(models×\\timesbenchmarks\), economic targetsTT, learnersLL, rungsRR\.foreach targett∈Tt\\in Tdo partition models into 5 outer foldsforeach outer fold \(train, test\)do fit standardisation and EFA on train only; form Thurstone weightsW=R−1​ΛW=R^\{\-1\}\\Lambdaproject test models throughWW; orient factors by train mean indexforeach rungr∈Rr\\in Rand learnerℓ∈L\\ell\\in Ldo tuneℓ\\ellby inner 5\-fold grid search on train; predict test, store out\-of\-fold errorselect best learner per rung by pooled out\-of\-fold RMSEcomputeΔ​MSE=MSE⁡\(rung ii\)−MSE⁡\(rung iv\)\\Delta\\mathrm\{MSE\}=\\mathrm\{MSE\}\(\\text\{rung ii\}\)\-\\mathrm\{MSE\}\(\\text\{rung iv\}\); bootstrapB=2000B=2000for 95% CIOutput:per\-target and pooledΔ​MSE\\Delta\\mathrm\{MSE\}with confidence intervals \(H4\)\.

Figure 4:The leave\-one\-benchmark\-out procedure with leakage\-safe factor projection and nested tuning\.
## Appendix CPositionality, directionality, and threats to validity

*Supports Section[4\.3](https://arxiv.org/html/2608.29420#S4.SS3)\.*A factor solution is identified only up to sign, so we state the direction of every factor explicitly\. We orient each factor to correlate positively with the model’s mean standardised score\. “Higher” then always means “more capable”, and every loading in Table[3](https://arxiv.org/html/2608.29420#A5.T3)reads in the same direction\. The choice is not cosmetic, since an un\-oriented solution could report the same economic factor with flipped signs and invite the opposite interpretation\. The direction of the date adjustment is equally deliberate\. We regress scores on release date and analyse the residuals, which measures how much benchmark co\-movement survives once the release date is known\. We do not run the reverse regression, as we make no claim that the general factor causes the calendar\. Finally, we fix the sign of the predictive contrast in advance, so thatΔ​MSE\\Delta\\mathrm\{MSE\}is the single\-index error minus thekk\-factor error, a positive value favours the richer representation, and H4 requires this positive interval to exclude zero\.

Two threats to validity follow from the study’s vantage point\. First, we observe only the models that vendors chose to submit to a public leaderboard\. The sample is therefore a convenience sample of frontier systems, and the claims generalise to models on this leaderboard alone\. Second, the economic benchmarks are newer and sparser, so any economic\-specific structure is estimated from fewer observations and is more sensitive to the particular models that carry those scores\. The wide bootstrap intervals in Table[1](https://arxiv.org/html/2608.29420#S5.T1)quantify that limitation\.

## Appendix DMathematical basis of the two tasks

*Supports Sections[4\.1](https://arxiv.org/html/2608.29420#S4.SS1)and[4\.2](https://arxiv.org/html/2608.29420#S4.SS2)\.*Each subsection states a component of the method, derives the quantity the main text reports, and gives the reason the component was chosen over its nearest alternative\.

#### The common\-factor model and the dominance statistic\.

Let𝐳∈ℝp\\mathbf\{z\}\\in\\mathbb\{R\}^\{p\}be the vector ofppper\-benchmark standardised scores for a model\. The common\-factor model writes

𝐳=Λ​𝐟\+𝜺,Σ=Λ​Φ​Λ⊤\+Ψ,\\mathbf\{z\}=\\Lambda\\mathbf\{f\}\+\\boldsymbol\{\\varepsilon\},\\qquad\\Sigma=\\Lambda\\Phi\\Lambda^\{\\top\}\+\\Psi,withΛ∈ℝp×k\\Lambda\\in\\mathbb\{R\}^\{p\\times k\}the loading matrix,𝐟∈ℝk\\mathbf\{f\}\\in\\mathbb\{R\}^\{k\}the latent factors withCov⁡\(𝐟\)=Φ\\mathrm\{Cov\}\(\\mathbf\{f\}\)=\\Phi, and𝜺\\boldsymbol\{\\varepsilon\}benchmark\-specific uniquenesses with diagonalCov⁡\(𝜺\)=Ψ\\mathrm\{Cov\}\(\\boldsymbol\{\\varepsilon\}\)=\\PsiandCov⁡\(𝐟,𝜺\)=𝟎\\mathrm\{Cov\}\(\\mathbf\{f\},\\boldsymbol\{\\varepsilon\}\)=\\mathbf\{0\}\. The decomposition ofΣ\\Sigmainto a common partΛ​Φ​Λ⊤\\Lambda\\Phi\\Lambda^\{\\top\}and a unique partΨ\\Psiis the reason we use factor analysis and not principal components for a validity question\. A principal component maximises total variance and so absorbs benchmark\-specific noise along with shared capability, whereas the common\-factor model attributes to the factors only the variance that benchmarks share\[[Thurstone, 1947](https://arxiv.org/html/2608.29420#bib.bib46)\]\. Parameters are estimated by maximum likelihood under multivariate normality\. The dominance statistic for H1 is computed on the unrotated solution, whereΦ=I\\Phi=Iand each factor’s contribution to common variance is its sum of squared loadings, themm\-th diagonal entry ofΛ⊤​Λ\\Lambda^\{\\top\}\\Lambda,

s1=∑j=1pλj​12∑m=1k∑j=1pλj​m2=\(Λ⊤​Λ\)11tr⁡\(Λ⊤​Λ\)\.s\_\{1\}=\\frac\{\\sum\_\{j=1\}^\{p\}\\lambda\_\{j1\}^\{2\}\}\{\\sum\_\{m=1\}^\{k\}\\sum\_\{j=1\}^\{p\}\\lambda\_\{jm\}^\{2\}\}=\\frac\{\(\\Lambda^\{\\top\}\\Lambda\)\_\{11\}\}\{\\mathrm\{tr\}\(\\Lambda^\{\\top\}\\Lambda\)\}\.On the complete\-case grid the three factors carry sums of squared loadings 7\.79, 2\.34, and 0\.32, sos1=7\.79/10\.45=74\.5%s\_\{1\}=7\.79/10\.45=74\.5\\%\. This is a factor\-level share, distinct from a single benchmark’s communality∑mλj​m2\\sum\_\{m\}\\lambda\_\{jm\}^\{2\}, a diagonal entry ofΛ​Λ⊤\\Lambda\\Lambda^\{\\top\}\.

#### Dimensionality by parallel analysis\.

Letℓ1≥⋯≥ℓp\\ell\_\{1\}\\geq\\dots\\geq\\ell\_\{p\}be the eigenvalues of the sample correlation matrixRR, and letℓm\(0\.95\)\\ell\_\{m\}^\{\(0\.95\)\}be the 95th percentile of themm\-th eigenvalue over correlation matrices of independent standard\-normal data of the same dimensionsn×pn\\times p\. Parallel analysis retains factormmwhenℓm\>ℓm\(0\.95\)\\ell\_\{m\}\>\\ell\_\{m\}^\{\(0\.95\)\}\[[Horn, 1965](https://arxiv.org/html/2608.29420#bib.bib19)\]\. The rule asks whether a factor explains more than chance sampling structure would, which is why it is more conservative than the Kaiser ruleℓm\>1\\ell\_\{m\}\>1\[[Kaiser, 1960](https://arxiv.org/html/2608.29420#bib.bib23)\]on strongly collinear data, where sampling eigenvalues of later factors fall below one and the Kaiser rule over\-retains\. Information criteria such as the BIC\[[Schwarz, 1978](https://arxiv.org/html/2608.29420#bib.bib43)\]compare likelihoods of nested models and reward any factor that improves fit, which on a near\-rank\-one correlation matrix means retaining factors that carry very little common variance; the BIC’s preference fork=4k=4here \(Section[4\.3](https://arxiv.org/html/2608.29420#S4.SS3)\) is an instance\. On our gridℓ1=9\.53\\ell\_\{1\}=9\.53againstℓ1\(0\.95\)=1\.79\\ell\_\{1\}^\{\(0\.95\)\}=1\.79, andℓ2=0\.91\\ell\_\{2\}=0\.91againstℓ2\(0\.95\)=1\.57\\ell\_\{2\}^\{\(0\.95\)\}=1\.57, so one factor is retained \(Figure[5](https://arxiv.org/html/2608.29420#A4.F5)\)\.

Figure 5:Dimensionality of the battery\.\(a\) The scree plot of principal\-component variance shares drops sharply after the first component \(79\.4%\)\. \(b\) Horn parallel analysis retains a single factor, as only the first observed eigenvalue \(9\.53\) exceeds its random\-data 95th percentile \(1\.79\), while the second \(0\.91\) already falls below \(1\.57\)\.
#### Oblique rotation\.

The maximum\-likelihood solution is identified only up to rotation\. Direct oblimin\[[Jennrich and Sampson, 1966](https://arxiv.org/html/2608.29420#bib.bib22)\]chooses the rotation that minimises

Q⁡\(Λ\)=∑m<l\(∑j=1pλj​m2​λj​l2−γp​∑j=1pλj​m2​∑j=1pλj​l2\),Q\(\\Lambda\)=\\sum\_\{m<l\}\\Bigl\(\\sum\_\{j=1\}^\{p\}\\lambda\_\{jm\}^\{2\}\\lambda\_\{jl\}^\{2\}\-\\frac\{\\gamma\}\{p\}\\sum\_\{j=1\}^\{p\}\\lambda\_\{jm\}^\{2\}\\sum\_\{j=1\}^\{p\}\\lambda\_\{jl\}^\{2\}\\Bigr\),withγ=0\\gamma=0giving the quartimin criterion we use, while allowingΦ≠I\\Phi\\neq I\. An orthogonal rotation such as varimax forcesΦ=I\\Phi=I; on data with a strong general factor this distributes shared variance across artificially independent axes, so that no single factor can carry the general variance and the dominance we set out to measure is obscured by construction\. The date\-adjusted factors correlate at 0\.67 and 0\.72 \(Section[5\.3](https://arxiv.org/html/2608.29420#S5.SS3)\), which confirms that the oblique choice was the right one\.

#### Factor scores and the leakage\-safe projection\.

Regression \(Thurstone\) factor scores are the linear predictor𝐟^=W⊤​𝐳\\hat\{\\mathbf\{f\}\}=W^\{\\top\}\\mathbf\{z\}that minimises𝔼​∥𝐟−W⊤​𝐳∥2\\mathbb\{E\}\\lVert\\mathbf\{f\}\-W^\{\\top\}\\mathbf\{z\}\\rVert^\{2\}\. Setting the derivative to zero givesW=Σ−1​Cov​\(𝐳,𝐟\)=R−1​Λ​ΦW=\\Sigma^\{\-1\}\\mathrm\{Cov\}\(\\mathbf\{z\},\\mathbf\{f\}\)=R^\{\-1\}\\Lambda\\Phi, whereRRis the sample correlation matrix,Λ\\Lambdathe pattern loadings, andΦ\\Phithe inter\-factor correlation matrix of the sorted solution, so thatW=R−1​ΛW=R^\{\-1\}\\Lambdawhen the factors are orthogonal\[[Thurstone, 1947](https://arxiv.org/html/2608.29420#bib.bib46)\]\. The structural factor scores behind H2 and the model clustering use the oblique form\. The prediction task uses the orthogonal\-form weightsW=R−1​ΛW=R^\{\-1\}\\Lambdabuilt from the pattern loadings; thereWWis estimated on the training fold only and applied to the held\-out fold, the leakage\-safe projection of Section[4\.2](https://arxiv.org/html/2608.29420#S4.SS2)\. The two score sets are an invertible reparameterisation of one another,𝐟^oblique=𝐟^pattern​Φ\\hat\{\\mathbf\{f\}\}\_\{\\mathrm\{oblique\}\}=\\hat\{\\mathbf\{f\}\}\_\{\\mathrm\{pattern\}\}\\Phi, so an unpenalised regression on either gives identical fitted values; the ridge learners of Task 2 are not invariant to the reparameterisation, and the reported ladder uses the pattern form throughout\. The point is thatWWis a function of the data\. If it were estimated on allnnmodels, every held\-out model’s score would have contributed to the weights used to predict it, and the predictive comparison in Task 2 would be contaminated by exactly the information it is meant to withhold\.

#### Date adjustment as removal of a rank\-one component\.

Each standardised benchmarkzjz\_\{j\}is regressed on release datedd\(and, in a subsample, additionally on log training compute\),

zj=β0​j\+β1​j​d\+rj,z\_\{j\}=\\beta\_\{0j\}\+\\beta\_\{1j\}\\,d\+r\_\{j\},and the factor model is re\-fitted on the residualsrjr\_\{j\}\. Writing𝐛=\(β11,…,β1​p\)⊤\\mathbf\{b\}=\(\\beta\_\{11\},\\dots,\\beta\_\{1p\}\)^\{\\top\}, the covariance of the residual vector isΣr=Σ−σd2​𝐛𝐛⊤\\Sigma\_\{r\}=\\Sigma\-\\sigma\_\{d\}^\{2\}\\,\\mathbf\{b\}\\mathbf\{b\}^\{\\top\}, so the adjustment removes from the benchmark covariance exactly the rank\-one component that lies along the date vector\. If the general factor were nothing but the shared time trend,Λ1\\Lambda\_\{1\}would be proportional to𝐛\\mathbf\{b\}and the first factor’s share would collapse after adjustment; if it were unrelated to date, the share would be unchanged\. The observed change,74\.5%→59\.6%74\.5\\%\\rightarrow 59\.6\\%, is the H2\(ii\) statistic and measures where the battery sits between those two cases\. The design mirrors the scaling\-law residualisation of[Kearns \[2026\]](https://arxiv.org/html/2608.29420#bib.bib25), with release date in place of parameter count as the nuisance variable\.

#### Nested leave\-one\-benchmark\-out estimator\.

For an economic targettt, the out\-of\-fold error of rungrris

MSEr=1n​∑i=1n\(zi​t−z^i​t\(−κ⁡\(i\)\)\)2,\\mathrm\{MSE\}\_\{r\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\bigl\(z\_\{it\}\-\\hat\{z\}\_\{it\}^\{\(\-\\kappa\(i\)\)\}\\bigr\)^\{2\},wherez^i​t\(−κ⁡\(i\)\)\\hat\{z\}\_\{it\}^\{\(\-\\kappa\(i\)\)\}is the prediction for modeliifrom a learner trained on the outer folds excluding the foldκ⁡\(i\)\\kappa\(i\)that containsii, with hyperparameters chosen by an inner five\-fold grid search on those training folds alone\. Selecting hyperparameters on the same folds that supply the reported error biases the estimate optimistically, and the bias grows with the size of the search space\[[Varma and Simon, 2006](https://arxiv.org/html/2608.29420#bib.bib49)\]; nesting removes it at the cost of a five\-fold increase in computation\. The predictive\-validity statistic contrasts the single\-index rung against thekk\-factor rung,Δ​MSE=MSEii−MSEiv\\Delta\\mathrm\{MSE\}=\\mathrm\{MSE\}\_\{\\mathrm\{ii\}\}\-\\mathrm\{MSE\}\_\{\\mathrm\{iv\}\}\.

#### Paired bootstrap forΔ​MSE\\Delta\\mathrm\{MSE\}\.

Leteiiie\_\{i\}^\{\\mathrm\{ii\}\}andeiive\_\{i\}^\{\\mathrm\{iv\}\}be modelii’s out\-of\-fold squared errors under the two rungs\. A 95% interval forΔ​MSE\\Delta\\mathrm\{MSE\}is obtained by resampling models with replacement, recomputing1n​∑i\(eiii−eiiv\)\\frac\{1\}\{n\}\\sum\_\{i\}\(e\_\{i\}^\{\\mathrm\{ii\}\}\-e\_\{i\}^\{\\mathrm\{iv\}\}\)on each ofB=2000B=2000resamples, and taking the 2\.5th and 97\.5th percentiles\[[Efron and Tibshirani, 1993](https://arxiv.org/html/2608.29420#bib.bib13)\]\. Resampling the paired differences, and not the two error vectors separately, matters because the two rungs are evaluated on the same models and their errors are strongly correlated; the paired design removes the between\-model variance that the two rungs share, so the interval reflects only the variance of the contrast\. H4 is supported when this interval excludes zero\. Because the resampling unit is a model, near\-duplicate configurations of one base model would be treated as independent draws; the deduplicated re\-run in Appendix[L](https://arxiv.org/html/2608.29420#A12), whose resampling unit is a distinct base model, is the check on that assumption\.

## Appendix EDate\-adjusted three\-factor loadings

*Supports Section[5\.3](https://arxiv.org/html/2608.29420#S5.SS3)\.*Table[3](https://arxiv.org/html/2608.29420#A5.T3)tabulates the loadings shown as a heatmap in Figure[2](https://arxiv.org/html/2608.29420#S5.F2)b\. F1 is an agentic and work\-realistic factor, F2 an academic\-knowledge factor, and F3 a physics and hard\-reasoning factor; all four economic benchmarks load highest on F1\.

Table 3:Date\-adjusted three\-factor oblimin loadings, economic benchmarks in bold\.
## Appendix FBenchmark battery

*Supports Section[3](https://arxiv.org/html/2608.29420#S3)\.*Table[4](https://arxiv.org/html/2608.29420#A6.T4)describes every candidate benchmark, the construct it targets, the type of score on which the leaderboard reports it, its coverage before and after the inclusion rule, and the role the rule assigned, with the source paper for each listed beneath\. The economic benchmarks differ from the academic ones in kind, not only in coverage\. They score multi\-step, tool\-using, or human\-preference\-judged task completion on work\-like tasks, whereas the academic benchmarks score single\-response accuracy on knowledge or reasoning items\. Three of the twelve \(AA\-Omniscience, AA\-LCR, and Terminal\-Bench Hard\) are constructed or subset by the leaderboard operator itself; all twelve are run by that operator under a single harness\[[Artificial Analysis, 2026](https://arxiv.org/html/2608.29420#bib.bib2)\]\.

Table 4:Candidate benchmarks by capability block, with the construct each targets, the type of score reported, the number of models scored on the raw 548\-configuration snapshot and among the 421 retained, and the role assigned by the inclusion rule\. Twelve benchmarks are primary; MMMU\-Pro is sensitivity\-only for sparse and multimodal coverage; APEX\-Agents falls below the sixty\-model floor\. Higher is better throughout; sources are listed below the table\.BenchmarkConstruct measuredScore typeRawRetainedRole*Economic*GDPval \(Elo\)Paid professional task qualitypairwise\-judged Elo117112primaryTerminal\-Bench v2\.1Terminal and agentic task completionpass rate121121primaryτ3\\tau^\{3\}\-BankingMulti\-step banking workflowstask success112112primaryτ2\\tau^\{2\}\-BenchTool\-use agentic taskstask success428413primaryAPEX\-AgentsAgentic task completiontask success26–dropped \(sparse\)*Academic*GPQA DiamondGraduate\-level science QAaccuracy513421primaryHLEFrontier academic knowledgeaccuracy509421primaryAA\-OmniscienceBroad knowledge, hallucination\-penalisedpenalised score418418primaryMMMU\-ProMultimodal understandingaccuracy204–sensitivity only*Scientific coding*SciCodeResearch\-level scientific codingpass rate507421primaryCritPtFrontier physics problem solvingaccuracy422419primaryTerminal\-Bench HardHard terminal coding taskspass rate421413primary*Long\-context / instruction*AA\-LCRLong\-context reasoningaccuracy441420primaryIFBenchInstruction followingaccuracy437417primary*Sources\.*GDPval\[[Patwardhan et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib32)\]; Terminal\-Bench v2\.1\[[Merrill et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib29)\];τ3\\tau^\{3\}\-Banking\[[Shi et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib44)\];τ2\\tau^\{2\}\-Bench\[[Barres et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib3)\]; GPQA Diamond\[[Rein et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib37)\]; HLE\[[Phan et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib34)\]; SciCode\[[Tian et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib47)\]; CritPt\[[Zhu et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib55)\]; IFBench\[[Pyatkin et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib35)\]; AA\-Omniscience, AA\-LCR, and Terminal\-Bench Hard are constructed or subset by the leaderboard operator\[[Artificial Analysis, 2026](https://arxiv.org/html/2608.29420#bib.bib2)\]\.

GDPval’s Elo rating comes from pairwise, head\-to\-head judgements of model outputs on economically valuable tasks, so a higher rating means more often preferred on a scale with no fixed maximum\. Theτ\\tau\-Bench family\[[Yao et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib52),[Barres et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib3)\]scores an agent on whether the final database state and required confirmations match a policy\-compliant goal after a multi\-turn conversation with a simulated user;τ3\\tau^\{3\}\-Banking is the banking\-knowledge domain of its third generation, in which agents must complete policy\-compliant state changes over a large document base\. The leaderboard labels its Terminal\-Bench column v2\.1; the cited paper presents the suite’s 2\.0 release\. MMMU\-Pro\[[Yue et al\., 2025](https://arxiv.org/html/2608.29420#bib.bib53)\]is demoted to a sensitivity\-only role and APEX\-Agents\[[Vidgen et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib50)\]is dropped for sparsity \(Appendix[G](https://arxiv.org/html/2608.29420#A7)\)\.

## Appendix GFull coverage audit

*Supports the inclusion rule of Section[3](https://arxiv.org/html/2608.29420#S3)\.*Table[4](https://arxiv.org/html/2608.29420#A6.T4)records every candidate benchmark’s coverage on the raw snapshot and the decision applied to it, and Figure[6](https://arxiv.org/html/2608.29420#A7.F6)shows that coverage together with the pairwise overlaps that set the effective sample size for every correlation in the study\. The sixty\-model floor is the smallest count at which the sparsest retained benchmark still overlaps every other benchmark on more than one hundred models \(Figure[6](https://arxiv.org/html/2608.29420#A7.F6)b\), which keeps each pairwise correlation estimable with a standard error below about 0\.1\.

![Refer to caption](https://arxiv.org/html/2608.29420v1/fig_coverage.png)Figure 6:Coverage audit behind the inclusion rule\.\(a\) Number of models carrying a score on each candidate benchmark on the raw 548\-configuration snapshot, coloured by capability block, with the sixty\-model inclusion floor marked\. Three economic benchmarks \(Terminal\-Bench v2\.1, GDPval, andτ3\\tau^\{3\}\-Banking\) are scored on only 112 to 121 models, whileτ2\\tau^\{2\}\-Bench and the academic core are each scored on four hundred or more; APEX\-Agents, with 26 models, falls below the floor\. \(b\) Pairwise counts of models scored on both benchmarks in a pair, which set the effective sample size for every correlation in the study and are the evidence for the two\-grid design\.
## Appendix HExploratory data analysis

*Supports Section[3](https://arxiv.org/html/2608.29420#S3)\.*We examined coverage and missingness, correlation structure, and temporal patterns before fitting any model \(Figure[7](https://arxiv.org/html/2608.29420#A8.F7)\)\. All standardisation and model fitting uses the training portion of each split only, so these views characterise the data without leaking test information into the models\.

![Refer to caption](https://arxiv.org/html/2608.29420v1/fig_eda.png)Figure 7:Exploratory data analysis of the twelve\-benchmark battery on the complete\-case grid \(n=96n=96\)\.\(a\) Pairwise Spearman rank correlations, with benchmark labels coloured by capability block; almost every pair is strongly positive\. \(b\) Per\-benchmark scores \(min–max scaled\) against model release date, showing capability rising steeply with time\. \(c\) Coverage per benchmark across the 421 retained models, with the four economic benchmarks \(red\) scored on far fewer models than the academic core\.Figure 8:Univariate benchmark score distributions\.\(a\) Raw score histograms for the twelve benchmarks, coloured by capability block;γ\\gammais the Fisher skewness\. \(b\) Skewness ordered across benchmarks\. Most benchmarks are bounded accuracies in\[0,1\]\[0,1\], and several are strongly right\-skewed \(CritPtγ=3\.8\\gamma=3\.8, HLEγ=1\.5\\gamma=1\.5,τ3\\tau^\{3\}\-Bankingγ=0\.9\\gamma=0\.9\), with scores concentrated near the floor\. GDPval \(Elo\) and AA\-Omniscience are on unbounded scales and are closer to symmetric\.The battery is highly collinear before any modelling\. The mean off\-diagonal Spearman correlation isρ=0\.79\\rho=0\.79, every benchmark pair is positive, and the economic benchmarks correlate with the academic core almost as strongly as with one another \(Figure[7](https://arxiv.org/html/2608.29420#A8.F7)a\)\. A dominant general factor would produce this near\-block structure, and the structure analysis must explain it, since a strong positive manifold is equally consistent with one underlying capability and with several capabilities that improve together over time\. Scores rise steeply and jointly with release date \(Figure[7](https://arxiv.org/html/2608.29420#A8.F7)b\), so release date confounds any claim that two benchmarks measure the same thing; the saturation literature finds saturation rising with benchmark age\[[Akhtar et al\., 2026](https://arxiv.org/html/2608.29420#bib.bib1)\], and we therefore treat date as a first\-class nuisance variable\.

Most benchmarks are bounded accuracies in\[0,1\]\[0,1\], and several are strongly right\-skewed, with CritPt at Fisher skewness 3\.8, HLE at 1\.5, andτ3\\tau^\{3\}\-Banking at 0\.9, their scores piling up near the floor \(Figure[8](https://arxiv.org/html/2608.29420#A8.F8)\)\. Two benchmarks are unbounded, GDPval on an Elo scale and AA\-Omniscience on a penalised score, and both sit closer to symmetric\. Bounded and skewed scales can manufacture or hide structure through the metric alone\[[Schaeffer et al\., 2023](https://arxiv.org/html/2608.29420#bib.bib42)\], which motivates the logit\-transform check in Appendix[L](https://arxiv.org/html/2608.29420#A12); that check confirms that the first\-factor share is no artefact of the raw scale\. The low\-scoring tails hold weak models with genuinely low scores, so we retain every row and handle the boundedness through the transform, trimming no outliers\.

The dimensionality evidence behind H1, the scree plot and the parallel analysis, is Figure[5](https://arxiv.org/html/2608.29420#A4.F5)in Appendix[D](https://arxiv.org/html/2608.29420#A4); the first principal component accounts for 79\.4% of total score variance, and only the first observed eigenvalue exceeds its random\-data 95th percentile\.

## Appendix IHyperparameter grid and selections

*Supports Section[4\.2](https://arxiv.org/html/2608.29420#S4.SS2)and the robustness claim of Section[5\.5](https://arxiv.org/html/2608.29420#S5.SS5)*that the H4 verdict does not hinge on a single tuning choice\. The inner grid searched ridgeα∈\{0\.03,0\.1,0\.3,1,3,10,30\}\\alpha\\in\\\{0\.03,0\.1,0\.3,1,3,10,30\\\}, elastic\-netα∈\{0\.03,0\.1,0\.3,1\}\\alpha\\in\\\{0\.03,0\.1,0\.3,1\\\}withℓ1\\ell\_\{1\}ratio in\{0\.2,0\.5,0\.8\}\\\{0\.2,0\.5,0\.8\\\}, random\-forest depth in\{4,none\}\\\{4,\\text\{none\}\\\}with 300 trees, and gradient\-boosting depth in\{2,3\}\\\{2,3\\\}with learning rate in\{0\.05,0\.1\}\\\{0\.05,0\.1\\\}\. The ridge and elastic\-net grids span three orders of magnitude of regularisation strength on a log scale, the standard practice for a regularisation path\[[Hastie et al\., 2009](https://arxiv.org/html/2608.29420#bib.bib16)\], and the tree grids bracket the shallow trees appropriate forn=96n=96andk\+4k\+4predictors\. Table[5](https://arxiv.org/html/2608.29420#A9.T5)reports the modal selection per rung; the modal ridgeα=0\.03\\alpha=0\.03sits at the weakly regularised end of the grid, and the H4 verdict is unchanged across the sweep ofkkin Appendix[L](https://arxiv.org/html/2608.29420#A12)\. The regression learners and cross\-validation use scikit\-learn\[[Pedregosa et al\., 2011](https://arxiv.org/html/2608.29420#bib.bib33)\]; the factor analysis uses thefactor\_analyzerpackage, and SHAP its reference implementation\[[Lundberg and Lee, 2017](https://arxiv.org/html/2608.29420#bib.bib28)\]\.

Table 5:Modal hyperparameters selected by the inner five\-fold grid search, reported as the plurality choice across all twelve leave\-one\-benchmark\-out targets with the agreement count\. The H4 verdict rests on fold\-averaged out\-of\-fold errors, not on any single setting\.
## Appendix JDeviations from the pre\-specified design

*Supports the reproducibility claims throughout\.*A study that fixes its analysis plan in advance owes the reader an explicit account of where the delivery departed from it\. Five departures matter\. First, the plan’s primary unit was one row per base model at its best default configuration\. The analysis instead uses all 421 configurations; the deduplicated base\-model dataset appears in Appendix[L](https://arxiv.org/html/2608.29420#A12)as planned check R1, and the configuration level serves as the primary\. Second, the plan’s factor\-count rule is parallel analysis, and it retains one factor\. The three\-factor solution that exhibits economic separation is a pre\-declared exploratory over\-extraction, and we report H3 as not supported under the plan in consequence\. Third, the taxonomy places Terminal\-Bench Hard in scientific coding, while the plan grouped it with Terminal\-Bench v2\.1 as economic\. The date\-adjusted loadings show Hard at 0\.82 on the economic factor, so the assignment is genuinely contestable, and we disclose the split openly\. Fourth, the delivered robustness checks depart from the planned suite\. We label them distinctly below, and we do not renumber a different battery onto the planned names\. Fifth, the plan set the two inclusion floors as indicative values to be confirmed at the coverage audit rather than as fixed constants; the audit confirmed both unchanged, and no score was used to choose them\. The plan also targeted 180 to 220 base models, where coverage on the economic block delivered 96 complete cases and 89 deduplicated base models\.

Two further differences are presentational and are recorded for readers comparing the paper with the deposited plan\. The plan organised the work under two questions, structure and prediction; the paper splits the structural question so that each hypothesis has its own research question, with no change to any hypothesis or test\. And the H2\(ii\) fifteen\-point threshold was calibrated to a reading of[Kearns \[2026\]](https://arxiv.org/html/2608.29420#bib.bib25)in which a scale correction lowered a dominant factor’s share from 72% to 41%\. That reading was mistaken, and Section[2](https://arxiv.org/html/2608.29420#S2)gives the corrected account; the threshold itself was fixed in advance and has not been moved\.

## Appendix KClustering of models and benchmarks

*Supports Section[4\.1](https://arxiv.org/html/2608.29420#S4.SS1)\.*A factor model can impose structure the raw data do not support, so we sought an independent check, and clustering corroborates the low dimensionality without assuming a factor model at all\. We cluster models in factor\-score space withkk\-means, Gaussian mixtures, and agglomerative Ward linkage, selecting the cluster count by silhouette\[[Rousseeuw, 1987](https://arxiv.org/html/2608.29420#bib.bib40)\]and the gap statistic\[[Tibshirani et al\., 2001](https://arxiv.org/html/2608.29420#bib.bib48)\]; three algorithms guard against a clustering that reflects one algorithm’s inductive bias\. All three preferk=2k=2by silhouette, at 0\.480 forkk\-means, 0\.499 for agglomerative, and 0\.467 for the Gaussian mixture \(Figure[9](https://arxiv.org/html/2608.29420#A11.F9)b\)\. The gap statistic increases monotonically inkkand finds no interior optimum, so it is inconclusive on data with a strong general factor and we set it aside\. The two clusters mark capability tiers along one axis of overall strength\. A larger, earlier, lower\-scoring group of 302 models, with median release 2025\-09 and mean Intelligence Index 13\.7, separates from a smaller, newer, higher\-scoring group of 107 models, with median release 2026\-03 and mean index 37\.1\. The split sorts models by how much capability they hold, along the same date\-driven axis \(Figure[9](https://arxiv.org/html/2608.29420#A11.F9)a\)\.

Figure 9:Model clustering on the dense grid \(n=409n=409\)\.\(a\) Models in the space of the first two factors, coloured by the two\-cluster agglomerative solution\. The factor axes are the dense\-grid solution fitted largely without the economic columns, so they track overall capability, and the two clusters mark tiers of it\. \(b\) Silhouette scores bykk; all three algorithms agree onk=2k=2\.Clustering the benchmarks separately by correlation distance1−ρ1\-\\rhounder average linkage places ten of twelve, and three of four economic benchmarks, in one block, withτ2\\tau^\{2\}\-Bench and IFBench branching separately \(Figure[10](https://arxiv.org/html/2608.29420#A11.F10)\)\.

Figure 10:Benchmark clustering by correlation distance\(1−ρ1\-\\rho, average linkage\)\. Ten of the twelve benchmarks, including three of four economic benchmarks, fall in one block, withτ2\\tau^\{2\}\-Bench and IFBench branching separately\.
## Appendix LRobustness checks

*Supports Section[5\.5](https://arxiv.org/html/2608.29420#S5.SS5)\.*We first report the one planned check that bears directly on the deduplication concern, then the additional checks actually delivered\.

#### Planned check R1 \(configuration versus deduplicated\)\.

We collapse the 421 configurations to one row per base model, keeping the best default configuration as the plan specified\. We operationalise it as the highest\-intelligence\-index row for each base model, after stripping reasoning, effort, and preview suffixes from the model name\. Of the resulting base models, 89 carry the complete battery\. On this deduplicated grid the first\-factor share rises to 90\.2% and the date\-adjustment drop rises to 24\.1 points, and the four economic loadings remain above 0\.40\. Deduplication therefore strengthens H1 and H2\(ii\) and leaves the exploratory economic pattern intact, and the configuration\-level analysis is the conservative choice for the structural headline\. We also re\-run the H4 predictive test on this deduplicated grid, where the resampling unit is a distinct base model\. The pooled economic\-block improvement of thekk\-factor representation over the single index is\+0\.038\+0\.038with a 95% interval of\[\+0\.020,\+0\.056\]\[\+0\.020,\+0\.056\]\. It matches the configuration\-level\+0\.037\+0\.037and again excludes zero, so the predictive gain does not depend on treating near\-duplicate configurations as exchangeable\.

#### Delivered checks C1–C7\.

\(C1\) A rank\-based factor analysis raises the first\-factor share to 94\.5%, and \(C2\) a logit transform of the bounded accuracy benchmarks raises it to 91\.8%, so H1 is not a Pearson artefact\. \(C3\) On the 58\-model compute\-known subsample, the date\-only adjustment drops the first\-factor share by 16\.5 points, and adding log\-compute jointly gives a smaller 9\.3\-point drop, so compute does not deepen the temporal correction beyond date on the same models\. \(C4\) The three clustering algorithms agree onk=2k=2\. \(C5\) The predictor\-ladder ordering holds on both the economic block and all twelve targets\. \(C6\) The dominant\-factor date fit is robust to the link function, since a logistic fit givesR2=0\.505R^\{2\}=0\.505against an ordinary\-least\-squaresR2=0\.477R^\{2\}=0\.477, and both clear 0\.30 with the first factor ranked strongest, so H2\(i\) does not depend on that modelling choice\. \(C7\) Thekk\-factor predictive advantage does not depend on the exploratory factor count\. Sweeping the rung\-\(iv\) predictor overk∈\{2,3,4,5\}k\\in\\\{2,3,4,5\\\}, the improvement over the single index is positive at every count with a bootstrap interval excluding zero, reading\+0\.021\+0\.021atk=2k=2,\+0\.037\+0\.037at thek=3k=3we use,\+0\.058\+0\.058at the criterion’sk=4k=4, and\+0\.059\+0\.059atk=5k=5, with pooled testR2R^\{2\}rising from 0\.79 to 0\.83 across the range\. The gain increases withkk, so thek=3k=3we adopt is conservative for this test relative to the criterion’sk=4k=4\.

The planned checks not run are R3 \(FIML and MICE missingness sensitivity\), R4 \(replication of the single\-factor foil then correction\), R5 \(reasoning\-effort variants as a second scale axis\), and R6 \(the taxonomy swap test\)\. Each was deferred for space; R5 and R6 bear on the effort and taxonomy departures above and are the natural next step\.

## Appendix MExplainability and error analysis

*Supports Section[5\.4](https://arxiv.org/html/2608.29420#S5.SS4)\.*To interpret the predictive gain, we open the models fitted to the largest economic target\. GDPval is predicted best by akk\-factor ridge with out\-of\-foldR2R^\{2\}0\.88 \(Figure[11](https://arxiv.org/html/2608.29420#A13.F11)\)\. SHAP attribution on the gradient\-boosting learner fitted at the same rung ranks the agentic factor F1 first, followed by the reasoning flag and the hard\-reasoning factor F3, with scale and open\-weights status contributing little \(Figure[3](https://arxiv.org/html/2608.29420#S5.F3)c\)\. The ordering is informative in the light of prior work, since[Ruan et al\. \[2024\]](https://arxiv.org/html/2608.29420#bib.bib41)find agentic performance forecastable from a low\-dimensional capability space, and[Ilić and Gignac \[2024\]](https://arxiv.org/html/2608.29420#bib.bib20)find parameter count correlated with the general factor at about 0\.6; here, once the factors are in the model, parameter count and open\-weights status add almost nothing, so the scale signal that prior work reports reaches GDPval through the factors and not alongside them\. The predicted\-versus\-observed plot lies on the identity line, and the residuals show no systematic trend \(Figure[11](https://arxiv.org/html/2608.29420#A13.F11)a,b\)\. The largest residuals are interpretable outliers\. They include a non\-reasoning variant of a frontier model that outperforms its economic prediction, Grok 4\.3 Non\-reasoning at\+1\.14\+1\.14, and two models that under\-perform theirs \(Figure[11](https://arxiv.org/html/2608.29420#A13.F11)c\)\. The pattern fits GDPval rewarding agentic behaviour that the other benchmarks capture only indirectly\.

Figure 11:Regression diagnostics for the GDPvalkk\-factor ridge model, out\-of\-fold\.\(a\) Predicted versus observed, on the identity line withR2=0\.88R^\{2\}=0\.88\. \(b\) Residuals versus predicted show no systematic trend\. \(c\) The eight largest residuals, named\. These are interpretable outliers, and none signals misspecification\.
## Appendix NData provenance and reproducibility

*Supports Section[3](https://arxiv.org/html/2608.29420#S3)and the reproducibility claims throughout\.*The primary data set is the Artificial Analysis model leaderboard\[[Artificial Analysis, 2026](https://arxiv.org/html/2608.29420#bib.bib2)\], captured 2026\-07\-06 as a raw HTML snapshot and pinned by the SHA\-256 hash recorded in the shipped manifest \(aa\_snapshot\_manifest\.json\),6f19f8f08befaaf14cb2d9952e379404c39f246b52d7ae4d1c2aa44795404de6\. The Data API refused an unauthenticated request, so we captured the public page and stored the exact bytes\. Every downstream number is therefore reproducible regardless of live\-page drift\[[Biderman et al\., 2024](https://arxiv.org/html/2608.29420#bib.bib6)\]\. The secondary data set is Epoch AI Notable AI Models\[[Epoch AI, 2026](https://arxiv.org/html/2608.29420#bib.bib14)\], released under CC\-BY 4\.0 and used only for the scale robustness check\. The analysis plan was fixed before the data were captured and before any analysis was run, and is deposited at[https://doi\.org/10\.17605/OSF\.IO/VD34J](https://doi.org/10.17605/OSF.IO/VD34J); the deposit is retrospective, and the note there records the timeline, the plan’s checksum, and the limits of that evidence\. Both snapshots, every result table, and the analysis notebook are released at[https://github\.com/louisyzhu/frontier\-ai\-economic\-validity](https://github.com/louisyzhu/frontier-ai-economic-validity)\. The pipeline is a single notebook that runs top to bottom from the pinned snapshots with no manual step, no live network call and no path outside the repository, and all four learners run by default\. It regenerates the figures that carry the results, Figures[2](https://arxiv.org/html/2608.29420#S5.F2),[3](https://arxiv.org/html/2608.29420#S5.F3),[7](https://arxiv.org/html/2608.29420#A8.F7),[9](https://arxiv.org/html/2608.29420#A11.F9)and[11](https://arxiv.org/html/2608.29420#A13.F11), while the remaining figures and twelve result tables listed in the repository are shipped rather than recomputed within it; the notebook prints every value it reads, so each remains checkable against its own output\. The SHAP attribution behind Figure[3](https://arxiv.org/html/2608.29420#S5.F3)c and the R1 deduplication with its H4 re\-run verify themselves against the tables cited here and reproduce exactly\.

The prediction ladder is refitted on every run and written alongside the canonical table rather than over it, so no value reported here can be altered by re\-execution\. The notebook then recomputes the quantities this paper prints, the per\-target and pooled economicΔ​MSE\\Delta\\mathrm\{MSE\}and the five economic\-block rows of Table[1](https://arxiv.org/html/2608.29420#S5.T1), and compares them with the canonical values at a tolerance of5×10−45\\times 10^\{\-4\}, half a unit at the third decimal and so at the precision reported here; it separately verifies that the best\-learner ranking behind the table’s learner column is unchanged\. Those quantities reproduce in every printed digit on both machines we have tested\. Individual ladder cells are reported rather than enforced, because one rung is not platform\-invariant\. The maximum\-likelihood solution has a weakly determined leading direction for this correlation structure, so the single\-factor rung can settle on a slightly different point under a different linear\-algebra backend, and gradient\-boosted trees carry that difference most visibly because a split threshold either flips or does not\. The largest deviation we observed was7\.2×10−47\.2\\times 10^\{\-4\}, reaching thekk\-factor rung only through tree learners on a non\-economic target, while the mean\-index andkk\-factor ridge fits from which every reported effect and interval is built agreed to about10−710^\{\-7\}\. The divergence is a property of the platform and of that rung, not of the results\.

Similar Articles

Two AI Metrics Diverged: Will it Make All the Difference?

arXiv cs.AI

This paper analyzes how different AI performance metrics (bounded vs unbounded) determine whether frontier AI capabilities remain concentrated among wealthy actors or diffuse to smaller models, with implications for regulation.

The Capability Frontier: Benchmarks Miss 82% of Model Performance

arXiv cs.AI

The paper introduces the Capability Frontier, a Pareto frontier over models that corrects for biases in single-model and single-run evaluations, showing that standard benchmarks miss up to 82% of model performance and that collective LLM capabilities are substantially underestimated.

Open-World Evaluations for Measuring Frontier AI Capabilities

arXiv cs.AI

This paper argues that traditional benchmarks both overestimate and underestimate frontier AI capabilities, and proposes 'open-world evaluations'—long-horizon, real-world tasks assessed qualitatively—as a complementary approach. The CRUX project is introduced, with a demonstration where an AI agent successfully published an iOS app to the App Store with minimal intervention.