When benchmark inferences do not compose: Projectibility in AI evaluation

arXiv cs.AI Papers

Summary

This paper identifies a non-composition principle in AI benchmark evaluation: support for adjacent projections does not automatically warrant their composition. It proposes a projectibility audit to diagnose unsupported joins in benchmark-to-use arguments, with a legal-research case study and simulations.

arXiv:2607.26159v1 Announce Type: new Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
Original Article
View Cached Full Text

Cached at: 07/31/26, 04:00 AM

# When Benchmark Inferences Do Not Compose: Projectibility in AI Evaluation
Source: [https://arxiv.org/html/2607.26159](https://arxiv.org/html/2607.26159)
\[ Numbers=OldStyle, Ligatures=TeX, BoldFont=EBGaramond\-Regular\.ttf, ItalicFont=EBGaramond\-Italic\.ttf, BoldItalicFont=EBGaramond\-Italic\.ttf, \] \[ BoldFont=CharisSIL\-Bold\.ttf, ItalicFont=CharisSIL\-Italic\.ttf, BoldItalicFont=CharisSIL\-BoldItalic\.ttf, \] \[ BoldFont=HaranoAjiMincho\-Bold\.otf, \] \[ BoldFont=Inconsolata\-Regular\.ttf, Scale=MatchLowercase, \]

###### Abstract

An AI benchmark result rarely reaches a consequential claim in one step\. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences\. Validity\-centred approaches require evidence for each claim\. This paper identifies a further epistemic problem: warranted links don’t automatically make a warranted chain\. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent\.

Projectibilityconcerns whether a bounded extension from observed to unobserved cases is warranted\. Goodman supplies the problem of rival extensions; argument\-based validity supplies an architecture for testing them\. The paper’s distinctive claim is a non\-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through\. A legal\-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel\. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires\. The resultingprojectibility auditdiagnoses unsupported joins in benchmark\-to\-use arguments\.

Keywords:AI evaluation; benchmark validity; projectibility; construct validity; generalization; AGI

AI use\.The large language models ChatGPT 5; Claude Sonnet 4\.5 and Opus 4\.8; Gemini 3\.1 Pro served as drafting and editing aids throughout the preparation of this paper\. I am responsible for all theoretical claims, arguments, errors, and interpretive choices\.

## 1The Missing Problem: Composition

A benchmark result rarely supports a consequential claim by itself\. Consider a common chain of reasoning\. A base model scores highly on a multidomain battery\. The score is interpreted as evidence of general reasoning\. A product built from that model is expected to perform a professional task\. Human review is expected to improve the final work\. That expectation, projected savings, and an error tolerance support authorization\.

The chain spans four empirical objects: benchmark responses, outputs from a tool\-using application, work after human review, and consequences under a policy\. A capability attribution concerns the base model and usually licenses the move from benchmark to application performance\. The score records only the benchmark responses, not the other three objects\.

Psychological and educational measurement provide a disciplined response to this problem\. On an argument\-based view, validity belongs to a proposed interpretation and use of scores, not to a test considered apart from any claim\[[undefaaq](https://arxiv.org/html/2607.26159#biba.bibx31),[undefaai](https://arxiv.org/html/2607.26159#biba.bibx23)\]\.

Recent work brings that view into AI evaluation\. Claim\-centred frameworks distinguish measurements, evaluations, and the assertions or decisions drawn from them; construct\-validity studies ask whether benchmark content, structure, and relations support the capabilities named; and benchmark\-epistemology accounts identify the assumptions needed to move from performance on a learning problem to a scientific conclusion\[[undefaav](https://arxiv.org/html/2607.26159#biba.bibx36),[undefap](https://arxiv.org/html/2607.26159#biba.bibx4),[undefaab](https://arxiv.org/html/2607.26159#biba.bibx16)\]\. These approaches substantially improve on treating a benchmark as valid or invalid without specifying what it’s supposed to warrant\.

They also make a further problem visible\. Evaluation arguments are usually chains, but support for their links doesn’t automatically compose\. A benchmark may support a claim about new items generated under its rules\. A firm may separately show that a particular human–AI procedure works on a sample of its own cases\. Both results can be warranted without the first supplying a premise for the second\. The local study may bypass the benchmark rather than extend it\.

Conversely, two studies may appear to meet at a shared noun such aslegal reasoning,draft quality, orhuman oversightwhile operationalizing different objects, populations, and outcomes\. In either situation, adding the studies together doesn’t produce the broad claim\. These aren’t only possibilities\. Legal\-research vendors advertised retrieval\-grounded tools as hallucination\-free on the strength of how those tools were built, and a preregistered evaluation later measured hallucination on more than 17% of queries\[[undefaan](https://arxiv.org/html/2607.26159#biba.bibx28)\]\.

The problem isn’t that machine outputs are fallible while human judgment isn’t\. Human judgments and institutions also fail\. AI makes the interface problem especially pressing because one evaluated component can be replicated at high volume, modified into many applications, and inserted into workflows whose relevant evidence is distributed among different actors\. An unsupported bridge can become a repeated, correlated failure while no single study observes the complete chain\.

LLMs can also make the remedy easier to apply by helping evaluators prepare typed source and target descriptions, trace assumptions and evidence, and flag possible mismatches for inspection\. Those outputs don’t certify the warrant they describe; they remain fallible contributions to a process answerable to independent evidence and accountable judgment\.

The central claim of this paper is a non\-composition principle:

Warr⁡\(C1\)∧Warr⁡\(C2\)⇏Warr⁡\(C2∘C1\)\.\\operatorname\{Warr\}\(C\_\{1\}\)\\ \\land\\ \\operatorname\{Warr\}\(C\_\{2\}\)\\ \\not\\Rightarrow\\ \\operatorname\{Warr\}\(C\_\{2\}\\\!\\circ C\_\{1\}\)\.\(1\)HereC1C\_\{1\}andC2C\_\{2\}are links presented as adjacent\. IfC1C\_\{1\}’s target isn’tC2C\_\{2\}’s source, composition is undefined without a bridge; Equation[1](https://arxiv.org/html/2607.26159#S1.E1)covers defined composition whose warrant fails to transmit\. Equation[1](https://arxiv.org/html/2607.26159#S1.E1)states a non\-composition principle, not a theory of warrant\.

WritingWarr\\operatorname\{Warr\}as a predicate is a notational convenience:Warr⁡\(C\)\\operatorname\{Warr\}\(C\)means warrant sufficient for the declared use\. This threshold notation doesn’t imply that support for a fixed claim is binary or reducible to one scale; warrant for the component links doesn’t extend to their composition\. Warranted composition requires alignment, or a separately warranted bridge, in object, population, conditions, outcome, and period; compatible assumptions and effect modifiers; and propagation of dependence and uncertainty\.

Projectibilityconcerns whether one bounded extension from observed to unobserved cases is warranted\. The term comes from Goodman’s treatment of induction, but the account developed here doesn’t offer a general solution to induction and doesn’t make entrenchment a sufficient condition\. It gives a place within a validity argument for a narrower question: exactly what is being projected, across which boundary, under which assumptions, and with what evidence against the differences that could defeat the extension? Aprojectibility auditrecords the answer link by link and diagnoses unsupported joins rather than treating separately warranted links as automatically composable\.

This focus separates the paper from neighbouring proposals without displacing them\. A nomological network, the set of relations a construct is predicted to bear to other constructs and observables, can make a capability claim more substantive\[[undefax](https://arxiv.org/html/2607.26159#biba.bibx12),[undefaaa](https://arxiv.org/html/2607.26159#biba.bibx15)\]\. An estimand can state precisely which quantity a study targets\[[undefaq](https://arxiv.org/html/2607.26159#biba.bibx5)\]\. A selection diagram can identify conditions for transporting a causal effect\[[undefaar](https://arxiv.org/html/2607.26159#biba.bibx32)\]\. Each may supply evidence or structure for a projection\. None by itself guarantees that the target of one analysis is the source of the next\. Projectibility concerns whether each boundary\-specific extension is warranted; the audit also examines the interfaces among analyses\.

Section[2](https://arxiv.org/html/2607.26159#S2)locates projectibility within validity theory and distinguishes it from construct meaning, prediction, and decision\. Section[3](https://arxiv.org/html/2607.26159#S3)states the interface requirements and separates composition from convergence and replacement\. Section[4](https://arxiv.org/html/2607.26159#S4)works through a legal\-research case whose hypothetical results support one projection, defeat another, and leave a third unresolved\. Section[5](https://arxiv.org/html/2607.26159#S5)shows how an upstream mean can erase the item\-level differences a downstream projection needs\. Sections[6](https://arxiv.org/html/2607.26159#S6)and[7](https://arxiv.org/html/2607.26159#S7)apply the result to general\-capability claims and divide the evidential work between developers and deployers\.

## 2Projectibility within a Validity Argument

### 2\.1 Rival extensions from the same observations

Goodman’s new riddle of induction begins with observations that fit incompatible extensions equally well\. Every emerald examined before timettis green\. The observations fit the familiar hypothesis that emeralds are green\. They also fit the rival hypothesis that emeralds aregrue\. Under this hypothesis, the observed emeralds qualify because they’re green, but an emerald first examined afterttwould qualify only if it were blue\. The hypotheses agree on every observed emerald and disagree on the next unexamined one\. Fit to the source observations doesn’t determine which predicate is fit for projection\[[undefaae](https://arxiv.org/html/2607.26159#biba.bibx19), 74–80\]\.

The benchmark analogue isn’t that evaluators literally invent time\-indexed predicates\. Many classifications of unobserved cases agree on the scored items\. Success can be grouped undergeneral legal reasoning, underanswering supplied\-text legal questions, underrecognizing patterns in questions drawn from familiar sources, or under still narrower descriptions\. Those predicates may be extensionally equivalent on the test set and diverge on a research request that requires current authority, live retrieval, source comparison, and a cited memo\. More observations sampled under the same narrow item\-generating rule may distinguish none of them\.

Goodman appeals to entrenchment, the history of successful use of a predicate in past projections\[[undefaae](https://arxiv.org/html/2607.26159#biba.bibx19), 84–98\]\. Other accounts place more weight on revisable background theory and causal structure\[[undefat](https://arxiv.org/html/2607.26159#biba.bibx8),[undefaak](https://arxiv.org/html/2607.26159#biba.bibx25)\]\. This account needs no single sufficient criterion\. Its more modest lesson is that the rule of extension has to be exposed\. A target declaration identifies the cases over which rival predicates disagree\. Background knowledge identifies differences that could matter\. Direct samples, controlled variations, and later observations test those differences\. This doesn’t remove induction; it makes a particular induction inspectable and revisable\.

### 2\.2 Definition and scope

Aprojectionextends an interpretation, prediction, or explanation from specified source observations to unobserved cases\.Projectibilityconcerns whether that bounded extension is warranted relative to a declared target and use\. Three restrictions follow\.

First, projectibility belongs to a bounded projective claim, not to a score, benchmark, or system considered in isolation\. The same benchmark result may provide substantial warrant for performance on further items generated under a registered template and little warrant for performance on open\-ended professional work\. Calling the benchmarkprojectiblewithout naming the target suppresses the point at issue\.

Second, projectibility isn’t a scalar property\. Support for a fixed projection may be stronger or weaker, but support, scope, and the kinds of evidence bearing on the claim shouldn’t be collapsed onto one scale\. A report should preserve heterogeneous results: a predictive estimate and interval, a defeated invariance assumption, an unresolved population gap, and a supported narrow use\. Collapsing these into a projectibility score recreates the aggregation problem\. The status should remain claim\-specific: supported under a stated evidential standard, defeated by specified evidence, or unresolved because relevant evidence is absent or imprecise\. Each status should name its limit, such as support for one request type or dependence on evidence that only another party can supply\.

Third, source fit is only part of the evidence\. Warrant for a projection commonly comes from four places\. Direct target sampling shows what happens in cases admitted by the target rule\. Deliberate variation tests features suspected of changing the result while preserving what the claim treats as invariant\. Background knowledge explains why some differences are plausible defeaters and others aren’t\. Replication across genuinely new levels \(another template family, system lineage, site, or period\) tests whether an earlier regularity survives the dimension being projected over\. None can be replaced by a bare assertion that source and target are similar\.

This account belongs inside an argument\-based approach to validity\. Kane represents a proposed interpretation and use as a sequence of inferences and assumptions, each requiring support commensurate with the ambition of the claim\[[undefaai](https://arxiv.org/html/2607.26159#biba.bibx23), 1–3, 22–25\]\. Messick’s unified account likewise treats validity as an evidential judgment about score meaning and use, drawing on content, response processes, internal structure, relations with external variables, and consequences\[[undefaaq](https://arxiv.org/html/2607.26159#biba.bibx31), 741–749\]\. Projectibility concerns whether those parts of the argument that extend beyond cases observed under the source design are warranted\.

It doesn’t subsume every validity question\. Whether a rubric mis\-scores the observed responses is a scoring problem before it’s a projection problem; whether a proposed capability has a coherent meaning is an interpretive problem that may condition a projection without being reducible to one\.

### 2\.3 What projectibility adds to neighbouring frameworks

Neighbouring work can be sorted by how far along the inference it reaches\. Some frameworks stop at what a benchmark collects; others follow a single claim from evidence to interpretation and use\. This paper asks what warrants the join between two such claims\.

\[[undefaat](https://arxiv.org/html/2607.26159#biba.bibx34)\]argue that broad benchmark claims routinely outrun the contextual tasks used to construct them, and\[[undefas](https://arxiv.org/html/2607.26159#biba.bibx7)\]set out how construct validity fails in language\-model evaluation\.\[[undefaam](https://arxiv.org/html/2607.26159#biba.bibx27)\]adapt evidence\-centred design from educational assessment, formalizing benchmark construction into modules that require designers to describe, justify, and support each choice\. Their subject is the evidence a benchmark collects about the capabilities it declares; movement beyond the benchmark isn’t their topic\.

Recent validity\-centred work on AI is closer still, so the difference has to be stated precisely\.\[[undefaav](https://arxiv.org/html/2607.26159#biba.bibx36)\]offer a claim\-aware framework that maps measurements and evaluations to claims, differentiates forms of validity, and recognizes that measurement producers and downstream claimants may be different stakeholders\.\[[undefaab](https://arxiv.org/html/2607.26159#biba.bibx16)\]treat predictive benchmarks as measurement tools and specify internal, external, content, consequential, and auxiliary conditions for scientific inferences\.\[[undefap](https://arxiv.org/html/2607.26159#biba.bibx4)\]document construct\-validity weaknesses across 445 language\-model benchmarks and give practical recommendations for benchmark development\. This paper accepts the shared premise: the inferential object is a particular claim, and stronger claims need more evidence\.

Its additional object is theinterface between claims\. A claim\-centred framework can tell us which evidence would support a benchmark interpretation and which evidence would support a deployment criterion\. A projectibility audit asks whether the conclusion of the first is actually a premise of the second\. If the benchmark study concerns a base model answering supplied\-text questions and the deployment study concerns final memos produced by a retrieval system and reviewed by lawyers, the studies don’t meet merely because both are described as legal evaluation\. Their evidence may converge on a narrow conclusion, one may replace the need for the other, or they may concern different objects altogether\. Composition is a further relation\.

\[[undefaaa](https://arxiv.org/html/2607.26159#biba.bibx15)\]argues that the inferential account provides insufficient resources for articulating the meaning of theoretical capability constructs, and that capability benchmarks should instead be embedded in nomological networks relating tasks, constructs, processes, and external criteria\. That objection concerns a use of argument\-based validity not made here\. This paper neither proposes a meaning forreasoningnor infers the possession of reasoning from benchmark scores\. It uses Kane’s framework for the task it directly addresses: representing a proposed interpretation or use as a sequence of inferences and assumptions whose evidential support can be assessed\. A substantive construct attribution, where one is proposed, needs whatever nomological, process, or causal support that attribution requires; the projectibility audit doesn’t supply it\.

Nor would successful nomological validation remove the composition problem\. A nomological network might support an inference from task scores to a reasoning construct, or establish a relation between that construct and an external criterion in a declared population\. It wouldn’t establish that the relation survives replacement of the benchmarked model by a retrieval\-augmented application, that lawyer review catches the application’s errors, or that reviewed outputs produce the consequences a deployment policy assumes\. Those are further projections over different objects, observations, and conditions\. The non\-composition result holds whether the construct attribution is warranted, unwarranted, or omitted\. It can be omitted because it isn’t an empirical endpoint\. The step it appears to license still needs criterion evidence the attribution doesn’t by itself supply\.

An estimand framework answers another indispensable question: what quantity is the analysis trying to estimate, over which population, under which data\-acquisition and aggregation rule?\[[undefaq](https://arxiv.org/html/2607.26159#biba.bibx5)\]show how benchmark conclusions become confused when those elements are left implicit\. A chain may require several estimands\. One study may estimate an expected score over benchmark items, another a post\-review defect probability over local requests, and a third the rate at which defects reach clients\. Each can be well specified while the chain remains unsupported because the first quantity doesn’t predict the second or the second wasn’t sampled under the conditions of the third\.

Finally, causal transportability supplies stronger formal results for a narrower kind of projection\. Selection diagrams encode differences between source and target domains and establish when a causal relation is identifiable in the target from source experiments and target observations\[[undefaar](https://arxiv.org/html/2607.26159#biba.bibx32)\]\. Projectibility is broader: it covers descriptive generalization, construct interpretation, prediction, explanation, and sociotechnical outcomes, and it claims no identification theorem\. Where the target claim is causal, a projectibility audit should defer to the corresponding causal requirements rather than treating generic robustness evidence as enough\.

Table[1](https://arxiv.org/html/2607.26159#S2.T1)summarizes the division of labour\. Its final column identifies the question that motivates this paper\.

Nothing in the non\-composition principle conflicts with Kane’s logic, and the objection that it adds nothing deserves a direct answer\. A completely specified interpretation\-and\-use argument would include the interface assumptions among its inferences, and support for the whole argument would require support for those assumptions\.

Kane’s own summary of chain reasoning shows where the remaining gap lies:“A chain of reasoning is only as strong as its weakest link, and strong evidence for part of an argument does not compensate for weaknesses in other parts of the argument”\[[undefaai](https://arxiv.org/html/2607.26159#biba.bibx23), 64\]\. That principle presupposes that the inferences form a chain\. Whether they do is a prior question, and where the answer is no, the individually strong links don’t warrant the endpoint by composition\.

Composition becomes a distinct problem because of how AI evidence is actually produced\. It arrives from separate studies, run by different actors, assembled later as though favourable conclusions were automatically adjacent\. Kane’s framework makes room for assumptions at the interfaces among inferences, but it doesn’t treat composition as a distinct object of validation\.

This paper isolates a recurrent failure mode in distributed AI evaluation: separate studies may warrant their own conclusions while differing in the object, population, outcome, conditions, or uncertainty those conclusions need in order to serve as adjacent premises\. The projectibility audit makes the interfaces explicit and states necessary conditions that a proposed composition has to satisfy\.

Table 1:Neighbouring frameworks and the composition question\.

## 3Why Warranted Links Don’t Automatically Compose

### 3\.1 Typed nodes and projection edges

The easiest way to conceal an inferential gap is to give both sides the same broad noun\. A benchmark containslegal tasks, and a firm performslegal tasks; the shared label is taken to show that the score transfers\. A developer teststhe model, and a deployer usesthe model; the shared label is taken to show continuity between the tested and deployed objects\.

I usetypehere in the type\-theoretic sense, specifically as arecord type: roughly, a schema whose instances assign values to labelled fields\[[undefaas](https://arxiv.org/html/2607.26159#biba.bibx33), sec\. 11\.8\]\. For this audit, an empirical node has five fields: the object evaluated, the population of cases, the conditions under which the evaluation occurs, the outcome recorded, and the period to which the claim applies\. These types are revisable representations for an audit, not claims that the objects share an essence or form a natural kind\.

A particular benchmark or deployment is represented as a node by filling those fields with values\. Here,objectnames a field, whereas a named model build is a possible value of that field\. Likewise,legal tasksmay be intended as a value for the population field, but it supplies no case\-generating or inclusion rule by which the benchmark and deployment populations could be shown to match\.

Write an empirical node as

𝒩=⟨object,population,conditions,outcome,period⟩\.\\mathcal\{N\}=\\langle\\text\{object\},\\ \\text\{population\},\\ \\text\{conditions\},\\ \\text\{outcome\},\\ \\text\{period\}\\rangle\.\(2\)The population is a case\-generating or inclusion rule, not just a list of cases, and the period covers the relevant system version as well as the span of time\.

In a benchmark node, the object might be a named model build; the population, held\-out items produced under specified templates; the conditions, the prompt, decoding procedure, tool access, and scorer; the outcome, item correctness; and the period, the evaluation date\. In a deployment node, the object might instead be a model, retrieval index, prompt stack, and lawyer\-review procedure; the population, a rule admitting two kinds of logged research request; the conditions, the firm’s database and review instructions; the outcome, a substantive defect remaining in the final memo; and the period, the next six months\.

Typing doesn’t guarantee that two nodes match\. It makes the proposed match inspectable\. The benchmark and deployment above might both be described as involvinglegal tasks, but their corresponding fields differ throughout\. Repetition of the broad label supplies no warrant for the projection\.

A capability attribution isn’t an empirical node of the form in Equation[2](https://arxiv.org/html/2607.26159#S3.E2)\. It’s an interpretive claim about a bearer, supported where appropriate by content, process, causal, and nomological evidence\.\[[undefax](https://arxiv.org/html/2607.26159#biba.bibx12), 290\]state the constraint on it exactly: a construct“is not‘reduced’to the observations, but only combined with other constructs in the net to make predictions about observables”\. It may be the terminal conclusion of one argument or a substantive premise in another\.

But it doesn’t by itself specify the downstream population of cases, the criterion outcome, or the conditions under which the attributed capability is expected to predict that outcome\. Where a capability attribution is invoked to license a step between empirical nodes, the audit asks which observable relation is required and whether that relation has been tested in the declared population\.

An empirical projection edge includes its source and target nodes, projective claimCeC\_\{e\}, assumptionsAeA\_\{e\}, and evidenceEeE\_\{e\}bearing on the claim and assumptions\. Writee=⟨𝒩s,𝒩t,Ce,Ae,Ee⟩e=\\langle\\mathcal\{N\}\_\{s\},\\mathcal\{N\}\_\{t\},C\_\{e\},A\_\{e\},E\_\{e\}\\rangle, alsoe:𝒩s↝𝒩te:\\mathcal\{N\}\_\{s\}\\rightsquigarrow\\mathcal\{N\}\_\{t\}\.

Generalizing from observed benchmark items to further items under the same generating rule is one empirical edge\. Extrapolating to a new task population, transporting a relation to a new model build, and predicting a post\-review outcome are further edges\. Interpretive inferences remain part of the broader validity argument, but they aren’t themselves sample\-bearing interfaces between empirical studies\. The notation forces no commitment that every argument has the same sequence\. It prevents an argument from treating a sequence as one undifferentiated leap\.

Figure[1](https://arxiv.org/html/2607.26159#S3.F1)shows a common evaluation chain\. The projection appears above each arrow and evidence that can answer it below\. No amount of evidence under one arrow automatically fills another\.

A\. Benchmark\-to\-use architecture

Observed benchmarkresponsesFurther test\-ruleperformanceApplication outputs oneligible target tasksFinal work afterregistered reviewConsequences underauthorized policyCapabilityattributionbacked by content, process, causal, and nomological evidencegeneralizationtemplate\-family holdouttask and system projectiontarget\-task and configuration samplereview projectionpaired draft and final scoringoutcome projectionexposure and outcome surveillanceinterpretationcriterion relation required

B\. Interface audit

Endpoint alignmente1:𝒩0↝𝒩1e2:𝒩1′↝𝒩2e\_\{1\}:\\mathcal\{N\}\_\{0\}\\rightsquigarrow\\mathcal\{N\}\_\{1\}\\qquad e\_\{2\}:\\mathcal\{N\}^\{\\prime\}\_\{1\}\\rightsquigarrow\\mathcal\{N\}\_\{2\}𝒩1=?𝒩1′\\mathcal\{N\}\_\{1\}\\stackrel\{\{\\scriptstyle?\}\}\{\{=\}\}\\mathcal\{N\}^\{\\prime\}\_\{1\}object⋅\\cdotpopulation⋅\\cdotconditions⋅\\cdotoutcome⋅\\cdotperiod𝒩1≠𝒩1′⇒\\mathcal\{N\}\_\{1\}\\neq\\mathcal\{N\}^\{\\prime\}\_\{1\}\\Rightarrowbridge owed; composition undefinedWarrant transmissionIf the endpoints align,Warr⁡\(e1\)∧Warr⁡\(e2\)⇏Warr⁡\(e2∘e1\)\\operatorname\{Warr\}\(e\_\{1\}\)\\land\\operatorname\{Warr\}\(e\_\{2\}\)\\not\\Rightarrow\\operatorname\{Warr\}\(e\_\{2\}\\\!\\circ e\_\{1\}\)compatible assumptions and effect modifiersdependence and uncertainty carried through

Figure 1:Architecture and interface failures in a benchmark\-to\-use argument\. Panel A matches evidence to projections between empirical endpoints\. A capability attribution can be supported while leaving the target cases and criterion relation needed for application performance untested\. Panel B distinguishes endpoint mismatch, which makes composition undefined without a bridge, from failure of warrant transmission across aligned endpoints\.
### 3\.2 Endpoint alignment and warrant transmission

Supposee1:𝒩0↝𝒩1e\_\{1\}:\\mathcal\{N\}\_\{0\}\\rightsquigarrow\\mathcal\{N\}\_\{1\}ande2:𝒩1′↝𝒩2e\_\{2\}:\\mathcal\{N\}\_\{1\}^\{\\prime\}\\rightsquigarrow\\mathcal\{N\}\_\{2\}are each supported\. Two questions arise in order\. First, do the links meet? This is an endpoint\-alignment question about whether the proposed composition is well formed\. Alignment requires continuity in object, population, conditions, outcome, and period, or a separately warranted bridge across the difference\. Second, if they meet, does warrant transmit across the join? That further requires compatibility of assumptions and effect modifiers, and propagation of evidential dependence and uncertainty\. The first failure makes the proposed composition undefined; the second leaves a defined composition unwarranted\.

The five alignment requirements read off Equation[2](https://arxiv.org/html/2607.26159#S3.E2); the two transmission requirements read off the assumptionsAeA\_\{e\}and the evidenceEeE\_\{e\}carried by the edges\. All seven presuppose empirical nodes at both endpoints, which the previous subsection secured\.

Endpoint alignmenthas five requirements\. Where one fails, the mismatched field identifies the bridge that’s owed\. A result about supplied\-text questions doesn’t arrive at a population of live\-retrieval requests\.

1. 1\.Object continuity requires a declared system or workflow projection when the object changes\. A base model, a model connected to a database, a complete application with citation validation, and a memo after lawyer review are different objects\. Product names and version families don’t establish continuity\. Shared training, distillation, retrieval components, or evaluation data may create dependence without preserving behaviour\.
2. 2\.Population alignment requires the cases reached by the upstream projection to have the distribution, support, and grouping assumed downstream\. A study on short, self\-contained questions doesn’t supply cases for a study whose source consists of open\-ended requests with disputed authorities\. Even within one task label, a downstream study may sample routine files while the upstream claim includes novel, urgent, or multilingual work\. The inclusion rule, not the label, determines alignment\.
3. 3\.Condition alignment requires the operating conditions of the shared endpoint to match on both sides, or their difference to be shown harmless for the target claim\. Two studies conducted under incompatible prompts, decoding settings, databases, retrieval configurations, or reviewer instructions don’t compose merely because both are described as evaluating the same system\.
4. 4\.Outcome and scale continuity requires the first edge to deliver the quantity the next uses\. Benchmark accuracy, a factor score, the probability of a citation defect in a draft, a defect after review, and a client loss aren’t interchangeable outcomes\. Putting each measure on a zero\-to\-one scale changes its numerical range, not what a difference means\. A change of \.10 in benchmark accuracy isn’t equivalent to a \.10 change in defect probability or client loss\. If the outcome changes, the relation between the measures is itself an inferential link\.
5. 5\.Temporal alignment treats the period as a field like any other\. A result established on one model build, database index, and work period doesn’t reach a later one because the product name is unchanged\. Where the downstream claim covers a span the upstream study didn’t sample, the difference is a further projection, answered by a chronological holdout or a declared retest trigger rather than by an assumption of continuity\.

Warrant transmissionhas two further requirements once the endpoints align\.

1. 1\.Assumptions and effect modifiers have to remain compatible across the join\. Assumptions that support one edge can fail under the conditions of the next\. Review may reduce one error class while increasing delay or inducing overreliance; retrieval may improve source coverage while introducing irrelevant but persuasive cases\. Relevant effect modifiers have to be carried to the interface rather than averaged away\. An effect modifier dropped at the join can leave both component estimates correct and the composed prediction wrong\.
2. 2\.Dependence and uncertainty have to be propagated rather than reset at the next edge\. A chain can look better supported than it is when its links reuse the same items, scorer, model lineage, or development decisions\. Twenty outputs from one build aren’t twenty systems, and thirty\-two comparisons among related models and shared benchmarks aren’t thirty\-two independent replications\. Where estimates are composed, their covariance, selection, and uncertainty should be represented\. Where a qualitative assumption is uncertain, the chain inherits that uncertainty\.

Suppose one study finds that systems with higher benchmark scores produce fewer defective drafts, and another finds that lawyer review removes a certain proportion of draft defects\. It’s tempting to combine the results to predict the quality of final work\. But that combination makes two claims that neither study establishes alone: that both results apply to the same target population, and that review works in the same way across systems once draft quality is held fixed\.

LetBBbe a benchmark result,DDa draft\-level outcome, andZZa final outcome after review\. In a target populationTT,

PT​\(Z∣B\)=∑dPT​\(Z∣D=d,B\)​PT​\(D=d∣B\)\.P\_\{T\}\(Z\\mid B\)=\\sum\_\{d\}P\_\{T\}\(Z\\mid D=d,B\)\\,P\_\{T\}\(D=d\\mid B\)\.\(3\)In words, consider each possible draft outcome\. Multiply its probability at a given benchmark result by the probability of the final outcome after review for that draft and benchmark result, then add across the possible draft outcomes\.

The first study may estimatePT​\(D∣B\)P\_\{T\}\(D\\mid B\), while the second estimatesPT​\(Z∣D\)P\_\{T\}\(Z\\mid D\)\. Equation[3](https://arxiv.org/html/2607.26159#S3.E3), though, requiresPT​\(Z∣D,B\)P\_\{T\}\(Z\\mid D,B\)\. Substituting the second quantity is warranted only if the benchmark result adds no information about review success once the draft outcome is known,Z⟂B∣DZ\\perp B\\mid D\.

That condition could fail if, for example, high\-scoring systems produce more fluent errors that reviewers are likelier to trust\. And if either study concerns a different population, it doesn’t supply the corresponding relation inTT\. Perfectly aligned endpoint descriptions still don’t ensure that warrant passes across the join\. Formal causal claims require more, but the problem arises even for descriptive prediction\.

### 3\.3 Composition, convergence, and replacement

Three relations among studies are easily conflated\. Incomposition, the conclusion of one study supplies a premise or population for the next\. A benchmark score predicts draft defects across systems; draft defects predict post\-review defects under a specified procedure; those links may be composed if their systems, cases, outcomes, and conditions align\.

Inconvergence, distinct evidence bears on the same claim without forming a chain\. A benchmark, expert assessment, and process intervention might each support a reasoning attribution\. Their agreement can strengthen the claim, but no result is passed as an input from one study to another\. Dependence still matters: three tests built from the same source materials don’t provide three independent routes\.

Inreplacement, direct target evidence makes an ambitious upstream projection unnecessary for a narrower purpose\. A firm that samples its own research requests and measures final reviewed defects can assess that workflow without first proving that the base model possesses general legal reasoning\. The local study doesn’t validate the generality of the benchmark\. It starts from new source observations closer to the target\. This is often epistemically preferable: direct target evidence can support a bounded policy while the capability question remains open\.

The distinctions matter because the phrasebenchmark plus local validationcan describe all three\. If the local study compares several systems and tests whether their benchmark profiles predict their defect rates on the same requests, it tests a composable predictive edge\. If it evaluates only one frozen application, the benchmark profile is constant across its requests and can’t explain which request fails\. The local study then replaces, rather than confirms, the benchmark\-to\-use projection\. If both studies are cited merely because they sound favourable, they provide neither composition nor well\-characterized convergence\.

### 3\.4 Two counterexamples

The two failures separated above look alike on the page and come apart under inspection\. The first is spurious adjacency, where the links never meet\. Consider two individually true claims\. First, a model answers 90% of held\-out supplied\-text legal questions correctly under the benchmark protocol\. Second, lawyers following a written review instruction catch 99% of fabricated citations in drafts generated for a set of routine research requests\. Neither result is defective on its own\. Their conjunction still doesn’t warrant that reviewed memos are substantively correct\.

The retrieval system may omit controlling authority without fabricating any citation\. The benchmark never tests retrieval, and the review study counts only fabricated citations\. Suppose the application omits a controlling case on 20% of requests, the omission isn’t visible from the citations that remain, and the lawyers’ procedure doesn’t require an independent update search\. Both premises remain true while one fifth of final memos contain a substantive defect\. The chain fails at two interfaces: the outcome changes from supplied\-text correctness to citation fabrication, and the review evidence doesn’t cover omitted authority\. Multiplying the two percentages obscures rather than repairs the gap\.

The counterexample also shows what positive evidence would look like\. The firm could prepare an authority list independently of the system, score omitted controlling authority as well as citation fabrication, and require reviewers to compare every final memo with that list or conduct a defined update search\. The revised study wouldn’t make the benchmark general\. It would supply observations for the missing workflow link\.

Something close to this has already happened\. Legal\-research vendors advertised retrieval\-grounded products as delivering“100% hallucination\-free linked legal citations”or as avoiding hallucination by relying on trusted content, and\[[undefaan](https://arxiv.org/html/2607.26159#biba.bibx28)\]found that no empirical evidence accompanied those claims\. Their preregistered evaluation reports hallucination on more than 17% of queries for both the LexisNexis and Thomson Reuters tools, and incomplete answers on more than 60% for Thomson Reuters\. The advertised inference runs from a component property, retrieval over an authoritative database, to what the product asserts, with no observations at the join\. Incompleteness is the omission failure isolated above rather than the fabrication failure the marketing addressed\.

The second failure is harder to see, because the links do meet\. Suppose several systems are evaluated on the same held\-out requests, and a benchmark score predicts the frequency of draft defects across them\. Suppose reviewers working from one instruction catch 90% of draft defects on those same requests\. The draft\-defect variable is now the same variable in both studies, over the same systems and the same requests, so every field of the shared endpoint matches and no bridge is owed\. The composed prediction can still be wrong\. If high\-scoring systems fail mostly by omitting an authority that no citation flags, while low\-scoring systems fail mostly by inventing a citation that any reviewer catches, then review effectiveness depends on which system produced the draft\. In the terms of Equation[3](https://arxiv.org/html/2607.26159#S3.E3),PT​\(Z∣D,B\)≠PT​\(Z∣D\)P\_\{T\}\(Z\\mid D,B\)\\neq P\_\{T\}\(Z\\mid D\), and a catch rate averaged over the whole set overstates what review achieves for exactly the systems the benchmark ranks highest\. Both component relations are well estimated, the interface is aligned, and the chain still fails\. Nothing here is repaired by checking that the two studies denote the same things; the defect is an effect modifier dropped at the join\.

## 4A Worked Projectibility Audit

### 4\.1 The source result and the proposed use

Suppose a developer reports a high legal\-reasoning score for a base model\. The battery presents each question with a fixed set of materials and marks the answer against a reference response\. Its held\-out design may support generalization to further questions produced under the same item rules\. The score doesn’t observe whether the model can find current law, distinguish controlling from merely similar authority, produce a linked citation, or work within a lawyer’s review procedure\.

An Ontario employment\-law group is considering a product built from the model\. The application receives a lawyer’s research request, searches the firm’s licensed legal database, and drafts a one\-page memo\. The firm proposes to authorize it only for two kinds of request: whether a termination clause is enforceable and what period of reasonable notice is likely\. An eligible request is written in English, concerns Ontario law, identifies the relevant dates and employment facts, and asks one of those two questions\. Requests concerning another jurisdiction, a different subject, missing facts that prevent research, or a novel constitutional issue are outside the target\.

Each memo must identify the issue, state the governing propositions, cite the current controlling authorities, link each proposition to a passage that supports it, and undergo line\-by\-line lawyer review before it leaves the firm\.

The object under consideration isn’t the benchmarked base model\. It’s a frozen configuration consisting of the model build, legal\-database index, retrieval settings, system prompt, drafting template, citation\-validation component, and review instruction\. The final outcome belongs to the complete workflow, not to the model alone\. That distinction is familiar in sociotechnical analyses of AI: component accuracy doesn’t by itself determine the quality of decisions or work produced after human interaction\[[undefay](https://arxiv.org/html/2607.26159#biba.bibx13),[undefav](https://arxiv.org/html/2607.26159#biba.bibx10),[undefaal](https://arxiv.org/html/2607.26159#biba.bibx26)\]\.

Those measured rates don’t estimate the defect rate of this application\. They do identify a plausible defeater and an outcome the firm should record\. The firm also needs to count unsupported propositions and omitted controlling authority\. A citation can resolve and accurately quote a real case while the memo remains wrong because the decisive case never appeared\.

### 4\.2 The declared chain

Table[2](https://arxiv.org/html/2607.26159#S4.T2)states the proposed links before any local result is known\. The source and target columns are deliberately repetitive: they show where an apparently continuous claim changes object, population, or outcome\.

Table 2:The legal\-assistant argument decomposed into projection links\.Link 2 isn’t established by showing that the battery contains several legal domains\. It requires observations from the firm’s task\. Link 3 isn’t established by assigning a lawyer to each output; review effectiveness is itself an empirical relation\. Link 5 isn’t established by averaging both request types\. Link 7 isn’t a projection from cases alone\. It combines empirical premises with values, feasible alternatives, professional obligations, and institutional authority\.

### 4\.3 A confirmatory local design

The firm develops its prompt, inclusion rule, reference\-answer procedure, and scoring rubric on older closed files\. Every request from one client file remains in the same split\. Confirmation uses a later untouched block containing 200 eligible termination\-clause requests and 200 eligible reasonable\-notice requests, with one request selected from each client file\. The counts are illustrative, but each number in the design has a function: 200 observations per declared stratum allow the firm to assess its chosen defect tolerance separately rather than letting the more frequent or easier request type dominate a pooled estimate\.

Before any output is generated, two lawyers who supervise these file types independently prepare, for each request, a list of controlling authorities and the legal propositions for which each authority is required\. They resolve disagreements before seeing the application’s memo\. The procedure avoids defining the reference standard around whatever the system happened to retrieve\.

For each request, the application produces one draft under the ordinary retrieval setting\. It also produces a second draft after topically similar but legally irrelevant cases are inserted near the top of the retrieved set\. The alteration leaves the correct legal answer unchanged and tests whether retrieval rank rather than legal relevance controls the memo\. The two drafts from one request are assigned to different reviewing lawyers, and no lawyer sees both versions\. Order is counterbalanced\. Blind scorers assess each draft, the assigned lawyers conduct the registered review, and different blind scorers assess the final memos\.

Asubstantive defectis present when at least one citation doesn’t resolve, a cited passage doesn’t support the proposition attached to it, or the memo omits an authority on the independently prepared controlling\-authority list\. The three defect types are also reported separately\. An unresolvable citation remaining in a final memo is a red\-line outcome; it can’t be offset by correct propositions elsewhere\. Review time is recorded in minutes, but it isn’t folded into the defect indicator\. It enters the later policy comparison in its natural unit\.

The firm sets the confirmatory rule before opening the later block\. For each request type and condition, support for mandatory\-review use requires both \(i\) a one\-sided 95% exact upper confidence bound below a post\-review substantive\-defect tolerance of 3% and \(ii\) no unresolvable citation in a final memo\. The claim is defeated if the corresponding lower bound exceeds 3% or if a red\-line citation remains\. Other results are unresolved\. The 3% tolerance isn’t supplied by the benchmark or proposed as a general legal standard\. In the example, the firm would have to justify it by comparing the consequences and costs of the available policies\.

Fixing the rule in advance narrows analyst discretion without eliminating it\. Adjudicating a borderline citation, pooling the two request types or reporting them apart, and handling a request that yields two outputs all remain open, and each can move the estimate\. The audit records those choices and treats unplanned alternatives as exploratory rather than as independent confirmations\[[undefaad](https://arxiv.org/html/2607.26159#biba.bibx18)\]\.

The exact bounds keep the constructed example transparent by treating the file\-level outcomes as exchangeable Bernoulli draws\. An actual study in which the same lawyers review many files should register a clustered or hierarchical uncertainty analysis and apply its support criterion to that estimate\. The inferential point doesn’t depend on the simple bound\.

This design supplies direct evidence for links 2–5\. It doesn’t test link 6 unless the later period or changed component is sampled, and it doesn’t settle link 7\. It also doesn’t test whether the developer’s benchmark score predicts local defects\. With one frozen application, the benchmark profile is constant across the 400 requests\. A request\-level prediction requires variables that vary by request, such as request type, the number and age of retrieved authorities, conflicting appellate treatment, or exposure to the registered alteration\. A benchmark\-to\-local predictive study would instead need several genuinely distinct systems evaluated on the same held\-out requests\.

### 4\.4 Hypothetical results and their inferential status

Table[3](https://arxiv.org/html/2607.26159#S4.T3)gives constructed results\. They aren’t observations about any commercial system\. Their purpose is to show that the same audit can support, defeat, and suspend judgment about different projections without converting those statuses into one overall verdict\.

Table 3:Constructed confirmatory results for the worked example\. Exact bounds are one\-sided 95% bounds for the post\-review defect probability\.The ordinary termination\-clause claim is supported under the registered rule\. The upper bound is below the firm’s tolerance and no red\-line citation survives\. The result warrants a narrow projection from the sampled later files to other requests admitted by the same rule under the same frozen configuration and review procedure\. It doesn’t warrant unreviewed drafting, another office, another request type, or a later build\.

The alteration result is unresolved, not a failure disguised as success and not evidence of robustness\. Its point estimate is low, but the registered upper bound remains above 3%\. More importantly, the unresolved status identifies the exact missing precision\. The firm can collect new alteration cases if the prospective value of that evidence justifies the cost; it can’t pool the altered and ordinary outputs to make the interval narrower for a claim about the altered condition\.

The reasonable\-notice claim is defeated\. Its lower bound exceeds the tolerance, and two unresolvable citations remain after review\. Success on termination\-clause requests can’t compensate for that result\. The failed claim can lead to a prespecified narrower policy \(for example, search suggestions only on reasonable\-notice files\) but not to a post hoc target defined around the requests that happened to pass\.

The future\-build claim remains unresolved because no future build or database index appears in the observations\. A chronological holdout from the current configuration doesn’t become evidence about a materially changed configuration merely because the product name remains the same\. Authorization can be conditioned on freezing the components or on retesting after a declared update trigger\.

Table[4](https://arxiv.org/html/2607.26159#S4.T4)returns these results to the complete chain\. The crucial row is the benchmark\-to\-local\-defects projection, link 2 in Table[2](https://arxiv.org/html/2607.26159#S4.T2)\. The local draft study estimates the frozen application’s performance on local requests, but it doesn’t establish that the developer’s benchmark score predicted that performance\. For the firm’s bounded decision, the direct local evidence replaces the benchmark\-to\-use projection\. The benchmark remains useful as a description of the source evaluation and perhaps as a screening instrument for choosing systems to test\. It doesn’t acquire external reach from the local result\.

Table 4:Warrant status after the constructed local study\.The example gives projectibility a positive role rather than using it as a label for skepticism\. One declared projection survives a test, one is contradicted, and one lacks enough evidence\. The statuses arise from observations tied to the exact boundary, not from the breadth of the nouns used in the claim\. The firm can authorize a specific mandatory\-review workflow for termination\-clause requests while declining or further testing the others\. It needn’t decide whether the base model possesses legal reasoning in general\.

### 4\.5 From empirical warrant to policy

Even the supported termination\-clause result is only one premise in a decision\. The firm has at least four feasible policies: no assistant, assistant\-generated search suggestions, drafting with the registered line\-by\-line review, and drafting with a lighter check\. It can compare them on final defects, unresolvable citations, omitted authorities, lawyer minutes, confidentiality exposure, and whether an error reaches a client\-facing memo or filed document\. A policy that reduces draft time while increasing final defects may be inferior; a policy with slightly more review time may be preferable if it eliminates red\-line outcomes\.

This separation matters because human review isn’t a fixed safety multiplier\. Studies of human oversight show that detection depends on base rates, signal quality, reviewer incentives, workload, and the information supplied with an output\[[undefaal](https://arxiv.org/html/2607.26159#biba.bibx26)\]\. More accurate component systems can also produce worse joint performance when people under\- or over\-rely on them\[[undefav](https://arxiv.org/html/2607.26159#biba.bibx10)\]\. The legal study records the workflow’s final result rather than multiplying a model accuracy by an assumed review rate\.

Values enter before the final authorization as well as at it\. Treating an unresolvable citation as a noncompensatory red line is a choice about what errors may trade off against other gains\. Choosing 3% as a tolerance and deciding whether lawyer minutes justify a narrower use are further choices\. Validity evidence can show which empirical premises are supported; it can’t choose the firm’s loss function or discharge its professional duties\[[undefaaq](https://arxiv.org/html/2607.26159#biba.bibx31), 747–749\]\.

## 5Aggregation Can Destroy Interface Evidence

The non\-composition problem isn’t caused only by changes of system or site\. It can arise because the source report discards distinctions that a later projection needs\. A total score may answer the developer’s descriptive question while making the deployer’s target question unanswerable\. This section isolates that information loss and gives a compact empirical illustration\. The statistical quantities aren’t proposed as a new universal scorecard\. They show why the target claim has to determine what the source report preserves\.

### 5\.1 One stable mean, several incompatible states

Consider the sameNNitems scored under a baseline condition0and an altered conditionpp\. Letsi​c∈\[0,1\]s\_\{ic\}\\in\[0,1\]be the score for itemiiunder conditioncc, and letδi=si​p−si​0\\delta\_\{i\}=s\_\{ip\}\-s\_\{i0\}\. Define the signed level changeLLas the mean item change:

L=1N​∑i=1Nδi\.L=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\delta\_\{i\}\.\(4\)A value near zero says that positive and negative changes balance on average\. It doesn’t say that items were stable\. To distinguish average stability from item movement and concentration, use the item instability \(INS\\mathrm\{INS\}\) and worst\-tail degradation \(WTD\\mathrm\{WTD\}\) summaries introduced by\[[undefaay](https://arxiv.org/html/2607.26159#biba.bibx39)\]\. Item instability is the mean absolute paired change:

INS=1N​∑i=1N\|δi\|,\\mathrm\{INS\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\|\\delta\_\{i\}\|,\(5\)and define its positive and negative directional components,F\+F^\{\+\}andF−F^\{\-\}, as the mean improvement and deterioration:

F\+=1N​∑i\(δi\)\+,F−=1N​∑i\(−δi\)\+\.F^\{\+\}=\\frac\{1\}\{N\}\\sum\_\{i\}\(\\delta\_\{i\}\)\_\{\+\},\\qquad F^\{\-\}=\\frac\{1\}\{N\}\\sum\_\{i\}\(\-\\delta\_\{i\}\)\_\{\+\}\.\(6\)The two pairs are related byL=F\+−F−L=F^\{\+\}\-F^\{\-\}andINS=F\+\+F−\\mathrm\{INS\}=F^\{\+\}\+F^\{\-\}\. Reporting both makes cancellation observable rather than hidden\.

A mean can also hide concentration\. For a prespecified tail fractionqq, order the paired changes from smallest to largest, letm=⌈q​N⌉m=\\lceil qN\\rceil, and define worst\-tail degradation as

WTDq=−1m​∑j=1mδ\(j\)\.\\mathrm\{WTD\}\_\{q\}=\-\\frac\{1\}\{m\}\\sum\_\{j=1\}^\{m\}\\delta\_\{\(j\)\}\.\(7\)A positive value records average deterioration in the worst\-changing fraction\. It still measureschange\. A serious error repeated in both conditions hasδi=0\\delta\_\{i\}=0and disappears fromLL,INS\\mathrm\{INS\}, andWTDq\\mathrm\{WTD\}\_\{q\}\.

The distinction is visible without simulation\. Suppose 100 items have a stable mean \(Figure[2](https://arxiv.org/html/2607.26159#S5.F2)\)\.

- –In state A, every item hasδi=0\\delta\_\{i\}=0\. ThenL=0L=0,INS=0\\mathrm\{INS\}=0, andWTD\.1=0\\mathrm\{WTD\}\_\{\.1\}=0\.
- –In state B, 50 items improve by \.20 and 50 deteriorate by \.20\. ThenL=0L=0, butINS=\.20\\mathrm\{INS\}=\.20andWTD\.1=\.20\\mathrm\{WTD\}\_\{\.1\}=\.20\.
- –In state C, every item is unchanged across conditions, but the same ten items carry an absolute substantive\-loss probability of \.80 in both\. Then all three change quantities are zero while the worst\-decile case\-risk tail is \.80\.

The states license different claims\. State A supports stability on the observed items\. State B defeats item stability despite the stable mean\. State C shows why stability itself isn’t safety: an unchanged serious failure remains serious\. No estimator can recover which state obtained fromLLalone\. The target claim determines which distinctions are required\. A claim about answer\-preserving context variation needs paired changes and their concentration; a deployment claim about final defects needs an absolute loss after the relevant review procedure\.

−\.20\-\.200\+\.20\+\.20State AL=0L=0,INS=0\\mathrm\{INS\}=0,WTD\.1=0\\mathrm\{WTD\}\_\{\.1\}=0absolute risk unspecified−\.20\-\.200\+\.20\+\.20State BL=0L=0,INS=\.20\\mathrm\{INS\}=\.20,WTD\.1=\.20\\mathrm\{WTD\}\_\{\.1\}=\.20absolute risk unspecified−\.20\-\.200\+\.20\+\.20State CL=0L=0,INS=0\\mathrm\{INS\}=0,WTD\.1=0\\mathrm\{WTD\}\_\{\.1\}=0case\-risk tail=\.80=\.80Figure 2:Three item\-level states behind one stable mean, constructed for illustration rather than measured\. Each panel gives the distribution of per\-item changeδi\\delta\_\{i\}; all three haveL=0L=0\. States A and C are indistinguishable in every change statistic, yet the ten worst items in C carry an absolute substantive\-loss probability of \.80 in both conditions\. A change statistic can’t stand in for an absolute risk measure, and a source report that publishes onlyLLleaves a downstream claim about final defects unanswerable\.For an absolute target loss, letℓi​R\(T\)∈\[0,1\]\\ell^\{\(T\)\}\_\{iR\}\\in\[0,1\]be the realized loss for itemiiunder residual response, scoring, and outcome variationRR, and letμi\(T\)=𝔼R​\(ℓi​R\(T\)∣i\)\\mu^\{\(T\)\}\_\{i\}=\\mathbb\{E\}\_\{R\}\(\\ell^\{\(T\)\}\_\{iR\}\\mid i\)\. Orderingμi\(T\)\\mu^\{\(T\)\}\_\{i\}from largest to smallest gives a case\-risk tail; ordering realized draws gives a different tail\. The first identifies cases expected to be risky, while the second includes residual bad luck within otherwise moderate\-risk cases\. A projectibility declaration has to say which object the decision concerns\. Neither tail should be smuggled into a behavioural\-change statistic\.

### 5\.2 Released\-output illustration

The logical possibilities occur in released model evaluations\.\[[undefaay](https://arxiv.org/html/2607.26159#biba.bibx39)\]add task\-irrelevant context to benchmark items and report paired baseline and altered responses\. Their released design crosses four benchmarks with eight models, yielding 32 model–benchmark comparisons with 20 baseline and 20 context trials per item\. The empirical companion to this paper reanalyses those responses rather than regenerating proprietary\-model outputs, pins the source version and data hashes, and provides the estimator code at[https://github\.com/BrettRey/benchmark\-inference\-composition](https://github.com/BrettRey/benchmark-inference-composition)\.

One comparison shows both cancellation and the need to separate tail selection from tail estimation\. It pairs gpt\-5\.4 with MMLU\-Pro, a more demanding version of the Massive Multitask Language Understanding benchmark\[[undefaax](https://arxiv.org/html/2607.26159#biba.bibx38)\]\. Baseline accuracy is \.8119 and altered\-condition accuracy \.7904, a signed decline of \.0215\. Its directional components areF\+=\.0229F^\{\+\}=\.0229andF−=\.0444F^\{\-\}=\.0444, so the mean hides \.0673 of two\-sided item movement\.

Raw worst\-decile degradation is \.3680\. Reusing the same finite response trials to select and estimate the tail can inflate its apparent magnitude\. A response\-half procedure instead selects items on one half and estimates their changes on the other; the resulting estimate is \.2911\. The result isn’t that 29% is a deployment risk\. A 2\.15\-percentage\-point mean decline coexists with much larger deterioration among a selected group of items\.

Across the 32 fixed comparisons, both directional components exceed twice the absolute signed change in 22 cells\. Raw worst\-decile degradation exceeds the response\-half estimate by \.0906 on average, ranging from \.0330 to \.1696\. Those summaries describe one crossed dataset with shared items and related model lineages, not 32 independent replications\. They also show why a tail selected on noisy observed changes needs either a model of latent item effects or disjoint information for selection and estimation\.

A known\-truth simulation makes the separation from absolute loss explicit\. In the stable\-poor scenario, the same low\-performing items have unchanged response probabilities under baseline and altered conditions\. Latent worst\-tail change is zero, while their worst\-decile conditional expected loss is \.80\. The response\-half estimate of the change tail is \.0013; the cross\-fitted case\-risk tail is \.7911\. The disagreement persists without ambiguity about the truth; it follows from the estimands\.

These results verify that source aggregates can erase cancellation, concentration, and persistent loss\. They don’t show that any benchmark quantity predicts legal research, another system, or a future period\. Indeed, the released data contain no research requests, database searches, citations, lawyer reviews, or client outcomes\. Their relevance is architectural: if an upstream evaluator publishes only a total, a downstream evaluator can’t inspect whether the source failures line up with the features that define its target\.

### 5\.3 Evidence preservation at the interface

A projectibility audit imposes an information requirement on source reports\. The developer needn’t publish every conceivable statistic, but it should retain the resolution at which target\-relevant distinctions can later be checked: item identifiers or auditable equivalents, item\-level outputs and scores, template or source clusters, prompts and tool conditions, repeated\-response structure, scorer decisions, model build and lineage, and registered alterations\. A domain profile is better than a total when domains matter, but it can still hide item\-level reversals and stable failures within each domain\.

The requirement isn’t maximal disaggregation for its own sake\. Raw data can be sensitive, proprietary, or too large to release\. The relevant standard is whether a declared downstream claim can be audited\. Access controls, secure evaluation environments, independent auditors, or sufficiently detailed derived data may answer the need\. What doesn’t answer it is an aggregate from which several target\-relevant states are observationally indistinguishable\.

The legal example illustrates the interface\. A high legal\-domain mean doesn’t reveal whether errors concentrate on questions requiring current authority or on outputs with plausible but unsupported citations\. A local evaluator needs those distinctions to decide what to sample and score\. If the developer can report item changes under retrieval\-like distractors, citation\-related errors, and build provenance, that evidence may inform the firm’s design\. It still doesn’t replace the firm’s observations of its own requests and review procedure\.

## 6What the Non\-Composition Principle Changes for AGI Evaluation

Multidomain AGI evaluations make the composition problem unusually acute\. The labelgeneralinvites a conclusion that reaches beyond every fixed set of tasks, systems, and conditions used to construct the score\. Adding domains increases coverage, but coverage and inferential reach aren’t the same\. Before a total or profile is read as evidence of general capability, three questions have to be separated: whether its components share a scale, whether the permitted tradeoffs are acceptable, and whether the resulting description projects to the target claim\.

### 6\.1 Commensurability, compensability, and projectibility

Commensurabilityasks whether differences on component scores have a common interpretation that makes addition meaningful\. Putting a legal score and a quantitative score on a common scale from zero to one doesn’t show that a \.05 increase represents the same amount of improvement in both\. Equal weights don’t solve the problem; components with different variances and covariances can make unequal statistical contributions even under equal nominal weights\[[undefau](https://arxiv.org/html/2607.26159#biba.bibx9), 305–308\]\.

Compensabilityasks whether gains on one component may offset failures on another for the interpretation or use at issue\. A weighted sum permits such substitutions over the reported ranges\. That may be acceptable for a descriptive index\. It’s unacceptable when the target imposes a bottleneck or red line\. In the legal case, a high quantitative score doesn’t compensate for an unresolvable citation, and success on routine termination\-clause research doesn’t compensate for failure on reasonable\-notice requests if both uses are to be authorized\. Value models can represent accepted tradeoffs, but the weights are range\-dependent preferences rather than context\-free measures of domain importance\[[undefaaj](https://arxiv.org/html/2607.26159#biba.bibx24)\]\.

Projectibilityasks whether the aggregate supports a specified claim beyond the cases and conditions used to construct it\. Even a commensurable, defensibly weighted index may provide little warrant for a projection to a new task, system, or period\. Conversely, a score can predict a bounded criterion without being an interpretable measure of a single latent capability\. The measurement, prediction, and decision roles need different evidence\.

This tripartite distinction prevents one objection from doing the work of another\. Showing that a total predicts an external outcome doesn’t establish that its components measure one construct\. Showing a coherent factor structure doesn’t establish that its compensatory total is suitable for a decision\. Showing that a battery samples ten domains doesn’t establish that its score projects to open\-ended general capability\.

### 6\.2 Four roles for a cognitive taxonomy

The CHC framework used in some AGI proposals can enter an evaluation in at least four roles\[[undefaw](https://arxiv.org/html/2607.26159#biba.bibx11),[undefaao](https://arxiv.org/html/2607.26159#biba.bibx29),[undefaaw](https://arxiv.org/html/2607.26159#biba.bibx37),[undefaaf](https://arxiv.org/html/2607.26159#biba.bibx20)\]\. First, it can organize benchmark content: evaluators select tasks under labels inherited from a theory of human cognitive abilities\. Second, correlations among task scores in a declared population of artificial systems can be summarized by a factor model\. Third, the resulting scores can be interpreted as evidence of capacities with a nomological or process\-based meaning\. Fourth, the score or profile can be used to predict an external criterion or guide a decision\. Evidence for one role doesn’t establish the next\.

A content taxonomy can ensure that a battery doesn’t consist entirely of one familiar task format\. It doesn’t establish that artificial systems reproduce the human covariance structure that motivated the taxonomy\. A positive manifold among model scores can support a statistical general factor in the sampled model–task matrix without identifying the factor as human\-like intelligence or a causally efficacious internal capacity\. A factor interpretation can become more substantive through a nomological network of expected task, process, and external relations, but those relations still need to be tested in the artificial\-system population\[[undefax](https://arxiv.org/html/2607.26159#biba.bibx12),[undefaaa](https://arxiv.org/html/2607.26159#biba.bibx15)\]\.

\[[undefaz](https://arxiv.org/html/2607.26159#biba.bibx14)\]distinguishesconstruct representation, an account of the processes, strategies, and knowledge involved in item responses, fromnomothetic span, the network of relations between scores and external measures, groups, and tasks\. The two rest on different units of variation\. Construct representation“is concerned with task variability rather than subject variability”, while nomothetic span“is assessed by individual differences data”, making it“possible to obtain strong support for one, but not for the other”\[[undefaz](https://arxiv.org/html/2607.26159#biba.bibx14), 180\]\.

Process evidence can support an account of how a system produced benchmark responses without showing that the score predicts another task\. External prediction can be strong while the process interpretation remains unsettled\. Treating both as one achievement under the labelvalidityobscures which conclusion the evidence supports\.

Empirical work on artificial\-system psychometrics illustrates the gap\. Across 591 nominally distinct models and 12 tests,\[[undefaag](https://arxiv.org/html/2607.26159#biba.bibx21), 4–5\]report a positive manifold and a strong general factor\. Their proposed four\-factor hierarchy beneath it wasn’t interpretable, and the task set was largely verbal; dependence among nominally distinct models also limits the system population represented\[[undefaag](https://arxiv.org/html/2607.26159#biba.bibx21), 7–9\]\. The study supports a factor model for that sampled matrix\. It doesn’t establish the CHC decomposition for a retrieval\-augmented system, a human–AI workflow, or future systems\.

Comparability across systems is another projection\. Interpreting score differences as differences in the same abilities requires evidence that items and score relations retain their meaning across model families\. Depending on the design, that can involve rubric checks, family\-specific item analyses, differential item functioning, and invariance models\[[undefaap](https://arxiv.org/html/2607.26159#biba.bibx30)\]\.\[[undefaah](https://arxiv.org/html/2607.26159#biba.bibx22)\]give a direct warning: across 17 language models, human psychometric instruments showed moderate reliability across item and prompt variations while their scores failed to align with, and sometimes ran opposite to, model behaviour on downstream tasks\. Reliability of the instrument under selected variations didn’t supply criterion validity\.

Philosophically informed benchmark construction can strengthen the content and interpretation links\. For example,\[[undefao](https://arxiv.org/html/2607.26159#biba.bibx3)\]develop a behavioural benchmark for scientific understanding around retrieval, explanation, and counterfactual performance\. Such work makes the target construct more explicit than a battery assembled by domain coverage alone\. It still leaves empirical questions about which artificial systems satisfy the proposed relations and whether benchmark performance reaches a different task or deployment setting\. The non\-composition principle isn’t an objection to theorized benchmarks\. It specifies the additional work their results require downstream\.

### 6\.3 Why accumulated coverage doesn’t entail unrestricted generality

A finite battery can define and measure an index over its construction sample\. It can also support bounded generalizations when its item\-generating process, sampling assumptions, and uncertainty are defensible\. The stronger claim begins when the index is read as evidence of capability across tasks and conditions not represented by that process\. More domains may make the source description broader, but a run of successful source descriptions doesn’t by itself identify the rule by which they extend\.

This is the Goodman point in practical form\. Passing tasks in law, mathematics, coding, and spatial reasoning fits a general\-capability hypothesis\. It can also fit a collection of narrower hypotheses under which performance depends on familiar interfaces, training overlap, static inputs, short horizons, or the absence of tool and user interactions\. The hypotheses agree on the battery and diverge elsewhere\. Domain count alone doesn’t select among them\. Evidence has to vary the suspected conditions or sample the target cases\.

Cross\-domain covariation is stronger evidence than a domain count\. Across 591 models and 12 tests, every score pair correlated positively, and a mathematics composite loaded most strongly on the general factor\[[undefaag](https://arxiv.org/html/2607.26159#biba.bibx21)\]\. The pattern makes one tested domain informative about others in that sampled matrix but doesn’t establish unrestricted reach\. Separate studies might support transfer from benchmark mathematics to unseen textbook problems, from coding exercises to repository\-level bug fixes, and from supplied\-text law questions to bounded research requests\. Their union says more than any one, but without evidence relating those targets it doesn’t establish arbitrary tasks, domain interactions, later systems, or novel tool environments\.

Bounded evidence can accumulate legitimately\. If the cases reached by one projection are sampled by the next, interface assumptions are tested, and dependence and uncertainty are preserved, the chain can extend\. A system’s benchmark profile might predict target\-task outcomes across independent builds; target\-task outcomes might predict final workflow results across sites; those results might remain stable in chronological holdouts\. The conclusion is as broad as the composed path and no broader\. The non\-composition principle blocks an automatic inference, not empirical progress\.

### 6\.4 Aggregates as descriptions, measures, predictors, and decision inputs

Many disputes about AGI scores are disputes about role\. A total can be a transparent descriptive index even when it isn’t a measure of one attribute\. It can be a useful predictor even when its construct interpretation is contested\. It can enter a decision model without deciding the weights, constraints, or alternatives\. Each role should be stated rather than allowed to slide into the next\.

As adescription, the score reports a declared operation on the construction sample\. Accuracy here is largely a matter of calculation, sampling, and scoring\. As ameasure, it represents a capability or attribute and requires content, structural, process, and external evidence appropriate to that interpretation\[[undefaaq](https://arxiv.org/html/2607.26159#biba.bibx31),[undefar](https://arxiv.org/html/2607.26159#biba.bibx6)\]\. As apredictor, it requires target\-matched out\-of\-sample performance at the independent unit named by the claim\. As adecision input, it also requires a defensible relation to consequences, feasible policies, and the people who bear their effects\.

A common error is to validate one role and write as though the others follow\. A high test–retest reliability supports score stability under the repeated protocol, not a capability ontology\. A factor loading records association in a fitted model, not an internal mechanism\. A predictive relation across systems doesn’t show that the predictor measures a causally efficacious AGI attribute\. A policy that uses the predictor successfully in one organization doesn’t show that the same rule is suitable elsewhere\.

The role distinctions also sharpen reporting\. An AGI index should be named as an index when it’s defined by its scoring rule\. A capability interpretation should state the nomological and process commitments that make the term more than a heading\. A predictive claim should name the target cases, systems, and loss\. A release decision should show why its tradeoffs and hard constraints are appropriate\. Projectibility is then assessed for each move rather than attributed to the total as a prestige property\.

## 7Procedure and Evidential Responsibility

A projectibility audit is useful only if it changes what evaluators record and test\. Its procedure begins before calculation, with a sentence that can be contradicted by observations and a typed description of the source and target\. It ends neither with a benchmark report nor with authorization\. Later outcomes and material system changes reopen the links they bear on\.

### 7\.1 Who can supply the evidence

The evidential burden is distributed because no single actor observes the whole chain\. A developer can describe the source evaluation in detail\. It can identify the exact model build and declared lineage; publish or provide controlled access to items, prompts, tool settings, scoring rules, item\-level outputs, and repeated\-response structure; document which templates or sources were held out; report registered alterations and known failures; and state which product changes invalidate the result\.

Psychometric analyses can improve the quality and interpretation of the instrument itself, as work on item\-response methods for NLP benchmarks illustrates\[[undefan](https://arxiv.org/html/2607.26159#biba.bibx2)\]\. Those analyses remain conditional on the systems, items, and score interpretations studied\.

A deployer observes different facts\. Only the law firm in the running example can define which requests it receives, sample them from its logs, prepare local authority lists, observe its lawyers’ review, compare the policies it can adopt, and record whether a defect remains internal or reaches a client or court\. A hospital, school, or public agency would have corresponding task, workflow, exposure, and outcome records\. The inability of a developer to foresee those details narrows the developer’s warranted claim\. It doesn’t turn a generic benchmark into evidence about every use\.

Some evidence belongs at the interface and requires cooperation\. The deployer needs notice of model, safety\-layer, retrieval, or API changes; enough configuration information to freeze the tested object; and access to outputs at the resolution required for local scoring\. The developer needs feedback about failures that source evaluation didn’t represent\. Independent evaluators may test both sides, but they can’t infer proprietary lineage or local workflow conditions from a public product label\. Where information remains unavailable, the corresponding projection remains conditional or unsupported rather than being filled by an assumption of continuity\.

Responsibility is divided but not dissolved\. Developers are responsible for the interpretations and uses they advertise and for foreseeable interface failures; deployers remain responsible for the decisions they make and the evidence available only in context\. A claim that no party can test may still be rhetorically attractive, but the absence of an evidential owner is itself a defect in the argument\.

### 7\.2 The projectibility declaration

For confirmatory use, the evaluator should complete aprojectibility declarationbefore opening the final holdout\. Exploratory work can use the same fields, but it should reserve new evidence for confirmation\. The declaration has ten parts\.

1. 1\.Write one source–target claim\.State what was observed and the new cases, interpretation, or outcome claimed\. Replacelegal capability transfers to practicewith a statement such as: final memos produced by the frozen application and registered review procedure will have a post\-review substantive\-defect probability below 3% on Ontario termination\-clause requests admitted by the written rule\.
2. 2\.Type the source and target\.Record the object, population, conditions, outcome, and period at each endpoint\. If the claim is a capability attribution, it isn’t an empirical endpoint; record instead the observable relation the downstream link needs and the evidence that would test it\. A shared label doesn’t count as endpoint identity\.
3. 3\.Name the extension\.Specify whether the claim generalizes to further items, interprets a score as a construct, extrapolates to another task, transports across systems or sites, predicts a later period, explains an outcome, or links evidence to a policy\. Composite claims should be split until each boundary is visible\.
4. 4\.List plausible defeaters\.Name actual differences that could change the result: live rather than supplied retrieval, added irrelevant cases, another database index, a different reviewer instruction, a new model lineage, conflicting authority, or a later work period\. The list is justified by prior failures, domain knowledge, process evidence, and stakeholder concerns, not by a generic field headedcontext\.
5. 5\.Assign evidence and ownership\.For each assumption or defeater, state the observation that bears on it and who can produce that observation\. If no actor can observe whether defects reach clients, a claim about client consequences isn’t ready for confirmation\.
6. 6\.Match the sampling and holdout unit to the claim\.Keep complete templates out for a template claim, client\-file clusters out for a request claim, alteration types out for a context claim, lineages out for a system claim, offices out for a site claim, and later periods out for a temporal claim\. Repeated outputs estimate variation within a request–condition pairing; they don’t create new requests, systems, or periods\.
7. 7\.Define the observation and loss\.State exactly what one row represents and what is scored\. Distinguish a request, a generated draft, a lawyer review, a final memo, a system build, and a work period\. Keep ordinal hazard classes separate unless a defensible cardinal scale exists\. Identify red\-line outcomes that may not be averaged away\.
8. 8\.Register support, defeat, and unresolved conditions\.Give the estimate, uncertainty requirement, minimum sample or tail count, multiplicity treatment, and any fallback claim before examining confirmatory results\. An inconclusive result remains inconclusive; it isn’t evidence of equivalence or robustness\.
9. 9\.Audit the interfaces\.For every adjacent link, ask whether the upstream target is the downstream source in object, population, condition, outcome, and time; whether assumptions are compatible; and how dependence and uncertainty are propagated\. Classify the relation as composition, convergence, replacement, or no evidential connection\.
10. 10\.Separate the empirical claim from the decision\.List feasible policies, affected parties, observed and prospective consequences, costs, hard constraints, and update triggers\. State which empirical results enter the choice and which values determine the boundary\. A validity judgment doesn’t select a policy by itself\.

Table[5](https://arxiv.org/html/2607.26159#S7.T5)gives a compact record\. The final column prevents the declaration from becoming a set of decorative headings: every entry resolves into a value, sampling rule, or named document\.

Table 5:Minimum projectibility\-declaration record\.
### 7\.3 Designs matched to common projections

The declaration changes the statistical design because the independent unit follows the claim\. To generalize from observed to unobserved responses under the same item rules, hold out item or template clusters rather than random response rows\. To extrapolate from benchmark questions to professional tasks, sample external tasks by a written rule and score them against independently prepared criteria\. To claim robustness to a new alteration family, keep one whole alteration type out of tuning\.

To claim transfer to another system, evaluate a genuinely different build or lineage; repeated drafts from one build only make the estimate for that build more precise\. To claim transfer to another office, run the complete task and review procedure there\. To predict later performance, use a chronological holdout at the stated horizon\.

Construct claims need another design\. A taxonomy needs content argument\. A factor interpretation needs a declared task and system population, adequate variation, and checks that the fitted relations aren’t artefacts of a narrow item set\. A nomological claim needs predicted convergent, discriminant, process, and criterion relations\. A causal explanation needs interventions or other identification conditions that discriminate plausible rivals\. More benchmark items can improve precision within the tested operation while leaving each of those obligations untouched\.

Predictive claims make an especially common unit error visible\. Suppose the question is whether benchmark profiles predict legal\-research defects across systems\. The independent units are systems or lineages, each with a benchmark profile and a defect estimate on the same held\-out requests\. With one system, its profile doesn’t vary across requests and can’t predict which request fails\. A request\-level model instead needs request\-level predictors\. Fitting the system profile to hundreds of request rows manufactures a sample size by repeating a constant\.

The same discipline applies to tails and groups\. A worst\-decile estimate selected after examining noisy item effects is subject to selection inflation; response splitting, cross\-fitting, or a hierarchical model can address that problem for a stated estimand\. None addresses transport to a new task type or system\. Sparse group samples can leave a group\-specific claim unresolved; a pooled estimate isn’t automatically a licensed substitute\. If the decision constrains the maximum defect tail across two request types, uncertainty has to account for selecting that maximum\.

### 7\.4 Reporting and updating

A projectibility report should preserve continuous estimates and uncertainty even when a policy uses a boundary\. The supported, defeated, or unresolved label summarizes the registered relation; it doesn’t replace the estimate\. Reports should also distinguish raw descriptive quantities from null\-referenced, split\-sample, cross\-fitted, or model\-based estimates and state what each targets\. A visually similar number can answer a different question\.

After authorization, the evidential record continues\. A shadow phase can run the application beside ordinary work without allowing it to determine the memo sent onward\. The firm records draft defects, corrections, review time, final defects, and any incident\. Continued use adds exposure and outcome records\. A model update, retrieval\-index change, revised prompt, new request type, changed review procedure, or shift in observed failures reopens the relevant links\. The update needn’t invalidate every part of the argument: an unchanged inclusion rule may remain useful while the system\-transport link requires new evidence\.

Frequent model releases don’t require every part of an evaluation to be repeated from scratch\. An update reopens the links whose object, conditions, or assumptions it changes\. Task definitions, inclusion rules, reference standards, scoring procedures, and evidence about unchanged parts of the workflow can often be reused; claims about model outputs require new evidence when the tested build changes materially\.

Version pinning, change logs, regression tests, rotated holdouts, and shadow evaluation can reduce the cost, while the extent of retesting should reflect the stakes, the opacity of the change, and the scope of the authorized claim\. If a provider doesn’t permit a build to be frozen or supply enough information to determine what changed, the result is a narrower or conditional authorization\. Sometimes rapid model turnover makes the proposed use impractical; it doesn’t make evidence about an earlier model applicable by default\.

Monitoring isn’t a substitute for predeployment evidence\. Allowing a high\-stakes system to produce consequences in order to learn whether it’s safe simply relocates the cost of uncertainty\. Nor is predeployment testing a substitute for monitoring, since later users, tasks, and systems can differ\. The two answer temporal links on opposite sides of authorization\.

The audit can end with a narrow claim\. Sparse or mismatched evidence doesn’t imply that the system is generally unreliable; it implies that a broader claim lacks support\. Conversely, a successful local study doesn’t prove general capability\. Precision about scope is an empirical result, not a rhetorical concession\.

## 8Limitations

Projectibility isn’t a complete theory of induction\. The framework requires evaluators to identify plausible defeaters, but observations never determine that list by themselves\. Domain theory, prior failures, causal knowledge, institutional experience, and affected stakeholders all shape which differences are investigated\. Those sources can be incomplete or contested\. A projectibility audit makes the commitments visible; it doesn’t guarantee that every important difference has been anticipated\.

Variation in support for a fixed projection doesn’t make projectibility a scalar property that can be calculated independently of a claim\. The supported, defeated, and unresolved statuses depend on a declared target, loss, and evidential standard\. That relativity can be abused by narrowing a claim until it becomes trivial\. Prospective declaration and independent scrutiny limit such rescue, but they don’t eliminate strategic framing\. The practical question is whether the narrow claim remains useful enough to justify the proposed policy\.

The interface requirements aren’t a replacement for specialized methods\. Causal projections require causal identification; psychometric interpretations require appropriate measurement models and construct evidence; predictive claims require target\-matched validation; and decisions require normative and institutional justification\. The framework’s contribution is to keep those analyses from being joined merely because their outputs share a label\.

The legal results are constructed\. They demonstrate how a complete audit assigns different statuses, not how any actual assistant performs\. The reanalysis and simulations establish statistical distinctions under the released and simulated designs\. They don’t validate the projectibility framework against deployment outcomes or show that the reported benchmark quantities predict another target\. The framework itself should be studied prospectively: evaluators could compare decisions, failure detection, and claim revision with and without an explicit interface audit\.

Some evidence will remain inaccessible\. Proprietary training overlap can make model lineages uncertain; privacy and privilege can limit release of local cases; rare harms can make direct estimation impractical; and important future conditions may be unforeseeable\. The response isn’t to infer continuity by default\. Evaluators can use secure audits, partial identification, stress tests, sensitivity analysis, and narrower authorization, while recording which links remain conditional\.

Finally, the procedure can be expensive\. Sampling target tasks, preparing independent reference standards, testing complete workflows, and retaining outcome records require expertise and time\. Those costs should be compared with the value and stakes of the proposed use\. A low\-consequence drafting aid may justify a lighter evidential standard than a system whose outputs enter court filings without independent verification\. Claim\-relative standards aren’t evidential relativism; they connect the cost of being wrong to the strength and kind of support required\.

## 9Conclusion

Validity\-centred AI evaluation has made an essential correction: validity belongs to a proposed interpretation and use, not to a benchmark in isolation\. This paper adds a non\-composition principle for benchmark\-to\-use arguments\. A benchmark result may support several individually warranted steps without warranting the chain assembled from them\. A capability attribution may support one step, but it doesn’t itself specify the downstream cases or criterion relation\.

Projectibility concerns whether one bounded extension is warranted\. Goodman shows why source fit alone can’t choose among rival extensions, and argument\-based validity locates each extension among explicit claims and assumptions\. A projectibility audit first asks whether links presented as adjacent meet in object, population, conditions, outcome, and period\. A mismatch makes the proposed composition undefined unless separately bridged\. If the fields align, warrant still fails to transmit when assumptions or effect modifiers conflict or when dependence and uncertainty are lost\.

The legal example shows the practical consequence\. A developer’s score on supplied\-text questions and a firm’s local study of reviewed memos can each be sound while remaining parallel\. The local study may support mandatory\-review use on one request type, defeat another, and leave a changed retrieval condition unresolved\. It supplies new observations for a bounded inference rather than causing the benchmark score to carry into practice\. A later build, another office, or another request type reopens a different link\.

Aggregation matters because joins require information at the resolution of the downstream claim\. One stable mean can conceal balanced changes, concentrated deterioration, or stable serious failures\. For AGI evaluation, broader domain coverage can likewise improve description without establishing commensurability, acceptable compensation, a common capability, or unrestricted reach\. The warranted conclusion extends only across the targets and joins actually supported\.

Developers should report what their evaluations observed and the boundaries they tested; deployers should sample the tasks, workflows, and consequences available only in context\. Supported bounded projections can accumulate when endpoints align and warrant transmits across each join\. A warranted benchmark\-to\-use argument is a declared, revisable chain with evidence for each link, aligned endpoints, compatible assumptions, and dependence and uncertainty carried through\.

## Appendix AStatistical Module for Paired\-Condition Reports

This appendix gives the fuller statistical module used in the empirical companion\. It’s optional: an evaluation should report the quantities its target claim requires rather than treating the module as a checklist\.

Let domains be indexed byg=1,…,Gg=1,\\ldots,G, items within domain byi=1,…,Ngi=1,\\ldots,N\_\{g\}, and conditions bycc, with baseline0and alterationpp\. Letsi​g​c∈\[0,1\]s\_\{igc\}\\in\[0,1\]be an item score andδi​g​p=si​g​p−si​g​0\\delta\_\{igp\}=s\_\{igp\}\-s\_\{ig0\}\. The domain mean is

ag​c=1Ng​∑i=1Ngsi​g​c,a\_\{gc\}=\\frac\{1\}\{N\_\{g\}\}\\sum\_\{i=1\}^\{N\_\{g\}\}s\_\{igc\},and the signed level change isLg​p=ag​p−ag​0L\_\{gp\}=a\_\{gp\}\-a\_\{g0\}\. An overall equal\-domain changeG−1​∑gLg​pG^\{\-1\}\\sum\_\{g\}L\_\{gp\}describes an equally weighted set of domains; it isn’t a deployment mixture unless the target distribution gives the domains those weights\.

Item instabilityINSg​p\\mathrm\{INS\}\_\{gp\}, the mean positive\-change componentFg​p\+F^\{\+\}\_\{gp\}, and the mean negative\-change componentFg​p−F^\{\-\}\_\{gp\}are

INSg​p=1Ng​∑i\|δi​g​p\|,Fg​p\+=1Ng​∑i\(δi​g​p\)\+,Fg​p−=1Ng​∑i\(−δi​g​p\)\+\.\\mathrm\{INS\}\_\{gp\}=\\frac\{1\}\{N\_\{g\}\}\\sum\_\{i\}\|\\delta\_\{igp\}\|,\\qquad F^\{\+\}\_\{gp\}=\\frac\{1\}\{N\_\{g\}\}\\sum\_\{i\}\(\\delta\_\{igp\}\)\_\{\+\},\\qquad F^\{\-\}\_\{gp\}=\\frac\{1\}\{N\_\{g\}\}\\sum\_\{i\}\(\-\\delta\_\{igp\}\)\_\{\+\}\.The identitiesLg​p=Fg​p\+−Fg​p−L\_\{gp\}=F^\{\+\}\_\{gp\}\-F^\{\-\}\_\{gp\}andINSg​p=Fg​p\+\+Fg​p−\\mathrm\{INS\}\_\{gp\}=F^\{\+\}\_\{gp\}\+F^\{\-\}\_\{gp\}make cancellation explicit\. For deterministic binary scores,F\+F^\{\+\}andF−F^\{\-\}are the incorrect\-to\-correct and correct\-to\-incorrect transition proportions\.

For a prespecifiedq∈\(0,1\]q\\in\(0,1\], letmg=⌈q​Ng⌉m\_\{g\}=\\lceil qN\_\{g\}\\rceiland order paired changes from smallest to largest\. Domain\-conditional worst\-tail degradation is

WTDq,g​p=−1mg​∑j=1mgδ\(j\)​g​p\.\\mathrm\{WTD\}\_\{q,gp\}=\-\\frac\{1\}\{m\_\{g\}\}\\sum\_\{j=1\}^\{m\_\{g\}\}\\delta\_\{\(j\)gp\}\.It’s an empirical expected\-shortfall analogue applied to the loss−δ\-\\delta\[[undefam](https://arxiv.org/html/2607.26159#biba.bibx1),[undefaau](https://arxiv.org/html/2607.26159#biba.bibx35)\]\. Becausemgm\_\{g\}rounds upward, the empirical tail mass ismg/Ngm\_\{g\}/N\_\{g\}\. A positive value indicates deterioration in the selected lower tail; a negative value means even that tail improved on average\.

Absolute loss requires another variable\. For targetTT, letRRcollect response, scorer, and downstream\-outcome variation within an item–condition pairing\. Letℓi​g​c​R\(T\)∈\[0,1\]\\ell^\{\(T\)\}\_\{igcR\}\\in\[0,1\]be realized loss andμi​g​c\(T\)=𝔼R​\(ℓi​g​c​R\(T\)∣i,g,c\)\\mu^\{\(T\)\}\_\{igc\}=\\mathbb\{E\}\_\{R\}\(\\ell^\{\(T\)\}\_\{igcR\}\\mid i,g,c\)conditional expected loss\. Orderingμ\(T\)\\mu^\{\(T\)\}identifies high\-risk cases; ordering realized draws ofℓ\(T\)\\ell^\{\(T\)\}identifies high\-loss realizations\. They coincide only under restrictive conditions such as no residual variation within a case\.

For a deployment\-facing target, letJ∼PTJ\\sim P\_\{T\}index a joint draw over every varying feature in scope: system, task or episode, request type, alteration, operator, affected group, and time\. WithQV,TQ\_\{V,T\}the quantile function of variableVVunderPTP\_\{T\}, define upper expected shortfall

ESq,T\+⁡\(V\)=1q​∫1−q1QV,T​\(u\)​du\.\\operatorname\{ES\}^\{\+\}\_\{q,T\}\(V\)=\\frac\{1\}\{q\}\\int\_\{1\-q\}^\{1\}Q\_\{V,T\}\(u\)\\,\\mathrm\{d\}u\.Then a deployment\-facing change tail can be written asESq,T\+⁡\(−δJ\)\\operatorname\{ES\}^\{\+\}\_\{q,T\}\(\-\\delta\_\{J\}\), a case\-risk tail asESq,T\+⁡\(μJ\(T\)\)\\operatorname\{ES\}^\{\+\}\_\{q,T\}\(\\mu^\{\(T\)\}\_\{J\}\), and a realized\-loss tail asESq,T\+⁡\(ℓJ​R\(T\)\)\\operatorname\{ES\}^\{\+\}\_\{q,T\}\(\\ell^\{\(T\)\}\_\{JR\}\)\. The target distribution, not the visual similarity of benchmark items, determines the probability mass being ordered\.

These formulas support three different readings\. With deterministic decoding and a fixed paired item set, they’re exact descriptions of that set\. Under a declared stochastic response and scoring protocol, they estimate latent item quantities from repeated trials and scorers\. Inference to a population of new items, contexts, systems, or periods requires sampling or holdout from that population\. Repeated responses from one item estimate within\-item variation; they don’t create a probability sample of items\.

Nonlinear quantities introduce selection problems\. Absolute values create a positive response\-sampling floor even under a null condition effect\. Selecting the worst observed tail and estimating it on the same responses selects noise and exaggerates magnitude, a Type M error in the sense of\[[undefaac](https://arxiv.org/html/2607.26159#biba.bibx17)\]\. A baseline\-only pseudo\-null resampling procedure estimates the nonlinear statistic expected under a simulation with no condition difference\. Subtracting that expectation yields a useful null\-referenced diagnostic, not a general bias correction: under a real effect, variances and selection can change nonadditively\.

When latent item changes are the target, a hierarchical model can partially pool noisy extremes before the tail is summarized\. A simpler response\-half estimator selects items on one half of the trials and estimates their changes on the other\. Cross\-fitting reverses the halves and combines the results\. This avoids reusing the same response noise, but its estimand is the latent effect of a set selected noisily, not necessarily the ideal tail selected from unobserved latent effects\. Reports should name that difference\.

Uncertainty and resampling units follow the projectibility declaration\. Preserve paired items, shared baselines, alteration realizations, scorers, client\-file clusters, and model lineages\. A bootstrap over individual prompts is wrong when several prompts come from one client file; the file cluster is the relevant unit\. Holding out one irrelevant document doesn’t test a new alteration family\. Generating additional drafts from one model doesn’t test a new system\. Response\-noise correction and projection across populations are separate problems\.

## Appendix BInterface Audit Questions

For each pair of adjacent links, the evaluator can use the following questions as a compact composition audit\.

1. 1\.Do the endpoints denote the same object?If the upstream target is a base model and the downstream source is a retrieval application or reviewed workflow, what evidence links them? If either endpoint is named as a capability, which tested score–criterion relation stands in for it?
2. 2\.Are the cases aligned?Does the downstream sampling rule admit the cases reached upstream, with the same exclusions, grouping, and target weights?
3. 3\.Is the outcome continuous?Does the first link deliver the variable the second consumes, or has accuracy become a factor score, defect probability, reviewed outcome, or consequence without a tested mapping?
4. 4\.Are operating conditions compatible?Which prompts, tools, databases, reviewers, sites, and periods differ, and which observations test those differences?
5. 5\.Are assumptions inherited consistently?Does the second link rely on an invariance, independence, or process claim contradicted or left open by the first?
6. 6\.Is the evidence as independent as claimed?Which items, templates, scorers, training runs, model lineages, or development choices are shared?
7. 7\.Has uncertainty been carried through?Are estimates, selection procedures, and unresolved qualitative assumptions propagated rather than replaced by point conclusions?
8. 8\.Is this actually composition?Would the downstream claim be supported if the upstream study didn’t exist? If so, the relation may be replacement or convergence rather than composition\.
9. 9\.What observation would defeat the composed claim?A chain with no stated defeater hasn’t yet exposed its empirical content\.
10. 10\.Does the conclusion stop at the reached target?Successful links warrant the declared path, not unspecified tasks, systems, or times beyond it\.

A completed audit needn’t show that every interface is secure\. Its first function is diagnostic\. A missing link can lead to a new study, a narrower claim, a conditional conclusion, or rejection of the proposed use\. What it can’t legitimately produce is scope by accumulation: several passed evaluations don’t add up to a claim whose interfaces were never tested\.

## References

- \[undef\]Carlo Acerbi and Dirk Tasche“On the Coherence of Expected Shortfall”In*Journal of Banking & Finance*26\.7, 2002, pp\. 1487–1503DOI:[10\.1016/S0378\-4266\(02\)00283\-2](https://dx.doi.org/10.1016/S0378-4266(02)00283-2)
- \[undefa\]Dominik Bachmann et al\.“fl\-IRT\-ing with Psychometrics to Improve NLP Bias Measurement”In*Minds and Machines*34\.37, 2024DOI:[10\.1007/s11023\-024\-09695\-9](https://dx.doi.org/10.1007/s11023-024-09695-9)
- \[undefb\]Kristian Gonzalez Barman, Sascha Caron, Tom Claassen and Henk Regt“Towards a Benchmark for Scientific Understanding in Humans and Machines”In*Minds and Machines*34\.6, 2024DOI:[10\.1007/s11023\-024\-09657\-1](https://dx.doi.org/10.1007/s11023-024-09657-1)
- \[undefc\]Andrew M\. Bean et al\.“Measuring what Matters: Construct Validity in Large Language Model Benchmarks” NeurIPS 2025 Track on Datasets and BenchmarksIn*arXiv preprint arXiv:2511\.04703*, 2025DOI:[10\.48550/arXiv\.2511\.04703](https://dx.doi.org/10.48550/arXiv.2511.04703)
- \[undefd\]Olivier Binette and Jerome P\. Reiter“Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework”, 2024arXiv:[https://arxiv\.org/abs/2406\.10366](https://arxiv.org/abs/2406.10366)
- \[undefe\]Denny Borsboom, Gideon J\. Mellenbergh and Jaap Heerden“The Concept of Validity”In*Psychological Review*111\.4, 2004, pp\. 1061–1071DOI:[10\.1037/0033\-295X\.111\.4\.1061](https://dx.doi.org/10.1037/0033-295X.111.4.1061)
- \[undeff\]Samuel R\. Bowman and George Dahl“What Will it Take to Fix Benchmarking in Natural Language Understanding?”In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, 2021, pp\. 4843–4855DOI:[10\.18653/v1/2021\.naacl\-main\.385](https://dx.doi.org/10.18653/v1/2021.naacl-main.385)
- \[undefg\]Richard Boyd“Realism, Anti\-Foundationalism and the Enthusiasm for Natural Kinds”In*Philosophical Studies*61\.1–2, 1991, pp\. 127–148DOI:[10\.1007/BF00385837](https://dx.doi.org/10.1007/BF00385837)
- \[undefh\]Robert L\. Brennan“Generalizability Theory”New York: Springer, 2001DOI:[10\.1007/978\-1\-4757\-3456\-0](https://dx.doi.org/10.1007/978-1-4757-3456-0)
- \[undefi\]Stefan Buijsman“Accuracy is not all you need\! The Reasons to Require AI Explainability”In*Minds and Machines*36\.14, 2026DOI:[10\.1007/s11023\-026\-09768\-x](https://dx.doi.org/10.1007/s11023-026-09768-x)
- \[undefj\]John B\. Carroll“Human cognitive abilities: A survey of factor\-analytic studies”Cambridge University Press, 1993DOI:[10\.1017/CBO9780511571312](https://dx.doi.org/10.1017/CBO9780511571312)
- \[undefk\]Lee J\. Cronbach and Paul E\. Meehl“Construct Validity in Psychological Tests”In*Psychological Bulletin*52\.4, 1955, pp\. 281–302DOI:[10\.1037/h0040957](https://dx.doi.org/10.1037/h0040957)
- \[undefl\]Roel Dobbe and Anouk Wolters“Toward Sociotechnical AI: Mapping Vulnerabilities for Machine Learning in Context”In*Minds and Machines*34\.12, 2024DOI:[10\.1007/s11023\-024\-09668\-y](https://dx.doi.org/10.1007/s11023-024-09668-y)
- \[undefm\]Susan E\. Embretson“Construct Validity: Construct Representation versus Nomothetic Span”In*Psychological Bulletin*93\.1, 1983, pp\. 179–197DOI:[10\.1037/0033\-2909\.93\.1\.179](https://dx.doi.org/10.1037/0033-2909.93.1.179)
- \[undefn\]Timo Freiesleben“Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks”In*arXiv preprint arXiv:2603\.15121*, 2026DOI:[10\.48550/arXiv\.2603\.15121](https://dx.doi.org/10.48550/arXiv.2603.15121)
- \[undefo\]Timo Freiesleben and Sebastian Zezulka“The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models” Submitted 27 October 2025In*arXiv preprint arXiv:2510\.23191*, 2025DOI:[10\.48550/arXiv\.2510\.23191](https://dx.doi.org/10.48550/arXiv.2510.23191)
- \[undefp\]Andrew Gelman and John Carlin“Beyond Power Calculations: Assessing Type S \(Sign\) and Type M \(Magnitude\) Errors”In*Perspectives on Psychological Science*9\.6, 2014, pp\. 641–651DOI:[10\.1177/1745691614551642](https://dx.doi.org/10.1177/1745691614551642)
- \[undefq\]Andrew Gelman and Eric Loken“The garden of forking paths: Why multiple comparisons can be a problem, even when there is no“fishing expedition”or“p\-hacking”and the research hypothesis was posited ahead of time” Unpublished manuscript, Columbia University, 2013URL:[https://sites\.stat\.columbia\.edu/gelman/research/unpublished/p\_hacking\.pdf](https://sites.stat.columbia.edu/gelman/research/unpublished/p_hacking.pdf)
- \[undefr\]Nelson Goodman“Fact, Fiction, and Forecast”Cambridge, MA: Harvard University Press, 1983
- \[undefs\]Dan Hendrycks et al\.“A Definition of AGI” v3, December 2025, 2025DOI:[10\.48550/arXiv\.2510\.18212](https://dx.doi.org/10.48550/arXiv.2510.18212)
- \[undeft\]David Ilić and Gilles E\. Gignac“Evidence of Interrelated Cognitive\-Like Capabilities in Large Language Models: Indications of Artificial General Intelligence or Achievement?”In*Intelligence*106, 2024, pp\. 101858DOI:[10\.1016/j\.intell\.2024\.101858](https://dx.doi.org/10.1016/j.intell.2024.101858)
- \[undefu\]Jana Jung, Marlene Lutz, Indira Sen and Markus Strohmaier“Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality”In*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*Association for Computational Linguistics, 2026, pp\. 8143–8173DOI:[10\.18653/v1/2026\.eacl\-long\.380](https://dx.doi.org/10.18653/v1/2026.eacl-long.380)
- \[undefv\]Michael T\. Kane“Validating the Interpretations and Uses of Test Scores”In*Journal of Educational Measurement*50\.1, 2013, pp\. 1–73DOI:[10\.1111/jedm\.12000](https://dx.doi.org/10.1111/jedm.12000)
- \[undefw\]Ralph L\. Keeney and Howard Raiffa“Decisions with Multiple Objectives: Preferences and Value Tradeoffs”Cambridge University Press, 1993DOI:[10\.1017/CBO9781139174084](https://dx.doi.org/10.1017/CBO9781139174084)
- \[undefx\]Muhammad Ali Khalidi“Natural Categories and Human Kinds: Classification in the Natural and Social Sciences”Cambridge: Cambridge University Press, 2013DOI:[10\.1017/CBO9780511998553](https://dx.doi.org/10.1017/CBO9780511998553)
- \[undefy\]Markus Langer, Kevin Baum and Nadine Schlicker“Effective Human Oversight of AI\-Based Systems: A Signal Detection Perspective on the Detection of Inaccurate and Unfair Outputs” Volume 35 \(2025\); Crossref records online\-first publication in 2024In*Minds and Machines*35\.1, 2025DOI:[10\.1007/s11023\-024\-09701\-0](https://dx.doi.org/10.1007/s11023-024-09701-0)
- \[undefz\]Yu Lu Liu et al\.“ECBD: Evidence\-Centered Benchmark Design for NLP”In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*Association for Computational Linguistics, 2024, pp\. 16349–16365DOI:[10\.18653/v1/2024\.acl\-long\.861](https://dx.doi.org/10.18653/v1/2024.acl-long.861)
- \[undefaa\]Varun Magesh et al\.“Hallucination\-Free? Assessing the Reliability of Leading AI Legal Research Tools”In*Journal of Empirical Legal Studies*22, 2025, pp\. 216–242DOI:[10\.1111/jels\.12413](https://dx.doi.org/10.1111/jels.12413)
- \[undefab\]Kevin S\. McGrew“CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research”In*Intelligence*37\.1, 2009, pp\. 1–10DOI:[10\.1016/j\.intell\.2008\.08\.004](https://dx.doi.org/10.1016/j.intell.2008.08.004)
- \[undefac\]William Meredith“Measurement Invariance, Factor Analysis and Factorial Invariance”In*Psychometrika*58\.4, 1993, pp\. 525–543DOI:[10\.1007/BF02294825](https://dx.doi.org/10.1007/BF02294825)
- \[undefad\]Samuel Messick“Validity of Psychological Assessment: Validation of Inferences from Persons’ Responses and Performances as Scientific Inquiry into Score Meaning”In*American Psychologist*50\.9, 1995, pp\. 741–749DOI:[10\.1037/0003\-066X\.50\.9\.741](https://dx.doi.org/10.1037/0003-066X.50.9.741)
- \[undefae\]Judea Pearl and Elias Bareinboim“External Validity: From Do\-Calculus to Transportability Across Populations”In*Statistical Science*29\.4, 2014, pp\. 579–595DOI:[10\.1214/14\-STS486](https://dx.doi.org/10.1214/14-STS486)
- \[undefaf\]Benjamin C\. Pierce“Types and Programming Languages”Cambridge, MA: MIT Press, 2002
- \[undefag\]Inioluwa Deborah Raji et al\.“AI and the Everything in the Whole Wide World Benchmark”In*Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks*1, 2021arXiv:[https://datasets\-benchmarks\-proceedings\.neurips\.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9\-Abstract\-round2\.html](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html)
- \[undefah\]R\. Rockafellar and Stanislav Uryasev“Conditional Value\-at\-Risk for General Loss Distributions”In*Journal of Banking & Finance*26\.7, 2002, pp\. 1443–1471DOI:[10\.1016/S0378\-4266\(02\)00271\-6](https://dx.doi.org/10.1016/S0378-4266(02)00271-6)
- \[undefai\]Olawale Salaudeen et al\.“Measurement to Meaning: A Validity\-Centered Framework for AI Evaluation” Submitted 13 May 2025In*arXiv preprint arXiv:2505\.10573*, 2025DOI:[10\.48550/arXiv\.2505\.10573](https://dx.doi.org/10.48550/arXiv.2505.10573)
- \[undefaj\]W\. Schneider and Kevin S\. McGrew“The Cattell\-Horn\-Carroll theory of cognitive abilities”In*Contemporary intellectual assessment: Theories, tests, and issues*The Guilford Press, 2018, pp\. 73–163
- \[undefak\]Yubo Wang et al\.“MMLU\-Pro: A More Robust and Challenging Multi\-Task Language Understanding Benchmark”In*Advances in Neural Information Processing Systems 37*, 2024, pp\. 95266–95290DOI:[10\.52202/079017\-3018](https://dx.doi.org/10.52202/079017-3018)
- \[undefal\]Yanzhe Zhang, Sanmi Koyejo and Diyi Yang“The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task\-Irrelevant Context” arXiv preprint, version 2, 2026DOI:[10\.48550/arXiv\.2607\.12963](https://dx.doi.org/10.48550/arXiv.2607.12963)

## References

- \[undefam\]Carlo Acerbi and Dirk Tasche“On the Coherence of Expected Shortfall”In*Journal of Banking & Finance*26\.7, 2002, pp\. 1487–1503DOI:[10\.1016/S0378\-4266\(02\)00283\-2](https://dx.doi.org/10.1016/S0378-4266(02)00283-2)
- \[undefan\]Dominik Bachmann et al\.“fl\-IRT\-ing with Psychometrics to Improve NLP Bias Measurement”In*Minds and Machines*34\.37, 2024DOI:[10\.1007/s11023\-024\-09695\-9](https://dx.doi.org/10.1007/s11023-024-09695-9)
- \[undefao\]Kristian Gonzalez Barman, Sascha Caron, Tom Claassen and Henk Regt“Towards a Benchmark for Scientific Understanding in Humans and Machines”In*Minds and Machines*34\.6, 2024DOI:[10\.1007/s11023\-024\-09657\-1](https://dx.doi.org/10.1007/s11023-024-09657-1)
- \[undefap\]Andrew M\. Bean et al\.“Measuring what Matters: Construct Validity in Large Language Model Benchmarks” NeurIPS 2025 Track on Datasets and BenchmarksIn*arXiv preprint arXiv:2511\.04703*, 2025DOI:[10\.48550/arXiv\.2511\.04703](https://dx.doi.org/10.48550/arXiv.2511.04703)
- \[undefaq\]Olivier Binette and Jerome P\. Reiter“Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework”, 2024arXiv:[https://arxiv\.org/abs/2406\.10366](https://arxiv.org/abs/2406.10366)
- \[undefar\]Denny Borsboom, Gideon J\. Mellenbergh and Jaap Heerden“The Concept of Validity”In*Psychological Review*111\.4, 2004, pp\. 1061–1071DOI:[10\.1037/0033\-295X\.111\.4\.1061](https://dx.doi.org/10.1037/0033-295X.111.4.1061)
- \[undefas\]Samuel R\. Bowman and George Dahl“What Will it Take to Fix Benchmarking in Natural Language Understanding?”In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, 2021, pp\. 4843–4855DOI:[10\.18653/v1/2021\.naacl\-main\.385](https://dx.doi.org/10.18653/v1/2021.naacl-main.385)
- \[undefat\]Richard Boyd“Realism, Anti\-Foundationalism and the Enthusiasm for Natural Kinds”In*Philosophical Studies*61\.1–2, 1991, pp\. 127–148DOI:[10\.1007/BF00385837](https://dx.doi.org/10.1007/BF00385837)
- \[undefau\]Robert L\. Brennan“Generalizability Theory”New York: Springer, 2001DOI:[10\.1007/978\-1\-4757\-3456\-0](https://dx.doi.org/10.1007/978-1-4757-3456-0)
- \[undefav\]Stefan Buijsman“Accuracy is not all you need\! The Reasons to Require AI Explainability”In*Minds and Machines*36\.14, 2026DOI:[10\.1007/s11023\-026\-09768\-x](https://dx.doi.org/10.1007/s11023-026-09768-x)
- \[undefaw\]John B\. Carroll“Human cognitive abilities: A survey of factor\-analytic studies”Cambridge University Press, 1993DOI:[10\.1017/CBO9780511571312](https://dx.doi.org/10.1017/CBO9780511571312)
- \[undefax\]Lee J\. Cronbach and Paul E\. Meehl“Construct Validity in Psychological Tests”In*Psychological Bulletin*52\.4, 1955, pp\. 281–302DOI:[10\.1037/h0040957](https://dx.doi.org/10.1037/h0040957)
- \[undefay\]Roel Dobbe and Anouk Wolters“Toward Sociotechnical AI: Mapping Vulnerabilities for Machine Learning in Context”In*Minds and Machines*34\.12, 2024DOI:[10\.1007/s11023\-024\-09668\-y](https://dx.doi.org/10.1007/s11023-024-09668-y)
- \[undefaz\]Susan E\. Embretson“Construct Validity: Construct Representation versus Nomothetic Span”In*Psychological Bulletin*93\.1, 1983, pp\. 179–197DOI:[10\.1037/0033\-2909\.93\.1\.179](https://dx.doi.org/10.1037/0033-2909.93.1.179)
- \[undefaaa\]Timo Freiesleben“Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks”In*arXiv preprint arXiv:2603\.15121*, 2026DOI:[10\.48550/arXiv\.2603\.15121](https://dx.doi.org/10.48550/arXiv.2603.15121)
- \[undefaab\]Timo Freiesleben and Sebastian Zezulka“The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models” Submitted 27 October 2025In*arXiv preprint arXiv:2510\.23191*, 2025DOI:[10\.48550/arXiv\.2510\.23191](https://dx.doi.org/10.48550/arXiv.2510.23191)
- \[undefaac\]Andrew Gelman and John Carlin“Beyond Power Calculations: Assessing Type S \(Sign\) and Type M \(Magnitude\) Errors”In*Perspectives on Psychological Science*9\.6, 2014, pp\. 641–651DOI:[10\.1177/1745691614551642](https://dx.doi.org/10.1177/1745691614551642)
- \[undefaad\]Andrew Gelman and Eric Loken“The garden of forking paths: Why multiple comparisons can be a problem, even when there is no“fishing expedition”or“p\-hacking”and the research hypothesis was posited ahead of time” Unpublished manuscript, Columbia University, 2013URL:[https://sites\.stat\.columbia\.edu/gelman/research/unpublished/p\_hacking\.pdf](https://sites.stat.columbia.edu/gelman/research/unpublished/p_hacking.pdf)
- \[undefaae\]Nelson Goodman“Fact, Fiction, and Forecast”Cambridge, MA: Harvard University Press, 1983
- \[undefaaf\]Dan Hendrycks et al\.“A Definition of AGI” v3, December 2025, 2025DOI:[10\.48550/arXiv\.2510\.18212](https://dx.doi.org/10.48550/arXiv.2510.18212)
- \[undefaag\]David Ilić and Gilles E\. Gignac“Evidence of Interrelated Cognitive\-Like Capabilities in Large Language Models: Indications of Artificial General Intelligence or Achievement?”In*Intelligence*106, 2024, pp\. 101858DOI:[10\.1016/j\.intell\.2024\.101858](https://dx.doi.org/10.1016/j.intell.2024.101858)
- \[undefaah\]Jana Jung, Marlene Lutz, Indira Sen and Markus Strohmaier“Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality”In*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*Association for Computational Linguistics, 2026, pp\. 8143–8173DOI:[10\.18653/v1/2026\.eacl\-long\.380](https://dx.doi.org/10.18653/v1/2026.eacl-long.380)
- \[undefaai\]Michael T\. Kane“Validating the Interpretations and Uses of Test Scores”In*Journal of Educational Measurement*50\.1, 2013, pp\. 1–73DOI:[10\.1111/jedm\.12000](https://dx.doi.org/10.1111/jedm.12000)
- \[undefaaj\]Ralph L\. Keeney and Howard Raiffa“Decisions with Multiple Objectives: Preferences and Value Tradeoffs”Cambridge University Press, 1993DOI:[10\.1017/CBO9781139174084](https://dx.doi.org/10.1017/CBO9781139174084)
- \[undefaak\]Muhammad Ali Khalidi“Natural Categories and Human Kinds: Classification in the Natural and Social Sciences”Cambridge: Cambridge University Press, 2013DOI:[10\.1017/CBO9780511998553](https://dx.doi.org/10.1017/CBO9780511998553)
- \[undefaal\]Markus Langer, Kevin Baum and Nadine Schlicker“Effective Human Oversight of AI\-Based Systems: A Signal Detection Perspective on the Detection of Inaccurate and Unfair Outputs” Volume 35 \(2025\); Crossref records online\-first publication in 2024In*Minds and Machines*35\.1, 2025DOI:[10\.1007/s11023\-024\-09701\-0](https://dx.doi.org/10.1007/s11023-024-09701-0)
- \[undefaam\]Yu Lu Liu et al\.“ECBD: Evidence\-Centered Benchmark Design for NLP”In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*Association for Computational Linguistics, 2024, pp\. 16349–16365DOI:[10\.18653/v1/2024\.acl\-long\.861](https://dx.doi.org/10.18653/v1/2024.acl-long.861)
- \[undefaan\]Varun Magesh et al\.“Hallucination\-Free? Assessing the Reliability of Leading AI Legal Research Tools”In*Journal of Empirical Legal Studies*22, 2025, pp\. 216–242DOI:[10\.1111/jels\.12413](https://dx.doi.org/10.1111/jels.12413)
- \[undefaao\]Kevin S\. McGrew“CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research”In*Intelligence*37\.1, 2009, pp\. 1–10DOI:[10\.1016/j\.intell\.2008\.08\.004](https://dx.doi.org/10.1016/j.intell.2008.08.004)
- \[undefaap\]William Meredith“Measurement Invariance, Factor Analysis and Factorial Invariance”In*Psychometrika*58\.4, 1993, pp\. 525–543DOI:[10\.1007/BF02294825](https://dx.doi.org/10.1007/BF02294825)
- \[undefaaq\]Samuel Messick“Validity of Psychological Assessment: Validation of Inferences from Persons’ Responses and Performances as Scientific Inquiry into Score Meaning”In*American Psychologist*50\.9, 1995, pp\. 741–749DOI:[10\.1037/0003\-066X\.50\.9\.741](https://dx.doi.org/10.1037/0003-066X.50.9.741)
- \[undefaar\]Judea Pearl and Elias Bareinboim“External Validity: From Do\-Calculus to Transportability Across Populations”In*Statistical Science*29\.4, 2014, pp\. 579–595DOI:[10\.1214/14\-STS486](https://dx.doi.org/10.1214/14-STS486)
- \[undefaas\]Benjamin C\. Pierce“Types and Programming Languages”Cambridge, MA: MIT Press, 2002
- \[undefaat\]Inioluwa Deborah Raji et al\.“AI and the Everything in the Whole Wide World Benchmark”In*Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks*1, 2021arXiv:[https://datasets\-benchmarks\-proceedings\.neurips\.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9\-Abstract\-round2\.html](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html)
- \[undefaau\]R\. Rockafellar and Stanislav Uryasev“Conditional Value\-at\-Risk for General Loss Distributions”In*Journal of Banking & Finance*26\.7, 2002, pp\. 1443–1471DOI:[10\.1016/S0378\-4266\(02\)00271\-6](https://dx.doi.org/10.1016/S0378-4266(02)00271-6)
- \[undefaav\]Olawale Salaudeen et al\.“Measurement to Meaning: A Validity\-Centered Framework for AI Evaluation” Submitted 13 May 2025In*arXiv preprint arXiv:2505\.10573*, 2025DOI:[10\.48550/arXiv\.2505\.10573](https://dx.doi.org/10.48550/arXiv.2505.10573)
- \[undefaaw\]W\. Schneider and Kevin S\. McGrew“The Cattell\-Horn\-Carroll theory of cognitive abilities”In*Contemporary intellectual assessment: Theories, tests, and issues*The Guilford Press, 2018, pp\. 73–163
- \[undefaax\]Yubo Wang et al\.“MMLU\-Pro: A More Robust and Challenging Multi\-Task Language Understanding Benchmark”In*Advances in Neural Information Processing Systems 37*, 2024, pp\. 95266–95290DOI:[10\.52202/079017\-3018](https://dx.doi.org/10.52202/079017-3018)
- \[undefaay\]Yanzhe Zhang, Sanmi Koyejo and Diyi Yang“The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task\-Irrelevant Context” arXiv preprint, version 2, 2026DOI:[10\.48550/arXiv\.2607\.12963](https://dx.doi.org/10.48550/arXiv.2607.12963)

Similar Articles

The Evaluation Trap: Benchmark Design as Theoretical Commitment

arXiv cs.AI

This paper identifies the 'evaluation trap' where AI benchmarks inadvertently stabilize dominant paradigms by narrowing what counts as progress, and introduces Epistematics, a meta-evaluative methodology to ensure evaluation criteria discriminate true capability from proxy behaviors.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders

arXiv cs.AI

This paper introduces Benchmarking-Cultures-25, a dataset analyzing how AI model builders selectively highlight benchmarks in press releases. It finds a fragmented evaluation landscape with limited cross-model comparability, arguing that benchmarks are used as narrative devices for market positioning rather than standardized scientific measurement.

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

arXiv cs.AI

The BEAMS Initiative presents a benchmark suite for evaluating AI tools in modeling and simulation, focusing on human-centered and responsible AI practices. Tests reveal variability across LLM-based engines, with better performance in qualitative tasks than causal reasoning.

Good Benchmarks

arXiv cs.AI

This paper from July 2026 defines principles for designing effective benchmark tasks for AI agents, drawing on experience from the Terminal Bench project. Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons, with an emphasis on real-world relevance and outcome-based verification.