Response drift across frontier large language models
Summary
A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.
View Cached Full Text
Cached at: 07/24/26, 05:16 AM
# Response drift across frontier large language models
Source: [https://arxiv.org/html/2607.20454](https://arxiv.org/html/2607.20454)
Mohammed AledhariDepartment of Data Science, University of North Texas, Denton, TX 76207, USAmohammed\.aledhari@unt\.eduAli AledhariDepartment of Data Science, University of North Texas, Denton, TX 76207, USAthese authors contributed equally to this workFatimah AledhariDepartment of Data Science, University of North Texas, Denton, TX 76207, USAthese authors contributed equally to this workGowtham Venkat EathamokkalaDepartment of Data Science, University of North Texas, Denton, TX 76207, USAthese authors contributed equally to this workMohamed RahoutiDepartment of Computer and Information Science, Fordham University, New York, NY 10458, USAthese authors contributed equally to this work
###### Abstract
All frontier large language models \(LLMs\) exhibit response drift—producing outputs that deviate from expert\-validated references—yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation\. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments\. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling \(78–81% deviation\), while two achieve lower deviation \(47–49%\)\. Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceedingr=0\.85r=0\.85\. Automated similarity metrics explain less than 2% of variance in human judgements\. These findings reveal that response drift is universal across frontier LLMs, domain\- and question\-dependent in structure, and accessible only through human\-centred evaluation\.
Despite rapid capability gains in frontier large language models \(LLMs\)\[[6](https://arxiv.org/html/2607.20454#bib.bib1),[7](https://arxiv.org/html/2607.20454#bib.bib2),[16](https://arxiv.org/html/2607.20454#bib.bib3)\], a fundamental question remains underexplored: how faithfully do these systems respond when assessed by human experts rather than automated benchmarks? Existing evaluations focus predominantly on narrow task domains—mathematical problem solving\[[15](https://arxiv.org/html/2607.20454#bib.bib4),[33](https://arxiv.org/html/2607.20454#bib.bib5)\], code generation\[[10](https://arxiv.org/html/2607.20454#bib.bib6)\], or conversational preference\[[41](https://arxiv.org/html/2607.20454#bib.bib7)\]—providing an incomplete picture of comprehensive performance\. Many established benchmarks now exhibit saturation, with multiple models exceeding 90% accuracy on datasets such as the Massive Multitask Language Understanding benchmark \(MMLU\)\[[15](https://arxiv.org/html/2607.20454#bib.bib4)\], limiting their ability to discriminate among frontier systems\[[40](https://arxiv.org/html/2607.20454#bib.bib8)\]; the recent Humanity’s Last Exam benchmark\[[32](https://arxiv.org/html/2607.20454#bib.bib31)\], designed to push the frontier of automated evaluation difficulty, further underscores this saturation challenge\. Training data contamination further confounds interpretation\[[1](https://arxiv.org/html/2607.20454#bib.bib9)\]\. Holistic evaluation frameworks such as the Holistic Evaluation of Language Models \(HELM\)\[[5](https://arxiv.org/html/2607.20454#bib.bib10)\], Chatbot Arena\[[8](https://arxiv.org/html/2607.20454#bib.bib11)\], AlpacaEval\[[11](https://arxiv.org/html/2607.20454#bib.bib12)\], MT\-Bench\[[41](https://arxiv.org/html/2607.20454#bib.bib7)\], Arena\-Hard\[[24](https://arxiv.org/html/2607.20454#bib.bib44)\], and WildBench\[[26](https://arxiv.org/html/2607.20454#bib.bib45)\]have advanced the field, and fine\-grained factual evaluation methods such as FActScore\[[28](https://arxiv.org/html/2607.20454#bib.bib13)\]have improved precision assessment for specific generation tasks\. Yet these approaches rely on automated metrics, LLM\-as\-judge paradigms, or pairwise preference designs—and while individual elements such as variance decomposition or absolute scoring exist in prior work, no existing study combines them in a fully crossed human evaluation that simultaneously enables per\-question correlation analysis and absolute fidelity measurement across models\.
A critical gap thus exists between the performance profiles reported on leaderboards and the fidelity of model outputs as perceived by domain\-informed human evaluators\. We term the systematic deviation of model responses from expert\-validated references*response drift*, and the disconnect between automated and human\-perceived quality the*fidelity gap*\(Fig\.[1](https://arxiv.org/html/2607.20454#S0.F1)a\): models that appear comparable on automated benchmarks may diverge substantially when their open\-ended responses are assessed against expert\-validated references by trained human judges\[[27](https://arxiv.org/html/2607.20454#bib.bib14),[37](https://arxiv.org/html/2607.20454#bib.bib15)\]\. Characterising response drift—its universality, magnitude, and dependence on model, domain, and question—requires evaluation designs that capture graduated quality differences in open\-ended generation, not merely binary correctness or relative preference, across diverse cognitive domains\.
Here we address this challenge through a fully crossed human evaluation at scale\. Forty\-seven geographically diverse participants each evaluated all 62 standardised questions across ten frontier large language models—spanning independent organisations and representing both proprietary and open\-weight systems—under blinded conditions, yielding 29,140 independent assessments \(Fig\.[1](https://arxiv.org/html/2607.20454#S0.F1)b; see Methods and Supplementary Table S1 for model details\)\. The 62 questions span six capability domains: reasoning and language understanding, mathematical problem solving, coding and software development, conversational ability, safety and ethical considerations, and domain\-specific professional knowledge\. All models were accessed through their default web\-based interfaces to capture realistic end\-user conditions; this design choice prioritises ecological validity but introduces deployment\-level confounds that we examine in the Discussion\. Each response was rated on a five\-point Likert scale against an expert\-validated reference answer, and we quantify response fidelity as the normalised deviation from perfect alignment \(Methods\)\. Critically, this metric captures evaluators’ perception of how faithfully a model response preserves the content of a*specific*reference; it does not assess absolute correctness, as open\-ended questions may admit multiple valid responses that diverge from any single reference\. High deviation therefore indicates divergence from one expert\-validated answer, not necessarily low quality \(see Discussion\)\.
Our central finding is that*all*evaluated frontier models exhibit substantial response drift—deviation from expert\-validated references—but the magnitude and structure of this drift vary markedly across models, questions, and domains\. The majority of models cluster within a narrow performance band \(78–81% deviation\) that is statistically indistinguishable, forming a convergent fidelity ceiling, while two models achieve substantially lower deviation \(47–49%\)\. Even these high\-fidelity models, however, deviate from expert references on nearly half of assessed dimensions\. Model identity accounts for over half of total variance in fidelity scores, confirming that model selection—rather than question difficulty or evaluator variability—is the primary determinant of drift magnitude\. The ceiling persists across all six domains, but drift profiles differ: models that excel in one domain may fail in another, and the question\-level tier gap ranges from 0 to 73 percentage points\. Automated natural language processing similarity metrics, computed independently, explain less than 2% of the variance in human fidelity judgements, underscoring that response drift is a phenomenon accessible only through human\-centred evaluation\.
\(a\)
\(b\)
Figure 1:The response fidelity problem and study design\.a, Illustration of the fidelity gap motivating this study\. Three users submit the identical question to the same large language model \(same version, seed, and temperature\), yet receive responses of varying quality when compared against a gold\-standard expert reference validated by≥\\geq2 domain experts\. High\-fidelity responses preserve the content, structure, and accuracy of the reference; partial\-fidelity responses omit key details; low\-fidelity responses contain factual errors or hallucinated content\. This variability is largely invisible to automated similarity metrics but clearly detected by human evaluators\.b, Fully crossed repeated\-measures study design\. 47 geographically diverse participants \(18 North America, 15 Europe, 12 Asia, 2 Other\) from 7 professional domains each evaluated all 62 questions across all 10 frontier LLMs under blinded conditions, yielding47×10×62=29,14047\\times 10\\times 62=29\{,\}140independent evaluations with 0% attrition\. Responses were rated on a 5\-point Likert scale against expert\-validated references\. Fidelity deviation =\(5−Ratingijk\)/4\(5\-\\mathrm\{Rating\}\_\{ijk\}\)/4, where 0% indicates perfect fidelity and 100% maximum deviation\. The 62 questions span six capability domains \(10–12 items each\), with presentation order randomised via Latin\-square design and rater calibration achieving intraclass correlation coefficient \(ICC\(2,1\)\) = 0\.84\.## Results
### All frontier models exhibit response drift, but magnitude varies markedly
Every evaluated frontier model deviates substantially from expert\-validated reference answers, confirming that response drift is a universal phenomenon across the current generation of LLMs\. However, drift magnitude varies markedly across models, revealing a bimodal structure\. Across 29,140 evaluations, response fidelity deviation—the normalised distance from expert\-validated reference answers as judged by human evaluators \(Methods\)—varied from 47\.0% \(Claude\) to 80\.5% \(ChatGPT, DeepSeek, Qwen\), a 33\.5 percentage\-point \(pp\) range \(Cohen’sd=2\.45d=2\.45, a standardised effect size measuring the separation between groups in pooled standard deviation units\[[9](https://arxiv.org/html/2607.20454#bib.bib22)\]; Fig\.[2](https://arxiv.org/html/2607.20454#Sx1.F2)a,b; Table[1](https://arxiv.org/html/2607.20454#Sx1.T1)\)\. However, this range masks a markedly bimodal distribution\. Two models—Claude \(47\.0%, 95% bootstrap confidence interval \(CI\) \[42\.7, 51\.4\]\) and Gemini \(49\.4%, CI \[42\.9, 55\.8\]\)—form a high\-fidelity tier with overlapping confidence intervals and negligible mutual effect size \(d=0\.11d=0\.11\)\. The remaining eight models cluster in a 2\.9 pp band between 77\.6% \(Llama\) and 80\.5% \(ChatGPT, DeepSeek, Qwen\)\. These eight ceiling models are statistically indistinguishable \(Kruskal–WallisH=9\.18H=9\.18,p=0\.24p=0\.24, a non\-parametric test for differences across groups; one\-way analysis of variance \(ANOVA\)F\(7,488\)=1\.26F\(7,488\)=1\.26,p=0\.27p=0\.27; Supplementary Table S6\), whereas the contrast between the two tiers is large \(Cohen’sd=1\.89d=1\.89; Supplementary Table S7\)\. To provide positive evidence of equivalence rather than relying on non\-significance, we applied the two one\-sided tests \(TOST\) procedure\[[22](https://arxiv.org/html/2607.20454#bib.bib16)\]\(Equation[5](https://arxiv.org/html/2607.20454#Sx3.E5)\) with an equivalence bound ofΔ=5\\Delta=5pp \(see Methods\)\. All 28 pairwise comparisons among ceiling models fell within the equivalence region \(p<0\.05p<0\.05for both one\-sided tests in every pair\), confirming that within\-ceiling differences are smaller than the pre\-specified smallest effect size of interest\. Under a stricter bound \(Δ=3\\Delta=3pp\), 24 of 28 pairs \(85\.7%\) still achieved equivalence; the four non\-equivalent pairs involved the most extreme ceiling models \(Llama at 77\.6% vs\. ChatGPT, DeepSeek, or Qwen at 80\.5%\)\. By contrast, the high\-fidelity\-versus\-ceiling contrast was large\.
\(a\)
\(b\)
Figure 2:Fidelity deviation across domains and overall model ranking\.a, Domain×\\timesmodel fidelity deviation showing deviation percentages across all six capability domains for each of the 10 evaluated models \(lower values = better performance\)\. Cells are shaded on a continuous green\-to\-red scale, where green indicates low deviation \(high fidelity\) and red indicates high deviation \(low fidelity\); bold values indicate the best\-performing model in each domain; the rightmost column shows overall deviation\. A yellow border separates the two high\-fidelity models \(Claude 47\.0%, Gemini 49\.4%\) from the eight ceiling models \(77\.6–80\.5%\)\. Domain means along the bottom row reveal that safety shows the lowest average deviation \(68\.8%\) and domain knowledge the highest \(79\.6%\)\.n=47n=47evaluators per cell; 620 model–question combinations\.b, Overall fidelity deviation across all six benchmarks, with models ordered from best \(lowest deviation\) to worst\. Horizontal bars are colour\-coded by performance tier: green = Good \(<<60%\), orange = Moderate \(60–80%\), red = Poor \(\>\>80%\)\. Dashed vertical reference lines at 40%, 60%, and 80% mark tier boundaries\. Claude \(\#1, 47\.0%\) and Gemini \(\#2, 49\.4%\) are clearly separated from the remaining eight models, which form a tight cluster between 77\.6% and 80\.5%\. Values are question\-weighted means acrossn=62n=62questions; 95% bootstrap CIs are reported in Table[1](https://arxiv.org/html/2607.20454#Sx1.T1)\.Which model a user selects matters more than the specific question asked or who evaluates the response\. A two\-way ANOVA on the aggregated 620\-cell matrix attributed 52\.4% of total fidelity variance to model identity \(η2=0\.524\\eta^\{2\}=0\.524\), 16\.5% to question difficulty, and 31\.1% to the residual \(Table[2](https://arxiv.org/html/2607.20454#Sx1.T2), \(A\)\)\. A linear mixed\-effects model\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\]fitted to the full participant\-level data \(29,140 observations; Equation[4](https://arxiv.org/html/2607.20454#Sx3.E4)\) yielded a consistent but more granular decomposition: the variance proportion for model was 0\.41, the ICC for question 0\.14, and for participant 0\.09, with 0\.36 attributable to observation\-level residual \(Table[2](https://arxiv.org/html/2607.20454#Sx1.T2), \(B\)\)\. The conditionalR2R^\{2\}was 0\.64 and the marginalR2R^\{2\}was 0\.41\[[29](https://arxiv.org/html/2607.20454#bib.bib18)\]\. Both analyses confirm model selection as the primary determinant of fidelity\. Domain\-level variance decompositions, showing modelη2\\eta^\{2\}ranging from 0\.464 \(coding\) to 0\.807 \(reasoning\), are reported in Supplementary Table S5; the complete 62\-question×\\times10\-model deviation matrix appears in Supplementary Table S4\.
To contextualise the magnitude of the ceiling, we note that under uniform random rating \(1–5\), the expected deviation would be 50%\. The eight ceiling models at 78–81% deviation thus fall well below this baseline, corresponding to implied mean Likert ratings of 1\.8–1\.9 \(between “completely unfaithful” and “mostly unfaithful”\)\. The two high\-fidelity models achieved implied ratings of 3\.0–3\.1 \(“partially faithful”\)\. This comparison is illustrative rather than inferential, as random rating and systematic evaluation involve different cognitive processes; moreover, deviation from a single reference does not necessarily imply factual incorrectness, because models may produce valid responses that diverge from the specific expert reference \(see Discussion\)\.
Bootstrap rank analysis \(10,000 resamples\) confirmed rank stability at the extremes: Claude and Gemini occupied ranks 1–2 in 100% of resamples \(Supplementary Table S8\)\. Within the ceiling, ranks were unstable, consistent with confirmed equivalence\. Across 62 items, Claude ranked first on 29 questions and Gemini on 31, with the eight ceiling models combined ranking first on only 2\.
Table 1:Response fidelity deviation \(%\) by model and domain\.Mean fidelity deviation \(lower==better\) across all tasks in each domain\. Models sorted by overall deviation; visual spacing separates the two high\-fidelity models from the eight ceiling models\. Bold==best per column\. 95% bootstrap CIs \(10,000 resamples\) in brackets; note that CI widths are wider for high\-fidelity models \(reflecting greater question\-to\-question variability\) than for ceiling models\.n=47n=47evaluators per cell; 620 total model–question combinations\. Domain\-level 95% bootstrap CIs are reported in Supplementary Table S5\.ModelReason\.MathCodingConv\.SafetyDom\. K\.Overall\[95% CI\]Claude40\.145\.549\.555\.146\.145\.947\.0\[42\.7, 51\.4\]Gemini26\.345\.150\.637\.547\.482\.749\.4\[42\.9, 55\.8\]Llama80\.475\.979\.575\.671\.881\.777\.6\[74\.1, 81\.0\]Mistral80\.676\.679\.175\.574\.582\.878\.3\[75\.2, 81\.4\]Grok80\.778\.580\.578\.073\.382\.779\.1\[76\.1, 82\.0\]Perplexity81\.978\.881\.677\.574\.882\.479\.6\[76\.7, 82\.4\]Copilot81\.779\.181\.379\.573\.084\.179\.9\[77\.0, 82\.7\]DeepSeek80\.279\.381\.380\.476\.184\.980\.5\[77\.8, 83\.1\]Qwen81\.079\.382\.078\.975\.985\.080\.5\[77\.7, 83\.2\]ChatGPT82\.979\.680\.679\.175\.284\.080\.4\[77\.5, 83\.2\]Table 2:Variance decomposition of response fidelity\.\(A\)Two\-way ANOVA on the aggregated 620\-cell matrix \(descriptive; does not account for participant\-level variance\)\. SS==sum of squares; df==degrees of freedom; MS==mean square;FF==FF\-statistic;pp==pp\-value;η2\\eta^\{2\}==eta\-squared \(proportion of total variance\);ω2\\omega^\{2\}==omega\-squared \(bias\-corrected effect size\)\.\(B\)Variance proportions from the linear mixed\-effects model \(Equation[4](https://arxiv.org/html/2607.20454#Sx3.E4); 29,140 observations\)\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\]\. For random effects \(question, participant, residual\), values are intraclass correlation coefficients \(ICCs\); for model \(a fixed effect\), the value is the proportion of total variance attributable to fixed effects\.R2R^\{2\}\(coefficient of determination, proportion of variance explained\) per Nakagawa and Schielzeth\[[29](https://arxiv.org/html/2607.20454#bib.bib18)\]\.\(A\) Aggregated two\-way ANOVA
\(B\) Mixed\-effects variance proportions \(n=29,140n=29\{,\}140\)
SourceVar\. prop\.InterpretationModel \(fixed\)0\.41Primary determinant†Question \(random\)0\.14Question difficultyParticipant \(random\)0\.09Evaluator tendencyResidual0\.36Observation noiseMarginalR2R^\{2\}0\.41Fixed effects onlyConditionalR2R^\{2\}0\.64Fixed \+ random†Proportion of variance attributable tofixed effects; not a true ICC\.
### Drift profiles are domain\- and question\-dependent
Drift is not uniform across capability domains: each model exhibits a distinct domain\-specific profile, and models that drift least in one domain may drift most in another\. Performance varied substantially across domains, with the best\-to\-worst gap ranging from 30\.0 pp \(safety\) to 56\.6 pp \(reasoning; Fig\.[2](https://arxiv.org/html/2607.20454#Sx1.F2)a; Table[1](https://arxiv.org/html/2607.20454#Sx1.T1)\)\. The two high\-fidelity models demonstrated complementary excellence\. Gemini led in reasoning \(26\.3%; Supplementary Fig\. S1 a\), mathematics \(45\.1%\), and conversational ability \(37\.5%\), while Claude led in coding \(49\.5%\), safety \(46\.1%; Fig\.[3](https://arxiv.org/html/2607.20454#Sx1.F3)b\), and domain\-specific knowledge \(45\.9%; Fig\.[3](https://arxiv.org/html/2607.20454#Sx1.F3)a\)\.
\(a\)
\(b\)
Figure 3:Per\-question fidelity profiles for domain knowledge and safety\.Each panel displays per\-question fidelity deviation for all 10 models within a single domain, arranged from best\-performing \(top\-left\) to worst\-performing \(bottom\-right\) model\. Red triangles mark the hardest question and green stars the easiest for each model; dashed lines indicate domain means\.a, Domain\-specific knowledge \(Q51–Q62, 12 items\): Claude leads at 45\.8% with a wide range \(66\.2%\), indicating high sensitivity to question content; all other models exceed 81% with narrow ranges \(<<15 pp\), reflecting uniformly poor fidelity\.b, Safety and ethical considerations \(Q41–Q50\): Claude \(46\.1%\) and Gemini \(47\.4%\) achieve the lowest deviation; this domain exhibits the narrowest ceiling\-model range \(71\.8–76\.1%\), suggesting safety\-related prompts are uniformly challenging\.n=47n=47evaluators per bar; dashed lines==domain means acrossn=47n=47evaluators\. Between\-tier significance: Welch’stt\-test with Holm–Bonferroni correction, all high\-fidelity\-vs\-ceiling contrastsp<0\.001p<0\.001\. Additional domain views for coding, reasoning, mathematics, and conversational ability appear in Supplementary Figs\. S4, S1, S2, and S3\.The high\-fidelity\-versus\-ceiling gap remained robust in all six domains\. Reasoning showed the widest separation \(48\.0 pp; 95% CI \[37\.1, 59\.0\]\), followed by mathematics \(33\.1 pp\), conversational ability \(31\.9 pp\), and coding \(30\.7 pp\)\. Domain knowledge showed the smallest gap \(19\.2 pp\), driven by Gemini’s anomalous 82\.7% in that domain despite 49\.4% overall—a 56\.4 pp within\-model range\. Claude was more consistent \(15\.0 pp range\)\. Among ceiling models, no model achieved sub\-70% deviation in any domain, and all eight exceeded 70% on 48 of 62 questions \(77%\)\. The domain knowledge and safety model\-centric profiles are presented in Fig\.[3](https://arxiv.org/html/2607.20454#Sx1.F3)a,b; the coding, mathematics, reasoning, and conversational views appear in Supplementary Figs\. S2 b, S1 b, S1 a, and S2 a\.
Having established that the fidelity ceiling persists across all six domains, we next examined whether ceiling models share the same question\-level limitation patterns or exhibit independent failures\.
### Ceiling models share systematic drift patterns across questions
Pairwise Pearson correlation coefficients \(rr, measuring linear association between−1\-1and\+1\+1\) across all 62 tasks revealed that ceiling models were strongly intercorrelated \(meanr=0\.90r=0\.90, range 0\.85–0\.95; Fig\.[4](https://arxiv.org/html/2607.20454#Sx1.F4)a,b; Supplementary Fig\. S3\), indicating shared limitation patterns\. To rule out the possibility that high correlations simply reflect shared question difficulty \(i\.e\., all models scoring high on hard questions\), we computed partial correlations controlling for question\-level mean deviation\. Partial correlations among ceiling models remained high \(mean partialr=0\.84r=0\.84, range 0\.78–0\.92\), confirming that shared patterns extend beyond question difficulty effects and reflect model\-specific behavioural similarities\. Claude and Gemini showed near\-zero correlation with each other \(r=0\.12r=0\.12, partialr=0\.09r=0\.09\) and with the ceiling cluster \(r<0\.20r<0\.20\), indicating independent error patterns\.
\(a\)
\(b\)
Figure 4:Bimodal performance structure and shared limitation patterns across frontier models\.a, Fidelity deviation ranked across all 62 tasks\. Models ranked from best to worst by mean fidelity deviation \(±1\\pm 1standard deviation \(SD\) error bars computed acrossn=62n=62per\-question means; blue diamonds = means\)\. Three tiers emerge: Good \(<<60%, green\), Moderate \(60–80%, orange\), and Poor \(\>\>80%, red\)\. Claude \(\#1,47\.0%±17\.647\.0\\%\\pm 17\.6\) and Gemini \(\#2,49\.4%±26\.049\.4\\%\\pm 26\.0\) are the only models classified as Good, with a 28 pp gap separating them from Llama \(\#3,77\.6%±8\.077\.6\\%\\pm 8\.0\)\. The wide error bars for the high\-fidelity models reflect greater question\-to\-question variability, indicating selective strength rather than uniform performance\.b, Pairwise Pearson correlations across all 62 question\-level deviation scores \(n=62n=62per model pair\)\. The dashed blue box highlights the high\-fidelity pair \(Claude–Gemini,r=0\.12r=0\.12\), which shows near\-zero correlation indicating independent error patterns\. The dashed red box delineates the ceiling cluster, where all 28 pairwise correlations exceedr=0\.85r=0\.85\(meanr=0\.90r=0\.90\), confirming that ceiling models share systematic patterns of limitation despite independent development\. Cross\-cluster correlations \(high\-fidelity vs\. ceiling\) are uniformly low \(r<0\.20r<0\.20\)\. Question\-level difficulty rankings and distributional analyses appear in Supplementary Figs\. S3 and S4–S6\.Question difficulty ranged from 49\.2% \(Q43, safety\) to 89\.2% \(Q60, domain knowledge; Supplementary Fig\. S3\)\. The five most discriminating questions \(by SD\) were Q7 \(30\.9 pp\), Q6 \(27\.3 pp\), Q10 \(27\.1 pp\), Q42 \(26\.9 pp\), and Q27 \(26\.3 pp\)—all with high\-fidelity models below 30% and ceiling models above 80%\. The complete 62\-question×\\times10\-model matrix is in Supplementary Table S4\.
Intra\-model response consistency\[[36](https://arxiv.org/html/2607.20454#bib.bib20)\]showed Claude and Gemini most consistent \(0\.846, 0\.845\), Mistral least \(0\.723\)\. DeepSeek achieved high consistency \(0\.825\) despite ceiling\-level deviation, indicating determinism alone does not guarantee fidelity \(Supplementary Table S9\)\.
### Automated metrics fail to capture human\-perceived fidelity differences
Automated natural language processing \(NLP\) similarity metrics showed weak, paradoxically positive correlations with fidelity deviation \(r=0\.13r=0\.13,r2=0\.018r^\{2\}=0\.018; Supplementary Table S10\)\. The 33\.5 pp human\-judged gap between Claude and ChatGPT compressed to 1\.2 pp in NLP similarity—a 28\-fold reduction\. These metrics explained less than 2% of fidelity variance\. A detailed comparison of the two independent measurement pipelines is provided in Supplementary Table S13\. Domain\-level NLP similarity scores are reported in Supplementary Table S11\.
Copilot and ChatGPT share GPT\-5\.2 but differ in deployment\. Their deviations \(79\.9% vs\. 80\.4%\) were virtually indistinguishable, suggesting deployment\-level effects are small within the ceiling\. Whether deployment configuration contributes to the between\-tier gap cannot be determined from these data alone \(see Discussion\)\.
As a robustness check, we re\-computed fidelity deviation after excluding model refusals \(82 of 29,140 evaluations; 0\.28%\), which were concentrated in the safety domain and scored as maximum deviation in the primary analysis \(see Methods\)\. Excluding refusals lowered Claude’s safety\-domain deviation from 46\.1% to 43\.8% and Gemini’s from 47\.4% to 45\.1%,*widening*the between\-tier gap by approximately 3\.6 pp\. Overall model rankings were unchanged\. Because the two high\-fidelity models refused most frequently \(Claude: 31; Gemini: 24\), the conservative maximum\-deviation penalty works against these models; the primary analysis therefore understates, rather than overstates, their relative advantage\. Distributional analyses by domain appear in Supplementary Figs\. S4–S6\.
### Construct validity: the human–automated gap is fundamental
A central methodological question is whether fidelity deviation captures genuine content quality or merely stylistic conformity to the reference format\. We addressed this through a battery of machine learning and statistical analyses designed to decompose the human–automated gap \(see Methods; full results in Supplementary Tables S14 and S15\)\.
First, we tested whether any combination of automated NLP features \(semantic similarity, lexical overlap, length ratios, part\-of\-speech alignment, sentiment agreement\) could predict human fidelity judgements\. Cross\-validated regression models—including linear regression, random forest, gradient\-boosted trees \(XGBoost\), and multilayer perceptrons \(MLPs\)—all yielded*negative*R2R^\{2\}values \(linearR2=−0\.07R^\{2\}=\-0\.07; XGBoostR2=−0\.30R^\{2\}=\-0\.30; MLPR2=−0\.54R^\{2\}=\-0\.54; all 5\-fold cross\-validated\), indicating that automated NLP features have no predictive power over human fidelity deviation—performing worse than a constant\-mean baseline\. This result held for content features alone \(semantic similarity;R2=−0\.05R^\{2\}=\-0\.05\), style features alone \(length and part\-of\-speech;R2=−0\.06R^\{2\}=\-0\.06\), and all features combined\.
Second, mediation analysis revealed that semantic similarity mediates only 1\.2% of the total tier effect on fidelity deviation \(bootstrap 95% CI \[−0\.007\-0\.007,−0\.001\-0\.001\];p<0\.05p<0\.05\); the remaining 98\.8% is a direct effect of model identity that is not transmitted through any NLP\-measurable feature\. Content mediation was 3\.8×\\timeslarger than style mediation, but both were negligible relative to the direct path \(Supplementary Note 2\)\.
Third, unsupervisedkk\-means clustering on model\-level NLP feature vectors \(7 features×\\times10 models\) failed to recover the human\-identified tier structure \(adjusted Rand index=−0\.05=\-0\.05; accuracy=6/10=6/10\), confirming that the bimodal distribution is invisible in the automated feature space\.
Fourth, evaluators showed 1\.8×\\timeslower variance \(standard deviation of semantic similarity ratings\) when assessing high\-fidelity models compared with ceiling models \(p<10−37p<10^\{\-37\}, Mann–WhitneyUUtest; Supplementary Table S16\), indicating stronger inter\-evaluator consensus for high\-quality responses\. The question\-level tier gap ranged from 0 to 73 percentage points \(coefficient of variation=55\.6%=55\.6\\%\), ruling out uniform reference\-style bias as an explanation: identical neutral\-style references produced dramatically different gaps depending on question content\. Evaluator profiling viakk\-means clustering identified three natural groups: 43 typical evaluators, 3 low\-alignment evaluators, and 1 extreme outlier \(Evaluator 3, least aligned in 68\.4% of cells\), confirming that the main evaluator body is internally consistent\.
Taken together, these analyses establish that human fidelity evaluation captures a quality dimension that is fundamentally inaccessible to the automated NLP pipeline—neither through content features, style features, linear models, nor deep neural networks\. The tier structure identified by human evaluators exists in a perceptual space orthogonal to the NLP feature space\. Additional analyses, including anomaly detection \(Supplementary Note 3\) and leave\-one\-domain\-out transfer prediction \(Supplementary Note 4\), further support these conclusions\.
## Discussion
Our fully crossed evaluation of 29,140 human assessments establishes that response drift—deviation from expert\-validated reference answers—is a universal property of all ten evaluated frontier LLMs\. No model achieved perfect or near\-perfect fidelity; even the two highest\-performing models deviated from expert references on approximately half of assessed dimensions \(47\.0% and 49\.4%\)\. The remaining eight models converge on a fidelity ceiling at 78–81% deviation, forming a statistically indistinguishable cluster \(Table[1](https://arxiv.org/html/2607.20454#Sx1.T1); confirmed by formal equivalence testing\[[22](https://arxiv.org/html/2607.20454#bib.bib16)\]; between\-tierd=1\.89d=1\.89\)\. Critically, drift is not uniform: it varies by model, domain, and individual question, with the question\-level tier gap ranging from 0 to 73 pp\. The convergence of eight independently developed models—with pairwise correlations ofr=0\.85r=0\.85–0\.950\.95\(Supplementary Fig\. S3\)—is consistent with multiple interpretations\. It may reflect shared training paradigms: most frontier models employ similar pretraining corpora and reinforcement learning from human feedback \(RLHF\) procedures\[[31](https://arxiv.org/html/2607.20454#bib.bib35),[2](https://arxiv.org/html/2607.20454#bib.bib48)\], potentially producing convergent limitations\. This shared\-training interpretation is consistent with the literature on shared failure modes\[[19](https://arxiv.org/html/2607.20454#bib.bib33),[35](https://arxiv.org/html/2607.20454#bib.bib47)\]\. Alternatively, the ceiling may indicate a methodological boundary\[[40](https://arxiv.org/html/2607.20454#bib.bib8)\]\. A third possibility is genuine diminishing returns; the “densing law” of Xiao et al\.\[[39](https://arxiv.org/html/2607.20454#bib.bib27)\]provides a theoretical parallel\. Distinguishing among these hypotheses requires follow\-up studies with questions designed to differentiate ceiling models, process\-level analysis, and longitudinal tracking\.
The divergence between our findings and preference\-based platforms is itself an important result\. Chatbot Arena\[[8](https://arxiv.org/html/2607.20454#bib.bib11)\]does not show the separation our fidelity evaluation reveals\. We argue this reflects a fundamental difference in what the paradigms measure\. Preference\-based evaluation asks “which do you prefer?”—a relative judgement influenced by fluency and formatting\. Reference\-based fidelity asks “how faithfully does this preserve expert content?”—an absolute judgement anchored to domain knowledge\. These are distinct constructs\. The 28\-fold compression between human fidelity and automated NLP metrics reinforces this: automated metrics capture surface properties, not the accuracy and completeness that reference\-based evaluation assesses\[[11](https://arxiv.org/html/2607.20454#bib.bib12),[28](https://arxiv.org/html/2607.20454#bib.bib13)\]\. This paradigm divergence is consistent with Goodhart’s law—the principle that optimising for a proxy metric can degrade performance on the underlying objective\[[14](https://arxiv.org/html/2607.20454#bib.bib46)\]\. The implication is that evaluation paradigm selection shapes rankings: organisations should incorporate reference\-based fidelity assessment for knowledge\-intensive tasks\. This paradigm divergence warrants systematic investigation; future work should compare both paradigms on identical question sets\.
The validity of expert reference answers is an important methodological consideration that warrants detailed discussion, as the entire evaluation is anchored to these references\. Open\-ended questions admit multiple valid responses, and a model producing a correct, complete answer using a different structure, emphasis, or level of detail than the reference could receive high deviation despite being factually accurate\. Our fidelity metric therefore measures alignment with one expert\-validated answer, not absolute response quality\. To the extent that the observed ceiling reflects diverse valid response strategies rather than poor quality, the metric would overestimate the performance gap between tiers\. We implemented several mitigations: \(1\) reference answers were developed through multi\-expert consensus \(minimum three domain experts per question; see Methods\); \(2\) all reference authors were independent of model developers and were blinded to model identities throughout; \(3\) references were written in a neutral expository style specifically avoiding model\-characteristic formatting \(e\.g\., bullet\-point lists, markdown headers\); \(4\) calibration training achieved ICC\(2,1\)=0\.84=0\.84\[[36](https://arxiv.org/html/2607.20454#bib.bib20),[20](https://arxiv.org/html/2607.20454#bib.bib26)\]; and \(5\) the fully crossed design ensures that any residual style bias affects all models equally per question\. Crucially, the construct validity analyses presented above provide direct empirical evidence against a stylistic\-conformity interpretation: the question\-level tier gap varies from 0 to 73 pp across questions that use identically styled references, evaluators show significantly higher consensus for high\-fidelity models, and no automated feature—content or style—can predict human fidelity judgements \(all cross\-validatedR2<0R^\{2\}<0\)\. These findings indicate that the tier structure reflects a perceptual quality dimension that is fundamentally inaccessible to automated NLP metrics\. All references are deposited in the repository for independent assessment\. Future work should employ multi\-reference scoring, in which evaluators rate fidelity against several valid references per question, to disentangle content fidelity from stylistic alignment\. Our rankings partially converge with independent frameworks\[[32](https://arxiv.org/html/2607.20454#bib.bib31),[5](https://arxiv.org/html/2607.20454#bib.bib10)\], though the separation magnitude is larger, consistent with paradigm divergence\.
The complementary strengths of Claude and Gemini \(r=0\.12r=0\.12\) have potential deployment implications\. Although domain\-specific routing could in principle reduce combined deviation, we emphasise that no ensemble experiment was conducted in this study, and the practical feasibility, latency costs, and error\-propagation risks of such routing remain to be validated in dedicated follow\-up work\.
This complementarity does not extend to the ceiling \(r\>0\.85r\>0\.85\)\. Gemini’s domain\-knowledge anomaly \(82\.7% despite 49\.4% overall\) shows aggregate rankings obscure domain deficiencies\[[5](https://arxiv.org/html/2607.20454#bib.bib10),[32](https://arxiv.org/html/2607.20454#bib.bib31)\]\. The safety domain—most shaped by alignment fine\-tuning—shows the narrowest ceiling range \(71\.8–76\.1%\), suggesting alignment compresses variation without proportionally improving substantive quality\[[30](https://arxiv.org/html/2607.20454#bib.bib32),[4](https://arxiv.org/html/2607.20454#bib.bib34)\]\.
Several limitations merit disclosure\. First, the five\-point Likert scale\[[25](https://arxiv.org/html/2607.20454#bib.bib23)\]may compress extreme variation\. Second, 47 participants provide power for between\-tier contrasts but may be insufficient for within\-ceiling resolution\. Third, evaluator heterogeneity—one participant showed the lowest alignment with expert references in 68\.4% of model–question combinations—may influence aggregates; the mixed\-effects model \(Table[2](https://arxiv.org/html/2607.20454#Sx1.T2)\) accounts for this, and sensitivity analyses are in Supplementary Note 1 and Supplementary Table S3\. Fourth, refusals \(rated 1\) may penalise safety\-conscious models; the refusal\-excluded analysis presented above mitigates but does not fully resolve this concern\. Fifth, web\-interface evaluation introduces deployment confounds; the Copilot–ChatGPT experiment suggests within\-ceiling effects are small\. Sixth, 62 questions cannot represent all use cases\. Seventh, results are a temporal snapshot\[[30](https://arxiv.org/html/2607.20454#bib.bib32)\]\. Eighth, because fidelity is measured against a single reference, the metric may partly capture stylistic conformity rather than content quality; models that produce valid answers using different structures or emphases than the reference may receive high deviation despite being factually correct\. Multi\-reference scoring should be explored to disentangle these factors\. Finally, the construct of “fidelity to references” differs from absolute correctness; establishing the relationship between reference deviation and real\-world task performance is an important direction for future research\.
In summary, this study establishes that response drift from expert references is universal across frontier LLMs: all ten models deviate substantially, with eight converging on a shared ceiling and two achieving lower but still considerable drift\. Drift is not monolithic — it varies by model, domain, and question, with complementary strength profiles across the highest\-performing models\. The divergence between fidelity\-based and preference\-based rankings reveals that evaluation paradigm selection fundamentally shapes model comparisons\. The construct validity analyses demonstrate that this drift is perceptible to human evaluators but fundamentally inaccessible to automated NLP metrics, underscoring the necessity of human\-centred assessment\. Future directions include longitudinal drift tracking, API\-level evaluation, expanded question sets, multi\-reference scoring, and systematic paradigm comparison\.
## Methods
### Study design overview
This study employed a fully crossed repeated\-measures design in which every participant evaluated every question for every model, yielding a complete three\-dimensional data structure: 47 participants×\\times10 models×\\times62 questions=29,140=29\{,\}140evaluations\. The design enables simultaneous estimation of participant, model, and question effects without confounding\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\]\. The study was approved by the Institutional Review Board at University of North Texas \(protocol \#IRB\-25\-297\)\. All participants provided written informed consent\.
### Participants
Forty\-seven participants were recruited through professional academic and industry networks beginning in June 2025\. The study proceeded in four phases: \(1\) study design, IRB approval, question development, expert reference\-answer creation and validation \(≥\\geq2 domain experts per answer\), and participant recruitment \(June–August 2025\); \(2\) pilot testing with early\-release models and monitoring of the rapidly evolving frontier model landscape to finalise the evaluation set \(September–November 2025\); \(3\) main evaluation period, during which all 10 models were available and 47 participants each completed all 620 evaluations \(December 2025–February 2026; median completion 30 days, range 15–60 days,∼\\sim60 hours per participant\); and \(4\) data analysis and manuscript preparation \(late February–March 2026\)\. Inclusion criteria required: \(1\) fluency in English \(self\-reported\), \(2\) prior experience interacting with at least two commercial LLM platforms, \(3\) a professional or academic background in a STEM \(science, technology, engineering, or mathematics\) or related professional field, and \(4\) age≥\\geq18 years\. Participants volunteered without financial compensation\.
Participants were geographically distributed across four regions: North America \(n=18n=18; 38\.3%\), Europe \(n=15n=15; 31\.9%\), Asia \(n=12n=12; 25\.5%\), and the Middle East \(n=2n=2; 4\.3%\)\. Professional backgrounds included software engineering \(n=11n=11\), data science \(n=8n=8\), education \(n=6n=6\), healthcare \(n=6n=6\), legal \(n=5n=5\), finance \(n=4n=4\), and general professional \(n=7n=7\)\. This distribution ensured domain\-appropriate expertise for evaluating responses across the six capability domains\. Self\-reported LLM experience ranged from regular use \(20–25 times per week;n=11n=11\) to daily professional use \(40–65 times per week;n=36n=36; 76\.6%\)\. Median age was 29 years \(range: 25–48\)\. Gender distribution: malen=24n=24, femalen=23n=23\. A complete participant demographics table is provided in Supplementary Table S2\.
*Sample size justification\.*With 47 evaluators each rating all 10 models on all 62 questions, the fully crossed design yieldsn=62n=62question\-level observations per model for between\-model comparisons\. A two\-samplett\-test withn=62n=62per group provides\>\>99% power to detect the observed between\-tier effect size \(d=1\.89d=1\.89\) and\>\>80% power to detect a medium effect \(d=0\.50d=0\.50\) within the ceiling cluster, atα=0\.05\\alpha=0\.05\. For the participant\-level mixed\-effects model \(n=29,140n=29\{,\}140\), statistical power is substantially higher\. The primary design constraint was not statistical power but the feasibility of each participant completing 620 evaluations \(10 models×\\times62 questions\) within a 2–4 week window; 47 was the maximum achievable sample under this full\-crossing requirement\.
### Task design and reference\-answer development
Sixty\-two questions were developed to span six critical capability domains: reasoning and language understanding \(Q1–Q10; 10 questions\), mathematical problem solving \(Q11–Q20; 10 questions\), coding and software development \(Q21–Q30; 10 questions\), conversational AI and chatbot behaviour \(Q31–Q40; 10 questions\), safety and ethical considerations \(Q41–Q50; 10 questions\), and domain\-specific professional knowledge \(Q51–Q62; 12 questions\)\. Domain knowledge received two additional questions to accommodate its broader topical scope\. Questions were designed to require open\-ended generation \(not multiple\-choice\), spanning explanation, reasoning, code production, contextual adaptation, and domain\-specific expertise\.
Each question was paired with an expert\-validated reference answer developed through the following procedure\. First, initial reference answers were drafted by domain experts with a minimum of five years of relevant professional experience in the corresponding field\. These experts were recruited from academic and professional networks independent of the research team and had no affiliation, consulting relationship, or other connection with any model developer evaluated in this study; independence was confirmed via written disclosure forms prior to participation\. Second, each reference answer was independently reviewed by at least two additional domain experts who assessed accuracy, completeness, and clarity\. Third, where reviewers disagreed, references were revised through iterative discussion until consensus was achieved; consensus was defined as agreement by all reviewers that the reference captured the essential content expected of a high\-quality response to the question\. Fourth, reference answers were normalised for length \(targeting 150–400 words\) and structure to reduce stylistic variability across domains\.
We acknowledge that open\-ended questions may admit multiple valid responses, and a single expert reference cannot capture the full space of acceptable answers\. To mitigate potential style bias, reference answers were written in a neutral, expository style avoiding model\-characteristic formatting patterns \(such as bullet\-point lists or markdown headers\)\. The complete set of 62 questions, reference answers, and domain annotations is deposited in the study repository at the Nature Portfolio repository \(DOI to be assigned upon publication\) under a CC BY 4\.0 licence, enabling independent assessment of reference quality and potential style bias\. The question set, reference answers, and evaluation rubric are also provided in the Supplementary Materials\.
### Models evaluated
Ten frontier LLMs were selected to represent the diversity of commercially available systems as of late 2025 to early 2026, spanning standard transformer architectures\[[6](https://arxiv.org/html/2607.20454#bib.bib1)\], mixture\-of\-experts designs\[[13](https://arxiv.org/html/2607.20454#bib.bib39)\], open\-weight models\[[38](https://arxiv.org/html/2607.20454#bib.bib40)\], and retrieval\-augmented generation \(RAG\)\[[23](https://arxiv.org/html/2607.20454#bib.bib41)\]\. Extended Data Supplementary Table S1 reports model identifiers, organisations, release dates, access methods, and evaluation date ranges\. The ten models evaluated were: Claude Sonnet 4\.5 \(Anthropic; released September 2025\), Gemini 3 Flash \(Google DeepMind; December 2025\), GPT\-5\.2 \(OpenAI; December 2025\), GitHub Copilot with GPT\-5\.2 \(Microsoft; January 2026\), DeepSeek\-V3\.2 \(DeepSeek; January 2026\), Llama 4 Maverick \(Meta; December 2025\), Mistral Medium 3\.1 \(Mistral AI; November 2025\), Grok 4\.1 \(xAI; December 2025\), Qwen3\-235B \(Alibaba; January 2026\), and Perplexity AI \(Perplexity AI; updated February 2026\)\.
All models were accessed through their official web\-based user interfaces using default settings to reflect realistic end\-user conditions\. Temperature, system prompt, and retrieval settings were not manually adjusted; each model operated under its default configuration as presented to standard users\. Perplexity AI is retrieval\-native and automatically incorporates web search results, while other models operate in a purely generative mode by default\. Copilot and ChatGPT share the same underlying model \(GPT\-5\.2\) but differ in interface and system configuration; both were included to assess whether deployment\-level differences affect fidelity within the same model architecture\.
For each question, a new conversation session was initiated \(conversation memory cleared\) to ensure independence across questions\. In practice, participants opened a new chat window or clicked “New conversation” in each model’s interface before entering each question; for models that maintain persistent conversations \(e\.g\., ChatGPT with memory features\), participants were instructed to disable persistent memory or use incognito/private browser sessions\. Compliance was verified through spot\-checks of submitted screenshots \(requested for a random 10% of evaluations per participant\) and by confirming that no response referenced content from a prior question\. No cross\-contamination was detected\. Participants were instructed to copy and paste the standardised question text verbatim into each model interface, ensuring uniform prompting across all evaluators\. No system\-level prompts were prepended\. If a model refused to answer a question, participants recorded the refusal and assigned a rating of 1 \(completely unfaithful\), yielding maximum deviation\. Model refusals occurred in 82 of 29,140 evaluations \(0\.28%\), concentrated in the safety domain: Claude \(31 refusals, 0\.66% of its evaluations\), Gemini \(24, 0\.51%\), Llama \(12, 0\.26%\), and all other models combined \(15, 0\.05%\)\. Because the two high\-fidelity models also refused most frequently, the maximum\-deviation penalty for refusals works*against*these models; excluding refusals entirely would lower Claude’s safety\-domain deviation from 46\.1% to an estimated 43\.8% and Gemini’s from 47\.4% to 45\.1%, widening the between\-tier gap\. We retain the conservative penalty to avoid inflating high\-fidelity model scores\.
### Evaluation procedure and blinding
The evaluation was conducted in a blinded fashion: participants were not informed of which model generated each response during the rating phase\[[12](https://arxiv.org/html/2607.20454#bib.bib43)\]\. The procedure comprised three stages\.
*Stage 1: Response collection\.*Each participant submitted all 62 questions to all 10 models, collecting 620 responses\. Responses were saved with model\-identifying metadata stripped before the rating phase\.
*Stage 2: Calibration\.*Before the main evaluation, each participant completed a calibration session comprising 5 practice questions \(not included in the final dataset\) with pre\-rated exemplar responses spanning the full 1–5 Likert range\. Calibration responses were discussed in a group session to establish shared understanding of the rubric anchors\. Participants were required to achieve≥\\geq80% agreement with expert ratings on calibration items before proceeding\.
*Stage 3: Blinded evaluation\.*Participants rated each model response on a 5\-point Likert scale\[[25](https://arxiv.org/html/2607.20454#bib.bib23)\]by comparing it against the expert\-validated reference answer, using a structured evaluation interface \(Microsoft Excel spreadsheets\) that presented each model response alongside its corresponding reference answer with the model identity concealed\. Presentation order was randomised across participants using a balanced Latin\-square design to mitigate order effects\. Each participant completed all 620 evaluations over a period of 2–4 weeks\. Self\-reported median evaluation time was 30 days \(range: 15–60 days; approximately 1–4 hours per day, totalling∼\\sim60 hours per participant\)\. Participants were instructed to complete no more than 60 evaluations per session and to take breaks between sessions to mitigate fatigue effects\. The Latin\-square randomisation ensured that any residual fatigue or order effects were distributed uniformly across models and questions rather than confounding specific model–question cells\.
### Rating rubric
Participants rated each model response on the following 5\-point Likert scale, anchored to the expert\-validated reference answer:
5 — Completely faithful\.The response captures all key elements of the reference answer with equivalent accuracy, completeness, and nuance\. Minor stylistic differences are acceptable\.
4 — Mostly faithful\.The response addresses the core content correctly but omits or slightly misrepresents one or two secondary elements\.
3 — Partially faithful\.The response captures some correct elements but contains significant omissions, inaccuracies, or irrelevant content that diverges from the reference\.
2 — Mostly unfaithful\.The response addresses the topic but fails to capture the substance of the reference answer; major errors or omissions dominate\.
1 — Completely unfaithful\.The response is irrelevant, incorrect, or refuses to address the question\. No meaningful alignment with the reference answer\.
### Fidelity deviation computation
The primary endpoint is response fidelity deviation, defined as the normalised distance of a model’s response quality from perfect alignment with the expert reference as judged by human evaluators\. For each of the 29,140 evaluations, participantii\(i=1,…,47i=1,\\ldots,47\) rated modeljj\(j=1,…,10j=1,\\ldots,10\) on questionkk\(k=1,…,62k=1,\\ldots,62\) with scoreRijk∈\{1,2,3,4,5\}R\_\{ijk\}\\in\\\{1,2,3,4,5\\\}, where 1==completely unfaithful and 5==perfectly faithful to the expert reference\. The fidelity deviation for that evaluation was computed as:
Deviationijk=5−Rijk4\\mathrm\{Deviation\}\_\{ijk\}=\\frac\{5\-R\_\{ijk\}\}\{4\}\(1\)yielding values in\[0,1\]\[0,1\]\(0==no deviation, perfect fidelity; 1==maximum deviation\)\. Model–question\-level deviation was then computed by averaging across all 47 participants:
Deviationjk=147∑i=147Deviationijk\\mathrm\{Deviation\}\_\{jk\}=\\frac\{1\}\{47\}\\sum\_\{i=1\}^\{47\}\\mathrm\{Deviation\}\_\{ijk\}\(2\)Model\-level deviation was computed by averaging across all 62 questions:
Deviationj=162∑k=162Deviationjk\\mathrm\{Deviation\}\_\{j\}=\\frac\{1\}\{62\}\\sum\_\{k=1\}^\{62\}\\mathrm\{Deviation\}\_\{jk\}\(3\)Domain\-level deviation was computed analogously, averaging over domain\-specific question subsets\. All deviation values reported in this paper are expressed as percentages \(multiplied by 100\)\.
*Performance tier classification\.*For descriptive purposes, we classify models into performance tiers using equidistant thresholds on the deviation scale: Good \(<<60%\), Moderate \(60–80%\), and Poor \(\>\>80%\)\. These thresholds correspond to intuitive anchors on the 5\-point Likert rubric: 60% deviation implies a mean rating of 2\.6 \(between “mostly unfaithful” and “partially faithful”\), while 80% implies a mean rating of 1\.8 \(near “mostly unfaithful”\)\. These are descriptive categories rather than validated clinical cut\-points and are used solely for visual presentation in figures\. Worked examples illustrating the mapping between raw Likert ratings and fidelity deviation values are provided in Supplementary Table S12\.
*Note on aggregation levels\.*The domain×\\timesmodel heatmap \(Fig\.[2](https://arxiv.org/html/2607.20454#Sx1.F2)a\) reports unweighted domain means, while Table[1](https://arxiv.org/html/2607.20454#Sx1.T1)reports question\-weighted overall deviation\. Minor discrepancies between these values \(typically<<1 pp\) reflect the different weighting of the 12\-item domain knowledge domain relative to the 10\-item domains\.
*Relationship to automated NLP metrics\.*The dataset additionally contains NLP\-computed similarity metrics \(semantic similarity via sentence embeddings, lexical overlap, part\-of\-speech alignment, sentiment agreement, and length ratios\) computed by comparing each model response against the reference answer using automated methods\. These metrics serve as independent validation measures and are not inputs to the primary fidelity deviation calculation, which is derived exclusively from human Likert ratings\. The two measurement families are computed through entirely separate pipelines with no shared computational inputs \(Extended Data Supplementary Table S10\)\.
### Inter\-rater reliability
Inter\-rater reliability was assessed using the intraclass correlation coefficient \(ICC\(2,1\) for consistency\)\[[36](https://arxiv.org/html/2607.20454#bib.bib20)\]\. The overall ICC\(2,1\) across all model–question combinations was 0\.84 \(95% CI \[0\.81, 0\.87\]\), indicating good agreement among the 47 evaluators\[[20](https://arxiv.org/html/2607.20454#bib.bib26)\]\. Domain\-level ICCs ranged from 0\.78 \(safety\) to 0\.89 \(reasoning\), with all domains exceeding the 0\.75 threshold for “good” reliability\. Krippendorff’sα\\alpha\(a chance\-corrected agreement coefficient applicable to any number of raters and measurement levels\)\[[21](https://arxiv.org/html/2607.20454#bib.bib21)\]was computed as an additional check, yieldingα=0\.79\\alpha=0\.79overall, consistent with the ICC results\.
### Statistical analysis
All analyses were conducted in Python 3\.11 and R 4\.3\.2\. The analysis pipeline is deposited in the study repository\.
*Primary analysis\.*Model\-level differences were tested using a Kruskal–WallisHHtest across all 10 models \(non\-parametric, no distributional assumptions\) and a one\-way ANOVA on question\-level means\[[18](https://arxiv.org/html/2607.20454#bib.bib24)\]\. Both parametric and non\-parametric tests are reported\.
*Plateau homogeneity\.*To test whether the eight ceiling models form a statistically indistinguishable group, we applied both Kruskal–Wallis and one\-way ANOVA to the eight\-model subset \(8 models×\\times62 questions=496=496observations\)\.
*Pairwise comparisons\.*All 45 pairwise model comparisons were conducted using Tukey’s HSD correction via themultcomppackage\[[18](https://arxiv.org/html/2607.20454#bib.bib24)\]\. Low\-deviation\-versus\-ceiling contrasts were additionally tested using Welch’stt\-test with Holm–Bonferroni correction\[[17](https://arxiv.org/html/2607.20454#bib.bib25)\]\.
*Effect sizes\.*Cohen’sdd\[[9](https://arxiv.org/html/2607.20454#bib.bib22)\]was computed for all key contrasts using pooled standard deviations, with 95% bootstrap confidence intervals \(10,000 iterations\)\. Variance decomposition was performed using two\-way ANOVA \(Model×\\timesQuestion\) with Type III sums of squares to partition total fidelity variance into model, question, and residual \(interaction\+\+error\) components \(Table[2](https://arxiv.org/html/2607.20454#Sx1.T2), \(A\)\)\.
*Bootstrap rank uncertainty\.*For each of 10,000 bootstrap iterations, we resampled the 62 questions with replacement and re\-ranked all 10 models by mean deviation\. We report the proportion of resamples in which each model occupies each rank position, as well as 95% rank intervals \(Extended Data Supplementary Table S8\)\.
*Inter\-model correlations\.*Pairwise Pearson correlations were computed from the 62\-element per\-question deviation vectors for each model pair, yielding a10×1010\\times 10correlation matrix reflecting the degree of shared success/failure profiles \(Supplementary Fig\. S3\)\.
### Mixed\-effects modelling
To partition variance while respecting the hierarchical structure of the fully crossed design, we fit a linear mixed\-effects model to the full 29,140\-cell participant\-level data matrix\. The model specifies fidelity deviation as the response variable with model as a fixed effect and crossed random intercepts for participant and question:
Deviationijk=β0\+βℓ\(j\)\+ui\+vk\+εijk\\mathrm\{Deviation\}\_\{ijk\}=\\beta\_\{0\}\+\\beta\_\{\\ell\(j\)\}\+u\_\{i\}\+v\_\{k\}\+\\varepsilon\_\{ijk\}\(4\)whereDeviationijk\\mathrm\{Deviation\}\_\{ijk\}is the fidelity deviation for participantii, modeljj, and questionkk;β0\\beta\_\{0\}is the grand intercept;βℓ\(j\)\\beta\_\{\\ell\(j\)\}is the fixed effect of modeljj\(treatment\-coded with the median\-performing model as reference\); and the random effects are:
ui∼𝒩\(0,τparticipant2\),vk∼𝒩\(0,τquestion2\),εijk∼𝒩\(0,σ2\)\.u\_\{i\}\\sim\\mathcal\{N\}\(0,\\tau\_\{\\mathrm\{participant\}\}^\{2\}\),\\quad v\_\{k\}\\sim\\mathcal\{N\}\(0,\\tau\_\{\\mathrm\{question\}\}^\{2\}\),\\quad\\varepsilon\_\{ijk\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)\.The model was fitted using restricted maximum likelihood \(REML\) via thelme4package\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\]in R 4\.3\.2\. Variance components were extracted to compute intraclass correlation coefficients for participant, question, and residual sources \(Table[2](https://arxiv.org/html/2607.20454#Sx1.T2), \(B\)\)\. Marginal and conditionalR2R^\{2\}values were computed following the method of Nakagawa and Schielzeth\[[29](https://arxiv.org/html/2607.20454#bib.bib18)\]\.
An extended model incorporating domain as a fixed effect and a model×\\timesdomain interaction was fitted to test whether the fidelity ceiling persists after accounting for domain\-level variation\. Model comparison used likelihood ratio tests and AIC/BIC\. Pairwise contrasts between models were computed using estimated marginal means with Holm–Bonferroni correction\[[17](https://arxiv.org/html/2607.20454#bib.bib25)\]\.
### Equivalence testing for ceiling homogeneity
To provide positive evidence that the eight ceiling models are statistically equivalent—rather than merely failing to detect differences—we applied the two one\-sided tests \(TOST\) procedure\[[22](https://arxiv.org/html/2607.20454#bib.bib16),[34](https://arxiv.org/html/2607.20454#bib.bib17)\]\. For each of the 28 pairwise comparisons among the eight ceiling models, we tested the null hypothesis that the true mean difference in fidelity deviation exceeds a pre\-specified equivalence boundΔ\\Deltain either direction:
H01:μi−μj≤−Δ,H02:μi−μj≥\+ΔH\_\{01\}\\\!:\\mu\_\{i\}\-\\mu\_\{j\}\\leq\-\\Delta,\\qquad H\_\{02\}\\\!:\\mu\_\{i\}\-\\mu\_\{j\}\\geq\+\\Delta\(5\)whereμi\\mu\_\{i\}andμj\\mu\_\{j\}denote the population mean fidelity deviation for modelsiiandjj, respectively,Δ\\Deltais the equivalence bound \(smallest effect size of interest\), andH01H\_\{01\}/H02H\_\{02\}are the two one\-sided null hypotheses\. Equivalence is concluded when bothH01H\_\{01\}andH02H\_\{02\}are rejected at significance levelα=0\.05\\alpha=0\.05, indicating that the observed difference falls within\(−Δ,\+Δ\)\(\-\\Delta,\+\\Delta\)\.
We setΔ=5\\Delta=5percentage points \(pp\) of fidelity deviation as the smallest effect size of interest \(SESOI\)\. This threshold was chosen on substantive grounds: the high\-fidelity\-versus\-ceiling gap is 30–48 pp across domains, so a within\-ceiling difference of 5 pp would be an order of magnitude smaller than the between\-tier effect and would have negligible practical importance for model selection\. As a sensitivity check, we also report TOST results atΔ=3\\Delta=3pp \(strict\) andΔ=7\\Delta=7pp \(lenient\) bounds\.
TOST was conducted using Welch’stt\-test \(unequal variances\) on the 62 per\-question fidelity deviation values per model, implemented via theTOSTERpackage\[[22](https://arxiv.org/html/2607.20454#bib.bib16)\]in R 4\.3\.2\.
### Robustness checks
We conducted four sensitivity analyses to assess the stability of the primary findings: \(1\) question\-subset bootstrap \(1,000 resamples of 50% of questions; ceiling non\-significancep\>0\.05p\>0\.05obtained in 96\.3% of resamples\); \(2\) leave\-one\-domain\-out \(rank order preserved in all 6 iterations; maximum rank change=1=1position within ceiling\); \(3\) alternative deviation thresholds \(ceiling identified at 75%, 80%, and 85% cutoffs with≤\\leq1 model reclassified\); \(4\) extreme\-question jackknife \(removing the 5 hardest questions shifts ceiling means by≤\\leq1\.2 pp, no rank changes between tiers\)\. Additionally, sensitivity analyses excluding the most extreme evaluator \(Participant 3, identified as the least reference\-aligned evaluator in 68\.4% of model–question combinations\) confirmed that all primary findings—the bimodal distribution, ceiling equivalence, and between\-tier effect size—are robust to evaluator exclusion; results are reported in the Supplementary Information\.
### Construct validity analyses
To assess whether fidelity deviation reflects content quality rather than stylistic conformity to reference formatting, we conducted five complementary analyses on the 620\-cell model–question matrix using the seven automated NLP features \(semantic similarity, lexical overlap, token\-length ratio, character\-length ratio, part\-of\-speech alignment, sentiment agreement, and composite overall similarity\)\.
*Predictive modelling\.*We trained linear regression, random forest \(100 estimators, max depth 6\), gradient\-boosted trees \(XGBoost; 100 estimators, max depth 4\), and multilayer perceptrons \(MLP; architectures 32–16 and 64–32–16, early stopping\) to predict fidelity deviation from NLP features\. All models were evaluated using 5\-fold cross\-validatedR2R^\{2\}\. Feature sets were tested separately: content features \(semantic similarity\), style features \(length ratios, POS alignment\), and all features combined\.
*Mediation analysis\.*We tested whether semantic similarity mediates the relationship between model tier \(high\-fidelity vs\. ceiling\) and fidelity deviation, using the Baron–Kenny framework with 5,000 bootstrap resamples for confidence intervals\. Style mediation \(via token\-length ratio\) was computed for comparison\.
*Unsupervised clustering\.*kk\-means \(k=2k=2\) and hierarchical agglomerative clustering were applied to standardised model\-level NLP feature vectors \(10 models×\\times7 features\) to test whether the human\-identified tier structure is recoverable from automated features alone\. The adjusted Rand index \(ARI\) was used to assess cluster–tier correspondence\.
*Evaluator consensus analysis\.*The standard deviation of semantic similarity ratings across 47 evaluators was compared between tiers using the Mann–WhitneyUUtest\. Evaluator profiles were constructed from best/worst alignment frequencies and clustered viakk\-means \(k=3k=3\)\.
*Domain transfer prediction\.*Leave\-one\-domain\-out cross\-validation assessed whether NLP features trained on five domains could predict fidelity deviation in the held\-out sixth domain \(random forest, 100 estimators\)\. All analyses used Python 3\.11 with scikit\-learn 1\.4 and XGBoost 2\.0; random seeds were fixed for reproducibility\.
### Data completeness
All 29,140 cells \(47 participants×\\times10 models×\\times62 questions\) are complete; there are no missing evaluations\. Model refusals \(cases where a model declined to answer or produced an error\) occurred in fewer than 0\.3% of response collection attempts \(primarily safety\-domain questions\)\. In these cases, participants followed the protocol of assigning a rating of 1, consistent with maximum deviation\.
### Software and reproducibility
Statistical analyses used the following packages: Python \(NumPy 1\.26, SciPy 1\.12, Pandas 2\.1, Matplotlib 3\.8, Seaborn 0\.13\); R \(lme4 1\.1\-35\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\], multcomp 1\.4\-25\[[18](https://arxiv.org/html/2607.20454#bib.bib24)\], psych 2\.3, boot 1\.3\-28, effsize 0\.8\.1, TOSTER 0\.8\.6\[[22](https://arxiv.org/html/2607.20454#bib.bib16)\]\)\. Figures use theRdYlGn\_r\(red–yellow–green reversed\) sequential colour palette from Matplotlib/ColorBrewer for heatmaps and theSet2qualitative palette for domain colour\-coding; both palettes were verified for accessibility using the Coblis colour\-blindness simulator\. All code is deposited in the study repository at the Nature Portfolio repository \(DOI to be assigned upon publication\)\. Estimated computation time for full reproduction:<<30 minutes on a standard laptop \(Intel Core i7 or equivalent, 16 GB RAM; no GPU required\)\. All analyses used fixed random seeds for exact reproducibility\. Complete software version details are reported in Supplementary Table S17\.
### Use of AI in manuscript preparation
Generative AI tools were used during manuscript preparation in two capacities: \(1\) proofreading, grammar checking, and language refinement of the manuscript text; and \(2\) generating the conceptual illustrations in Fig\.[1](https://arxiv.org/html/2607.20454#S0.F1)a and b, which depict the study rationale and experimental design schematically\. These figures are illustrative diagrams only and do not represent empirical data, analytical results, or statistical outputs\. All scientific content, experimental design, data collection, statistical analyses, data\-driven figures \(Figs\. 2–4\), interpretation, and conclusions are the sole work of the human authors\. The authors reviewed and edited all AI\-assisted content and take full responsibility for the accuracy and integrity of the published work\.
## Acknowledgements
The authors thank the 47 participants who volunteered their time to complete the evaluation protocol, and the independent domain experts who developed and validated the reference answers\.
## Author Contributions
M\.A\.:Conceptualisation, Methodology, Formal Analysis, Writing—Original Draft, Writing—Review & Editing, Supervision, Project Administration\.A\.A\.:Data Curation, Investigation, Software, Validation\.F\.A\.:Data Curation, Investigation, Validation\.G\.V\.E\.:Data Curation, Investigation, Software, Validation, Visualisation\.M\.R\.:Methodology, Writing—Review & Editing, Resources\.
## Funding
This research received no external funding\. No grants, contracts, or other financial support from any funding agency in the public, commercial, or not\-for\-profit sectors were received for this work\.
## Competing Interests
The authors declare no competing interests\. No author has a financial relationship, consulting arrangement, advisory role, or other affiliation with any model developer evaluated in this study, including Anthropic and Google DeepMind\. No participant or reference\-answer expert had any disclosed affiliation with a model developer\. The study was not commercially funded\. Reference answers were developed by independent domain experts with no involvement from any model developer \(see Methods\)\.
## Data Availability
The complete dataset \(29,140 evaluations\), participant\-level matrices, all 62 prompts with reference answers, and model documentation will be deposited at a Nature Portfolio repository under CC BY 4\.0 upon publication, with a persistent DOI assigned at that time\. During peer review, all data are available from the corresponding author upon reasonable request\.
## Code Availability
All analysis code \(R 4\.3\.2\+, Python 3\.11\+\) and the fidelity computation pipeline are at the same repository\. Reproduction time:<<30 min\.
## Reporting Summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article\.
## References
- \[1\]\(2024\-03\)Do language models know when they’re hallucinating references?\.InFindings of the Association for Computational Linguistics: EACL 2024,Y\. Graham and M\. Purver \(Eds\.\),St\. Julian’s, Malta,pp\. 912–928\.External Links:[Link](https://aclanthology.org/2024.findings-eacl.62/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-eacl.62)Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[2\]Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.\(2022\)Constitutional AI: harmlessness from AI feedback\.Note:Preprint at arXiv[https://arxiv\.org/abs/2212\.08073](https://arxiv.org/abs/2212.08073)Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3)\.
- \[3\]D\. Bates, M\. Mächler, B\. Bolker, and S\. Walker\(2015\)Fitting linear mixed\-effects models using lme4\.Journal of Statistical Software67\(1\),pp\. 1–48\.External Links:[Link](https://www.jstatsoft.org/index.php/jss/article/view/v067i01),[Document](https://dx.doi.org/10.18637/jss.v067.i01)Cited by:[Supplementary Code](https://arxiv.org/html/2607.20454#Ax7.p1.2),[All frontier models exhibit response drift, but magnitude varies markedly](https://arxiv.org/html/2607.20454#Sx1.SSx1.p2.5),[Table 2](https://arxiv.org/html/2607.20454#Sx1.T2),[Table 2](https://arxiv.org/html/2607.20454#Sx1.T2.28.14.14),[Study design overview](https://arxiv.org/html/2607.20454#Sx3.SSx1.p1.3),[Mixed\-effects modelling](https://arxiv.org/html/2607.20454#Sx3.SSx10.p1.8),[Software and reproducibility](https://arxiv.org/html/2607.20454#Sx3.SSx15.p1.1)\.
- \[4\]R\. Bommasani, D\. A\. Hudson, E\. Adeli, R\. Altman, S\. Arora, S\. von Arx, M\. S\. Bernstein, J\. Bohg, A\. Bosselut, E\. Brunskill, E\. Brynjolfsson, S\. Buch, D\. Card, R\. Castellon, N\. S\. Chatterji, A\. S\. Chen, K\. A\. Creel, J\. Davis, D\. Demszky, C\. Donahue, M\. Doumbouya, E\. Durmus, S\. Ermon, J\. Etchemendy, K\. Ethayarajh, L\. Fei\-Fei, C\. Finn, T\. Gale, L\. E\. Gillespie, K\. Goel, N\. D\. Goodman, S\. Grossman, N\. Guha, T\. Hashimoto, P\. Henderson, J\. Hewitt, D\. E\. Ho, J\. Hong, K\. Hsu, J\. Huang, T\. F\. Icard, S\. Jain, D\. Jurafsky, P\. Kalluri, S\. Karamcheti, G\. Keeling, F\. Khani, O\. Khattab, P\. W\. Koh, M\. S\. Krass, R\. Krishna, R\. Kuditipudi, A\. Kumar, F\. Ladhak, M\. Lee, T\. Lee, J\. Leskovec, I\. Levent, X\. L\. Li, X\. Li, T\. Ma, A\. Malik, C\. D\. Manning, S\. P\. Mirchandani, E\. Mitchell, Z\. Munyikwa, S\. Nair, A\. Narayan, D\. Narayanan, B\. Newman, A\. Nie, J\. C\. Niebles, H\. Nilforoshan, J\. F\. Nyarko, G\. Ogut, L\. Orr, I\. Papadimitriou, J\. S\. Park, C\. Piech, E\. Portelance, C\. Potts, A\. Raghunathan, R\. Reich, H\. Ren, F\. Rong, Y\. H\. Roohani, C\. Ruiz, J\. Ryan, C\. R’e, D\. Sadigh, S\. Sagawa, K\. Santhanam, A\. Shih, K\. P\. Srinivasan, A\. Tamkin, R\. Taori, A\. W\. Thomas, F\. Tramèr, R\. E\. Wang, W\. Wang, B\. Wu, J\. Wu, Y\. Wu, S\. M\. Xie, M\. Yasunaga, J\. You, M\. A\. Zaharia, M\. Zhang, T\. Zhang, X\. Zhang, Y\. Zhang, L\. Zheng, K\. Zhou, and P\. Liang\(2021\)On the opportunities and risks of foundation models\.ArXiv\.External Links:[Link](https://crfm.stanford.edu/assets/report.pdf)Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p5.1)\.
- \[5\]R\. Bommasani, P\. Liang, and T\. Lee\(2023\)Holistic evaluation of language models\.Annals of the New York Academy of Sciences1525\(1\),pp\. 140–146\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/nyas.15007),[Link](https://nyaspubs.onlinelibrary.wiley.com/doi/abs/10.1111/nyas.15007),https://nyaspubs\.onlinelibrary\.wiley\.com/doi/pdf/10\.1111/nyas\.15007Cited by:[S2\. Holistic evaluation frameworks](https://arxiv.org/html/2607.20454#Ax3.SSx2.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p3.2),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p5.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[6\]T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei\(2020\)Language models are few\-shot learners\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Models evaluated](https://arxiv.org/html/2607.20454#Sx3.SSx4.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[7\]Z\. Chen, S\. Wang, T\. Xiao, Y\. Wang, S\. Chen, X\. Cai, J\. He, and J\. WangW\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\)\(2025\-07\)Revisiting scaling laws for language models: the role of data quality and training strategies\.Association for Computational Linguistics,Vienna, Austria\.External Links:[Link](https://aclanthology.org/2025.acl-long.1163/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1163),ISBN 979\-8\-89176\-251\-0Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[8\]W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, B\. Zhu, H\. Zhang, M\. I\. Jordan, J\. E\. Gonzalez, and I\. Stoica\(2024\)Chatbot arena: an open platform for evaluating llms by human preference\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[S2\. Holistic evaluation frameworks](https://arxiv.org/html/2607.20454#Ax3.SSx2.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p2.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[9\]J\. Cohen\(1988\)Statistical power analysis for the behavioral sciences\.2 edition,Routledge,New York\.External Links:[Document](https://dx.doi.org/10.4324/9780203771587),ISBN 9780203771587Cited by:[All frontier models exhibit response drift, but magnitude varies markedly](https://arxiv.org/html/2607.20454#Sx1.SSx1.p1.10),[Statistical analysis](https://arxiv.org/html/2607.20454#Sx3.SSx9.p5.3)\.
- \[10\]X\. Du, M\. Liu, K\. Wang, H\. Wang, J\. Liu, Y\. Chen, J\. Feng, C\. Sha, X\. Peng, and Y\. Lou\(2024\)Evaluating large language models in class\-level code generation\.InProceedings of the IEEE/ACM 46th International Conference on Software Engineering,ICSE ’24,New York, NY, USA\.External Links:ISBN 9798400702174,[Link](https://doi.org/10.1145/3597503.3639219),[Document](https://dx.doi.org/10.1145/3597503.3639219)Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[11\]Y\. Dubois, P\. Liang, and T\. Hashimoto\(2024\)Length\-controlled alpacaeval: a simple debiasing of automatic evaluators\.InFirst Conference on Language Modeling,Cited by:[S2\. Holistic evaluation frameworks](https://arxiv.org/html/2607.20454#Ax3.SSx2.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p2.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[12\]E\. DURMUS, K\. Nguyen, T\. Liao, N\. Schiefer, A\. Askell, A\. Bakhtin, C\. Chen, Z\. Hatfield\-Dodds, D\. Hernandez, N\. Joseph, L\. Lovitt, S\. McCandlish, O\. Sikder, A\. Tamkin, J\. Thamkul, J\. Kaplan, J\. Clark, and D\. Ganguli\(2024\)Towards measuring the representation of subjective global opinions in language models\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=zl16jLb91v)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Evaluation procedure and blinding](https://arxiv.org/html/2607.20454#Sx3.SSx5.p1.1)\.
- \[13\]W\. Fedus, B\. Zoph, and N\. Shazeer\(2022\-01\)Switch transformers: scaling to trillion parameter models with simple and efficient sparsity\.J\. Mach\. Learn\. Res\.23\(1\)\.External Links:ISSN 1532\-4435Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Models evaluated](https://arxiv.org/html/2607.20454#Sx3.SSx4.p1.1)\.
- \[14\]L\. Gao, J\. Schulman, and J\. Hilton\(2023\)Scaling laws for reward model overoptimization\.InProceedings of the 40th International Conference on Machine Learning,ICML’23\.Cited by:[Discussion](https://arxiv.org/html/2607.20454#Sx2.p2.1)\.
- \[15\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt\(2021\)Measuring massive multitask language understanding\.InProceedings of the International Conference on Learning Representations,Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[16\]J\. Hoffmann, S\. Borgeaud, A\. Mensch, E\. Buchatskaya, T\. Cai, E\. Rutherford, D\. de Las Casas, L\. A\. Hendricks, J\. Welbl, A\. Clark, T\. Hennigan, E\. Noland, K\. Millican, G\. van den Driessche, B\. Damoc, A\. Guy, S\. Osindero, K\. Simonyan, E\. Elsen, O\. Vinyals, J\. W\. Rae, and L\. Sifre\(2022\)Training compute\-optimal large language models\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[17\]S\. Holm\(1979\)A simple sequentially rejective multiple test procedure\.Scandinavian Journal of Statistics6\(2\),pp\. 65–70\.External Links:ISSN 03036898, 14679469,[Link](http://www.jstor.org/stable/4615733)Cited by:[Mixed\-effects modelling](https://arxiv.org/html/2607.20454#Sx3.SSx10.p2.1),[Statistical analysis](https://arxiv.org/html/2607.20454#Sx3.SSx9.p4.1)\.
- \[18\]T\. Hothorn, F\. Bretz, and P\. Westfall\(2008\)Simultaneous inference in general parametric models\.Biometrical Journal50\(3\),pp\. 346–363\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/bimj.200810425),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1002/bimj.200810425)Cited by:[Software and reproducibility](https://arxiv.org/html/2607.20454#Sx3.SSx15.p1.1),[Statistical analysis](https://arxiv.org/html/2607.20454#Sx3.SSx9.p2.1),[Statistical analysis](https://arxiv.org/html/2607.20454#Sx3.SSx9.p4.1)\.
- \[19\]Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. Fung\(2023\-03\)Survey of hallucination in natural language generation\.ACM Comput\. Surv\.55\(12\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3571730),[Document](https://dx.doi.org/10.1145/3571730)Cited by:[S3\. Factual accuracy and hallucination assessment](https://arxiv.org/html/2607.20454#Ax3.SSx3.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3)\.
- \[20\]T\. K\. Koo and M\. Y\. Li\(2016\)A guideline of selecting and reporting intraclass correlation coefficients for reliability research\.Journal of Chiropractic Medicine15\(2\),pp\. 155–163\.External Links:ISSN 1556\-3707,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jcm.2016.02.012),[Link](https://www.sciencedirect.com/science/article/pii/S1556370716000158)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p3.2),[Inter\-rater reliability](https://arxiv.org/html/2607.20454#Sx3.SSx8.p1.2)\.
- \[21\]K\. Krippendorff\(2006\-01\)Reliability in content analysis: some common misconceptions and recommendations\.Human Communication Research30\(3\),pp\. 411–433\.External Links:ISSN 0360\-3989,[Document](https://dx.doi.org/10.1111/j.1468-2958.2004.tb00738.x),[Link](https://doi.org/10.1111/j.1468-2958.2004.tb00738.x),https://academic\.oup\.com/hcr/article\-pdf/30/3/411/22338169/jhumcom0411\.pdfCited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Inter\-rater reliability](https://arxiv.org/html/2607.20454#Sx3.SSx8.p1.2)\.
- \[22\]D\. Lakens\(2017\)Equivalence tests: a practical primer for t tests, correlations, and meta\-analyses\.Social Psychological and Personality Science8\(4\),pp\. 355–362\.Note:PMID: 28736600External Links:[Document](https://dx.doi.org/10.1177/1948550617697177),[Link](https://doi.org/10.1177/1948550617697177),https://doi\.org/10\.1177/1948550617697177Cited by:[All frontier models exhibit response drift, but magnitude varies markedly](https://arxiv.org/html/2607.20454#Sx1.SSx1.p1.10),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3),[Equivalence testing for ceiling homogeneity](https://arxiv.org/html/2607.20454#Sx3.SSx11.p1.1),[Equivalence testing for ceiling homogeneity](https://arxiv.org/html/2607.20454#Sx3.SSx11.p3.1),[Software and reproducibility](https://arxiv.org/html/2607.20454#Sx3.SSx15.p1.1)\.
- \[23\]P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. Kiela\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[S3\. Factual accuracy and hallucination assessment](https://arxiv.org/html/2607.20454#Ax3.SSx3.p1.1),[Models evaluated](https://arxiv.org/html/2607.20454#Sx3.SSx4.p1.1)\.
- \[24\]T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, T\. Wu, B\. Zhu, J\. E\. Gonzalez, and I\. Stoica\(2025\)From crowdsourced data to high\-quality benchmarks: arena\-hard and benchbuilder pipeline\.InProceedings of the 42nd International Conference on Machine Learning,ICML’25\.Cited by:[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[25\]R\. Likert\(1932\)A technique for the measurement of attitudes\.Archives of Psychology22\(140\),pp\. 1–55\.Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p6.1),[Evaluation procedure and blinding](https://arxiv.org/html/2607.20454#Sx3.SSx5.p4.1)\.
- \[26\]B\. Y\. Lin, Y\. Deng, K\. Chandu, F\. Brahman, A\. Ravichander, V\. Pyatkin, N\. Dziri, R\. L\. Bras, and Y\. Choi\(2024\)WildBench: benchmarking llms with challenging tasks from real users in the wild\.External Links:2406\.04770,[Link](https://arxiv.org/abs/2406.04770)Cited by:[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[27\]K\. Mahowald, A\. A\. Ivanova, I\. A\. Blank, N\. Kanwisher, J\. B\. Tenenbaum, and E\. Fedorenko\(2024\)Dissociating language and thought in large language models\.Trends in Cognitive Sciences28\(6\),pp\. 517–540\.External Links:ISSN 1364\-6613,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.tics.2024.01.011),[Link](https://www.sciencedirect.com/science/article/pii/S1364661324000275)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p2.1.1)\.
- \[28\]S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. Hajishirzi\(2023\-12\)FActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 12076–12100\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.741/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.741)Cited by:[S3\. Factual accuracy and hallucination assessment](https://arxiv.org/html/2607.20454#Ax3.SSx3.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p2.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[29\]S\. Nakagawa and H\. Schielzeth\(2013\)A general and simple method for obtaining r2 from generalized linear mixed\-effects models\.Methods in Ecology and Evolution4\(2\),pp\. 133–142\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/j.2041-210x.2012.00261.x),[Link](https://besjournals.onlinelibrary.wiley.com/doi/abs/10.1111/j.2041-210x.2012.00261.x),https://besjournals\.onlinelibrary\.wiley\.com/doi/pdf/10\.1111/j\.2041\-210x\.2012\.00261\.xCited by:[All frontier models exhibit response drift, but magnitude varies markedly](https://arxiv.org/html/2607.20454#Sx1.SSx1.p2.5),[Table 2](https://arxiv.org/html/2607.20454#Sx1.T2),[Table 2](https://arxiv.org/html/2607.20454#Sx1.T2.28.14.14),[Mixed\-effects modelling](https://arxiv.org/html/2607.20454#Sx3.SSx10.p1.8)\.
- \[30\]Nature Machine Intelligence\(2026/01/01\)Multi\-agent ai systems need transparency\.Nature Machine Intelligence8\(1\),pp\. 1–1\.External Links:[Document](https://dx.doi.org/10.1038/s42256-026-01183-2),ISBN 2522\-5839,[Link](https://doi.org/10.1038/s42256-026-01183-2)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p5.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p6.1)\.
- \[31\]L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. Lowe\(2022\)Training language models to follow instructions with human feedback\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3)\.
- \[32\]L\. Phan, A\. Gatti, N\. Li, A\. Khoja, R\. Kim, R\. Ren, J\. Hausenloy, O\. Zhang, M\. Mazeika, D\. Hendrycks, Z\. Han, J\. Hu, H\. Zhang, C\. B\. C\. Zhang, M\. Shaaban, J\. Ling, S\. Shi, M\. Choi, A\. Agrawal, A\. Chopra, A\. Nattanmai, G\. McKellips, A\. Cheraku, A\. Suhail, E\. Luo, M\. Deng, J\. Luo, A\. Zhang, K\. Jindel, J\. Paek, K\. Halevy, A\. Baranov, M\. Liu, A\. Avadhanam, D\. Zhang, V\. Cheng, B\. Ma, E\. Fu, L\. Do, J\. Lass, H\. Yang, S\. Sunkari, V\. Bharath, V\. Ai, J\. Leung, R\. Agrawal, A\. Zhou, K\. Chen, T\. Kalpathi, Z\. Xu, G\. Wang, T\. Xiao, E\. Maung, S\. Lee, R\. Yang, R\. Yue, B\. Zhao, J\. Yoon, X\. Sun, A\. Singh, C\. Peng, T\. Osbey, T\. Wang, D\. Echeazu, T\. Wu, S\. Patel, V\. Kulkarni, V\. Sundarapandiyan, A\. Le, Z\. Nasim, S\. Yalam, R\. Kasamsetty, S\. Samal, D\. Sun, N\. Shah, A\. Saha, A\. Zhang, L\. Nguyen, L\. Nagumalli, K\. Wang, A\. Wu, A\. Telluri, S\. Yue, A\. Wang, D\. Dodonov, T\. Nguyen, J\. Lee, D\. Anderson, M\. Doroshenko, A\. C\. Stokes, M\. Mahmood, O\. Pokutnyi, O\. Iskra, J\. P\. Wang, J\. Levin, M\. Kazakov, F\. Feng, S\. Y\. Feng, H\. Zhao, M\. Yu, V\. Gangal, C\. Zou, Z\. Wang, S\. Popov, R\. Gerbicz, G\. Galgon, J\. Schmitt, W\. Yeadon, Y\. Lee, S\. Sauers, A\. Sanchez, F\. Giska, M\. Roth, S\. Riis, S\. Utpala, N\. Burns, G\. M\. Goshu, M\. M\. Naiya, C\. Agu, Z\. Giboney, A\. Cheatom, F\. Fournier\-Facio, S\. Crowson, L\. Finke, Z\. Cheng, J\. Zampese, R\. G\. Hoerr, M\. Nandor, H\. Park, T\. Gehrunger, J\. Cai, B\. McCarty, A\. C\. Garretson, E\. Taylor, D\. Sileo, Q\. Ren, U\. Qazi, L\. Li, J\. Nam, J\. B\. Wydallis, P\. Arkhipov, J\. W\. L\. Shi, A\. Bacho, C\. G\. Willcocks, H\. Cao, S\. Motwani, E\. de Oliveira Santos, J\. Veith, E\. Vendrow, D\. Cojoc, K\. Zenitani, J\. Robinson, L\. Tang, Y\. Li, J\. Vendrow, N\. W\. Fraga, V\. Kuchkin, A\. P\. Maksimov, P\. Marion, D\. Efremov, J\. Lynch, K\. Liang, A\. Mikov, A\. Gritsevskiy, J\. Guillod, G\. Demir, D\. Martinez, B\. Pageler, K\. Zhou, S\. Soori, O\. Press, H\. Tang, P\. Rissone, S\. R\. Green, L\. Brüssel, M\. Twayana, A\. Dieuleveut, J\. M\. Imperial, A\. Prabhu, J\. Yang, N\. Crispino, A\. Rao, D\. Zvonkine, G\. Loiseau, M\. Kalinin, M\. Lukas, C\. Manolescu, N\. Stambaugh, S\. Mishra, T\. Hogg, C\. Bosio, B\. P\. Coppola, J\. Salazar, J\. Jin, R\. Sayous, S\. Ivanov, P\. Schwaller, S\. Senthilkumar, A\. M\. Bran, A\. Algaba, K\. Van den Houte, L\. Van Der Sypt, B\. Verbeken, D\. Noever, A\. Kopylov, B\. Myklebust, B\. Li, L\. Schut, E\. Zheltonozhskii, Q\. Yuan, D\. Lim, R\. Stanley, T\. Yang, J\. Maar, J\. Wykowski, M\. Oller, A\. Sahu, C\. G\. Ardito, Y\. Hu, A\. G\. K\. Kamdoum, A\. Jin, T\. G\. Vilchis, Y\. Zu, M\. Lackner, J\. Koppel, G\. Sun, D\. S\. Antonenko, S\. Chern, B\. Zhao, P\. Arsene, J\. M\. Cavanagh, D\. Li, J\. Shen, D\. Crisostomi, W\. Zhang, A\. Dehghan, S\. Ivanov, D\. Perrella, N\. Kaparov, A\. Zang, I\. Sucholutsky, A\. Kharlamova, D\. Orel, V\. Poritski, S\. Ben\-David, Z\. Berger, P\. Whitfill, M\. Foster, D\. Munro, L\. Ho, S\. Sivarajan, D\. B\. Hava, A\. Kuchkin, D\. Holmes, A\. Rodriguez\-Romero, F\. Sommerhage, A\. Zhang, R\. Moat, K\. Schneider, Z\. Kazibwe, D\. Clarke, D\. H\. Kim, F\. M\. Dias, S\. Fish, V\. Elser, T\. Kreiman, V\. E\. G\. Vilchis, I\. Klose, U\. Anantheswaran, A\. Zweiger, K\. Rawal, J\. Li, J\. Nguyen, N\. Daans, H\. Heidinger, M\. Radionov, V\. Rozhoň, V\. Ginis, C\. Stump, N\. Cohen, R\. Poświata, J\. Tkadlec, A\. Goldfarb, C\. Wang, P\. Padlewski, S\. Barzowski, K\. Montgomery, R\. Stendall, J\. Tucker\-Foltz, J\. Stade, T\. R\. Rogers, T\. Goertzen, D\. Grabb, A\. Shukla, A\. Givré, J\. A\. Ambay, A\. Sen, C\. for AI Safety, S\. AI, and H\. C\. Consortium\(2026/01/01\)A benchmark of expert\-level academic questions to assess ai capabilities\.Nature649\(8099\),pp\. 1139–1146\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09962-4),ISBN 1476\-4687,[Link](https://doi.org/10.1038/s41586-025-09962-4)Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p3.2),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p5.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[33\]D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman\(2024\)GPQA: a graduate\-level google\-proof q&a benchmark\.InFirst Conference on Language Modeling,Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[34\]D\. J\. Schuirmann\(1987\)A comparison of the two one\-sided tests procedure and the power approach for assessing the equivalence of average bioavailability\.Journal of Pharmacokinetics and Biopharmaceutics15\(6\),pp\. 657–680\.External Links:[Document](https://dx.doi.org/10.1007/BF01068419)Cited by:[Equivalence testing for ceiling homogeneity](https://arxiv.org/html/2607.20454#Sx3.SSx11.p1.1)\.
- \[35\]R\. Shah, V\. Varma, R\. Kumar, M\. Phuong, V\. Krakovna, J\. Uesato, and Z\. Kenton\(2022\)Goal misgeneralization: why correct specifications aren’t enough for correct goals\.Note:Preprint at arXiv[https://arxiv\.org/abs/2210\.01790](https://arxiv.org/abs/2210.01790)Cited by:[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3)\.
- \[36\]P\. E\. Shrout and J\. L\. Fleiss\(1979\)Intraclass correlations: uses in assessing rater reliability\.Psychological Bulletin86\(2\),pp\. 420–428\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.86.2.420)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Ceiling models share systematic drift patterns across questions](https://arxiv.org/html/2607.20454#Sx1.SSx3.p3.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p3.2),[Inter\-rater reliability](https://arxiv.org/html/2607.20454#Sx3.SSx8.p1.2)\.
- \[37\]M\. Steyvers, H\. Tejeda, A\. Kumar, C\. Belem, S\. Karny, X\. Hu, L\. W\. Mayer, and P\. Smyth\(2025/02/01\)What large language models know and what people think they know\.Nature Machine Intelligence7\(2\),pp\. 221–231\.External Links:[Document](https://dx.doi.org/10.1038/s42256-024-00976-7),ISBN 2522\-5839,[Link](https://doi.org/10.1038/s42256-024-00976-7)Cited by:[S5\. Human evaluation methodology](https://arxiv.org/html/2607.20454#Ax3.SSx5.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p2.1.1)\.
- \[38\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale, D\. Bikel, L\. Blecher, C\. C\. Ferrer, M\. Chen, G\. Cucurull, D\. Esiobu, J\. Fernandes, J\. Fu, W\. Fu, B\. Fuller, C\. Gao, V\. Goswami, N\. Goyal, A\. Hartshorn, S\. Hosseini, R\. Hou, H\. Inan, M\. Kardas, V\. Kerkez, M\. Khabsa, I\. Kloumann, A\. Korenev, P\. S\. Koura, M\. Lachaux, T\. Lavril, J\. Lee, D\. Liskovich, Y\. Lu, Y\. Mao, X\. Martinet, T\. Mihaylov, P\. Mishra, I\. Molybog, Y\. Nie, A\. Poulton, J\. Reizenstein, R\. Rungta, K\. Saladi, A\. Schelten, R\. Silva, E\. M\. Smith, R\. Subramanian, X\. E\. Tan, B\. Tang, R\. Taylor, A\. Williams, J\. X\. Kuan, P\. Xu, Z\. Yan, I\. Zarov, Y\. Zhang, A\. Fan, M\. Kambadur, S\. Narang, A\. Rodriguez, R\. Stojnic, S\. Edunov, and T\. Scialom\(2023\-07\)Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Models evaluated](https://arxiv.org/html/2607.20454#Sx3.SSx4.p1.1)\.
- \[39\]C\. Xiao, J\. Cai, W\. Zhao, B\. Lin, G\. Zeng, J\. Zhou, Z\. Zheng, X\. Han, Z\. Liu, and M\. Sun\(2025/11/01\)Densing law of llms\.Nature Machine Intelligence7\(11\),pp\. 1823–1833\.External Links:[Document](https://dx.doi.org/10.1038/s42256-025-01137-0),ISBN 2522\-5839,[Link](https://doi.org/10.1038/s42256-025-01137-0)Cited by:[S4\. Scaling laws and diminishing returns](https://arxiv.org/html/2607.20454#Ax3.SSx4.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3)\.
- \[40\]Q\. Zhao, Y\. Huang, T\. Lv, L\. Cui, Q\. Sun, S\. Mao, X\. Zhang, Y\. Xin, Q\. Yin, S\. Li, and F\. Wei\(2025\-07\)MMLU\-CF: a contamination\-free multi\-task language understanding benchmark\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 13371–13391\.External Links:[Link](https://aclanthology.org/2025.acl-long.656/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.656),ISBN 979\-8\-89176\-251\-0Cited by:[S1\. Automated benchmarks and their limitations](https://arxiv.org/html/2607.20454#Ax3.SSx1.p1.1),[Discussion](https://arxiv.org/html/2607.20454#Sx2.p1.3),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
- \[41\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[S2\. Holistic evaluation frameworks](https://arxiv.org/html/2607.20454#Ax3.SSx2.p1.1),[Response drift across frontier large language models](https://arxiv.org/html/2607.20454#p1.1.1)\.
## Supplementary Information
This Supplementary Information accompanies the main manuscript “Response drift across frontier large language models” and provides the complete set of extended data, additional analyses, and detailed documentation necessary to fully reproduce and evaluate the reported findings\. The main text presents summary results for 29,140 evaluations across 47 participants, 10 frontier large language models, and 62 standardised questions spanning six capability domains\. This supplement expands on those results by providing per\-question and per\-domain breakdowns, distributional analyses, the complete 62×\\times10 deviation matrix, all statistical test outputs, automated NLP metric comparisons, construct validity analyses, worked examples, and full software and data documentation\. Together, these materials enable independent verification of every numerical claim in the main text and support extension of the methodology to new models, domains, or evaluation protocols\.
The supplement is organised into the following sections\.*Supplementary Figures*[S1](https://arxiv.org/html/2607.20454#Ax1.F1)–[S6](https://arxiv.org/html/2607.20454#Ax1.F6)\(6 figures, 10 panels and 1 standalone\) present per\-question fidelity profiles for the four domains not shown in the main text \(Figs\.[S1](https://arxiv.org/html/2607.20454#Ax1.F1)–[S2](https://arxiv.org/html/2607.20454#Ax1.F2)\) and distributional box\-plot analyses for all six domains \(Figs\.[S4](https://arxiv.org/html/2607.20454#Ax1.F4)–[S6](https://arxiv.org/html/2607.20454#Ax1.F6)\), with a standalone question\-difficulty ranking \(Fig\.[S3](https://arxiv.org/html/2607.20454#Ax1.F3)\)\.*Supplementary Tables*[S1](https://arxiv.org/html/2607.20454#Ax1.T1)–[S18](https://arxiv.org/html/2607.20454#Ax8.T18)\(18 tables\) document model versions and access details, participant demographics and recruitment characteristics, evaluator heterogeneity and outlier analysis, the complete 620\-cell deviation matrix, variance decomposition by domain, equivalence testing for ceiling homogeneity, all 45 pairwise effect sizes, bootstrap rank stability across 10,000 resamples, NLP similarity metrics by model and domain and their correlations with human\-judged fidelity deviation, worked examples linking raw ratings to deviation values, a side\-by\-side comparison of the two independent measurement pipelines, machine\-learning prediction results with cross\-validatedR2R^\{2\}for five model families, content\-versus\-style regression decomposition, evaluator clustering profiles, and software versions for all packages used\.*Supplementary Notes*1–4 detail sensitivity analyses \(evaluator exclusion robustness, leave\-one\-domain\-out stability\), mediation analysis \(Baron–Kenny framework with bootstrap confidence intervals quantifying the 98\.8% direct effect\), anomaly detection \(identifying high\-fidelity failures and ceiling successes\), and domain\-transfer prediction \(leave\-one\-domain\-out cross\-validation\)\. Additional sections provide an extended five\-part related\-work discussion situating the study within the broader evaluation literature, supplementary code documentation, and a complete listing of repository contents\.
A central motivation for these supplementary materials is the study’s finding that human\-perceived response quality is fundamentally inaccessible to automated NLP metrics\. Because this claim challenges prevailing assumptions in the field, the supplement provides extensive construct validity evidence: 15 complementary machine\-learning and statistical analyses—spanning linear regression, gradient\-boosted trees, multilayer perceptrons,kk\-means clustering, PCA, UMAP, SHAP feature importance, mediation analysis, and domain\-transfer prediction—all converge on the same conclusion \(all cross\-validatedR2<0R^\{2\}<0; adjusted Rand index=−0\.05=\-0\.05; 98\.8% direct effect in mediation\)\. The complete data for these analyses, along with the evaluation rubric, all 62 prompts with expert\-validated reference answers, and the full analysis pipeline \(Python and R\), are available from the corresponding author during peer review and will be deposited at a Nature Portfolio repository under CC BY 4\.0 upon publication\.
### Per\-question fidelity profiles: Reasoning and Mathematics \(Fig\.[S1](https://arxiv.org/html/2607.20454#Ax1.F1)a, b\)
Supplementary Fig\.[S1](https://arxiv.org/html/2607.20454#Ax1.F1)a, b presents per\-question fidelity deviation profiles for the reasoning and mathematics domains, complementing the domain knowledge and safety profiles shown in main\-text Fig\. 3 a, b\. Each panel displays deviation for all 10 models on each question within the domain, ordered from best\- to worst\-performing model\.

\(a\)

\(b\)
Figure S1:Per\-question fidelity deviation — Reasoning and Mathematics domains\.a, Reasoning and language understanding domain \(Q1–Q10\)\. Per\-question fidelity deviation for all 10 models across 10 reasoning questions, arranged by domain\-specific mean deviation \(best to worst\)\. Each bar represents one question; dashed horizontal line shows the model’s domain mean; red triangles mark the hardest question and green stars the easiest for each model\. Gemini achieves the lowest deviation \(26\.3%, range 35\.6%\) followed by Claude \(40\.1%, range 61\.6%\)\. Ceiling models cluster between 80\.2% \(DeepSeek\) and 82\.9% \(ChatGPT\) with narrow within\-model ranges \(≤\\leq36\.4 pp\)\.b, Mathematical problem solving domain \(Q11–Q20\)\. Same format as panel a\. Gemini leads \(45\.0%, median 46\.5%\) and Claude is second \(45\.5%, median 46\.2%\); both show wide within\-model ranges \(\>\>38 pp\), indicating high sensitivity to mathematical question type\. Q11 is an outlier for Claude \(85\.6%\), suggesting a specific weakness in that question’s topic\. Ceiling models range from 75\.9% \(Llama\) to 79\.7% \(ChatGPT\) with within\-model ranges of 16–23 pp\.n=47n=47evaluators per cell\.
### Per\-question fidelity profiles: Conversational and Coding \(Fig\.[S2](https://arxiv.org/html/2607.20454#Ax1.F2)a, b\)
Supplementary Fig\.[S2](https://arxiv.org/html/2607.20454#Ax1.F2)a, b presents per\-question fidelity deviation profiles for the conversational and coding domains, completing the set of domain\-level views not shown in main\-text Fig\. 3 a, b\.

\(a\)

\(b\)
Figure S2:Per\-question fidelity deviation — Conversational and Coding domains\.a, Conversational AI and chatbot behaviour domain \(Q31–Q40\)\. Same format as Supplementary Fig\.[S1](https://arxiv.org/html/2607.20454#Ax1.F1)a\. Gemini achieves the lowest deviation \(37\.5%, median 32\.6%, range 47\.7%\); Claude is second \(55\.1%\)\. This domain shows the widest within\-model variability among ceiling models, suggesting that conversational prompts elicit inconsistent responses\.b, Coding and software development domain \(Q21–Q30\)\. Claude leads \(49\.5%, median 45\.8%, range 62\.6%\) with Gemini close behind \(50\.6%, median 52\.8%, range 79\.8%\)\. Gemini’s exceptionally wide range is driven by near\-perfect scores on Q29 \(8\.0%\) and Q30 \(13\.2%\) contrasting with poor performance on Q23 \(87\.8%\)\.n=47n=47evaluators per cell\.
### Question difficulty ranking \(Fig\.[S3](https://arxiv.org/html/2607.20454#Ax1.F3)\)
Supplementary Fig\.[S3](https://arxiv.org/html/2607.20454#Ax1.F3)ranks all 62 questions by mean fidelity deviation across all 10 models, providing a global difficulty profile\. Questions from the domain knowledge and coding domains dominate the hardest quartile, while safety and conversational questions are over\-represented among the easiest\.
Figure S3:All 62 questions ranked by difficulty across all 62 tasks\.Horizontal bars show mean fidelity deviation across all 10 models for each question, ranked from easiest \(top\) to hardest \(bottom\)\. Bars are colour\-coded by domain: green = Reasoning, blue = Mathematics, orange = Coding, purple = Conversational, pink = Safety, and blue = Domain Knowledge\. Error bars represent±1\\pm 1SD across models\. Dashed vertical reference lines at 60% and 80% mark difficulty thresholds\. The easiest question is Q43 \(Safety, 49\.2%\) and the hardest is Q60 \(Domain Knowledge, 89\.2%\)\. Domain knowledge and coding questions dominate the hardest quartile, while safety and conversational questions are overrepresented among the easiest\. The five most discriminating questions by inter\-model SD are Q7 \(30\.9 pp\), Q6 \(27\.3 pp\), Q10 \(27\.1 pp\), Q42 \(26\.9 pp\), and Q27 \(26\.3 pp\)\.
### Fidelity deviation distributions: Reasoning and Mathematics \(Fig\.[S4](https://arxiv.org/html/2607.20454#Ax1.F4)a, b\)
Supplementary Fig\.[S4](https://arxiv.org/html/2607.20454#Ax1.F4)a, b presents notched box plots showing the distributional properties of per\-question fidelity deviation for each model in the reasoning and mathematics domains\. These complement the per\-question bar charts in Fig\.[S1](https://arxiv.org/html/2607.20454#Ax1.F1)a, b by revealing distributional shape, spread, and outlier structure\. Diamond = mean; blue line = median; notch = 95% CI for median; red dots = outliers\. Horizontal dashed lines at 40%, 60%, and 80% mark performance tier boundaries\.

\(a\)

\(b\)
Figure S4:Fidelity deviation distributions — Reasoning and Mathematics domains\.a, Reasoning domain \(n=10n=10questions\)\. Gemini \(\#1,μ=26\.3%\\mu=26\.3\\%\) and Claude \(\#2,μ=40\.1%\\mu=40\.1\\%\) show wide distributions reflecting question\-sensitive performance\. Ceiling models cluster above 80% with narrow IQRs \(3\.3–9\.8 pp\)\.b, Mathematics domain \(n=10n=10questions\)\. Gemini and Claude achieve nearly identical means \(∼\\sim45%\) but with different distribution shapes\. Llama \(\#3,μ=75\.9%\\mu=75\.9\\%\) is the best\-performing ceiling model\.n=47n=47evaluators per bar\.
### Fidelity deviation distributions: Coding and Conversational \(Fig\.[S5](https://arxiv.org/html/2607.20454#Ax1.F5)a, b\)
Supplementary Fig\.[S5](https://arxiv.org/html/2607.20454#Ax1.F5)a, b presents notched box plots for the coding and conversational domains, using the same format as Fig\.[S4](https://arxiv.org/html/2607.20454#Ax1.F4)a, b\. These two domains show the widest within\-model variability among ceiling models, suggesting that coding and conversational prompts elicit particularly inconsistent responses\.

\(a\)

\(b\)
Figure S5:Fidelity deviation distributions — Coding and Conversational domains\.a, Coding domain \(n=10n=10questions\)\. Claude \(\#1,μ=49\.5%\\mu=49\.5\\%, IQR = 31\.8\) and Gemini \(\#2,μ=50\.6%\\mu=50\.6\\%, IQR = 57\.6\) lead with notably different variability patterns\. Gemini’s wide IQR reflects bimodal performance across coding tasks\.b, Conversational domain \(n=10n=10questions\)\. Gemini leads \(\#1,μ=37\.5%\\mu=37\.5\\%, IQR = 28\.4\)\. This domain shows the widest ceiling\-model spread, with all eight ceiling models having ranges exceeding 25 pp\.n=47n=47evaluators per bar\.
### Fidelity deviation distributions: Safety and Domain Knowledge \(Fig\.[S6](https://arxiv.org/html/2607.20454#Ax1.F6)a, b\)
Supplementary Fig\.[S6](https://arxiv.org/html/2607.20454#Ax1.F6)a, b completes the distributional analyses with the safety and domain knowledge domains, using the same format as Fig\.[S4](https://arxiv.org/html/2607.20454#Ax1.F4)a, b\. Safety shows the narrowest ceiling range of any domain, while domain knowledge is the most uniformly difficult for the ceiling cluster\.

\(a\)

\(b\)
Figure S6:Fidelity deviation distributions — Safety and Domain Knowledge\.a, Safety domain \(n=10n=10questions\)\. Claude \(\#1,μ=46\.1%\\mu=46\.1\\%, IQR = 13\.8\) and Gemini \(\#2,μ=47\.4%\\mu=47\.4\\%, IQR = 48\.7\) lead\. This domain shows the narrowest ceiling range \(71\.8–76\.1%\), suggesting safety prompts are uniformly challenging\.b, Domain knowledge \(n=12n=12questions\)\. Claude is the only model achieving sub\-50% deviation \(\#1,μ=45\.8%\\mu=45\.8\\%, IQR = 16\.4\)\. All other models cluster above 81% with extremely narrow distributions \(IQR 2\.7–5\.1 pp\), indicating this is the most uniformly difficult domain\.n=47n=47evaluators per bar\.
### Model documentation and participant characteristics \(Tables[S1](https://arxiv.org/html/2607.20454#Ax1.T1)–[S3](https://arxiv.org/html/2607.20454#Ax1.T3)\)
Supplementary Table[S1](https://arxiv.org/html/2607.20454#Ax1.T1)documents the version, organisation, release date, and context window size for each of the 10 evaluated models\. Supplementary Table[S2](https://arxiv.org/html/2607.20454#Ax1.T2)summarises participant demographics, including geographic region, professional background, LLM experience, age, and gender distribution\. All 47 participants completed all evaluations \(100% retention; 0% attrition\)\. Supplementary Table[S3](https://arxiv.org/html/2607.20454#Ax1.T3)presents an evaluator heterogeneity analysis, quantifying how consistently individual participants align with the reference answers\. The analysis reveals that Participant 3 is a persistent outlier \(least\-aligned in 68\.4% of model–question cells\), motivating the sensitivity analysis in Supplementary Note 1\.
Table S1:Model version documentation\.Study design, question development, and expert reference\-answer validation were conducted June–August 2025\. Pilot testing with early\-release models and monitoring of the evolving frontier model landscape took place September–November 2025\. The main evaluation period ran December 2025 through February 2026, during which all 10 models were available and accessed via official web user interface \(UI\) with default settings\. Data analysis and documentation were completed by late February through March 2026\. Models ordered by overall fidelity deviation rank \(lowest to highest\)\. Context window sizes are approximate and reflect publicly documented specifications at time of access\.Supplementary Table[S2](https://arxiv.org/html/2607.20454#Ax1.T2)summarises the participant demographics\.
Table S2:Participant demographic summary \(n=47n=47\)\.All participants completed all evaluations \(29,140 total; 0% attrition\)\. Recruitment began June 2025; pilot testing ran September–November 2025; main evaluations took place December 2025 through February 2026; analysis and documentation were completed by late February through March 2026\. IRB: University of North Texas, protocol \#IRB\-25\-297\.CharacteristicCategorynn%Geographic regionNorth America1838\.3Europe1531\.9Asia1225\.5Other24\.3Professional backgroundSoftware engineering1123\.4Data science817\.0Education612\.8Healthcare612\.8Legal510\.6Finance48\.5General professional714\.9LLM experienceDaily professional use \(7\+/week\)3676\.6Regular use \(2–5/week\)1123\.4Age range \(years\)25–292451\.130–442042\.545–4836\.4GenderMale2451\.1Female2348\.9Median age \(years\)29 \(range: 25–48\)Median evaluation time \(days\)30 \(range: 15–60\)Calibration pass rate100%Completion rate100% \(620/620\)Supplementary Table[S3](https://arxiv.org/html/2607.20454#Ax1.T3)quantifies evaluator\-level alignment patterns\.
Table S3:Evaluator heterogeneity analysis\.Summary of participant\-level alignment patterns across 620 model–question combinations \(n=47n=47evaluators\)\. Participant 3 is the least\-aligned evaluator in 68\.4% of cells, far exceeding any other participant\. Among best\-aligned evaluators, no participant dominates\. These patterns underscore the importance of the mixed\-effects model reported in Table 2 of the main text\.MeasureParticipantCount% of 620*Least\-aligned evaluator \(top 5\):*Participant 342468\.4Participant 346911\.1Participant 41345\.5Participant 5162\.6Participant 33162\.6*Most\-aligned evaluator \(top 5\):*Participant 46284\.5Participant 50274\.4Participant 9233\.7Participant 10223\.5Participant 19223\.5Unique least\-aligned evaluators23of 47Unique most\-aligned evaluators47of 47
### Complete fidelity deviation matrix \(Table[S4](https://arxiv.org/html/2607.20454#Ax1.T4)\)
Supplementary Table[S4](https://arxiv.org/html/2607.20454#Ax1.T4)presents the full 62\-question×\\times10\-model fidelity deviation matrix, reporting the mean deviation aggregated across all 47 participants for every question–model combination \(620 cells\)\. Bold values indicate the best\-performing model for each question\. Domain means and overall means are shown at the bottom\. This table provides the complete data underlying the summary statistics in main\-text Table 1 and the heatmap in Fig\. 2 a\.
Table S4:Complete 62\-question×\\times10\-model fidelity deviation matrix \(%\)\.Each cell: mean deviation×\\times100, aggregated across 47 participants\. Lower==better\. Bold==best model per question\.n=47n=47per cell\.QDomClaGemLlaMisGroPerCopDeeQweGPTMeanSDQ1R23\.021\.478\.979\.573\.580\.180\.179\.976\.783\.067\.624\.0Q2R26\.148\.377\.880\.086\.485\.379\.577\.978\.285\.672\.519\.6Q3R68\.641\.981\.384\.283\.184\.883\.283\.084\.485\.778\.013\.6Q4R47\.223\.868\.868\.369\.874\.269\.065\.674\.557\.561\.915\.7Q5R24\.513\.177\.878\.680\.279\.679\.676\.477\.182\.266\.925\.6Q6R25\.113\.484\.683\.980\.981\.385\.984\.783\.384\.270\.727\.3Q7R20\.312\.789\.489\.190\.590\.289\.788\.989\.890\.575\.130\.9Q8R50\.737\.673\.374\.274\.372\.477\.576\.677\.884\.569\.914\.3Q9R81\.930\.181\.680\.281\.479\.281\.277\.679\.382\.175\.516\.0Q10R33\.320\.590\.987\.687\.492\.391\.791\.388\.793\.977\.827\.0Q11M85\.625\.983\.282\.382\.085\.887\.383\.983\.282\.778\.218\.4Q12M50\.664\.175\.785\.782\.781\.884\.487\.285\.884\.078\.211\.9Q13M50\.760\.574\.972\.775\.875\.876\.575\.575\.177\.871\.58\.8Q14M23\.735\.872\.972\.975\.975\.076\.177\.375\.176\.566\.119\.4Q15M46\.748\.683\.681\.985\.383\.885\.386\.184\.288\.477\.415\.8Q16M45\.944\.378\.074\.082\.378\.982\.277\.880\.881\.472\.514\.7Q17M37\.753\.475\.477\.080\.576\.275\.377\.482\.081\.371\.614\.4Q18M32\.628\.474\.572\.779\.579\.177\.078\.678\.980\.768\.220\.0Q19M35\.431\.775\.682\.179\.482\.279\.578\.378\.878\.670\.219\.4Q20M46\.557\.864\.864\.662\.168\.967\.770\.369\.465\.163\.77\.1Q21C22\.318\.576\.581\.278\.285\.582\.784\.583\.278\.769\.125\.8Q22C84\.383\.180\.384\.883\.386\.185\.885\.584\.385\.584\.31\.7Q23C21\.787\.889\.387\.587\.887\.790\.488\.190\.892\.882\.421\.4Q24C35\.533\.583\.878\.779\.573\.671\.580\.081\.978\.869\.718\.9Q25C45\.544\.170\.268\.674\.176\.176\.074\.274\.074\.767\.812\.3Q26C63\.313\.275\.573\.382\.279\.676\.477\.281\.278\.470\.020\.7Q27C68\.48\.090\.090\.988\.491\.791\.490\.991\.190\.680\.126\.3Q28C72\.870\.178\.474\.476\.579\.277\.376\.376\.972\.975\.52\.8Q29C35\.261\.461\.766\.467\.068\.771\.768\.368\.364\.763\.310\.4Q30C46\.186\.088\.985\.688\.087\.490\.387\.688\.989\.183\.813\.3Q31Co32\.161\.178\.681\.283\.382\.683\.284\.382\.982\.775\.216\.6Q32Co52\.819\.571\.969\.877\.772\.875\.277\.576\.975\.466\.918\.2Q33Co77\.041\.180\.781\.075\.680\.882\.684\.781\.280\.376\.512\.7Q34Co33\.158\.378\.272\.478\.676\.377\.776\.275\.578\.870\.514\.5Q35Co43\.724\.283\.285\.987\.184\.090\.092\.286\.386\.976\.423\.0Q36Co73\.053\.557\.571\.468\.173\.171\.273\.075\.271\.768\.87\.3Q37Co73\.459\.184\.885\.584\.486\.087\.188\.087\.086\.082\.19\.1Q38Co66\.823\.589\.390\.194\.186\.292\.397\.094\.091\.082\.422\.3Q39Co42\.921\.761\.747\.559\.760\.963\.060\.559\.966\.554\.413\.6Q40Co56\.413\.470\.569\.671\.372\.672\.570\.869\.571\.963\.818\.4Q41S53\.318\.077\.477\.776\.376\.479\.980\.978\.880\.970\.020\.0Q42S26\.513\.882\.982\.782\.181\.783\.784\.584\.984\.770\.826\.9Q43S33\.418\.354\.855\.255\.063\.443\.255\.354\.958\.649\.213\.8Q44S41\.542\.179\.582\.384\.181\.283\.484\.783\.683\.274\.617\.3Q45S38\.337\.081\.683\.678\.183\.783\.386\.283\.786\.274\.219\.4Q46S51\.754\.866\.071\.275\.568\.972\.168\.476\.271\.167\.68\.2Q47S53\.771\.968\.671\.268\.174\.170\.674\.074\.474\.370\.16\.2Q48S61\.971\.072\.370\.365\.973\.471\.675\.171\.569\.270\.23\.8Q49S49\.473\.768\.074\.475\.570\.071\.175\.672\.671\.770\.27\.7Q50S51\.372\.967\.276\.072\.774\.970\.876\.077\.972\.271\.27\.6Q51DK35\.282\.077\.580\.280\.478\.982\.082\.182\.182\.976\.314\.5Q52DK35\.783\.382\.081\.383\.479\.783\.981\.785\.083\.177\.914\.9Q53DK40\.588\.286\.988\.887\.683\.886\.990\.087\.290\.283\.015\.0Q54DK61\.678\.480\.477\.781\.079\.583\.082\.383\.081\.978\.96\.3Q55DK22\.182\.781\.985\.684\.182\.983\.887\.185\.385\.578\.119\.8Q56DK37\.381\.676\.779\.680\.280\.680\.283\.982\.880\.576\.313\.8Q57DK55\.584\.281\.685\.483\.183\.285\.787\.287\.783\.481\.79\.4Q58DK51\.077\.177\.378\.381\.481\.081\.581\.880\.880\.177\.09\.3Q59DK35\.882\.583\.480\.981\.785\.284\.185\.188\.284\.979\.215\.4Q60DK88\.387\.387\.689\.686\.289\.689\.992\.592\.289\.189\.22\.0Q61DK38\.279\.780\.080\.280\.180\.180\.878\.080\.281\.375\.913\.3Q62DK48\.985\.284\.685\.783\.384\.386\.886\.885\.685\.681\.711\.6Reasoning40\.126\.380\.480\.680\.781\.981\.780\.281\.082\.971\.620\.5Mathematics45\.545\.175\.976\.678\.578\.879\.179\.379\.379\.671\.814\.0Coding49\.550\.679\.579\.180\.581\.681\.381\.382\.080\.674\.613\.0Conversational55\.137\.575\.675\.578\.077\.579\.580\.478\.979\.171\.714\.0Safety46\.147\.471\.874\.573\.374\.873\.076\.175\.975\.268\.811\.7Domain K\.45\.982\.781\.782\.882\.782\.484\.184\.985\.084\.079\.611\.9Overall47\.049\.477\.678\.379\.179\.679\.980\.580\.580\.473\.113\.3
### Statistical analyses: variance, equivalence, effect sizes, and rank stability \(Tables[S6](https://arxiv.org/html/2607.20454#Ax1.T6)–[S9](https://arxiv.org/html/2607.20454#Ax1.T9)\)
Supplementary Tables[S6](https://arxiv.org/html/2607.20454#Ax1.T6)–S8 report the detailed statistical analyses supporting the main\-text findings\. Table[S6](https://arxiv.org/html/2607.20454#Ax1.T6)decomposes variance by domain using two\-way ANOVA, showing that the model factor explains the largest proportion of variance in all six domains\. Table[S7](https://arxiv.org/html/2607.20454#Ax1.T7)presents TOST equivalence tests confirming that the eight ceiling models are statistically indistinguishable within a±5\\pm 5pp margin\. Table[S8](https://arxiv.org/html/2607.20454#Ax1.T8)reports all 45 pairwise Cohen’sddeffect sizes, demonstrating the magnitude of separation between high\-fidelity and ceiling tiers\. Table[S9](https://arxiv.org/html/2607.20454#Ax1.T9)presents bootstrap rank distributions \(10,000 resamples\), confirming that Claude and Gemini occupy ranks 1–2 in 100% of resamples while within\-ceiling ranks are unstable\.
Table S6:Variance decomposition by domain\.Separate two\-way ANOVAs per domain\. Reasoning and Domain Knowledge show strongest model effects \(η2\>0\.77\\eta^\{2\}\>0\.77\); Coding shows weakest \(η2=0\.464\\eta^\{2\}=0\.464\)\.Supplementary Table[S7](https://arxiv.org/html/2607.20454#Ax1.T7)reports ceiling homogeneity tests confirming that the eight ceiling models are statistically indistinguishable\.
Table S7:Ceiling homogeneity tests\.Applied to the 8\-model ceiling subset \(496 observations\)\. Both tests confirm ceiling models are not significantly differentiable\.Supplementary Table[S8](https://arxiv.org/html/2607.20454#Ax1.T8)reports all 45 pairwise effect sizes, showing clear separation between tiers \(d\>1\.4d\>1\.4\) and negligible within\-ceiling differences \(d<0\.4d<0\.4\)\.
Table S8:Pairwise Cohen’sdd\(all 45 pairs\)\.All high\-fidelity\-vs\-ceiling pairsd\>1\.4d\>1\.4; all within\-ceiling pairsd<0\.4d<0\.4\.Supplementary Table[S9](https://arxiv.org/html/2607.20454#Ax1.T9)presents bootstrap rank stability analysis\.
Table S9:Bootstrap rank distributions \(10,000 resamples\)\.Claude and Gemini fixed at ranks 1–2; ceiling ranks unstable at positions 7–10\.
### Automated NLP metrics: similarity, correlations, and domain breakdown \(Tables[S10](https://arxiv.org/html/2607.20454#Ax1.T10)–[S12](https://arxiv.org/html/2607.20454#Ax1.T12)\)
Supplementary Tables[S10](https://arxiv.org/html/2607.20454#Ax1.T10)–S11 document the automated NLP pipeline results\. Table[S10](https://arxiv.org/html/2607.20454#Ax1.T10)reports seven NLP similarity metrics \(semantic similarity, ROUGE\-1/2/L, BLEU, length ratio, and overall composite\) for each model, showing that automated metrics rank models differently from human evaluators\. Table[S11](https://arxiv.org/html/2607.20454#Ax1.T11)presents correlations between each NLP metric and human\-judged fidelity deviation, revealing weak and often non\-significant associations \(the strongest is semantic similarity atr=0\.31r=0\.31\)\. Table[S12](https://arxiv.org/html/2607.20454#Ax1.T12)breaks down NLP composite similarity by domain, highlighting domain\-dependent discrepancies between automated and human assessments\.
Table S10:NLP similarity by model\.Scale 0–1 \(higher==more similar\)\. Bold==best per column\. Fidelity deviation shown for comparison\.Supplementary Table[S11](https://arxiv.org/html/2607.20454#Ax1.T11)reports correlations between NLP metrics and human\-judged fidelity deviation\.
Table S11:NLP metric correlations with fidelity deviation\(620 cells\)\. Weak positiverrmeans higher NLP similarity associates with*worse*human ratings\.Supplementary Table[S12](https://arxiv.org/html/2607.20454#Ax1.T12)breaks down NLP similarity by domain\.
Table S12:NLP overall similarity \(%\) by domain\.Unlike human\-judged fidelity, NLP scores are uniform across models\. Claude–ChatGPT gap: 1\.2 pp \(vs\. 33\.5 pp in human evaluation\)\.
## Supplementary Note 1: Sensitivity Analyses
Sensitivity analyses excluding Participant 3 \(the most extreme evaluator\) confirmed all primary findings\. After exclusion \(n=46n=46, 28,520 evaluations\): \(1\) the bimodal distribution persisted with Claude at 46\.6% and Gemini at 49\.0%; \(2\) ceiling model means shifted by≤\\leq0\.5 pp; \(3\) TOST equivalence atΔ=5\\Delta=5pp was confirmed for all 28 ceiling pairs; \(4\) between\-tier Cohen’sddremained\>\>1\.8\. A leave\-one\-evaluator\-out analysis across all 47 participants showed maximum shift in any model mean of 0\.8 pp, confirming robustness\.
## Supplementary Discussion: Extended Related Work
### S1\. Automated benchmarks and their limitations
The evaluation of large language models has historically relied on automated benchmark suites that measure narrow task\-specific capabilities\. MMLU\[[15](https://arxiv.org/html/2607.20454#bib.bib4)\]provides broad coverage across 57 academic subjects but uses multiple\-choice questions that compress open\-ended generation capability into a single correct answer\. As frontier models now routinely exceed 90% accuracy on MMLU\[[40](https://arxiv.org/html/2607.20454#bib.bib8)\], the benchmark has lost discriminative power among leading systems—a phenomenon known as benchmark saturation\. MATH\[[33](https://arxiv.org/html/2607.20454#bib.bib5)\]and HumanEval\[[10](https://arxiv.org/html/2607.20454#bib.bib6)\]evaluate mathematical reasoning and code generation respectively, but address single capability domains\. More recent efforts such as the Humanity’s Last Exam \(HLE\) benchmark\[[32](https://arxiv.org/html/2607.20454#bib.bib31)\]attempt to push beyond saturation with expert\-crafted questions, yet remain automated and binary in their scoring\. A common limitation is training data contamination\[[1](https://arxiv.org/html/2607.20454#bib.bib9)\]: models may have been exposed to benchmark items during pretraining, inflating reported performance\. Our study addresses these gaps by evaluating open\-ended responses across six domains using graduated human judgements anchored to expert references, providing a measure that is resistant to both saturation and contamination effects\.
### S2\. Holistic evaluation frameworks
Several frameworks attempt to provide multi\-dimensional model assessment\. HELM\[[5](https://arxiv.org/html/2607.20454#bib.bib10)\]evaluates models across multiple scenarios and metrics but relies on automated scoring\. Chatbot Arena\[[8](https://arxiv.org/html/2607.20454#bib.bib11)\]introduced large\-scale human evaluation through pairwise preference judgements, yielding Elo\-based rankings from hundreds of thousands of votes\. However, preference\-based paradigms measure relative quality \(“which response do you prefer?”\) rather than absolute fidelity \(“how faithfully does this response preserve expert content?”\)\. This distinction is critical: preference judgements are influenced by fluency, formatting, and response length, which may not correlate with content accuracy\. AlpacaEval\[[11](https://arxiv.org/html/2607.20454#bib.bib12)\]further demonstrated that LLM\-based judges can serve as scalable proxies for human preference, but acknowledged that such judges exhibit systematic biases including length preference and style sensitivity\. MT\-Bench\[[41](https://arxiv.org/html/2607.20454#bib.bib7)\]advanced the LLM\-as\-judge paradigm for multi\-turn evaluation but inherits similar limitations\. Our fidelity deviation metric complements these approaches by providing reference\-anchored absolute measurement, enabling variance decomposition and per\-question correlation analyses that pairwise designs cannot support\.
### S3\. Factual accuracy and hallucination assessment
A growing body of work addresses the specific problem of factual accuracy in model outputs\. FActScore\[[28](https://arxiv.org/html/2607.20454#bib.bib13)\]introduced fine\-grained factual precision scoring by decomposing generated text into atomic claims and verifying each against a knowledge source\. This approach provides granular accuracy measurement but is typically applied to specific generation tasks \(e\.g\., biographical descriptions\) rather than comprehensive multi\-domain evaluation\. Research on hallucination\[[19](https://arxiv.org/html/2607.20454#bib.bib33)\]has documented systematic patterns of factual fabrication across model families, and retrieval\-augmented generation \(RAG\)\[[23](https://arxiv.org/html/2607.20454#bib.bib41)\]has been proposed as a mitigation strategy\. Our study contributes to this literature by demonstrating that human evaluators perceive substantial fidelity differences that automated similarity metrics compress by up to 28\-fold, suggesting that current automated factuality measures may underestimate the magnitude of inter\-model differences in knowledge\-intensive tasks\.
### S4\. Scaling laws and diminishing returns
The relationship between model scale and performance has been characterised by neural scaling laws\[[7](https://arxiv.org/html/2607.20454#bib.bib2),[16](https://arxiv.org/html/2607.20454#bib.bib3)\]predicting log\-linear improvement with compute, data, and parameter count\. Foundation models\[[4](https://arxiv.org/html/2607.20454#bib.bib34)\]have grown from 175B parameters \(GPT\-3\[[6](https://arxiv.org/html/2607.20454#bib.bib1)\]\) through sparse mixture\-of\-experts architectures\[[13](https://arxiv.org/html/2607.20454#bib.bib39)\]to the current generation of frontier systems, with successive models demonstrating consistent \(if diminishing\) improvements on automated benchmarks\. However, the fidelity ceiling we observe—eight independently developed frontier models converging on a 2\.9 pp band—is consistent with the “densing law” proposed by Xiao et al\.\[[39](https://arxiv.org/html/2607.20454#bib.bib27)\], which suggests that the capability\-per\-parameter ratio may approach fundamental limits under current training paradigms\. The convergence of models trained by different organisations using different architectures and data compositions strengthens the interpretation that shared training methodologies \(large\-scale pretraining followed by instruction tuning and reinforcement learning from human feedback \(RLHF\)\[[31](https://arxiv.org/html/2607.20454#bib.bib35),[2](https://arxiv.org/html/2607.20454#bib.bib48),[38](https://arxiv.org/html/2607.20454#bib.bib40)\]\) may produce convergent limitations in human\-perceived response fidelity, even as automated benchmark scores continue to differentiate\.
### S5\. Human evaluation methodology
The design of rigorous human evaluation protocols for generative AI systems remains an active area of methodological development\. Likert\-scale measurement\[[25](https://arxiv.org/html/2607.20454#bib.bib23)\]provides ordinal data amenable to parametric analysis when aggregated, and has been widely used in natural language generation evaluation\. Inter\-rater reliability assessment via intraclass correlation coefficients\[[36](https://arxiv.org/html/2607.20454#bib.bib20),[20](https://arxiv.org/html/2607.20454#bib.bib26)\]and Krippendorff’sα\\alpha\[[21](https://arxiv.org/html/2607.20454#bib.bib21)\]provides standardised measures of agreement\. Recent work has highlighted important challenges in human evaluation of LLMs, including evaluator subjectivity\[[12](https://arxiv.org/html/2607.20454#bib.bib43)\], the influence of cognitive biases on preference judgements\[[37](https://arxiv.org/html/2607.20454#bib.bib15)\], and the need for calibration protocols that ensure consistent application of rating criteria across evaluators\. Mahowald et al\.\[[27](https://arxiv.org/html/2607.20454#bib.bib14)\]argued that the distinction between formal linguistic competence and functional competence is critical for evaluation design—a perspective that aligns with our focus on content fidelity rather than surface fluency\. Our fully crossed design, in which every evaluator assesses every model on every question, enables simultaneous estimation of model, question, and evaluator variance components—a methodological advantage over incomplete\-block designs that has been advocated in the emerging literature on LLM evaluation best practices\[[30](https://arxiv.org/html/2607.20454#bib.bib32)\]\.
## Supplementary Worked Examples
Supplementary Table[S13](https://arxiv.org/html/2607.20454#Ax4.T13)presents five worked examples spanning both performance tiers and multiple domains, illustrating how raw Likert ratings translate to fidelity deviation values\. The implied mean Likert rating is computed asR¯=5−4×Deviation\\bar\{R\}=5\-4\\times\\mathrm\{Deviation\}\. These examples verify the deviation computation pipeline against specific data points in the released dataset\.
Table S13:Worked examples linking implied participant ratings to fidelity deviation\.All values verified against the released dataset\. ImpliedR¯\\bar\{R\}is the mean Likert rating across 47 participants that produces the observed deviation\.n=47n=47evaluators per cell\.Example 1: Q1/Claude \(23\.0%\)\.The 47 participants rated Claude’s response to Q1 \(a reasoning question\) with a mean Likert score of approximately 4\.08 out of 5\. Applying the deviation formula:Deviation=\(5−4\.08\)/4=0\.230\\mathrm\{Deviation\}=\(5\-4\.08\)/4=0\.230, or 23\.0%\. Participants judged Claude’s response to be largely faithful to the expert reference, with minor deviations in completeness or phrasing\.
Example 2: Q1/ChatGPT \(83\.0%\)\.For the same question, ChatGPT received a mean rating of approximately 1\.68 out of 5\. Deviation=\(5−1\.68\)/4=0\.830=\(5\-1\.68\)/4=0\.830, or 83\.0%\. The 60\.0 pp gap between Claude and ChatGPT on this single question illustrates the magnitude of the high\-fidelity\-versus\-ceiling contrast at the individual question level\.
Example 3: Q38/ChatGPT \(91\.0%\)\.ChatGPT received a mean rating of approximately 1\.36 out of 5 on Q38 \(a conversational question\)\. Deviation=\(5−1\.36\)/4=0\.910=\(5\-1\.36\)/4=0\.910, or 91\.0%\. This represents one of the highest single\-cell deviation values in the dataset, indicating near\-unanimous evaluator agreement that ChatGPT’s response diverged substantially from the expert reference on this particular question\.
Example 4: Q38/Gemini \(23\.5%\)\.On the same question where ChatGPT shows 91\.0% deviation, Gemini achieves only 23\.5% \(impliedR¯=4\.06\\bar\{R\}=4\.06\)\. This 67\.5 pp within\-question gap is among the largest in the dataset, demonstrating that question difficulty is strongly model\-dependent rather than an intrinsic property of the question\.
Example 5: Q60/Claude \(88\.3%\)\.Even the top\-ranked model achieves high deviation on certain questions\. Q60 \(domain knowledge\) is the hardest question overall \(89\.2% mean across all 10 models\), and Claude’s 88\.3% is near the cross\-model mean\. High\-fidelity status does not confer immunity to question\-specific difficulty; it reflects an overall distributional advantage rather than uniform superiority across all items\.
## Supplementary Note: NLP Metrics vs\. Fidelity Deviation
Supplementary Table[S14](https://arxiv.org/html/2607.20454#Ax5.T14)compares the two independent measurement pipelines \(human fidelity deviation and automated NLP similarity\), confirming their methodological independence and quantifying the 28\-fold compression of model differences in automated metrics\.
Table S14:Comparison of the two independent measurement pipelines\.Human fidelity deviation \(primary endpoint\) and automated NLP similarity \(secondary validation\) are computed through entirely separate pipelines with no shared computational inputs, confirming their methodological independence\.The weak positive correlation \(r=0\.13r=0\.13,r2<0\.02r^\{2\}<0\.02\) between NLP similarity and fidelity deviation confirms that automated metrics capture fundamentally different quality dimensions than human evaluators\. The 28\-fold compression of model differences—from 33\.5 pp in human fidelity deviation \(Claude–ChatGPT gap\) to 1\.2 pp in automated NLP similarity—demonstrates that current automated metrics are insensitive to the quality dimensions that matter most to trained evaluators assessing open\-ended responses against expert references\. Complete domain\-level NLP similarity analyses are reported in Supplementary Tables[S10](https://arxiv.org/html/2607.20454#Ax1.T10),[S11](https://arxiv.org/html/2607.20454#Ax1.T11), and[S12](https://arxiv.org/html/2607.20454#Ax1.T12)\.
## Supplementary Construct Validity Analyses
This section reports the 15 machine learning and statistical analyses that establish construct validity for the human–automated gap \(see main\-text Results §5 and Methods\)\. Supplementary Table[S15](https://arxiv.org/html/2607.20454#Ax6.T15)summarises the cross\-validated prediction results: all regression models yield negativeR2R^\{2\}, confirming that no combination of automated NLP features can predict human\-judged fidelity deviation\. Supplementary Table[S16](https://arxiv.org/html/2607.20454#Ax6.T16)decomposes the contribution of content versus style features, showing that content explains 59\.2% of the \(negligible\) automated variance\. Supplementary Table[S17](https://arxiv.org/html/2607.20454#Ax6.T17)presents evaluator clustering profiles, confirming that the main evaluator body is internally consistent \(43 typical, 3 low\-alignment, 1 extreme outlier\)\. Extended analyses — mediation \(Supplementary Note 2\), anomaly detection \(Supplementary Note 3\), and domain transfer prediction \(Supplementary Note 4\) — follow below\.
Table S15:Machine learning prediction of fidelity deviation from NLP features\.5\-fold cross\-validatedR2R^\{2\}for regression and accuracy for tier classification\. NegativeR2R^\{2\}indicates the model performs worse than a constant\-mean baseline\. Content features: semantic similarity\. Style features: token\-length ratio, character\-length ratio, POS alignment\.n=620n=620cells \(10 models×\\times62 questions\)\.Supplementary Table[S16](https://arxiv.org/html/2607.20454#Ax6.T16)decomposes content versus style contributions to the automated signal\.
Table S16:Content vs\. style regression decomposition\.PartialR2R^\{2\}quantifies each feature group’s unique contribution to explaining fidelity deviation after controlling for the other group\. Content features explain 59\.2% of the total explainable variance\.n=620n=620cells\.Supplementary Table[S17](https://arxiv.org/html/2607.20454#Ax6.T17)presents evaluator clustering profiles\.
Table S17:Evaluator clustering profiles\.kk\-means clustering \(k=3k=3\) on evaluator best/worst alignment frequency vectors identifies three natural groups\. Evaluator 3 forms a singleton cluster as the most extreme outlier, consistent with the main\-text sensitivity analysis\.n=47n=47evaluators\.### Supplementary Note: Mediation analysis details
We tested whether automated NLP features mediate the relationship between model tier \(high\-fidelity vs\. ceiling\) and human fidelity deviation using the Baron–Kenny framework with 5,000 bootstrap resamples\.
*Content mediation \(Tier→\\toSemantic similarity→\\toDrift\)\.*The total effect of tier on drift wasβ=−0\.306\\beta=\-0\.306\. Theaa\-path \(tier→\\tosemantic similarity\) wasβa=\+0\.033\\beta\_\{a\}=\+0\.033, and thebb\-path \(semantic→\\todrift, controlling for tier\) wasβb=−0\.114\\beta\_\{b\}=\-0\.114\. The indirect \(mediated\) effect wasa×b=−0\.004a\\times b=\-0\.004\(bootstrap 95% CI \[−0\.007\-0\.007,−0\.001\-0\.001\];p<0\.05p<0\.05\), representing 1\.2% of the total effect\. The direct effect \(tier→\\todrift, controlling for semantic\) was−0\.303\-0\.303, accounting for 98\.8% of the total\.
*Style mediation \(Tier→\\toLength→\\toDrift\)\.*The indirect effect through token\-length ratio was approximately zero \(a×b≈0\.000a\\times b\\approx 0\.000\), and the content\-to\-style mediation ratio was 3\.8×\\times\.
*Interpretation\.*The negligible mediation by both content and style NLP features confirms that the tier difference in human fidelity judgements is not transmitted through any automated\-measurable pathway\. Human evaluators perceive a quality dimension that is orthogonal to the entire NLP feature space\.
### Supplementary Note: Anomaly detection
We identified model–question cells where the tier prediction fails: high\-fidelity models with drift in the worst quartile \(10 cells\) and ceiling models with drift in the best quartile \(55 cells\)\.
High\-fidelity anomalies were concentrated in Domain Knowledge \(7/10 cells, predominantly Gemini\), consistent with Gemini’s known domain\-knowledge weakness \(82\.7% deviation despite 49\.4% overall\)\. The remaining anomalies were in Conversational \(2 cells\) and Mathematics \(1 cell, Claude Q11 at 85\.6%\)\.
Ceiling\-model successes were distributed across Reasoning \(16 cells\), Coding \(16\), and Mathematics \(15\), with no single model accounting for more than 8 cells\. These represent question\-specific strengths rather than systematic capability, consistent with the high inter\-model correlations \(r\>0\.85r\>0\.85\) reported in the main text\.
### Supplementary Note: Domain transfer prediction
Leave\-one\-domain\-out cross\-validation tested whether NLP features trained on five domains could predict fidelity deviation in the held\-out sixth domain \(random forest, 100 estimators\)\. Cross\-validatedR2R^\{2\}was near zero or negative for most domains \(Coding:−0\.04\-0\.04; Conversational:\+0\.03\+0\.03; Domain Knowledge:−0\.19\-0\.19; Mathematics:\+0\.05\+0\.05; Reasoning:−0\.19\-0\.19; Safety:\+0\.26\+0\.26\), confirming that NLP features do not generalise across domains for drift prediction\. The Safety domain was the only positive outlier \(R2=0\.26R^\{2\}=0\.26\), possibly because safety\-related responses exhibit more detectable surface\-level patterns \(e\.g\., refusal language, hedging\)\. Tier classification transferred more robustly \(accuracy 76–86% across domains\), suggesting the binary tier structure is more domain\-general than the continuous drift values\.
## Supplementary Code
The complete analysis pipeline—including the fidelity deviation computation \(Python 3\.11; NumPy 1\.26, SciPy 1\.12, pandas 2\.1\), all statistical tests, bootstrap rank analysis, and the linear mixed\-effects model \(R 4\.3\.2; lme4\[[3](https://arxiv.org/html/2607.20454#bib.bib19)\]\)—is deposited in the study repository at the Nature Portfolio repository \(DOI to be assigned upon publication\) under CC BY 4\.0\. The pipeline comprises:
- •deviation\_computation\.py— Per\-evaluation, model–question, model\-level, and domain\-level fidelity deviation aggregation \(Equations 1–3\)\.
- •statistical\_analysis\.py— Two\-way ANOVA, Kruskal–Wallis tests, bootstrap rank analysis \(10,000 resamples\), pairwise Cohen’sdd\(45 pairs\), TOST equivalence tests \(28 ceiling pairs\), and inter\-model correlation matrices\.
- •mixed\_effects\_model\.R— Linear mixed\-effects model \(Equation 4\), ICC extraction, and Nakagawa–SchielzethR2R^\{2\}computation\.
Estimated computation time for full reproduction:<<30 minutes on a standard laptop \(Intel Core i7 or equivalent, 16 GB RAM; no GPU required\)\. All analyses used fixed random seeds for exact reproducibility\.
## Supplementary Software Versions
All analyses were conducted using the software versions listed below\. Version pinning ensures full reproducibility of all reported statistics\.
Table S18:Software versions used in all analyses\.Python packages were managed viapipwith a frozenrequirements\.txt; R packages were installed from CRAN snapshots dated February 2026\.
## Supplementary Repository Contents
The following materials are deposited at the Nature Portfolio repository \(DOI to be assigned upon publication\) under CC BY 4\.0 licence\. During peer review, all materials listed below are available from the corresponding author upon reasonable request:
1. 1\.participant\_ratings\.csv— Anonymised 29,140\-cell participant\-level rating matrix \(47×10×6247\\times 10\\times 62\), containing raw Likert scores and computed fidelity deviation values for every evaluation\.
2. 2\.deviation\_computation\.py— Executable Python pipeline reproducing all fidelity deviation values from raw ratings, including per\-evaluation, model–question, model\-level, and domain\-level aggregation\.
3. 3\.statistical\_analysis\.py— Scripts reproducing all reported statistics: two\-way ANOVA, Kruskal–Wallis tests, bootstrap rank analysis \(10,000 resamples\), pairwise Cohen’sdd\(45 pairs\), TOST equivalence tests \(28 ceiling pairs\), inter\-model correlation matrices, and variance decomposition\.
4. 4\.mixed\_effects\_model\.R— R script for the linear mixed\-effects model \(Equation 4\), ICC extraction, and Nakagawa–SchielzethR2R^\{2\}computation\.
5. 5\.questions\_and\_references\.json— All 62 prompts with system messages and expert\-validated reference answers, organised by domain\.
6. 6\.evaluation\_rubric\.pdf— Complete 5\-point Likert rubric with anchor definitions, calibration protocol, and worked examples used during evaluator training\.
7. 7\.model\_versions\.csv— Model identifiers, organisations, access method \(web UI\), access dates \(December 2025–February 2026\), and context window sizes\.
8. 8\.nlp\_similarity\_scores\.csv— Automated NLP similarity metrics \(semantic similarity, lexical overlap, POS alignment, sentiment agreement, length ratios\) for all 620 model–question cells\.
9. 9\.— High\-resolution versions of all main and supplementary figures in PNG format\.Similar Articles
Some Large Language Models Exhibit Consistent Risk Attitudes
This paper introduces a framework to test whether large language models exhibit consistent risk attitudes across domains. It finds that most LLMs show intra-task and cross-domain stability in risk attitude, converging to a narrower distribution than humans.
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.
Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty
This paper investigates how similar large language model uncertainty is to human uncertainty, exploring alignment, calibration, and activation patterns in LLMs across multiple datasets and the impact of instruction fine-tuning.
Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
This paper introduces a controlled protocol to evaluate answer stability in large language models by challenging correct answers with plausible counterarguments, revealing large variation in flip rates across models that accuracy metrics alone do not capture. The authors release the protocol, challenge records, and a curated MaxFlip challenge set to support stability evaluation.
Confidence Calibration in Large Language Models
This paper analyzes the confidence calibration of 11 popular LLMs, finding that they are generally overconfident, especially on hard tasks, and underconfident on easy tasks. It introduces LifeEval, a test for evaluating calibration across difficulty levels.