How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks (1 minute read)

TLDR AI Papers

Summary

Expert re-grading of physics benchmarks reveals that frontier AI models perform better than previously evaluated, indicating broken assessments and near-saturation on closed-ended tasks, which underscores the need for more rigorous evaluations.

Those scary physics scores were often the test's fault. Experts rechecked six popular benchmarks and found wrong answer keys, fuzzy questions, and grader bugs behind most model “fails.” Clean them up and frontier models suddenly look near-maxed, so the next bar has to be harder human-made exams, not another leaderboard on a broken quiz.
Original Article
View Cached Full Text

Cached at: 09/14/26, 02:19 PM

# How Good Are Frontier Models at Physics?Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Source: [https://arxiv.org/html/2609.13009](https://arxiv.org/html/2609.13009)
Ali AnsariHaoran SunAndy Zeyi LiuMark JabbourYongshan DingSteven GirvinYu HeSohrab Ismail\-Beigi Aleksander KubicaOwen D\. MillerCorey O’HernVidvuds Ozolins David PolandA\. Douglas StoneFrank C\. van den BoschLogan WrightNavid AkbariSantanu AntuKangle CaiAndrew Calabrese\-DayMateo Cárdenes WuttigMeng ChengBarry T\. ChiangAli GhorashiShouzhen GuHaoyang HuangZhibo KangLukas Kienesberger Hantian LiuCharles LombaZhongling LuWenchao Ma Rohin E\. McIntoshEvan McKinneyIvan RojkovXulei Sun Yarone Meir TokayerNaveen Balaji UmasankarMira Varma Leda WangQimin WangTyler WangHaoyu WeiJinming Yang Jinchen ZhaoSherlock Tingrui ZhaoQinyuan ZhengJay S\. ZouLucas BakerArman CohanJohn SousYale University Jump Trading Group University of Cambridge University of Southern California†\\daggerCore contributors\.α\\alphaPhysics advisors\.β\\betaData auditors\. See the[Author Contributions statement](https://arxiv.org/html/2609.13009#A7)for details\.Correspondence:john\.sous@yale\.edu

###### Abstract

Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index \(2026\), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem\-solving abilities\. Yet this impression does not always align with domain experts’ experiences using these models in their work\. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text\-only problems with verifiable final answers\. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions\. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models’ physics reasoning\. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions\. We find that GPT\-5\.6\-Sol’s measuredmean@4rises from47\.347\.3% to78\.778\.7% on HLE\-Physics and from61\.061\.0% to87\.287\.2% on CMT\-Benchmark, while its correctedpass@4reaches94\.494\.4% on the 54 retained CritPt challenges\. Corrected scores are computed on the retained evaluation subsets following expert review\. Scores on the audited subsets of UGPhysics, PRISM\-Physics, and PHYBench also rise substantially after correction\. These findings suggest that current benchmarks substantially understate frontier models’ ability to solve well\-posed physics problems\. Near\-saturation on these closed\-ended tasks highlights the need for more demanding, expert\-validated evaluations\.

Figure 1:Pre\-audit and validated/repaired performance on six physics benchmarksfor GPT\-5\.6\-Sol, Fable 5, and Gemini 3\.1 Pro\. All scores usemean@4, except pre\-audit CritPt scores, which use Artificial Analysis’smean@5\. Light and solid bars show pre\-audit and validated/repaired performance, respectively\. Validated/repaired performance is measured after correcting evaluation errors and repairing or excluding flawed questions\.## 1Introduction

Large language models \(LLMs\) have made rapid progress in mathematical reasoning\. A substantial body of work has focused on assessing these reasoning abilities through mathematics benchmarks\. As model performance on these benchmarks has improved, evaluation efforts have shifted toward physics, a foundational domain for testing LLMs’ scientific reasoning and problem\-solving capabilities\. Over the past two years, numerous physics benchmarks have emerged, ranging from undergraduate exercises to expert\-curated challenges\([Xu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib1);[Qiu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib2);[Feng et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib8);[Center for AI Safety et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib10);[Pan et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib4);[Zhu et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib5)\)\. These benchmarks commonly report that even frontier models perform substantially worse than expert physicists\. For example, GPT\-5\.6\-Sol scores amean@5of32\.332\.3% on CritPt\([Zhu et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib5)\)at Max reasoning effort and amean@4of47\.347\.3% at High reasoning effort on the physics component of Humanity’s Last Exam \(HLE\)\([Center for AI Safety et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib10)\), which we henceforth call HLE\-Physics\. Because physics combines modeling, mathematics, and computation, these scores suggest limitations that extend to other quantitative applications in finance and technology\.

These conclusions, however, warrant closer scrutiny in light of recent advances in AI\-assisted mathematical research on open problems\. Frontier models and agents built around them have contributed to new proofs and solutions to open research questions\([Bubeck et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib13);[Sothanaphan, 2026](https://arxiv.org/html/2609.13009#bib.bib14);[Feng et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib15)\), including the resolution of longstanding conjectures and substantial improvements to established bounds\([Alon et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib16);[Tao, 2026](https://arxiv.org/html/2609.13009#bib.bib17);[OpenAI, 2026c](https://arxiv.org/html/2609.13009#bib.bib18);[Alpöge and Furman, 2026](https://arxiv.org/html/2609.13009#bib.bib19);[OpenAI, 2026a](https://arxiv.org/html/2609.13009#bib.bib20)\)\. Such achievements do not on their own rule out a genuine gap between mathematical and physical reasoning, since physics requires skills beyond mathematics, such as modeling, abstraction, and deciding on appropriate assumptions\. The poor scores on physics benchmarks, however, appear to be at odds with the experience of physicists, who report strong model capabilities in practice\([Schwartz, 2026b](https://arxiv.org/html/2609.13009#bib.bib21);[Guevara et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib23)\)\. This apparent discrepancy motivates our central question:

Do frontier models*really*struggle with physics?

To study this question, we conduct an expert audit of several leading physics benchmarks\. We find that frontier models can now solve a broad range of well\-defined, problem\-set\-style physics questions with near\-perfect accuracy, contrary to the low scores reported by Artificial Analysis\([Artificial Analysis, 2026](https://arxiv.org/html/2609.13009#bib.bib6)\)\. We recruit a team of physics experts made up of faculty members and their graduate researchers to audit text\-only, closed\-ended physics questions in their respective subfields\. Reviewers examine problem statements, reference solutions, and model responses to separate genuine errors by the model under test from grader errors and flaws in the benchmark materials\. When the materials are at fault, reviewers identify ill\-defined questions and missing assumptions, correct flawed problem statements and reference solutions where possible, and independently derive missing solutions\. After correcting evaluation errors and repairing or excluding flawed questions, scores rise substantially \(Figure[1](https://arxiv.org/html/2609.13009#S0.F1)\)\. On HLE\-Physics, GPT\-5\.6\-Sol’smean@4rises from47\.347\.3% to78\.778\.7%\. On CMT\-Benchmark\([Pan et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib4)\), it rises from61\.061\.0% to87\.287\.2%\. On CritPt, themean@5over the 70 evaluated challenges is32\.332\.3% before the audit, and themean@4over the 54 retained challenges is87\.587\.5% after it, with apass@4of94\.494\.4%\. The apparent gap to near\-perfect performance on these benchmarks is therefore an artifact of flawed benchmark materials and evaluation procedures, not evidence of genuine limitations in frontier models’ physics reasoning\.

## 2Methodology: Diagnosing Errors in Physics Evaluation

Our method combines benchmark evaluation with expert audits to separate genuine errors by the model under test from grader errors and flaws in the benchmark materials\. We first measure performance using the original benchmarks and available evaluation procedures, then review problem statements, reference solutions, and model responses to identify the sources of the reported errors\. Using these findings, we correct the evaluation procedures and reference solutions, repair or exclude flawed questions, and reassess model performance\.

### 2\.1Benchmark and Evaluation Suite

#### 2\.1\.1Benchmarks

We evaluate six widely used physics benchmarks spanning a broad range of subjects and difficulty levels, organized into two groups primarily according to question provenance\. The first group comprises benchmarks whose questions are drawn or adapted from existing physics exercises and examinations: UGPhysics\([Xu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib1)\)focuses on undergraduate physics, PHYBench\([Qiu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib2)\)includes problems extending to Physics Olympiad difficulty,111PHYBench’s questions were proposed by physics students\. We include PHYBench in this group because its problems are comparable in difficulty to those of the other two\.and PRISM\-Physics\([Zhao et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib3)\)contains advanced physics problems with both final\-answer and process\-level evaluation, of which we use only the former\. These benchmarks draw on publicly available source material, creating a potential route for training\-data contamination\. The second group comprises benchmarks constructed from original questions contributed by*domain experts*\. HLE\-Physics is the physics subset of Humanity’s Last Exam \(HLE\)\([Center for AI Safety et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib10)\), CMT\-Benchmark\([Pan et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib4)\)focuses on advanced condensed matter theory, and CritPt\([Zhu et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib5)\)contains advanced questions curated by experts across most frontier areas of physics\. All six are closed\-ended benchmarks\. Each problem is intended to have a definite final answer that the model must obtain, which the benchmark supplies as a reference solution\. Some benchmarks also supply a worked solution, a derivation of that answer\. A response need not follow a prescribed derivation so long as its final answer, in any equivalent mathematical form, is correct\.

#### 2\.1\.2Pre\-Audit Evaluation

We evaluate three frontier models, GPT\-5\.6\-Sol, Claude Fable 5, and Gemini 3\.1 Pro, on text\-only physics questions\. Appendix[A](https://arxiv.org/html/2609.13009#A1)summarizes tool access and reasoning settings\. We first score each model’s responses with the original benchmark materials and, where available, the benchmark’s own evaluator, before making any corrections to either\. We call these the pre\-audit scores\. CMT\-Benchmark has no public evaluator, so we adapt the HLE evaluation pipeline, including its system prompt and LLM\-judge prompt, to assess responses against the provided reference solutions\. We report pre\-audit accuracy asmean@4, the average accuracy over four attempts, except for CritPt, where the pre\-audit scores are themean@5values as reported by Artificial Analysis\.222Artificial Analysis reports no CritPt score for Fable 5 at High reasoning effort, so Fable 5’s pre\-audit CritPt score is its Max reasoning effort score, while its corrected score uses High\. GPT\-5\.6\-Sol uses the Max setting for both its pre\-audit and corrected CritPt scores\.Figure[1](https://arxiv.org/html/2609.13009#S0.F1)and Table[1](https://arxiv.org/html/2609.13009#S3.T1)report the pre\-audit accuracy of the three models\. The corresponding error rates lump together genuine model errors, grader errors, and flaws in the benchmark materials\. We therefore conduct expert audits, described below, to identify the source of each error, and then report corrected accuracy\.

### 2\.2Error Attribution

The primary goal of our audit is to identify the source of each apparent model error\. We assign each audited case to one of three categories\.

Model error\.The problem is well posed and the reference solution is correct, but the model under test gives an incorrect answer\. Only these count as genuine model errors\.

Grader error\.The problem is well posed, the reference solution is correct, and the model under test gives a correct answer, but the evaluator marks it incorrect\. This happens mainly with rule\-based evaluators, which can fail to recognize a correct answer written in an equivalent mathematical form or in a different convention\.

Benchmark error\.The problem statement or reference solution is defective\. This includes an incorrect reference solution, inconsistent conditions, ambiguity, or a missing assumption\. We classify a missing assumption as a benchmark error when it is necessary to determine the intended answer and cannot be inferred unambiguously from the problem statement\.

Breakdown of Errors Across Benchmarks

Figure 2:Benchmark defects are common, and grader errors are especially prevalent with rule\-based evaluators\.Grader errors dominate the audited cases on PHYBench and PRISM\-Physics, which use rule\-based evaluators\. Bars show the share of each benchmark’s audit set: rejected answers for the first four benchmarks \(250 in total\), and all audited questions for CMT\-Benchmark and CritPt\. Grader errors cannot be determined for CMT\-Benchmark or CritPt because per\-question judgments from their original evaluators are unavailable\.
### 2\.3Audit Data Collection

To reduce the burden of expert review for HLE\-Physics, PHYBench, PRISM\-Physics, and UGPhysics, we restrict the audit to questions for which all of GPT\-5\.6\-Sol High’s attempts were evaluated as incorrect\. The runs used for the audit are separate from those used to obtain the pre\-audit evaluation results reported in Table[1](https://arxiv.org/html/2609.13009#S3.T1)\. Appendix[B\.3](https://arxiv.org/html/2609.13009#A2.SS3)provides details of these runs\.

### 2\.4Expert Audit

We conduct the audit with a team of physicists, primarily based at Yale University\. We match each problem to an auditor with expertise in its subfield\. Assignments follow the protocol described in Appendix[F](https://arxiv.org/html/2609.13009#A6)\.

Auditors first assess whether the problem statement is well posed and the reference solution is correct\. If not, the case is classified as a benchmark error\. We exclude such questions from HLE\-Physics, PHYBench, PRISM\-Physics, and UGPhysics\. For CritPt and CMT\-Benchmark, auditors instead repair the problem statement or reference solution where possible, for example by adding a missing boundary condition or clarifying a convention needed to determine the intended answer \(Figure[6](https://arxiv.org/html/2609.13009#A6.F6)\)\. We retain repaired questions and exclude those for which no defensible repair was found\. We label the resulting evaluations “validated/repaired” and refer to them as “corrected” evaluations elsewhere in the paper\. For the remaining cases, auditors assess whether the model under test gives a correct answer, allowing for equivalent mathematical expressions and alternative conventions consistent with the problem statement, and classify the case as a model error or a grader error accordingly\. Each audited case receives one label\.

## 3Results

A Representative Example of a Rule\-Based Grader Error

PHYBench, problem 140: equivalent expressions for the same rope tensionProblem statement\.Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other\. A rope is wrapped around the spheres at the height of their centers, tying them together\. A fourth identical sphere is placed on top of the three spheres\. Find the tensionTTin the rope\. It is given that the weight of each sphere isPP\.Model final answerReference final answerT=P3​6T=\\dfrac\{P\}\{3\\sqrt\{6\}\}T=618​PT=\\dfrac\{\\sqrt\{6\}\}\{18\}PWhy the answers are equivalent\.Multiplying the numerator and denominator of the model’s answer by6\\sqrt\{6\}gives exactly the reference solution:P3​6=P​63×6=618​P\.\\frac\{P\}\{3\\sqrt\{6\}\}=\\frac\{P\\sqrt\{6\}\}\{3\\times 6\}=\\frac\{\\sqrt\{6\}\}\{18\}P\.Pre\-audit grader error: EED score 0\.0, binary score 0Audit: grader errorReviewer note\.“Expressions are algebraically the same\.”

Figure 3:A correct answer in an equivalent form receives zero credit in PHYBench’s Expression Edit Distance \(EED\) evaluation\. The model answer is from GPT\-5\.6\-Sol High; only final answers are shown\. The reference and model give the same rope tension, yet the stored score and binary decision are both zero\. Rationalizing the denominator makes the equivalence explicit\. Appendix[D\.2](https://arxiv.org/html/2609.13009#A4.SS2)gives a different PHYBench example involving a change in factor order\.As discussed earlier, we distinguish benchmarks that draw on publicly available questions and solutions, and are therefore susceptible to training\-data contamination, from those constructed from original, expert\-authored questions\. Our analysis focuses primarily on the latter group, as its lower presumed contamination risk makes model performance a more informative signal of physics problem\-solving ability\. Audit coverage varies with benchmark size and access to reference solutions\. Appendix[B](https://arxiv.org/html/2609.13009#A2)gives the sampling procedure, audit coverage, and repair or exclusion decisions for each benchmark\.

All corrected evaluations use a common pipeline adapted from HLE, with its system prompt for response generation and its judge prompt for grading\([Center for AI Safety et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib10)\)\. In our audit, this evaluator has the lowest grader error rate of all the evaluators we audited \(4\.08%\)\. Since all six benchmarks are closed\-ended, we organize their questions and reference solutions in a shared format and use the same procedure to assess answer equivalence\. As in the pre\-audit evaluation, we reportmean@4, except for the pre\-audit CritPt scores described below\. We also reportpass@4, the fraction of questions solved in at least one of four attempts\. Pre\-auditpass@4is unavailable for CritPt\.

Table[1](https://arxiv.org/html/2609.13009#S3.T1)summarizes the pre\-audit and validated/repaired results\. Corrected evaluations use the retained or repaired question sets \(Appendix[B](https://arxiv.org/html/2609.13009#A2)\), so pre\-audit and corrected scores are not always computed on the same questions\.

Table 1:Pre\-audit→\\rightarrowvalidated/repaired accuracy \(%\)\. Question counts are shown before and after validation/repair\. Dashes indicate unavailable results\. All scores usemean@4, except pre\-audit CritPt scores\*\. For all six benchmarks,pass@4is also shown\. Appendix[B](https://arxiv.org/html/2609.13009#A2)gives the selection and correction procedures for each benchmark\.Benchmark\# QuestionsMetricAccuracy \(%\)Pre\-auditValidated/RepairedFable 5High†GPT\-5\.6\-SolHigh†Gemini 3\.1 ProHigh \(no tools\)Benchmarks drawn from public sourcesPHYBench10087mean@439\.50→87\.6439\.50\\rightarrow 87\.6426\.50→90\.2326\.50\\rightarrow 90\.2346\.50→89\.9446\.50\\rightarrow 89\.94pass@447\.00→91\.9547\.00\\rightarrow 91\.9534\.00→95\.4034\.00\\rightarrow 95\.4050\.00→94\.2550\.00\\rightarrow 94\.25PRISM\-Physics10074mean@47\.50→84\.807\.50\\rightarrow 84\.8013\.00→94\.5913\.00\\rightarrow 94\.5911\.25→87\.8411\.25\\rightarrow 87\.84pass@420\.00→90\.5420\.00\\rightarrow 90\.5424\.00→95\.9524\.00\\rightarrow 95\.9524\.00→94\.5924\.00\\rightarrow 94\.59UGPhysics10082mean@478\.25→87\.8078\.25\\rightarrow 87\.8083\.00→92\.0783\.00\\rightarrow 92\.0786\.50→90\.8586\.50\\rightarrow 90\.85pass@482\.00→92\.6882\.00\\rightarrow 92\.6885\.00→93\.9085\.00\\rightarrow 93\.9088\.00→93\.9088\.00\\rightarrow 93\.90Expert\-authored benchmarksHLE\-Physics202116mean@447\.03→75\.6547\.03\\rightarrow 75\.6547\.28→78\.6647\.28\\rightarrow 78\.6640\.97→64\.8740\.97\\rightarrow 64\.87pass@453\.47→83\.6253\.47\\rightarrow 83\.6255\.94→91\.3855\.94\\rightarrow 91\.3848\.51→74\.1448\.51\\rightarrow 74\.14CMT\-Benchmark5049mean@453\.50→85\.2053\.50\\rightarrow 85\.2061\.00→87\.2461\.00\\rightarrow 87\.2450\.50→78\.0650\.50\\rightarrow 78\.06pass@460\.00→93\.8860\.00\\rightarrow 93\.8872\.00→97\.9672\.00\\rightarrow 97\.9658\.00→87\.7658\.00\\rightarrow 87\.76CritPt\*7054mean@4\*28\.57→78\.2428\.57\\rightarrow 78\.2432\.29→87\.5032\.29\\rightarrow 87\.5017\.71→54\.6317\.71\\rightarrow 54\.63pass@4–→90\.74\\text\{\-\-\}\\rightarrow 90\.74–→94\.44\\text\{\-\-\}\\rightarrow 94\.44–→68\.52\\text\{\-\-\}\\rightarrow 68\.52
†\\daggerThe pre\-audit and validated/repaired CritPt scores for GPT\-5\.6\-Sol use the Max setting\. The pre\-audit CritPt score for Fable 5 also uses the Max setting\.

\*Pre\-audit CritPt scores are Artificial Analysis’smean@5on the 70 challenges\. Validated/repaired scores aremean@4on the 54 challenges retained from the 56 that were audited\. Pre\-auditpass@4is unavailable for all models\.

### 3\.1Benchmarks Drawn From Public Sources

On the evaluated public\-source subsets, Fable 5 High’s pre\-audit and correctedmean@4are39\.5039\.50% and87\.6487\.64% on PHYBench,7\.507\.50% and84\.8084\.80% on PRISM\-Physics, and78\.2578\.25% and87\.8087\.80% on UGPhysics\. GPT\-5\.6\-Sol High’s corresponding scores are26\.5026\.50% and90\.2390\.23% on PHYBench,13\.0013\.00% and94\.5994\.59% on PRISM\-Physics, and83\.0083\.00% and92\.0792\.07% on UGPhysics\. Gemini 3\.1 Pro High’s corresponding scores are46\.5046\.50% and89\.9489\.94% on PHYBench,11\.2511\.25% and87\.8487\.84% on PRISM\-Physics, and86\.5086\.50% and90\.8590\.85% on UGPhysics\. The corrected scores are on the retained subsets, after excluding flawed questions\.

##### Sources of Reported Errors\.

Across the three public\-source audit sets, 148 of 152 cases \(97\.37%\) are attributed to benchmark or grader errors, and 4 \(2\.63%\) to model errors\. The benchmark\-level breakdown appears in Table[2](https://arxiv.org/html/2609.13009#A3.T2)and Figure[2](https://arxiv.org/html/2609.13009#S2.F2)\. Grader errors account for the largest share on PHYBench and PRISM\-Physics\. On UGPhysics, where the benchmark’s evaluator already includes an auxiliary LLM judge, 18 of 22 audited cases are benchmark errors, 3 are grader errors, and 1 is a model error\. Figure[3](https://arxiv.org/html/2609.13009#S3.F3)shows a correct answer in an equivalent form that PHYBench’s evaluator rejected\.

### 3\.2Expert\-Authored Benchmarks

#### 3\.2\.1HLE\-Physics

With tools enabled, correctedmean@4reaches78\.6678\.66% for GPT\-5\.6\-Sol High and75\.6575\.65% for Fable 5 High, compared with pre\-audit scores of47\.2847\.28% and47\.0347\.03%, respectively\. Their correctedpass@4scores are91\.3891\.38% and83\.6283\.62%\. Most errors identified in the HLE\-Physics audit are benchmark errors rather than model errors\. Appendix[B\.2\.1](https://arxiv.org/html/2609.13009#A2.SS2.SSS1)details the audit coverage and error attribution\.

#### 3\.2\.2CMT\-Benchmark

On CMT\-Benchmark, expert correction raises GPT\-5\.6\-Sol High’smean@4with tools from61\.0061\.00% to87\.2487\.24%, and itspass@4from72\.0072\.00% to97\.9697\.96%\. The corrected evaluation retains 49 of the 50 original questions, 29 of them after expert repair, and excludes one\. The same HLE\-adapted evaluator is used before and after correction\. The corrections are to the benchmark materials only\. Appendix[B\.2\.2](https://arxiv.org/html/2609.13009#A2.SS2.SSS2)describes the audit and repair procedures\.

#### 3\.2\.3CritPt

For CritPt, the pre\-auditmean@5reported by Artificial Analysis\([Artificial Analysis, 2026](https://arxiv.org/html/2609.13009#bib.bib6)\)is32\.2932\.29% for GPT\-5\.6\-Sol Max on 70 challenges\. Fable 5’s pre\-audit CritPt result uses the Max setting\. Gemini 3\.1 Pro High’s pre\-audit score is17\.7117\.71%\. On the 54 challenges retained from the 56 audited, correctedmean@4reaches87\.5087\.50% for GPT\-5\.6\-Sol Max and78\.2478\.24% for Fable 5 High, while correctedpass@4reaches94\.4494\.44% and90\.7490\.74%, respectively\. Pre\-auditpass@4is unavailable for all models, since Artificial Analysis does not report per\-challenge judgments\. Because the official reference solutions are unavailable to us, the corrected evaluation uses reference solutions derived independently by our auditors\. Appendix[B\.2\.3](https://arxiv.org/html/2609.13009#A2.SS2.SSS3)details the audit process and the composition of the pre\-audit and corrected sets\.

Figure[4](https://arxiv.org/html/2609.13009#S3.F4)compares an original CritPt problem with the auditor’s repaired version and explains why the repair was needed\. Appendix[E](https://arxiv.org/html/2609.13009#A5)discusses representative questions on the expert\-authored benchmarks that remain unsolved after correction\.

CritPt challenge 44: original and expert\-corrected problemOriginal problemProblem setupConsider the Kitaev honeycomb model at the isotropic limit \(assumingJx=Jy=Jz=1J\_\{x\}=J\_\{y\}=J\_\{z\}=1\) on a 3x2 Bravais lattice with periodic boundary conditions\.Main problemHow many degenerate ground states are there? How many of them are in the flux\-free sector? Compute the energy of the ground states with three decimal precision\.Expert\-corrected problemProblem setupConsider the Kitaev honeycomb model at the isotropic limit \(assumingJx=Jy=Jz=1J\_\{x\}=J\_\{y\}=J\_\{z\}=1\) on a 3x2 Bravais lattice with periodic boundary conditions\.To fix the normalization and sign convention unambiguously, takeH=\\displaystyle H=\{\}−∑⟨i​j⟩xσixσjx−∑⟨i​j⟩yσiyσjy\\displaystyle\-\\sum\_\{\\langle ij\\rangle\_\{x\}\}\\sigma\_\{i\}^\{x\}\\sigma\_\{j\}^\{x\}\-\\sum\_\{\\langle ij\\rangle\_\{y\}\}\\sigma\_\{i\}^\{y\}\\sigma\_\{j\}^\{y\}−∑⟨i​j⟩zσizσjz,\\displaystyle\-\\sum\_\{\\langle ij\\rangle\_\{z\}\}\\sigma\_\{i\}^\{z\}\\sigma\_\{j\}^\{z\},whereσiα\\sigma\_\{i\}^\{\\alpha\}denotes the standard2×22\\times 2Pauli matrix acting on the local spin\-12\\tfrac\{1\}\{2\}Hilbert space at siteii\.Main problemHow many degenerate ground states are there? How many of them are in the flux\-free sector? Compute the energy of the ground states with three decimal precision\.Expert comment\.The Hamiltonian is not stated explicitly, so it is unclear whetherJα=1J\_\{\\alpha\}=1multiplies Pauli operatorsσiα​σjα\\sigma\_\{i\}^\{\\alpha\}\\sigma\_\{j\}^\{\\alpha\}or spin\-12\\tfrac\{1\}\{2\}operatorsSiα​SjαS\_\{i\}^\{\\alpha\}S\_\{j\}^\{\\alpha\}, whereSα=σα/2S^\{\\alpha\}=\\sigma^\{\\alpha\}/2\. These two conventions give energies that differ by a factor of four, so the numerical ground\-state energy is not uniquely defined\. The correction makes the problem well posed and gives it a unique answer\.Figure 4:CritPt challenge 44 before and after expert repair\. Text added by the expert is shown in teal; the expert comment explains why the added convention is necessary\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)gives the corresponding benchmark\-error entry\.

## 4Related Work

##### Performance of Frontier Models on Physics Beyond Standardized Benchmarks\.

Several recent studies report physics work done with frontier models outside standardized benchmarks\. OpenAI’s 2025 science report describes GPT\-5 Pro reconstructing a hiddenSL⁡\(2,ℝ\)\\mathrm\{SL\}\(2,\\mathbb\{R\}\)symmetry algebra after a simpler warm\-up problem\([Bubeck et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib13)\); that algebra underpins Lupsasca’s analysis of vanishing black\-hole Love numbers\([Lupsasca, 2025](https://arxiv.org/html/2609.13009#bib.bib27)\)\.[Schwartz \(2026b\)](https://arxiv.org/html/2609.13009#bib.bib21)reports guiding Claude through an extended theoretical\-physics project that produced a new factorization theorem and a resummed C\-parameter calculation\([Schwartz, 2026a](https://arxiv.org/html/2609.13009#bib.bib22)\)\.[Brenner et al\. \(2026\)](https://arxiv.org/html/2609.13009#bib.bib28)pair Gemini Deep Think with tree search and numerical feedback to derive exact analytic results for cosmic\-string radiation\. None of these is a score on held\-out questions\. Each puts a human expert or an external checker in the loop, over hours or weeks, on a calculation with no reference solution to grade against\. While they do not in themselves rigorously quantify frontier models’ capability in physics, they suggest that such models are quite capable, as we systematically demonstrate in this work\.

##### Benchmark Flaws Matter More as Models Improve\.

This work highlights problems with physics benchmarks and their evaluation and shows that correcting them dramatically changes measured performance\. Related issues have been exposed in prior work outside physics\.[Northcutt et al\. \(2021\)](https://arxiv.org/html/2609.13009#bib.bib11)identify label errors throughout widely used test sets in computer vision, natural language, and audio, and show that correcting them can change model rankings\.[Bowman and Dahl \(2021\)](https://arxiv.org/html/2609.13009#bib.bib29)highlight that most natural\-language benchmarks fail to meet the standard for which they were created, and propose four criteria for an adequate benchmark: validity, reliable annotation, adequate statistical power, and disincentives for biased models\. Two of these bear directly on our setting\. First, these authors argue that reliable annotation requires distinguishing items that are simply mislabeled from items that have no clear right answer, either because the question is underspecified or because competent people read it differently\. Second, they note that a fixed benchmark loses statistical power at high accuracy: going from, for example, 98% to 98\.1% removes the same fraction of the error as going from 80% to 81% but takes roughly an order of magnitude more evaluation data to detect\. These concerns become more severe as model capability improves\. Grader error follows the same pattern\. Poor evaluation is masked when models are weak because a faulty grader errs mainly by marking correct solutions wrong, and a weak model produces few of them, so the measured score stays close to the true one\. As capability grows, correct solutions become common, the grader’s errors accumulate, and the gap between measured and true performance widens\.

##### Rule\-Based Evaluators, LLM Judges, and Reference Solution Accuracy\.

We identified rule\-based evaluators, LLM judges, and reference solution accuracy as three sources of error in evaluation pipelines\. It is widely appreciated that rule\-based evaluators, whether exact\-match or symbolic, are reproducible but brittle, rejecting equivalent answers that differ in normalization, convention, etc\. This has not deterred their use, partly because some pipelines fall back on an LLM judge when they fail\. LLM judges grade more flexibly\([Zheng et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib9)\)but at the cost of the judge’s own inaccuracies as a grader, with failures ranging from position and verbosity bias to weak accuracy on objective reasoning tasks\([Zheng et al\., 2023](https://arxiv.org/html/2609.13009#bib.bib12);[Chen et al\., 2024](https://arxiv.org/html/2609.13009#bib.bib30);[Tan et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib31)\)\. These issues are more severe in older pipelines using older frontier models as judges, and are gradually diminishing as more capable models are used as judges\. The choice of evaluator is separate from the accuracy of the reference solutions, so we audit the questions and reference solutions separately from the answer evaluation\.

##### Benchmark Failures Beyond Physics\.

The same pattern has appeared in software engineering\. SWE\-bench was built from 2,294 real GitHub issues with executable tests\([Jimenez et al\., 2024](https://arxiv.org/html/2609.13009#bib.bib32)\), and expert review of those tasks motivated the smaller Verified subset\([OpenAI, 2024](https://arxiv.org/html/2609.13009#bib.bib24)\)\. A later audit of 138 Verified tasks found material problems in 59\.4% of them, and a follow\-up estimated that roughly 30% of SWE\-bench Pro tasks were broken\([OpenAI, 2026d](https://arxiv.org/html/2609.13009#bib.bib25);[OpenAI, 2026b](https://arxiv.org/html/2609.13009#bib.bib26)\)\. SWE\-rebench makes a related argument from staleness and contamination, rebuilding tasks continuously instead of fixing a set\([Badertdinov et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib33)\)\. Defects survived two rounds of curation in a domain where every task ships with an executable test\. Physics benchmarks are graded against written reference solutions and have no comparable check, so there is no reason to expect them to be cleaner, and no way to find their defects short of re\-deriving each answer\.

## 5Discussion

In this work, we examine frontier model capability in solving physics problems\. We find that, contrary to the scores reported by Artificial Analysis and by leading benchmarks, frontier models perform strongly on all benchmarks considered in our study, which constitute a considerable subset of available benchmarks\. We attribute the discrepancy to broken benchmarks and their evaluation pipelines\. Every benchmark has some fraction of defective items, including wrong reference solutions, ambiguous questions, and grading mistakes\. A model is marked wrong on those no matter how it performs, so no model’s measured error rate can fall below the defect rate\. While a model still gets many valid items wrong, the defects add only a little to its error\. Once its true error rate drops below the defect rate, most of the errors on its scorecard are the benchmark’s, not the model’s, and this is when benchmark defects become detrimental to measuring capability\. Our results suggest that model capability has grown so much that the true error rate of frontier models is now far below the defect rate of available benchmarks\. This is supported by an expert audit of 250 cases across four benchmark subsets, in which expert review attributes 12 to the model\. The other 238 are defects in the questions or the graders; the full breakdown appears in Table[2](https://arxiv.org/html/2609.13009#A3.T2)and Appendix[C](https://arxiv.org/html/2609.13009#A3)\. Frontier models are not failing these physics problems\. The benchmarks are failing to pose them\. Quantitatively, the audit shows that on the PRISM\-Physics sample GPT\-5\.6\-Sol makes no errors among the audited cases, so every audited rejection is a defect; on UGPhysics, there is only 1 model error\. HLE\-Physics and PHYBench have not reached that point, though their raw rejection rates overstate the model component by roughly an order of magnitude\.

Having established near\-saturation of leading benchmarks by frontier models, we close with some comments\. The results suggest that frontier models are now so capable that almost no closed\-ended problem of the kind used in physics problem sets lies beyond their reach\. Our analysis suggests that small fixes, such as adding missing context to a question or allowing the model more attempts, would repair most of the reported failures\. While this suggests that models may have reached a tipping point in physics, commonly regarded as the most fundamental of the natural sciences, it by no means follows that frontier models are capable of end\-to\-end physics research\. For example, we used GPT\-based agentic harnesses of the kind that have been used successfully on open mathematics conjectures to attack several open problems in theoretical physics, and found that the agents make considerably less progress on physics than on mathematics\. As of now, we have not been able to fully solve a single one of these open problems this way\. This, together with our results, calls for a new style of benchmark for frontier models, built from new hard physics tasks and adequately verified\. Such a benchmark will likely require substantial financial resources to achieve high\-quality curation, which raises the question of how to set up data collection that is fair, fast, and of high quality, so as to measure model capability objectively\. Some of this may be done by nonprofit organizations through a more elaborate process that ensures rigor, or through open competitions in the style recently adopted in mathematics\.

##### Note added\.

During the final stages of this work, we became aware that Anthropic’s Claude Fable 5\.1 and Claude Mythos 5\.1 system card\([Anthropic, 2026](https://arxiv.org/html/2609.13009#bib.bib7)\)reports a separate expert correction of CritPt\. Anthropic obtained expert revisions to 31 of the 71 problem statements and evaluated Fable 5\.1 on the resulting internal benchmark, CritPt\-Corrected\. Using 16 max\-effort, tool\-enabled attempts per problem, with Claude Opus 4\.8 as the judge, Fable 5\.1 achieves an averagepass@1of 88\.4%; in our notation, this corresponds tomean@16\. This independently supports our finding that expert correction substantially changes the measured performance on CritPt, and their corrected score is broadly in line with our correctedmean@4of87\.587\.5% for GPT\-5\.6\-Sol Max and78\.278\.2% for Fable 5 High, with differences arising from the different corrected question sets, models, attempt budgets, and judges, which make the two evaluations not directly comparable\.

## 6Acknowledgments

We acknowledge useful conversations with Adam Brown and Amirhossein Tajdini\. This research was supported in part by the Yale Office of the Provost AI Initiatives and by a gift from Jump Trading Group\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not reflect the views of Jump Trading Group\.

## References

- N\. Alon, T\. F\. Bloom, W\. T\. Gowers, D\. Litt, W\. Sawin, A\. Shankar, J\. Tsimerman, V\. Wang, and M\. M\. WoodRemarks on the disproof of the unit distance conjecture\.External Links:2605\.20695,[Link](https://arxiv.org/abs/2605.20695)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Alpöge and Furman \(2026\)L\. Alpöge and R\. FurmanMore than two thirds of the zeta zeros are simple and on the critical line\.External Links:2608\.13637,[Link](https://arxiv.org/abs/2608.13637)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Anthropic \(2026\)AnthropicClaude Fable 5\.1 and Claude Mythos 5\.1 system card\.External Links:[Link](https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card)Cited by:[§B\.2\.3](https://arxiv.org/html/2609.13009#A2.SS2.SSS3.p3.1),[§5](https://arxiv.org/html/2609.13009#S5.SS0.SSS0.Px1.p1.1)\.
- Artificial Analysis \(2026\)Artificial AnalysisCritPt benchmark leaderboard\.External Links:[Link](https://artificialanalysis.ai/evaluations/critpt)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p4.1),[§3\.2\.3](https://arxiv.org/html/2609.13009#S3.SS2.SSS3.p1.1)\.
- Badertdinovet al\.\(2025\)I\. Badertdinov, A\. Golubev, M\. Nekrashevich, A\. Shevtsov, S\. Karasik, A\. Andriushchenko, M\. Trofimova, D\. Litvintseva, and B\. YangelSWE\-rebench: an automated pipeline for task collection and decontaminated evaluation of software engineering agents\.External Links:2505\.20411,[Link](https://arxiv.org/abs/2505.20411)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px4.p1.1)\.
- Bowman and Dahl \(2021\)S\. R\. Bowman and G\. E\. DahlWhat will it take to fix benchmarking in natural language understanding?\.External Links:2104\.02145,[Link](https://arxiv.org/abs/2104.02145)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px2.p1.1)\.
- Brenneret al\.\(2026\)M\. P\. Brenner, V\. Cohen\-Addad, and D\. WoodruffSolving an open problem in theoretical physics using AI\-assisted discovery\.External Links:2603\.04735,[Link](https://arxiv.org/abs/2603.04735)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px1.p1.1)\.
- Bubecket al\.\(2025\)S\. Bubeck, C\. Coester, R\. Eldan, T\. Gowers, Y\. T\. Lee, A\. Lupsasca, M\. Sawhney, R\. Scherrer, M\. Sellke, B\. K\. Spears, D\. Unutmaz, K\. Weil, S\. Yin, and N\. ZhivotovskiyEarly science acceleration experiments with GPT\-5\.External Links:2511\.16072,[Link](https://arxiv.org/abs/2511.16072)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1),[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px1.p1.1)\.
- Center for AI Safetyet al\.\(2026\)Center for AI Safety, Scale AI, and HLE Contributors ConsortiumA benchmark of expert\-level academic questions to assess AI capabilities\.External Links:2501\.14249,[Link](https://arxiv.org/abs/2501.14249)Cited by:[§B\.2\.1](https://arxiv.org/html/2609.13009#A2.SS2.SSS1.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1),[§3](https://arxiv.org/html/2609.13009#S3.p2.1)\.
- Chenet al\.\(2024\)G\. H\. Chen, S\. Chen, Z\. Liu, F\. Jiang, and B\. WangHumans or LLMs as the judge? A study on judgement biases\.External Links:2402\.10669,[Link](https://arxiv.org/abs/2402.10669)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px3.p1.1)\.
- Fenget al\.\(2025\)K\. Feng, Y\. Zhao, Y\. Liu, T\. Yang, C\. Zhao, J\. Sous, and A\. CohanPHYSICS: benchmarking foundation models on university\-level physics problem solving\.External Links:2503\.21821,[Link](https://arxiv.org/abs/2503.21821)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p1.1)\.
- Fenget al\.\(2026\)T\. Feng, T\. Trinh, G\. Bingham, J\. Kang, S\. Zhang, S\. Kim, K\. Barreto, C\. Schildkraut, J\. Jung, J\. Seo, C\. Pagano, Y\. Chervonyi, D\. Hwang, K\. Hou, S\. Gukov, C\. Tsai, H\. Choi, Y\. Jin, W\. Li, H\. Wu, R\. Shiu, Y\. Shih, Q\. V\. Le, and T\. LuongSemi\-autonomous mathematics discovery with Gemini: a case study on the Erdős problems\.External Links:2601\.22401,[Link](https://arxiv.org/abs/2601.22401)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Guevaraet al\.\(2026\)A\. Guevara, A\. Lupsasca, D\. Skinner, A\. Strominger, and K\. WeilSingle\-minus gluon tree amplitudes are nonzero\.External Links:2602\.12176,[Link](https://arxiv.org/abs/2602.12176)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSWE\-bench: can language models resolve real\-world GitHub issues?\.External Links:2310\.06770,[Link](https://arxiv.org/abs/2310.06770)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px4.p1.1)\.
- Lupsasca \(2025\)A\. LupsascaWhy there is no Love in black holes\.External Links:2506\.05298,[Link](https://arxiv.org/abs/2506.05298)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px1.p1.1)\.
- Northcuttet al\.\(2021\)C\. G\. Northcutt, A\. Athalye, and J\. MuellerPervasive label errors in test sets destabilize machine learning benchmarks\.External Links:2103\.14749,[Link](https://arxiv.org/abs/2103.14749)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px2.p1.1)\.
- OpenAI \(2024\)OpenAIIntroducing SWE\-bench Verified\.External Links:[Link](https://openai.com/index/introducing-swe-bench-verified/)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px4.p1.1)\.
- OpenAI \(2026a\)OpenAIFinite time blowup for Navier–Stokes\.External Links:[Link](https://cdn.openai.com/pdf/32d9f210-8b73-45e0-91bc-82a30aef8a9a/navier-stokes.pdf)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- OpenAI \(2026b\)OpenAISeparating signal from noise in coding evaluations\.External Links:[Link](https://openai.com/index/separating-signal-from-noise-coding-evaluations/)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px4.p1.1)\.
- OpenAI \(2026c\)OpenAITen advances in mathematics and theoretical computer science\.External Links:[Link](https://cdn.openai.com/pdf/ten-proofs-oai.pdf)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- OpenAI \(2026d\)OpenAIWhy SWE\-bench Verified no longer measures frontier coding capabilities\.External Links:[Link](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px4.p1.1)\.
- Panet al\.\(2026\)H\. Pan, J\. V\. Roggeveen, E\. Berg, J\. Carrasquilla, D\. Chowdhury, S\. Ganguli, F\. Ghimenti, J\. Hasik, H\. Hunt, H\. Jiang, M\. Kamb, Y\. Kao, E\. Khatami, M\. J\. Lawler, D\. Luo, T\. Neupert, X\. Qi, M\. P\. Brenner, and E\. KimCMT\-Benchmark: a benchmark for condensed matter theory built by expert researchers\.External Links:2510\.05228,[Link](https://arxiv.org/abs/2510.05228)Cited by:[§B\.2\.2](https://arxiv.org/html/2609.13009#A2.SS2.SSS2.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p4.1),[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1)\.
- Qiuet al\.\(2025\)S\. Qiu, S\. Guo, Z\. Song, Y\. Sun, Z\. Cai, J\. Wei, T\. Luo, Y\. Yin, H\. Zhang, Y\. Hu, C\. Wang, C\. Tang, H\. Chang, Q\. Liu, Z\. Zhou, T\. Zhang, J\. Zhang, Z\. Liu, M\. Li, Y\. Zhang, B\. Jing, X\. Yin, Y\. Ren, Z\. Fu, J\. Ji, W\. Wang, X\. Tian, A\. Lv, L\. Man, J\. Li, F\. Tao, Q\. Sun, Z\. Liang, Y\. Mu, Z\. Li, J\. Zhang, S\. Zhang, X\. Li, X\. Xia, J\. Lin, Z\. Shen, J\. Chen, Q\. Xiong, B\. Wang, F\. Wang, Z\. Ni, B\. Zhang, F\. Cui, C\. Shao, Q\. Cao, M\. Luo, Y\. Yang, M\. Zhang, and H\. X\. ZhuPHYBench: holistic evaluation of physical perception and reasoning in large language models\.External Links:2504\.16074,[Link](https://arxiv.org/abs/2504.16074)Cited by:[§B\.1\.1](https://arxiv.org/html/2609.13009#A2.SS1.SSS1.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1)\.
- Schwartz \(2026a\)M\. D\. SchwartzResummation of the C\-parameter Sudakov shoulder using effective field theory\.External Links:2601\.02484,[Link](https://arxiv.org/abs/2601.02484)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px1.p1.1)\.
- Schwartz \(2026b\)M\. D\. SchwartzVibe physics: the AI grad student\.External Links:[Link](https://www.anthropic.com/research/vibe-physics)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1),[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px1.p1.1)\.
- Sothanaphan \(2026\)N\. SothanaphanResolution of Erdős Problem \#728: a writeup of Aristotle’s Lean proof\.External Links:2601\.07421,[Link](https://arxiv.org/abs/2601.07421)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Tanet al\.\(2025\)S\. Tan, S\. Zhuang, K\. Montgomery, W\. Y\. Tang, A\. Cuadron, C\. Wang, R\. A\. Popa, and I\. StoicaJudgeBench: a benchmark for evaluating LLM\-based judges\.External Links:2410\.12784,[Link](https://arxiv.org/abs/2410.12784)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px3.p1.1)\.
- Tao \(2026\)T\. TaoA digestion of the Jacobian conjecture counterexample\.External Links:[Link](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/)Cited by:[§1](https://arxiv.org/html/2609.13009#S1.p2.1)\.
- Xuet al\.\(2025\)X\. Xu, Q\. Xu, T\. Xiao, T\. Chen, Y\. Yan, J\. Zhang, S\. Diao, C\. Yang, and Y\. WangUGPhysics: a comprehensive benchmark for undergraduate physics reasoning with large language models\.External Links:2502\.00334,[Link](https://arxiv.org/abs/2502.00334)Cited by:[§B\.1\.3](https://arxiv.org/html/2609.13009#A2.SS1.SSS3.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1)\.
- Zhaoet al\.\(2025\)W\. Zhao, Q\. Ma, J\. Shi, S\. Wu, J\. Han, Y\. Xiao, S\. Chen, X\. Luo, L\. Schmidt, and J\. ZouPRISM\-Physics: causal DAG\-based process evaluation for physics reasoning\.External Links:2510\.03185,[Link](https://arxiv.org/abs/2510.03185)Cited by:[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.External Links:2306\.05685,[Link](https://arxiv.org/abs/2306.05685)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px3.p1.1)\.
- Zhenget al\.\(2025\)S\. Zheng, C\. Huang, F\. Yu, J\. Yao, J\. Ye, T\. Chen, Y\. Luo, N\. Ding, L\. Bai, G\. Cui, and P\. YeSCI\-Verifier: scientific verifier with thinking\.External Links:2509\.24285,[Link](https://arxiv.org/abs/2509.24285)Cited by:[§4](https://arxiv.org/html/2609.13009#S4.SS0.SSS0.Px3.p1.1)\.
- Zhuet al\.\(2026\)M\. Zhu, M\. Tian, X\. Yang, T\. Zhou, L\. Yuan, P\. Zhu, E\. Chertkov, S\. Liu, Y\. Du, Z\. Ji, I\. Das, Q\. Chen, J\. Cao, Y\. Du, J\. Yu, P\. Wu, J\. He, Y\. Su, Y\. Jiang, Y\. Zhang, C\. Liu, Z\. Huang, W\. Jia, Y\. Wang, F\. Jafarpour, Y\. Zhao, X\. Chen, J\. Shelton, A\. W\. Young, J\. Bartolotta, W\. Xu, Y\. Sun, A\. Chu, V\. Colussi, C\. Akers, N\. Brooks, W\. Fu, J\. Zhao, M\. Qi, A\. Mu, Y\. Yang, A\. Zang, Y\. Lyu, P\. Mai, C\. Wilson, X\. Guo, J\. Zhou, D\. Inafuku, C\. Xue, L\. Gao, Z\. Yang, Y\. Hein, Y\. Kahn, K\. Zhou, D\. Luo, J\. D\. Wilson, J\. T\. Reilly, D\. Bandak, O\. Press, L\. Yang, X\. Wang, H\. Tong, N\. Chia, E\. Huerta, and H\. PengProbing the critical point \(CritPt\) of AI reasoning: a frontier physics research benchmark\.External Links:2509\.26574,[Link](https://arxiv.org/abs/2509.26574)Cited by:[§B\.2\.3](https://arxiv.org/html/2609.13009#A2.SS2.SSS3.p1.1),[§1](https://arxiv.org/html/2609.13009#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2609.13009#S2.SS1.SSS1.p1.1)\.

## Appendix

## Appendix AEvaluation Protocol

For our evaluations on the public\-source benchmarks \(PHYBench, PRISM\-Physics, and UGPhysics\), all models answer without tools\. On the expert\-authored benchmarks \(HLE\-Physics, CMT\-Benchmark, and CritPt\), GPT\-5\.6\-Sol and Fable 5 use Codex and Claude Code, respectively, with tools enabled\. Gemini 3\.1 Pro uses no tools throughout\. Unless otherwise stated, we use High reasoning effort for both answer generation and LLM judging\. On CritPt, GPT\-5\.6\-Sol uses Max reasoning effort for answer generation\. The separately reported pre\-audit CritPt scores come from Artificial Analysis\.

## Appendix BBenchmark Selection and Expert\-Audit Details

This section documents details about benchmarks, the evaluation subsets, audit data collection procedure, and benchmark corrections underlying the results in Section[3](https://arxiv.org/html/2609.13009#S3)\. The four pooled audit sets contain all questions rejected in the runs of Section[2\.3](https://arxiv.org/html/2609.13009#S2.SS3): 98 for HLE\-Physics, 56 for PHYBench, 74 for PRISM\-Physics, and 22 for UGPhysics\. These runs, selection rules, and response budgets differ from those of the pre\-audit evaluation, and were chosen to reduce the number of cases requiring expert audit\. CritPt and CMT\-Benchmark are excluded from the pooled attribution because their audits differ\. All questions in their audit sets are audited, regardless of the model’s pre\-audit response\.

### B\.1Benchmarks Drawn From Public Sources

#### B\.1\.1PHYBench

PHYBench contains 500 questions, including high\-school, undergraduate, and Physics Olympiad problems\([Qiu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib2)\)\. Reference solutions and worked solutions are publicly available for only 100 questions, to which we restrict our analysis\.333We requested the solutions from the authors, but they were not made available to us\.The authors report both a partial\-credit EED score and an accuracy, which counts an answer as correct only if its EED score is 100\. Partial credit can hide rejections of equivalent answers \(Figure[3](https://arxiv.org/html/2609.13009#S3.F3)shows an equivalent answer with EED score 0\), and we want to count every such rejection, so we use accuracy: acceptance requires an EED score of 100\. After up to five attempts, 56 questions remain rejected\. The audit attributes 13 to benchmark errors, 40 to grader errors, and 3 to model errors\. Excluding the 13 benchmark\-error questions leaves 87 questions for re\-evaluation\. GPT\-5\.6\-Sol High’smean@4increases from26\.5026\.50% to90\.2390\.23%, and itspass@4increases from34\.0034\.00% to95\.4095\.40% after correction\. Figure[3](https://arxiv.org/html/2609.13009#S3.F3)shows a correct answer in an equivalent form rejected by the EED evaluator\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)and Appendix[D\.2](https://arxiv.org/html/2609.13009#A4.SS2)give representative examples of benchmark and grader errors, respectively\.

#### B\.1\.2PRISM\-Physics

PRISM\-Physics contains 1,401 questions categorized into Easy, Medium, and Hard difficulty levels across seven physics domains\. To focus on text\-based reasoning, we exclude 549 image\-dependent questions and 19 with formatting or data\-loading issues, leaving 833 text\-only problems\. From this set, we randomly sample 100 questions\. The sampled subset contains 26 accepted and 74 rejected questions\. The audit attributes 26 rejections to benchmark errors, 48 to grader errors, and none to model errors\. Excluding the 26 benchmark\-error questions leaves 74 questions for re\-evaluation\. GPT\-5\.6\-Sol High’smean@4increases from13\.0013\.00% to94\.5994\.59%, and itspass@4increases from24\.0024\.00% to95\.9595\.95% after correction\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)and Appendix[D\.2](https://arxiv.org/html/2609.13009#A4.SS2)give representative examples of benchmark and grader errors, respectively\.

#### B\.1\.3UGPhysics

The English text\-only portion of UGPhysics contains 5,520 questions\([Xu et al\., 2025](https://arxiv.org/html/2609.13009#bib.bib1)\), from which we randomly sample 100\. Its evaluation pipeline combines a rule\-based SymPy evaluator with an optional auxiliary LLM judge\. We include this auxiliary judge in the pre\-audit evaluation, replacing the originalgpt\-4o\-2024\-08\-06judge with a frontier model to improve recognition of equivalent answers\. The audit attributes 18 of the 22 cases from the audit run to benchmark errors, 3 to grader errors, and 1 to model errors\. Excluding the benchmark\-error questions leaves 82 questions for re\-evaluation\. GPT\-5\.6\-Sol High’smean@4increases from83\.0083\.00% to92\.0792\.07%, and itspass@4increases from85\.0085\.00% to93\.9093\.90% after correction\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)gives a representative benchmark error\.

### B\.2Expert\-Authored Benchmarks

#### B\.2\.1HLE\-Physics

HLE\-Physics contains 230 physics questions from Humanity’s Last Exam, an expert\-authored benchmark designed to assess graduate\-level expertise and specialized academic knowledge\([Center for AI Safety et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib10)\)\. Questions use multiple\-choice or short\-answer formats\. We exclude 28 multimodal questions, leaving 202 text\-only questions\. The audit run yields 98 rejected questions\. The audit attributes 86 to benchmark errors, 4 to grader errors, and 8 to model errors\. Excluding the benchmark\-error questions leaves 116 questions for re\-evaluation\. GPT\-5\.6\-Sol High’smean@4increases from47\.2847\.28% to78\.6678\.66%, and itspass@4increases from55\.9455\.94% to91\.3891\.38% after correction\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)and Appendix[D\.2](https://arxiv.org/html/2609.13009#A4.SS2)give examples of benchmark and grader errors, respectively\. Appendix[E](https://arxiv.org/html/2609.13009#A5)presents representative questions that remain unsolved by GPT\-5\.6\-Sol High\.

#### B\.2\.2CMT\-Benchmark

CMT\-Benchmark contains 50 expert\-authored questions in condensed matter theory, with answers expressed as numerical values, multiple\-choice selections, algebraic expressions, or non\-commuting operator expressions\([Pan et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib4)\)\. We evaluate all 50 questions\. Although the authors provide reference solutions, their automatic grading code was not available to us\. We therefore use the HLE\-adapted evaluator for both the pre\-audit and corrected evaluations\. Experts in condensed matter theory review each problem statement for completeness and consistency, then check its reference solution and make corrections where needed\. This review identifies benchmark errors in 30 of the 50 questions\. Of these, 29 are repaired and one is excluded, leaving 49 questions in the corrected benchmark\. There are also two model\-error cases on the original valid questions; grader\-error counts are unavailable\. GPT\-5\.6\-Sol High’smean@4increases from61\.0061\.00% to87\.2487\.24%, and itspass@4increases from72\.0072\.00% to97\.9697\.96% after correction\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)gives a representative benchmark error, with the corresponding repair shown in Figure[6](https://arxiv.org/html/2609.13009#A6.F6)\.

#### B\.2\.3CritPt

CritPt contains 71 challenges\([Zhu et al\., 2026](https://arxiv.org/html/2609.13009#bib.bib5)\)\. The available CritPt evaluation interface reports aggregate scores but does not provide per\-question correctness judgments\. We therefore review every question in the selected subset, regardless of the model’s pre\-audit response\. Each question is assigned to a specialist in the relevant subfield, either a faculty member or a researcher whose participation is endorsed by their faculty supervisor\. Reviewers first assess whether each problem is sufficiently specified and self\-contained, and repair it where necessary and feasible\. Because the official reference solutions are not publicly available, reviewers independently solve the questions to establish reference solutions, then compare model responses against them\.

Experts review a 56\-challenge subset and identify benchmark errors in 21 of them\. They repair 19 and exclude the remaining two, producing a corrected set of 54 challenges with reference solutions derived by the reviewers\. The audit also identifies five model errors; grader\-error counts are unavailable\. GPT\-5\.6\-Sol Max’s pre\-audit score is32\.2932\.29%mean@5on the 70 challenges evaluated by Artificial Analysis\. On the corrected set, GPT\-5\.6\-Sol Max achieves amean@4of87\.5087\.50% and apass@4of94\.4494\.44%\. Because only aggregate pre\-audit scores are available, pre\-auditpass@4is unavailable\. Figure[4](https://arxiv.org/html/2609.13009#S3.F4)shows a representative benchmark error, the expert’s assessment, and the resulting repair\. Appendix[D\.1](https://arxiv.org/html/2609.13009#A4.SS1)gives further details, and Appendix[E](https://arxiv.org/html/2609.13009#A5)describes a representative remaining model error\.

*Comparison with a separate audit\.*Our findings are consistent with the separate audit described in Anthropic’s Claude Fable 5\.1 and Claude Mythos 5\.1 system card\([Anthropic, 2026](https://arxiv.org/html/2609.13009#bib.bib7)\)\. Anthropic reports obtaining expert corrections to 31 problem statements and evaluating an internal version of the benchmark called CritPt\-Corrected\. On that version, Fable 5\.1 achieves amean@16of 88\.4%\. Their evaluation uses a different model, corrected question set, attempt budget, and judge, so the scores are not directly comparable to ours\.

### B\.3Audit Data Collection

All four audit runs use GPT\-5\.6\-Sol at High reasoning effort\. For each question, the audit set stores one response and the original evaluator’s final binary decision\.

##### HLE\-Physics\.

The audit collection run uses up to five attempts, with tools enabled in the fifth attempt, and stops as soon as the problem is graded correct\. It then collects questions rejected in all attempts\. Among the 230 questions available, we exclude 28 multimodal questions, leaving us with 202 text\-only questions\.

##### PHYBench\.

The audit run gives each of the 100 answer\-bearing questions up to five attempts without tools, stopping after the first answer with an Expression Edit Distance \(EED\) score of 100\. In total 44 questions are accepted, leaving 56 rejected on all five attempts\. The collected response used for audit for each question is the attempt with its highest EED score\.

##### PRISM\-Physics and UGPhysics\.

The audit collection run uses a single attempt per problem for these benchmarks\. The 100\-question evaluation subsets contain 26 accepted and 74 rejected questions for PRISM\-Physics, and 78 accepted and 22 rejected for UGPhysics\.

*Retained evaluation subsets\.*After the audit, we exclude the 13, 26, and 18 benchmark\-error questions from the 100\-question PHYBench, PRISM\-Physics, and UGPhysics evaluation subsets, leaving 87, 74, and 82 questions\. The corrected scores are obtained by running the HLE\-adapted evaluation pipeline on these retained subsets\. For HLE\-Physics, the audit excludes 86 benchmark\-error questions and retains 116\. Appendix[C](https://arxiv.org/html/2609.13009#A3)gives the attribution counts\.

## Appendix CComplete Counts for the Four Pooled Audits

The four audit runs of Appendix[B\.3](https://arxiv.org/html/2609.13009#A2.SS3)cover 502 questions: 252 accepted and 250 rejected and sent for review\. After conflict resolution, 143 \(57\.20%\) are benchmark errors, 95 \(38\.00%\) grader errors, and 12 \(4\.80%\) model errors\. In total, 238 of the 250 audited rejections \(95\.20%\) are benchmark or grader errors\.

Table 2:Attribution in the four processed audit sets, after conflict resolution\. Counts are followed by percentages within each set\. CritPt and CMT\-Benchmark have different audit coverage and are excluded\. These counts are not a reconstruction of the scores in Table[1](https://arxiv.org/html/2609.13009#S3.T1)\.BenchmarkRejectionsBenchmark error \(QQ\)Grader \(GG\)Model \(MM\)HLE\-Physics9886 \(87\.76%\)4 \(4\.08%\)8 \(8\.16%\)PHYBench5613 \(23\.21%\)40 \(71\.43%\)3 \(5\.36%\)PRISM\-Physics7426 \(35\.14%\)48 \(64\.86%\)0 \(0\.00%\)UGPhysics2218 \(81\.82%\)3 \(13\.64%\)1 \(4\.55%\)Pooled250143 \(57\.20%\)95 \(38\.00%\)12 \(4\.80%\)
## Appendix DRepresentative Benchmark and Grader Errors

The following examples illustrate the two kinds of non\-model error: a defective problem statement or reference solution \(benchmark error\), and a correct response rejected by the evaluator \(grader error\)\. The final\-answer comparisons omit derivations and normalize mathematical typography for readability\. Unless otherwise stated, model answers are from the GPT\-5\.6\-Sol High audit\-collection responses to the original problems\. Expert references for repaired problems are identified explicitly\.

### D\.1Benchmark Errors

This subsection gives examples of ill\-posed questions and incorrect reference solutions\.

PHYBench: a reference expression that always vanishesProblem statement\.Two spacecraft are traveling in a space medium that is flowing uniformly at a constant velocityuuwith respect to an inertial frameSS\. The spacecraft are moving through the medium with equal relative velocitiesvvwith respect to the medium\. Neither of the velocities is known\. Ultimately, the velocities of the two spacecraft with respect to the inertial frameSSarev1v\_\{1\}andv2v\_\{2\}, and the angle between the directions of these velocities is an acute angleα\\alpha\. These three quantities are given\. The speed of light iscc\. Considering relativistic effects, determine the minimum possible value ofuu\.Model final answerOriginal reference final answerumin=c2​\|γ1−γ2\|γ12​v12\+γ22​v22−2​γ1​γ2​v1​v2​cos⁡α\\displaystyle u\_\{\\min\}=\\frac\{c^\{2\}\|\\gamma\_\{1\}\-\\gamma\_\{2\}\|\}\{\\sqrt\{\\gamma\_\{1\}^\{2\}v\_\{1\}^\{2\}\+\\gamma\_\{2\}^\{2\}v\_\{2\}^\{2\}\-2\\gamma\_\{1\}\\gamma\_\{2\}v\_\{1\}v\_\{2\}\\cos\\alpha\}\}umin=0u\_\{\\min\}=0\(the printed expression simplifies to zero\)Hereγi=\(1−vi2/c2\)−1/2\\gamma\_\{i\}=\(1\-v\_\{i\}^\{2\}/c^\{2\}\)^\{\-1/2\}abbreviates the factors in the model’s final answer\.Reviewer note\.“Reference answer equals 0, which is incorrect\.”

PRISM\-Physics: a time of flight with no timing or distance dataProblem statement\.An electron is emitted and then detected in a time\-of\-flight measurement\. Lettft\_\{f\}denote the time of flight from emission to detection, expressed in nanoseconds\. Select the correct value oftft\_\{f\}from the options below\.\(a\)​330​ns\(b\)​66​ns\(c\)​33​ns\.\\text\{\(a\) \}330\\,\\mathrm\{ns\}\\qquad\\text\{\(b\) \}66\\,\\mathrm\{ns\}\\qquad\\text\{\(c\) \}33\\,\\mathrm\{ns\}\.Model final answerOriginal reference final answerCannot be determined from the given information; no option can be uniquely selected\.tf=33​nst\_\{f\}=33\\,\\mathrm\{ns\}\(option c\)Reviewer note\.“Not enough information is given in the problem\.”

UGPhysics: reaching equilibrium versus settling without overshootProblem statement\.A particle with massmmmoves under the influence of a restoring force−k​x\-kxand a resistive force−r​x˙\-r\\dot\{x\}\(where bothkkandrrare positive constants\), wherexxis the displacement of the particle from its equilibrium position\. Ifr<r0r<r\_\{0\}, is it possible for the particle to return to the equilibrium position more quickly for certain initial conditions compared to whenr=r0r=r\_\{0\}?Model final answerOriginal reference final answerTrue \(yes\)NoReviewer note\.“In critical damping, we have the fastest decay without overshooting, but not necessarily the fastest possible way to get to the equilibrium position\. Perhaps the question meant ‘without overshoot,’ but it didn’t state that, which would be another issue with the problem/reference\.”

HLE\-Physics: an arithmetic error in the reference answerProblem statement\.In one frame of reference, an observer sees light from four distant starsS1S\_\{1\},S2S\_\{2\},S3S\_\{3\}, andS4S\_\{4\}so that the apparent angle between any pair of stars is equal\. In another frame of reference, another observer sees starsS1S\_\{1\}andS2S\_\{2\}at a right angle, andS3S\_\{3\}at an angle of3​π/43\\pi/4to both\. Ifθi​j\\theta\_\{ij\}is the angle between starsSiS\_\{i\}andSjS\_\{j\}, find the value of\(1−cos⁡\(θ14\)\)/\(1−cos⁡\(θ34\)\)\(1\-\\cos\(\\theta\_\{14\}\)\)/\(1\-\\cos\(\\theta\_\{34\}\)\)\.Model final answerOriginal reference final answer2−22\-\\sqrt\{2\}−2\-\\sqrt\{2\}Reviewer note\.“The final arithmetic is performed incorrectly, despite having the correct algebraic expression\. They should obtain2−22\-\\sqrt\{2\}instead\.”

CritPt: an unspecified spin normalizationChallenge 44\.Problem statement\.Consider the Kitaev honeycomb model at the isotropic limit \(assumingJx=Jy=Jz=1J\_\{x\}=J\_\{y\}=J\_\{z\}=1\) on a3×23\\times 2Bravais lattice with periodic boundary conditions\.How many degenerate ground states are there? How many of them are in the flux\-free sector? Compute the energy of the ground states with three decimal precision\.Expert assessment \(summary\)\.The expert notes that the statement does not write the Hamiltonian or distinguish Pauli matricesσα\\sigma^\{\\alpha\}from spin operatorsSα=σα/2S^\{\\alpha\}=\\sigma^\{\\alpha\}/2\. Both conventions are common and yield energies differing by a factor of four at the same numerical coupling\. Figure[4](https://arxiv.org/html/2609.13009#S3.F4)quotes the expert’s comment and shows the explicit Hamiltonian added in the repair\.

CMT\-Benchmark: missing assumptions in an Ising\-model questionProblem statement\.Consider a quantum Ising model at zero temperature with the following nearest\-neighbor Hamiltonian:H=−∑i,jσziσzj\+h∑iσxi\+g∑iσziH=\-\\sum\_\{i,j\}\\sigma^\{z\}\_\{i\}\\sigma^\{z\}\_\{j\}\+h\\sum\_\{i\}\\sigma^\{x\}\_\{i\}\+g\\sum\_\{i\}\\sigma^\{z\}\_\{i\}\. Which of the following statements are correct? Indicate all that apply\.\(a\)Ath=1h=1andg=0g=0, the excitation energy vanishes\.\(b\)There is a symmetry breaking transition at a finite value ofhh\.\(c\)Ath=10h=10andg=1g=1,⟨σiz​σjz⟩\\langle\\sigma^\{z\}\_\{i\}\\sigma^\{z\}\_\{j\}\\ranglevanishes exponentially as a function of\|i−j\|\|i\-j\|\.\(d\)Ath=0\.1h=0\.1andg=1g=1,⟨σix​σjx⟩\\langle\\sigma^\{x\}\_\{i\}\\sigma^\{x\}\_\{j\}\\ranglevanishes as a power\-law for a large\|i−j\|\|i\-j\|\.Model final answerOriginal reference final answera;b\\boxed\{a;b\}a;c\\boxed\{a;c\}The model answer is from GPT\-5\.6\-Sol High with tools, first attempt on the original problem\. After repair, both the model’s first answer and the corrected reference area;b;c\\boxed\{a;b;c\}; Figure[6](https://arxiv.org/html/2609.13009#A6.F6)shows both versions\.Recorded audit assessment \(summary\)\.The valueh=1h=1assumes an unstated one\-dimensional normalization\. Forg≠0g\\neq 0, the ordinary correlator in \(c\) approaches a nonzero product of one\-point functions; exponential decay applies to the connected correlator\. The reference also omits theg=0g=0symmetry\-breaking transition in \(b\)\. The documented repair specifies these assumptions and changes the answer from \(a; c\) to \(a; b; c\); see Figure[6](https://arxiv.org/html/2609.13009#A6.F6)\.

### D\.2Grader Errors

The audits record grader errors on PHYBench, PRISM\-Physics, HLE\-Physics, and UGPhysics\. We give examples from the first three, showing the model’s final answer and the reference solution’s final answer, followed by an explanation of why they are equivalent\. Derivations are omitted from both, and mathematical typography is normalized for readability\. All model answers below are from GPT\-5\.6\-Sol High\. CritPt and CMT\-Benchmark have no per\-question pre\-audit judgments, so grader errors cannot be identified for them\.

PHYBench: changing the order of factors does not change the answerProblem statement\.A small bug with a mass ofmmcrawls on a disk with a radius of2​R2R\. Relative to the disk, its crawling trajectory is a circle of radiusRRthat passes through the center of the disk\. The disk rotates with a constant angular velocityω\\omegaabout an axis passing through its center and perpendicular to the plane of the disk\. The bug’s angular crawling velocity relative to the disk is in the same direction as the disk’s angular velocity and has the same magnitude\. Solve for the maximum forceFm​a​xF\_\{max\}between the bug and the disk required to maintain this motion \(neglecting gravity\)\.Model final answerReference final answerFmax=5​m​R​ω2F\_\{\\max\}=5mR\\omega^\{2\}Fmax=5​m​ω2​RF\_\{\\max\}=5m\\omega^\{2\}RWhy the answers are equivalent\.The same scalar factors appear in a different order:m​R​ω2=m​ω2​RmR\\omega^\{2\}=m\\omega^\{2\}R\. Both expressions therefore give the same maximum force\. Subscript typography is normalized for readability\.Pre\-audit grader error:EED score0\.00\.0, binary score00\.

PRISM\-Physics: matching answers to a diving problemProblem statement\.An Olympic diver of massmmbegins a descent from a diving board of heighth=10​mh=10\\,\\mathrm\{m\}with zero initial velocity\. Take gravitational accelerationg=9\.8​m/s2g=9\.8\\,\\mathrm\{m/s^\{2\}\}\. After entering the water, assume the buoyant force balances the diver’s weight so the net gravitational\-buoyant force is zero, and the only retarding force is a viscous dragFd=b​v2F\_\{d\}=bv^\{2\}acting upward, wherebbis a positive constant andvvis the downward speed\. Letxxdenote the vertical depth below the water surface measured downward withx=0x=0at the surface\. Letttdenote time measured from the instant of impact with the water, sot=0t=0atx=0x=0\. LetTTdenote the elapsed time in air from the dive until impact\. LetV0V\_\{0\}be the speed on impact with the water \(atx=0x=0\)\. LetV⁡\(x\)V\(x\)denote the speed at depthxxunder water, withV⁡\(x\)=d​xd​tV\(x\)=\\frac\{dx\}\{dt\}fort\>0t\>0\.\(a\)Calculate the velocityV0V\_\{0\}on impact with the water and the approximate elapsed timeTTfrom the dive until impact\. Use any method you choose\.\(b\)Set up the equation of motion for vertical descent of the diver through the water\. Solve for the velocityV⁡\(x\)V\(x\)as a function of the depthxxunder water and impose the boundary conditionV⁡\(0\)=V0V\(0\)=V\_\{0\}\.\(c\)Ifb/m=25​m−1b/m=\\frac\{2\}\{5\}\\,\\mathrm\{m^\{\-1\}\}, estimate the depthxxat whichV⁡\(x\)=V010V\(x\)=\\frac\{V\_\{0\}\}\{10\}\.\(d\)Solve for the vertical depthx⁡\(t\)x\(t\)of the diver under water in terms of the timettunder water, withx=0x=0att=0t=0\.Model final answerReference final answerFinal results only; derivations omitted from both\.\(a\)V0=14​m/sV\_\{0\}=14\\,\\mathrm\{m/s\},T≈1\.43​sT\\approx 1\.43\\,\\mathrm\{s\}V0=14​m/sV\_\{0\}=14\\,\\mathrm\{m/s\},T=1\.43​sT=1\.43\\,\\mathrm\{s\}\(b\)m​d​Vd​t=−b​V2m\\dfrac\{dV\}\{dt\}=\-bV^\{2\}m​d2​xd​t2=−b​\(d​xd​t\)2m\\dfrac\{d^\{2\}x\}\{dt^\{2\}\}=\-b\\left\(\\dfrac\{dx\}\{dt\}\\right\)^\{2\}V\(x\)=V0e−bx/mV\(x\)=V\_\{0\}e^\{\-bx/m\}V⁡\(x\)=V0​e−bm​xV\(x\)=V\_\{0\}e^\{\-\\frac\{b\}\{m\}x\}\(c\)x≈5\.76​mx\\approx 5\.76\\,\\mathrm\{m\}x=52​ln⁡10​m=5\.76​mx=\\dfrac\{5\}\{2\}\\ln 10\\,\\mathrm\{m\}=5\.76\\,\\mathrm\{m\}\(d\)x⁡\(t\)=mb​ln⁡\(1\+b​V0m​t\)x\(t\)=\\dfrac\{m\}\{b\}\\ln\\left\(1\+\\dfrac\{bV\_\{0\}\}\{m\}t\\right\)x=mb​ln⁡\(1\+b​V0​tm\)x=\\dfrac\{m\}\{b\}\\ln\\left\(1\+\\dfrac\{bV\_\{0\}t\}\{m\}\\right\)Why the answers are equivalent\.Parts \(a\) and \(c\) give the same numerical results to the stated precision; the reference uses equality signs for rounded values\. In \(b\),V=d​x/d​tV=dx/dtmakes the two equations of motion identical, andb​x/m=\(b/m\)​xbx/m=\(b/m\)xmakes the velocity expressions identical\. In \(d\),\(b​V0/m\)​t=b​V0​t/m\(bV\_\{0\}/m\)t=bV\_\{0\}t/m; writingx⁡\(t\)x\(t\)simply makes the time dependence explicit\. Both solutions satisfyV⁡\(0\)=V0V\(0\)=V\_\{0\}andx⁡\(0\)=0x\(0\)=0\.Reviewer note\.“The two responses are identical\.”Pre\-audit grader error:Binary score00\.

HLE\-Physics: the same force law in the requested limitProblem statement\.In the freely jointed chain model of a polymer there arennidentical mass points joined by massless struts that constrain the distances between successive mass points to beℓ\\ell\. There are no forces on the mass points other than the constraint forces provided by the struts\.When the polymer is in thermal equilibrium with a reservoir at temperatureTTthere is a force of attraction between the polymer ends proportional toTTand linear in the separation of the ends,xx, valid whenxxis small and changes slowly\. What is the force law between the polymer ends when the polymer is thermally isolated, i\.e\. not in contact with a reservoir? Your answer may involvexx,ℓ\\ell,nnand the kinetic energy of the polymer at zero extension,E⁡\(0\)E\(0\)\. You may assumennis large\.Model final answer \(derivation omitted\)\.F⁡\(x\)=−3​E​\(0\)​xn2​ℓ2​exp⁡\(3​x22​n2​ℓ2\),F\(x\)=\-\\frac\{3E\(0\)x\}\{n^\{2\}\\ell^\{2\}\}\\exp\\\!\\left\(\\frac\{3x^\{2\}\}\{2n^\{2\}\\ell^\{2\}\}\\right\),directed toward the other end\. For very smallxx,F≃−3​E​\(0\)n2​ℓ2​x\.F\\simeq\-\\frac\{3E\(0\)\}\{n^\{2\}\\ell^\{2\}\}x\.Reference final answer \(derivation omitted\)\.F⁡\(x\)=3​E​\(0\)​x\(n​ℓ\)2\.F\(x\)=\\frac\{3E\(0\)x\}\{\(n\\ell\)^\{2\}\}\.Why the answers are equivalent\.In the stated small\-xxlimit, the model’s exponential factor tends to one, giving the linear force law also explicitly included in its answer\. Since\(n​ℓ\)2=n2​ℓ2\(n\\ell\)^\{2\}=n^\{2\}\\ell^\{2\}, this has the same magnitude as the reference\. The model’s minus sign denotes attraction toward the other end; the reference reports the attractive\-force magnitude\. The agreement is in this limit, not at arbitrary extension\.Reviewer note\.“Answers are the same, after the limit is taken\. Here, force law is only defined up to a sign convention\.”Pre\-audit grader error:Binary score00\.

## Appendix EExamples of Unsolved Questions

### E\.1CritPt: An Unjustified Real\-Polarizability Assumption

In CritPt Challenge 18, all four GPT\-5\.6\-Sol Max attempts assume that the particle polarizabilities are real, although the problem imposes no such condition\. Their expressions agree with the reference only in this special case and omit the polarizability\-phase contributions present in the general result\. Figure[5](https://arxiv.org/html/2609.13009#A5.F5)summarizes the discrepancy\.

CritPt Challenge 18Problem SetupTwo dielectric nanoparticles are deeply trapped in two Gaussian optical traps that propagate along thezz\-axis, both characterized by the wave vectorkkand the Rayleigh rangezRz\_\{R\}\. Suppose the focal planes of these traps are located atz=0z=0, and the nanoparticles are located atz=z1z=z\_\{1\}andz=z2z=z\_\{2\}respectively, wherez1,z2≪zR\{z\_\{1\}\},\{z\_\{2\}\}\\ll\{z\_\{R\}\}\. Let the distance between the two nanoparticles bedd, satisfying the far\-field conditionk​d≫1kd\\gg 1\. The polarizabilities of the two nanoparticles areα1\\alpha\_\{1\}andα2\\alpha\_\{2\}, respectively\. Both the tweezers have identical polarization, the electric field amplitudes areE1E\_\{1\}andE2E\_\{2\}, and the phases at the focal planes areϕ1\\phi\_\{1\}andϕ2\\phi\_\{2\}, respectively\.Main problemAssume that, at equilibrium, the distance vector between the two spheres isd0=\(d0,0,0\)d\_\{0\}=\(d\_\{0\},0,0\)\. The angle between the laser polarization and the particle\-connecting axis isπ/2\\pi/2\. Derivek1k\_\{1\}andk2k\_\{2\}in the following equations of motion along thezz\-direction for the two nanospheres:m​z¨1=\\displaystyle m\{\{\\ddot\{z\}\}\_\{1\}\}=\{\}−m​Ω12​z1−\(k1\+k2\)​z1\+\(k1\+k2\)​z2,\\displaystyle\-m\\Omega\_\{1\}^\{2\}\{z\_\{1\}\}\-\(\{k\_\{1\}\}\+\{k\_\{2\}\}\)\{z\_\{1\}\}\+\(\{k\_\{1\}\}\+\{k\_\{2\}\}\)\{z\_\{2\}\},m​z¨2=\\displaystyle m\{\{\\ddot\{z\}\}\_\{2\}\}=\{\}−m​Ω22​z2−\(k1−k2\)​z2\+\(k1−k2\)​z1\.\\displaystyle\-m\\Omega\_\{2\}^\{2\}\{z\_\{2\}\}\-\(\{k\_\{1\}\}\-\{k\_\{2\}\}\)\{z\_\{2\}\}\+\(\{k\_\{1\}\}\-\{k\_\{2\}\}\)\{z\_\{1\}\}\.Final\-answer comparison\.GPT\-5\.6\-Sol Max with tools, first attempt on the corrected evaluation set, versus the expert reference\. Derivations are omitted\.Expert reference final answer\.k1=\\displaystyle k\_\{1\}=\{\}E1​E2​k28​π​ε0​d0​\(k−1zR\)2​Re⁡\[α1​α2​ei​k​d0\]​cos⁡\(ϕ1−ϕ2\),\\displaystyle\\frac\{E\_\{1\}E\_\{2\}k^\{2\}\}\{8\\pi\\varepsilon\_\{0\}d\_\{0\}\}\\left\(k\-\\frac\{1\}\{z\_\{R\}\}\\right\)^\{2\}\\operatorname\{Re\}\\\!\\left\[\\alpha\_\{1\}\\alpha\_\{2\}e^\{ikd\_\{0\}\}\\right\]\\cos\(\\phi\_\{1\}\-\\phi\_\{2\}\),k2=\\displaystyle k\_\{2\}=\{\}E1​E2​k28​π​ε0​d0​\(k−1zR\)2​Im⁡\[α1​α2​ei​k​d0\]​sin⁡\(ϕ1−ϕ2\)\.\\displaystyle\\frac\{E\_\{1\}E\_\{2\}k^\{2\}\}\{8\\pi\\varepsilon\_\{0\}d\_\{0\}\}\\left\(k\-\\frac\{1\}\{z\_\{R\}\}\\right\)^\{2\}\\operatorname\{Im\}\\\!\\left\[\\alpha\_\{1\}\\alpha\_\{2\}e^\{ikd\_\{0\}\}\\right\]\\sin\(\\phi\_\{1\}\-\\phi\_\{2\}\)\.Model final answer\.k1=\\displaystyle k\_\{1\}=\{\}α1​α2​E1​E2​k28​π​ε0​d0​\(k−1zR\)2​cos⁡\(k​d0\)​cos⁡\(ϕ1−ϕ2\),\\displaystyle\\frac\{\\alpha\_\{1\}\\alpha\_\{2\}E\_\{1\}E\_\{2\}k^\{2\}\}\{8\\pi\\varepsilon\_\{0\}d\_\{0\}\}\\left\(k\-\\frac\{1\}\{z\_\{R\}\}\\right\)^\{2\}\\cos\(kd\_\{0\}\)\\cos\(\\phi\_\{1\}\-\\phi\_\{2\}\),k2=\\displaystyle k\_\{2\}=\{\}α1​α2​E1​E2​k28​π​ε0​d0​\(k−1zR\)2​sin⁡\(k​d0\)​sin⁡\(ϕ1−ϕ2\)\.\\displaystyle\\frac\{\\alpha\_\{1\}\\alpha\_\{2\}E\_\{1\}E\_\{2\}k^\{2\}\}\{8\\pi\\varepsilon\_\{0\}d\_\{0\}\}\\left\(k\-\\frac\{1\}\{z\_\{R\}\}\\right\)^\{2\}\\sin\(kd\_\{0\}\)\\sin\(\\phi\_\{1\}\-\\phi\_\{2\}\)\.

Figure 5:A representative model error on CritPt\. The model assumes real polarizabilities and therefore gives only a special case of the general result\.

## Appendix FAudit Protocol and Example Correction

### F\.1CritPt and CMT\-Benchmark

For CritPt and CMT\-Benchmark, domain experts reviewed the problem statements and reference solutions using the following procedure\.

1. 1\.Assignment\.Reviewers selected problems matching their expertise\. Each problem was assigned to a single reviewer\.
2. 2\.Initial assessment\.Before reading the reference solution, reviewers were asked to outline their own approach to the problem\.
3. 3\.Verification\.Reviewers checked the problem statement, reference solution, and final answer for correctness and consistency\. They also checked numerical calculations when applicable\.
4. 4\.Correction and ground truth\.If a problem defect admitted a defensible repair, the reviewer repaired the statement and continued the evaluation\. Otherwise, the problem was excluded\. For each retained problem, the reviewer supplied a verified reference solution and final answer, correcting the original when necessary\. These expert\-verified materials served as the ground truth for the corrected evaluations\.
5. 5\.Submission\.Reviewers submitted an evaluation form and any revised files for review by the project team\. AI tools could be used for assistance, but reviewers were responsible for independently verifying the submitted results\.

Figure[6](https://arxiv.org/html/2609.13009#A6.F6)shows a CMT\-Benchmark question whose statement and reference solution both changed in the audit\. The original question omitted several assumptions needed to determine which answer choices were correct\.

Original itemQuestion\.Consider a quantum Ising model at zero temperature with the nearest\-neighbor HamiltonianH=−∑i,jσizσjz\+h∑iσix\+g∑iσiz\.H=\-\\sum\_\{i,j\}\\sigma\_\{i\}^\{z\}\\sigma\_\{j\}^\{z\}\+h\\sum\_\{i\}\\sigma\_\{i\}^\{x\}\+g\\sum\_\{i\}\\sigma\_\{i\}^\{z\}\.Which of the following statements are correct?\(a\)Ath=1h=1andg=0g=0, the excitation energy vanishes\.\(b\)There is a symmetry\-breaking transition at a finite value ofhh\.\(c\)Ath=10h=10andg=1g=1,⟨σiz​σjz⟩\\langle\\sigma\_\{i\}^\{z\}\\sigma\_\{j\}^\{z\}\\ranglevanishes exponentially with\|i−j\|\|i\-j\|\.\(d\)Ath=0\.1h=0\.1andg=1g=1,⟨σix​σjx⟩\\langle\\sigma\_\{i\}^\{x\}\\sigma\_\{j\}^\{x\}\\ranglevanishes as a power law at large\|i−j\|\|i\-j\|\.Reference final answer\.a;c\\boxed\{a;c\}Model final answer\.a;b\\boxed\{a;b\}Expert\-corrected itemQuestion\.Consider aone\-dimensionalquantum Isingchainat zero temperature\. In the Hamiltonian below,each nearest\-neighbor bond is counted once:H=−∑i,jσizσjz\+h∑iσix\+g∑iσiz\.H=\-\\sum\_\{i,j\}\\sigma\_\{i\}^\{z\}\\sigma\_\{j\}^\{z\}\+h\\sum\_\{i\}\\sigma\_\{i\}^\{x\}\+g\\sum\_\{i\}\\sigma\_\{i\}^\{z\}\.The critical valueh=1h=1in statement \(a\) depends on both the spatial dimension and the bond\-counting normalization\.Unless stated otherwise, the correlators below are ordinary correlators\.Which of the following statements are correct?\(a\)Ath=1h=1andg=0g=0, the excitation energy vanishes\.\(b\)Atg=0g=0,there is a symmetry\-breaking transition at a finite value ofhh\.A nonzero longitudinal fieldggexplicitly breaks the Ising symmetry\.\(c\)Ath=10h=10andg=1g=1, theconnectedcorrelator⟨σiz​σjz⟩−⟨σiz⟩​⟨σjz⟩\\langle\\sigma\_\{i\}^\{z\}\\sigma\_\{j\}^\{z\}\\rangle\{\\color\[rgb\]\{0\.75,0,0\}\-\\langle\\sigma\_\{i\}^\{z\}\\rangle\\langle\\sigma\_\{j\}^\{z\}\\rangle\}vanishes exponentially with\|i−j\|\|i\-j\|\.Forg≠0g\\neq 0, the ordinary correlator approaches the nonzero disconnected contribution⟨σiz⟩​⟨σjz⟩\\langle\\sigma\_\{i\}^\{z\}\\rangle\\langle\\sigma\_\{j\}^\{z\}\\rangle\.\(d\)Ath=0\.1h=0\.1andg=1g=1,⟨σix​σjx⟩\\langle\\sigma\_\{i\}^\{x\}\\sigma\_\{j\}^\{x\}\\ranglevanishes as a power law at large\|i−j\|\|i\-j\|\.Reference final answer\.a;b;c\\boxed\{a;\{\\color\[rgb\]\{0\.75,0,0\}b\};c\}Model final answer\.a;b;c\\boxed\{a;b;c\}

Figure 6:CMT\-Benchmark problem 31 before and after expert repair\. The correction specifies the spatial dimension and bond\-counting convention, restricts statement \(b\) tog=0g=0, replaces the ordinary correlator in statement \(c\) with its connected counterpart, and updates the reference solution\. Corrections are shown in red; dark\-green annotations explain why they are required\. Identical answer\-format instructions are omitted\.

### F\.2HLE\-Physics, PHYBench, PRISM\-Physics, and UGPhysics

For the 250 questions rejected in the runs of Appendix[B\.3](https://arxiv.org/html/2609.13009#A2.SS3), physics PhD students reviewed the problem statement, the reference solution, the model response, and an AI\-generated preliminary review\. Reviewers selected an area of expertise and were assigned questions they had not previously reviewed\. The interface requested 30 reviews per contributor and allowed up to 15 skips for questions outside their expertise\. Reviewers classified each case into one of three categories: benchmark error, grader error, or model error\. Benchmark errors take precedence when the statement or reference solution is defective\.

##### Review coverage and conflict resolution\.

The review produced 446 annotations\. There are 196 questions reviewed by two distinct reviewers and 54 with one review, due to limited human resources\. The twice\-reviewed questions have 140 matching label pairs \(71\.43%\) and 56 disagreements \(28\.57%\)\. All 56 disagreements were resolved by a third review\. Table[3](https://arxiv.org/html/2609.13009#A6.T3)gives the per\-benchmark counts\.

Table 3:Review coverage and conflict resolution\.SingleandDoublecount questions reviewed by one and two reviewers, respectively\.AgreeandDisagreecount twice\-reviewed questions whose labels matched or differed before conflict resolution\.BenchmarkSingleDoubleAgreeDisagreeHLE\-Physics14846024PHYBench9473710PRISM\-Physics23513219UGPhysics814113Total5419614056Table[4](https://arxiv.org/html/2609.13009#A6.T4)gives the disagreeing label pairs before conflict resolution\.

Table 4:Conflicting label pairs in the audit annotations, before conflict resolution\. Each item is counted once, regardless of which reviewer assigned which label\. The 54 items with only one annotation are excluded\.Conflicting label pairHLE\-PhysicsPHYBenchPRISM\-PhysicsUGPhysicsTotalBenchmark error↔\\leftrightarrowgrader error11417234Benchmark error↔\\leftrightarrowmodel error1341119Grader error↔\\leftrightarrowmodel error02103Total241019356

## Appendix GAuthor Contributions

##### Project advising\.

The project advisors designed the project and guided the research\.

Authors:Lucas Baker, Arman Cohan, and John Sous\.

##### Core team\.

The core team carried out the study\.

Authors:Ali Ansari, Haoran Sun, Andy Zeyi Liu, and Mark Jabbour\.

##### Physics advisors\.

The physics advisors discussed the physics in their areas, endorsed graduate students as auditors, and advised on the audits\.

Authors:Steven Girvin, Yu He, Sohrab Ismail\-Beigi, Yongshan Ding, Aleksander Kubica, Owen Miller, Corey O’Hern, Vidvuds Ozolins, David Poland, A\. Douglas Stone, Frank C\. van den Bosch, and Logan Wright\.

##### Data auditors\.

The data auditors conducted the expert audits\. Where necessary, they established or corrected reference solutions to provide ground truth for evaluation, in particular for CritPt and CMT\-Benchmark\.

Authors:Navid Akbari1, Santanu Antu1, Kangle Cai1, Andrew Calabrese\-Day1, Mateo Cárdenes Wuttig1, Meng Cheng1, Barry T\. Chiang1, Ali Ghorashi1, Shouzhen Gu1, Yu He1, Haoyang Huang3, Sohrab Ismail\-Beigi1, Zhibo Kang1, Lukas Kienesberger1, Aleksander Kubica1, Hantian Liu1, Andy Zeyi Liu1, Charles Lomba1, Zhongling Lu1, Wenchao Ma1, Rohin E\. McIntosh1, Evan McKinney1, Vidvuds Ozolins1, David Poland1, Ivan Rojkov1, Haoran Sun1, Xulei Sun1, Yarone Meir Tokayer1, Naveen Balaji Umasankar1, Mira Varma1, Leda Wang1, Qimin Wang4, Tyler Wang1, Haoyu Wei1, Jinming Yang1, Jinchen Zhao1, Sherlock Tingrui Zhao1, Qinyuan Zheng1, Jay S\. Zou1\.

Similar Articles

Evaluating AI’s ability to perform scientific research tasks

OpenAI Blog

OpenAI introduces FrontierScience, a new benchmark for measuring expert-level AI scientific capabilities across physics, chemistry, and biology, with GPT-5.2 achieving 77% on olympiad-style tasks and 25% on research-style tasks. The paper presents early evidence that GPT-5 meaningfully accelerates real scientific workflows, shortening work from weeks to hours while establishing metrics for tracking progress toward AI-accelerated science.