The AI-Enabled Scientific Frontier

arXiv cs.AI Papers

Summary

This paper evaluates AI's performance against traditional statistics and scientific computing in 27 scientific disciplines, finding that AI often outperforms statistics at higher computational cost but increasingly outperforms computing at lower cost, reshaping the scientific frontier.

arXiv:2609.16258v1 Announce Type: new Abstract: As artificial intelligence's capabilities improve, it is increasingly viewed as a general scientific method. But how true are these claims? Does AI outperform all techniques, or only some, and how is this changing? To assess the claims, we assemble a corpus of 2,507 head-to-head comparisons between AI and other scientific analysis techniques across 27 scientific disciplines from papers published between 2000 and early 2025. We find a profound dichotomy. Relative to traditional statistics, AI often outperforms, but at a significantly higher computational cost. But there are also nearly a quarter of cases where AI is both more expensive and performs worse than traditional statistical techniques and this fraction has been stable for a decade. Relative to scientific computing, AI often underperforms, but at lower computational cost. This has begun to change: since 2020, AI's performance against scientific computing has notably strengthened and it now outperforms on more than half of comparisons. These patterns suggest that AI is therefore not a universal replacement for existing methods, but rather a valuable -- and improving -- part of a new AI-enabled scientific frontier.
Original Article
View Cached Full Text

Cached at: 09/16/26, 08:58 AM

# The AI-Enabled Scientific Frontier
Source: [https://arxiv.org/html/2609.16258](https://arxiv.org/html/2609.16258)
Emma FuAffiliation:MIT CSAIL, MIT FutureTech∗Corresponding author\. Email: neil\_t@mit\.eduNeil Thompson

###### Abstract

As artificial intelligence’s capabilities improve, it is increasingly viewed as a general scientific method\. But how true are these claims? Does AI outperform all techniques, or only some, and how is this changing? To assess the claims, we assemble a corpus of 2,507 head\-to\-head comparisons between AI and other scientific analysis techniques across 27 scientific disciplines from papers published between 2000 and early 2025\. We find a profound dichotomy\. Relative to traditional statistics, AI often outperforms, but at a significantly higher computational cost\. But there are also nearly a quarter of cases where AI is both more expensive and performs worse than traditional statistical techniques and this fraction has been stable for a decade\. Relative to scientific computing, AI oftenunderperforms, but at lower computational cost\. This has begun to change: since 2020, AI’s performance against scientific computing has notably strengthened and it now outperforms on more than half of comparisons\. These patterns suggest that AI is therefore not a universal replacement for existing methods, but rather a valuable – and improving – part of a new AI\-enabled scientific frontier\.

One Sentence Summary:AI is reshaping scientific analysis by opening an AI\-enabled scientific frontier, where gains over traditional statistics often come at higher computational cost, but advances over scientific computing increasingly come at far lower cost\.

## Introduction

The future of science will be shaped by the analytical methods researchers choose, and by the performance they are willing to buy with computation\[[39](https://arxiv.org/html/2609.16258#bib.bib1)\]\. One view of this future is that AI is emerging as a general scientific method and may become a dominant analytical approach across science\[[5](https://arxiv.org/html/2609.16258#bib.bib10),[37](https://arxiv.org/html/2609.16258#bib.bib11),[42](https://arxiv.org/html/2609.16258#bib.bib13)\]\. Another holds that AI should be treated as a “normal” technology whose advantages are substantial but context dependent\[[27](https://arxiv.org/html/2609.16258#bib.bib14),[28](https://arxiv.org/html/2609.16258#bib.bib15),[26](https://arxiv.org/html/2609.16258#bib.bib16)\]\. The evidence assembled here supports neither simple universality nor simple skepticism, and is thus more in line with the second view\. More importantly, it shows that AI is increasingly reshaping the scientific frontier itself\.

Over the past six decades, advances in scientific computing have transformed how scientists interrogate nature\. Fast Fourier transforms accelerated spectral analysis by orders of magnitude\[[11](https://arxiv.org/html/2609.16258#bib.bib17)\], finite element methods transformed continuum mechanics\[[10](https://arxiv.org/html/2609.16258#bib.bib18)\], and molecular dynamics enabled atomic\-resolution simulations\[[31](https://arxiv.org/html/2609.16258#bib.bib19)\]\. Today, many anticipate another revolution, this time enabled by artificial intelligence\. Deep neural networks and large foundation models already perform tasks such as protein structure prediction, medium\-range weather forecasting, and broader Earth\-system prediction at or beyond long\-standing analytical pipelines, often with dramatically different computational requirements\[[19](https://arxiv.org/html/2609.16258#bib.bib12),[1](https://arxiv.org/html/2609.16258#bib.bib20),[22](https://arxiv.org/html/2609.16258#bib.bib21),[30](https://arxiv.org/html/2609.16258#bib.bib22),[6](https://arxiv.org/html/2609.16258#bib.bib23),[32](https://arxiv.org/html/2609.16258#bib.bib24)\]\.

These successes have encouraged expansive claims that AI may function as a general\-purpose scientific method rather than a domain\-specific tool\. At the same time, there are reasons for caution\. Success in one scientific setting need not translate to another; theoretical and empirical work has highlighted no\-free\-lunch limits, failures under distribution shift, overfitting to narrow benchmarks, and the risk of illusory understanding\[[43](https://arxiv.org/html/2609.16258#bib.bib25),[21](https://arxiv.org/html/2609.16258#bib.bib26),[23](https://arxiv.org/html/2609.16258#bib.bib27),[26](https://arxiv.org/html/2609.16258#bib.bib16)\]\. The computational budgets required to train frontier models have also risen steeply\[[12](https://arxiv.org/html/2609.16258#bib.bib28),[40](https://arxiv.org/html/2609.16258#bib.bib29)\], suggesting that cost concerns may also restrict AI’s usage\.

Individual cases highlight the variety of experiences that scientific disciplines are having with AI\. A climate modeler may point to neural weather systems that rival or exceed leading operational baselines at far lower runtime cost\[[22](https://arxiv.org/html/2609.16258#bib.bib21),[30](https://arxiv.org/html/2609.16258#bib.bib22),[6](https://arxiv.org/html/2609.16258#bib.bib23)\], whereas a researcher working with tabular biomedical or economic data may find that deep models improve accuracy only inconsistently, or only after substantial tuning and compute\[[35](https://arxiv.org/html/2609.16258#bib.bib30),[16](https://arxiv.org/html/2609.16258#bib.bib31),[25](https://arxiv.org/html/2609.16258#bib.bib32),[24](https://arxiv.org/html/2609.16258#bib.bib33)\]\. A notable single\-domain precedent comes from clinical prediction, where a systematic review found no overall performance benefit of neural networks over logistic regression\[[9](https://arxiv.org/html/2609.16258#bib.bib34)\]\. Recent reviews and benchmark initiatives have begun to map out scientific machine learning within particular domains\[[38](https://arxiv.org/html/2609.16258#bib.bib35)\], but what is missing is a cross\-domain empirical map of where AI methods sit in the broader cost–performance landscape defined by traditional statistical learning and physics\-based scientific computing\. Without such a map, researchers risk overgeneralizing from their own experience\.

In our analysis, we focus on AI’s analytic capabilities, specifically its ability to replace other techniques from statistical analysis or scientific computing\. This is distinct from work examining AI systems designed to assist or automate tasks currently performed by scientists themselves\[[20](https://arxiv.org/html/2609.16258#bib.bib36),[7](https://arxiv.org/html/2609.16258#bib.bib37),[36](https://arxiv.org/html/2609.16258#bib.bib38)\]\. Our focus matters because analytical techniques often set both the quality ceiling and the computational bottleneck of scientific work\. If AI systematically changes that trade\-off, it changes not only scientific methods, but the pace and scope of scientific discovery itself\[[4](https://arxiv.org/html/2609.16258#bib.bib2),[18](https://arxiv.org/html/2609.16258#bib.bib9)\]\.

The productive question, then, is not whether AI is universally superior\. It is where AI sits relative to existing alternatives, how it performs against them, and how that position and performance are changing as AI \(and the other techniques\) evolve over time\. To answer these questions, we construct a dataset of 2,507 direct comparisons drawn from studies published between 2000 and early 2025, each reporting predictive performance and computational cost for at least two methods evaluated on the same task and dataset\. We classify methods into three families—traditional statistics, scientific computing, and AI—based on how they learn from data\[[8](https://arxiv.org/html/2609.16258#bib.bib39)\], and we group 27 scientific disciplines into seven clusters based on their dominant computational methods\. This structure lets us test whether AI’s effectiveness depends on what it replaces\.

Historically, scientific analysis has been anchored by two method families: traditional statistics, which usually offer low computational cost and solid baseline performance, and scientific computing, which offers higher\-fidelity representations of complex systems but often at far greater computational expense\. We hypothesize that AI can alter this landscape in two distinct ways: by outperforming traditional statistical methods at additional computational cost, or by enabling \(better or worse\) scientific\-computing results at far lower cost\. Figure[1](https://arxiv.org/html/2609.16258#Sx1.F1)illustrates this logic and introduces the reference\-class structure that underlies the rest of the paper\.

Five findings organize the discussion\. The central one is that AI is best understood as creating a new point on the scientific frontier, rather than merely an intermediate compromise\. Second, when AI replaces traditional statistical methods, the modal outcome is higher performance at higher computational cost; when AI replaces scientific\-computing baselines, it provides a lower cost option that sometimes improves performance and sometimes does not\. Third, the impact of AI adoption is highly domain dependent\. Fourth, AI’s success at replacing other methods has accelerated since 2020\. And, fifth,allscientific methods are evolving over time, and thus AI’s relative position on the scientific frontier is shifting compared to other techniques\.

![Refer to caption](https://arxiv.org/html/2609.16258v1/panel1-sub-v2.png)Fig\. 1:The AI\-enabled scientific frontier—hypothesis\.Conceptual illustration of how AI can extend the historical compute–performance trade\-off defined by traditional statistics and scientific computing\. In this interpretation, AI is not merely an intermediate option: it opens up a new position on the frontier that offers higher performance than traditional statistics at moderate additional cost, or much lower cost than scientific computing at a reduced level of performance\.### Analysis techniques across scientific domains

Whether AI appears transformative or incremental depends strongly on what it is being compared to\. If a neural weather model is judged against an operational ensemble forecast that consumes millions of core\-hours per day, even modest accuracy improvements at much lower runtime will seem consequential\. If a similar neural architecture is benchmarked against a well\-regularized random forest on a small tabular dataset, the relevant question is different: whether the performance gain, if any, justifies the additional computational burden\. Figure[2](https://arxiv.org/html/2609.16258#Sx1.F2)makes this reference\-class problem concrete\. For our analysis, we infer the reference class based on the comparisons made by the literature – that is, we consider a comparison technique to be relevant if a study incorporates and documents a comparison analysis as an integral element of its empirical or analytical framework\.

Fig\. 2:AI’s reference class across scientific fields\.\(A\) Most common traditional\-statistics and scientific\-computing baselines used in head\-to\-head AI comparisons\. \(B\) Share of AI comparisons against each baseline family across domain clusters\.On the statistical side, AI is usually compared not with trivial models but with strong workhorses such as support vector machines, random forests, logistic regression, decision trees, andkk\-nearest neighbors\. On the scientific\-computing side, the baselines include stochastic simulations, geophysical system models, partial differential equation solvers, analytic solutions, and quantum or atomistic simulations\. These methods are often the backbone of high\-fidelity scientific workflows and can be extremely computationally demanding\. Figure[2](https://arxiv.org/html/2609.16258#Sx1.F2)A therefore shows that claims that “AI performs better” can refer to very different kinds of replacement, depending on the incumbent method family\.

That heterogeneity is not incidental; it is a structural feature of the corpus\. Figure[2](https://arxiv.org/html/2609.16258#Sx1.F2)B shows that in Earth & Space Sciences and Physical Modeling Sciences, AI is evaluated almost equally often against traditional statistical and scientific\-computing baselines\. By contrast, Life Sciences, Engineering, Environmental & Agricultural Sciences, and the heterogeneous “Other” cluster remain overwhelmingly statistics\-heavy\.

Figure[2](https://arxiv.org/html/2609.16258#Sx1.F2)therefore reveals two important facts\. First, AI’s baselines are not uniform\. In some fields it is mainly trying to displace traditional statistics, whereas in others it is competing directly with mature numerical solvers\. Second, where AI is evaluated against scientific computing, the baseline methods are typically high\-fidelity models rather than toy simulations, implying both substantial headroom for computational savings and high standards for fidelity\.

### What happens when AI replaces existing methods?

When researchers adopt AI in place of an existing method, what happens to performance and cost? Because these measures are reported in field\-specific units \(from accuracy and energy\-per\-atom to Courant\-number violations and core\-hours\), we harmonize comparisons using two encoding schemes\. With our first encoding scheme, we assign each replacement to a discrete outcome category based on whether AI improves or degrades*performance*and*cost*\. An outcome iswin\-winwhen AI achieves both better performance and lower cost,performance\-prioritizationwhen AI improves accuracy but at higher cost,efficiency prioritizationwhen AI reduces cost but hurts performance, andlose\-losewhen AI is both less accurate and more expensive than the baseline it replaces\. Over all the comparisons, we find that AI achieves win\-win and lose\-lose outcomes in nearly equal proportions \(21\.1% and 19\.3%\)\. Whereas, AI often allows performance prioritization \(45\.6%\) and only sometimes efficiency prioritization \(13\.9%\)\. An even clearer pattern emerges when we disaggregate these results based on whether AI is being compared to traditional statistics or scientific computing, as shown in Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)\.

Four examples from our corpus of comparisons illustrate these scenarios:

- •Win\-Win: Scientific computing→\\toAI\.Tubiana et al\. predict protein–protein binding sites using a geometric deep learning model in place of a structural homology baseline, improving AUCPR from 0\.613 to 0\.694 while reducing computation time from≈30\{\\approx\}30days to≈1\.5\{\\approx\}1\.5hours\[[41](https://arxiv.org/html/2609.16258#bib.bib6)\]\.
- •Performance prioritization: Traditional statistics→\\toAI\.Davagdorj et al\. predict patient risk of non\-communicable diseases from national health survey data using a neural network in place ofkk\-nearest neighbors, improving classification accuracy from 85\.0% to 95\.0% but at≈10\.5×\{\\approx\}10\.5\\timeslonger execution time\[[13](https://arxiv.org/html/2609.16258#bib.bib3)\]\.
- •Efficiency prioritization: Scientific computing→\\toAI\.George and Huerta estimate gravitational\-wave parameters with a convolutional neural network \(CNN\) rather than a matched\-filter pipeline, reducing inference time by≈10,000×\{\\approx\}10\{,\}000\\timesbut increasing mean relative error from 18% to 22%\[[15](https://arxiv.org/html/2609.16258#bib.bib5)\]\.
- •Lose\-Lose: Traditional statistics→\\toAI\.Ebiwonjumi et al\. predict the decay heat of spent light\-water\-reactor fuel assemblies with a neural network instead of Gaussian process regression, finding that mean absolute error rose from 4\.28 to 5\.56 while total training time increased by≈32×\{\\approx\}32\\times\[[14](https://arxiv.org/html/2609.16258#bib.bib4)\]\.

When numerical values are available, as in these examples, we compute our second encoding for outcomes: log\-ratio measures that capture the magnitude of change and locate each comparison in a common cost–performance plane by anchoring the baseline technique at the origin and orienting so that positive values indicate improvement\. Full details of corpus construction, taxonomy definitions, metric harmonization, and robustness checks are provided in the Supplementary Materials\.

![Refer to caption](https://arxiv.org/html/2609.16258v1/panel3-sub-v8.png)Fig\. 3:AI outcomes when replacing existing methods\.\(A,B\) Outcome categories based on whether AI improves performance and/or computational cost relative to the baseline it replaces\. \(C,D\) Quantitative comparisons plotted in normalized performance–cost space\. \(E\) Outcome distributions across scientific domain clusters\.The results differ sharply depending on what AI replaces\. When AI replaces traditional statistical methods \(Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)A\), the dominant effect is to improve performance, but at higher computational cost\. The modal outcome is*performance prioritization*; in 59\.6% of comparisons, AI performs better but at higher computational cost\. Another 13\.9% are win\-win, where AI is both better and less computationally expensive\. Lose\-lose cases occur in 23\.6% of comparisons, and efficiency prioritization is rare \(2\.8%\)\. Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)C shows the quantitative distribution of these replacements, where the average performance gain is \+23\.2% \(Δ¯perf=0\.090\\bar\{\\Delta\}\_\{\\text\{perf\}\}=0\.090\) and computational cost increases by one order of magnitude on average \(Δ¯cost=0\.996\\bar\{\\Delta\}\_\{\\text\{cost\}\}=0\.996, or≈10×\\approx 10\\times\), with both calculated as geometric means\. This reinforces the typical trade\-off in science in which better performance can be achieved at the cost of additional computation\.

When AI replaces scientific computing \(Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)B\), the dominant effect is to reduce computational cost but about half the time this comes with a reduction in performance\. More specifically, AI is win\-win \(better performance, lower cost\) for 41\.1% of comparisons and efficiency\-prioritizing for 44\.8%\. Lose\-lose cases are rare \(7\.4%\), as is performance prioritization \(6\.7%\)\. Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)D shows that when we can quantify these outcomes they are in the upper half of the plane, consistent with learned surrogates that improve performance by \+26\.4% on average \(Δ¯perf=0\.102\\bar\{\\Delta\}\_\{\\text\{perf\}\}=0\.102\) while sharply reducing cost by approximately 2\.8 orders of magnitude on average \(Δ¯cost=−2\.788\\bar\{\\Delta\}\_\{\\text\{cost\}\}=\-2\.788, or≈614×\\approx 614\\timescheaper\)\. These results reflect AI’s ability to preserve the incumbent model’s fidelity while delivering near\-simulation\-level performance without the associated computational burden\[[3](https://arxiv.org/html/2609.16258#bib.bib40)\]\.

AI’s value therefore depends on what it replaces, but not because it usually fails against traditional statistics and only succeeds against simulation\. Relative to traditional statistics, AI frequently opens a performance\-for\-compute trade\-off\. Relative to scientific computing, AI is roughly split between being a cheaper\-but\-imperfect surrogate and being strictly better than the comparison technique\.

### Domain\-specific effectiveness and the migration of effort

Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)E shows that AI’s gains are not uniform across science\. Earth & Space Sciences show the clearest broad gains, with win\-win outcomes account for 38\.3% of comparisons\. This is followed by Environmental & Agricultural Sciences \(27\.3%\) and Economics \(24\.6%\)\. Conversely, Life Sciences has the lowest share of win\-win outcomes \(8\.0%\) followed by Engineering \(15\.3%\), which is similar to the Other category \(14\.5%\)\. By contrast, lose\-lose cases are most common in Engineering \(39\.0%\) followed by Life Sciences \(23\.5%\), Physical Modeling Sciences \(21\.7%\), and Economics Sciences \(21\.6%\)\.

A natural question is how much of this variation reflects genuine domain\-specific differences in AI’s effectiveness, as opposed to differences in what AI is being compared against\. Because replacements of traditional statistics and of scientific computing have sharply different outcome profiles \(Figures[3](https://arxiv.org/html/2609.16258#Sx1.F3)A,B\), fields with more scientific\-computing baselines will mechanically produce more win\-win and efficiency prioritization outcomes regardless of any domain\-specific effect\. To disentangle the two, we weight the corpus\-wide outcome rates for each baseline family by each domain’s baseline mix \(Figure[2](https://arxiv.org/html/2609.16258#Sx1.F2)B\) to obtain expected outcome shares under the null hypothesis that AI performs identically everywhere, with variation arising solely from what it replaces\. Expressed as the deviation between observed and expected rates, the domains where AI outperforms its baselines more often than composition alone would predict are Environmental & Agricultural Sciences \(\+21\.0\+21\.0pp\), Other \(\+16\.9\+16\.9pp\), Earth & Space Sciences \(\+8\.4\+8\.4pp\), and Economic Sciences \(\+4\.5\+4\.5pp\)\. Life Sciences is near parity \(\+0\.5\+0\.5pp\), while Physical Modeling Sciences \(−9\.3\-9\.3pp\) and Engineering \(−17\.0\-17\.0pp\) fall below expectation\. On the computational side, Earth & Space Sciences \(\+9\.3\+9\.3pp\) and Environmental & Agricultural Sciences \(\+8\.6\+8\.6pp\) produce cheaper\-than\-expected AI models, whereas Life Sciences \(−5\.9\-5\.9pp\) and the remaining fields show modest negative deviations\.

As these deviations show, several domains deviate substantially from the null\. The cleanest test comes from Earth & Space Sciences and Physical Modeling Sciences, which share nearly identical baseline compositions \(≈55%\{\\approx\}55\\%traditional statistics,≈45%\{\\approx\}45\\%scientific computing\) yet exhibit sharply divergent outcomes \(χ2=25\.8\\chi^\{2\}=25\.8,d​f=3df=3,p=1\.0×10−5p=1\.0\\times 10^\{\-5\}\)\. Earth & Space Sciences achieves a win\-win rate of 38\.3%—twelve percentage points above its composition\-predicted rate of 26\.2%—while Physical Modeling Sciences falls below the same expected rate at 21\.7%, with correspondingly elevated lose\-lose outcomes \(21\.7% observed vs\. 16\.4% expected\)\. Because baseline composition is effectively held constant in this comparison, the divergence overwhelmingly reflects domain\-specific factors: differences in the structure of prediction tasks, the maturity of incumbent methods, or the intrinsic learnability of the underlying data\. A second, independent line of evidence points in the same direction\. Even when restricted to traditional\-statistics\-only comparisons—removing composition effects entirely—Life Sciences achieves a win\-win rate of just 6\.2%, less than half the corpus\-wide traditional\-statistics average of 13\.9%, while its performance prioritization rate \(66\.9%\) exceeds the same average by seven percentage points\. These patterns confirm that the heterogeneity in Figure[3](https://arxiv.org/html/2609.16258#Sx1.F3)E is not merely a compositional artifact\. Some disciplines may feature problems whose underlying structure is intrinsically more amenable to neural approximation—for instance, spatiotemporal fields with strong physical regularities—while others may present data manifolds that are harder to compress into the continuous, low\-dimensional representations that neural architectures favor\[[29](https://arxiv.org/html/2609.16258#bib.bib7),[2](https://arxiv.org/html/2609.16258#bib.bib8)\]\.

Overall, AI’s usage in most disciplines leads to performance prioritization, with higher costs but also higher performance\. This is consistent with their wide\-scale usage of baseline techniques from traditional statistics and the higher computational cost that is usually required to improve it\.

### How AI’s advantage has changed over time

Figure[4](https://arxiv.org/html/2609.16258#Sx1.F4)shows that AI’s relative advantage has not been static\. It has changed substantially over time, and the form of that change depends on what AI is replacing\.

Fig\. 4:How AI’s advantage has changed over time\.\(A,B\) Outcome distributions across time when AI replaces traditional statistics and scientific computing\. \(C\) Average normalized positions of traditional statistics, scientific computing, and AI, showing a shift from an intermediate position to AI outperforming scientific computing on average\.When AI replaces traditional statistical methods \(Figure[4](https://arxiv.org/html/2609.16258#Sx1.F4)A\), the early evidence is limited and mixed\. In the pre\-2012 sample, the observed cases consist entirely of performance prioritization \(58\.3%\) and lose\-lose \(41\.7%\) outcomes, indicating that early AI systems did not provide cheaper substitutions for traditional statistics\. In 2012–2015, results are mixed: lose\-lose outcomes account for 43\.6% of cases, while performance prioritization and efficiency prioritization each account for 20\.5%, and win\-win outcomes remain limited \(15\.4%\)\. After 2015, however, the pattern seems to stabilize\. In 2016–2019, performance prioritization becomes the dominant outcome \(56\.1%\), with lose\-lose outcomes falling to 21\.9%\. By 2020–2024, this pattern strengthens further: 63\.4% of substitutions are performance prioritization, whereas 22\.5% are lose\-lose\. The mature pattern is that AI delivers higher performance at higher computational cost*and that this pattern has been stable for nearly a decade\.*

The temporal pattern looks different when AI replaces scientific computing \(Figure[4](https://arxiv.org/html/2609.16258#Sx1.F4)B\)\. In the earliest period, favorable substitutions are already the norm, but they arise mainly through efficiency gains: 75\.6% of pre\-2012 cases are efficiency prioritization and 22\.2% are win\-win\. The small 2012–2015 sample shows the same basic structure\. In 2016–2019, the distribution broadens, but favorable outcomes still dominate overall: 49\.2% of cases are efficiency prioritization and 30\.5% are win\-win, even though lose\-lose outcomes rise to 16\.9%\. By 2020–2024, the balance shifts again, this time toward stronger dominance: 49\.5% of substitutions are win\-win, 34\.8% are efficiency prioritization, and lose\-lose outcomes fall to just 6\.5%\. Thus, relative to scientific computing, AI begins primarily as a cheaper\-but\-imperfect surrogate and over time some of these become strict improvements on both performance and cost\.

Figure[4](https://arxiv.org/html/2609.16258#Sx1.F4)C integrates these temporal changes by locating traditional statistics, scientific computing, and AI in the normalized cost–performance plane across broad periods\. Recall that in Figure[1](https://arxiv.org/html/2609.16258#Sx1.F1), we hypothesized that AI would occupy an intermediate position on the performance–cost curve\. This is indeed the pattern prior to 2020, but we can also see a new pattern emerging post 2020\. More specifically, in the pre\-2012 period, AI occupies an intermediate position between the two historical anchors of scientific analysis: traditional statistics are cheaper but worse\-performing, whereas scientific computing is more expensive but better\-performing\. This same broad ordering persists in 2012–2019, but with much smaller differences in performance between the techniques\. The computational cost difference in this period between AI and scientific computing is particularly pronounced\. The clearest reordering appears in the 2020\+ period\. Traditional statistics remain cheaper but worse\-performing than AI, while scientific computing becomes both more expensive and less performant on average\. In this sense, the figure does not merely show AI filling the space between existing approaches\. It shows AI moving from an intermediate position to one that increasingly dominates some implementations of scientific computing while maintaining a performance advantage over traditional statistics\. Here, we emphasize that it is only*some*implementations because even in the 2020\+ period,41%41\\%of AI implementations perform worse than the scientific computing method they are compared to \(and thus remain well represented by Figure[1](https://arxiv.org/html/2609.16258#Sx1.F1)\)\.

### The AI\-enabled scientific frontier

The previous section documented how AI’s relative position has shifted over time\. Figure[5](https://arxiv.org/html/2609.16258#Sx1.F5)instead takes a cross\-sectional view, asking how the three method families compare when all observations are considered together in the cost–performance plane\.

Traditional statistical methods occupy a low\-cost, lower\-performance region of the plane\. Scientific computing lies at the opposite extreme, achieving higher performance but at computational costs≈3,300​x\\approx 3\{,\}300\\text\{x\}greater\. AI occupies a distinct position\. Its performance is comparable to, and on average slightly exceeds, that of scientific computing, while≈330​x\\approx 330\\text\{x\}cheaper\. In other words, AI delivers near–simulation\-level performance without simulation\-level cost\. Rather than simply filling the gap between traditional statistics and scientific computing, it introduces a new frontier point in the cost–performance landscape\.

![Refer to caption](https://arxiv.org/html/2609.16258v1/panel5-sub.png)Fig\. 5:The AI\-enabled scientific frontier—quantified\.Average positions of traditional statistics, AI, and scientific computing in normalized cost–performance space\.Figure[6](https://arxiv.org/html/2609.16258#Sx1.F6)extends this perspective by adding additional head\-to\-head technique comparisonswithineach method family \(e\.g\. an older AI technique compared to a newer one\)\. It therefore shows how each method family is evolving, with arrows indicating quadrant averages for how each family is extending the scientific frontier\. The picture that emerges is of an ever\-deepening set of analytical techniques that offer a broad range of cost\-performance options for science\.

Taken together, these results clarify how the AI\-enabled scientific frontier is evolving\. AI does not simply replace existing approaches, nor does it uniformly dominate across the entire cost–performance space\. Instead, it expands the set of attainable trade\-offs in regions where simulation\-based methods have historically dominated\. The resulting picture is not one of wholesale replacement but of a reconfigured frontier: traditional statistics, AI, and scientific computing occupy adjoining segments of a shared efficiency boundary, and scientific progress increasingly comes from matching problems to the segment whose inductive biases and computational profile best suit their data and accuracy requirements\.

![Refer to caption](https://arxiv.org/html/2609.16258v1/panel6-sub.png)Fig\. 6:Broadening regions of possibility in modern scientific methods\.Distributions of observed comparisons within each method family in normalized cost–performance space\.

## Conclusion

Our results suggest that AI is reshaping scientific analysis, but not in a single uniform way\. Relative to traditional statistical methods, AI most often improves performance by drawing on additional computation\. Relative to scientific computing, AI more often preserves or improves performance while sharply reducing cost\. These are distinct mechanisms\. One reflects scientists’ willingness to spend more computation to obtain better results; the other reflects AI’s growing ability to act as a high\-performing surrogate for computationally intensive pipelines\.

This distinction helps explain why AI’s impact varies so strongly across fields\. In simulation\-heavy domains, especially Earth & Space Sciences and related physical\-modeling settings, AI frequently delivers clear two\-dimensional gains: lower cost with comparable or better performance\. In statistics\-heavy domains, the more common pattern is different\. AI often improves performance, but only at higher computational cost, making its value depend on whether those gains justify the added burden of training, tuning, and deployment\[[35](https://arxiv.org/html/2609.16258#bib.bib30),[16](https://arxiv.org/html/2609.16258#bib.bib31),[25](https://arxiv.org/html/2609.16258#bib.bib32),[24](https://arxiv.org/html/2609.16258#bib.bib33),[34](https://arxiv.org/html/2609.16258#bib.bib41)\]\. Engineering remains more mixed, underscoring that AI’s scientific value depends not only on model class but also on the structure of the underlying problem\.

The temporal evidence suggests that these patterns are not static\. Before 2020, AI often occupied an intermediate position between traditional statistics and scientific computing\. In the most recent period, AI more often exceeds scientific computing on average while maintaining a performance advantage over traditional statistics\. This does not establish an inexorable trend, nor does it imply that all scientific domains will move in the same direction\. It does, however, indicate that AI’s role in science is becoming more expansionary than marginal\.

Our study has several limitations\. Only about half of the comparisons in our dataset report numerical cost metrics, and cost is measured in diverse units; we mitigate this with log\-ratio encodings and robustness checks but cannot fully correct for reporting bias\. The corpus largely reflects communities with strong benchmarking and model\-comparison practices; fields with weaker norms around sharing baselines or reporting computational budgets are correspondingly less visible, which limits how far our conclusions can be generalized beyond the domains represented here\. Our method taxonomy simplifies a continuum of architectures, and our domain clustering aggregates heterogeneous subfields\. We also focus on predictive performance and computational cost, not on other dimensions of scientific value such as interpretability, calibration, robustness, or the capacity to generate novel hypotheses\.

These limitations point to clear priorities for future work\. First, journals and conferences could require standardized reporting of computational budgets and hardware for all methods, including baselines, enabling more systematic cost\-aware evaluation\[[33](https://arxiv.org/html/2609.16258#bib.bib42),[17](https://arxiv.org/html/2609.16258#bib.bib43)\]\. Second, richer datasets are needed to track hybrid workflows in which AI accelerates or guides parts of numerical pipelines rather than replacing them outright\. Third, predictive models linking problem characteristics, such as dimensionality, data volume, physical structure, and distribution shift, to the relative advantage of different computational approaches would help researchers and funders identify where AI is most likely to expand the scientific frontier\.

Taken together, our findings support a more precise view of AI in science than either triumphalist or skeptical accounts allow\. AI is neither a universal replacement for existing methods nor merely another tool within an unchanged landscape\. It is a method family that increasingly opens new regions of the scientific cost–performance frontier: by purchasing performance relative to traditional statistics, and by reducing cost relative to simulation\-heavy incumbents\. The central question, therefore, is not whether AI will replace existing approaches everywhere, but where its inductive biases, computational profile, and surrogate capabilities allow it to create genuinely new possibilities for scientific analysis\.

## Acknowledgments

We acknowledge support from MIT’s UROP Program and the contributions of the master’s students who assisted with early\-stage data collection: Caroline C\. Warren and Haley Nakamura\.

##### Funding:

This work was supported by Open Philanthropy/Good Ventures\.

##### Data and materials availability:

The full dataset and analysis code are available in the[GitHub](https://github.com/MIT-FutureTech/ia-enabled-scientific-frontier)repository\. All data needed to evaluate the conclusions in the paper are present in the paper or the Supplementary Materials\.

## References and Notes

- \[1\]J\. Abramson, J\. Adler, J\. Dunger, R\. Evans, T\. Green, A\. Pritzel, O\. Ronneberger, L\. Willmore, A\. J\. Ballard, J\. Bambrick,et al\.\(2024\)Accurate structure prediction of biomolecular interactions with alphafold 3\.Nature630\(8016\),pp\. 493–500\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07487-w)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[2\]A\. Ansuini, A\. Laio, J\. H\. Macke, and D\. Zoccolan\(2019\)Intrinsic dimension of data representations in deep neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.32\.External Links:[Link](https://arxiv.org/abs/1905.12784)Cited by:[Domain\-specific effectiveness and the migration of effort](https://arxiv.org/html/2609.16258#Sx1.SSx3.p3.1)\.
- \[3\]K\. Azizzadenesheli, N\. Kovachki, Z\. Li, M\. Liu\-Schiaffini, J\. Kossaifi, and A\. Anandkumar\(2024\)Neural operators for accelerating scientific simulations and design\.Nature Reviews Physics6\(5\),pp\. 320–328\.External Links:[Document](https://dx.doi.org/10.1038/s42254-024-00712-5),[Link](https://doi.org/10.1038/s42254-024-00712-5)Cited by:[What happens when AI replaces existing methods?](https://arxiv.org/html/2609.16258#Sx1.SSx2.p5.1)\.
- \[4\]T\. Besiroglu, N\. Emery\-Xu, and N\. Thompson\(2024\)Economic impacts of ai\-augmented r&d\.Research Policy53\(7\),pp\. 105037\.External Links:ISSN 0048\-7333,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.respol.2024.105037),[Link](https://www.sciencedirect.com/science/article/pii/S0048733324000866)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p5.1)\.
- \[5\]S\. Bianchini, M\. Müller, and P\. Pelletier\(2022\)Artificial intelligence in science: an emerging general method of invention\.Research Policy51\(10\),pp\. 104604\.External Links:[Document](https://dx.doi.org/10.1016/j.respol.2022.104604)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[6\]C\. Bodnar, W\. P\. Bruinsma, A\. Lucic, M\. Stanley, A\. Allen, J\. Brandstetter, P\. Garvan, M\. Riechert, J\. A\. Weyn, H\. Dong, J\. K\. Gupta, K\. Thambiratnam, A\. T\. Archibald, C\. Wu, E\. Heider, M\. Welling, R\. E\. Turner, P\. Perdikaris,et al\.\(2025\)A foundation model for the earth system\.Nature641\(8065\),pp\. 1180–1187\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09005-y)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1)\.
- \[7\]D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes\(2023\)Autonomous chemical research with large language models\.Nature624\(7992\),pp\. 570–578\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06792-0)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p5.1)\.
- \[8\]L\. Breiman\(2001\)Statistical modeling: the two cultures \(with comments and a rejoinder by the author\)\.Statistical Science16\(3\),pp\. 199–231\.External Links:[Document](https://dx.doi.org/10.1214/ss/1009213726),[Link](https://doi.org/10.1214/ss/1009213726)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p6.1)\.
- \[9\]E\. Christodoulou, J\. Ma, G\. S\. Collins, E\. W\. Steyerberg, J\. Y\. Verbakel, and B\. Van Calster\(2019\)A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models\.Journal of Clinical Epidemiology110,pp\. 12–22\.External Links:[Document](https://dx.doi.org/10.1016/j.jclinepi.2019.02.004),[Link](https://doi.org/10.1016/j.jclinepi.2019.02.004)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1)\.
- \[10\]R\. W\. Clough\(1960\)The finite element method in plane stress analysis\.InProceedings of the 2nd ASCE Conference on Electronic Computation,Pittsburgh, PA,pp\. 345–378\.Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[11\]J\. W\. Cooley and J\. W\. Tukey\(1965\)An algorithm for the machine calculation of complex fourier series\.Mathematics of Computation19\(90\),pp\. 297–301\.External Links:[Document](https://dx.doi.org/10.1090/S0025-5718-1965-0178586-1)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[12\]B\. Cottier, R\. Rahman, L\. Fattorini, N\. Maslej, T\. Besiroglu, and D\. Owen\(2024\)The rising costs of training frontier ai models\.arXiv preprint arXiv:2405\.21015\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.21015),[Link](https://arxiv.org/abs/2405.21015)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.
- \[13\]K\. Davagdorj, J\. Bae, V\. Pham, N\. Theera\-Umpon, and K\. H\. Ryu\(2021\)Explainable artificial intelligence based framework for non\-communicable diseases prediction\.IEEE Access9,pp\. 123672–123688\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2021.3110336)Cited by:[2nd item](https://arxiv.org/html/2609.16258#Sx1.I1.i2.p1.1)\.
- \[14\]B\. Ebiwonjumi, A\. Cherezov, S\. Dzianisau, and D\. Lee\(2021\)Machine learning of LWR spent nuclear fuel assembly decay heat measurements\.Nuclear Engineering and Technology53\(11\),pp\. 3563–3579\.External Links:[Document](https://dx.doi.org/10.1016/j.net.2021.05.037)Cited by:[4th item](https://arxiv.org/html/2609.16258#Sx1.I1.i4.p1.1)\.
- \[15\]D\. George and E\. A\. Huerta\(2018\)Deep neural networks to enable real\-time multimessenger astrophysics\.Physical Review D97\(4\),pp\. 044039\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevD.97.044039)Cited by:[3rd item](https://arxiv.org/html/2609.16258#Sx1.I1.i3.p1.1)\.
- \[16\]L\. Grinsztajn, E\. Oyallon, and G\. Varoquaux\(2022\)Why do tree\-based models still outperform deep learning on tabular data?\.arXiv preprint arXiv:2207\.08815\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2207.08815),[Link](https://arxiv.org/abs/2207.08815)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p2.1)\.
- \[17\]P\. Henderson, J\. Hu, J\. Romoff, E\. Brunskill, D\. Jurafsky, and J\. Pineau\(2020\)Towards the systematic reporting of the energy and carbon footprints of machine learning\.Journal of Machine Learning Research21\(248\),pp\. 1–43\.External Links:[Link](https://jmlr.org/papers/v21/20-312.html)Cited by:[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p5.1)\.
- \[18\]B\. F\. Jones\(2025\)Artificial intelligence in research and development\.Working PaperTechnical Report34312,National Bureau of Economic Research\.External Links:[Document](https://dx.doi.org/10.3386/w34312)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p5.1)\.
- \[19\]J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko, A\. Bridgland, C\. Meyer, S\. A\. A\. Kohl, A\. J\. Ballard, A\. Cowie, B\. Romera\-Paredes, S\. Nikolov, R\. Jain, J\. Adler, T\. Back, S\. Petersen, D\. Reiman, E\. Clancy, M\. Zielinski, M\. Steinegger, M\. Pacholska, T\. Berghammer, S\. Bodenstein, D\. Silver, O\. Vinyals, A\. W\. Senior, K\. Kavukcuoglu, P\. Kohli, and D\. Hassabis\(2021\)Highly accurate protein structure prediction with AlphaFold\.Nature596\(7873\),pp\. 583–589\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03819-2),[Link](https://doi.org/10.1038/s41586-021-03819-2)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[20\]A\. Karpathy\(2026\)Autoresearch\.Note:GitHub repositoryExternal Links:[Link](https://github.com/karpathy/autoresearch)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p5.1)\.
- \[21\]P\. W\. Koh, S\. Sagawa, H\. Marklund, S\. M\. Xie, M\. Zhang, A\. Balsubramani, W\. Hu, M\. Yasunaga, R\. L\. Phillips, S\. Beery,et al\.\(2021\)WILDS: a benchmark of in\-the\-wild distribution shifts\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 5637–5664\.External Links:[Link](https://proceedings.mlr.press/v139/koh21a.html)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.
- \[22\]R\. Lam, Á\. Sanchez\-Gonzalez, M\. Willson, P\. Wirnsberger, M\. Fortunato, F\. Alet, S\. Ravuri, T\. Ewalds, Z\. Eaton\-Rosen, W\. Hu, A\. Merose,et al\.\(2023\)Learning skillful medium\-range global weather forecasting\.Science382\(6677\),pp\. 1416–1421\.External Links:[Document](https://dx.doi.org/10.1126/science.adi2336)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1)\.
- \[23\]T\. Liao, R\. Taori, I\. D\. Raji, and L\. Schmidt\(2021\)Are we learning yet? a meta review of evaluation failures across machine learning\.InThirty\-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track \(Round 2\),External Links:[Link](https://openreview.net/forum?id=mPducS1MsEK)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.
- \[24\]S\. Makridakis, E\. Spiliotis, V\. Assimakopoulos, A\. A\. Semenoglou, G\. Mulder, and K\. Nikolopoulos\(2023\)Statistical, machine learning and deep learning forecasting methods: comparisons and ways forward\.Journal of the Operational Research Society74\(3\),pp\. 840–859\.External Links:[Document](https://dx.doi.org/10.1080/01605682.2022.2118629)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p2.1)\.
- \[25\]S\. Makridakis, E\. Spiliotis, and V\. Assimakopoulos\(2018\)Statistical and machine learning forecasting methods: concerns and ways forward\.PLOS ONE13\(3\),pp\. e0194889\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0194889)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p2.1)\.
- \[26\]L\. Messeri and M\. J\. Crockett\(2024\)Artificial intelligence and illusions of understanding in scientific research\.Nature627\(8002\),pp\. 49–58\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07146-0)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.
- \[27\]A\. Narayanan and S\. Kapoor\(2025\)AI as normal technology: an alternative to the vision of ai as a potential superintelligence\.Note:Knight First Amendment Institute essayExternal Links:[Link](https://knightcolumbia.org/content/ai-as-normal-technology)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[28\]A\. Narayanan and S\. Kapoor\(2025\)Why an overreliance on ai\-driven modelling is bad for science\.Nature640\(8058\),pp\. 312–314\.External Links:[Document](https://dx.doi.org/10.1038/d41586-025-01067-2),[Link](https://www.nature.com/articles/d41586-025-01067-2)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[29\]P\. Pope, C\. Zhu, A\. Abdelkader, M\. Goldblum, and T\. Goldstein\(2021\)The intrinsic dimension of images and its impact on learning\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=XJk19XzGq2J)Cited by:[Domain\-specific effectiveness and the migration of effort](https://arxiv.org/html/2609.16258#Sx1.SSx3.p3.1)\.
- \[30\]I\. Price, Á\. Sanchez\-Gonzalez, F\. Alet, T\. R\. Andersson, A\. El\-Kadi, D\. Masters, T\. Ewalds, J\. Stott, S\. Mohamed, P\. Battaglia, R\. Lam, M\. Willson,et al\.\(2025\)Probabilistic weather forecasting with machine learning\.Nature637\(8044\),pp\. 84–90\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08252-9)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1)\.
- \[31\]A\. Rahman\(1964\)Correlations in the motion of atoms in liquid argon\.Physical Review136\(2A\),pp\. A405–A411\.External Links:[Document](https://dx.doi.org/10.1103/PhysRev.136.A405)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[32\]M\. Reichstein, G\. Camps\-Valls, B\. Stevens, M\. Jung, J\. Denzler, N\. Carvalhais, and Prabhat\(2019\)Deep learning and process understanding for data\-driven earth system science\.Nature566\(7743\),pp\. 195–204\.External Links:[Document](https://dx.doi.org/10.1038/s41586-019-0912-1),[Link](https://doi.org/10.1038/s41586-019-0912-1)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p2.1)\.
- \[33\]R\. Schwartz, J\. Dodge, N\. A\. Smith, and O\. Etzioni\(2020\)Green ai\.Communications of the ACM63\(12\),pp\. 54–63\.External Links:[Document](https://dx.doi.org/10.1145/3381831)Cited by:[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p5.1)\.
- \[34\]D\. Sculley, G\. Holt, D\. Golovin, E\. Davydov, T\. Phillips, D\. Ebner, V\. Chaudhary, M\. Young, J\. Crespo, and D\. Dennison\(2015\)Hidden technical debt in machine learning systems\.InAdvances in Neural Information Processing Systems,Vol\.28,pp\. 2503–2511\.External Links:[Link](https://proceedings.neurips.cc/paper/5656-hidden-technical-debt-in-machine-learning-sy)Cited by:[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p2.1)\.
- \[35\]R\. Shwartz\-Ziv and A\. Armon\(2022\)Tabular data: deep learning is not all you need\.Information Fusion81,pp\. 84–90\.External Links:[Document](https://dx.doi.org/10.1016/j.inffus.2021.11.011)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1),[Conclusion](https://arxiv.org/html/2609.16258#Sx2.p2.1)\.
- \[36\]N\. J\. Szymanski, B\. Rendy, Y\. Fei, R\. E\. Kumar, T\. He, D\. Milsted, M\. J\. McDermott, M\. C\. Gallant, E\. D\. Cubuk, A\. Merchant, H\. Kim, A\. Jain, C\. J\. Bartel, K\. A\. Persson, Y\. Zeng, and G\. Ceder\(2023\)An autonomous laboratory for the accelerated synthesis of inorganic materials\.Nature624\(7990\),pp\. 86–91\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06734-w)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p5.1)\.
- \[37\]The Royal Society\(2024\)Science in the age of ai: how artificial intelligence is changing the nature and method of scientific research\.Technical reportThe Royal Society,London, UK\.External Links:[Link](https://royalsociety.org/-/media/policy/projects/science-in-the-age-of-ai/science-in-the-age-of-ai-report.pdf)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[38\]J\. Thiyagalingam, M\. Shankar, G\. Fox, and T\. Hey\(2022\)Scientific machine learning benchmarks\.Nature Reviews Physics4\(6\),pp\. 413–420\.External Links:[Document](https://dx.doi.org/10.1038/s42254-022-00441-7)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p4.1)\.
- \[39\]N\. C\. Thompson, S\. Ge, and G\. F\. Manso\(2022\)The importance of \(exponentially more\) computing power\.External Links:2206\.14007,[Link](https://arxiv.org/abs/2206.14007)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[40\]N\. C\. Thompson, K\. Greenewald, K\. Lee, and G\. F\. Manso\(2023\)The computational limits of deep learning\.InProceedings of the 2023 ACM Conference on Computing and Sustainable Societies \(LIMITS ’23\),External Links:[Document](https://dx.doi.org/10.21428/bf6fb269.1f033948),[Link](https://doi.org/10.21428/bf6fb269.1f033948)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.
- \[41\]J\. Tubiana, D\. Schneidman\-Duhovny, and H\. J\. Wolfson\(2022\)ScanNet: an interpretable geometric deep learning model for structure\-based protein binding site prediction\.Nature Methods19\(6\),pp\. 730–739\.External Links:[Document](https://dx.doi.org/10.1038/s41592-022-01490-7)Cited by:[1st item](https://arxiv.org/html/2609.16258#Sx1.I1.i1.p1.1)\.
- \[42\]H\. Wang, T\. Fu, Y\. Du, W\. Gao, K\. Huang, Z\. Liu, P\. Chandak, S\. Liu, P\. Van Katwyk, A\. Deac, A\. Anandkumar, K\. Bergen, C\. P\. Gomes, S\. Ho, P\. Kohli, J\. Lasenby, J\. Leskovec, T\. Liu, A\. Manrai, D\. S\. Marks, B\. Ramsundar, L\. Song, J\. Sun, J\. Tang, P\. Velickovic, M\. Welling, L\. Zhang, C\. W\. Coley, Y\. Bengio, and M\. Zitnik\(2023\)Scientific discovery in the age of artificial intelligence\.Nature620\(7972\),pp\. 47–60\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06221-2)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p1.1)\.
- \[43\]D\. H\. Wolpert and W\. G\. Macready\(1997\)No free lunch theorems for optimization\.IEEE Transactions on Evolutionary Computation1\(1\),pp\. 67–82\.External Links:[Document](https://dx.doi.org/10.1109/4235.585893)Cited by:[Introduction](https://arxiv.org/html/2609.16258#Sx1.p3.1)\.

## Supplementary Materials for The AI\-Enabled Scientific Frontier

Gabriel Manso1, Emma Fu1, Neil Thompson1∗ 1MIT CSAIL, MIT FutureTech ∗Corresponding author\. Email: neil\_t@mit\.edu

#### This PDF file includes:

S1\.Corpus Construction and Literature Search\.[1](https://arxiv.org/html/2609.16258#S1) S2\.Data Extraction and Validation\.[2](https://arxiv.org/html/2609.16258#S2) S3\.Method Taxonomy\.[3](https://arxiv.org/html/2609.16258#S3) S4\.Domain Consolidation\.[4](https://arxiv.org/html/2609.16258#S4) S5\.Metric Harmonization\.[5](https://arxiv.org/html/2609.16258#S5) S6\.Data Completeness\.[6](https://arxiv.org/html/2609.16258#S6) S7\.Statistical Analyses\.[7](https://arxiv.org/html/2609.16258#S7) S8\.Sensitivity and Robustness\.[8](https://arxiv.org/html/2609.16258#S8) Tables S1–S12

## 1Corpus Construction and Literature Search

We assembled a corpus of published studies that reported direct same\-task comparisons between computational methods in scientific applications\. The central inclusion principle was comparability\. Each retained study had to evaluate at least two methods on a shared problem, dataset, or experimental setting, thereby permitting a meaningful assessment of relative performance and computational cost\. By restricting the corpus to explicit head\-to\-head evaluations within the same task environment, we isolate the empirical contexts in which AI is not merely proposed, but directly tested against incumbent computational approaches\.

### 1\.1Databases and Sources

Candidate papers were identified through systematic queries of five databases: IEEE Xplore, ACM Digital Library, arXiv, Google Scholar, and Scopus\. Each query combined three components: domain anchor terms, AI method terms, and comparison terms\. For each of the 27 disciplines, two to five query variants were constructed to capture differences in field\-specific vocabulary while preserving a common search logic\. Representative examples are shown in Table[S1](https://arxiv.org/html/2609.16258#S1.T1)\.

Table S1:Representative search queries used during corpus construction\.DomainRepresentative QueryFluid Dynamics\(“CFD” OR “turbulence modeling”\) AND \(“neural network” OR “surrogate model” OR “deep learning”\) AND \(“benchmark” OR “speedup” OR “comparison”\)Medicine\(“medical imaging” OR “clinical prediction”\) AND \(“deep learning” OR “convolutional neural network” OR “CNN”\) AND \(“versus” OR “benchmark” OR “outperform” OR “underperform”\)Atmospheric Science\(“weather prediction” OR “numerical weather prediction”\) AND \(“neural network” OR “transformer” OR “machine learning”\) AND \(“forecast skill” OR “computational cost” OR “comparison”\)Materials Science\(“molecular dynamics” OR “density functional theory” OR “DFT”\) AND \(“neural network” OR “machine learning potential” OR “ML potential”\) AND \(“accuracy” OR “speedup” OR “benchmark”\)Finance\(“financial forecasting” OR “stock prediction”\) AND \(“deep learning” OR “LSTM” OR “neural network”\) AND \(“versus” OR “ARIMA” OR “benchmark”\)
### 1\.2Inclusion and Exclusion Criteria

A paper entered the screened corpus \(N=880N=880\) if it satisfied two core criteria\. First, it had to present at least one explicit comparison between two or more computational methods on the same task\. Second, the paper had to address a scientific or scientific\-computing task and be available through a peer\-reviewed venue or recognized preprint server\.

Papers were promoted to the final dataset \(N=268N=268papers;N=2,507N=2\{,\}507comparisons\) if they additionally reported at least one performance metric for both methods and at least one computational cost indicator, whether quantitative or qualitative\.

Papers were excluded if the comparison was purely theoretical, if only a single method was evaluated, if the task was outside the scope of scientific analysis, if the reported evidence was insufficient even for directional classification, or if the paper was a review rather than an original empirical study\.

## 2Data Extraction and Validation

### 2\.1Extraction Protocol

The canonical unit of observation is one method pair evaluated on one dataset or experimental setting\. This definition prevents metric multiplication within the same experiment while preserving meaningful variation across datasets, benchmarks, or environmental conditions\. When a paper reported multiple metrics for the same comparison, the metric most standard to the relevant domain was selected as the primary performance indicator\. The same pair of methods evaluated on distinct datasets or settings was treated as separate observations\.

For each comparison, the extraction template recorded paper metadata, domain classification, baseline and proposed method names, method\-family assignments, performance metric name and values, cost metric name and values, directional labels for performance and cost, the source of any qualitative cost judgment, and free\-text notes for edge cases\.

To maximize consistency, information was extracted according to a fixed evidence hierarchy\. Highest priority was given to tables containing numerical values, followed by explicit numerical statements in the main text, then figures, and finally qualitative author statements used only for directional classification when numerical cost data were unavailable\.

### 2\.2LLM\-Assisted Validation

To reduce extraction errors at scale while avoiding automation bias, we implemented amanual\-first, LLM\-secondvalidation pipeline\. All comparisons were initially identified and extracted by human reviewers\. The LLM system \(GPT\-4\) was used only as a secondary verification layer and never introduced new observations or replaced human judgments without explicit review\.

The validation process proceeded in six phases\. First, the model mapped the structure of each paper, identifying relevant tables, figures, and result sections\. Second, it enumerated candidate method comparisons within those sections\. Third, it extracted the supporting evidence associated with each comparison, including reported metrics and contextual statements\. Fourth, numerical values were cross\-checked against the original source to verify consistency with the recorded entries\. Fifth, the model classified whether its interpretation agreed with the human extraction\. Finally, any disagreements were reviewed and adjudicated by a human annotator before the comparison was retained in the dataset\.

## 3Method Taxonomy

All computational methods in the corpus were classified into one of three families, corresponding to the broad methodological categories discussed in the main text and positioned along the cost–performance frontier in Figure 1\.

- •Traditional statistics \(TS\)\.Algorithms that extract patterns from empirical data without relying on neural network architectures\. Representative examples include support vector machines, random forests, gradient boosting, logistic regression,kk\-nearest neighbors, Gaussian processes, and principal components methods\. These approaches are often computationally lighter than modern neural models, although performance depends strongly on the fit between their inductive biases and the underlying problem structure\.
- •Scientific computing \(SC\)\.Methods whose algorithmic structure is derived from embedded mathematical, physical, or mechanistic principles\. Examples include Navier–Stokes solvers, finite element methods, molecular dynamics, computational fluid dynamics simulations, and quantum chemistry calculations\. These approaches are often computationally demanding but are typically grounded in explicit model structure, constraints, or convergence properties\.
- •AI\.Methods whose primary representational mechanism is a neural network architecture\. Examples include convolutional, recurrent, and transformer networks, graph neural networks, physics\-informed neural networks, neural operators, and related generative or surrogate\-model architectures\. These methods typically rely on large numbers of learned parameters optimized through gradient\-based training\.

### 3\.1Replacement vs\. Augmentation

Beyond the family labels, cross\-family AI comparisons were further classified according to whether AI replaced an incumbent method or augmented it\.

- •Replacement\.The proposed AI method operates independently and is evaluated as a substitute for the incumbent method\.
- •Augmentation\.The incumbent computational pipeline remains in place, but an AI component is introduced to improve or accelerate one part of the workflow\.

This distinction is important for interpretation\. Replacement cases align most directly with the main paper’s central question of whether AI occupies a different position on the cost–performance frontier than the methods it displaces, whereas augmentation cases more often reflect targeted optimization within an existing pipeline\.

Of the2,5072\{,\}507comparisons in the full dataset,1,2011\{,\}201are cross\-family AI comparisons\. Among these,1,1221\{,\}122\(93\.4%\) are classified asreplacementand7979\(6\.6%\) asaugmentation\. The replacement set comprises825825TS→\\toAI comparisons and297297SC→\\toAI comparisons\. The remaining observations are within\-family comparisons or other non\-AI cross\-family substitutions\. The full distribution is shown in Table[S2](https://arxiv.org/html/2609.16258#S3.T2)\.

Table S2:Distribution of comparison types in the full dataset \(N=2,507N=2\{,\}507\)\.Comparison TypeCount \(%\)Cross\-family AI replacementTraditional Statistics→\\toAI825 \(32\.9%\)Scientific Computing→\\toAI297 \(11\.8%\)Other cross\-familyScientific Computing→\\toTraditional Statistics149 \(5\.9%\)Within\-familyAI→\\toAI661 \(26\.4%\)Traditional Statistics→\\toTraditional Statistics392 \(15\.6%\)Scientific Computing→\\toScientific Computing43 \(1\.7%\)Augmentation \(hybrid\)SC→\\toSC\+AI65 \(2\.6%\)SC→\\toSC\+TS29 \(1\.2%\)SC\+AI→\\toSC\+AI21 \(0\.8%\)SC\+TS→\\toSC\+AI14 \(0\.6%\)SC\+TS→\\toSC\+TS11 \(0\.4%\)Total2,507 \(100\.0%\)

## 4Domain Consolidation: 27 Disciplines to 7 Clusters

To support cross\-domain comparison without collapsing all scientific areas into a single undifferentiated group, we consolidated the 27 fine\-grained disciplines into seven broader clusters\. This mapping preserved meaningful domain structure while aligning fields that share similar computational tasks, data modalities, or methodological baselines \(see Table[S3](https://arxiv.org/html/2609.16258#S4.T3)\)\. The resulting clustering was used throughout the descriptive analyses in the supplement and main text\.

Table S3:Domain consolidation into seven clusters with representative characteristics and observation counts\.ClusterConstituent DomainsShared CharacteristicsNNPhysical ModelingSciencesPhysics, Chemistry, Materials Science, Fluid Dynamics, Nuclear Engineering, Energy Systems, Optics and PhotonicsFirst\-principles simulations, continuous optimization, high\-dimensional regression970LifeSciencesBiology, Bioinformatics, Healthcare, MedicineBiological sequences, clinical records, imaging, heterogeneous biological data520Earth & SpaceSciencesAstronomy, Astrophysics, Atmospheric Science, Climate Science, Earth Science, Geophysics, HydrologyObservational time series, remote sensing, spatiotemporal modeling, physical constraints328EconomicSciencesFinance, EconomicsFinancial time series, nonstationarity, strategic or adversarial dynamics285EngineeringControl Systems, Robotics, Industrial Engineering, ManufacturingControl signals, sensor streams, real\-time decision systems166Environmental &Agricultural SciencesAgricultural and Food Science, Environmental ScienceField\-scale observations, ecosystem modeling, resource management problems78OtherInterdisciplinary and cross\-domain studiesHeterogeneous computational problems not fitting the primary disciplinary categories160Total2,507
## 5Metric Harmonization

The dataset spans a highly heterogeneous measurement landscape, with more than 284 distinct performance metrics and 239 distinct computational cost metrics reported across 27 scientific domains\. Because these metrics differ widely in scale, interpretation, and reporting convention, direct aggregation across studies is not generally meaningful\. To preserve comparability while retaining information about the direction and, when available, the magnitude of change, we employ two complementary encoding systems\.

### 5\.1Binary Encoding

For every comparison, both performance and cost are assigned directional labels:

Performance binary=\{\+1new method achieves better predictive performance0equivalent within a negligible margin−1new method achieves worse predictive performance\\text\{Performance binary\}=\\begin\{cases\}\+1&\\text\{new method achieves better predictive performance\}\\\\ \\phantom\{\+\}0&\\text\{equivalent within a negligible margin\}\\\\ \-1&\\text\{new method achieves worse predictive performance\}\\end\{cases\}\(S1\)
Cost binary=\{\+1new method requires greater computational resources0equivalent−1new method requires fewer computational resources\\text\{Cost binary\}=\\begin\{cases\}\+1&\\text\{new method requires greater computational resources\}\\\\ \\phantom\{\+\}0&\\text\{equivalent\}\\\\ \-1&\\text\{new method requires fewer computational resources\}\\end\{cases\}\(S2\)
This directional encoding covers all 2,507 comparisons\. To ensure consistency, we validated all qualitative binary labels against their corresponding quantitative log\-ratios\. For ’lower\-is\-better’ performance cases, where a numerical decrease denotes a performance improvement, we inverted the corresponding log\-ratios to allow for direct comparison\.

### 5\.2Log\-Ratio Encoding

When numerical values are available, we compute oriented log\-ratio encodings:

ΔperfX:Y=log10\(PXPY\),ΔcostX:Y=log10\(CXCY\)\\Delta^\{X:Y\}\_\{\\text\{perf\}\}=\\log\_\{10\}\\\!\\left\(\\frac\{P\_\{X\}\}\{P\_\{Y\}\}\\right\),\\qquad\\Delta^\{X:Y\}\_\{\\text\{cost\}\}=\\log\_\{10\}\\\!\\left\(\\frac\{C\_\{X\}\}\{C\_\{Y\}\}\\right\)\(S3\)
whereYYdenotes the baseline method,XXthe proposed method,PPa performance metric oriented so that positive values indicate improvement, andCCa computational cost metric oriented so that positive values indicate increased cost\.

### 5\.3Outcome Categories

Each comparison is also assigned to one of four mutually exclusive quadrants defined by the joint binary performance and cost outcomes:

- •Win\-Win \(WW\)\.Performance is equal or better and cost is equal or lower\.
- •Performance Prioritization \(PP\)\.Performance improves, but cost increases\.
- •Efficiency Prioritization \(EP\)\.Performance worsens, but cost decreases or remains unchanged\.
- •Lose\-Lose \(LL\)\.Performance worsens or remains unchanged while cost increases\.

Ties \(binary=0=0\) are grouped with favorable outcomes where they arise: performance ties with cost reduction \(N=36N=36\) count as Win\-Win; performance ties with cost increases \(N=16N=16\) count as Lose\-Lose; joint ties \(00/00;N=2N=2\) count as Win\-Win\.

## 6Data Completeness and Selective Reporting

Because the corpus draws on heterogeneous literatures with different reporting conventions, the availability of quantitative information varies across comparisons\. Predictive performance metrics are usually reported numerically, whereas computational cost is often omitted or described only qualitatively\. Table[S4](https://arxiv.org/html/2609.16258#S6.T4)summarizes this difference for both directional and quantitative encodings\.

Table S4:Data completeness by metric type \(N=2,507N=2\{,\}507\)\.Metric TypeCountPercentageBinary \(directional\) comparisonsPerformance2,507100\.0%Cost2,507100\.0%Quantitative \(log\-ratio\) comparisonsPerformance2,17086\.6%Cost1,13545\.3%Both99539\.7%The 41\.3\-percentage\-point gap between quantitative performance and quantitative cost reporting is substantively important for interpretation\. This gap means that analyses based on quantitative cost ratios necessarily draw on a more restricted subset of the literature than the directional analysis emphasized in the main text\.

### 6\.1Cost Reporting Rates by Domain Cluster

The availability of quantitative cost data also varies across scientific domains\. Some areas routinely report training time, runtime, or hardware usage, whereas others provide only qualitative descriptions of computational burden\. Table[S5](https://arxiv.org/html/2609.16258#S6.T5)summarizes the share of comparisons with numerical cost measurements within each domain cluster\.

Table S5:Quantitative cost reporting rates by domain cluster \(full dataset,N=2,507N=2\{,\}507\)\.ClusterNtotalN\_\{\\text\{total\}\}Nquant\. costN\_\{\\text\{quant\.\\ cost\}\}RateLife Sciences52027152\.1%Engineering1668651\.8%Physical Modeling Sci\.97046848\.2%Environ\. & Agri\. Sci\.782734\.6%Earth & Space Sci\.32810832\.9%Economic Sciences2853612\.6%Other16013986\.9%Overall2,5071,13545\.3%These rates suggest that quantitative cost analysis is more feasible in some literatures than in others\. This heterogeneity does not invalidate cross\-domain comparison, but it does mean that numerical cost ratios should be interpreted with particular care in underreported fields\.

### 6\.2Selective Cost Reporting Test

A potential concern in compiled datasets where computational cost is not universally reported is selective reporting\. In this context, authors might be more likely to report quantitative cost metrics when the proposed AI method appears computationally favorable, while omitting such values when models impose substantial computational burdens\.

To assess this possibility, we tested whether the availability of quantitative cost data depends on the reported performance outcome of the proposed AI method\. Using the AI replacement sample \(N=1,122N=1\{,\}122\), we conducted a chi\-squared test of independence between AI performance outcome \(superior versus inferior/equivalent\) and the presence of numerical cost data\.

The test yieldsχ2=0\.055\\chi^\{2\}=0\.055\(d​f=1df=1,p=0\.814p=0\.814\)\. This result provides no evidence that the availability of quantitative cost measurements depends on the reported performance outcome of the proposed method\. It does not eliminate all possible reporting biases, but it suggests that the quantitative cost subsample is unlikely to be systematically biased along this particular dimension\.

## 7Statistical Analyses

The analyses in this section focus primarily on the subset of comparisons classified asAI replacement\(N=1,122N=1\{,\}122\), isolating settings in which a neural network–based method is evaluated as a direct substitute for an incumbent computational approach belonging to either Traditional Statistics \(TS\) or Scientific Computing \(SC\)\. Restricting attention to this subset removes hybrid augmentation cases and allows the empirical structure of direct methodological substitution to be examined more cleanly\. The section is primarily descriptive, with formal inferential summaries used only to characterize association structure in the compiled literature rather than to support causal claims\. Throughout, we connect the statistical patterns reported here to the conceptual argument developed in the main text while recognizing that the underlying evidence comes from heterogeneous published studies rather than a controlled benchmark environment\.

### 7\.1Baseline Outcome Proportions

We begin with a descriptive summary of the outcome distribution\. Table[S6](https://arxiv.org/html/2609.16258#S7.T6)reports the four outcome quadrants together with 95% Wilson score confidence intervals, which provide more reliable coverage for bounded proportions than normal approximations\.

Table S6:AI replacement outcome proportions \(N=1,122N=1\{,\}122\) with 95% Wilson confidence intervals\.OutcomennNNProportion95% CI lower95% CI upperWin\-Win2371,12221\.1%18\.8%23\.6%Performance Prioritization5121,12245\.6%42\.7%48\.6%Efficiency Prioritization1561,12213\.9%12\.0%16\.1%Lose\-Lose2171,12219\.3%17\.1%21\.8%Across the corpus, the most frequent outcome when AI replaces an incumbent computational method isPerformance Prioritization, in which predictive accuracy improves while computational cost increases \(45\.6% \[42\.7%, 48\.6%\]\)\. Strict improvements in both dimensions \(Win\-Win\) occur in roughly one fifth of cases \(21\.1% \[18\.8%, 23\.6%\]\)\. The remaining observations are divided betweenEfficiency Prioritizationoutcomes \(13\.9%\) andLose\-Loseoutcomes \(19\.3%\)\.

This descriptive pattern is consistent with the main text’s broader claim that AI does not occupy a single fixed position on the cost–performance plane\. In the aggregate replacement sample, neural architectures most often improve predictive capability, but these gains frequently come with added computational burden rather than universal two\-dimensional improvement\.

For contrast, we also summarize the much smaller set of augmentation cases, in which AI components are incorporated into existing workflows rather than replacing them outright\.

Table S7:AI augmentation outcome proportions \(N=79N=79\) with 95% Wilson confidence intervals\.OutcomennNNProportion95% CI lower95% CI upperWin\-Win557969\.6%58\.8%78\.7%Performance Prioritization157919\.0%11\.9%29\.0%Efficiency Prioritization6797\.6%3\.5%15\.6%Lose\-Lose3793\.8%1\.3%10\.6%Augmentation scenarios show a markedly different profile\. Nearly seventy percent of cases fall into the Win\-Win category, suggesting that AI components are often introduced to alleviate localized bottlenecks within established computational pipelines\. This contrast does not imply that augmentation is universally preferable, since the sample is small and structurally different from the replacement set, but it does reinforce the distinction drawn in the main text between substituting an incumbent method and selectively enhancing one\.

### 7\.2Outcome Structure by Baseline Type

A central claim of the main paper is that the empirical role of AI depends strongly on the type of baseline method it replaces\. We evaluate that claim directly by stratifying the replacement sample according to whether the incumbent method belongs to Traditional Statistics or Scientific Computing\.

Table S8:AI replacement outcomes stratified by baseline method type, with 95% Wilson confidence intervals\.BaselineWin\-WinPerf\.\-Prior\.Eff\.\-Prior\.Lose\-LoseTS \(N=825N=825\)13\.9% \[11\.7, 16\.5\]59\.6% \[56\.3, 62\.9\]2\.8% \[1\.9, 4\.1\]23\.6% \[20\.9, 26\.7\]SC \(N=297N=297\)41\.1% \[35\.6, 46\.8\]6\.7% \[4\.4, 10\.2\]44\.8% \[39\.2, 50\.5\]7\.4% \[4\.9, 11\.0\]Two distinct substitution regimes emerge\. When AI replaces traditional statistical methods, outcomes are dominated by the Performance Prioritization quadrant \(59\.6%\), indicating that neural models often achieve higher predictive performance at increased computational cost\. By contrast, when AI replaces scientific simulation methods, the majority of observations fall into the Win\-Win \(41\.1%\) or Efficiency Prioritization \(44\.8%\) regions\.

These descriptive contrasts closely match the interpretation advanced in the main text\. Relative to TS baselines, AI most often behaves as a more computationally intensive predictive model\. Relative to SC baselines, AI more often functions as an empirical surrogate for computationally expensive simulation pipelines\. At the same time, the nontrivial presence of Lose\-Lose outcomes in both groups indicates that neither regime is universal and that substitution remains application dependent\.

### 7\.3Cost and Performance Effect Sizes by Baseline

The quadrant summaries above describe directional trade\-offs\. We next examine the magnitude of those changes using the quantitative log\-ratio encodings available for the subset of comparisons with numerical performance and cost values\.

Table S9:Quantitative effect sizes by baseline type \(AI replacement, log10\-ratio encoding\)\.Baselinenperfn\_\{\\text\{perf\}\}MeanΔ¯perf\\bar\{\\Delta\}\_\{\\text\{perf\}\}ncostn\_\{\\text\{cost\}\}MeanΔ¯cost\\bar\{\\Delta\}\_\{\\text\{cost\}\}MedianΔ~cost\\tilde\{\\Delta\}\_\{\\text\{cost\}\}SC176\+0\.102\+0\.102\(\+26\.4%\+26\.4\\%\)137−2\.788\-2\.788−2\.518\-2\.518TS794\+0\.090\+0\.090\(\+23\.2%\+23\.2\\%\)298\+0\.996\+0\.996\+1\.000\+1\.000Against TS baselines, AI yields average performance improvements of roughly 23%, accompanied by a median computational cost increase of approximately one order of magnitude \(101\.000=1010^\{1\.000\}=10\)\. Against SC baselines, the cost pattern reverses\. The mean log10cost shift of−2\.788\-2\.788corresponds to an average acceleration factor of approximately102\.788≈61410^\{2\.788\}\\approx 614, and the median shift of−2\.518\-2\.518implies a reduction of roughly102\.518≈33010^\{2\.518\}\\approx 330\.

These values should be interpreted with appropriate caution\. Computational cost metrics differ substantially across studies and include heterogeneous quantities such as runtime, wall\-clock execution, simulation time, or hardware\-linked resource usage\. Accordingly, the numerical translations above are best understood as indicating large order\-of\-magnitude differences in the published literature rather than as universal benchmarks\. Even with that caveat, the quantitative effect sizes reinforce the main paper’s central asymmetry between AI substituting TS and AI substituting SC\.

### 7\.4Outcome Variation Across Scientific Domains

The main text also argues that AI’s position on the cost–performance frontier depends on domain context\. To assess whether outcome distributions vary systematically across fields, we cross\-tabulate domain cluster and outcome quadrant\.

Table S10:Contingency table for domain cluster by outcome quadrant in the AI replacement sample \(N=1,122N=1\{,\}122\)\.ClusterWWPPEPLLTotalEarth & Space Sci\.74604217193Economic Sciences3371129134Engineering92252359Environ\. & Agri\. Sci\.9211233Life Sciences201621059251Other12622783Physical Modeling Sci\.801149580369Total2375121562171,122The null hypothesis of independence is strongly rejected by a Pearson chi\-squared test \(χ2=244\.96\\chi^\{2\}=244\.96,d​f=18df=18,p=8\.63×10−42p=8\.63\\times 10^\{\-42\}\)\. The corresponding Cramér’sV=0\.270V=0\.270indicates a moderate association between domain cluster and outcome type\. This result supports the main paper’s claim that AI’s empirical role differs across scientific settings rather than following a single uniform replacement pattern\.

Substantively, the table suggests that simulation\-heavy areas such as Earth & Space Sciences and Physical Modeling Sciences contribute disproportionately to Win\-Win and efficiency prioritization outcomes, whereas Life Sciences and Economic Sciences are more heavily concentrated in the performance prioritization quadrant\. Engineering remains notably mixed, with both Performance\-Prioritization and Lose\-Lose outcomes substantial\. These domain\-level patterns are descriptive rather than causal, but they provide an important bridge between the aggregate evidence and the field\-specific discussion in the main text\.

## 8Sensitivity and Robustness Analyses

The main text describes systematic differences in the cost–performance trade\-offs associated with AI substitution across scientific settings\. Because the dataset aggregates comparisons from heterogeneous literature, it is important to verify that the reported patterns are not driven disproportionately by domain composition or by differences in reporting convention\. In this section we therefore present two general robustness checks on the AI replacement sample\.

### 8\.1Leave\-One\-Cluster\-Out Reliability

One possible concern is that a single domain cluster could dominate the aggregate results\. This is particularly relevant for simulation\-heavy areas such as Physical Modeling Sciences, which contain many comparisons involving computationally intensive mechanistic baselines\. To assess this possibility, we recompute the outcome distribution after sequentially removing each of the seven domain clusters from the AI replacement sample\.

Table S11:Leave\-one\-cluster\-out sensitivity for the AI replacement sample\. Each row removes the named cluster and reports the outcome distribution for the remaining comparisons, where WW = win\-win \(better performance, lower cost\), PP = performance\-prioritization \(better performance, higher cost\), EP = efficiency prioritization \(worse performance, lower cost\), and LL = lose\-lose \(worse performance, higher cost\)\. Stable proportions across rows indicate that no single cluster drives the aggregate pattern\.Cluster DroppedNNWW%PP%EP%LL%Earth & Space Sci\.92917\.548\.712\.321\.5Economic Sciences98820\.644\.615\.719\.0Engineering1,06321\.446\.114\.218\.3Environ\. & Agri\. Sci\.1,08920\.945\.114\.219\.7Life Sciences87124\.940\.216\.818\.1Other1,03921\.743\.314\.820\.2Physical Modeling Sci\.75320\.852\.98\.118\.2Full sample1,12221\.145\.613\.919\.3Across all exclusion scenarios the relative ordering of the outcome categories remains broadly similar\. Performance\-Prioritization continues to represent the largest share of outcomes, while Win\-Win and Lose\-Lose remain intermediate and Efficiency Prioritization remains the least common category\. This indicates that the aggregate patterns reported in the main text are not solely driven by any single domain cluster\. At the same time, the magnitude of the proportions varies across exclusion scenarios, which underscores that domain composition still influences the precise balance of outcomes\.

### 8\.2Quantitative\-Cost\-Only Check

A further concern is that some studies report computational cost only qualitatively\. Because such reports may be less comparable than explicit numerical measurements, we repeat the quadrant analysis using only the subset of comparisons for which numerical cost values are available\.

Table S12:Outcome distribution restricted to comparisons with quantitative cost measurements \(N=435N=435\)\.OutcomeFull SampleQuant\. Cost OnlyΔ\\DeltaWin\-Win21\.1%24\.4%\+3\.3\+3\.3ppPerf\.\-Prioritization45\.6%44\.1%−1\.5\-1\.5ppEfficiency Prioritization13\.9%20\.7%\+6\.8\+6\.8ppLose\-Lose19\.3%10\.8%−8\.5\-8\.5ppThe restricted sample produces outcome proportions that remain broadly similar to those of the full dataset, although efficiency prioritization becomes somewhat more common and Lose\-Lose outcomes become less frequent\. Because studies reporting quantitative cost metrics often provide more detailed benchmarking overall, this subset may represent a somewhat different population of experiments\. Even so, the central structure of the outcome distribution remains intact, suggesting that the main descriptive findings are not driven solely by qualitative cost reporting\.

Similar Articles

Evaluating AI’s ability to perform scientific research tasks

OpenAI Blog

OpenAI introduces FrontierScience, a new benchmark for measuring expert-level AI scientific capabilities across physics, chemistry, and biology, with GPT-5.2 achieving 77% on olympiad-style tasks and 25% on research-style tasks. The paper presents early evidence that GPT-5 meaningfully accelerates real scientific workflows, shortening work from weeks to hours while establishing metrics for tracking progress toward AI-accelerated science.

AI outperforms mathematicians

Reddit r/singularity

AI has progressed to the point of contributing to original mathematical research, outperforming human mathematicians and potentially reducing demand for the profession, though human-AI teams may ultimately excel.

AI Science & Economy: Systems Map

Reddit r/artificial

This article argues that while AI excels at pattern recognition and hypothesis generation, scientific and economic progress requires grounded interaction with reality and institutional execution, emphasizing the need for human-AI collaboration.