The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

arXiv cs.LG Papers

Summary

This perspective paper develops a conceptual and methodological framework for evaluating evidence-licensed claims in AI-assisted research, emphasizing calibration as a mechanism for managing scientific assertion rights and distinguishing between different AI research routes.

arXiv:2606.31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them. This Perspective-style paper develops a conceptual and methodological framework for evidence-licensed claims in AI-assisted research. Motivated by representative routes including specialized scientific foundation models, LLM research assistants, multi-agent co-scientists, AI Scientist pipelines, mathematical discovery agents, and self-driving laboratories, it represents AI-assisted research as five operators: hypothesis generation, model-mediated consequence derivation, external validation, belief update, and claim calibration. The central claim is that calibration is not merely cautious wording but a mechanism for managing scientific assertion rights: evidence licenses some forms of speech and withholds others. The paper distinguishes linguistic, consequence-based, interventional, and evidence-licensed semantics; defines the claim-evidence gap and epistemic debt; and treats minimal structural reconstruction across heterogeneous outputs as an upward form of claim calibration. AISim-Cal is included as an illustrative synthetic dynamics exercise, not as an empirical forecast or benchmark. The resulting principles are: no claim without license, validation does not determine claim level, and automation amplifies the need for calibration. Reliable AI-assisted research is therefore evaluated as a loop that generates hypotheses, derives testable consequences, accepts independent adjudication, updates beliefs, and outputs only evidence-licensed claims.
Original Article
View Cached Full Text

Cached at: 07/01/26, 05:35 AM

# The Calibration Turn in AI-Assisted Research A Conceptual and Methodological Framework for Evidence-Licensed Claims
Source: [https://arxiv.org/html/2606.31273](https://arxiv.org/html/2606.31273)
Hongmin Li1,2 1School of Life Science and Technology, Institute of Science Tokyo, 2\-12\-1 Ookayama, Meguro\-ku, Tokyo 152\-8550, Japan 2Department of Computational Biology and Medical Sciences, Graduate School of Frontier Sciences, The University of Tokyo, 5\-1\-5 Kashiwanoha, Kashiwa\-shi, Chiba 277\-8561, JapanAuthor note: Hongmin Li is a Researcher at the School of Life Science and Technology, Institute of Science Tokyo, and a Guest Researcher at the Department of Computational Biology and Medical Sciences, Graduate School of Frontier Sciences, The University of Tokyo\. ORCID:[0000\-0003\-0228\-0600](https://orcid.org/0000-0003-0228-0600)\. Correspondence:lihongmin@edu\.k\.u\-tokyo\.ac\.jp\.

\(June 16, 2026\)

###### Abstract

AI\-assisted research has entered a stage in which the central question is no longer only whether AI systems can generate hypotheses, run experiments, or produce manuscripts, but whether their final scientific claims are calibrated to the evidence that supports them\. This Perspective\-style paper develops a conceptual and methodological framework for evaluating evidence\-licensed claims in AI\-assisted research\. The framework is motivated by a methodological comparison of representative routes, including specialized scientific foundation models, human\-in\-the\-loop LLM research assistants, multi\-agent co\-scientist systems, end\-to\-end AI Scientist pipelines, algorithmic and mathematical discovery agents, and self\-driving laboratories\. It represents AI\-assisted research as a composition of five operators: hypothesis generation, model\-mediated consequence derivation, external validation, belief update, and claim calibration\. The central claim is that calibration is not merely cautious wording\. It is a mechanism for managing scientific assertion rights: evidence licenses some forms of speech and withholds others\. The paper distinguishes linguistic, consequence\-based, interventional, and evidence\-licensed semantics; defines the claim\-evidence gap and epistemic debt; and treats minimal structural reconstruction across heterogeneous outputs as a special, upward form of claim calibration\. It argues that AI science routes differ not only in consequence\-derivation capacity or automation, but also in whether their evaluators are independent, reliable, world\-grounded, and translated into bounded scientific claims\. AISim\-Cal is included only as an illustrative synthetic dynamics exercise, not as an empirical forecast or benchmark\. The resulting calibration principles are: no claim without license, validation does not determine claim level, and automation amplifies the need for calibration\. This framework therefore evaluates reliable AI\-assisted research as an AI\-augmented loop that generates hypotheses, derives testable consequences, accepts independent adjudication, updates beliefs, and outputs only evidence\-licensed claims\.

Keywords:AI for science; scientific claims; evidential calibration semantics; epistemology; multi\-agent systems; self\-driving laboratories; external validation

## 1Introduction: From Scientific Automation to Claim Licensing

Discussions of AI scientific exploration often focus on whether a system has produced a sufficiently substantive scientific contribution or a genuine discovery\. This issue is broader than model capability or manuscript generation, but it remains incomplete unless contribution and discovery are constrained by evidence\. Scientific knowledge is not primarily a matter of novelty labels or textual productivity; it is a matter of testable constraint on the world\. A chapter by King and Zenil in the OECD report on artificial intelligence in science frames the goal of science as the construction of models that predict what happens in the real world, especially in experiments, and treats predictive performance on experiments as a natural objective for AI systems in science\(OECD,,[2023](https://arxiv.org/html/2606.31273#bib.bib21); King and Zenil,,[2023](https://arxiv.org/html/2606.31273#bib.bib14)\)\. This paper adopts that predictive framing where it applies, but treats prediction as an important special case of a broader consequence relation\. Some sciences derive future\-oriented predictions; others derive retrodictive traces, diagnostic features, proof obligations, classification constraints, or measurement signatures\. This framing is also continuous with older ideas about pragmatic meaning and inquiry, falsifiability, research programmes, intervention, severe testing, and the relation between claims, data, and warrants\(Peirce,,[1878](https://arxiv.org/html/2606.31273#bib.bib24); Popper,,[1959](https://arxiv.org/html/2606.31273#bib.bib25); Lakatos,,[1970](https://arxiv.org/html/2606.31273#bib.bib16); Hacking,,[1983](https://arxiv.org/html/2606.31273#bib.bib12); Mayo,,[1996](https://arxiv.org/html/2606.31273#bib.bib18); Toulmin,,[1958](https://arxiv.org/html/2606.31273#bib.bib31)\)\. This provides the starting point for the present paper: an AI scientific system is reliable only insofar as it can turn hypotheses into domain\-appropriate testable consequences, update those hypotheses through independent adjudication, and calibrate its final claims to the evidence produced by that adjudication\.

As of June 2026, this Perspective uses six non\-exhaustive route families as epistemological comparators\. AlphaFold 3\(Abramson et al\.,,[2024](https://arxiv.org/html/2606.31273#bib.bib1)\)and GNoME\(Merchant et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib19)\)represent specialized scientific foundation models\. AlphaFold 3 uses a substantially updated diffusion\-based architecture to predict the joint structure of biomolecular complexes that may include proteins, nucleic acids, small molecules, ions, and modified residues\(Abramson et al\.,,[2024](https://arxiv.org/html/2606.31273#bib.bib1)\)\. GNoME uses deep learning to expand the search space for materials and generate large\-scale computational candidates for stable crystal structures\(Merchant et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib19)\)\. GPT\-5 and related frontier models are used by experts for literature synthesis, proof sketches, mechanistic hypotheses, and experimental suggestions, but the published case reports explicitly describe curated examples rather than systematic samples and emphasize the need for expert validation\(Bubeck et al\.,,[2025](https://arxiv.org/html/2606.31273#bib.bib4); OpenAI,,[2025](https://arxiv.org/html/2606.31273#bib.bib22)\)\. Google Co\-Scientist and FutureHouse Robin organize hypothesis generation, critique, ranking, data analysis, and experimental planning into multi\-agent systems\(Gottweis et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib9); Ghareeb et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib7)\)\. The Sakana AI Scientist line attempts to automate idea generation, code execution, figures, manuscripts, and review in machine learning research\(Lu et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib17)\)\. AlphaEvolve, AlphaProof, and AlphaGeometry 2 show that when candidate objects can be evaluated by program tests, benchmarks, or formal proof checkers, AI systems can obtain much stronger epistemic warrants\(Novikov et al\.,,[2025](https://arxiv.org/html/2606.31273#bib.bib20); Hubert et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib13)\)\. Boiko et al\.’s Coscientist chemistry system and A\-Lab connect AI systems to robotic experimentation and real data feedback, while also exposing the high cost of experimental engineering, interpretation, and reproducibility review\(Boiko et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib3); Szymanski et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib27); Chawla,,[2026](https://arxiv.org/html/2606.31273#bib.bib5)\)\. The evidential status of these sources differs: peer\-reviewed articles, arXiv or white\-paper claims, official reports, and journalism\-supported corrections are treated here as epistemological comparators, not as independently verified field rankings\.

The central argument of this paper is that these routes differ less in their model families than in their implicit theory of scientific meaning\. Some systems locate meaning in prediction distributions, some in textual hypotheses, some in research artifacts, some in programs or proofs, and some in experimental responses\. Across routes, the shared failure mode is claim drift: the migration from a licensed weaker claim to an unlicensed stronger one\. The key issue is therefore not only the identity and independence of the adjudicator, but also the strength of scientific assertion licensed after adjudication\.

The paper addresses three questions\. First, how does an AI\-generated hypothesis become a scientific claim? Second, what level of claim is licensed by a given external validation result? Third, when evidence is insufficient, how should the system cancel the resulting epistemic debt by collecting more evidence, weakening the claim, narrowing its boundary, or changing its modal status?

The contributions are fourfold\. First, the paper introduces an evidential calibration vocabulary for AI scientific exploration, moving the analysis from hypothesis generation and external validation to evidence\-licensed assertion\. Second, it formalizes scientific assertion rights at a schematic level through a domain\-sensitive license relationE⊩𝖣,VCE\\Vdash\_\{\\mathsf\{D\},V\}C, together with claim orders and claim\-relative evidence preorders\. Third, it defines the claim\-evidence gap and epistemic debt as diagnostic measures of mismatch between the evidential strength required by a claim and the support actually provided by the evidence\. Fourth, it uses representative AI science routes as methodological comparators to argue that the central risk is not simply lack of validation, but the overinterpretation of experimental actionability, artifact completeness, or paper form as stronger scientific discovery than the evidence permits\.

This paper develops the methodological and epistemological framework\. It does not introduce a new AI\-assisted science system or report domain\-specific scientific results\.

This is a Perspective\-style theoretical synthesis and methodological framework rather than a systematic literature review\. The routes discussed below are selected as prominent epistemological comparators for AI\-assisted research, not as an exhaustive census of all systems\.

In this paper, “AI\-assisted research” is the umbrella term, while “AI scientific exploration” refers to the subset of AI\-assisted research systems that generate, evaluate, or communicate scientific hypotheses and claims\.

Article type and scope\.The primary contribution of this paper is a conceptual and methodological framework for calibrating scientific claims in AI\-assisted research\. The route comparison motivates and stress\-tests the framework across representative paradigms; it is not a systematic field map\. AISim\-Cal is an illustrative synthetic dynamics exercise; it is not an empirical forecast, a benchmark, or evidence that any route is superior\. The phrase “calibration turn” names the evaluative shift proposed here, not a claim that the field has already converged on a settled paradigm\.

Every calibrated scientific output should therefore specify both the claim licensed by the evidence and the stronger paired non\-claim that remains unlicensed\.

Core thesis\.AI scientific exploration should be evaluated not only by whether a system generates hypotheses or obtains validation, but by whether it manages scientific assertion rights\. We formalize this with the license relationE⊩𝖣,VCE\\Vdash\_\{\\mathsf\{D\},V\}C, the claim\-evidence gapΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\), the epistemic debtDebt𝖣,Vℓ⁡\(C,E\)\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C,E\), and the calibration operatorCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}, which maps a raw claim, evidence, evaluator, and domain context to a maximal evidence\-licensed claim frontier\. A selected calibrated claim is thenC∗∈Cal𝖣⁡\(C,E,V\)C^\{\\ast\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\. Experimental actionability, artifact completeness, or benchmark success becomes scientifically reliable only when it is converted into an evidence\-licensed claim\.

The framework is designed to connect philosophical semantics with AI\-system analysis\. Definitions are repeated when they become operationally relevant, and key symbols are explained both in the notation table and near the formulas in which they are used\. This redundancy supports comparisons across LLM\-based scientific agents, AI Scientist pipelines, autonomous laboratories, formal proof agents, and specialized scientific models\. In each case, the same decomposition is required: which subsystem generates hypotheses, which subsystem derives consequences, which mechanism functions as an evaluator, and where the system’s language must be downgraded when evidence is insufficient\.

The same question governs all routes considered below: what claim does the evidence license? This question remains invariant across language models, formal proof systems, robotic laboratories, benchmark\-driven machine learning, and domain\-specific foundation models, even though their artifacts, evaluators, costs, and validation schedules differ substantially\.

Knowledge, data, and research question\(𝒦,Dt,q\)\(\\mathcal\{K\},D\_\{t\},q\)GG: hypothesis generationcandidate hypothesishhM𝖣M\_\{\\mathsf\{D\}\}: consequence derivationχh​\(a,x\)∈𝒬𝖣\\chi\_\{h\}\(a,x\)\\in\\mathcal\{Q\}\_\{\\mathsf\{D\}\}VV: external adjudicationexperiment, simulation, proof, benchmark, or statisticEvidence bodyEEand belief updateUUupdated state of knowledgeCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}: assertion\-right calibrationC∗∈Cal𝖣⁡\(C,E,V\)⊆ℒ𝖣,V​\(E\)C^\{\\ast\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\\subseteq\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)Raw claimCCepistemic debtDebt𝖣,Vℓ⁡\(C,E\)\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C,E\)Figure 1:The calibrated AI science loop\. Hypothesis generationGGproposes candidates from knowledge, accumulated dataDtD\_\{t\}, and a research question\. ModelingM𝖣M\_\{\\mathsf\{D\}\}turns a hypothesis into domain\-appropriate testable consequences\. Interventional predictionPh​\(Y∣do⁡\(a\),x\)P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),x\)is one important special case; other domains may derive retrodictive traces, diagnostic features, proof obligations, classification constraints, or measurement signatures\. ValidationVVadjudicates the consequence through experiment, simulation, proof checking, benchmarking, measurement, or statistical testing\. Belief updateUUincorporates the result into the current state of knowledge\. Claim calibrationCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}maps the raw claimCC, evidenceEE, evaluatorVV, and domain context𝖣\\mathsf\{D\}to a maximal licensed frontier, from which a selected calibrated claimC∗C^\{\\ast\}must belong toℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\. WithoutCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}, an AI science loop can obtain validation while still producing overstrong claims and epistemic debt\.Table 1:Licensed claims and paired non\-claims across representative AI science routes
## 2Notation and Core Symbols

Because the paper compares several AI science routes, the same mathematical section contains variables for worlds, hypotheses, claims, evaluators, and workflow artifacts\. TableLABEL:tab:notationgives a consolidated glossary: the left column gives the symbol, the middle column gives its formal role, and the right column gives the plain interpretation used in the argument\.

Three conventions reduce symbol overload\. First, uppercaseCCis reserved for scientific claims in expressions such asClaim⁡\(h,s,b,m\)\\operatorname\{Claim\}\(h,s,b,m\),Δ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\), orCal𝖣⁡\(C,E,V\)\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\. Costs are written asCost⁡\(⋅\)\\operatorname\{Cost\}\(\\cdot\), notC​\(⋅\)C\(\\cdot\)\. Second, accumulated data are written asDtD\_\{t\}, whereas the domain context is written as𝖣\\mathsf\{D\}when it indexes a calibration or licensing relation\. Third, claim requirements and evidential support are written asReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)andSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\), not as genericRRandSS\. This separation is important becauseCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}is the paper’s central operator\.

The most important symbols for the whole argument arehh,CC,EE,VV, andCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\. The hypothesishhis what an AI system proposes or models\. The claimCCis what the system, paper, or user\-facing report says about that hypothesis\. The evidenceEEis what has actually been observed, tested, proven, simulated, or replicated\. The evaluatorVVis the mechanism that judges whether the evidence supports the candidate\. The calibration operatorCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}converts the raw claim into a bounded claimC∗C^\{\\ast\}that does not exceed the evidence\. The core dependency is:

hypothesish⟶claimC⟶evidenceE⟶calibrated claimC∗\.\\boxed\{\\text\{hypothesis \}h\\longrightarrow\\text\{claim \}C\\longrightarrow\\text\{evidence \}E\\longrightarrow\\text\{calibrated claim \}C^\{\\ast\}\.\}
Table 2:Main notation used in the manuscriptSymbolFormal rolePlain interpretation in this paperΩ\\OmegaSpace of possible worlds or mechanismsThe set of causal, physical, biological, material, or mathematical structures that could explain the data\.ω∗\\omega^\{\\ast\}True but unknown element ofΩ\\OmegaThe real mechanism that science tries to approximate but cannot observe directly\.DtD\_\{t\}Data accumulated up to timettEverything observed so far: literature, experiments, simulations, benchmark results, proof attempts, and database records\.𝖣\\mathsf\{D\}Domain contextThe field\-specific context that determines admissible evidence, claim strength, and calibration norms\.\(ai,yi,ci\)\(a\_\{i\},y\_\{i\},c\_\{i\}\)Action–outcome–context tripleTheiith test or query: what was done, what happened, and under which conditions\.a,ata,a\_\{t\}Scientific action or testAn experiment, intervention, simulation, query, proof attempt, benchmark run, or robotic laboratory action\.𝒜\\mathcal\{A\}Action spaceThe set of interventions, experiments, simulations, proof attempts, or queries available to the system\.𝒳\\mathcal\{X\}Context spaceThe set of biological, chemical, computational, or experimental contexts; this avoids usingCCfor contexts\.y,Yy,YObserved or random outcomeThe measured response produced by an action;YYis the random variable andyyis a realized observation\.ccContext conditionThe biological, chemical, computational, or experimental setting in which an action is evaluated\.𝒦\\mathcal\{K\}Prior knowledge baseLiterature, databases, theory, expert background, and previously accepted constraints\.qqResearch questionThe problem statement or scientific task that guides hypothesis generation\.hhScientific hypothesisA proposed mechanism, relation, conjecture, explanation, or candidate scientific statement before claim calibration\.HtH\_\{t\}Hypothesis set at timettThe current population of hypotheses available to the AI or scientist\.MtM\_\{t\}Current modelThe predictive, mechanistic, statistical, or simulation model currently used to derive consequences\.Bt​\(ω\)B\_\{t\}\(\\omega\)Belief distribution overΩ\\OmegaThe current degree of belief that mechanismω\\omegais the correct one after seeingDtD\_\{t\}\.bt​\(ω\)b\_\{t\}\(\\omega\)Belief state in the SDL sectionSame idea asBt​\(ω\)B\_\{t\}\(\\omega\), used in the self\-driving\-laboratory active\-learning notation\.StS\_\{t\}Inquiry stateThe full state\(Dt,𝒦,Ht,Mt,Bt\)\(D\_\{t\},\\mathcal\{K\},H\_\{t\},M\_\{t\},B\_\{t\}\)used by a scientific policy\.π\\piScientific policyA rule that chooses the next actionat\+1a\_\{t\+1\}from the current stateStS\_\{t\}\.Φ​\(at,St\)\\Phi\(a\_\{t\},S\_\{t\}\)Local scientific utilityThe immediate value of actionata\_\{t\}, combining prediction, causality, explanation, novelty, cost, and risk\.Δ​Pred\\Delta\\mathrm\{Pred\}Predictive improvementHow much better the system predicts observations after a proposed action or update\.Δ​Causal\\Delta\\mathrm\{Causal\}Causal\-identification improvementHow much the action clarifies intervention\-relevant causal structure\.Δ​Explanation\\Delta\\mathrm\{Explanation\}Explanatory improvementHow much the action improves mechanism\-level interpretation\.Δ​Novelty\\Delta\\mathrm\{Novelty\}Novelty improvementHow much the action expands beyond known hypotheses or artifacts\.Cost⁡\(at\)\\operatorname\{Cost\}\(a\_\{t\}\)Cost of actionata\_\{t\}A local cost term, kept distinct from the scientific claimCC\.Risk⁡\(at\)\\operatorname\{Risk\}\(a\_\{t\}\)Risk of actionata\_\{t\}A local safety or feasibility penalty, kept distinct from the requirement profileReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\.\[\[h\]\]\\left\[\\\!\\left\[h\\right\]\\\!\\right\]Model\-theoretic denotationThe set of possible worlds in which hypothesishhis true\.Sem⁡\(h\)\\operatorname\{Sem\}\(h\)Scientific semantics ofhhThe domain\-appropriate consequences that make a hypothesis scientifically testable\.Conseq𝖣⁡\(h\)\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\)Consequence profile underhhThe set of testable consequences that hypothesishhimplies in domain𝖣\\mathsf\{D\}\.χh​\(a,x\)\\chi\_\{h\}\(a,x\)Local consequence underhhA predicted outcome, retrodictive trace, proof obligation, diagnostic feature, measurement signature, or constraint induced byhh\.𝒬𝖣\\mathcal\{Q\}\_\{\\mathsf\{D\}\}Domain consequence spaceThe space containing the kinds of consequences admissible in domain𝖣\\mathsf\{D\}\.Ph​\(Y∣do⁡\(a\),x\)P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),x\)Interventional prediction underhhThe special case in which the consequence is a probability distribution over outcomes after actionaain contextxx\.do⁡\(a\)\\operatorname\{do\}\(a\)Intervention operatorThe act of doing or enforcing actionaa, rather than merely observing correlation\.𝒯​\(h\)\\mathcal\{T\}\(h\)Test set forhhThe experiments, simulations, benchmarks, or proof obligations that can evaluatehh\.𝒱​\(h\)\\mathcal\{V\}\(h\)Verification or refutation ruleThe rule that says what would count as support, failure, or refutation for hypothesishh\.ℬ​\(h\)\\mathcal\{B\}\(h\)Claim boundary forhhThe range of assertions abouthhthat current evidence is allowed to support\.CCScientific claimA sentence\-level assertion made in a paper, report, or system output\.𝒞𝖣\\mathcal\{C\}\_\{\\mathsf\{D\}\}Domain\-specific claim spaceThe set of claims that can be made in domain context𝖣\\mathsf\{D\}\.⪯𝖣\\preceq\_\{\\mathsf\{D\}\}Claim\-strength orderC1⪯𝖣C2C\_\{1\}\\preceq\_\{\\mathsf\{D\}\}C\_\{2\}means thatC1C\_\{1\}is no stronger thanC2C\_\{2\}in domain𝖣\\mathsf\{D\}\.Claim⁡\(h,s,b,m\)\\operatorname\{Claim\}\(h,s,b,m\)Structured claim representationA claim built from a hypothesishh, strengthss, boundarybb, and modal statusmm\.ssClaim strengthHow strong the assertion is: possible, supported, demonstrated, general, translational, and so on\.bbClaim boundaryThe domain, population, assay, benchmark, material class, or context where the claim is asserted to hold\.mmModal statusThe claim’s epistemic mode: possible, supported, demonstrated, explained, discovered, or translated\.EEEvidence bodyThe experiments, simulations, proofs, benchmarks, databases, or replications available for a claim\.ℰ𝖣\\mathcal\{E\}\_\{\\mathsf\{D\}\}Domain\-specific evidence spaceThe set of evidence bodies admissible in domain𝖣\\mathsf\{D\}\.𝒲𝖣\\mathcal\{W\}\_\{\\mathsf\{D\}\}Warrant\-profile spaceThe ordered space in which evidential requirements and support profiles are compared\.E1⊑𝖣,V,𝒞0E2E\_\{1\}\\sqsubseteq\_\{\\mathsf\{D\},V,\\mathcal\{C\}\_\{0\}\}E\_\{2\}Claim\-relative evidence orderEvidenceE2E\_\{2\}is no weaker thanE1E\_\{1\}for every claim in the comparison family𝒞0\\mathcal\{C\}\_\{0\}under evaluatorVV\.E⊩𝖣,VCE\\Vdash\_\{\\mathsf\{D\},V\}CLicense relationUnder domain𝖣\\mathsf\{D\}and evaluatorVV, evidenceEElicenses scientific assertionCC; this is warrant, not truth\.Req𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)Requirement profile for claimCCThe evidential warrant needed before claimCCis legitimate in domain context𝖣\\mathsf\{D\}\.Sup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)Support profile provided byEEforCCThe evidential support that the available evidence and evaluator give to the claim\.Sup𝖣,V⪰𝒲,𝖣Req𝖣\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Req\}\_\{\\mathsf\{D\}\}Evidence\-licensing relationThe support profileSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)covers the requirement profileReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\.ℱ\\mathcal\{F\}Family of heterogeneous claimsA comparison family\{C1,…,Ck\}\\\{C\_\{1\},\\ldots,C\_\{k\}\\\}whose members may share a higher\-order structure\.CstrC\_\{\\mathrm\{str\}\}Candidate structural claimA proposed family\-level minimal reconstruction that captures a candidate common structure acrossℱ\\mathcal\{F\}; it is still a raw claim until calibrated\.Eℱ,VℱE\_\{\\mathcal\{F\}\},V\_\{\\mathcal\{F\}\}Family\-level evidence and evaluator recordThe evidence and evaluator provenance used to decide whether a structural reconstruction is licensed\.ℓ𝖣\\ell\_\{\\mathsf\{D\}\}Scalarization mapAn order\-preserving map from warrant profiles into the scalar diagnosticΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)\.Δ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)Domain\-indexed claim\-evidence gapA diagnostic projection ofReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)minusSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\); positive values mean the claim is stronger than the evidence under domain context𝖣\\mathsf\{D\}\.Debt𝖣,Vℓ⁡\(C,E\)\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C,E\)Scalar epistemic debtThe positive part of the scalarized claim\-evidence gap, or the unmet warrant owed by an overstrong claim under scalarizationℓ𝖣\\ell\_\{\\mathsf\{D\}\}\.ℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)Licensed claim setAll claims licensed by evidenceEEunder domain context𝖣\\mathsf\{D\}and evaluatorVV; a lower set in the claim poset\.⊥𝖣\\bot\_\{\\mathsf\{D\}\}Bottom claimA fallback claim stating that no scientific assertion beyond proposal is licensed\.↓C\\downarrow CDown\-set of a claimAll claims no stronger thanCC:\{C′:C′⪯𝖣C\}\\\{C^\{\\prime\}:C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C\\\}\.C′⪯𝖣CC^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}CClaim weakening relationClaimC′C^\{\\prime\}is no stronger than the raw or intended claimCCin domain context𝖣\\mathsf\{D\}\.Strength⁡\(C′\)\\operatorname\{Strength\}\(C^\{\\prime\}\)Strength score of a claimA ranking of how strong or informative a licensed claim is\.C∗C^\{\\ast\}Selected calibrated claimA reportable claim selected from the maximal evidence\-licensed frontierCal𝖣⁡\(C,E,V\)\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\.Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}Claim\-calibration operatorThe assertion\-right operator: it maps a raw claim, evidence, evaluator, and domain context to a maximal licensed claim frontier; a domain\-specific selection rule may chooseC∗C^\{\\ast\}from this frontier\.CliteralC\_\{\\mathrm\{literal\}\}Literal claimThe exact sentence or formal claim written in a paper, report, title, or system output\.CimpliedC\_\{\\mathrm\{implied\}\}Implied claimThe stronger or broader claim induced by framing, title, abstract, figure, press release, or downstream communication\.Shearr⁡\(t\)\\operatorname\{Shear\}\_\{r\}\(t\)Epistemic shearA route\-level imbalance between generation/consequence\-derivation capacity and adjudication/calibration capacity\.Qr,𝖣cal​\(t\)Q^\{\\mathrm\{cal\}\}\_\{r,\\mathsf\{D\}\}\(t\)Route\-level calibration qualityA scalar AISim\-Cal simulation parameter for how well routerrpulls claims back toward the licensed frontier in domain𝖣\\mathsf\{D\}; kept distinct from the operatorCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\.Γ𝖣\\Gamma\_\{\\mathsf\{D\}\}Goal clarity in domain𝖣\\mathsf\{D\}How well a domain or task has specified its objective, scoring rule, and claim\-value target\.τΓ\\tau\_\{\\Gamma\}Goal\-clarity thresholdThe minimum clarity level used in AISim\-Cal before a route ordering is reported\.Ar,𝖣A\_\{r,\\mathsf\{D\}\}Route\-domain applicabilityHow compatible routerris with domain𝖣\\mathsf\{D\}; it prevents strong evaluators from implying universal route superiority\.d,i,r,e,cid,ξ,nd,i,r,e,c\_\{\\mathrm\{id\}\},\\xi,nSupport dimensionsDirectness, independence, reproducibility, effect size, causal identification, external validity, and negative evidence\.F​\(⋅\)F\(\\cdot\)Support aggregation or profiling functionThe domain\-specific rule that combines evidential dimensions intoSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\.L​0L0–L​6L6Claim ladder levelsA practical scale from hypothesis generation to translational or application\-ready assertion\.GGHypothesis\-generation operatorMaps knowledge, data, and question to a hypothesis set\.MMModel\-mediated consequence\-derivation operatorMaps hypothesis, action, and context to a domain\-appropriate consequence profile\.VVValidation or evaluator operatorJudges whether a hypothesis, consequence, program, proof, or experiment survived a test\.UUBelief\-update operatorUpdates the belief state after new evidence is observed\.W​\(h,Ch,Eh\)W\(h,C\_\{h\},E\_\{h\}\)Discovery objectiveA score for hypotheses and intended claims that rewards evidence and penalizes overclaiming\.WtextW\_\{\\mathrm\{text\}\}Textual warrantTextual plausibility or continuation\-based support\.WconseqW\_\{\\mathrm\{conseq\}\}Consequence warrantEvidence that a hypothesis yields domain\-appropriate testable consequences\.WadjudicatedW\_\{\\mathrm\{adjudicated\}\}Adjudicated warrantEvidence that derived consequences have survived an external evaluator, such as a test, proof checker, benchmark, or assay\.WlicenseW\_\{\\mathrm\{license\}\}Evidence\-licensed assertionThe strongest level here: a claim is insideℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\.V​\(p\)V\(p\)Program or algorithm evaluatorRoute\-specific evaluator for algorithmic discovery; higher values mean better candidate programs\.Check⁡\(π\)\\operatorname\{Check\}\(\\pi\)Proof checkerA strict verifier that accepts or rejects a proof certificateπ\\pi\.z=\(h,c,e,r,p\)z=\(h,c,e,r,p\)AI Scientist artifact tupleIdea, code, experimental output, analysis, and manuscript in an end\-to\-end pipeline\.RCSC​\(A′\)R\_\{\\mathrm\{CSC\}\}\(A^\{\\prime\}\)Typed\-artifact residualA complementary measure of whether a new artifact exceeds what is transported from a prior schema\.A′A^\{\\prime\},It\+1′I^\{\\prime\}\_\{t\+1\},ρA′\\rho\_\{A^\{\\prime\}\}Typed\-artifact termsNew artifact, new interpretation, and schema\-transport map in the CategoryScienceClaw comparison\.α,β,γ,δ,η\\alpha,\\beta,\\gamma,\\delta,\\etaPositive\-value weightsCoefficients for prediction, causality, explanation, novelty, information gain, or related benefits\.ρ,λ,κ,μ,ν\\rho,\\lambda,\\kappa,\\mu,\\nuPenalty or auxiliary weightsCoefficients for novelty, cost, risk, complexity, and the overclaim penalty in the objective functions\.
## 3A Unified Abstraction: Scientific Exploration as Policy Optimization

Let the real world contain an unknown mechanism

ω∗∈Ω,\\omega^\{\\ast\}\\in\\Omega,\(1\)whereΩ\\Omegadenotes the space of possible worlds, mechanisms, causal structures, physical laws, biological processes, or material stability rules\. Scientists cannot directly observeω∗\\omega^\{\\ast\}\. They obtain data through literature, observation, experiment, simulation, theorem proving, and databases:

Dt=\{\(ai,yi,ci\)\}i=1t\.D\_\{t\}=\\\{\(a\_\{i\},y\_\{i\},c\_\{i\}\)\\\}\_\{i=1\}^\{t\}\.\(2\)Hereaia\_\{i\}denotes an experimental action, computational action, query, proof attempt, simulation condition, or intervention;yiy\_\{i\}denotes the observed outcome; andcic\_\{i\}denotes contextual conditions such as cell type, temperature, solvent, genetic background, material composition, or benchmark setting\.

Thus, throughout the paper,DtD\_\{t\}denotes the system’s accumulated evidential state rather than a single dataset\. The triple\(ai,yi,ci\)\(a\_\{i\},y\_\{i\},c\_\{i\}\)keeps three things separate: what the system did, what it observed, and where that observation is valid\.

Scientific inquiry can be written as a policy:

π:St→at\+1,\\pi:S\_\{t\}\\rightarrow a\_\{t\+1\},\(3\)with state

St=\(Dt,𝒦,Ht,Mt,Bt\)\.S\_\{t\}=\(D\_\{t\},\\mathcal\{K\},H\_\{t\},M\_\{t\},B\_\{t\}\)\.\(4\)𝒦\\mathcal\{K\}is the prior knowledge base,HtH\_\{t\}is the set of hypotheses,MtM\_\{t\}is the current model, andBt​\(ω\)=P​\(ω∣Dt,𝒦\)B\_\{t\}\(\\omega\)=P\(\\omega\\mid D\_\{t\},\\mathcal\{K\}\)is the current belief distribution over possible mechanisms\. An idealized AI science system may be written as:

π∗\\displaystyle\\pi^\{\\ast\}=arg​maxπ⁡𝔼​\[∑tγt​Φ​\(at,St\)\],\\displaystyle=\\operatorname\*\{arg\\,max\}\_\{\\pi\}\\mathbb\{E\}\\left\[\\sum\_\{t\}\\gamma^\{t\}\\Phi\(a\_\{t\},S\_\{t\}\)\\right\],\(5\)Φ​\(at,St\)\\displaystyle\\Phi\(a\_\{t\},S\_\{t\}\)=α​Δ​Pred\+β​Δ​Causal\+η​Δ​Explanation\+ρ​Δ​Novelty\\displaystyle=\\alpha\\Delta\\mathrm\{Pred\}\+\\beta\\Delta\\mathrm\{Causal\}\+\\eta\\Delta\\mathrm\{Explanation\}\+\\rho\\Delta\\mathrm\{Novelty\}\+ζ​Δ​Info\\displaystyle\\quad\+\\zeta\\Delta\\mathrm\{Info\}−λ​Cost⁡\(at\)−κ​Risk⁡\(at\)\.\\displaystyle\\quad\-\\lambda\\operatorname\{Cost\}\(a\_\{t\}\)\-\\kappa\\operatorname\{Risk\}\(a\_\{t\}\)\.The system should maximize predictive improvement, causal identification, explanatory power, novelty, and information gain under constraints of cost, safety, and feasibility\.

In this objective,π\\piis not a language model prompt; it is the whole inquiry policy that chooses the next test\. The functionΦ\\Phiis the local scientific value of a test, and the coefficientsα,β,η,ρ,ζ,λ,κ\\alpha,\\beta,\\eta,\\rho,\\zeta,\\lambda,\\kappaindicate how much the system values prediction, causality, explanation, novelty, information, cost, and risk\. The explicit functionsCost\\operatorname\{Cost\}andRisk\\operatorname\{Risk\}avoid overloadingCC, which is reserved for scientific claims\.

This expression is a schematic policy objective rather than a fully specified MDP or POMDP\. A concrete transition kernelP​\(St\+1∣St,at\)P\(S\_\{t\+1\}\\mid S\_\{t\},a\_\{t\}\), observation model, and belief\-update rule must be supplied by the domain\. The schematic objective does not imply that every AI system must become a complete scientist\. It allows us to identify which part of the scientific objective each route optimizes\. Specialized models primarily optimize predictive improvement\. LLM assistants expand the hypothesis space\. Multi\-agent systems optimize generation and critique\. AI Scientist systems optimize research\-artifact production\. Formal systems optimize verifiable objects\. Self\-driving laboratories optimize information gain through real\-world interventions\.

## 4Scientific Semantics: From Text Continuation to World Constraint

A scientific hypothesis cannot be reduced to a natural\-language sentence\. Consider:

h=“gene​g​activates pathway​p​\.”h=\\text\{\`\`gene \}g\\text\{ activates pathway \}p\\text\{\.''\}At the linguistic level, this is just a sentence\. At the scientific level, it must constrain which worlds, interventions, and observations are possible\. Its model\-theoretic meaning can be written as

\[\[h\]\]=\{ω∈Ω:ω⊧h\}\.\\left\[\\\!\\left\[h\\right\]\\\!\\right\]=\\\{\\omega\\in\\Omega:\\omega\\models h\\\}\.\(6\)Scientific meaning, however, also requires testable consequences:

Sem𝖣⁡\(h\)=Conseq𝖣⁡\(h\)=\{χh​\(a,x\):a∈𝒜𝖣,x∈𝒳𝖣\},χh​\(a,x\)∈𝒬𝖣\.\\operatorname\{Sem\}\_\{\\mathsf\{D\}\}\(h\)=\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\)=\\\{\\chi\_\{h\}\(a,x\):a\\in\\mathcal\{A\}\_\{\\mathsf\{D\}\},x\\in\\mathcal\{X\}\_\{\\mathsf\{D\}\}\\\},\\qquad\\chi\_\{h\}\(a,x\)\\in\\mathcal\{Q\}\_\{\\mathsf\{D\}\}\.\(7\)In words,Conseq𝖣⁡\(h\)\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\)is the profile of domain\-appropriate consequences that hypothesishhimplies under possible actions, queries, observations, or proof attempts\. The space𝒬𝖣\\mathcal\{Q\}\_\{\\mathsf\{D\}\}may contain probability distributions, deterministic functions, constraint sets, proof obligations, counterexample conditions, diagnostic features, measurement signatures, classification rules, or retrodictive traces\. In interventional experimental science, an important special case is:

χh​\(a,x\)=Ph​\(Y∣do⁡\(a\),x\)\.\\chi\_\{h\}\(a,x\)=P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),x\)\.\(8\)The notationPh​\(Y∣do⁡\(a\),x\)P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),x\)follows the interventionist convention that the consequence is derived under the assumption thathhis true and that the actionaais performed rather than merely observed\(Pearl,,[2009](https://arxiv.org/html/2606.31273#bib.bib23)\)\. The context variable is written asx∈𝒳𝖣x\\in\\mathcal\{X\}\_\{\\mathsf\{D\}\}here so that the symbolCCremains reserved for scientific claims\. For example, the hypothesish:g→ph:g\\rightarrow pshould be expandable into:

𝔼\[Yp∣do\(g↑\),c\]−𝔼\[Yp∣do\(g↓\),c\]\>δ\.\\mathbb\{E\}\[Y\_\{p\}\\mid\\operatorname\{do\}\(g\\uparrow\),c\]\-\\mathbb\{E\}\[Y\_\{p\}\\mid\\operatorname\{do\}\(g\\downarrow\),c\]\>\\delta\.\(9\)The meaning of a scientific hypothesis is therefore not whether a language model can explain it fluently, but whether the hypothesis can be translated into a structure that constrains observable, formal, diagnostic, measurement, or retrodictive consequences\.

This paper defines the meaning of a scientific hypothesis as:

ScientificMeaning𝖣​\(h\)=\(\[\[h\]\],Conseq𝖣⁡\(h\),𝒯𝖣​\(h\),𝒱𝖣​\(h\),ℬ𝖣​\(h\)\)\\boxed\{\\mathrm\{ScientificMeaning\}\_\{\\mathsf\{D\}\}\(h\)=\\left\(\\left\[\\\!\\left\[h\\right\]\\\!\\right\],\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{T\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{V\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{B\}\_\{\\mathsf\{D\}\}\(h\)\\right\)\}\(10\)where𝒯𝖣​\(h\)\\mathcal\{T\}\_\{\\mathsf\{D\}\}\(h\)is the set of tests, observations, simulations, benchmarks, measurements, or proof obligations;𝒱𝖣​\(h\)\\mathcal\{V\}\_\{\\mathsf\{D\}\}\(h\)is the rule for verification, support, failure, or refutation; andℬ𝖣​\(h\)\\mathcal\{B\}\_\{\\mathsf\{D\}\}\(h\)is the claim boundary: the range of assertions abouthhthat the current evidence permits\. Withoutℬ𝖣​\(h\)\\mathcal\{B\}\_\{\\mathsf\{D\}\}\(h\), scientific semantics remains incomplete\. A hypothesis may have clear testable consequences while still leaving open whether the evidence licenses the language of possibility, preliminary support, mechanistic establishment, general law, or application\-ready discovery\.

The boxed tuple therefore states that a hypothesis has scientific meaning only when five components are specified: the worlds it allows, the consequences it implies, the tests it induces, the rule that evaluates those tests, and the boundary on what may be claimed after evaluation\.

Ordinary LLMs mainly learnP​\(textt\+1∣text≤t\)P\(\\mathrm\{text\}\_\{t\+1\}\\mid\\mathrm\{text\}\_\{\\leq t\}\)\. Scientific semantics requires a family of consequencesP​\(Y∣do⁡\(a\),h,x\)P\(Y\\mid\\operatorname\{do\}\(a\),h,x\)across actions and contexts\. Evidential calibration semantics further requires a boundary on what can be asserted after validation\. The first is linguistic continuation; the second is constraint on the world; the third is permission to make a bounded scientific claim\.

## 5Evidential Calibration Semantics

The preceding section treats the hypothesishhas the central semantic object\. Scientific papers, however, do not end with hypotheses\. They end with claims\. A hypothesis may be proposed, tested, supported, weakened, or rejected, but the manuscript must decide what it is permitted to say\. This makes claim calibration a distinct epistemic operation rather than a stylistic afterthought\.

The deeper point is that a scientific claim is not an ordinary sentence\. It is an assertion act that requires warrant\. Evidence does not merely increase confidence; it licenses certain forms of scientific speech and withholds others\. A proof certificate, an in vitro assay, a simulation, a benchmark improvement, and an independent replication differ not only in strength, but also in the kinds of claims they authorize\. Calibration semantics is therefore a theory of warranted assertability: under a given domain context, what assertion right has the system acquired?

The intended contribution is not to rename philosophical warrant, Toulmin\-style warrants, or evidence\-grading frameworks\(Toulmin,,[1958](https://arxiv.org/html/2606.31273#bib.bib31); Guyatt et al\.,,[2008](https://arxiv.org/html/2606.31273#bib.bib11)\)\. It is to make warrant operational for AI\-assisted research systems whose outputs can shift between hypotheses, artifacts, figures, titles, abstracts, press releases, and user\-facing claims\. The additional object is an output contract: given a raw or implied claim, an evidence body, an evaluator, and a domain context, the system should expose the licensed frontier and the residual epistemic debt rather than merely produce plausible scientific prose\.

For AI\-system analysis, this section defines an output\-contract layer for scientific agents\. A scientific agent may generate many hypotheses, rank them, run code, search literature, call a simulator, submit jobs to a robotic laboratory, or write a manuscript\. These capabilities still leave one question unresolved: which claim is licensed once evaluation is complete? Evidential calibration semantics answers that question by separating three objects that are often conflated in AI systems: the hypothesis being considered, the evidence that has been collected, and the claim that the system makes to the user\.

This distinction matters because AI systems tend to produce fluent assertions even when their evidence is heterogeneous\. A benchmark improvement, a proof certificate, a simulation, a single wet\-lab assay, and a multi\-site replication are all forms of evidence, but they do not license the same language\. A system that treats all successful validation as equivalent may report a claim that is too strong\. The purpose of the calibration layer is to make this failure mode visible and correctable\. In this sense,Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}is not a module for making sentences sound more cautious\. It is the mechanism that determines whether an AI system has the epistemic right to assert a scientific sentence in domain context𝖣\\mathsf\{D\}\.

This section has three layers\. First, it separates hypotheses, claims, literal wording, implied framing, and evidence\. Second, it formalizes evidence\-licensed assertion through a domain\-indexed license relation\. Third, it defines calibration as a projection from raw claims to the maximal evidence\-licensed frontier\. The notation is deliberately repeated near each definition because the section is the theoretical center of the paper\.

### 5\.1Claims Are Not Hypotheses

Let a scientific claim be written as:

C=Claim⁡\(h,s,b,m\),C=\\operatorname\{Claim\}\(h,s,b,m\),\(11\)wherehhis the hypothesis,ssis claim strength,bbis the boundary of applicability, andmmis the modal status of the assertion: possible, supported, demonstrated, explained, discovered, or translated\. The same hypothesis can license very different claims:

- •a weak claim: an AI system proposed an experimentally testable candidate mechanism;
- •a moderate claim: the candidate mechanism received in vitro support under specified cell\-line and assay conditions;
- •a strong claim: the mechanism establishes a general disease principle;
- •an overstrong claim: the system discovered a clinically validated therapy\.

These are not equivalent formulations\. They differ in the evidence required for their assertion\.

In an AI system, the claim objectCCshould be treated as a first\-class output, not as a prose afterthought\. The system may internally manipulate embeddings, programs, proofs, experimental protocols, or molecular structures, but the external scientific product is usually a claim expressed in language\. The representationClaim⁡\(h,s,b,m\)\\operatorname\{Claim\}\(h,s,b,m\)forces that claim to expose four components: what hypothesis it refers to, how strong it is, where it applies, and what modal verb or epistemic status it uses\.

Calibration is therefore not merely weakening\. It is the joint adjustment of strength, scope, modality, and mechanistic content\. In structured form:

Cal𝖣⁡\(Claim⁡\(h,s,b,m\),E,V\)=Claim⁡\(h′,s′,b′,m′\),\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(\\operatorname\{Claim\}\(h,s,b,m\),E,V\)=\\operatorname\{Claim\}\(h^\{\\prime\},s^\{\\prime\},b^\{\\prime\},m^\{\\prime\}\),\(12\)where the formal requirement is that the whole calibrated claim satisfies

Claim⁡\(h′,s′,b′,m′\)⪯𝖣Claim⁡\(h,s,b,m\)\.\\operatorname\{Claim\}\(h^\{\\prime\},s^\{\\prime\},b^\{\\prime\},m^\{\\prime\}\)\\preceq\_\{\\mathsf\{D\}\}\\operatorname\{Claim\}\(h,s,b,m\)\.This claim order may be induced by weaker strength, narrower boundary, weaker modal status, or a less mechanistic hypothesis object, but the paper does not require every component to share a single primitive order\. For example, a raw claim about a causal mechanism may become a calibrated claim about a predictive association; a raw translational claim may become a claim about selected cell\-line settings; a raw discovery claim may become a claim about generated candidates consistent with current evidence\.

Calibration should also apply to implied claims, not only literal sentences\. LetCliteralC\_\{\\mathrm\{literal\}\}denote the exact sentence in a paper, report, title, or system output, and letCimpliedC\_\{\\mathrm\{implied\}\}denote the stronger or broader claim induced by the title, abstract, figure design, framing, press release, or downstream communication\. It is possible that

Cliteral∈ℒ𝖣,V​\(E\)butCimplied∉ℒ𝖣,V​\(E\)\.C\_\{\\mathrm\{literal\}\}\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\\quad\\text\{but\}\\quad C\_\{\\mathrm\{implied\}\}\\notin\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\.\(13\)This is why calibration is not only a sentence\-level editing rule\. It is a publication\-level and system\-output\-level discipline\.

### 5\.2License Relations and Evidence Preorders

Let the domain\-specific claim space be a partially ordered set:

\(𝒞𝖣,⪯𝖣\),\(\\mathcal\{C\}\_\{\\mathsf\{D\}\},\\preceq\_\{\\mathsf\{D\}\}\),\(14\)whereC1⪯𝖣C2C\_\{1\}\\preceq\_\{\\mathsf\{D\}\}C\_\{2\}means thatC1C\_\{1\}is no stronger thanC2C\_\{2\}in domain context𝖣\\mathsf\{D\}\. Strength is not only a matter of adjective choice\. It depends on the claim’s boundary, modal status, intended population, mechanism, intervention, benchmark, material class, or formalization\.

Let the evidence space be:

ℰ𝖣,\\mathcal\{E\}\_\{\\mathsf\{D\}\},\(15\)the set of admissible evidence bodies in domain𝖣\\mathsf\{D\}\. Evidence is not assumed to have a single global strength order\. Instead, evidence comparisons are claim\-relative\. For a comparison family𝒞0⊆𝒞𝖣\\mathcal\{C\}\_\{0\}\\subseteq\\mathcal\{C\}\_\{\\mathsf\{D\}\}, write

E1⊑𝖣,V,𝒞0E2⟺∀C∈𝒞0,Sup𝖣,V⁡\(E2,C\)⪰𝒲,𝖣Sup𝖣,V⁡\(E1,C\)\.E\_\{1\}\\sqsubseteq\_\{\\mathsf\{D\},V,\\mathcal\{C\}\_\{0\}\}E\_\{2\}\\quad\\Longleftrightarrow\\quad\\forall C\\in\\mathcal\{C\}\_\{0\},\\ \\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E\_\{2\},C\)\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E\_\{1\},C\)\.This says thatE2E\_\{2\}is no weaker thanE1E\_\{1\}for the claims being compared\. Formally,⊑𝖣,V,𝒞0\\sqsubseteq\_\{\\mathsf\{D\},V,\\mathcal\{C\}\_\{0\}\}is a preorder on evidence bodies, not necessarily a partial order: two distinct evidence packages can induce the same support profile over𝒞0\\mathcal\{C\}\_\{0\}\. If a partial order is needed, one can quotient by the equivalence relationE1∼E2E\_\{1\}\\sim E\_\{2\}whenE1⊑E2E\_\{1\}\\sqsubseteq E\_\{2\}andE2⊑E1E\_\{2\}\\sqsubseteq E\_\{1\}\. One independent replication may dominate many non\-independent benchmark runs for a biological mechanism claim; a proof certificate may dominate extensive informal argument for a formal theorem; neither necessarily dominates the other for all possible claims\.

The support and requirement profiles live in a warrant\-profile space:

\(𝒲𝖣,⪰𝒲,𝖣\)\.\(\\mathcal\{W\}\_\{\\mathsf\{D\}\},\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\)\.\(16\)The requirement map and support map are typed as

Req𝖣:𝒞𝖣→𝒲𝖣,Sup𝖣,V:ℰ𝖣×𝒞𝖣→𝒲𝖣\.\\operatorname\{Req\}\_\{\\mathsf\{D\}\}:\\mathcal\{C\}\_\{\\mathsf\{D\}\}\\rightarrow\\mathcal\{W\}\_\{\\mathsf\{D\}\},\\qquad\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}:\\mathcal\{E\}\_\{\\mathsf\{D\}\}\\times\\mathcal\{C\}\_\{\\mathsf\{D\}\}\\rightarrow\\mathcal\{W\}\_\{\\mathsf\{D\}\}\.\(17\)HereVVappears in the support function because evaluator provenance, independence, reliability, and grounding affect what the evidence supports\. Throughout this manuscript, the evaluator is kept explicit in the license relation unless a domain\-specific implementation encodes evaluator provenance directly inside the evidence object\.

The central relation is:

E⊩𝖣,VC,E\\Vdash\_\{\\mathsf\{D\},V\}C,\(18\)This relation means that, under domain context𝖣\\mathsf\{D\}and evaluatorVV, evidenceEElicenses claimCC\. It is a license relation, not a truth relation\. A claim may be true but not yet licensed by the available evidence; a claim may be licensed by current evidence and later withdrawn when stronger evidence appears\. Scientific practice does not directly possess truth\. It acquires defensible assertion rights under fallible but auditable evidential norms\.

### 5\.3The Claim\-Evidence Gap and Epistemic Debt

Given a body of evidenceEE, an evaluatorVV, and a claimCC, first define the support\-coverage condition as:

Sup𝖣,V⁡\(E,C\)⪰𝒲,𝖣Req𝖣⁡\(C\)\.\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\.\(19\)HereReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)is the evidential requirement profile imposed by the claim,Sup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)is the support profile that the available evidence and evaluator provide for that claim, and⪰𝒲,𝖣\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}means that the support profile covers the requirement profile\. The primary licensing relation is therefore not a universal scalar score, but a domain\-specific comparison between evidential profiles:

E⊩𝖣,VC⟺Sup𝖣,V⁡\(E,C\)⪰𝒲,𝖣Req𝖣⁡\(C\)\.E\\Vdash\_\{\\mathsf\{D\},V\}C\\quad\\Longleftrightarrow\\quad\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\.\(20\)IfSup𝖣,V⁡\(E,C\)⋡𝒲,𝖣Req𝖣⁡\(C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\\not\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\), the claim exceeds the evidence\.

For engineering use, the profile comparison can be projected to a scalar diagnostic:

Δ𝖣,V​\(C,E\)=ℓ𝖣​\(Req𝖣⁡\(C\)\)−ℓ𝖣​\(Sup𝖣,V⁡\(E,C\)\),\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)=\\ell\_\{\\mathsf\{D\}\}\(\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\)\-\\ell\_\{\\mathsf\{D\}\}\(\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\),\(21\)whereℓ𝖣:𝒲𝖣→ℝ\\ell\_\{\\mathsf\{D\}\}:\\mathcal\{W\}\_\{\\mathsf\{D\}\}\\rightarrow\\mathbb\{R\}is an order\-preserving scalarization used only for diagnostics:

w1⪰𝒲,𝖣w2⇒ℓ𝖣​\(w1\)≥ℓ𝖣​\(w2\)\.w\_\{1\}\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}w\_\{2\}\\Rightarrow\\ell\_\{\\mathsf\{D\}\}\(w\_\{1\}\)\\geq\\ell\_\{\\mathsf\{D\}\}\(w\_\{2\}\)\.Under such a projection,Δ𝖣,V​\(C,E\)\>0\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)\>0means that the claim is stronger than the evidence;Δ𝖣,V​\(C,E\)≤0\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)\\leq 0means that the chosen scoring map treats the evidence as sufficient for that level of assertion\. The scalar gap is useful for audits and dashboards, but it should not be mistaken for a universal theory of evidence\.

The formula has a right\-to\-left interpretation\. First ask what evidence the claim requires,Req𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\. Then ask what the actual evidence and evaluator provide,Sup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\. The gapΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)is the overclaiming pressure that the paper tries to make explicit after a domain has chosen how to score or order evidence\.

This distinction is essential because validation does not have a single semantic value\. An in vitro result, a benchmark improvement, a formal proof, a simulation, and an independent cohort replication all count as validation in some sense, but they license different claims\. The central question becomes: what does this evidence permit the system to say?

For implementation,Δ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)can be interpreted as a diagnostic signal\. IfΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)is positive, the generated claim should be revised before being shown as a scientific conclusion\. IfΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)is near zero, the system is operating near the edge of what the evidence licenses and should report boundary conditions clearly\. IfΔ𝖣,V​\(C,E\)\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)is negative, the evidence is at least sufficient for the current level of assertion, although it may still be insufficient for a stronger claim\.

Overclaiming is not merely an error in wording\. It is epistemic debt:

Debt𝖣,Vℓ⁡\(C,E\)=max⁡\{0,Δ𝖣,V​\(C,E\)\}\.\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C,E\)=\\max\\\{0,\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)\\\}\.\(22\)This is scalarized epistemic debt, because it depends on the diagnostic mapℓ𝖣\\ell\_\{\\mathsf\{D\}\}\. A profile\-level debt object,Debt𝖣,Vprofile⁡\(C,E\)\\operatorname\{Debt\}^\{\\mathrm\{profile\}\}\_\{\\mathsf\{D\},V\}\(C,E\), can also be defined by each domain as the unmet part ofReq𝖣⁡\(C\)\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)after comparison withSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\), but there is no domain\-independent subtraction operation on warrant profiles\. A scientific system can pay this debt by collecting more evidence, such as a new experiment, replication, proof certificate, benchmark, or independent cohort\. It can also cancel the debt by weakening the claim, narrowing its boundary, changing its modal status, or revising the hypothesis object\. A system that does neither accumulates epistemic debt while presenting a stronger assertion than it has earned\.

### 5\.4Licensed Claim Sets and the Calibration Operator

Define the set of claims licensed by evidenceEEas:

ℒ𝖣,V​\(E\)=\{C∈𝒞𝖣:E⊩𝖣,VC\}\.\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)=\\\{C\\in\\mathcal\{C\}\_\{\\mathsf\{D\}\}:E\\Vdash\_\{\\mathsf\{D\},V\}C\\\}\.\(23\)This set should be downward closed in the claim poset, provided both requirement monotonicity and claim\-relative support monotonicity hold\. Requirement monotonicity says:

C′⪯𝖣C⇒Req𝖣⁡\(C′\)⪯𝒲,𝖣Req𝖣⁡\(C\),C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C\\quad\\Rightarrow\\quad\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C^\{\\prime\}\)\\preceq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\),\(24\)where⪯𝒲,𝖣\\preceq\_\{\\mathcal\{W\},\\mathsf\{D\}\}is the reverse of⪰𝒲,𝖣\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}and means “requires no more warrant than”\. Claim\-relative support monotonicity says:

C′⪯𝖣C⇒Sup𝖣,V⁡\(E,C′\)⪰𝒲,𝖣Sup𝖣,V⁡\(E,C\)\.C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C\\quad\\Rightarrow\\quad\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C^\{\\prime\}\)\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\.\(25\)This means that the same evidence is not less supportive of a weaker version of the same claim\. Equivalently, a domain may take downward license monotonicity as a primitive norm:

C∈ℒ𝖣,V​\(E\),C′⪯𝖣C⇒C′∈ℒ𝖣,V​\(E\)\.C\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\),\\quad C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C\\quad\\Rightarrow\\quad C^\{\\prime\}\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\.\(26\)In words: if the evidence licenses a strong claim, it also licenses all weaker claims, assuming weaker claims require no stronger warrant than stronger claims and the same evidence is not less supportive of a weaker assertion\. Thusℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)is a lower set of the domain\-specific claim space\.

To make calibration total, assume that each domain has a bottom claim⊥𝖣∈𝒞𝖣\\bot\_\{\\mathsf\{D\}\}\\in\\mathcal\{C\}\_\{\\mathsf\{D\}\}, such as “no scientific claim is licensed beyond proposal\.” The bottom claim satisfies⊥𝖣⪯𝖣C\\bot\_\{\\mathsf\{D\}\}\\preceq\_\{\\mathsf\{D\}\}Cfor allC∈𝒞𝖣C\\in\\mathcal\{C\}\_\{\\mathsf\{D\}\}and is always licensed at the minimal proposal level,⊥𝖣∈ℒ𝖣,V\(E\)\\bot\_\{\\mathsf\{D\}\}\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\. This prevents the calibration set from being empty even when the raw claim is entirely unsupported\.

Let the down\-set of the raw claim be:

↓C=\{C′∈𝒞𝖣:C′⪯𝖣C\}\.\\downarrow C=\\\{C^\{\\prime\}\\in\\mathcal\{C\}\_\{\\mathsf\{D\}\}:C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C\\\}\.\(27\)The task of claim calibration is to find the maximal licensed weakening frontier of the raw claim:

Cal𝖣⁡\(C,E,V\)=Max⪯𝖣\(↓C∩ℒ𝖣,V​\(E\)\)\.\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)=\\operatorname\{Max\}\_\{\\preceq\_\{\\mathsf\{D\}\}\}\\left\(\\downarrow C\\cap\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\\right\)\.\(28\)ThusCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}is set\-valued in general:

Cal𝖣:𝒞𝖣×ℰ𝖣×𝒱𝖣→𝒫​\(𝒞𝖣\)\.\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}:\\mathcal\{C\}\_\{\\mathsf\{D\}\}\\times\\mathcal\{E\}\_\{\\mathsf\{D\}\}\\times\\mathcal\{V\}\_\{\\mathsf\{D\}\}\\rightarrow\\mathcal\{P\}\(\\mathcal\{C\}\_\{\\mathsf\{D\}\}\)\.Because⊥𝖣∈↓C∩ℒ𝖣,V\(E\)\\bot\_\{\\mathsf\{D\}\}\\in\\downarrow C\\cap\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\), the intersection is non\-empty\. If it has a unique maximum, the frontier contains one calibrated claim\. If there are several incomparable maximal claims,Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}returns a frontier rather than a single sentence\. A domain\-specific selection ruleσ𝖣\\sigma\_\{\\mathsf\{D\}\}may then choose a reportable claim,

C∗=σ𝖣​\(Cal𝖣⁡\(C,E,V\)\),C^\{\\ast\}=\\sigma\_\{\\mathsf\{D\}\}\(\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\),according to domain convention, audience, and reporting purpose, but it should not step outside the frontier\.

In practical reporting settings, the claim space is implemented as a finite or discretized reporting lattice, such as the L0–L6 ladder introduced below, so maximal licensed weakenings exist\. More abstractly, the framework assumes that each feasible set considered byCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}has at least one maximal element\.

The important output is the licensed frontier, and then the selected claimC∗C^\{\\ast\}if the system or author needs a single sentence\.C∗C^\{\\ast\}is allowed to be weaker than the original claimCC; in fact, downgrading an overstrong claim is one function ofCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\. ButCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}can also narrow the boundary, change the assertion modality, or revise the mechanistic object\. Calibration is therefore a semantic transformation, not a conservative language filter\.

#### Structural compression and minimal reconstruction\.

A less obvious form of calibration is not a downward weakening but an upward change of level\. Several heterogeneous outputs may each support only a local claim, yet together suggest a common structure\. In model selection, the minimum\-description\-length tradition treats learning as the extraction of regularity that supports a shorter description of data\(Rissanen,,[1978](https://arxiv.org/html/2606.31273#bib.bib26); Grünwald and Roos,,[2019](https://arxiv.org/html/2606.31273#bib.bib10)\)\. In philosophy of science, unification accounts treat explanation as reducing heterogeneous phenomena to a smaller family of explanatory patterns\(Friedman,,[1974](https://arxiv.org/html/2606.31273#bib.bib6); Kitcher,,[1981](https://arxiv.org/html/2606.31273#bib.bib15)\)\. In mathematics, understanding often consists not only in checking isolated propositions, but in finding reusable methods, definitions, and ways of thinking that reconstruct many results at once\(Thurston,,[1994](https://arxiv.org/html/2606.31273#bib.bib29); Avigad,,[2006](https://arxiv.org/html/2606.31273#bib.bib2)\)\. These traditions suggest a special claim\-calibration problem: the strongest licensed contribution may be neither a narrower local claim nor a broad unqualified discovery claim, but a minimal structural reconstruction of what multiple local claims have in common\. In this sense, structural compression is claim calibration across heterogeneous contents: it asks what common claim, if any, can be asserted over a family rather than over each member separately\.

Formally, letℱ=\{C1,…,Ck\}⊆𝒞𝖣\\mathcal\{F\}=\\\{C\_\{1\},\\ldots,C\_\{k\}\\\}\\subseteq\\mathcal\{C\}\_\{\\mathsf\{D\}\}be a comparison family of claims, and letEℱE\_\{\\mathcal\{F\}\}andVℱV\_\{\\mathcal\{F\}\}record the evidence and evaluator provenance for that family\. A candidate compressed claim is a minimal structural reconstruction when it is a minimal upper claim for the family:

Cstr∈Min⪯𝖣⁡\{C∈𝒞𝖣:∀i,Ci⪯𝖣C\}\.C\_\{\\mathrm\{str\}\}\\in\\operatorname\{Min\}\_\{\\preceq\_\{\\mathsf\{D\}\}\}\\\{C\\in\\mathcal\{C\}\_\{\\mathsf\{D\}\}:\\forall i,\\ C\_\{i\}\\preceq\_\{\\mathsf\{D\}\}C\\\}\.\(29\)It becomes assertable only through the ordinary calibration operator:

Cstr∗∈Cal𝖣⁡\(Cstr,Eℱ,Vℱ\)\.C\_\{\\mathrm\{str\}\}^\{\\ast\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C\_\{\\mathrm\{str\}\},E\_\{\\mathcal\{F\}\},V\_\{\\mathcal\{F\}\}\)\.\(30\)Here “upper” names the level of reconstruction relative to the comparison family, not an automatic increase in evidential authority: the reconstructed claim may speak at a broader structural level only to the extent thatEℱE\_\{\\mathcal\{F\}\}andVℱV\_\{\\mathcal\{F\}\}license the shared structure and leave route\-specific residuals explicit\. Non\-amplification applies afterCstrC\_\{\\mathrm\{str\}\}has been posed as a new family\-level raw claim; it does not permitCal𝖣⁡\(Ci,Ei,Vi\)\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C\_\{i\},E\_\{i\},V\_\{i\}\)to strengthen an individual local claim without additional family\-level evidence\. This is an upward or structural use of calibration\. Compression may organize heterogeneous licensed claims, but any additional generality, causal interpretation, cross\-context scope, or discovery language introduced by the reconstruction must itself be licensed\. A shorter common structure is therefore not automatically a truer or stronger scientific claim; it is a candidate structural claim whose assertion rights depend on the family\-level evidence and the residual differences that the reconstruction leaves outside its scope\.

This is whyCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}is separated fromVV\. The evaluatorVVanswers whether a test was passed, a benchmark improved, a proof checked, or an assay produced the expected response\. The calibration operatorCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}answers a different question: given that result, what assertion right has the system acquired? A single positive test may justify “is consistent with”, “supports”, or “demonstrates under these conditions”, but not necessarily “discovers”, “establishes”, or “translates to clinical use”\.

#### Minimal desiderata for calibration\.

A calibration operator should satisfy at least four conditions\. First, it should be sound:

Cal𝖣⁡\(C,E,V\)⊆ℒ𝖣,V​\(E\),\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\\subseteq\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\),\(31\)Second, it should be non\-amplifying:

∀C′∈Cal𝖣⁡\(C,E,V\),C′⪯𝖣C,\\forall C^\{\\prime\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\),\\quad C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C,\(32\)so calibration may weaken an overstrong claim but should not strengthen it without additional evidence\. Third, it should be monotone with respect to evidential support:

\[∀C′⪯𝖣C,Sup𝖣,V⁡\(E′,C′\)⪰𝒲,𝖣Sup𝖣,V⁡\(E,C′\)\]⇒CE∗⪯𝖣CE′∗,\\left\[\\forall C^\{\\prime\}\\preceq\_\{\\mathsf\{D\}\}C,\\ \\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E^\{\\prime\},C^\{\\prime\}\)\\succeq\_\{\\mathcal\{W\},\\mathsf\{D\}\}\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C^\{\\prime\}\)\\right\]\\Rightarrow C\_\{E\}^\{\\ast\}\\preceq\_\{\\mathsf\{D\}\}C\_\{E^\{\\prime\}\}^\{\\ast\},\(33\)whereCE∗=σ𝖣​\(Cal𝖣⁡\(C,E,V\)\)C\_\{E\}^\{\\ast\}=\\sigma\_\{\\mathsf\{D\}\}\(\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\)andCE′∗=σ𝖣​\(Cal𝖣⁡\(C,E′,V\)\)C\_\{E^\{\\prime\}\}^\{\\ast\}=\\sigma\_\{\\mathsf\{D\}\}\(\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E^\{\\prime\},V\)\)are selected maximal elements under a fixed domain selection rule\. That is, if a later evidence body provides uniformly stronger support for every weakening of the same raw claim, the selected calibrated claim should not become weaker\. In the set\-valued version, the downward closure of the licensed frontier should not shrink\. Stronger evidence may license the same or a stronger calibrated claim, but should not force a downgrade unless the domain context or evaluator changes\. Fourth, it should preserve boundaries: the calibrated claim must state the domain, assay, benchmark, population, material class, or formalization under which the evidence was obtained\. These conditions makeCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}an assertion\-right layer rather than a stylistic rewriting module\.

When the maximal frontier is not unique, the monotonicity statement assumes a fixed selection rule compatible withStrength\\operatorname\{Strength\}\. Equivalently, the set\-valued interpretation is that stronger evidence should not shrink the downward closure of the licensed frontier for the same raw claim and domain context\.

The resulting theory has four nested levels\. The semantic level specifies

ScientificMeaning𝖣​\(h\)=\(\[\[h\]\],Conseq𝖣⁡\(h\),𝒯𝖣​\(h\),𝒱𝖣​\(h\),ℬ𝖣​\(h\)\)\.\\mathrm\{ScientificMeaning\}\_\{\\mathsf\{D\}\}\(h\)=\(\\left\[\\\!\\left\[h\\right\]\\\!\\right\],\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{T\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{V\}\_\{\\mathsf\{D\}\}\(h\),\\mathcal\{B\}\_\{\\mathsf\{D\}\}\(h\)\)\.The license level asks whetherE⊩𝖣,VCE\\Vdash\_\{\\mathsf\{D\},V\}C\. The calibration level computes

Cal𝖣⁡\(C,E,V\)=Max⪯𝖣\(↓C∩ℒ𝖣,V​\(E\)\)\.\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)=\\operatorname\{Max\}\_\{\\preceq\_\{\\mathsf\{D\}\}\}\(\\downarrow C\\cap\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\)\.The dynamical level, introduced later, asks whether generation and consequence\-derivation capacity grow faster than adjudication and calibration capacity\. This four\-layer structure is the bridge from semantic theory to AI\-system design\.

### 5\.5Dimensions of Evidential Support

The support profileSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)is not one\-dimensional\. It may be decomposed as:

Sup𝖣,V⁡\(E,C\)=F𝖣,V​\(d,i,r,e,cid,ξ,n\),\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)=F\_\{\\mathsf\{D\},V\}\(d,i,r,e,c\_\{\\mathrm\{id\}\},\\xi,n\),\(34\)whereddis directness of the test,iiis independence from the generating system,rris reproducibility,eeis effect size,cidc\_\{\\mathrm\{id\}\}is causal identification,ξ\\xiis external validity, andnncaptures negative evidence, failed cases, or counterexamples\. This decomposition turns claim calibration from a vague philosophical caution into an auditable checklist\.

ThusSup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)is not just a confidence score\. A claim can have a large effect size but weak external validity, or strong computational support but little independence from the generating system\. In this paper,Sup𝖣,V⁡\(E,C\)\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)is treated as a domain\- and evaluator\-specific warrant functional rather than a universal scalar\. Different domains may instantiateF𝖣,VF\_\{\\mathsf\{D\},V\}as a checklist, ordinal rubric, Bayesian score, proof\-verification predicate, benchmark protocol, or expert\-audited evidence profile\. The calibration operator should see these differences before it licenses a claim\.

The ladder is not intended as a universal hierarchy of all sciences\. It is a reporting template: each domain must instantiate what counts as support, replication, causality, generalization, and application\. A formal proof, a wet\-lab assay, a benchmark result, and a materials synthesis campaign occupy different evidential geometries even when they are all called validation\.

Table 3:Domain\-parametric claim ladder for AI\-driven scientific assertionsMuch of the evidence currently reported in AI science systems lies around L1–L3\. Public\-facing language, however, can drift toward L4–L6\. The role of evidential calibration semantics is to prevent this drift by forcing the final claim to remain insideℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\.

The ladder is domain\-parametric rather than universal\. In mathematics, a proof accepted by a trusted formal checker may license theorem\-level claims without experimental replication\. In biology, in vitro support does not license a translational or clinical claim\. In machine learning, a benchmark improvement may license a constrained performance claim, but not necessarily a mechanistic discovery claim\. In materials science, synthesis of a target compound is not automatically equivalent to discovery of a new material unless phase identification, novelty, and reproducibility have also been established\. The same ladder therefore functions as a template whose evidential thresholds must be instantiated by domain\-specific norms, similar in spirit to evidence\-grading frameworks in medicine\(Guyatt et al\.,,[2008](https://arxiv.org/html/2606.31273#bib.bib11)\)\.

### 5\.6Worked Example: Co\-Scientist and AML Drug Repurposing

Consider an overstrong reading: Co\-Scientist discovered new AML treatments\. The evidence reported for Co\-Scientist is more specific: the system generated drug\-repurposing candidates and synergistic combination hypotheses for acute myeloid leukemia, with selected candidates validated in vitro\(Gottweis et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib9)\)\. A calibrated claim is therefore:

> Co\-Scientist generated drug\-repurposing candidates for AML, and selected candidates showed in vitro anti\-tumor activity in specific AML cell\-line settings under expert\-guided validation\.

This claim is stronger than mere hypothesis generation, because it includes reported in vitro support beyond model output\. It is weaker than saying that Co\-Scientist discovered a clinically validated AML therapy\. The difference is not rhetorical\. It is the difference between within\-study laboratory validation under specified assay conditions and a translational or clinically validated therapeutic claim\.

### 5\.7Worked Example: AI Scientist and Research Artifacts

Consider a strong raw claim: The AI Scientist autonomously performs scientific discovery\. The evidence reported for AI Scientist systems is more constrained: the system can generate machine\-learning ideas, write code, execute experiments, produce figures, draft manuscripts, and obtain workshop\-level review outcomes under a restricted computational research setting\(Lu et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib17)\)\. A calibrated claim is therefore:

> The AI Scientist automates substantial portions of machine\-learning research artifact generation and can produce workshop\-level manuscripts under constrained benchmark and review settings\.

This calibrated claim recognizes the system’s contribution without upgrading artifact production into general autonomous science\. In this case,Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}does not deny the system’s automation capability; it prevents workflow completion and workshop\-level review from being redescribed as general scientific autonomy\. The claim may be strong at the level of workflow automation, but it remains weaker than saying that the system has established a reliable general mechanism for producing novel scientific knowledge across domains\.

### 5\.8Worked Example: A\-Lab and Materials Claims

Consider a commonly circulated strong claim: A\-Lab discovered 43 new materials\. After the correction and subsequent discussion, the more careful evidential reading is narrower: the autonomous laboratory realized a set of target compounds under specified solid\-state synthesis and characterization criteria, while claims of novelty require phase identification, database deduplication, and independent review\(Szymanski et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib27),[2026](https://arxiv.org/html/2606.31273#bib.bib28); Chawla,,[2026](https://arxiv.org/html/2606.31273#bib.bib5)\)\. A calibrated claim is therefore:

> A\-Lab demonstrated an automated workflow for proposing, synthesizing, and characterizing inorganic target compounds over a short campaign, while the strength of new\-material discovery claims depends on corrected phase assignment, novelty checks, and independent materials review\.

This example shows why world contact is not itself sufficient for claim licensing\. A robotic laboratory can produce real experimental evidence, but the final scientific claim still depends on how the evidence is classified, compared with prior databases, and bounded by independent expertise\.

### 5\.9Relation to Typed Artifact Discovery

Frameworks such as CategoryScienceClaw, as proposed in the arXiv preprint on self\-revising discovery systems byWang and Buehler, \([2026](https://arxiv.org/html/2606.31273#bib.bib33)\), solve a complementary problem: they formalize typed artifacts, provenance, verifiers, and regime transitions in agentic scientific discovery systems\. In such a typed\-artifact view, discovery is modeled as a regime transition in which scientific artifacts, verifiers, provenance, and schema revisions are preserved and audited\. A schematic residual can be written as:

RCSC​\(A′\)=It\+1′​\(A′\)∖im⁡\(ρA′\),R\_\{\\mathrm\{CSC\}\}\(A^\{\\prime\}\)=I^\{\\prime\}\_\{t\+1\}\(A^\{\\prime\}\)\\setminus\\operatorname\{im\}\(\\rho\_\{A^\{\\prime\}\}\),\(35\)which measures whether a new artifact exceeds what can be transported from the previous schema\. The present framework instead asks whether a scientific claim exceeds what the evidence licenses:

Δ𝖣,V​\(C,E\)=ℓ𝖣​\(Req𝖣⁡\(C\)\)−ℓ𝖣​\(Sup𝖣,V⁡\(E,C\)\)\.\\Delta\_\{\\mathsf\{D\},V\}\(C,E\)=\\ell\_\{\\mathsf\{D\}\}\(\\operatorname\{Req\}\_\{\\mathsf\{D\}\}\(C\)\)\-\\ell\_\{\\mathsf\{D\}\}\(\\operatorname\{Sup\}\_\{\\mathsf\{D\},V\}\(E,C\)\)\.\(36\)The former is an artifact\-regime residual; the latter is a claim\-evidence gap\. A new artifact is not yet a licensed claim\. Typed provenance can show where a claim came from, but not by itself how strong the claim is allowed to be\. Typed\-artifact systems address how scientific workflows can be typed, preserved, and audited\. Evidential semantics addresses how scientific claims are permitted, restricted, or downgraded by evidence\.

## 6Six Routes and Their Epistemological Structure

### 6\.1Specialized Scientific Foundation Models

Specialized scientific foundation models are not autonomous scientists\. They are strong predictors for particular scientific problems\. AlphaFold 3 targets biomolecular complex structure prediction\. The AlphaFold Protein Structure Database is an open database of predicted protein structures and covers more than 200 million predicted entries\(Varadi et al\.,,[2024](https://arxiv.org/html/2606.31273#bib.bib32); Google DeepMind and EMBL\-EBI,,[2026](https://arxiv.org/html/2606.31273#bib.bib8)\)\. In materials search,Merchant et al\., \([2023](https://arxiv.org/html/2606.31273#bib.bib19)\)used GNoME to computationally identify 2\.2 million structures stable with respect to the Materials Project, of which 381,000 lie on the updated convex hull\.

The basic form is supervised learning or generative modeling:

θ∗=arg​minθ⁡𝔼\(x,y\)∼𝒟train​\[ℓ​\(fθ​\(x\),y\)\],\\theta^\{\\ast\}=\\operatorname\*\{arg\\,min\}\_\{\\theta\}\\mathbb\{E\}\_\{\(x,y\)\\sim\\mathcal\{D\}\_\{\\mathrm\{train\}\}\}\[\\ell\(f\_\{\\theta\}\(x\),y\)\],\(37\)or probabilisticallypθ​\(y∣x\)p\_\{\\theta\}\(y\\mid x\)\. In materials discovery one may write:

x∗=arg​maxx⁡Pθ​\(stable∣x\)\.x^\{\\ast\}=\\operatorname\*\{arg\\,max\}\_\{x\}P\_\{\\theta\}\(\\mathrm\{stable\}\\mid x\)\.\(38\)The implicit epistemology is:

to know≈to predict structure, energy, or property accurately\.\\boxed\{\\text\{to know\}\\approx\\text\{to predict structure, energy, or property accurately\}\.\}The advantage is high\-throughput screening:\|X\|≫\|Xcandidate\|\|X\|\\gg\|X\_\{\\mathrm\{candidate\}\}\|\. The risk is:

pθ​\(y∣x\)≠Pworld​\(y∣do⁡\(x\)\)\.p\_\{\\theta\}\(y\\mid x\)\\neq P\_\{\\mathrm\{world\}\}\(y\\mid\\operatorname\{do\}\(x\)\)\.\(39\)A prediction is not an explanation, and a candidate is not a validation\. These models are best understood as search\-space compressors, not as final scientific justifiers\.

### 6\.2Human\-Led LLM Research Assistants

An LLM research assistant may be written as:

h∼Pϕ​\(h∣𝒦,q,Dt\),h\\sim P\_\{\\phi\}\(h\\mid\\mathcal\{K\},q,D\_\{t\}\),\(40\)where𝒦\\mathcal\{K\}is the literature and knowledge base,qqis the research question,DtD\_\{t\}is available data accumulated up to timett, andhhis a hypothesis, explanation, experimental suggestion, or proof sketch\. A rough optimization form is:

h∗=arg​maxh⁡\[log⁡Pϕ​\(h∣𝒦,q,Dt\)\+λ​Scoherence​\(h\)\+μ​Snovelty​\(h\)\]\.h^\{\\ast\}=\\operatorname\*\{arg\\,max\}\_\{h\}\\left\[\\log P\_\{\\phi\}\(h\\mid\\mathcal\{K\},q,D\_\{t\}\)\+\\lambda S\_\{\\mathrm\{coherence\}\}\(h\)\+\\mu S\_\{\\mathrm\{novelty\}\}\(h\)\\right\]\.\(41\)The GPT\-5 science acceleration report presents cases in which frontier LLMs helped experts connect literature, draft proofs, explore computations, and generate testable hypotheses\. The same report stresses that the examples are curated, not a systematic sample, and that models may still hallucinate citations, mechanisms, or proofs\(Bubeck et al\.,,[2025](https://arxiv.org/html/2606.31273#bib.bib4); OpenAI,,[2025](https://arxiv.org/html/2606.31273#bib.bib22)\)\.

The underlying view of science is:

scientific discovery≈recombining high\-value explanations from prior knowledge\.\\boxed\{\\text\{scientific discovery\}\\approx\\text\{recombining high\-value explanations from prior knowledge\}\.\}The strength of this route is𝒦→h\\mathcal\{K\}\\rightarrow h: the generation of hypotheses from a knowledge base\. Its weakness ish→Conseq𝖣⁡\(h\)h\\rightarrow\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\)\. LLM semantics is closer toSemLLM⁡\(h\)=embedding​\(h,𝒦\)\\operatorname\{Sem\}\_\{\\mathrm\{LLM\}\}\(h\)=\\mathrm\{embedding\}\(h,\\mathcal\{K\}\), whereas scientific semantics requires a domain\-appropriate consequence profile such asConseq𝖣⁡\(h\)\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\), withPh​\(Y∣do⁡\(a\),c\)P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),c\)as an interventional special case\. LLM assistants should therefore be treated as human\-in\-the\-loop exploration expanders, not as independent bearers of truth\.

### 6\.3Multi\-Agent Co\-Scientist Systems

Multi\-agent co\-scientist systems treat science as a cycle of hypothesis generation, critique, rewriting, ranking, and experimental planning\. LetHt=\{h1,…,hn\}H\_\{t\}=\\\{h\_\{1\},\\ldots,h\_\{n\}\\\}be the current hypothesis population\. Define a generatorGgenG\_\{\\mathrm\{gen\}\}, criticCcritC\_\{\\mathrm\{crit\}\}, reviserRrevR\_\{\\mathrm\{rev\}\}, and selectorSselS\_\{\\mathrm\{sel\}\}:

Ht\+1=Ssel∘Rrev∘Ccrit∘Ggen​\(Ht,𝒦,q,Dt\)\.H\_\{t\+1\}=S\_\{\\mathrm\{sel\}\}\\circ R\_\{\\mathrm\{rev\}\}\\circ C\_\{\\mathrm\{crit\}\}\\circ G\_\{\\mathrm\{gen\}\}\(H\_\{t\},\\mathcal\{K\},q,D\_\{t\}\)\.\(42\)A hypothesis score can be written as:

s​\(h\)=λ1​N​\(h\)\+λ2​P​\(h∣𝒦,Dt\)\+λ3​T​\(h\)\+λ4​F​\(h\)\+λ5​I​\(h;Y\)−λ6​Ctest​\(h\),s\(h\)=\\lambda\_\{1\}N\(h\)\+\\lambda\_\{2\}P\(h\\mid\\mathcal\{K\},D\_\{t\}\)\+\\lambda\_\{3\}T\(h\)\+\\lambda\_\{4\}F\(h\)\+\\lambda\_\{5\}I\(h;Y\)\-\\lambda\_\{6\}C\_\{\\mathrm\{test\}\}\(h\),\(43\)whereNNis novelty,TTis testability,FFis feasibility, andI​\(h;Y\)I\(h;Y\)is expected information gain\.

Gottweis et al\., \([2026](https://arxiv.org/html/2606.31273#bib.bib9)\)describe Google Co\-Scientist as a Gemini\-based multi\-agent system that continuously generates, critiques, refines, and evolves hypotheses using tournament evolution\. Google Co\-Scientist is distinct from Boiko et al\.’s earlier Coscientist chemistry system, which is discussed below under self\-driving laboratories\. They evaluate Google Co\-Scientist in biomedical applications including drug repurposing, target discovery, and antimicrobial resistance explanation, with AML drug\-repurposing candidates and synergistic combinations tested in vitro\. FutureHouse Robin integrates literature\-search and data\-analysis agents for experimental biology, including hypothesis generation, experimental suggestions, interpretation of results, and hypothesis updates\(Ghareeb et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib7)\)\.

This route combines abductive inference, Popperian falsifiability, and social epistemology\. Its importance is not that it merely simulates debate inside a scientific community, but that it pushes multi\-agent hypothesis generation toward experimental actionability\. It can generate, criticize, rank, and translate hypotheses into tests\. The more precise epistemological question is therefore not whether there is any validation, but which level of claim the validation supports\.

If evidence shows that selected AI\-generated candidates are effective under specified in vitro conditions, the licensed claim is that the system generated candidates that received in vitro support in expert\-guided settings\. This does not automatically license stronger claims, such as autonomous scientific discovery in general or clinically translatable therapy\. Co\-scientist systems therefore expose a calibration problem: experimental actionability can be mistaken for a broader scientific discovery unlessCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}explicitly maps evidence to a bounded claim\.

### 6\.4End\-to\-End AI Scientist Pipelines

The AI Scientist route turns a research trajectory into:

z=\(h,c,e,r,p\),z=\(h,c,e,r,p\),\(44\)wherehhis an idea,ccis code,eeis experimental output,rris analysis, andppis the manuscript\. The objective can be written as:

J​\(z\)=λ1​Sbenchmark​\(e\)\+λ2​Snovelty​\(h\)\+λ3​Spaper​\(p\)\+λ4​Sreview​\(p\)−λ5​C​\(z\)\.J\(z\)=\\lambda\_\{1\}S\_\{\\mathrm\{benchmark\}\}\(e\)\+\\lambda\_\{2\}S\_\{\\mathrm\{novelty\}\}\(h\)\+\\lambda\_\{3\}S\_\{\\mathrm\{paper\}\}\(p\)\+\\lambda\_\{4\}S\_\{\\mathrm\{review\}\}\(p\)\-\\lambda\_\{5\}C\(z\)\.\(45\)Lu et al\., \([2026](https://arxiv.org/html/2606.31273#bib.bib17)\)describe The AI Scientist as a system that creates research ideas, writes code, runs experiments, plots and analyzes data, writes manuscripts, and performs its own peer review\. One generated manuscript exceeded the average human acceptance threshold in the first round of review for an ICLR workshop, but the authors explicitly note that the workshop had a lower bar than a main conference track\.

The implicit epistemology is:

scientific research≈an executable workflow that can be written and reviewed\.\\boxed\{\\text\{scientific research\}\\approx\\text\{an executable workflow that can be written and reviewed\}\.\}The advantage is speed and process automation, especially in machine learning, simulation, and low\-cost computational experiments\. The risk is Goodhart’s law:

max⁡Spaper⇏max⁡Truth\.\\max S\_\{\\mathrm\{paper\}\}\\not\\Rightarrow\\max\\mathrm\{Truth\}\.\(46\)If the system optimizes paper form, small benchmark improvements, and automatic review scores, it may generate research that is formally complete but epistemically weak\.

### 6\.5Algorithmic and Mathematical Discovery Agents

For AlphaEvolve\-like systems, the object is not a textual hypothesis but a program, algorithm, or proof\. Letp∈𝒫p\\in\\mathcal\{P\}andV:𝒫→ℝV:\\mathcal\{P\}\\rightarrow\\mathbb\{R\}\. The target is:

p∗=arg​maxp∈𝒫⁡V​\(p\)\.p^\{\\ast\}=\\operatorname\*\{arg\\,max\}\_\{p\\in\\mathcal\{P\}\}V\(p\)\.\(47\)For proof tasks, the target is a certificateπ\\pisuch that:

∃π:Verifier⁡\(π,T\)=1\.\\exists\\pi:\\operatorname\{Verifier\}\(\\pi,T\)=1\.\(48\)AlphaEvolve combines Gemini\-based generation, automated evaluator feedback, and evolutionary search for selected machine\-gradable tasks\. Its white paper describes applications to data\-center scheduling, circuit simplification for hardware accelerators, LLM training efficiency, matrix multiplication, and problems in mathematics and computer science\(Novikov et al\.,,[2025](https://arxiv.org/html/2606.31273#bib.bib20)\)\. AlphaProof and AlphaGeometry 2 solved four of six IMO 2024 problems, scoring 28 out of 42 points, in the silver\-medal range; the reported problem\-specific reasoning substantially exceeded human contest time limits\(Hubert et al\.,,[2026](https://arxiv.org/html/2606.31273#bib.bib13)\)\.

The scientific view is:

to know≈to construct an object accepted by a verifier\.\\boxed\{\\text\{to know\}\\approx\\text\{to construct an object accepted by a verifier\}\.\}If the evaluator is reliable,V​\(p\)=1V\(p\)=1orCheck⁡\(π,T\)=1\\operatorname\{Check\}\(\\pi,T\)=1provides a strong epistemic warrant\. The limitation is equally clear:

science​domain⊆domains​with​evaluators\.\\mathrm\{science\\ domain\}\\subseteq\\mathrm\{domains\\ with\\ evaluators\}\.Algorithms, proofs, code, compiler optimization, and benchmarks suit this route\. Complex biology, clinical medicine, ecosystems, and social science often lack cheap, clear, low\-noise evaluators\.

### 6\.6Self\-Driving Laboratories

Self\-driving laboratories come closest to a scientific closed loop\. They are naturally expressed as Bayesian active learning or a partially observable Markov decision process\. Givenbt​\(ω\)=P​\(ω∣Dt\)b\_\{t\}\(\\omega\)=P\(\\omega\\mid D\_\{t\}\), the next experiment is:

at=arg​maxa∈A⁡\[𝔼y∼P​\(y∣a,bt\)​U​\(y,a\)\+η​I​\(ω;Ya∣Dt\)−λ​Cost⁡\(a\)−κ​Risk⁡\(a\)\]\.a\_\{t\}=\\operatorname\*\{arg\\,max\}\_\{a\\in A\}\\left\[\\mathbb\{E\}\_\{y\\sim P\(y\\mid a,b\_\{t\}\)\}U\(y,a\)\+\\eta I\(\\omega;Y\_\{a\}\\mid D\_\{t\}\)\-\\lambda\\operatorname\{Cost\}\(a\)\-\\kappa\\operatorname\{Risk\}\(a\)\\right\]\.\(49\)After the experiment,yt∼P​\(y∣do⁡\(at\),ω∗\)y\_\{t\}\\sim P\(y\\mid\\operatorname\{do\}\(a\_\{t\}\),\\omega^\{\\ast\}\)and:

bt\+1​\(ω\)=P​\(yt∣do⁡\(at\),ω\)​bt​\(ω\)∫ΩP​\(yt∣do⁡\(at\),ω′\)​bt​\(ω′\)​𝑑ω′\.b\_\{t\+1\}\(\\omega\)=\\frac\{P\(y\_\{t\}\\mid\\operatorname\{do\}\(a\_\{t\}\),\\omega\)b\_\{t\}\(\\omega\)\}\{\\int\_\{\\Omega\}P\(y\_\{t\}\\mid\\operatorname\{do\}\(a\_\{t\}\),\\omega^\{\\prime\}\)b\_\{t\}\(\\omega^\{\\prime\}\)d\\omega^\{\\prime\}\}\.\(50\)
Boiko et al\.’s Coscientist chemistry system demonstrates a GPT\-4\-driven workflow that combines internet and documentation search, code execution, and experimental automation to design, plan, and execute chemistry experiments, including palladium\-catalyzed cross\-coupling tasks\(Boiko et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib3)\)\. A review of self\-driving laboratories describes SDLs as systems that combine automated experimental workflows with autonomous experimental planning and hold potential for accelerating chemistry and materials discovery\(Tom et al\.,,[2024](https://arxiv.org/html/2606.31273#bib.bib30)\)\. A\-Lab integrates computation, literature\-derived historical data, machine learning, active learning, and robotic solid\-state synthesis\. After Nature’s correction boundary, it should be described as realizing target compounds under corrected structural\-characterization criteria, not as a settled discovery of 43 scientifically new materials\(Szymanski et al\.,,[2023](https://arxiv.org/html/2606.31273#bib.bib27),[2026](https://arxiv.org/html/2606.31273#bib.bib28)\)\.

The scientific view is:

scientific knowledge≈a stable regularity under intervention\.\\boxed\{\\text\{scientific knowledge\}\\approx\\text\{a stable regularity under intervention\}\.\}This route is expensive and difficult to scale\. Experiments cost far more than text generation; results are noisy,y=f​\(a,ω\)\+ϵy=f\(a,\\omega\)\+\\epsilon; and the action space is constrained by instruments, robots, safety, and protocols\. The A\-Lab controversy also shows that contact with the physical world does not remove interpretive risk\. It increases the need for experimental audit, data standards, phase identification, novelty checks, and independent replication\(Chawla,,[2026](https://arxiv.org/html/2606.31273#bib.bib5)\)\.

## 7Comparison Across Routes

Table 4:Epistemological comparison of AI\-driven science routes and artifact frameworksThe table shows that these routes do not solve the same problem\. Specialized predictive models place meaning in output distributions\. LLMs place meaning in corpus relations\. Co\-scientist systems place meaning in a hypothesis competition network\. AI Scientist systems place meaning in executable research artifacts\. AlphaEvolve and formal proof systems place meaning in objects accepted by evaluators\. Self\-driving laboratories place meaning in experimentally observed interventions\. Typed artifact frameworks place meaning in provenance\-preserving transitions between regimes\. Evidential semantics asks a different but complementary question: which of these outputs licenses which scientific claim?

The shared failure mode is claim drift: the language, framing, or social interpretation of a result migrates toward a stronger claim while the evidence profile has not comparably strengthened\. LetDrift𝖣,V⁡\(Cliteral,Cimplied,E\)\>0\\operatorname\{Drift\}\_\{\\mathsf\{D\},V\}\(C\_\{\\mathrm\{literal\}\},C\_\{\\mathrm\{implied\}\},E\)\>0denote the case in which the implied claim is stronger than the literal or licensed claim and is not contained inℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\. The drift takes different forms in different routes:

Table 5:Claim\-drift modes across AI science routesWe can therefore distinguish four epistemic stages\. They are not all forms of cognitive understanding in the same sense; rather, they are increasing levels of warrant for scientific assertion:

Wtext​\(h\)\\displaystyle W\_\{\\mathrm\{text\}\}\(h\)=Pϕ​\(continuation∣h,𝒦\),\\displaystyle=P\_\{\\phi\}\(\\mathrm\{continuation\}\\mid h,\\mathcal\{K\}\),\(51\)Wconseq​\(h\)\\displaystyle W\_\{\\mathrm\{conseq\}\}\(h\)=𝟏​\{Conseq𝖣⁡\(h\)​is specified and testable in​𝖣\},\\displaystyle=\\mathbf\{1\}\\\{\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\)\\ \\text\{is specified and testable in \}\\mathsf\{D\}\\\},\(52\)Wadjudicated​\(h,E,V\)\\displaystyle W\_\{\\mathrm\{adjudicated\}\}\(h,E,V\)=𝟏​\{derived consequences of​h​survive​V​in​E\},\\displaystyle=\\mathbf\{1\}\\\{\\text\{derived consequences of \}h\\text\{ survive \}V\\text\{ in \}E\\\},\(53\)Wlicense​\(C,E,V\)\\displaystyle W\_\{\\mathrm\{license\}\}\(C,E,V\)=𝟏​\{C∈ℒ𝖣,V​\(E\)\}\.\\displaystyle=\\mathbf\{1\}\\\{C\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\\\}\.\(54\)HereWlicenseW\_\{\\mathrm\{license\}\}is deliberately written as a membership test: a claim counts as evidence\-licensed only if it belongs to the set of claims that the evidence permits\. The consequence and adjudication warrants are also written schematically because their concrete forms are domain\-specific: predictive warrant and causal warrant are important special cases in experimental and interventional sciences\. This warrant notation also avoids confusing these stages with the belief\-update operatorUUintroduced later\. These objects are not all scalars of the same type\. The following relation is therefore a qualitative hierarchy of warrant, not a numerical inequality:

Wtext≺warrantWconseq≺warrantWadjudicated≺warrantWlicense\.W\_\{\\mathrm\{text\}\}\\prec\_\{\\mathrm\{warrant\}\}W\_\{\\mathrm\{conseq\}\}\\prec\_\{\\mathrm\{warrant\}\}W\_\{\\mathrm\{adjudicated\}\}\\prec\_\{\\mathrm\{warrant\}\}W\_\{\\mathrm\{license\}\}\.This does not make linguistic plausibility useless\. It means that its epistemic warrant is weaker\. A scientific system must move textual hypotheses into consequence\-bearing, externally adjudicated, and claim\-calibrated structure\.

## 8Adjudication and Claim Calibration

The most important comparison is not model size but feedback quality\. LetVVbe the evaluator\. HereVVmeans the system, test, proof checker, benchmark, experiment, or review process that judges a candidate; it should not be confused with𝒱​\(h\)\\mathcal\{V\}\(h\), the hypothesis\-specific verification rule in the semantic tuple\. Then:

EpistemicStrength​\(V\)∝Independence​\(V\)×Reliability​\(V\)×WorldGrounding​\(V\)\.\\mathrm\{EpistemicStrength\}\(V\)\\propto\\mathrm\{Independence\}\(V\)\\times\\mathrm\{Reliability\}\(V\)\\times\\mathrm\{WorldGrounding\}\(V\)\.\(55\)The weakness of LLM self\-evaluation is thatVLLMV\_\{\\mathrm\{LLM\}\}is highly correlated withGLLMG\_\{\\mathrm\{LLM\}\}\. Its independence is low\. Proof checkers, program tests, experiments, and statistical tests are more independent and therefore provide stronger epistemic warrant\. But strong evaluation is still not enough\. A validation result becomes scientifically meaningful only after it is converted into a calibrated claim:

ClaimQuality​\(C,E,V\)∝EpistemicStrength​\(V\)⋅exp⁡\[−λ​Debt𝖣,Vℓ⁡\(C,E\)\]\.\\mathrm\{ClaimQuality\}\(C,E,V\)\\propto\\mathrm\{EpistemicStrength\}\(V\)\\cdot\\exp\[\-\\lambda\\,\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C,E\)\]\.\(56\)Thus, the final epistemic bottleneck is not only evaluator strength, but claim calibration after evaluation\. For a sound calibrated outputC∗∈Cal𝖣⁡\(C,E,V\)C^\{\\ast\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\), membership inℒ𝖣,V​\(E\)\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)is already guaranteed by definition; the quality penalty is therefore applied to the raw or intended claim before calibration\.

Table 6:Evaluator strength across AI science routesReliability does not increase monotonically with automation\. End\-to\-end automation can optimize the wrong proxy\. Formal systems may be less general but more trustworthy because their evaluators are strict\. Self\-driving laboratories touch the world but require heavy engineering, statistical, and replication governance\. Across all routes, the final claim should be set by the pair\(V,Cal𝖣\)\(V,\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\): the evaluator determines what survived, and the calibration operator determines what may be asserted\.

## 9Toward an AI\-Augmented Calibrated Scientific Loop

All AI science systems can be decomposed into five operators\. These operators answer five different questions: what to propose, what follows from it, whether it survives a test, how beliefs change, and what may finally be claimed\.

Table 7:Five operators in the calibrated AI science loopIn formula form, the same five\-operator loop is:

G\\displaystyle G:\(𝒦,Dt,q\)→H,\\displaystyle:\(\\mathcal\{K\},D\_\{t\},q\)\\rightarrow H,\(57\)M𝖣\\displaystyle M\_\{\\mathsf\{D\}\}:\(h,a,x\)→χh​\(a,x\)∈𝒬𝖣,\\displaystyle:\(h,a,x\)\\rightarrow\\chi\_\{h\}\(a,x\)\\in\\mathcal\{Q\}\_\{\\mathsf\{D\}\},\(58\)V\\displaystyle V:\(h,a,y\)→\{0,1\}​or​ℝ,\\displaystyle:\(h,a,y\)\\rightarrow\\\{0,1\\\}\\ \\text\{or\}\\ \\mathbb\{R\},\(59\)U\\displaystyle U:\(Bt,at,yt\)→Bt\+1,\\displaystyle:\(B\_\{t\},a\_\{t\},y\_\{t\}\)\\rightarrow B\_\{t\+1\},\(60\)Cal𝖣\\displaystyle\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}:\(C,E,V\)→Max⪯𝖣\(↓C∩ℒ𝖣,V​\(E\)\)\.\\displaystyle:\(C,E,V\)\\rightarrow\\operatorname\{Max\}\_\{\\preceq\_\{\\mathsf\{D\}\}\}\(\\downarrow C\\cap\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\)\.\(61\)In interventional or probabilistic settings, the local consequence may take the familiar formχh​\(a,x\)=P​\(Y∣h,do⁡\(a\),x\)\\chi\_\{h\}\(a,x\)=P\(Y\\mid h,\\operatorname\{do\}\(a\),x\)\. In the general case,MMis not restricted to future\-oriented numerical prediction\. It may derive retrodictions, diagnostic features, proof obligations, classification constraints, measurement signatures, or expected traces\.GGgenerates hypotheses\.MMturns hypotheses into observable, formal, or measurement\-level consequences\.VVadjudicates whether a hypothesis survives\.UUupdates beliefs\.Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}calibrates the final claim to the evidence\. The ideal system is therefore:

AIScientificExploration=\(G\+M\+V\+U\)\+Cal𝖣\.\\boxed\{\\mathrm\{AI\\ Scientific\\ Exploration\}=\(G\+M\+V\+U\)\+\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\.\}\(62\)Equivalently, the first four operators form the evidence\-producing loop, whileCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}manages assertion rights:

AIScientificExploration=G\+M\+V\+U⏟evidence\-producing loop\+Cal𝖣⏟assertion\-right management\.\\boxed\{\\mathrm\{AI\\ Scientific\\ Exploration\}=\\underbrace\{G\+M\+V\+U\}\_\{\\text\{evidence\-producing loop\}\}\+\\underbrace\{\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\}\_\{\\text\{assertion\-right management\}\}\.\}\(63\)The mature form of AI science may therefore not be a machine that autonomously writes papers, but a system that manages assertion rights: it proposes hypotheses, exposes them to adjudication, records evidence, and outputs only those claims for which it has acquired warrant\.

This also defines a route\-level failure pressure\. Let:

Shearr⁡\(t\)=Grnorm​\(t\)\+Mrnorm​\(t\)Vrnorm​\(t\)\+Qr,𝖣cal,norm​\(t\)\+ϵ\.\\operatorname\{Shear\}\_\{r\}\(t\)=\\frac\{G^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\+M^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\}\{V^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\+Q^\{\\mathrm\{cal,norm\}\}\_\{r,\\mathsf\{D\}\}\(t\)\+\\epsilon\}\.\(64\)WhenShearr⁡\(t\)≫1\\operatorname\{Shear\}\_\{r\}\(t\)\\gg 1, the route can generate or predict more candidates than it can adjudicate and calibrate\. The expected failure modes are hypothesis inflation, artifact inflation, overclaiming, and trust collapse\. Foundation\-model progress can improveGGandMMfaster than it improvesVVandCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\. Stronger models can therefore increase epistemic risk unless adjudication and calibration scale with them\.

The ratio is not proposed as a directly measured universal quantity\. It is a diagnostic abstraction for simulation and system design: all four capacities are normalized indices on a common synthetic scale, andϵ\>0\\epsilon\>0prevents division by zero\. The diagnostic asks whether the production side of a route is scaling faster than the adjudication and assertion\-right side\.

A more appropriate discovery objective is:

W​\(h,Ch,Eh\)\\displaystyle W\(h,C\_\{h\},E\_\{h\}\)=α​log⁡\(P​\(Dtest∣h\)P​\(Dtest∣h0\)\)\\displaystyle=\\alpha\\log\\left\(\\frac\{P\(D\_\{\\mathrm\{test\}\}\\mid h\)\}\{P\(D\_\{\\mathrm\{test\}\}\\mid h\_\{0\}\)\}\\right\)\(65\)\+β​I​\(ω;Y𝒯​\(h\)\)\+γ​CausalScope​\(h\)\\displaystyle\\quad\+\\beta I\(\\omega;Y\_\{\\mathcal\{T\}\(h\)\}\)\+\\gamma\\mathrm\{CausalScope\}\(h\)\+δ​Novelty​\(h\)−λ​Complexity​\(h\)−μ​Cost​\(h\)\\displaystyle\\quad\+\\delta\\mathrm\{Novelty\}\(h\)\-\\lambda\\mathrm\{Complexity\}\(h\)\-\\mu\\mathrm\{Cost\}\(h\)−ν​Debt𝖣,Vℓ⁡\(Ch,Eh\)\.\\displaystyle\\quad\-\\nu\\operatorname\{Debt\}^\{\\ell\}\_\{\\mathsf\{D\},V\}\(C\_\{h\},E\_\{h\}\)\.HereChC\_\{h\}is the claim the system intends to make about hypothesishh, andEhE\_\{h\}is the evidence collected for that hypothesis\. This is a better target than simply maximizing paper score or LLM plausibility because it penalizes claims that exceed evidence\.

In practice, a robust pipeline should:

1. 1\.generate hypotheses with an LLM or co\-scientist system,ht∼PLLM​\(h∣𝒦,Dt,q\)h\_\{t\}\\sim P\_\{\\mathrm\{LLM\}\}\(h\\mid\\mathcal\{K\},D\_\{t\},q\);
2. 2\.derive domain\-appropriate testable consequences using a model, simulator, theory, measurement framework, or formal system: χh​\(a,x\)∈𝒬𝖣;\\chi\_\{h\}\(a,x\)\\in\\mathcal\{Q\}\_\{\\mathsf\{D\}\};In interventional experimental settings, this may specialize toχh​\(a,x\)=Ph​\(Y∣do⁡\(a\),x\)\\chi\_\{h\}\(a,x\)=P\_\{h\}\(Y\\mid\\operatorname\{do\}\(a\),x\);
3. 3\.select the highest information\-per\-cost test: at=arg​maxa⁡I​\(h;Ya∣Dt\)/Cost⁡\(a\);a\_\{t\}=\\operatorname\*\{arg\\,max\}\_\{a\}I\(h;Y\_\{a\}\\mid D\_\{t\}\)/\\operatorname\{Cost\}\(a\);
4. 4\.adjudicate with experiment, simulation, proof checker, benchmark, or statistical test: V​\(h,at,yt\);V\(h,a\_\{t\},y\_\{t\}\);
5. 5\.update beliefs: P​\(h∣Dt\+1\)∝P​\(yt∣h,at,xt,Dt\)​P​\(h∣Dt\)\.P\(h\\mid D\_\{t\+1\}\)\\propto P\(y\_\{t\}\\mid h,a\_\{t\},x\_\{t\},D\_\{t\}\)P\(h\\mid D\_\{t\}\)\.
6. 6\.output the strongest evidence\-licensed claim or maximal licensed claim frontier: C∗∈Cal𝖣⁡\(C,E,V\)\.C^\{\\ast\}\\in\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\(C,E,V\)\.

Without the sixth step, an AI science loop can still produce overstrong language\. A system may generate a hypothesis, derive a testable consequence, run an experiment or proof check, and update a belief while still reporting a claim that exceeds the evidence\. The goal is therefore not to replace science with a single autonomous AI scientist\. It is to build an AI\-augmented calibrated scientific loop: domain models derive consequences, LLM agents expand the hypothesis space, automated platforms execute, formal and experimental systems adjudicate,Cal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}calibrates the claim, and human scientists define the problem, boundaries, and final interpretation\.

## 10Illustrative Synthetic Dynamics

AISim\-Cal is introduced here as a computational thought experiment and synthetic sensitivity model for the calibration semantics developed above\. Its purpose is not to forecast the future productivity of AI science, to rank institutions, or to predict which route will dominate a field\. The simulation is included to illustrate the internal behavior of the proposed concepts, not to validate the framework or the superiority of any route\. Rather, it turns the paper’s conceptual variables into a stylized dynamical system so that assumptions about generation, evaluation, calibration, cost, and claim strength can be inspected explicitly\. In this sense, the simulation is closer to an epistemological model of possible mechanism relations than to an empirical measurement program\(Winsberg,,[2010](https://arxiv.org/html/2606.31273#bib.bib34)\)\. All numerical quantities in this section are dimensionless synthetic quantities under a chosen parameterization, and the reported values should be read as illustrative methodological diagnostics rather than route performance measurements\.

The model represents each route as a policy\-like composition of the same operators used in the calibrated loop\. A route has generation capacityGr​\(t\)G\_\{r\}\(t\), consequence\-modeling capacityMr​\(t\)M\_\{r\}\(t\), evaluator strengthVr​\(t\)V\_\{r\}\(t\), belief\-update capacityUr​\(t\)U\_\{r\}\(t\), and route\-level calibration qualityQr,𝖣cal​\(t\)Q^\{\\mathrm\{cal\}\}\_\{r,\\mathsf\{D\}\}\(t\)\. The symbolQr,𝖣cal​\(t\)Q^\{\\mathrm\{cal\}\}\_\{r,\\mathsf\{D\}\}\(t\)is a scalar simulation parameter; it is kept distinct from the set\-valued calibration operatorCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}\. The simulated state variables include route capability, evaluator strength, calibration quality, cost burden, update\-loop maturity, goal clarity, true licensed claim production, final claim level, licensed claim level, and overclaim burden\. The seven route families are specialized foundation models, human\-led LLM assistants, multi\-agent co\-scientist systems, end\-to\-end AI Scientist pipelines, algorithmic and mathematical discovery agents, self\-driving laboratories, and a*Hybrid\-Cal*loop\. The last item is not treated as an empirically established route already proven superior\. It is a design hypothesis: a deliberately calibrated architecture that combines broad hypothesis generation with stronger external evaluators, explicit update loops, and conservative claim calibration\. The default run uses 128 Monte Carlo draws, random seed 4317, and a 2026–2035 synthetic horizon\.

Table 8:AISim\-Cal default route parameters used for the illustrative base settingThe table reports dimensionless design parameters, not measurements\. The route ordering in later figures is therefore a consequence of this illustrative parameterization and must not be interpreted as an empirical ranking of real systems\.

AISim\-Cal also includes a domain\-level goal\-clarity parameterΓ𝖣∈\[0,1\]\\Gamma\_\{\\mathsf\{D\}\}\\in\[0,1\]\. This parameter records whether the target of inquiry is sufficiently specified to support a meaningful comparison of routes\. HighΓ𝖣\\Gamma\_\{\\mathsf\{D\}\}corresponds to settings in which a domain has a clear objective, scoring rule, admissible evidence type, and claim\-value target; examples include formal proof, benchmarked algorithmic improvement, or a well\-defined screening objective\. LowΓ𝖣\\Gamma\_\{\\mathsf\{D\}\}corresponds to open\-ended problem formulation, where the central task is still to decide what should count as a good question, a meaningful consequence relation, a valid evaluator, or a licensed claim\. AISim\-Cal uses a thresholdτΓ=0\.45\\tau\_\{\\Gamma\}=0\.45: whenΓ𝖣<τΓ\\Gamma\_\{\\mathsf\{D\}\}<\\tau\_\{\\Gamma\}, the model still reports diagnostic scores but does not license a route ordering\. It marks the corresponding row as*no licensed ordering*\. The threshold is illustrative rather than empirically estimated, and the companion simulation exposes it as a configurable parameter\. In that regime, route comparison moves upstream to problem definition and evaluator construction rather than downstream to route selection\.

In addition to route\-level parameters and goal clarity, AISim\-Cal applies a route\-domain applicability factorAr,𝖣∈\[0,1\]A\_\{r,\\mathsf\{D\}\}\\in\[0,1\]\. A route with strong evaluators is not automatically applicable to every scientific domain\. Formal agents can dominate evaluator\-rich formal domains, but they should not be interpreted as generally superior in wet\-lab biology, chemical synthesis, or materials characterization unless their operators are domain\-compatible\. The reported utility is therefore computed as

LicensedUtilityr,𝖣=Ar,𝖣⋅RawLicensedUtilityr,𝖣,\\mathrm\{LicensedUtility\}\_\{r,\\mathsf\{D\}\}=A\_\{r,\\mathsf\{D\}\}\\cdot\\mathrm\{RawLicensedUtility\}\_\{r,\\mathsf\{D\}\},\(66\)whereRawLicensedUtilityr,𝖣\\mathrm\{RawLicensedUtility\}\_\{r,\\mathsf\{D\}\}is the synthetic utility before route\-domain coverage is applied\. This additional factor makes the heatmap a conditional route comparison rather than a claim that evaluator strictness alone determines scientific usefulness\.

Table 9:AISim\-Cal route\-domain applicability matrixAr,𝖣A\_\{r,\\mathsf\{D\}\}used in the illustrative runAISim\-Cal decomposes evaluator strength into three multiplicative components:

EvaluatorStrengthr​\(t\)=\[Independencer​\(t\)×Reliabilityr​\(t\)×Groundingr​\(t\)\]1/3\.\\mathrm\{EvaluatorStrength\}\_\{r\}\(t\)=\\left\[\\mathrm\{Independence\}\_\{r\}\(t\)\\times\\mathrm\{Reliability\}\_\{r\}\(t\)\\times\\mathrm\{Grounding\}\_\{r\}\(t\)\\right\]^\{1/3\}\.\(67\)Independence records how far the evaluator is separated from the generator\. Reliability records whether repeated evaluation would give stable and auditable outcomes\. Grounding records whether the evaluator is connected to a domain\-relevant constraint, such as an experiment, proof checker, benchmark, simulator, statistical test, or structured external review\. The geometric mean makes the evaluator bottleneck visible: a route cannot obtain high evaluator strength by excelling in only one component while the others remain weak\. This decomposition preserves a central claim of the manuscript: a fluent internal critic can increase review\-like structure without supplying the same epistemic warrant as an evaluator that is independent, reliable, and grounded in the target domain\.

The simulated claim ladder follows the manuscript’s L0–L6 scale\. L0 denotes hypothesis generation\. L1 denotes computational or model support\. L2 denotes internal experimental support\. L3 denotes external replication or independent support\. L4 denotes causal mechanism support\. L5 denotes generalizable scientific knowledge within stated boundaries\. L6 denotes translational, application\-ready, or otherwise high\-stakes assertion\. AISim\-Cal then applies three calibration modes\. In the simulation implementation, evaluator and evidence\-side variables determine the licensed claim level, while calibration quality controls how strongly the final asserted claim is pulled back toward that licensed frontier\. In the no\-calibration mode, the final claim can exceed the licensed claim level\. In the imperfect\-calibration mode, claims are partially pulled back toward the licensed frontier but may still overshoot when evaluator strength or calibration quality is weak\. In the perfect\-calibration mode, the final claim is forced not to exceed the licensed claim level\. Perfect calibration is included as a limiting reference condition, not as a realistic assumption about current AI science systems\.

Four diagnostics summarize each run\. The synthetic licensed\-utility diagnostic measures useful claim production after the claim has been restricted to the licensed frontier and adjusted by route\-domain applicabilityAr,𝖣A\_\{r,\\mathsf\{D\}\}\. The overclaim gap measures how far the final asserted level exceeds the licensed level\. False\-discovery burden combines unsupported claim production with overclaiming pressure\. The evidence bottleneck index measures the mismatch between generation/modeling capacity and the combined strength of evaluation, updating, and calibration\. These metrics are intentionally formal rather than empirical\. They make visible the paper’s claim that a route can look productive when judged by raw outputs while being weak when judged by evidence\-licensed assertion\.

The score\-landscape diagnostic has two additional preconditions\. First, a route comparison can be interpreted as an ordering only relative to a specified objective and scoring rule\. If the domain is still asking what the right objective should be, then a numerical score is useful as an exploratory diagnostic but should not be promoted into a dominance statement\. Second, the route must have sufficient domain applicabilityAr,𝖣A\_\{r,\\mathsf\{D\}\}for the comparison to be meaningful\. For this reason, the score landscape still reports utility values for low\-Γ𝖣\\Gamma\_\{\\mathsf\{D\}\}settings, but it marks them as unordered rather than assigning route dominance under an arbitrary scalarization\. It also treats strict formal evaluators as strong primarily in domains where the route itself is applicable\. The full route–domain–scenario landscape is provided in Appendix[A](https://arxiv.org/html/2606.31273#A1); the main text emphasizes the route utility and calibration\-ablation diagnostics\.

![Refer to caption](https://arxiv.org/html/2606.31273v1/figures/route_licensed_utility.png)Figure 2:Synthetic licensed\-utility diagnostic across AISim\-Cal routes\. All quantities are dimensionless synthetic quantities and not empirical forecasts\. Lines show means over the 128 Monte Carlo draws in the illustrative parameterization; propagated standard\-error bands are small in this run and are provided in the companion outputs\. The curves include goal clarity and route\-domain applicability factors; they are conditional diagnostics of the calibration framework, not predictions of future scientific output\. The ordering of routes is not a claim about real\-world dominance; it is a consequence of the illustrative parameter setting\.Under this parameterization, the route\-level outputs illustrate the difference between apparent productivity and a synthetic licensed\-utility diagnostic\. Routes with strict evaluators, especially the algorithmic and mathematical discovery route, obtain high licensed utility where the synthetic evaluator is comparatively independent, reliable, cheap to apply, and domain\-compatible\. The Hybrid\-Cal loop also ranks highly in several scenarios because its parameters encode strong calibration quality, evaluator integration, and broader route\-domain coverage\. This is a conditional consequence of the design assumptions, not evidence that such a route has already been demonstrated\. No inference should be drawn about real\-world superiority of Hybrid\-Cal or any route\. Conversely, routes with high throughput but weaker calibration can accumulate larger overclaim burdens, especially when automatic artifact production or benchmark optimization is not matched by independent adjudication and claim weakening\.

![Refer to caption](https://arxiv.org/html/2606.31273v1/figures/calibration_ablation.png)Figure 3:Calibration ablation in AISim\-Cal\. All quantities are dimensionless synthetic quantities and not empirical forecasts\. The no\-calibration, imperfect\-calibration, and perfect\-calibration conditions show how the synthetic overclaim gap changes when the same route dynamics are passed through different claim\-calibration regimes\.The calibration ablation is the most direct synthetic illustration of the manuscript’s core semantics\. In the generated outputs, the mean final\-over\-licensed claim\-level gap is largest when calibration is absent, smaller under imperfect calibration, and zero under the perfect\-calibration reference condition\. This pattern is stylized and conditional, but it expresses the formal point ofCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}: external validation alone does not determine the permissible claim level\. A system must still translate evidence into a licensed claim, and a weaker calibration operator leaves residual epistemic debt even when some validation exists\.

The companion source repository contains CSV summaries, result notes, and generated figures for the AISim\-Cal illustration\. These artifacts are included to make the synthetic parameterization and figure generation auditable; they are not new empirical experimental data\.

## 11AI\-System Design Checklist

The framework above can be translated into a concrete design checklist for evaluating AI scientific exploration systems\. The checklist does not require every system to contain all modules at the same level of maturity\. A language\-model assistant, a theorem\-proving system, and a robotic laboratory will have very different engineering constraints\. The purpose is instead to make the epistemic contract explicit\. A valid claim of AI scientific exploration or discovery should specify the hypothesis object, consequence object, evaluator, update rule, and calibrated claim\.

Table 10:Checklist for evaluating AI scientific exploration systemsThis checklist also clarifies the evidential structure of AI science claims\. A system can be impressive atGGbut weak atVV: it may generate creative hypotheses but lack independent validation\. A system can be strong atVVbut weak atCal𝖣\\operatorname\{Cal\}\_\{\\mathsf\{D\}\}: it may obtain a valid benchmark result yet describe it with language that implies a broader scientific discovery\. A system can be strong atMMbut weak atUU: it may derive accurate consequence profiles without updating beliefs about mechanisms\. These distinctions are operational rather than merely terminological\. They are the difference between a useful AI research tool, an auditable scientific discovery engine, and a paper\-writing system that produces plausible but under\-validated claims\.

The checklist is intentionally compatible with different AI research cultures\. In machine learning,VVmay be a benchmark suite, ablation protocol, or held\-out dataset\. In formal mathematics,VVmay be a proof assistant or proof checker\. In materials science,VVmay be a density\-functional\-theory calculation, phase\-stability criterion, synthesis attempt, or structural characterization\. In biology,VVmay be an in vitro assay, perturbation experiment, animal model, independent cohort, or clinical study\. The notation is shared, but the evidential threshold is domain\-specific\. This is why the calibration operator depends not only onCCandEE, but also onVVand𝖣\\mathsf\{D\}: the same evidence can license different claims in different domains\.

For manuscript reporting, the framework suggests a simple discipline\. A manuscript should state the raw AI output, the testable consequence derived from that output, the external or semi\-external evaluator, the evidence actually obtained, the claim level on the claim ladder, and the reason stronger claims are not yet licensed\. This discipline does not weaken the contribution\. It makes evidential strengths and remaining validation needs explicit\.

## 12Practical Implications

For data\-rich experimental domains, including life science, chemistry, materials science, and AI for biology, a stable configuration is to use specialized models to generate candidates, LLM or multi\-agent systems for literature integration and mechanistic hypotheses, and then strict verification through database back\-testing, negative controls, statistical tests, wet\-lab experiments, or independent cohorts\. The final claim should be reported at the correct level of the claim ladder: computational candidate, in vitro support, independent replication, causal mechanism, or translational claim\. AI may generate proposals at scale, but scientists should retain control over problem definition, evaluation criteria, and final claim boundaries\.

For AI and machine learning research itself, stronger automation is more defensible because experiments are cheap and feedback is fast\. AI Scientist systems, code agents, automatic baseline search, and benchmark evaluation can be useful\. The danger is that automatic tuning plus automatic writing may be mistaken for theory\. Baselines, statistical robustness, ablations, data leakage, and novelty require explicit audit, and benchmark improvements should be calibrated against the claim they are used to support\.

For mathematics, algorithms, and theoretical computer science, a particularly auditable configuration is to use LLMs for conjectures or proof sketches and connect them to Lean, Coq, Isabelle, SAT/SMT solvers, program verifiers, or strong benchmarks\. The reason is not that the route is universally superior, but that candidate objects can be judged by strict external evaluators and then reported within explicit proof, program, or benchmark boundaries\.

For chemistry, materials, and robotic laboratories, the long\-term direction is the self\-driving experimental loop\. Short\-term claims must not underestimate instruments, protocols, samples, data standards, safety, and replication governance\. The A\-Lab correction is a warning: even when AI enters the physical world, humans must still audit experimental definitions, phase identification, database deduplication, and novelty claims\. World contact strengthens evidence, but it does not by itself determine whether the licensed claim is synthesis success, new phase discovery, generalizable materials rule, or application\-ready material\.

## 13Conclusion

AI scientific exploration is uneven across evaluator regimes\. Specialized predictive models illustrate high consequence\-derivation capacity without full autonomy\. Research assistants illustrate broad hypothesis and artifact support under human responsibility\. Multi\-agent systems illustrate structured hypothesis search\. End\-to\-end AI Scientist systems illustrate workflow automation under constrained computational settings\. Self\-driving laboratories illustrate world\-grounded loops with high engineering and validation costs\. Mathematics and algorithms illustrate the special epistemic value of strict evaluators\. These are conditional contrasts about evidence and claim licensing, not forecasts about which route will dominate\. Across all routes, the central issue is not only whether the system generates hypotheses or obtains validation, but whether it converts validation into an appropriately bounded scientific claim\.

The framework can be summarized by three calibration laws\. First, no claim without license:

C​is publishable as a scientific assertion only if​C∈ℒ𝖣,V​\(E\)\.C\\ \\text\{is publishable as a scientific assertion only if\}\\ C\\in\\mathcal\{L\}\_\{\\mathsf\{D\},V\}\(E\)\.\(68\)Second, validation does not determine claim level:

E⊩𝖣,VC1,C1⪯𝖣C2⇏E⊩𝖣,VC2\.E\\Vdash\_\{\\mathsf\{D\},V\}C\_\{1\},\\quad C\_\{1\}\\preceq\_\{\\mathsf\{D\}\}C\_\{2\}\\quad\\not\\Rightarrow\\quad E\\Vdash\_\{\\mathsf\{D\},V\}C\_\{2\}\.\(69\)The same evidence may license a weak claim without licensing a stronger claim\. Third, automation amplifies the need for calibration:

Shearr\(t\)=Grnorm​\(t\)\+Mrnorm​\(t\)Vrnorm​\(t\)\+Qr,𝖣cal,norm​\(t\)\+ϵ↑⇒Riskoverclaim↑\.\\operatorname\{Shear\}\_\{r\}\(t\)=\\frac\{G^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\+M^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\}\{V^\{\\mathrm\{norm\}\}\_\{r\}\(t\)\+Q^\{\\mathrm\{cal,norm\}\}\_\{r,\\mathsf\{D\}\}\(t\)\+\\epsilon\}\\uparrow\\quad\\Rightarrow\\quad\\mathrm\{Risk\}\_\{\\mathrm\{overclaim\}\}\\uparrow\.\(70\)When generation and consequence\-derivation capacities scale faster than adjudication and calibration, AI systems may become more productive while also accumulating more epistemic debt\.

Overclaiming is epistemic debt: it must be paid by additional evidence or cancelled by weakening the claim, narrowing its boundary, changing its modal status, or revising the hypothesis object\.

The compressed thesis is:

LLMs expand hypotheses, consequence models give structure,experiments or proof checkers give adjudication,and evidence calibration gives licensed claims\.\\boxed\{\\begin\{gathered\}\\text\{LLMs expand hypotheses, consequence models give structure,\}\\\\ \\text\{experiments or proof checkers give adjudication,\}\\\\ \\text\{and evidence calibration gives licensed claims\.\}\\end\{gathered\}\}More strictly:

scientific semantics=linguistic hypothesis\+testable consequence\+external validation rule\+claim boundary\.\\boxed\{\\begin\{gathered\}\\text\{scientific semantics\}=\\text\{linguistic hypothesis\}\\\\ \+\\text\{testable consequence\}\+\\text\{external validation rule\}\\\\ \+\\text\{claim boundary\.\}\\end\{gathered\}\}The main bottleneck in AI science is not the ability to generate hypotheses\. It is not only the transition fromPlanguage​\(h\)P\_\{\\mathrm\{language\}\}\(h\)toConseq𝖣⁡\(h\)\\operatorname\{Conseq\}\_\{\\mathsf\{D\}\}\(h\), withPworld​\(Y∣do⁡\(a\),h\)P\_\{\\mathrm\{world\}\}\(Y\\mid\\operatorname\{do\}\(a\),h\)as one interventional special case\. It is also the transition from validation result to licensed assertion\. An AI system may generate a plausible hypothesis, derive a testable consequence, and even obtain experimental support\. It must still answer: what strength of claim does this evidence permit, and what epistemic debt remains? On this framework, a reliable AI scientist would not merely write papers or run experiments\. It would manage scientific assertion rights by transforming candidate hypotheses into evidence\-licensed, boundary\-explicit, and strength\-calibrated scientific claims\.

The practical purpose of the framework is comparative\. It does not claim that AI science has already converged on one architecture, nor that any current system has solved autonomous discovery in full\. It proposes a language for comparing systems that are otherwise difficult to compare: LLM assistants, co\-scientists, AI Scientist pipelines, formal proof agents, specialized scientific models, and self\-driving laboratories\. The comparison becomes sharper when every system is asked the same questions: what does it generate, what consequence does it derive, who evaluates it, how is belief updated, and what final claim is licensed by the evidence?

The same discipline also explains why minimal reconstruction is difficult and valuable\. Claim calibration is not always a move toward narrower language\. Sometimes it is a move toward a broader structural assertion that reconstructs a licensed common pattern from many heterogeneous local results\. Such a claim is scientifically powerful precisely because it is rare: it must preserve the local evidence, expose the shared structure, and leave the uncompressed residuals visible rather than converting them into unsupported universality\.

The practical unit of AI\-assisted scientific reporting is therefore not a claim alone, but a licensed claim paired with an explicit non\-claim\.

## Limitations

This paper is a theoretical synthesis and methodological comparison\. It does not introduce new empirical experimental data\. The AISim\-Cal section adds a synthetic simulation and generated figure/CSV artifacts for sensitivity\-style illustration of the calibration framework; those outputs are not empirical forecasts and should not be interpreted as measurements of future scientific productivity\. Several 2025–2026 systems are still changing quickly, and their results should be interpreted within the boundaries stated by the original papers or official reports\. This is especially important for GPT\-5 science case studies, AI Scientist workshop review results, AlphaEvolve infrastructure claims, and the A\-Lab materials controversy\. The cases are used here as epistemological comparators, not as proof that any route has reached mature autonomous science\.

## Data Availability

No new empirical experimental data were generated\. The accompanying clean preprint source files contain the AISim\-Cal synthetic outputs, including CSV summaries and generated figures\. External source caches used for citation checking are not part of the clean preprint source release; external materials should be accessed through the public papers, official pages, or news reports cited in the reference list\.

## Code Availability

The AISim\-Cal synthetic simulation scripts, configuration files, generated CSV summaries, and generated figures are maintained with the clean preprint source files\. They are illustrative research artifacts for inspecting the calibration framework, not empirical forecasts of future scientific productivity\.

## Ethics Statement

This paper does not involve human participants, animal experiments, or sensitive personal data\.

## Competing Interests

The author declares no financial competing interests\. The analysis is based on public sources, cited literature, and the evidential boundaries explicitly stated in the manuscript\.

## Funding Statement

No dedicated funding was received for this paper\.

## Author Contributions

Hongmin Li developed the conceptual framework, argument structure, and claim\-calibration interpretation\. AI tools assisted with English drafting, structural organization, LaTeX preparation, reference formatting, and compilation checks\. The author is responsible for final accuracy, citation completeness, and academic integrity\.

## AI Use Statement

AI tools were used for drafting assistance, structural organization, LaTeX formatting, bibliography preparation, and compilation checks\. The author reviewed all generated text, verified the argument structure, and is responsible for all citations, mathematical definitions, interpretations, and final claims\.

## Appendix AAISim\-Cal Score Landscape

![[Uncaptioned image]](https://arxiv.org/html/2606.31273v1/x1.png)

Figure 4:AISim\-Cal route–domain–scenario score landscape\. Each panel corresponds to one scenario and calibration mode; each cell reports the synthetic licensed\-utility diagnostic for one route in one domain\. The values are dimensionless scores under the chosen synthetic parameterization, including goal clarity and route\-domain applicabilityAr,𝖣A\_\{r,\\mathsf\{D\}\}; they are not empirical forecasts\. Hatched rows show low\-goal\-clarity regimes: the scores remain visible as exploratory diagnostics, but they do not license a route ordering because the scientific task is still problem formulation, evaluator construction, and claim\-boundary definition\.

## References

- Abramson et al\., \(2024\)Abramson, J\., Adler, J\., Dunger, J\., Evans, R\., Green, T\., Pritzel, A\., Ronneberger, O\., Willmore, L\., Ballard, A\. J\., Bambrick, J\., Bodenstein, S\. W\., Evans, D\. A\., Hung, C\.\-C\., O’Neill, M\., Reiman, D\., Tunyasuvunakool, K\., Wu, Z\., Zidek, A\., et al\. \(2024\)\.Accurate structure prediction of biomolecular interactions with AlphaFold 3\.Nature, 630:493–500\.
- Avigad, \(2006\)Avigad, J\. \(2006\)\.Mathematical method and proof\.Synthese, 153:105–159\.
- Boiko et al\., \(2023\)Boiko, D\. A\., MacKnight, R\., Kline, B\., and Gomes, G\. \(2023\)\.Autonomous chemical research with large language models\.Nature, 624:570–578\.
- Bubeck et al\., \(2025\)Bubeck, S\., Coester, C\., Eldan, R\., Gowers, T\., Lee, Y\. T\., Lupsasca, A\., Sawhney, M\., Scherrer, R\., Sellke, M\., Spears, B\. K\., Unutmaz, D\., Weil, K\., Yin, S\., and Zhivotovskiy, N\. \(2025\)\.Early science acceleration experiments with GPT\-5\.arXiv preprint arXiv:2511\.16072\.
- Chawla, \(2026\)Chawla, D\. S\. \(2026\)\.Nature robot chemist paper corrected, but some questions remain unanswered\.Chemical & Engineering News\.Chemical & Engineering News report, accessed 2026\-06\-16\.
- Friedman, \(1974\)Friedman, M\. \(1974\)\.Explanation and scientific understanding\.The Journal of Philosophy, 71\(1\):5–19\.
- Ghareeb et al\., \(2026\)Ghareeb, A\. E\., Chang, B\., Mitchener, L\., et al\. \(2026\)\.A multi\-agent system for automating scientific discovery\.Nature\.Advance online publication, published 2026\-05\-19\.
- Google DeepMind and EMBL\-EBI, \(2026\)Google DeepMind and EMBL\-EBI \(2026\)\.AlphaFold protein structure database\.[https://alphafold\.ebi\.ac\.uk/](https://alphafold.ebi.ac.uk/)\.Accessed 2026\-06\-16\.
- Gottweis et al\., \(2026\)Gottweis, J\., Weng, W\.\-H\., Daryin, A\., et al\. \(2026\)\.Accelerating scientific discovery with Co\-Scientist\.Nature\.Advance online publication, published 2026\-05\-19\.
- Grünwald and Roos, \(2019\)Grünwald, P\. and Roos, T\. \(2019\)\.Minimum description length revisited\.International Journal of Mathematics for Industry, 11\(1\):1930001\.
- Guyatt et al\., \(2008\)Guyatt, G\. H\., Oxman, A\. D\., Vist, G\. E\., Kunz, R\., Falck\-Ytter, Y\., Alonso\-Coello, P\., and Schunemann, H\. J\. \(2008\)\.GRADE: An emerging consensus on rating quality of evidence and strength of recommendations\.BMJ, 336\(7650\):924–926\.
- Hacking, \(1983\)Hacking, I\. \(1983\)\.Representing and Intervening: Introductory Topics in the Philosophy of Natural Science\.Cambridge University Press, Cambridge\.
- Hubert et al\., \(2026\)Hubert, T\., Mehta, R\., Sartran, L\., et al\. \(2026\)\.Olympiad\-level formal mathematical reasoning with reinforcement learning\.Nature, 651:607–613\.
- King and Zenil, \(2023\)King, R\. D\. and Zenil, H\. \(2023\)\.A framework for evaluating the AI\-driven automation of science\.InArtificial Intelligence in Science: Challenges, Opportunities and the Future of Research\. OECD Publishing\.Accessed 2026\-06\-16\.
- Kitcher, \(1981\)Kitcher, P\. \(1981\)\.Explanatory unification\.Philosophy of Science, 48\(4\):507–531\.
- Lakatos, \(1970\)Lakatos, I\. \(1970\)\.Falsification and the methodology of scientific research programmes\.In Lakatos, I\. and Musgrave, A\., editors,Criticism and the Growth of Knowledge, pages 91–196\. Cambridge University Press, Cambridge\.
- Lu et al\., \(2026\)Lu, C\., Lu, C\., Lange, R\. T\., Yamada, Y\., Hu, S\., Foerster, J\., Ha, D\., and Clune, J\. \(2026\)\.Towards end\-to\-end automation of AI research\.Nature, 651:914–919\.
- Mayo, \(1996\)Mayo, D\. G\. \(1996\)\.Error and the Growth of Experimental Knowledge\.University of Chicago Press, Chicago\.
- Merchant et al\., \(2023\)Merchant, A\., Batzner, S\., Schoenholz, S\. S\., Aykol, M\., Cheon, G\., Cubuk, E\. D\., et al\. \(2023\)\.Scaling deep learning for materials discovery\.Nature, 624:80–85\.
- Novikov et al\., \(2025\)Novikov, A\., Vu, N\., Eisenberger, M\., Dupont, E\., Huang, P\.\-S\., Wagner, A\. Z\., Shirobokov, S\., Kozlovskii, B\., Ruiz, F\. J\. R\., Mehrabian, A\., Kumar, M\. P\., See, A\., Chaudhuri, S\., Holland, G\., Davies, A\., Nowozin, S\., Kohli, P\., and Balog, M\. \(2025\)\.AlphaEvolve: A coding agent for scientific and algorithmic discovery\.arXiv preprint arXiv:2506\.13131\.Technical white paper\.
- OECD, \(2023\)OECD \(2023\)\.Artificial intelligence in science: Challenges, opportunities and the future of research\.See the chapter “A framework for evaluating the AI\-driven automation of science”\.
- OpenAI, \(2025\)OpenAI \(2025\)\.Early experiments in accelerating science with GPT\-5\.[https://openai\.com/index/accelerating\-science\-gpt\-5/](https://openai.com/index/accelerating-science-gpt-5/)\.Official web report, accessed 2026\-06\-16\.
- Pearl, \(2009\)Pearl, J\. \(2009\)\.Causality: Models, Reasoning, and Inference\.Cambridge University Press, Cambridge, 2 edition\.
- Peirce, \(1878\)Peirce, C\. S\. \(1878\)\.How to make our ideas clear\.Popular Science Monthly, 12:286–302\.
- Popper, \(1959\)Popper, K\. R\. \(1959\)\.The Logic of Scientific Discovery\.Hutchinson, London\.
- Rissanen, \(1978\)Rissanen, J\. \(1978\)\.Modeling by shortest data description\.Automatica, 14\(5\):465–471\.
- Szymanski et al\., \(2023\)Szymanski, N\. J\., Rendy, B\., Fei, Y\., et al\. \(2023\)\.An autonomous laboratory for the accelerated synthesis of inorganic materials\.Nature, 624:86–91\.
- Szymanski et al\., \(2026\)Szymanski, N\. J\., Rendy, B\., Fei, Y\., Kumar, R\. E\., He, T\., Milsted, D\., McDermott, M\. J\., Gallant, M\., Cubuk, E\. D\., Merchant, A\., Kim, H\., Jain, A\., Bartel, C\. J\., Persson, K\., Zeng, Y\., and Ceder, G\. \(2026\)\.Author correction: An autonomous laboratory for the accelerated synthesis of inorganic materials\.Nature, 650:E1–E1\.
- Thurston, \(1994\)Thurston, W\. P\. \(1994\)\.On proof and progress in mathematics\.Bulletin of the American Mathematical Society, 30\(2\):161–178\.
- Tom et al\., \(2024\)Tom, G\., Schmid, S\. P\., Baird, S\. G\., Cao, Y\., Darvish, K\., Hao, H\., Lo, S\., Pablo\-Garcia, S\., Rajaonson, E\. M\., Skreta, M\., Yoshikawa, N\., Corapi, S\., Akkoc, G\. D\., Strieth\-Kalthoff, F\., Seifrid, M\., and Aspuru\-Guzik, A\. \(2024\)\.Self\-driving laboratories for chemistry and materials science\.Chemical Reviews, 124:9633–9732\.
- Toulmin, \(1958\)Toulmin, S\. E\. \(1958\)\.The Uses of Argument\.Cambridge University Press, Cambridge\.
- Varadi et al\., \(2024\)Varadi, M\., Bertoni, D\., Magana, P\., Paramval, U\., Pidruchna, I\., Radhakrishnan, M\., Tsenkov, M\., Nair, S\., Mirdita, M\., Yeo, J\., Kovalevskiy, O\., Tunyasuvunakool, K\., Laydon, A\., Zidek, A\., Green, T\., Jumper, J\., Birney, E\., Steinegger, M\., Hassabis, D\., and Velankar, S\. \(2024\)\.AlphaFold protein structure database in 2024: Providing structure coverage for over 214 million protein sequences\.Nucleic Acids Research, 52\(D1\):D368–D375\.
- Wang and Buehler, \(2026\)Wang, F\. Y\. and Buehler, M\. J\. \(2026\)\.Self\-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence\.arXiv preprint arXiv:2606\.01444\.
- Winsberg, \(2010\)Winsberg, E\. \(2010\)\.Science in the Age of Computer Simulation\.University of Chicago Press, Chicago\.

Similar Articles

Evidence-Ledger Adjudication for Claim-Evidence Traceability

arXiv cs.AI

This paper introduces evidence-ledger adjudication, a workflow for claim-evidence traceability in AI-assisted writing, evaluated on a blind benchmark from AVeriTeC, CLIMATE-FEVER, and SciFact, showing agent-based methods outperform baselines.

CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

arXiv cs.CL

The paper introduces CALIBER, a method for calibrating confidence in reasoning language models by eliciting confidence estimates both before and after reasoning, with supervision targets matched to the information state. It achieves significant reductions in Expected Calibration Error (up to 52.5%) and strong Brier scores and AUROC across multiple benchmarks.

CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models

arXiv cs.CL

The paper presents CalBrief, a pilot diagnostic benchmark of 16 evidence packages and 96 human-verified takeaways for evaluating whether large language models can generate evidence-calibrated scientific briefings. The study finds that structured organization improves reasoning but explicit strength-calibration policies are overly conservative, with most conservatism arising from expanded label spaces rather than signal injection.