NOVA: Fundamental Limits of Knowledge Discovery Through AI

arXiv cs.AI Papers

Summary

The NOVA framework models the 'generate, verify, accumulate, retrain' loop as an adaptive sampling process over a knowledge space, identifying failure modes and proving a scaling law for cumulative generation cost under Zipf-like discovery distributions.

arXiv:2605.15219v1 Announce Type: new Abstract: Can AI systems discover genuinely new knowledge through iterative self improvement, and if so, at what cost? We introduce the NOVA framework, which models the common ``generate, verify, accumulate, retrain'' loop as an adaptive sampling process over a knowledge space. We identify sufficient conditions under which accumulated genuine knowledge eventually covers a finite domain, and show how their violations produce distinct failure modes: contamination, forgetting, exploration failure, and acceptance failure. We then analyze imperfect verification and identify a contamination trap: as easy-to-find knowledge is exhausted, the model mass assigned to new valid artifacts shrinks, so even small false-positive rates can cause invalid artifacts to enter the knowledge base faster than genuine discoveries. We clarify that Good--Turing estimation is a local batch-diversity diagnostic, not an estimator of the historically undiscovered valid mass that governs long-term discovery. Under a separate tail-equivalence assumption relating the model's effective discovery distribution to a Zipf law with exponent $\alpha>1$, we prove that the cumulative generation cost required to obtain $D$ distinct genuine discoveries satisfies $R_{\mathrm{cum}}(D)=\Theta(c_{\mathrm{gen}}D^\alpha)$, where $c_{\mathrm{gen}}$ is the per-candidate generation cost. This scaling law quantifies asymptotic diminishing returns as the discovery frontier advances. Finally, we formalize human amplification through guidance, generation, and verification, explaining why expert input is most valuable near autonomous exploration barriers.
Original Article
View Cached Full Text

Cached at: 05/18/26, 06:31 AM

# NOVA: Fundamental Limits of Knowledge Discovery Through AI
Source: [https://arxiv.org/html/2605.15219](https://arxiv.org/html/2605.15219)
Salman Avestimehr University of Southern California avestime@usc\.edu &Ken Duffy Northeastern University k\.duffy@northeastern\.edu &Muriel Médard Massachusetts Institute of Technology medard@mit\.edu

###### Abstract

Can AI systems discover genuinely new knowledge through iterative self improvement, and if so, at what cost? We introduce the NOVA framework, which models the common “generate, verify, accumulate, retrain” loop as an adaptive sampling process over a knowledge space\. We identify sufficient conditions under which accumulated genuine knowledge eventually covers a finite domain, and show how their violations produce distinct failure modes: contamination, forgetting, exploration failure, and acceptance failure\. We then analyze imperfect verification and identify a contamination trap: as easy\-to\-find knowledge is exhausted, the model mass assigned to new valid artifacts shrinks, so even small false\-positive rates can cause invalid artifacts to enter the knowledge base faster than genuine discoveries\. We clarify that Good–Turing estimation is a local batch\-diversity diagnostic, not an estimator of the historically undiscovered valid mass that governs long\-term discovery\. Under a separate tail\-equivalence assumption relating the model’s effective discovery distribution to a Zipf law with exponentα\>1\\alpha\>1, we prove that the cumulative generation cost required to obtainDDdistinct genuine discoveries satisfiesRcum​\(D\)=Θ​\(cgen​Dα\)R\_\{\\mathrm\{cum\}\}\(D\)=\\Theta\(c\_\{\\mathrm\{gen\}\}D^\{\\alpha\}\), wherecgenc\_\{\\mathrm\{gen\}\}is the per\-candidate generation cost\. This scaling law quantifies asymptotic diminishing returns as the discovery frontier advances\. Finally, we formalize human amplification through guidance, generation, and verification, explaining why expert input is most valuable near autonomous exploration barriers\.

## 1Introduction

Can artificial intelligence discover genuinely new knowledge, or is it fundamentally limited to recombining what it has already seen? Recent AI systems have begun to answer this question empirically: AlphaProof achieved silver\-medal performance at the International Mathematical Olympiad by generating and verifying formal proofs\(Hubertet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib26)\), DeepSeek\-Prover\-V2 advanced formal mathematical reasoning through reinforcement learning and subgoal decomposition\(Renet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib28)\), and self\-training methods such as STaR have shown that models can bootstrap their own reasoning capabilities\(Zelikmanet al\.,[2022](https://arxiv.org/html/2605.15219#bib.bib23)\)\. These successes share a common architecture: a model generates candidate artifacts, a verifier filters for correctness, and the verified outputs are fed back to improve the model\. This “generate, verify, accumulate, retrain” loop is emerging as the dominant paradigm for AI\-driven knowledge discovery\.

Yet, despite these empirical advances, we lack a theoretical understanding of the fundamental limits of this process\. How fast can an AI system discover new knowledge, and at what cost? Does the rate of discovery inevitably slow down, and if so, how severely? Under what conditions does the process converge to the target knowledge domain, and when does it collapse? What role does verification quality play, and what happens when verification is imperfect? And perhaps most importantly: can AI systems bootstrap themselves to discover all knowledge autonomously, or is human guidance fundamentally necessary?

### 1\.1Overview and Contributions

This paper introduces the NOVA \(Navigating theOrigins andVerification ofAI Knowledge\) framework to study when AI systems can discover genuinely new knowledge through iterative self\-improvement\. NOVA models discovery as an adaptive sampling loop over a knowledge space: a model generates candidates, a verifier accepts or rejects them, accepted candidates are accumulated, and the model is retrained\. This view separates the core bottlenecks of AI\-driven discovery: reachability, verification, retention, and the thinning of the unknown frontier\. It also connects the problem to species estimation, occupancy laws, and support\-limited sampling, yielding convergence guarantees, failure modes, and cost laws\. Our main contributions are:

NOVA formalization and convergence\.We formalize AI\-driven discovery over an ambient candidate space containing both valid and invalid artifacts\. We give sufficient conditions for almost\-sure coverage of a finite knowledge domain and show how violating these conditions produces distinct failure modes: forgetting, exploration failure, acceptance failure, and contamination\. This separates successful discovery from collapse, support loss, rejection of valid artifacts, and accumulation of invalid ones\.

Imperfect verification and contamination\.We analyze how genuine and invalid retained artifacts grow under an imperfect verifier\. This yields a local contamination threshold: as the system exhausts easy discoveries, new genuine artifacts become rarer, so the tolerable false\-positive rate must also shrink\. Thus a fixed false\-positive rate can become unsafe near the frontier unless invalid mass shrinks comparably or verification improves\.

Missing mass and discovery\-cost scaling\.We distinguish Good–Turing batch unseen mass from the historically undiscovered valid mass that drives long\-term discovery\. Under a tail\-equivalence assumption relating the effective discovery frontier to a Zipf law with exponentα\>1\\alpha\>1, we show that the expected cumulative generation cost of discoveringDDdistinct genuine artifacts scales asRcum​\(D\)=Θ​\(cgen​Dα\)R\_\{\\rm cum\}\(D\)=\\Theta\(c\_\{\\rm gen\}D^\{\\alpha\}\), wherecgenc\_\{\\rm gen\}is the per\-candidate generation cost\. This quantifies the diminishing returns that arise as high\-probability discoveries are exhausted\.

Human amplification\.We formalize how human experts amplify NOVA through guidance, generation, and verification\. Human input can increase the mass assigned to new valid artifacts, improve acceptance of valid candidates, add expert\-generated candidates, and expand the reachable support\. This explains why human guidance is especially valuable near exploration barriers, where autonomous sampling assigns vanishing probability to new valid artifacts\.

### 1\.2Related Work

NOVA sits at the intersection of recursive synthetic\-data training, verified AI reasoning, self\-improvement, and classical missing\-mass estimation\. These areas study different pieces of the generate–verify–accumulate–retrain loop\. Our goal is to formalize the loop as a unified discovery process and characterize when it converges, when it fails, and how its cost scales\.

Model collapse\.A growing line of work studies the risks of recursive training on model\-generated data\.Shumailovet al\.\([2024](https://arxiv.org/html/2605.15219#bib.bib20)\)showed that training on recursively generated synthetic data can cause model collapse\.Gerstgrasseret al\.\([2024](https://arxiv.org/html/2605.15219#bib.bib21)\)showed that accumulating real data alongside synthetic data can mitigate this effect, whileSeddiket al\.\([2024](https://arxiv.org/html/2605.15219#bib.bib22)\)provided a statistical analysis of collapse mechanisms\. These works focus on distributional degradation under recursive training\. NOVA studies the broader discovery loop, identifying when iterative generate–verify–retrain processes converge or fail through forgetting, exploration failure, acceptance failure, and contamination \(Section[3](https://arxiv.org/html/2605.15219#S3)\)\.

AI theorem proving\.AlphaProof achieved IMO silver\-medal performance using Lean verification\(Hubertet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib26)\), while DeepSeek\-Prover\-V1\.5 reached 63\.5% on miniF2F\-test\(Xinet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib27)\), and DeepSeek\-Prover\-V2 advanced formal reasoning through reinforcement learning and subgoal decomposition\(Renet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib28)\)\. These systems instantiate NOVA in a near\-perfect\-verification regime, where false positives are mechanically controlled\.

Self\-training and self\-play\.STaR bootstraps reasoning\(Zelikmanet al\.,[2022](https://arxiv.org/html/2605.15219#bib.bib23)\), ReST applies reinforced self\-training\(Gulcehreet al\.,[2023](https://arxiv.org/html/2605.15219#bib.bib24)\), and AlphaGo Zero demonstrated superhuman performance through pure self\-play\(Silveret al\.,[2017](https://arxiv.org/html/2605.15219#bib.bib25)\)\. NOVA abstracts such generate–filter–retrain loops as adaptive sampling and derives conditions for convergence, collapse, and discovery\-cost scaling\.

Species estimation\.Good–Turing estimation\(Good,[1953](https://arxiv.org/html/2605.15219#bib.bib1)\)and its extensions\(Efron and Thisted,[1976](https://arxiv.org/html/2605.15219#bib.bib3); Orlitskyet al\.,[2016](https://arxiv.org/html/2605.15219#bib.bib4); McAllester and Schapire,[2000](https://arxiv.org/html/2605.15219#bib.bib5); Painsky,[2021](https://arxiv.org/html/2605.15219#bib.bib13),[2022](https://arxiv.org/html/2605.15219#bib.bib6),[2023](https://arxiv.org/html/2605.15219#bib.bib7); Lee and Boehme,[2025](https://arxiv.org/html/2605.15219#bib.bib8); Hanet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib14)\), including results for specific distributional models\(Chandraet al\.,[2024](https://arxiv.org/html/2605.15219#bib.bib16); Wolfer and Kontorovich,[2021](https://arxiv.org/html/2605.15219#bib.bib19)\), provide the classical foundation for missing\-mass estimation\. Error and convergence analyses are closely related\(Palet al\.,[2026](https://arxiv.org/html/2605.15219#bib.bib15); Skorski,[2021](https://arxiv.org/html/2605.15219#bib.bib12); Chandra and Thangaraj,[2024](https://arxiv.org/html/2605.15219#bib.bib11); Cohenet al\.,[2022](https://arxiv.org/html/2605.15219#bib.bib10); Acharyaet al\.,[2018](https://arxiv.org/html/2605.15219#bib.bib9)\)\. NOVA uses these tools as local diagnostics of diversity within a fixed generation batch\. Cumulative discovery, however, depends on how much probability the system assigns to valid artifacts that have not yet been discovered\. Our scaling laws therefore require an assumption on the shape of this remaining discovery frontier, not convergence of the model’s full sampling distribution to an ideal target\.

## 2Problem Formalization: The NOVA Framework

Let𝒦\\mathcal\{K\}be the set of valid knowledge artifacts, and let𝒳⊇𝒦\\mathcal\{X\}\\supseteq\\mathcal\{K\}be the ambient candidate space, which may also contain invalid candidates\. The ideal knowledge distributionPPis a distribution over𝒦\\mathcal\{K\}representing the intrinsic difficulty of discovering valid artifacts: artifacts with largerP​\(k\)P\(k\)are easier to encounter under an ideal generator\.

At iterationtt, the modelℳt\\mathcal\{M\}\_\{t\}has a retained set𝒦^t⊆𝒳\\widehat\{\\mathcal\{K\}\}\_\{t\}\\subseteq\\mathcal\{X\}of accepted candidates accumulated from previous iterations\. Its genuine component is𝒦t\+:=𝒦^t∩𝒦\\mathcal\{K\}\_\{t\}^\{\+\}:=\\widehat\{\\mathcal\{K\}\}\_\{t\}\\cap\\mathcal\{K\}, and its retained invalid component is𝒦^t∖𝒦\\widehat\{\\mathcal\{K\}\}\_\{t\}\\setminus\\mathcal\{K\}\. The modelℳt\\mathcal\{M\}\_\{t\}induces an actual sampling distributionQtQ\_\{t\}over𝒳\\mathcal\{X\}\. The distinction betweenQtQ\_\{t\}andPPis important:QtQ\_\{t\}determines what the model actually generates, whilePPis used later to characterize idealized difficulty and tail behavior\.

#### The NOVA Loop\.

At each iterationt=0,1,2,…t=0,1,2,\\ldots, NOVA executes the following steps\.

1. 1\.Generate:The modelℳt\\mathcal\{M\}\_\{t\}generatesNNartifactsct,1,…,ct,N∈𝒳c\_\{t,1\},\\ldots,c\_\{t,N\}\\in\\mathcal\{X\}i\.i\.d\.\\mathrm\{i\.i\.d\.\}fromQtQ\_\{t\}\.
2. 2\.Verify:Apply a verifierVVto each candidatect,ic\_\{t,i\}\. For a generic candidateCt∼QtC\_\{t\}\\sim Q\_\{t\}, define the true\-positive rate on new valid candidates asrt:=Pr⁡\[V​\(Ct\)=1∣Ct∈𝒦∖𝒦t\+\]r\_\{t\}:=\\Pr\[V\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\], and false\-positive rate on invalid candidates asδt:=Pr⁡\[V​\(Ct\)=1∣Ct∈𝒳∖𝒦\]\\delta\_\{t\}:=\\Pr\[V\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{X\}\\setminus\\mathcal\{K\}\]\.
3. 3\.Accumulate:Let𝒜t:=\{ct,i:V​\(ct,i\)=1\}\\mathcal\{A\}\_\{t\}:=\\\{c\_\{t,i\}:V\(c\_\{t,i\}\)=1\\\}be the accepted candidates, and update the retained set as𝒦^t\+1=𝒦^t∪𝒜t\\widehat\{\\mathcal\{K\}\}\_\{t\+1\}=\\widehat\{\\mathcal\{K\}\}\_\{t\}\\cup\\mathcal\{A\}\_\{t\}, with updated genuine component𝒦t\+1\+:=𝒦^t\+1∩𝒦\\mathcal\{K\}\_\{t\+1\}^\{\+\}:=\\widehat\{\\mathcal\{K\}\}\_\{t\+1\}\\cap\\mathcal\{K\}\.
4. 4\.Retrain:Model is updated using𝒦^t\+1\\widehat\{\\mathcal\{K\}\}\_\{t\+1\}, yielding the next sampling distributionQt\+1Q\_\{t\+1\}\.

Appendix[A](https://arxiv.org/html/2605.15219#A1)illustrates these components through three motivating settings: formal proof discovery, molecular discovery, and scientific hypothesis generation\.

#### Interpretation\.

NOVA separates three objects that are often conflated\. The setKKis the target knowledge domain: the valid artifacts that could in principle be discovered\. The distributionPPdescribes an idealized difficulty profile over this domain, assigning larger mass to artifacts that are intrinsically easier to encounter\. The model distributionQtQ\_\{t\}, however, describes what the current AI system actually samples\. Thus discovery is governed not only by what is valid, but by the interaction between the generator’s reachable support, the verifier’s acceptance behavior, and the retained data used for retraining\.

#### Key quantities\.

The central quantity for discovery is the model mass assigned to historically undiscovered valid artifacts:

Mtnew=∑k∈K∖Kt\+Qt​\(k\)\.M\_\{t\}^\{\\rm new\}=\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)\.WhenMtnewM\_\{t\}^\{\\rm new\}is large, the current model frequently generates new valid artifacts; when it is small, generation mostly repeats already discovered artifacts or produces invalid ones\. We decompose the remaining probability mass as

At=∑k∈Kt\+Qt​\(k\),Ut=∑x∈𝒳∖KQt​\(x\),A\_\{t\}=\\sum\_\{k\\in K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\),\\qquad U\_\{t\}=\\sum\_\{x\\in\\mathcal\{X\}\\setminus K\}Q\_\{t\}\(x\),so thatMtnew\+At\+Ut=1M\_\{t\}^\{\\rm new\}\+A\_\{t\}\+U\_\{t\}=1\. HereAtA\_\{t\}is the mass on already discovered genuine artifacts, whileUtU\_\{t\}is the mass on invalid candidates\.

Let

St=\{k∈𝒦∖𝒦t\+:k​is generated at least once at iteration​t​and accepted\}\.S\_\{t\}=\\\{k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}:k\\text\{ is generated at least once at iteration \}t\\text\{ and accepted\}\\\}\.Under a uniform true\-positive ratertr\_\{t\},

𝔼​\[\|St\|\]=rt​∑k∈𝒦∖𝒦t\+\(1−\(1−Qt​\(k\)\)N\)=N​rt​Mtnew−N​\(N−1\)2​rt​∑k∈𝒦∖𝒦t\+Qt​\(k\)2\+⋯\.\\mathbb\{E\}\[\|S\_\{t\}\|\]=r\_\{t\}\\sum\_\{k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\}\\bigl\(1\-\(1\-Q\_\{t\}\(k\)\)^\{N\}\\bigr\)=Nr\_\{t\}M\_\{t\}^\{\\rm new\}\-\\frac\{N\(N\-1\)\}\{2\}r\_\{t\}\\sum\_\{k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)^\{2\}\+\\cdots\.We refer to the settingN​Qt​\(k\)≪1NQ\_\{t\}\(k\)\\ll 1for allk∈𝒦∖𝒦t\+k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}as thesparse regime\. In this regime, duplicate discoveries within a batch are negligible, and

𝔼​\[\|St\|\]≈N​rt​Mtnew\.\\mathbb\{E\}\[\|S\_\{t\}\|\]\\approx Nr\_\{t\}M\_\{t\}^\{\\rm new\}\.
A separate notion of missing mass arises within a single generation batch\. GivenX1,…,XN∼QtX\_\{1\},\\ldots,X\_\{N\}\\sim Q\_\{t\}, define the ambient batch unseen mass

Mt,𝒳batch:=∑x∈𝒳Qt​\(x\)​𝟏​\[x∉\{X1,…,XN\}\]\.M^\{\\rm batch\}\_\{t,\\mathcal\{X\}\}:=\\sum\_\{x\\in\\mathcal\{X\}\}Q\_\{t\}\(x\)\\mathbf\{1\}\[x\\notin\\\{X\_\{1\},\\ldots,X\_\{N\}\\\}\]\.Good–Turing\(Good,[1953](https://arxiv.org/html/2605.15219#bib.bib1)\)estimates this local batch\-unseen mass, which diagnoses repetition or loss of diversity under the current generator\. It is distinct fromMtnewM\_\{t\}^\{\\rm new\}, the historically undiscovered valid mass that drives cumulative discovery; Appendix[B](https://arxiv.org/html/2605.15219#A2)formalizes the distinction and recalls Good–Toulmin forecasting\(Good and Toulmin,[1956](https://arxiv.org/html/2605.15219#bib.bib2)\)\.

## 3Coverage, Collapse, and Imperfect Verification

We now study when NOVA succeeds or fails by giving conditions for almost\-sure coverage, identifying exploration barriers, and analyzing contamination under imperfect verification\. For an undiscovered valid artifactk∈K∖Kt\+k\\in K\\setminus K\_\{t\}^\{\+\}, letrt,k:=Pr⁡\[V​\(k\)=1∣ℱt\]r\_\{t,k\}:=\\Pr\[V\(k\)=1\\mid\\mathcal\{F\}\_\{t\}\], denote its conditional acceptance probability if generated at iterationtt\.

### 3\.1Sufficient Conditions for Convergence

Letℱt\\mathcal\{F\}\_\{t\}denote theσ−\\sigma\-algebra generated by the history before generation at iterationtt, including𝒦t\+,𝒦^t,Qt,rt,δt\\mathcal\{K\}\_\{t\}^\{\+\},\\widehat\{\\mathcal\{K\}\}\_\{t\},Q\_\{t\},r\_\{t\},\\delta\_\{t\}, and all randomness from previous iterations\.

###### Theorem 1\(Sufficient Conditions for Almost\-Sure Coverage\)\.

AssumeQtQ\_\{t\}andrtr\_\{t\}areℱt\\mathcal\{F\}\_\{t\}\-measurable, and conditional onℱt\\mathcal\{F\}\_\{t\}, the next batch is sampled i\.i\.d\. fromQtQ\_\{t\}\. Suppose\|𝒦\|<∞\|\\mathcal\{K\}\|<\\inftyand the following hold:

1. C1Monotone accumulation:𝒦t\+⊆𝒦t\+1\+\\mathcal\{K\}\_\{t\}^\{\+\}\\subseteq\\mathcal\{K\}\_\{t\+1\}^\{\+\}for alltt\.
2. C2Persistent pre\-discovery exposure:For eachk∈𝒦k\\in\\mathcal\{K\}, on any sample path wherekkis never discovered, ∑t:k∉Kt\+\(1−\(1−Qt​\(k\)\)N\)=∞\.\\sum\_\{t:\\,k\\notin K\_\{t\}^\{\+\}\}\\left\(1\-\(1\-Q\_\{t\}\(k\)\)^\{N\}\\right\)=\\infty\.
3. C3Artifact\-wise nondegenerate acceptance:There existsrmin\>0r\_\{\\min\}\>0such that, for everyk∈Kk\\in K, on any sample path wherek∉Kt\+k\\notin K\_\{t\}^\{\+\},rt,k≥rmin\.r\_\{t,k\}\\geq r\_\{\\min\}\.\.
4. C4No false positives:δt=0\\delta\_\{t\}=0for alltt\.

Then𝒦t\+→𝒦\\mathcal\{K\}\_\{t\}^\{\+\}\\to\\mathcal\{K\}almost surely\.

The proof is given in Appendix[C](https://arxiv.org/html/2605.15219#A3); the main idea is to apply an artifact\-wise recurrence argument\. C1 prevents forgetting of already discovered genuine artifacts\. Ifk∉Kt\+k\\notin K\_\{t\}^\{\+\}, then conditional on the history, its probability of being generated and accepted at iterationttis at leastrmin​\(1−\(1−Qt​\(k\)\)N\)r\_\{\\min\}\(1\-\(1\-Q\_\{t\}\(k\)\)^\{N\}\)\. C2 ensures divergent aggregate exposure for every undiscovered artifact, while C3 prevents exposed valid artifacts from being rejected too often\. C4 rules out false positives, but does not requirert=1r\_\{t\}=1: false negatives can slow discovery, whereas false positives corrupt the retained knowledge base\. The assumption\|𝒦\|<∞\|\\mathcal\{K\}\|<\\inftymakes artifact\-wise discovery imply eventual full coverage\. Appendix[H](https://arxiv.org/html/2605.15219#A8)discusses the infinite\-domain case\.

Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)is qualitative: it does not implyQt→PQ\_\{t\}\\to Por give a discovery rate\. It only states that finite\-domain coverage occurs almost surely under persistent pre\-discovery exposure and retention of accepted genuine artifacts\. The next corollary makes explicit the corresponding exploration barrier\.

###### Corollary 2\(Exploration Barrier\)\.

Let𝒦∞\+=limt→∞𝒦t\+\\mathcal\{K\}\_\{\\infty\}^\{\+\}=\\lim\_\{t\\to\\infty\}\\mathcal\{K\}\_\{t\}^\{\+\}denote the asymptotic knowledge base\. Assume retraining is support\-preserving:

supp⁡\(Qt\+1\)⊆supp⁡\(Qt\)for all​t\.\\operatorname\{supp\}\(Q\_\{t\+1\}\)\\subseteq\\operatorname\{supp\}\(Q\_\{t\}\)\\quad\\text\{for all \}t\.Then

𝒦∞\+⊆supp⁡\(Q0\)∩𝒦\.\\mathcal\{K\}\_\{\\infty\}^\{\+\}\\subseteq\\operatorname\{supp\}\(Q\_\{0\}\)\\cap\\mathcal\{K\}\.

The corollary follows by induction: support preservation givessupp⁡\(Qt\)⊆supp⁡\(Q0\)\\operatorname\{supp\}\(Q\_\{t\}\)\\subseteq\\operatorname\{supp\}\(Q\_\{0\}\)for alltt\. Hence anyk∉supp⁡\(Q0\)k\\notin\\operatorname\{supp\}\(Q\_\{0\}\)is never generated and cannot enter𝒦t\+\\mathcal\{K\}\_\{t\}^\{\+\}\. Autonomous NOVA therefore cannot discover valid artifacts outside the initial generative support unless retraining, composition, human guidance, or another mechanism expands support\. For neural generators with broad literal support, the relevant notion is effective support: artifacts that can be generated with non\-negligible probability under the available compute budget\.

Together, Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)and Corollary[2](https://arxiv.org/html/2605.15219#Thmtheorem2)separate two requirements that are easy to conflate\. Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)says that coverage is possible when every valid artifact receives persistent pre\-discovery exposure and accepted discoveries are retained\. Corollary[2](https://arxiv.org/html/2605.15219#Thmtheorem2)says that such exposure cannot arise for artifacts outside the model’s initial support unless the support itself expands\. In this sense, autonomous discovery is limited not only by verification, but also by geometry of generator’s reachable support\.

### 3\.2Imperfect Verification and the Contamination Trap

When verification is imperfect, the retained set may contain both genuine and invalid artifacts\. LetGt=\|𝒦t\+\|G\_\{t\}=\|\\mathcal\{K\}\_\{t\}^\{\+\}\|andBt=\|𝒦^t∖𝒦\|B\_\{t\}=\|\\widehat\{\\mathcal\{K\}\}\_\{t\}\\setminus\\mathcal\{K\}\|\. We compare their one\-step increments under true\-positive ratertr\_\{t\}and false\-positive rateδt\\delta\_\{t\}\.

###### Proposition 3\(One\-step contamination increments\)\.

Assume that, conditional onℱt\\mathcal\{F\}\_\{t\}, theNNcandidates at iterationttare drawn independently fromQtQ\_\{t\}, and that verifier decisions are conditionally independent given each candidate\. Letrt,kr\_\{t,k\}denote the conditional acceptance probability of an undiscovered valid artifactk∈K∖Kt\+k\\in K\\setminus K\_\{t\}^\{\+\}\. Then the expected number of newly accepted genuine artifacts is

𝔼​\[Δ​Gt∣ℱt\]=∑k∈K∖Kt\+\(1−\(1−rt,k​Qt​\(k\)\)N\)\.\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}\\left\(1\-\(1\-r\_\{t,k\}Q\_\{t\}\(k\)\)^\{N\}\\right\)\.If invalid false positives are counted per accepted candidate, then

𝔼​\[Δ​Bt∣ℱt\]=N​δt​Ut\.\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=N\\delta\_\{t\}U\_\{t\}\.In the sparse regime, whereN​Qt​\(k\)≪1NQ\_\{t\}\(k\)\\ll 1for relevant undiscovered artifactskk, and whenrt,k=rtr\_\{t,k\}=r\_\{t\}over this region,𝔼​\[Δ​Gt∣ℱt\]≈N​rt​Mtnew\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\\approx Nr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}, and hence

𝔼​\[Δ​Bt∣ℱt\]𝔼​\[Δ​Gt∣ℱt\]≈δt​Utrt​Mtnew\.\\frac\{\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\{\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\\approx\\frac\{\\delta\_\{t\}U\_\{t\}\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\}\.

The proof is in Appendix[D](https://arxiv.org/html/2605.15219#A4)\. HereΔ​Bt\\Delta B\_\{t\}counts invalid false positives per accepted candidate; deduplication only changes the exact finite\-batch expression, as detailed in the appendix\.

###### Corollary 4\(Local contamination threshold\)\.

Define the marginal contamination fraction

ftmarg=𝔼​\[Δ​Bt∣ℱt\]𝔼​\[Δ​Gt∣ℱt\]\+𝔼​\[Δ​Bt∣ℱt\]\.f\_\{t\}^\{\\mathrm\{marg\}\}=\\frac\{\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\{\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\+\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\.In the sparse regime,

ftmarg≈δt​Utrt​Mtnew\+δt​Ut\.f\_\{t\}^\{\\mathrm\{marg\}\}\\approx\\frac\{\\delta\_\{t\}U\_\{t\}\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\+\\delta\_\{t\}U\_\{t\}\}\.Therefore, to keepftmarg≤fcriticalf\_\{t\}^\{\\mathrm\{marg\}\}\\leq f\_\{\\mathrm\{critical\}\}, it is sufficient that

δt≤δt∗:=rt​Mtnew​fcriticalUt​\(1−fcritical\)\.\\delta\_\{t\}\\leq\\delta\_\{t\}^\{\*\}:=\\frac\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}f\_\{\\mathrm\{critical\}\}\}\{U\_\{t\}\(1\-f\_\{\\mathrm\{critical\}\}\)\}\.

Proposition[3](https://arxiv.org/html/2605.15219#Thmtheorem3)and Corollary[4](https://arxiv.org/html/2605.15219#Thmtheorem4)identify a verification bottleneck at the discovery frontier\. AsMtnew→0M\_\{t\}^\{\\mathrm\{new\}\}\\to 0, genuine discoveries become increasingly rare\. If the false\-positive rateδt\\delta\_\{t\}remains fixed, then invalid candidates can enter the retained set faster than new genuine artifacts\. This is the contamination trap: progress reduces the new\-valid mass, which in turn tightens the verification requirement\. In particular, the safe false\-positive threshold satisfies

δt∗=rt​Mtnew​fcriticalUt​\(1−fcritical\),\\delta\_\{t\}^\{\*\}=\\frac\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}f\_\{\\mathrm\{critical\}\}\}\{U\_\{t\}\(1\-f\_\{\\mathrm\{critical\}\}\)\},so maintaining a bounded marginal contamination fraction requiresδt=O​\(Mtnew/Ut\)\\delta\_\{t\}=O\(M\_\{t\}^\{\\mathrm\{new\}\}/U\_\{t\}\)\. Thus, unless the invalid massUtU\_\{t\}shrinks comparably, verification must become increasingly precise, exactly when genuine discoveries are hardest to find\. Figure[1](https://arxiv.org/html/2605.15219#S3.F1.6)illustrates this effect\.

Appendix[E](https://arxiv.org/html/2605.15219#A5)extends the contamination analysis by adding verification cost\. It shows how finite per\-iteration budgets constrain the feasible batch size, how discovery becomes verification\-limited when checking candidates is much more expensive than generating them, and why reducing false positives requires increasing verification effort near the discovery frontier\.

Figure 1:Local contamination trap\. The curves show the fraction of newly accepted artifacts that are invalid,ftmarg≈δt​Ut/\(rt​Mtnew\+δt​Ut\)f\_\{t\}^\{\\mathrm\{marg\}\}\\approx\\delta\_\{t\}U\_\{t\}/\(r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\+\\delta\_\{t\}U\_\{t\}\), as a function of the false\-positive rateδt\\delta\_\{t\}\. As the new\-valid massMtnewM\_\{t\}^\{\\mathrm\{new\}\}decreases, even small false\-positive rates can make invalid artifacts dominate the accepted increments\.![Refer to caption](https://arxiv.org/html/2605.15219v1/x1.png)

## 4Discovery Rate and Cost Scaling

Section[3](https://arxiv.org/html/2605.15219#S3)gave conditions for coverage and failure\. We now ask a quantitative question: how many samples are needed to obtainDDgenuine discoveries? No universal rate law can hold without assumptions on how probability mass is distributed over the remaining undiscovered artifacts\. If the model concentrates on a finite subset, discovery saturates\. If it spreads mass across a large tail, discovery can continue much longer\. We therefore introduce a tail\-equivalence assumption linking the model’s effective discovery distribution to the ideal difficulty distributionPP\.

#### Effective discovery distribution\.

WhenMtnew\>0M\_\{t\}^\{\\mathrm\{new\}\}\>0, define the model’s effective discovery distribution over undiscovered valid artifacts byQ~t​\(k\)=Qt​\(k\)Mtnew,k∈𝒦∖𝒦t\+\.\\widetilde\{Q\}\_\{t\}\(k\)=\\frac\{Q\_\{t\}\(k\)\}\{M\_\{t\}^\{\\mathrm\{new\}\}\},\\quad k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\.Similarly, define the ideal conditional tail distribution byP~t​\(k\)=P​\(k\)P​\(𝒦∖𝒦t\+\),k∈𝒦∖𝒦t\+\.\\widetilde\{P\}\_\{t\}\(k\)=\\frac\{P\(k\)\}\{P\(\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\)\},\\quad k\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\.IfMtnew=0M\_\{t\}^\{\\mathrm\{new\}\}=0, autonomous discovery has reached the exploration barrier: the current generator assigns no mass to undiscovered valid artifacts\.

###### Assumption 1\(Uniform Tail\-Equivalent Occupancy Approximation\)\.

At iterationtt, rank the undiscovered artifacts𝒦∖𝒦t\+\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}askt,1,kt,2,…k\_\{t,1\},k\_\{t,2\},\\ldotsin decreasingP~t\\widetilde\{P\}\_\{t\}\-probability\. We assume that, for all ranksjjin the relevant tail regime, there exist constants0<c1≤c2<∞0<c\_\{1\}\\leq c\_\{2\}<\\infty, independent ofttand of the rank index, such that

c1​P~t​\(kt,j\)≤Q~t​\(kt,j\)≤c2​P~t​\(kt,j\)\.c\_\{1\}\\widetilde\{P\}\_\{t\}\(k\_\{t,j\}\)\\leq\\widetilde\{Q\}\_\{t\}\(k\_\{t,j\}\)\\leq c\_\{2\}\\widetilde\{P\}\_\{t\}\(k\_\{t,j\}\)\.

We now address the central question of*how much computation is needed to obtain a target number of discoveries*\. The next proposition recalls the Zipf occupancy law for distinct discoveries, and Theorem[6](https://arxiv.org/html/2605.15219#Thmtheorem6)transfers it to NOVA via Assumption[1](https://arxiv.org/html/2605.15219#Thmassumption1)\.

###### Proposition 5\(Zipf Occupancy\(see, e\.g\., Karlin,[1967](https://arxiv.org/html/2605.15219#bib.bib17); Ben\-Hamouet al\.,[2017](https://arxiv.org/html/2605.15219#bib.bib18)\)\)\.

ForNNi\.i\.d\. draws from a Zipf distribution with exponentα\>1\\alpha\>1\(countably infinite, or finite in the pre\-saturation regime\), letDND\_\{N\}denote the number of distinct values observed\. Then

𝔼​\[DN\]=Θ​\(N1/α\)\.\\mathbb\{E\}\[D\_\{N\}\]=\\Theta\(N^\{1/\\alpha\}\)\.

###### Theorem 6\(Cumulative Cost under Tail Equivalence\)\.

Assume that, over the discovery horizon, the adaptive sequence of effective discovery distributions is uniformly comparable, up to constant factors, to sequential occupancy sampling from a Zipf tail with exponentα\>1\\alpha\>1, as in Assumption[1](https://arxiv.org/html/2605.15219#Thmassumption1)\. Assume also thatrtr\_\{t\}is bounded above and below by positive constants\. LetDT:=\|𝒦T\+\|D\_\{T\}:=\|\\mathcal\{K\}\_\{T\}^\{\+\}\|denote the cumulative number of genuine discoveries afterTTiterations, and letETnew:=∑t=0T−1N​rt​MtnewE\_\{T\}^\{\\mathrm\{new\}\}:=\\sum\_\{t=0\}^\{T\-1\}Nr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}denote the cumulative effective exposure to the undiscovered valid region\. Then

𝔼​\[DT\]=Θ​\(\(ETnew\)1/α\)\.\\mathbb\{E\}\[D\_\{T\}\]=\\Theta\\\!\\left\(\(E\_\{T\}^\{\\mathrm\{new\}\}\)^\{1/\\alpha\}\\right\)\.Equivalently, the effective discovery exposure required in expectation to obtainDDdistinct genuine discoveries scales asΘ​\(Dα\)\\Theta\(D^\{\\alpha\}\)\. If each unit of effective exposure costsΘ​\(cgen\)\\Theta\(c\_\{\\mathrm\{gen\}\}\), then the corresponding generation cost scales as

Rcum​\(D\)=Θ​\(cgen​Dα\)\.R\_\{\\mathrm\{cum\}\}\(D\)=\\Theta\(c\_\{\\mathrm\{gen\}\}D^\{\\alpha\}\)\.\(1\)

*Proof sketch\.*By Assumption[1](https://arxiv.org/html/2605.15219#Thmassumption1), conditional on reaching the undiscovered valid region, the adaptive NOVA discovery process is comparable, up to constant factors, to sequential occupancy sampling from a Zipf tail with exponentα\\alpha\. The cumulative effective exposure to this region isETnew=∑t=0T−1N​rt​MtnewE\_\{T\}^\{\\mathrm\{new\}\}=\\sum\_\{t=0\}^\{T\-1\}Nr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}, so Proposition[5](https://arxiv.org/html/2605.15219#Thmtheorem5)gives𝔼​\[DT\]=Θ​\(\(ETnew\)1/α\)\\mathbb\{E\}\[D\_\{T\}\]=\\Theta\(\(E\_\{T\}^\{\\mathrm\{new\}\}\)^\{1/\\alpha\}\)\. Inverting this expectation\-level scaling yields an effective exposure requirementΘ​\(Dα\)\\Theta\(D^\{\\alpha\}\)for obtainingDDdiscoveries in expectation\. Multiplying by the cost per unit of effective exposure givesRcum​\(D\)=Θ​\(cgen​Dα\)R\_\{\\mathrm\{cum\}\}\(D\)=\\Theta\(c\_\{\\mathrm\{gen\}\}D^\{\\alpha\}\)\. ∎

## 5Human–AI Collaborative NOVA

The preceding sections identify three barriers for autonomous NOVA: reachable support limits what can be discovered \(Corollary[2](https://arxiv.org/html/2605.15219#Thmtheorem2)\), imperfect verification becomes more dangerous as new\-valid mass shrinks \(Corollary[4](https://arxiv.org/html/2605.15219#Thmtheorem4)\), and Zipf occupancy makes cumulative discovery costs grow superlinearly \(Theorem[6](https://arxiv.org/html/2605.15219#Thmtheorem6)\)\. Human experts can intervene at these bottlenecks by expanding support, redirecting probability mass, proposing novel candidates, and providing high\-precision verification when formal verifiers are unavailable\.

This formalizes the qualitative observation emphasized byKlowden and Tao \([2026](https://arxiv.org/html/2605.15219#bib.bib29)\): in frontier mathematical practice, AI systems are most effective when embedded in expert\-guided workflows rather than treated as fully autonomous discovery engines\.

In the augmented NOVA loop, detailed in Appendix[G](https://arxiv.org/html/2605.15219#A7), a human expert modifiesQtQ\_\{t\}to a guided distributionQt′Q\_\{t\}^\{\\prime\}, the AI generatesNAIN\_\{\\mathrm\{AI\}\}candidates fromQt′Q\_\{t\}^\{\\prime\}, the expert generatesNHN\_\{H\}candidates from their own distributionPHP\_\{H\}, and the combined candidates are verified\. Let

Mtnew,guided:=Qt′\(𝒦∖𝒦t\+\),Mtnew,H:=PH\(𝒦∖𝒦t\+\)M\_\{t\}^\{\\mathrm\{new,guided\}\}:=Q\_\{t\}^\{\\prime\}\(\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\),\\qquad M\_\{t\}^\{\\mathrm\{new\},H\}:=P\_\{H\}\(\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\)denote the new\-valid mass under the guided AI distribution and the human proposal distribution, respectively\. Letreff,tr\_\{\\mathrm\{eff\},t\}denote the effective true\-positive rate for new valid AI\-generated candidates after human review:

reff,t:=Pr⁡\[Veff​\(Ct\)=1∣Ct∈𝒦∖𝒦t\+,Ct∼Qt′\]\.r\_\{\\mathrm\{eff\},t\}:=\\Pr\[V\_\{\\mathrm\{eff\}\}\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\},\\,C\_\{t\}\\sim Q\_\{t\}^\{\\prime\}\]\.Define the*human amplification factor*as

AH=𝔼​\[\|StH\+AI\|\]𝔼​\[\|StAI\|\]\.A\_\{H\}=\\frac\{\\mathbb\{E\}\[\|S\_\{t\}^\{H\+\\mathrm\{AI\}\}\|\]\}\{\\mathbb\{E\}\[\|S\_\{t\}^\{\\mathrm\{AI\}\}\|\]\}\.
###### Theorem 7\(Sparse\-Regime Human Amplification\)\.

In the sparse regime, neglecting duplicate discoveries between human\- and AI\-generated candidates, the amplification factor admits the first\-order decomposition

AH=Aguide⋅Averify⋅Agen,A\_\{H\}=A\_\{\\mathrm\{guide\}\}\\cdot A\_\{\\mathrm\{verify\}\}\\cdot A\_\{\\mathrm\{gen\}\},where

Aguide=Mtnew,guidedMtnew,Averify=reff,trt,Agen=1\+NH​ρH,t​Mtnew,HNAI​reff,t​Mtnew,guided\.A\_\{\\mathrm\{guide\}\}=\\frac\{M\_\{t\}^\{\\mathrm\{new,guided\}\}\}\{M\_\{t\}^\{\\mathrm\{new\}\}\},\\qquad A\_\{\\mathrm\{verify\}\}=\\frac\{r\_\{\\mathrm\{eff\},t\}\}\{r\_\{t\}\},\\qquad A\_\{\\mathrm\{gen\}\}=1\+\\frac\{N\_\{H\}\\rho\_\{H,t\}M\_\{t\}^\{\\mathrm\{new\},H\}\}\{N\_\{\\mathrm\{AI\}\}r\_\{\\mathrm\{eff\},t\}M\_\{t\}^\{\\mathrm\{new,guided\}\}\}\.Herertr\_\{t\}is the autonomous true\-positive rate,reff,tr\_\{\\mathrm\{eff\},t\}is the effective true\-positive rate for guided AI\-generated candidates after human review, andρH,t\\rho\_\{H,t\}is acceptance rate for valid human\-generated candidates\.

Theorem[7](https://arxiv.org/html/2605.15219#Thmtheorem7)separates three human contributions: guidance raises the new\-valid mass reached by AI generation, verification improves acceptance of valid AI artifacts, and generation adds human\-proposed candidates\. Together, these factors suggest a copilot regime in which AI performs high\-volume search within guided regions while the human provides frontier guidance, hypotheses, and difficult verification\. Guidance is especially central because it can change not only discovery probability within the current support, but also the support itself\. Next theorem formalizes this effect\.

###### Theorem 8\(Human\-Guided Support Expansion\)\.

If human guidance changesQtQ\_\{t\}toQt′Q\_\{t\}^\{\\prime\}such thatK∩supp⁡\(Qt\)⊊K∩supp⁡\(Qt′\)K\\cap\\operatorname\{supp\}\(Q\_\{t\}\)\\subsetneq K\\cap\\operatorname\{supp\}\(Q\_\{t\}^\{\\prime\}\), then the reachable valid set strictly expands\. Hence guidance can break the autonomous exploration barrier wheneverQt′Q\_\{t\}^\{\\prime\}assigns non\-negligible mass to valid artifacts outside the previous effective support while preserving the previously reachable valid support\.

Appendix[G](https://arxiv.org/html/2605.15219#A7)gives the details of augmented NOVA loop, proof of Theorems[7](https://arxiv.org/html/2605.15219#Thmtheorem7)and[8](https://arxiv.org/html/2605.15219#Thmtheorem8), and a simple effort\-allocation principle for deciding when human effort should be spent\.

## 6Conclusions, Limitations, Broader Impacts, and Future Directions

We formalized the generate–verify–accumulate–retrain loop to study when AI systems can reliably discover genuinely new knowledge, when they fail, and what limits their cost\. NOVA shows that AI\-driven discovery is a systems\-level phenomenon, not merely generating more candidates\. Discovery succeeds only when exploration, verification, accumulation, and retraining remain aligned\. Otherwise, the loop can stall, forget, or contaminate itself\. Autonomous discovery is constrained by what the model can reach, what the verifier can accept, and how quickly the frontier thins as easy discoveries are exhausted\. Human guidance is therefore not just additional sampling\. It can redirect search, expand reachable support, and reshape the discovery frontier\. This provides a foundation for understanding when AI can discover autonomously and why the strongest discovery systems are likely to be human\-guided, verification\-aware, and designed around the geometry of the unknown\.

Limitations\.NOVA is a stylized framework, and its results should be interpreted accordingly\. First, the coverage theorem gives sufficient, not necessary, conditions for finite\-domain discovery; it isolates failure modes but does not imply rapid discovery or practical feasibility\. Second, the cost lawRcum​\(D\)=Θ​\(cgen​Dα\)R\_\{\\rm cum\}\(D\)=\\Theta\(c\_\{\\rm gen\}D^\{\\alpha\}\)is not universal\. It depends on a Zipf\-like effective discovery tail that remains comparable across the relevant horizon; if retraining reshapes the frontier or induces a different tail, the appropriate occupancy law must replace the Zipf idealization\. Third, verification is modeled abstractly through true\- and false\-positive behavior\. This captures local contamination effects, but not the full complexity of proof checking, simulation, experimentation, human judgment, correlated errors, or verification latency\. Finally, NOVA treats discoveries as discrete artifacts sampled from a knowledge space\. This abstraction does not capture all aspects of scientific progress, such as conceptual reframing, new measurements, causal experimentation, etc\. Thus NOVA should be viewed as a foundation for analyzing discovery loops, rather than a complete theory of discovery\.

Broader impacts\.On the positive side, a theory of generate–verify–accumulate–retrain loops may help design more reliable AI\-assisted discovery systems for mathematics, science, and engineering by making failure modes, verification requirements, and human guidance explicit\. On the negative side, the same loops could be misused in harmful domains, or could lead to over\-reliance on autonomous systems when verification is weak\. In particular, false positives and contaminated retained artifacts may cause systems to amplify invalid discoveries\. Our analysis highlights verification, contamination control, support limits, and human oversight as important safeguards\.

Open problems\.There are several immediate future directions\.\(1\) Collaborative discovery:How does effective missing mass scale withmmmodels with diverse supports, and can composition yield superlinear gains?\(2\) Verification difficulty:Can we formalize how the difficulty of checking candidates \(from mechanical proof checking to code testing, noisy experiments, and subjective evaluation\) limits the knowledge level achievable by autonomous NOVA?\(3\) Information limits:For the subset𝒦∗⊆𝒦\\mathcal\{K\}^\{\*\}\\subseteq\\mathcal\{K\}discoverable from the initial data, model, and allowed NOVA operations, can\|𝒦∗\|\|\\mathcal\{K\}^\{\*\}\|be bounded by the information that𝒟0\\mathcal\{D\}\_\{0\}contains about𝒦\\mathcal\{K\}, e\.g\.,\|𝒦∗\|≲2I​\(𝒟0;𝒦\)\|\\mathcal\{K\}^\{\*\}\|\\lesssim 2^\{I\(\\mathcal\{D\}\_\{0\};\\mathcal\{K\}\)\}?

## References

- J\. Acharya, Y\. Bao, Y\. Kang, and Z\. Sun \(2018\)Improved bounds for minimax risk of estimating missing mass\.In2018 IEEE International Symposium on Information Theory \(ISIT\),pp\. 326–330\.External Links:[Document](https://dx.doi.org/10.1109/ISIT.2018.8437620)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- A\. Ben\-Hamou, S\. Boucheron, and M\. I\. Ohannessian \(2017\)Concentration inequalities in the infinite urn scheme for occupancy counts and the missing mass, with applications\.Bernoulli23\(1\),pp\. 249–287\.Cited by:[Proposition 5](https://arxiv.org/html/2605.15219#Thmtheorem5)\.
- P\. Chandra, A\. Thangaraj, and N\. Rajaraman \(2024\)How good is good\-turing for markov samples?\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=KokkP2nQ24)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- P\. Chandra and A\. Thangaraj \(2024\)Missing g\-mass: investigating the missing parts of distributions\.IEEE Transactions on Information Theory70\(10\),pp\. 7049–7065\.External Links:[Document](https://dx.doi.org/10.1109/TIT.2024.3440661)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- S\. Cohen, T\. Routtenberg, and L\. Tong \(2022\)Non\-bayesian parametric missing\-mass estimation\.IEEE Transactions on Signal Processing70,pp\. 3709–3725\.External Links:[Document](https://dx.doi.org/10.1109/TSP.2022.3186176)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- B\. Efron and R\. Thisted \(1976\)Estimating the number of unseen species: how many words did Shakespeare know?\.Biometrika63\(3\),pp\. 435–447\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- M\. Gerstgrasser, R\. Schaeffer, A\. Dey, R\. Rafailov, T\. Korbak, H\. Sleight, R\. Agrawal, J\. Hughes, D\. B\. Pai, A\. Gromov, D\. A\. Roberts, D\. Yang, D\. L\. Donoho, and S\. Koyejo \(2024\)Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data\.InCOLM,Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p2.1)\.
- I\. J\. Good and G\. H\. Toulmin \(1956\)The number of new species, and the increase in population coverage, when a sample is increased\.Biometrika43,pp\. 45–63\.Cited by:[§2](https://arxiv.org/html/2605.15219#S2.SS0.SSS0.Px3.p3.2),[Theorem 11](https://arxiv.org/html/2605.15219#Thmtheorem11)\.
- I\. J\. Good \(1953\)The population frequencies of species and the estimation of population parameters\.Biometrika40\(3\-4\),pp\. 237–264\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1),[§2](https://arxiv.org/html/2605.15219#S2.SS0.SSS0.Px3.p3.2),[Theorem 9](https://arxiv.org/html/2605.15219#Thmtheorem9)\.
- C\. Gulcehre, T\. L\. Paine, S\. Srinivasan, K\. Konyushkova, L\. Weerts, A\. Sharma, A\. Siddhant, A\. Ahern, M\. Wang, C\. Gu, W\. Macherey, A\. Doucet, O\. Firat, and N\. de Freitas \(2023\)Reinforced self\-training \(ReST\) for language modeling\.arXiv preprint arXiv:2308\.08998\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p4.1)\.
- Y\. Han, J\. Niles\-Weed, Y\. Shen, and Y\. Wu \(2025\)Besting good–turing: optimality of non\-parametric maximum likelihood for distribution estimation\.External Links:2509\.07355,[Link](https://arxiv.org/abs/2509.07355)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- T\. Hubert, R\. Mehta, L\. Sartran, M\. Z\. Horváth, G\. Žužić, E\. Wieser, A\. Huang, J\. Schrittwieser, Y\. Schroecker, H\. Masoom, O\. Bertolli, T\. Zahavy, A\. Mandhane, J\. Yung, I\. Beloshapka, B\. Ibarz, V\. Veeriah, L\. Yu, O\. Nash, P\. Lezeau, S\. Mercuri, C\. Sönne, B\. Mehta, A\. Davies, D\. Zheng, F\. Pedregosa, Y\. Li, I\. von Glehn, M\. Rowland, S\. Albanie, A\. Velingker, S\. Schmitt, E\. Lockhart, E\. Hughes, H\. Michalewski, N\. Sonnerat, D\. Hassabis, P\. Kohli, and D\. Silver \(2025\)Olympiad\-level formal mathematical reasoning with reinforcement learning\.Nature\.Cited by:[Appendix A](https://arxiv.org/html/2605.15219#A1.SS0.SSS0.Px1.p1.5),[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p3.1),[§1](https://arxiv.org/html/2605.15219#S1.p1.1)\.
- J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko, A\. Bridgland, C\. Meyer, S\. A\. A\. Kohl, A\. J\. Ballard, A\. Cowie, B\. Romera\-Paredes, S\. Nikolov, R\. Jain, J\. Adler, T\. Back, S\. Petersen, D\. Reiman, E\. Clancy, M\. Zielinski, M\. Steinegger, M\. Pacholska, T\. Berghammer, S\. Bodenstein, D\. Silver, O\. Vinyals, A\. W\. Senior, K\. Kavukcuoglu, P\. Kohli, and D\. Hassabis \(2021\)Highly accurate protein structure prediction with AlphaFold\.Nature596\(7873\),pp\. 583–589\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03819-2)Cited by:[Appendix A](https://arxiv.org/html/2605.15219#A1.SS0.SSS0.Px2.p1.3)\.
- S\. Karlin \(1967\)Central limit theorems for certain infinite urn schemes\.Journal of Mathematics and Mechanics17\(4\),pp\. 373–401\.Cited by:[Proposition 5](https://arxiv.org/html/2605.15219#Thmtheorem5)\.
- T\. Klowden and T\. Tao \(2026\)Mathematical methods and human thought in the age of AI\.arXiv preprint arXiv:2603\.26524\.Cited by:[§5](https://arxiv.org/html/2605.15219#S5.p2.1)\.
- S\. Lee and M\. Boehme \(2025\)How much is unseen depends chiefly on information about the seen\.InICLR \(Spotlight\),Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- D\. McAllester and R\. E\. Schapire \(2000\)On the convergence rate of Good\-Turing estimators\.InCOLT,Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1),[Theorem 9](https://arxiv.org/html/2605.15219#Thmtheorem9)\.
- A\. Merchant, S\. Batzner, S\. S\. Schoenholz, M\. Aykol, G\. Cheon, and E\. D\. Cubuk \(2023\)Scaling deep learning for materials discovery\.Nature624,pp\. 80–85\.Cited by:[Appendix A](https://arxiv.org/html/2605.15219#A1.SS0.SSS0.Px2.p1.3)\.
- A\. Orlitsky, A\. T\. Suresh, and Y\. Wu \(2016\)Optimal prediction of the number of unseen species\.Proceedings of the National Academy of Sciences113\(47\),pp\. 13283–13288\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- A\. Painsky \(2021\)Refined convergence rates of the good\-turing estimator\.In2021 IEEE Information Theory Workshop \(ITW\),pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ITW48936.2021.9611389)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- A\. Painsky \(2022\)Convergence guarantees for the Good\-Turing estimator\.Journal of Machine Learning Research23\(279\),pp\. 1–37\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- A\. Painsky \(2023\)Generalized Good\-Turing improves missing mass estimation\.Journal of the American Statistical Association118\(543\),pp\. 1890–1899\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- B\. Pal, S\. Bhattacharya, and M\. Singh \(2026\)Blind\-spot mass: a good\-turing framework for quantifying deployment coverage risk in machine learning systems\.External Links:2604\.05057,[Link](https://arxiv.org/abs/2604.05057)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- Z\.Z\. Ren, Z\. Shao, J\. Song, H\. Xin, H\. Wang, W\. Zhao, L\. Zhang, Z\. Fu, Q\. Zhu, D\. Yang, Z\.F\. Wu, Z\. Gou, S\. Ma, H\. Tang, Y\. Liu, W\. Gao, D\. Guo, and C\. Ruan \(2025\)DeepSeek\-Prover\-V2: advancing formal mathematical reasoning via reinforcement learning for subgoal decomposition\.arXiv preprint arXiv:2504\.21801\.Cited by:[Appendix A](https://arxiv.org/html/2605.15219#A1.SS0.SSS0.Px1.p1.5),[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p3.1),[§1](https://arxiv.org/html/2605.15219#S1.p1.1)\.
- M\. E\. A\. Seddik, S\. Chen, S\. Hayou, P\. Youssef, and M\. Debbah \(2024\)How bad is training on synthetic data? A statistical analysis of language model collapse\.InCOLM,Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p2.1)\.
- I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. Gal \(2024\)AI models collapse when trained on recursively generated data\.Nature631,pp\. 755–759\.Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p2.1)\.
- D\. Silver, J\. Schrittwieser, K\. Simonyan, I\. Antonoglou, A\. Huang, A\. Guez, T\. Hubert, L\. Baker, M\. Lai, A\. Bolton, Y\. Chen, T\. Lillicrap, F\. Hui, L\. Sifre, G\. van den Driessche, T\. Graepel, and D\. Hassabis \(2017\)Mastering the game of Go without human knowledge\.Nature550\(7676\),pp\. 354–359\.External Links:[Document](https://dx.doi.org/10.1038/nature24270)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p4.1)\.
- M\. Skorski \(2021\)Mean\-squared accuracy of good\-turing estimator\.In2021 IEEE International Symposium on Information Theory \(ISIT\),pp\. 2846–2851\.External Links:[Document](https://dx.doi.org/10.1109/ISIT45174.2021.9518169)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- G\. Wolfer and A\. Kontorovich \(2021\)Statistical estimation of ergodic Markov chain kernel over discrete state space\.Bernoulli27\(1\),pp\. 532–553\.External Links:[Document](https://dx.doi.org/10.3150/20-BEJ1248)Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p5.1)\.
- H\. Xin, Z\.Z\. Ren, J\. Song, Z\. Shao, W\. Zhao, H\. Wang, B\. Liu, L\. Zhang, X\. Lu, Q\. Du, W\. Gao, H\. Zhang, Q\. Zhu, D\. Yang, Z\. Gou, Z\.F\. Wu, F\. Luo, and C\. Ruan \(2025\)DeepSeek\-Prover\-V1\.5: harnessing proof assistant feedback for reinforcement learning and Monte\-Carlo tree search\.InICLR,Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p3.1)\.
- E\. Zelikman, Y\. Wu, J\. Mu, and N\. D\. Goodman \(2022\)STaR: bootstrapping reasoning with reasoning\.InNeurIPS,Cited by:[§1\.2](https://arxiv.org/html/2605.15219#S1.SS2.p4.1),[§1](https://arxiv.org/html/2605.15219#S1.p1.1)\.

## Appendix AMotivating Examples

#### Example 1: Discovering new mathematical proofs\.

Consider𝒦\\mathcal\{K\}as the set of valid formal proofs of mathematical conjectures\. A language modelℳt\\mathcal\{M\}\_\{t\}generates candidate proofs in a formal language \(e\.g\., Lean 4\)\. The verification step is performed by a proof assistant: the Lean type\-checker mechanically checks whether each candidate is a well\-typed proof of the stated theorem\. In this setting, verification is effectively perfect relative to the formal specification: false positives are mechanically ruled out, soδt=0\\delta\_\{t\}=0, while artifact\-wise acceptance satisfiesrt,k=1r\_\{t,k\}=1for any correctly generated formal proofkk\. This is the ideal verification regime for NOVA\. Recent systems such as AlphaProof\[Hubertet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib26)\]and DeepSeek\-Prover\-V2\[Renet al\.,[2025](https://arxiv.org/html/2605.15219#bib.bib28)\]operate in this regime\.

#### Example 2: Discovering new molecules and materials\.

Consider𝒦\\mathcal\{K\}as the set of molecules with a desired functional property\. A generative model proposes candidate molecular structures, and verification is performed through wet\-lab experiments or computational simulations\. Here, verification is stochastic: artifact\-wise true\-positive ratesrt,kr\_\{t,k\}may be less than one and false\-positive ratesδt\\delta\_\{t\}may be nonzero\. This places the system in the imperfect verification regime analyzed in Section[3](https://arxiv.org/html/2605.15219#S3), where the contamination dynamics of Proposition[3](https://arxiv.org/html/2605.15219#Thmtheorem3)become relevant\. Recent AI\-driven discovery platforms for proteins\[Jumperet al\.,[2021](https://arxiv.org/html/2605.15219#bib.bib30)\]and materials\[Merchantet al\.,[2023](https://arxiv.org/html/2605.15219#bib.bib31)\]exemplify this setting\.

#### Example 3: Discovering new scientific hypotheses\.

At the frontier of knowledge, consider𝒦\\mathcal\{K\}as the set of true scientific hypotheses\. Verification may be infeasible: testing a hypothesis may require experiments that take years or new instrumentation\. In the extreme case,rt,k≈0r\_\{t,k\}\\approx 0for the most novel artifacts, making autonomous NOVA essentially impossible\. It is precisely in this regime that human expert guidance \(Section[5](https://arxiv.org/html/2605.15219#S5)\) becomes indispensable\.

## Appendix BMissing Mass Estimation Details

This appendix clarifies the different missing\-mass quantities that appear in NOVA\. The classical Good–Turing estimator applies to the batch unseen mass under the current generatorQtQ\_\{t\}: the probability that another sample fromQtQ\_\{t\}would produce an artifact not seen in the current batch\. This is distinct from the historical new\-valid massMtnewM\_\{t\}^\{\\rm new\}, which is the probability that the generator produces a valid artifact not yet accumulated inKt\+K\_\{t\}^\{\+\}\. We formalize this distinction, give the exact one\-step discovery rate and its sparse\-regime approximation, and recall the Good–Toulmin forecasting formula for fixed\-generator batch extrapolation\.

### B\.1Good–Turing Estimator

###### Theorem 9\(Good–Turing Estimator for Batch Missing Mass –Good \[[1953](https://arxiv.org/html/2605.15219#bib.bib1)\]andMcAllester and Schapire \[[2000](https://arxiv.org/html/2605.15219#bib.bib5)\]\)\.

LetX1,…,XNX\_\{1\},\\ldots,X\_\{N\}be i\.i\.d\. from a discrete distributionQQ, and let

MN=∑xQ​\(x\)​𝟏​\[x∉\{X1,…,XN\}\]M\_\{N\}=\\sum\_\{x\}Q\(x\)\\mathbf\{1\}\[x\\notin\\\{X\_\{1\},\\ldots,X\_\{N\}\\\}\]be the missing mass after the batch\. The Good–Turing estimatorf1/Nf\_\{1\}/N, wheref1f\_\{1\}is the number of species observed exactly once, is the classical estimator forMNM\_\{N\}, the probability that the next draw belongs to a species unseen in the current batch\.

###### Proposition 10\(Good–Turing Estimates Ambient Batch Unseen Mass\)\.

In the NOVA framework, at iterationttthe model generatesNNcandidates i\.i\.d\. fromQtQ\_\{t\}over𝒳\\mathcal\{X\}\. SinceQtQ\_\{t\}may assign positive mass to invalid candidates \(Ut\>0U\_\{t\}\>0\), the Good–Turing estimatorf1\(t\)/Nf\_\{1\}^\{\(t\)\}/Nestimates the ambient batch unseen massMt,𝒳batchM\_\{t,\\mathcal\{X\}\}^\{\\mathrm\{batch\}\}, providing an upper bound on the valid\-artifact component:Mt,Kbatch≤Mt,𝒳batchM\_\{t,K\}^\{\\mathrm\{batch\}\}\\leq M\_\{t,\\mathcal\{X\}\}^\{\\mathrm\{batch\}\}\.

### B\.2Estimating New\-Valid Mass

If𝒦t\+\\mathcal\{K\}\_\{t\}^\{\+\}is known explicitly, a practical unbiased estimator from one batch is:

M^tnew,MC=1N​∑i=1N𝟏​\[Xi∈𝒦∖𝒦t\+\]\\widehat\{M\}\_\{t\}^\{\\mathrm\{new,MC\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[X\_\{i\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\]\(2\)This is unbiased forMtnewM\_\{t\}^\{\\mathrm\{new\}\}whenXi∼QtX\_\{i\}\\sim Q\_\{t\}, since𝔼​\[𝟏​\[Xi∈𝒦∖𝒦t\+\]\]=Mtnew\\mathbb\{E\}\[\\mathbf\{1\}\[X\_\{i\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\]\]=M\_\{t\}^\{\\mathrm\{new\}\}\. The indicator requires checking both thatXiX\_\{i\}is valid \(Xi∈𝒦X\_\{i\}\\in\\mathcal\{K\}\) and not already discovered \(Xi∉𝒦t\+X\_\{i\}\\notin\\mathcal\{K\}\_\{t\}^\{\+\}\)\.

### B\.3Good–Toulmin Forecasting

###### Theorem 11\(Good–Toulmin Forecasting \(Batch Extrapolation\) –Good and Toulmin \[[1956](https://arxiv.org/html/2605.15219#bib.bib2)\]\)\.

At fixed iterationtt, supposeNNcandidates are sampled i\.i\.d\. from a fixed generatorQtQ\_\{t\}, and letfr\(t\)f\_\{r\}^\{\(t\)\}be the number of artifacts observed exactlyrrtimes\. For an additional sample of sizes​NsN, the Good–Toulmin formal series estimates the expected number of new species by

∑r≥1\(−s\)r\+1​fr\(t\)\.\\sum\_\{r\\geq 1\}\(\-s\)^\{r\+1\}f\_\{r\}^\{\(t\)\}\.\(3\)

Note that this is a fixed\-generator batch extrapolation formula\. OnceQtQ\_\{t\}changes across iterations due to retraining, it no longer applies directly without additional stability assumptions\.

## Appendix CConvergence Proof

###### Proof of Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)\.

Monotone growth:Under C1, already discovered genuine artifacts are not lost, so\|𝒦t\+\|\|\\mathcal\{K\}\_\{t\}^\{\+\}\|is non\-decreasing\. Under C4, accepted artifacts are genuine, so false positives do not enter the retained set\.

Eventual discovery \(Borel–Cantelli\):Letℱt\\mathcal\{F\}\_\{t\}denote the history before generation at iterationtt, includingKt\+K\_\{t\}^\{\+\},QtQ\_\{t\}, the artifact\-wise acceptance probabilities, and all randomness from previous iterations\. Fix anyk∈𝒦k\\in\\mathcal\{K\}\. Conditional onℱt\\mathcal\{F\}\_\{t\}, ifk∉Kt\+k\\notin K\_\{t\}^\{\+\}, the probability thatkkis generated at least once in the batch is

1−\(1−Qt​\(k\)\)N\.1\-\(1\-Q\_\{t\}\(k\)\)^\{N\}\.By C3, if generated,kkis accepted with probability at leastrminr\_\{\\min\}\. Therefore,

Pr⁡\[k​is discovered at iteration​t∣ℱt\]≥rmin​\(1−\(1−Qt​\(k\)\)N\)\.\\Pr\\\!\\left\[k\\text\{ is discovered at iteration \}t\\mid\\mathcal\{F\}\_\{t\}\\right\]\\geq r\_\{\\min\}\\bigl\(1\-\(1\-Q\_\{t\}\(k\)\)^\{N\}\\bigr\)\.By C2, the aggregate pre\-discovery exposure tokkdiverges on any sample path wherekkremains undiscovered\. Hence

∑tPr⁡\[k​is discovered at iteration​t∣ℱt\]=∞\\sum\_\{t\}\\Pr\\\!\\left\[k\\text\{ is discovered at iteration \}t\\mid\\mathcal\{F\}\_\{t\}\\right\]=\\inftyon that event\. By the conditional Borel–Cantelli lemma \(Lévy’s extension\),kkis eventually discovered almost surely\.

Convergence for finite𝒦\\mathcal\{K\}:Since\|𝒦\|<∞\|\\mathcal\{K\}\|<\\infty, the artifact\-wise conclusion holds simultaneously for allk∈𝒦k\\in\\mathcal\{K\}by a finite union bound\. Thus every valid artifact is eventually discovered almost surely, and therefore𝒦t\+→𝒦\\mathcal\{K\}\_\{t\}^\{\+\}\\to\\mathcal\{K\}a\.s\. Monotonicity ensures that once an artifact is discovered, it remains in the genuine retained set\.

Extension to countably infinite𝒦\\mathcal\{K\}:For countably infinite𝒦\\mathcal\{K\}, the same artifact\-wise argument implies that everykkwith divergent aggregate pre\-discovery exposure is discovered almost surely\. Equivalently, coverage holds pointwise over such artifacts, although no finite\-time full\-coverage statement is possible for an infinite domain\. ∎

## Appendix DContamination Analysis Proof

###### Proof of Proposition[3](https://arxiv.org/html/2605.15219#Thmtheorem3)\.

Condition onℱt\\mathcal\{F\}\_\{t\}\. For each undiscovered valid artifactk∈K∖Kt\+k\\in K\\setminus K\_\{t\}^\{\+\}, a single generated candidate equalskkand is accepted with probabilityrt,k​Qt​\(k\)r\_\{t,k\}Q\_\{t\}\(k\)\. Since theNNcandidates and verifier decisions are conditionally independent, the probability thatkkis accepted at least once during iterationttis

1−\(1−rt,k​Qt​\(k\)\)N\.1\-\(1\-r\_\{t,k\}Q\_\{t\}\(k\)\)^\{N\}\.Therefore, by linearity of expectation,

𝔼​\[Δ​Gt∣ℱt\]=∑k∈K∖Kt\+\(1−\(1−rt,k​Qt​\(k\)\)N\)\.\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}\\left\(1\-\(1\-r\_\{t,k\}Q\_\{t\}\(k\)\)^\{N\}\\right\)\.
For invalid candidates, each draw falls in𝒳∖K\\mathcal\{X\}\\setminus Kwith probability

Ut=∑x∈𝒳∖KQt​\(x\)\.U\_\{t\}=\\sum\_\{x\\in\\mathcal\{X\}\\setminus K\}Q\_\{t\}\(x\)\.Conditional on being invalid, it is falsely accepted with probabilityδt\\delta\_\{t\}\. Hence each draw contributes an accepted invalid candidate with probabilityδt​Ut\\delta\_\{t\}U\_\{t\}, and linearity of expectation gives

𝔼​\[Δ​Bt∣ℱt\]=N​δt​Ut\.\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=N\\delta\_\{t\}U\_\{t\}\.
For the sparse\-regime expansion, use

1−\(1−rt,k​q\)N=N​rt,k​q\+O​\(N2​q2\),1\-\(1\-r\_\{t,k\}q\)^\{N\}=Nr\_\{t,k\}q\+O\(N^\{2\}q^\{2\}\),where0≤rt,k≤10\\leq r\_\{t,k\}\\leq 1, uniformly for smallN​qNq\. Substitutingq=Qt​\(k\)q=Q\_\{t\}\(k\)and summing overk∈K∖Kt\+k\\in K\\setminus K\_\{t\}^\{\+\}gives

𝔼​\[Δ​Gt∣ℱt\]=N​∑k∈K∖Kt\+rt,k​Qt​\(k\)\+O​\(N2​∑k∈K∖Kt\+Qt​\(k\)2\)\.\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=N\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}r\_\{t,k\}Q\_\{t\}\(k\)\+O\\\!\\left\(N^\{2\}\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)^\{2\}\\right\)\.Whenrt,k=rtr\_\{t,k\}=r\_\{t\}over the relevant undiscovered region, this becomes

𝔼​\[Δ​Gt∣ℱt\]=N​rt​∑k∈K∖Kt\+Qt​\(k\)\+O​\(N2​∑k∈K∖Kt\+Qt​\(k\)2\)\.\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=Nr\_\{t\}\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)\+O\\\!\\left\(N^\{2\}\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)^\{2\}\\right\)\.Since

Mtnew=∑k∈K∖Kt\+Qt​\(k\),M\_\{t\}^\{\\mathrm\{new\}\}=\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\),we obtain

𝔼​\[Δ​Gt∣ℱt\]=N​rt​Mtnew\+O​\(N2​∑k∈K∖Kt\+Qt​\(k\)2\)\.\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]=Nr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\+O\\\!\\left\(N^\{2\}\\sum\_\{k\\in K\\setminus K\_\{t\}^\{\+\}\}Q\_\{t\}\(k\)^\{2\}\\right\)\.Dividing the first\-order expressions for𝔼​\[Δ​Bt∣ℱt\]\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]and𝔼​\[Δ​Gt∣ℱt\]\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]yields

𝔼​\[Δ​Bt∣ℱt\]𝔼​\[Δ​Gt∣ℱt\]≈δt​Utrt​Mtnew\.\\frac\{\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\{\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\\approx\\frac\{\\delta\_\{t\}U\_\{t\}\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\}\.∎

#### Deduplicated invalid artifacts\.

The main text counts invalid false positives per accepted candidate\. If instead the retained set deduplicates invalid artifacts, then the expected number of distinct invalid artifacts falsely accepted at iterationttis

𝔼​\[Δ​Btdedup∣ℱt\]=∑x∈𝒳∖K\(1−\(1−δt​Qt​\(x\)\)N\)\.\\mathbb\{E\}\[\\Delta B\_\{t\}^\{\\rm dedup\}\\mid\\mathcal\{F\}\_\{t\}\]=\\sum\_\{x\\in\\mathcal\{X\}\\setminus K\}\\left\(1\-\(1\-\\delta\_\{t\}Q\_\{t\}\(x\)\)^\{N\}\\right\)\.In the sparse regime,

1−\(1−δt​Qt​\(x\)\)N=N​δt​Qt​\(x\)\+O​\(N2​Qt​\(x\)2\),1\-\(1\-\\delta\_\{t\}Q\_\{t\}\(x\)\)^\{N\}=N\\delta\_\{t\}Q\_\{t\}\(x\)\+O\(N^\{2\}Q\_\{t\}\(x\)^\{2\}\),where0≤δt≤10\\leq\\delta\_\{t\}\\leq 1\. Therefore,

𝔼​\[Δ​Btdedup∣ℱt\]=N​δt​Ut\+O​\(N2​∑x∈𝒳∖KQt​\(x\)2\)\.\\mathbb\{E\}\[\\Delta B\_\{t\}^\{\\rm dedup\}\\mid\\mathcal\{F\}\_\{t\}\]=N\\delta\_\{t\}U\_\{t\}\+O\\\!\\left\(N^\{2\}\\sum\_\{x\\in\\mathcal\{X\}\\setminus K\}Q\_\{t\}\(x\)^\{2\}\\right\)\.Thus the deduplicated and per\-candidate definitions have the same first\-order sparse\-regime limit, and the contamination ratio in the main text is unchanged to first order\.

## Appendix EVerification Cost Analysis

Section[3\.2](https://arxiv.org/html/2605.15219#S3.SS2)analyzes how imperfect verification affects the composition of newly accepted artifacts\. This appendix adds a simple cost layer to that analysis\. The goal is not to derive a universal optimal allocation rule, since such a rule would require a model of candidate ranking, verifier accuracy as a function of compute, and domain\-specific verification costs\. Instead, we record first\-order consequences of finite verification budgets and show how they interact with the contamination threshold\.

Recall that, at iterationtt, the verifier has true\-positive ratertr\_\{t\}on new valid artifacts and false\-positive rateδt\\delta\_\{t\}on invalid candidates:

rt=Pr⁡\[V​\(Ct\)=1∣Ct∈𝒦∖𝒦t\+\],δt=Pr⁡\[V​\(Ct\)=1∣Ct∈𝒳∖𝒦\]\.r\_\{t\}=\\Pr\[V\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\],\\qquad\\delta\_\{t\}=\\Pr\[V\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{X\}\\setminus\\mathcal\{K\}\]\.Letτ​\(c\)\\tau\(c\)denote the cost of verifying candidatecc, and letτ¯=𝔼​\[τ​\(Ct\)\]\\bar\{\\tau\}=\\mathbb\{E\}\[\\tau\(C\_\{t\}\)\]be the average verification cost under the current generation distribution\. For example, one may use the stylized model

τ​\(c\)=τ0​ℓ​\(c\)β,β≥1,\\tau\(c\)=\\tau\_\{0\}\\ell\(c\)^\{\\beta\},\\qquad\\beta\\geq 1,whereℓ​\(c\)\\ell\(c\)is the candidate length,τ0\\tau\_\{0\}is a base verification cost, andβ\\betacaptures possible superlinear scaling of verification effort with length\.

###### Proposition 12\(Feasible batch size under fixed verification cost\)\.

Suppose each generated candidate is verified, generation costscgenc\_\{\\mathrm\{gen\}\}per candidate, and verification costsτ¯\\bar\{\\tau\}per candidate on average\. Under a per\-iteration budgetBB, the feasible batch size is

N∗=⌊Bcgen\+τ¯⌋\.N^\{\*\}=\\left\\lfloor\\frac\{B\}\{c\_\{\\mathrm\{gen\}\}\+\\bar\{\\tau\}\}\\right\\rfloor\.In the sparse regime, the expected number of new genuine discoveries scales as

𝔼​\[\|St\|\]≈Bcgen\+τ¯​rt​Mtnew\.\\mathbb\{E\}\[\|S\_\{t\}\|\]\\approx\\frac\{B\}\{c\_\{\\mathrm\{gen\}\}\+\\bar\{\\tau\}\}\\,r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\.

*Proof sketch\.*Each candidate costscgen\+τ¯c\_\{\\mathrm\{gen\}\}\+\\bar\{\\tau\}in expectation, so the largest feasible batch size isN∗=⌊B/\(cgen\+τ¯\)⌋N^\{\*\}=\\lfloor B/\(c\_\{\\mathrm\{gen\}\}\+\\bar\{\\tau\}\)\\rfloor\. Substituting this into the sparse\-regime one\-step discovery approximation𝔼​\[\|St\|\]≈N​rt​Mtnew\\mathbb\{E\}\[\|S\_\{t\}\|\]\\approx Nr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}gives the result\. ∎

###### Corollary 13\(Verification\-dominated regime\)\.

Ifτ¯≫cgen\\bar\{\\tau\}\\gg c\_\{\\mathrm\{gen\}\}, then

N∗≈Bτ¯,𝔼​\[\|St\|\]≈Bτ¯​rt​Mtnew\.N^\{\*\}\\approx\\frac\{B\}\{\\bar\{\\tau\}\},\\qquad\\mathbb\{E\}\[\|S\_\{t\}\|\]\\approx\\frac\{B\}\{\\bar\{\\tau\}\}\\,r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\.Thus, when verification is much more expensive than generation, the one\-step discovery rate is limited primarily by verification throughput rather than by generation cost\.

### E\.1Cost\-Dependent Verification

The previous calculation treats verifier quality as fixed\. In many settings, however, additional verification compute can reduce the false\-positive rate\. To capture this tradeoff, consider the stylized model

δ​\(w\)=δ0​\(ww0\)−a,a\>0,\\delta\(w\)=\\delta\_\{0\}\\left\(\\frac\{w\}\{w\_\{0\}\}\\right\)^\{\-a\},\\qquad a\>0,wherewwis verification compute per candidate,w0w\_\{0\}is a reference compute level, andaameasures how quickly false positives decrease with additional verification effort\.

A useful first\-order proxy is the cost per reliable marginal discovery\. In the sparse regime, new genuine discoveries scale asrt​Mtnewr\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}, while false positives scale asδ​\(w\)​Ut\\delta\(w\)U\_\{t\}\. If false positives are treated as a penalty against reliable progress, one is led to objectives of the form

cgen\+wrt​Mtnew−δ​\(w\)​Ut,\\frac\{c\_\{\\mathrm\{gen\}\}\+w\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\-\\delta\(w\)U\_\{t\}\},up to problem\-dependent constants\. This proxy is intentionally stylized, but it captures the basic tension: increasingwwreduces contamination, while decreasingwwincreases the number of candidates that can be processed\.

###### Proposition 14\(Stylized verification allocation\)\.

Under the cost\-dependent false\-positive model

δ​\(w\)=δ0​\(ww0\)−a,\\delta\(w\)=\\delta\_\{0\}\\left\(\\frac\{w\}\{w\_\{0\}\}\\right\)^\{\-a\},and in the generation\-dominated approximationw≪cgenw\\ll c\_\{\\mathrm\{gen\}\}, the stationary verification effort for the first\-order proxy above scales as

w∗=\(a​δ0​Ut​w0a​cgenrt​Mtnew\)1/\(a\+1\)\.w^\{\*\}=\\left\(\\frac\{a\\,\\delta\_\{0\}\\,U\_\{t\}\\,w\_\{0\}^\{a\}\\,c\_\{\\mathrm\{gen\}\}\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\}\\right\)^\{1/\(a\+1\)\}\.

*Interpretation\.*The expression is not meant as a universal optimizer\. Rather, it shows the direction of the tradeoff\. Verification effort should increase when the invalid massUtU\_\{t\}is large, when the baseline false\-positive rateδ0\\delta\_\{0\}is large, or when the new\-valid massMtnewM\_\{t\}^\{\\mathrm\{new\}\}is small\. In particular, asMtnew→0M\_\{t\}^\{\\mathrm\{new\}\}\\to 0, the formula predicts increasing verification effort, matching the contamination\-trap intuition from Section[3\.2](https://arxiv.org/html/2605.15219#S3.SS2)\.

### E\.2Connection to the Contamination Threshold

The main text defines the marginal contamination fraction

ftmarg=𝔼​\[Δ​Bt∣ℱt\]𝔼​\[Δ​Gt∣ℱt\]\+𝔼​\[Δ​Bt∣ℱt\]\.f\_\{t\}^\{\\mathrm\{marg\}\}=\\frac\{\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\{\\mathbb\{E\}\[\\Delta G\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\+\\mathbb\{E\}\[\\Delta B\_\{t\}\\mid\\mathcal\{F\}\_\{t\}\]\}\.In the sparse regime,

ftmarg≈δt​Utrt​Mtnew\+δt​Ut\.f\_\{t\}^\{\\mathrm\{marg\}\}\\approx\\frac\{\\delta\_\{t\}U\_\{t\}\}\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\+\\delta\_\{t\}U\_\{t\}\}\.Therefore, to keepftmarg≤fcriticalf\_\{t\}^\{\\mathrm\{marg\}\}\\leq f\_\{\\mathrm\{critical\}\}, it is sufficient that

δt≤δt∗:=rt​Mtnew​fcriticalUt​\(1−fcritical\)\.\\delta\_\{t\}\\leq\\delta\_\{t\}^\{\*\}:=\\frac\{r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}f\_\{\\mathrm\{critical\}\}\}\{U\_\{t\}\(1\-f\_\{\\mathrm\{critical\}\}\)\}\.
This threshold partitions the local behavior into three regimes:

1. 1\.Safe discovery:δt<δt∗\\delta\_\{t\}<\\delta\_\{t\}^\{\*\}andMtnewM\_\{t\}^\{\\mathrm\{new\}\}is not too small\. Newly accepted artifacts are dominated by genuine discoveries\.
2. 2\.Contamination\-limited discovery:δt<δt∗\\delta\_\{t\}<\\delta\_\{t\}^\{\*\}, butMtnewM\_\{t\}^\{\\mathrm\{new\}\}is small\. Genuine discoveries are rare, so the system becomes increasingly sensitive to false positives\.
3. 3\.Local contamination collapse:δt≥δt∗\\delta\_\{t\}\\geq\\delta\_\{t\}^\{\*\}\. Invalid accepted artifacts exceed the allowed marginal contamination fraction, so additional accepted data can degrade the retained set unless verification is improved or invalid massUtU\_\{t\}is reduced\.

AsMtnew→0M\_\{t\}^\{\\mathrm\{new\}\}\\to 0, the safe thresholdδt∗→0\\delta\_\{t\}^\{\*\}\\to 0unlessUtU\_\{t\}shrinks comparably\. Thus the critical issue is not a fixed nonzero verification threshold, but the collapse of the tolerated false\-positive rate near the discovery frontier\. Verification cost matters because maintainingδt<δt∗\\delta\_\{t\}<\\delta\_\{t\}^\{\*\}may require increasing verification effort precisely when genuine discoveries are becoming hardest to find\.

## Appendix FCompute Growth and Sustainability

Section[4](https://arxiv.org/html/2605.15219#S4)derives the cumulative generation\-cost law

Rcum​\(D\)=Θ​\(cgen​Dα\)R\_\{\\mathrm\{cum\}\}\(D\)=\\Theta\(c\_\{\\mathrm\{gen\}\}D^\{\\alpha\}\)under the tail\-equivalence assumption\. This appendix discusses a simple consequence of that law: whether improvements in hardware, systems efficiency, and deployment scale can sustain a desired discovery trajectory over calendar time\. The discussion is not needed for the main NOVA guarantees; it is a scenario analysis that compares the compute demanded by a target discovery trajectory with the compute supplied by hardware and systems improvements\.

LetD​\(t\)D\(t\)be a target cumulative discovery trajectory over calendar time, and letC​\(t\)C\(t\)denote the effective generation compute available to the NOVA process by timett\. This available compute may reflect hardware efficiency, systems optimization, inference batching, deployment scale, and budget\. Verification costs are not included inC​\(t\)C\(t\); accounting for them would only make the sustainability requirement stronger\.

Theorem[6](https://arxiv.org/html/2605.15219#Thmtheorem6)immediately gives the necessary condition

C​\(t\)≳cgen​D​\(t\)α\.C\(t\)\\gtrsim c\_\{\\mathrm\{gen\}\}D\(t\)^\{\\alpha\}\.\(4\)Thus compute sustainability is a comparison between the desired discovery trajectory and the available effective compute supply\.

For illustration, suppose the effective compute supply grows exponentially over a limited horizon,

C​\(t\)=C0​2t/τ,C\(t\)=C\_\{0\}2^\{t/\\tau\},whereτ\\tauis an effective doubling time that aggregates accelerator improvements, systems optimization, deployment scale, and available budget\. Equivalently, if chip\-level efficiency and deployed compute scale contribute multiplicatively, one may write

C​\(t\)=η0​B0​2t/τchip​2t/τscale,1τ=1τchip\+1τscale\.C\(t\)=\\eta\_\{0\}B\_\{0\}2^\{t/\\tau\_\{\\mathrm\{chip\}\}\}2^\{t/\\tau\_\{\\mathrm\{scale\}\}\},\\qquad\\frac\{1\}\{\\tau\}=\\frac\{1\}\{\\tau\_\{\\mathrm\{chip\}\}\}\+\\frac\{1\}\{\\tau\_\{\\mathrm\{scale\}\}\}\.The specific values of these parameters are empirical and time\-dependent, so the model should be read as illustrative rather than predictive\.

Combining this supply model with \([4](https://arxiv.org/html/2605.15219#A6.E4)\) gives

D​\(t\)≲\(C0cgen\)1/α​2t/\(α​τ\)\.D\(t\)\\lesssim\\left\(\\frac\{C\_\{0\}\}\{c\_\{\\mathrm\{gen\}\}\}\\right\)^\{1/\\alpha\}2^\{t/\(\\alpha\\tau\)\}\.Thus exponential compute growth can support an exponentially growing discovery trajectory, but with the exponent reduced by a factor ofα\\alpha\. Largerα\\alphatherefore makes the same target trajectory harder to sustain\. Conversely, whenα\\alphais close to one, the discovery\-cost curve is closer to linear, so hardware and systems improvements translate more directly into additional discoveries\.

#### Marginal sustainability\.

The cumulative condition above concerns total compute required to reachD​\(t\)D\(t\)\. The corresponding marginal condition follows from the Section[4](https://arxiv.org/html/2605.15219#S4)scaling

d​Rcumd​D=Θ​\(cgen​Dα−1\)\.\\frac\{dR\_\{\\mathrm\{cum\}\}\}\{dD\}=\\Theta\(c\_\{\\mathrm\{gen\}\}D^\{\\alpha\-1\}\)\.Thus the incremental compute required for the next discovery grows asDα−1D^\{\\alpha\-1\}\. For any fixed maximum effective computeMmaxM\_\{\\max\}available per marginal discovery, the sustainable frontier must satisfy

cgen​Dα−1≲Mmax,c\_\{\\mathrm\{gen\}\}D^\{\\alpha\-1\}\\lesssim M\_\{\\max\},or equivalently

D≲\(Mmaxcgen\)1/\(α−1\)\.D\\lesssim\\left\(\\frac\{M\_\{\\max\}\}\{c\_\{\\mathrm\{gen\}\}\}\\right\)^\{1/\(\\alpha\-1\)\}\.This marginal threshold is highly sensitive toα\\alpha\. Whenα\\alphais close to one, the exponent1/\(α−1\)1/\(\\alpha\-1\)is large, so the marginal wall can be far away\. Whenα\\alphais larger, marginal costs rise much more quickly\.

#### Interpretation\.

The sustainability question is not simply whether compute improves over time\. It is whether effective compute supply grows fast enough relative to the desired discovery trajectory and the domain difficulty exponentα\\alpha\. Concentrated domains with largerα\\alphaexhibit sharper diminishing returns: easy discoveries are found early, and additional discoveries become increasingly expensive\. Heavier\-tailed domains withα\\alphaclose to one have more persistent long\-tail opportunities and milder marginal cost growth\. However, wheneverα\>1\\alpha\>1, marginal cost still grows withDD, so fixed compute budgets eventually face diminishing returns\.

### F\.1Illustrative Domain Regimes

Table[1](https://arxiv.org/html/2605.15219#A6.T1)illustrates how the cumulative and marginal cost laws vary with the Zipf exponent\. The assignedα\\alphavalues are heuristic placeholders; they are meant to show the qualitative dependence on tail heaviness, not to provide empirical estimates for these domains\.

Table 1:Illustrative discovery\-cost regimes under different Zipf exponents\. Theα\\alphavalues are heuristic and are intended only to show how NOVA’s scaling laws depend on the heaviness of the discovery tail\.The table highlights the role ofα\\alphaas a domain\-difficulty parameter\. Largerα\\alphameans that probability mass is concentrated on relatively few easier artifacts, so the system rapidly exhausts high\-probability discoveries and marginal costs rise sharply\. Smallerα\>1\\alpha\>1corresponds to a heavier tail: undiscovered mass decays more slowly, and marginal costs grow more mildly\. Thus the key distinction is not whether diminishing returns exist, but how quickly they appear\.

## Appendix GHuman–AI Collaboration Details

This appendix formalizes the human\-augmented NOVA loop used in Section[5](https://arxiv.org/html/2605.15219#S5)\. The goal is to isolate three ways in which human experts can amplify discovery: guidance, generation, and verification\.

###### Definition 1\(Human Expert\)\.

A human expertHHis characterized by four components:

- •a proposal distributionPHP\_\{H\}over candidate knowledge artifacts;
- •a generation rateλH\\lambda\_\{H\}, measured in candidates per unit human time;
- •a verification procedureVHV\_\{H\}, with acceptance rateρH,t\\rho\_\{H,t\}on valid human\-generated candidates;
- •a guidance functionGHG\_\{H\}, which modifies the model distribution fromQtQ\_\{t\}to a guided distributionQt′Q\_\{t\}^\{\\prime\}\.

###### Definition 2\(Augmented NOVA Loop\)\.

At iterationtt, the human\-augmented NOVA loop proceeds as follows:

1. 1\.Human guidance:the expert modifies the AI generator fromQtQ\_\{t\}toQt′=GH​\(Qt\)Q\_\{t\}^\{\\prime\}=G\_\{H\}\(Q\_\{t\}\)\.
2. 2\.AI generation:the AI generatesNAIN\_\{\\mathrm\{AI\}\}candidates fromQt′Q\_\{t\}^\{\\prime\}\.
3. 3\.Human generation:the expert generatesNH=λH​TgenN\_\{H\}=\\lambda\_\{H\}T\_\{\\mathrm\{gen\}\}candidates fromPHP\_\{H\}, whereTgenT\_\{\\mathrm\{gen\}\}is the human time allocated to direct generation\.
4. 4\.Combined verification:AI and human\-generated candidates are verified using the available formal, AI\-assisted, or human verification procedures\.
5. 5\.Accumulate and retrain:accepted candidates are added to the retained set, and the model is updated as in the standard NOVA loop\.

Recall that as defined in Section[5](https://arxiv.org/html/2605.15219#S5),

Mtnew,guided:=Qt′​\(𝒦∖𝒦t\+\),Mtnew,H:=PH​\(𝒦∖𝒦t\+\),M\_\{t\}^\{\\mathrm\{new,guided\}\}:=Q\_\{t\}^\{\\prime\}\(\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\),\\qquad M\_\{t\}^\{\\mathrm\{new\},H\}:=P\_\{H\}\(\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\}\),whereMtnew,guidedM\_\{t\}^\{\\mathrm\{new,guided\}\}is the new\-valid mass under the guided AI distribution andMtnew,HM\_\{t\}^\{\\mathrm\{new\},H\}is the new\-valid mass under the human proposal distribution\. Also

reff,t:=Pr\[Veff\(Ct\)=1∣Ct∈𝒦∖𝒦t\+,Ct∼Qt′\],r\_\{\\mathrm\{eff\},t\}:=\\Pr\[V\_\{\\mathrm\{eff\}\}\(C\_\{t\}\)=1\\mid C\_\{t\}\\in\\mathcal\{K\}\\setminus\\mathcal\{K\}\_\{t\}^\{\+\},\\,C\_\{t\}\\sim Q\_\{t\}^\{\\prime\}\],the effective true\-positive rate for new valid AI\-generated candidates after human review\.

### G\.1Amplification Decomposition

###### Proof of Theorem[7](https://arxiv.org/html/2605.15219#Thmtheorem7)\.

In the sparse regime, the autonomous baseline discovery rate is

𝔼​\[\|StAI\|\]≈NAI​rt​Mtnew\.\\mathbb\{E\}\[\|S\_\{t\}^\{\\mathrm\{AI\}\}\|\]\\approx N\_\{\\mathrm\{AI\}\}r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\.After human guidance, the AI distribution changes fromQtQ\_\{t\}toQt′Q\_\{t\}^\{\\prime\}, so the new\-valid mass changes fromMtnewM\_\{t\}^\{\\mathrm\{new\}\}toMtnew,guidedM\_\{t\}^\{\\mathrm\{new,guided\}\}\. If human review changes the effective true\-positive rate for new valid AI\-generated candidates fromrtr\_\{t\}toreff,tr\_\{\\mathrm\{eff\},t\}, then the guided AI contribution becomes

𝔼​\[\|StAI,guided\|\]≈NAI​reff,t​Mtnew,guided\.\\mathbb\{E\}\[\|S\_\{t\}^\{\\mathrm\{AI,guided\}\}\|\]\\approx N\_\{\\mathrm\{AI\}\}r\_\{\\mathrm\{eff\},t\}M\_\{t\}^\{\\mathrm\{new,guided\}\}\.The human expert also contributesNHN\_\{H\}candidates fromPHP\_\{H\}\. Neglecting duplicate discoveries between human\- and AI\-generated candidates, the expected number of accepted new valid human\-generated candidates is

𝔼​\[\|StH\|\]≈NH​ρH,t​Mtnew,H\.\\mathbb\{E\}\[\|S\_\{t\}^\{H\}\|\]\\approx N\_\{H\}\\rho\_\{H,t\}M\_\{t\}^\{\\mathrm\{new\},H\}\.Therefore,

AH=NAI​reff,t​Mtnew,guided\+NH​ρH,t​Mtnew,HNAI​rt​Mtnew\.A\_\{H\}=\\frac\{N\_\{\\mathrm\{AI\}\}r\_\{\\mathrm\{eff\},t\}M\_\{t\}^\{\\mathrm\{new,guided\}\}\+N\_\{H\}\\rho\_\{H,t\}M\_\{t\}^\{\\mathrm\{new\},H\}\}\{N\_\{\\mathrm\{AI\}\}r\_\{t\}M\_\{t\}^\{\\mathrm\{new\}\}\}\.Factoring the right\-hand side gives

AH=Mtnew,guidedMtnew⋅reff,trt⋅\(1\+NH​ρH,t​Mtnew,HNAI​reff,t​Mtnew,guided\)\.A\_\{H\}=\\frac\{M\_\{t\}^\{\\mathrm\{new,guided\}\}\}\{M\_\{t\}^\{\\mathrm\{new\}\}\}\\cdot\\frac\{r\_\{\\mathrm\{eff\},t\}\}\{r\_\{t\}\}\\cdot\\left\(1\+\\frac\{N\_\{H\}\\rho\_\{H,t\}M\_\{t\}^\{\\mathrm\{new\},H\}\}\{N\_\{\\mathrm\{AI\}\}r\_\{\\mathrm\{eff\},t\}M\_\{t\}^\{\\mathrm\{new,guided\}\}\}\\right\)\.Thus,

Aguide=Mtnew,guidedMtnew,Averify=reff,trt,Agen=1\+NH​ρH,t​Mtnew,HNAI​reff,t​Mtnew,guided,A\_\{\\mathrm\{guide\}\}=\\frac\{M\_\{t\}^\{\\mathrm\{new,guided\}\}\}\{M\_\{t\}^\{\\mathrm\{new\}\}\},\\qquad A\_\{\\mathrm\{verify\}\}=\\frac\{r\_\{\\mathrm\{eff\},t\}\}\{r\_\{t\}\},\\qquad A\_\{\\mathrm\{gen\}\}=1\+\\frac\{N\_\{H\}\\rho\_\{H,t\}M\_\{t\}^\{\\mathrm\{new\},H\}\}\{N\_\{\\mathrm\{AI\}\}r\_\{\\mathrm\{eff\},t\}M\_\{t\}^\{\\mathrm\{new,guided\}\}\},and hence

AH=Aguide⋅Averify⋅Agen\.A\_\{H\}=A\_\{\\mathrm\{guide\}\}\\cdot A\_\{\\mathrm\{verify\}\}\\cdot A\_\{\\mathrm\{gen\}\}\.The generation term is measured relative to the already guided and human\-reviewed AI discovery rate; the guidance and verification effects are therefore accounted for separately\. ∎

### G\.2Support Expansion

###### Proof of Theorem[8](https://arxiv.org/html/2605.15219#Thmtheorem8)\.

By definition, an artifact can be generated with positive probability under a distribution only if it lies in that distribution’s support\. Hence the reachable valid set under the autonomous generatorQtQ\_\{t\}is𝒦∩supp⁡\(Qt\)\\mathcal\{K\}\\cap\\operatorname\{supp\}\(Q\_\{t\}\), while after human guidance it becomes𝒦∩supp⁡\(Qt′\)\\mathcal\{K\}\\cap\\operatorname\{supp\}\(Q\_\{t\}^\{\\prime\}\)\. If

𝒦∩\(supp⁡\(Qt′\)∖supp⁡\(Qt\)\)≠∅,\\mathcal\{K\}\\cap\\bigl\(\\operatorname\{supp\}\(Q\_\{t\}^\{\\prime\}\)\\setminus\\operatorname\{supp\}\(Q\_\{t\}\)\\bigr\)\\neq\\emptyset,then there exists a valid artifact that is reachable underQt′Q\_\{t\}^\{\\prime\}but not underQtQ\_\{t\}\. Therefore,

𝒦∩supp⁡\(Qt\)⊊𝒦∩supp⁡\(Qt′\),\\mathcal\{K\}\\cap\\operatorname\{supp\}\(Q\_\{t\}\)\\subsetneq\\mathcal\{K\}\\cap\\operatorname\{supp\}\(Q\_\{t\}^\{\\prime\}\),so human guidance strictly expands the reachable valid set\. In particular, ifQt′Q\_\{t\}^\{\\prime\}assigns non\-negligible mass to valid artifacts outside the previous effective support, those artifacts are no longer blocked by the autonomous exploration barrier\. ∎

### G\.3Human Effort Allocation

###### Proposition 15\(Human Effort Allocation Principle\)\.

Given limited human timeTHT\_\{H\}, the appropriate use of human effort depends on the system’s bottleneck:

1. 1\.Far from the exploration barrier\(Mtnew\(M\_\{t\}^\{\\mathrm\{new\}\}large\): prioritize human generation, because many valid artifacts remain reachable\.
2. 2\.Near the exploration barrier\(Mtnew\(M\_\{t\}^\{\\mathrm\{new\}\}small\): prioritize guidance, because redirecting or expanding support is more valuable than generating more samples from a depleted region\.
3. 3\.When formal verification is unavailable\(rt≈0\)\(r\_\{t\}\\approx 0\): prioritize human verification, because valid artifacts may be generated but cannot be reliably accumulated\.

## Appendix HInfinite Knowledge Spaces

The main convergence theorem \(Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)\) assumes\|𝒦\|<∞\|\\mathcal\{K\}\|<\\infty, so artifact\-wise discovery implies eventual coverage of the whole knowledge domain\. Many natural domains, however, are effectively infinite: valid theorems, algorithms, programs, designs, and scientific hypotheses do not come from a small finite catalog\. This appendix records how the NOVA conclusions should be interpreted in countably infinite spaces\.

###### Definition 3\(Infinite NOVA\)\.

An infinite NOVA system operates over a countably infinite knowledge space𝒦=\{k1,k2,…\}\\mathcal\{K\}=\\\{k\_\{1\},k\_\{2\},\\ldots\\\}, with an ideal difficulty distributionPPsatisfying∑j≥1P​\(kj\)=1\\sum\_\{j\\geq 1\}P\(k\_\{j\}\)=1\.

In this setting, the almost\-sure coverage conclusion of Theorem[1](https://arxiv.org/html/2605.15219#Thmtheorem1)must be interpreted artifact\-wise\. If each fixedk∈𝒦k\\in\\mathcal\{K\}receives divergent aggregate pre\-discovery exposure and nondegenerate acceptance, then each suchkkis eventually discovered almost surely\. However, no finite time can cover an infinite𝒦\\mathcal\{K\}, and global coverage must be replaced by growth of the discovered set and decay of the remaining mass\.

###### Corollary 16\(Infinite knowledge horizon\)\.

In a countably infinite knowledge space, complete discovery is asymptotic rather than finite\-time: NOVA can continue expanding𝒦t\+\\mathcal\{K\}\_\{t\}^\{\+\}, but cannot exhaust all of𝒦\\mathcal\{K\}in finite computation\. Under Zipf tail equivalence withα\>1\\alpha\>1, the same occupancy scaling gives sublinear growth in distinct discoveries and a decaying undiscovered mass, so progress continues but with diminishing returns\.

Similar Articles

AI research tools are still too eager to turn public signals into certainty

Reddit r/artificial

The author critiques AI research tools for overconfidence in weak signals, praising Komo AI's rapid discovery and source-attached summaries but highlighting the need for better uncertainty and contradiction handling. They describe a workflow that splits discovery, verification, and structured checking across multiple AI tools.

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Hugging Face Daily Papers

Discovery Foundation Models are proposed as general-purpose systems for enabling open-ended scientific discovery through iterative problem formulation, hypothesis testing, and evidence-based revision across dry and wet lab settings. The paper introduces a framework with capabilities like problem discovery and continual improvement, instantiated with systems like Zetema and GALILEO.