Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory

arXiv cs.LG Papers

Summary

This paper critiques the pessimistic meta-inductive argument against scientific realism by leveraging convergence concepts from frequentist statistics and machine learning, showing that ordinary induction achieves convergence while meta-induction fails.

arXiv:2608.17213v1 Announce Type: new Abstract: This paper challenges the pessimistic meta-inductive argument against scientific realism by undermining its inductive step rather than its historical premise. Although related challenges already exist, I develop a new one. Drawing on a general epistemology of scientific inference developed in frequentist statistics, machine learning, and formal epistemology, I evaluate induction in terms of convergence to the truth. I argue that ordinary enumerative induction can achieve everywhere convergence, whereas meta-induction fails even to achieve almost everywhere convergence. Indeed, in the problem context where meta-induction arises, the failure is deeper: no inference method whatsoever achieves almost everywhere convergence.
Original Article
View Cached Full Text

Cached at: 08/19/26, 10:24 AM

# Lessons from Frequentist Statistics and Machine Learning Theory
Source: [https://arxiv.org/html/2608.17213](https://arxiv.org/html/2608.17213)
## Pessimistic Meta\-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory

Hanti LinThanks:AI use statement: AI tools were used in the preparation of this paper only for English editing, not for generating any content or images\.Affiliation:University of California, DavisEmail:[ika@ucdavis\.edu](mailto:[email protected])

###### Abstract

This paper challenges the pessimistic meta\-inductive argument against scientific realism by undermining its inductive step rather than its historical premise\. Although related challenges already exist, I develop a new one\. Drawing on a general epistemology of scientific inference developed in frequentist statistics, machine learning, and formal epistemology, I evaluate induction in terms of convergence to the truth\. I argue that ordinary enumerative induction can achieve everywhere convergence, whereas meta\-induction fails even to achieve almost everywhere convergence\. Indeed, in the problem context where meta\-induction arises, the failure is deeper: no inference method whatsoever achieves almost everywhere convergence\.

## 1Introduction

The pessimistic meta\-inductive argument against scientific realism is meant to show, roughly, that since all or most past scientific theories turned out to be not even approximately true, our current scientific theory is no exception \(Laudan 1981, Putnam 1978\)\.111Also see Wray \(2015\) for a reconstruction of different types of the meta\-inductive argument\.Many replies challenge the historical premise that most past theories have indeed been shown to be not even approximately true; see, for example, Devitt \(1984\), Kitcher \(1993\), Psillos \(1994\), Worrall \(1994\), and Leplin \(1997\)\. Here, however, I pursue a complementary strategy: undermining the inductive step itself\. Relatively few works do so\. Notable examples include Lewis \(2001\) and Magnus and Callender \(2004\), who appeal to considerations related to the base\-rate fallacy to explain why the meta\-inductive step does not amount to justified induction\.

Without criticizing those allies of mine, I propose a new way to challenge the inductive step—one based on ageneralapproach to the epistemology of scientific inference, now thriving in many branches of frequentist statistics, especially nonparametric statistics, and central to the theoretical foundations of machine learning\. These are, after all, fields devoted to the systematic study of scientific inference—by scientists and for scientists\. The core idea is to evaluate inference methods by their properties ofconvergence to the truth—whether in ordinary scientific reasoning, statistical procedures, or machine learning algorithms, and whether or not the data are generated stochastically\. Although this idea has a long history, traceable to C\. S\. Peirce \(1902\), its core theses have only recently been articulated by philosophers with clarity and generality; see, for example, Lin \(2025\)\. My aim is to use this framework to explain why, despite their superficial syntactic similarity, ordinary induction and meta\-induction differ sharply:

- •In the context of ordinary induction—for example, when we ask whether the ravens to be observed in the future will all be black—enumerative induction is justified because it achieves the relatively strong standard ofeverywhere convergence, i\.e\., convergence to the true answer ateverypossible state of the world on the table\.
- •In the context of meta\-induction—for example, when we ask whether the scientific theories to be tested in the future will all be refuted by data—enumerative induction is not justified because it fails to achieve even the weaker standard ofalmost everywhere convergence, i\.e\., convergence to the true answer atalmost allpossible states of the world on the table\.

The relevant notion of “almost all” will be defined rigorously below, though it has already been used in machine learning to evaluate algorithms for causal structure learning \(Lin and Zhang 2020\)\. In fact, the case against meta\-induction is stronger still\. If a problem context is defined by a set of competing hypotheses as potential answers to the question posed, then in the context in which meta\-induction is applied, no inference method whatsoever achieves almost everywhere convergence—this is the main mathematical result of the present paper\.

I begin with a simple mathematical model of ordinary induction and meta\-induction, open to future refinement \(Section[2](https://arxiv.org/html/2608.17213#S2)\)\. Though somewhat toy\-like, it captures an appealing version of meta\-induction and is therefore worth undermining on its own terms\. Its simplicity also helps reveal the underlying structure of the problem and suggests how the result might be extended\. Along the way, I explain the epistemological ideas in use, especially the concept of convergence to the truth, while avoiding some of its more traditional and misleading formulations \(Section[3](https://arxiv.org/html/2608.17213#S3)\)\. I conclude by emphasizing that this epistemology is a very general framework for scientific inference—one developed largely by scientists, with help from a few philosophers, and broad enough to encompass much of frequentist statistics and machine learning theory \(Section[4](https://arxiv.org/html/2608.17213#S4)\)\.

## 2Setting

Consider the data sequence:

001 000001 000000000,\\mathtt\{001\}\\;\\mathtt\{000001\}\\;\\mathtt\{000000000\}\\,,where each𝟶\\mathtt\{0\}means nothing bad—our current scientific theory passes a new test—and each𝟷\\mathtt\{1\}means something bad—our current theory fails a test and must be replaced\. The first block,𝟶𝟶𝟷\\mathtt\{001\}, shows that our first theory failed quickly, after just three tests\. We then moved to a second theory, which also failed but survived longer, as represented by the longer block𝟶𝟶𝟶𝟶𝟶𝟷\\mathtt\{000001\}\. Our third theory, by contrast, is still doing well: the third block has not \(yet\) ended with𝟷\\mathtt\{1\}\.

More generally, letTnT\_\{n\}be thenn\-th theory we would develop to maturity if we were diligent enough \(and lived long enough\)\. Thenn\-th occurrence of𝟷\\mathtt\{1\}in the data sequence marks the downfall ofTnT\_\{n\}and its replacement byTn\+1T\_\{n\+1\}\.

Now imagine that the following data sequence represents our current state of evidence:

000001 00000001⋯⋯𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟷𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶𝟶⏞Tn\+1has passed many tests\.\\mathtt\{000001\}\\;\\mathtt\{00000001\}\\cdots\\cdots\\mathtt\{0000000001\}\\;\\overbrace\{\\mathtt\{00000000000000000000000000000000000\}\}^\{\\text\{$T\_\{n\+1\}$ has passed many tests\.\}\}\\,​↑\\uparrow↑\\uparrow​↑\\uparrowT1T\_\{1\}fails\.T2T\_\{2\}fails\. ​TnT\_\{n\}fails, wherennis large\.

Confronted with this data sequence, we might be tempted to draw one of two types of inductive inference\. The first is:

Ordinary Induction\.Consider the annotationsabovethe data sequence: the current theoryTn\+1T\_\{n\+1\}has undergone many tests and survived them all\. We infer—inductively and ordinarily—thatTn\+1T\_\{n\+1\}would never suffer a downfall, no matter how many further tests it faces\.

The second is suggested by Putnam \(1978\) and Laudan \(1981\):

Pessimistic Meta\-Induction\.Consider the annotationsbelowthe data sequence\. The firstnntheories,T1,…,TnT\_\{1\},\\ldots,T\_\{n\}, have failed, andnnis very large\. We infer—inductively and pessimistically—that the current theoryTn\+1T\_\{n\+1\}, along with all its successors, would also fail if tested sufficiently\.

Now, which type of induction should we adopt? This is not easy to answer, since both instantiate the same syntactic template:

Syntactic Template of Induction *Premise 1*\. Many*F*s have been observed\. *Premise 2*\. All of them are*G*s\. —————————————————— *Conclusion*\. So, all*F*s are*G*s\.

To break the symmetry between ordinary induction and pessimistic meta\-induction, we must look beyond their syntactic form\. If I am right, there is an important asymmetry:

The Disparity Thesis\(To Be Refined and Defended Below\)\.•In the context where the question is whether the current theory would fail if sufficiently tested, the ordinary inductive method makes evidence an indicator of the true answer\.•But in the context where the question is whether every scientific theory in the pipeline would fail if sufficiently tested,everynon\-deductive method—including the pessimistic meta\-inductive method—failsto make evidence an indicator of the true answer\.

Certain concepts must be defined rigorously before this disparity thesis can be refined into a theorem\. I will appeal to the concept of convergence and to a topological conception of “almost everywhere”\.

But even setting aside the technical details, the disparity should not be surprising at an intuitive level\. Let infinite binary data sequences represent possible states of the world\. These states correspond, more or less, to real numbers in the unit interval, expressed in their binary expansions\. Let us ignore the minor redundancy of binary representation \(for example, both𝟷​𝟶¯\\mathtt\{1\\bar\{0\}\}and𝟶​𝟷¯\\mathtt\{0\\bar\{1\}\}represent the same real number, namely0\.50\.5\)\. In one problem context, the question is whether the current theory would fail\. Equivalently, we ask whether the actual state of the world is a particular real numberrrin the unit interval, whererris the available data sequence concatenated with an infinite block of𝟶\\mathtt\{0\}s\. In a different context, the question is whether every scientific theory in the pipeline would fail if sufficiently tested\. The true answer is “yes” iff there are infinitely many11s in the actual state of the world, that is, iff the actual state is an irrational number in the unit interval\.

So the two problem contexts are very different\. One asks whether the actual real number is a specific one; the other asks whether it is irrational\. As more digits of data arrive, we get more precise information about where the actual state lies in the continuum of the unit interval—an ever\-shrinking nonempty interval\[an,bn\]\[a\_\{n\},b\_\{n\}\]of real numbers\. Thus, in the first case, more data can still refute at least one of the two candidate answers\. In the second case, however, no possible data can refute either of the two candidate answers—the two potential answers, “Rational Number” vs\. “Irrational Number”, are too closely intertwined on the real line\. That is, any nonempty interval overlaps both\. So, the latter question is intuitively harder\. The formal development below vindicates this intuition\.

## 3Formal Development

It is time to turn the disparity thesis stated above into a theorem\.

### 3\.1Epistemic Scenarios and Convergence

Aninference methodis a mathematical function that takes a finite data sequence as evidential input and outputs a proposition as a conclusion, subject to revision as more data arrive\. A non\-deductive method is justified only if it can, in some sense, serve as a good indicator of truth—and the challenge is to make this idea precise\. An initial thought is that a good indicator of truth should point to atruthacross a certainrangeofepistemic scenarios\. But what truth? What scenarios? And what range? I will explain in turn\.

Let us begin with truth\. In any problem context, a question is posed, whose potential answers serve as the competing hypotheses\. The truth pursued is the \(unknown\) true answer to that question in context\.

Now turn to scenarios\. Anepistemic scenariocan be modeled by an ordered pair\(s,n\)\(s,n\)\. The first componentssis astate of the world—a coarse\-grained possibility that determines the truth value of each competing hypothesis in context\. The second componentnnis a*n*umber \(a positive integer\) representing an amount of evidence\. Accordingly,\(s,n\)\(s,n\)is the epistemic scenario one would be in if the actual state of the world weressand the amount of available evidence werenn\.

Then imagine a particular scenario\(s,n\)\(s,n\)as a point in a two\-dimensional space, wheressis theXX\-coordinate andnntheYY\-coordinate\. TheXX\-axis consists of possible states of the world—specifically those compatible with thebackground assumptionsof the problem context\. TheYY\-axis consists of positive integers\.

Different problem contexts may require different mathematical modelings of states of the world\. For the inductive problems at hand, a particularly simple definition suffices\. Let a state of the worldssbe an infinite binary sequence\. Here is an example:

s∗=010101⋯\(repeating the pattern of01\)⋯s^\{\*\}\\;=\\;\\mathtt\{010101\}\\cdots\\text\{\(repeating the pattern of \{01\}\)\}\\cdotsThe state is formalized as an infinite sequence, but this is not meant to represent a state in which we are immortal and endlessly accumulate data\. Rather, the infinite sequence is a convenient device for encoding counterfactuals: if the amount of evidencewerenn, the evidencewouldbe the initial segment of that sequence of lengthnn\. In this article, the states of the worldon the tableare all the infinite binary sequences—we make no background assumptions that rule out any sequence\. \(In other contexts of inquiry, we might have strong background assumptions, which leave few states of the world on the table\.\)

Consider, for example, this epistemic scenario:\(s∗,4\)\(s^\{\*\},4\), wheres∗=𝟶𝟷𝟶𝟷𝟶𝟷⋯s^\{\*\}=\\mathtt\{010101\}\\cdots; here the available evidence is𝟶𝟷𝟶𝟷\\mathtt\{0101\}\(four data points\)\. HereT2T\_\{2\}has just failed its second test, and worse, every theory would fail if sufficiently tested—specifically, if tested at least twice\. If the question posed is whether every theory would fail if sufficiently tested, then the true answer in states∗s^\{\*\}is𝚈𝚎𝚜\\mathtt\{Yes\}, and thus an inference methodMMoutputs the truth at this scenario\(s∗,4\)\(s^\{\*\},4\)iffM⁡\(𝟶𝟷𝟶𝟷\)=𝚈𝚎𝚜M\(\\mathtt\{0101\}\)=\\mathtt\{Yes\}\.

Now, what could count as a good indicator of truth? This is not easy to state precisely\. But a guiding intuition is that a non\-deductive method counts as a good indicator of truth in a problem context only if it outputs the true answer at each of a “wide range” of epistemic scenarios\. At a minimum, such a range must include epistemic scenarios\(s,n\)\(s,n\)with very largenn, i\.e\., with very large amounts of evidence\. The idea is that a good indicator of truth must output the truth at least in evidentially favorable scenarios\(s,n\)\(s,n\)—with very largenn—and, ideally, at many statessson the table, if not all\.

One tentative way to picture this idea—before formalizing it—is as follows: a non\-deductive methodMMcounts as good only if we can draw a bar across the two\-dimensional plane of epistemic scenarios, cutting through every vertical line and dividing the plane into upper and lower regions, such thatMMoutputs the truth across all scenarios in the upper region\. This can be formalized as follows:

Definition \(Everywhere Convergence to the Truth\)\.An inference methodMMis said to achieve the standard ofeverywhere convergence to the truthin a problem contextcciff, in that contextcc,for each statessof the worldon the table\(i\.e\. on theXX\-axis\), there exists a finite amount of evidenceNN\(on theYY\-axis\) such that, MMoutputs the trueanswerat epistemic scenario\(s,n\)\(s,n\)for eachn≥Nn\\geq N\.

The two underlines serve as a reminder of context\-sensitivity: what counts as a state of the world on the table is one compatible with the background assumptions in the problem contextccat hand, and what counts as a potential answer depends on the question posed in the problem contextccat hand\.

Housekeeping\.The evaluative standard just defined is not particularly high; it is often calledpointwise convergence\. A higher standard is obtained by swapping the first two quantifiers—‘for eachss’ and ‘there exists anNN’—yieldinguniform convergence to the truth\. Intermediate standards can also be defined\. It is natural to require that a non\-deductive inference method be justified only if it achieves the highest achievable standard, and I will adopt this requirement later\. For now, however, I focus on a minimum qualification for an indicator of truth, so the bar is not raised too high\. Indeed, I think the standard just defined—everywhere convergence—is still too strong to serve as a minimum qualification\. To define a genuine minimum, the quantifier ‘for each state’ must be weakened to ‘for almost all states’ in a rigorous sense\. Nevertheless, I begin with everywhere convergence to sketch the big picture before refining it\.

It is important to keep in mind that we are not assessing inference methods as in decision theory\. We are not choosing an inference method as a course of action based on its possible outcomes in the remote future, beyond our lifespans\. Such a decision\-theoretic approach would be pointless, since in the long run we are all dead\.

What we do here is instead a form ofmodal epistemology\. We seek a minimum qualification—a necessary condition—for a non\-deductive method to be justified, by evaluating its truth\-seeking performanceacross a range of epistemic scenarios, including those in which the evidence is extremely favorable\. Anything that counts as an indicator of truth should, if possible, point to the truthat leastin such favorable cases\. Accordingly, everywhere convergence, or variants of it, will be used to state a necessary condition for a non\-deductive method to be justified\.

Let me reiterate: here we are doing modal epistemology, not decision theory\. The use of convergence in epistemology is not new\. Reichenbach \(1938\) called it apragmaticapproach, and Schulte \(1999\) ameans\-endsepistemology\. Both labels are misleading, as they suggest a decision\-theoretic approach\. If this were decision theory, it would be a bad one\. Instead, it is modal epistemology: whether an inference method is justified depends on its truth\-seeking performanceacross a certain range of possible scenarios\.

### 3\.2A Sketch of the Main Result

The ordinary inductive problem and the meta\-inductive problem are two very different contexts\. The highest achievable standards in these contexts can be illustrated by the following rough diagram, though more precise statements will be developed as we proceed:

Disparity Theorem \(An Informal Sketch\)\(Strictly Stronger Conditions\)⋮\\qquad\\quad\\vdots←\\leftarrowthe highest achievable in the context⋮\\qquad\\quad\\vdotsof the ordinary inductive problemConvergence to the TruthEverywhere\|\\qquad\\quad\|Convergence to the Truth≈\\approxthe minimum qualificationQ∗Q^\{\*\}Almost Everywhere⋮\\qquad\\quad\\vdots←\\leftarrowthe highest achievable in the context⋮\\qquad\\quad\\vdotsof the meta\-inductive problem\(Strictly Weaker Conditions\)

Again, an intuitively appealing conception of “almost everywhere” will be defined rigorously below\. What matters for now is that this diagram displays a hierarchy of standards for assessing inference methods, using various modes of convergence as high or low standards for a good indicator of truth\. It also indicates where a problem contextcc“lies” along the hierarchy\. This is marked by arrows ‘←\\leftarrow’\. Each arrow corresponds to a mathematical theorem indicating the highest standard achievable in a given context\. Note thevertical differencebetween the two arrows—between the highest achievable standards in the two contexts\. This difference is as stark as night and day if the minimum qualificationQ∗Q^\{\*\}for a justified non\-deductive inference method falls between them\.

This choice ofQ∗Q^\{\*\}—the minimum qualification for justified inference—is notad hoc\. For nonparametric regression and model selection in statistics, and for supervised learning in machine learning, many algorithms are justified by showing that they achieve \(a stochastic version of\) everywhere convergence to the truth, where the truth is the true answer to the question posed in context\. More on this in the final section\.

So, here is my reply to Laudan\. It is mistaken to assume that if ordinary induction is justified, then meta\-induction must be as well, merely because of their syntactic similarity\. Context matters: ordinary induction is justified in its own context, while meta\-induction isnotjustified in its own context\. More specifically,even ifLaudan is right that most empirically successful past theories have later been shown to fail \(being not even approximately true\), andeven ifwe strengthen this premise of Laudan’s meta\-induction by replacing “most” with “all”, his non\-deductive inference is still not justified in its context—because every non\-deductive inference is doomed to achieve only a very low standard and is thus unjustified in that context\.

### 3\.3Defining “Almost Everywhere”

It is time to define “almost everywhere” rigorously\.

Think about possible states of the world as points in a space\. Consider the following set, wheresns\_\{n\}is a state identical to the actual one except that the Planck constant in the Schrödinger equation differs from the actual value by1/n\\nicefrac\{\{1\}\}\{\{n\}\}\(in units of your choice\):

S=\{sn:n​is a positive integer\}\.S=\\big\\\{s\_\{n\}:n\\text\{ is a positive integer\}\\big\\\}\\,\.Arguably, this setSSof states comes arbitrarily close to the actual state of the world\.

Let us return to our setting of possible states as infinite binary sequences\. What would be a natural relation of arbitrary closeness for such sequences? Consider the set consisting of the following binary sequences:

011111⋯\\cdots001111⋯\\cdots000111⋯\\cdots000011⋯\\cdots⋮\\vdots\\qquadThis set seems to contain enough initial segments𝟶,𝟶𝟶,𝟶𝟶𝟶,…\\mathtt\{0\},\\mathtt\{00\},\\mathtt\{000\},\\ldotsto approximate the following binary sequence as closely as we wish:

More generally:

Definition \(Cantor Space Topology\)\.On any setXXof infinite binary sequences, theCantor space topologyprovides a natural relation of arbitrary closeness, defined as follows: a subsetS⊆XS\\subseteq Xgets arbitrarily close to a points=e1e2⋯∈Xs=e\_\{1\}e\_\{2\}\\cdots\\in Xiff every initial segmente1⋯ene\_\{1\}\\cdots e\_\{n\}ofssappears as the initial segment of some sequence inSS\.

Here, the primitive topological concept in use is not that of open sets, but an equivalent one—that of arbitrary closeness \(Arkhangel’skii and Fedorchuk 1990\)\. The Cantor space topology will be the one in use for the rest of this paper\.

Using arbitrary closeness as the core concept, it is very easy to define a geometric conception of “almost everywhere” \(much easier than the standard textbook treatment\):

Definition \(Almost Everywhere, Almost All\)\.LetXXbe a set of states of the world\. A set is said to coveralmost everywhereinXXiff it is big enough to include a subsetSSofXXthat stands in the following asymmetric relations to its complementS¯\\overline\{S\}withinXX:1\.SScomes arbitrarily close to every point inS¯\\overline\{S\};2\.S¯\\overline\{S\}comes arbitrarily close to no point inSS\.A property is said to apply toalmost allelements ofXXiff it applies to every element of a set that covers almost everywhere inXX\.

This definition should look intuitively appealing\. Given clauses 1 and 2, the setSSdoes seem to cover almost everywhere\. \(In fact, clause 1 says thatSSis dense, and clause 2 says thatSSis open\.\) IfSScovers almost everywhere, then so do its supersets\. One more definition:222See Belot \(2013\) and Lin \(2022\) for existing epistemological applications of almost\-everywhere convergence and related concepts\.

Definition \(Almost Everywhere Convergence to the Truth\)\.Within a problem contextcc, an inference methodMMis said to converge to the truth almost everywhere iff, for each competing hypothesishhas a potential answer to the question posed in contextcc,MMconverges to the truth at almost allhh\-states, where•anhh\-stateis a state on the table at whichhhis true,•MM’sconvergence to the truthat a statessmeans that there exists a positive integerNNsuch thatMMoutputs the true answer at epistemic scenario\(s,n\)\(s,n\)for eachn≥Nn\\geq N\.

Then we have the main result:

Disparity Theorem \(Official Version\)\.In the context of the ordinary inductive problem, there exists an inference method that achieves the standard of everywhere convergence to the truth\. In the context of the meta\-inductive problem, there exists no inference method that achieves even the lower standard of almost everywhere convergence to the truth\.

This refines the disparity thesis stated above in response to the pessimistic argument\. See the appendix for a proof\.

## 4Closing: Toward a General Account of Scientific Inference

This paper assumes an epistemological view calledachievabilism:

Two Principles of Achievabilism\(1\)There exist justified non\-deductive methods in a problem contextccif, and only if, the minimum qualificationQ∗Q^\{\*\}for a non\-deductive indicator of truth is achievable in that contextcc—pending a specification of the minimum qualificationQ∗Q^\{\*\}\.\(2\)A non\-deductive methodMMis justified in a problem contextcconly ifMMachieves not merelyQ∗Q^\{\*\}but also the highest standard achievable in that contextccfor a good non\-deductive indicator of truth \(provided that such a highest achievable standard exists uniquely incc\)—pending a specification of the correct hierarchy of standards\.

These two principles provide only a framework, with two moving parts marked by the ‘pending’ clauses\. They leave open both the minimum qualification and the correct hierarchy of standards\.

Now, if the minimum qualificationQ∗Q^\{\*\}for justified non\-deductive inference is everywhere or almost everywhere convergence to the truth, then the Disparity Theorem of this paper suggests the following\. In using meta\-induction to argue against scientific realism, Laudan aimed to place us in a skeptical context where scientific anti\-realism appears plausible\. But he unintentionally created a far more skeptical context, one in which all non\-deductive methods are unjustified—including his meta\-induction\.

The first achievabilist principle \(1\) is, as far as I know, new\. The second achievabilist principle \(2\) is not: variants were recently stated by Lin \(2022, 2025\), though those formulations mistakenly omit reference to the minimum qualificationQ∗Q^\{\*\}\. Lin \(2025\) traces the idea of achievabilism to Putnam \(1965\), but an earlier version already appears in Neyman and Pearson \(1936\)\. They distinguish two evaluative standards for hypothesis\-testing procedures and argue that, in the so\-called two\-sided problem, only the lower standard should be applied because the higher one is unachievable, though the higher standard should still be applied when achievable\.

Convergence as a source of evaluative standards is employed across many scientific fields concerned with scientific inference\. Begin with frequentist point estimation, which asks: what is the value of a quantity of interest? A stochastic version of everywhere convergence to the true answer—estimation consistency—has long been regarded, under Fisher’s influence, as a basic requirement for a good estimator \(Fisher 1925\)\. See Lehmann and Casella \(2006\) for a textbook presentation\.

In frequentist nonparametric regression, the question becomes: which curvey=f⁡\(x\)y=f\(x\)on theX​YXY\-plane, within a function class𝒞\\mathscr\{C\}, has the highest predictive accuracy? Here too, a stochastic version of everywhere convergence to the true answer—regression consistency—is treated as a basic requirement of a good method\. This extends point estimation naturally: the target is now a point in a much larger space, namely a curve in a large function class\. See Györfi et al\. \(2002\) for a textbook presentation\.

Now modify theX​YXY\-plane: letXXbe a class of images andYYthe set of two categories “Yes, it is an image of a cat” and “No, it is not\.” Then nonparametric regression becomes classification in machine learning\. The question remains structurally the same: which “curve”y=f⁡\(x\)y=f\(x\)on this “X​YXY\-plane”, within a function class𝒞\\mathscr\{C\}, has the highest predictive accuracy? Again, a stochastic version of everywhere convergence to the true answer is often taken as a basic requirement of a good learning algorithm\. This is a pointwise mode of convergence; the stronger uniform mode corresponds to the well\-known criterion of PAC learning\. See Shalev\-Shwartz and Ben\-David \(2014\) for a textbook presentation\.

In the non\-stochastic setting adopted in this paper, everywhere convergence to the truth has long been studied in machine learning theory, beginning with Putnam \(1965\) and Gold \(1967\)\. See Case and Jain \(2017\) for a survey\.

In another area of machine learning, causal discovery, the question is: which causal model on the table is true, assuming one of them is? Here too, modes of convergence to the truth are among the standard criteria for evaluating learning methods \(Spirtes et al\. 2000\)\. Everywhere convergence to the truth is calledmodel\-selection consistency, and almost everywhere convergence has also been studied and applied in recent years \(Lin and Zhang 2020\)\.

So my appeal to convergence in replying to the pessimistic meta\-inductive argument is not a parochial use of an arcane mathematical tool\. Instead, it seeks to learn from an epistemological tradition that traces back to Peirce \(1902\), has taken firm root in the sciences, and is now emerging as a general account of scientific inference\.

## Appendix: Proof of the Disparity Theorem

The first part, concerning the ordinary inductive problem, is elementary in formal learning theory; a pictorial proof is given in Lin \(2025\)\.

For the second part, the key ingredients are the Baire Category Theorem and Belot’s \(2013\) main theorem, which he states for a problem context isomorphic to the meta\-inductive one\. Belot’s result implies that any open\-minded inference method converges to the truth only on a meager set\. Although Belot formulates the theorem for Bayesian methods, the proof extends straightforwardly to the broader class of inference methods considered here, including those with qualitative outputs\.

Now suppose, forreductio, that almost\-everywhere convergence to the truth is achievable, by some inference methodMM\. By the first clause in the definition of “almost everywhere”,MMmust converge to the truth on a dense subset of the set ofhh\-states, for each hypothesishh\. But denseness for both hypotheses immediately yields Belot’s open\-mindedness condition\. So, by Belot’s theorem,MMconverges to the truth only on a meager domain\. This contradicts the Baire Category Theorem, since in Cantor space no meager set can cover almost everywhere\. Therefore, almost\-everywhere convergence is not achievable\. Q\.E\.D\.

## References

Arkhangel’skii, A\. V\., & Fedorchuk, V\. V\. \(1990\)\. The basic concepts and constructions of general topology\. In A\. V\. Arkhangel’skii & L\. S\. Pontryagin \(Eds\.\),General topology I: Basic concepts and constructions\. Dimension theory\(pp\. 1–55\)\. Springer\-Verlag\.

Belot, G\. \(2013\)\. Bayesian orgulity\.Philosophy of Science, 80\(4\), 483–503\.

Case, J\., & Jain, S\. \(2017\)\. Connections between inductive inference and machine learning\. In C\. Sammut & G\. I\. Webb \(Eds\.\),Encyclopedia of machine learning and data mining\(pp\. 261–272\)\. Springer\.

Devitt, M\. \(1984\)\.Realism and truth\. Princeton University Press\.

Fisher, R\. A\. \(1925\)\.Statistical methods for research workers\. Oliver & Boyd\.

Gold, E\. M\. \(1967\)\. Language identification in the limit\.Information and Control, 10\(5\), 447–474\.

Györfi, L\., Kohler, M\., Krzyżak, A\., & Walk, H\. \(2002\)\.A distribution\-free theory of nonparametric regression\. Springer\.

Kitcher, P\. \(1993\)\.The advancement of science\. Oxford University Press\.

Laudan, L\. \(1981\)\. A confutation of convergent realism\.Philosophy of Science, 48\(1\), 19–49\.

Lehmann, E\. L\., & Casella, G\. \(2006\)\.Theory of point estimation\. Springer Science & Business Media\.

Leplin, J\. \(1997\)\.A novel defense of scientific realism\. Oxford University Press\.

Lewis, P\. J\. \(2001\)\. Why the pessimistic induction is a fallacy\.Synthese, 129\(3\), 371–380\.

Lin, H\. \(2022\)\. Modes of convergence to the truth: Steps toward a better epistemology of induction\.The Review of Symbolic Logic, 15\(2\), 277–310\.

Lin, H\. \(2025\)\. Convergence to the truth\. In K\. Sylvan, E\. Sosa, J\. Dancy, & M\. Steup \(Eds\.\),The Blackwell companion to epistemology\(3rd ed\.\)\. Wiley Blackwell\.

Lin, H\., & Zhang, J\. \(2020\)\. On learning causal structures from non\-experimental data without any faithfulness assumption\.Proceedings of Machine Learning Research, 117, 554–582\.

Lyons, T\. D\. \(2002\)\. Scientific realism and the pessimistic meta\-modus tollens\. In S\. Clarke & T\. D\. Lyons \(Eds\.\),Recent themes in the philosophy of science: Scientific realism and commonsense\(pp\. 63–90\)\. Springer\.

Magnus, P\. D\., & Callender, C\. \(2004\)\. Realist ennui and the base rate fallacy\.Philosophy of Science, 71\(3\), 320–338\.

Neyman, J\., & Pearson, E\. S\. \(1936\)\. Contributions to the theory of testing statistical hypotheses: Part I\.Statistical Research Memoirs, 1, 1–37\.

Peirce, C\. S\. \(1902\)\. Validity\. In J\. M\. Baldwin \(Ed\.\),Dictionary of philosophy and psychology\. Macmillan\.

Psillos, S\. \(1994\)\. A philosophical study of the transition from the caloric theory of heat to thermodynamics: Resisting the pessimistic meta\-induction\.Studies in History and Philosophy of Science, 25\(2\), 159–190\.

Putnam, H\. \(1965\)\. Trial and error predicates and the solution to a problem of Mostowski\.Journal of Symbolic Logic, 30\(1\), 49–57\.

Putnam, H\. \(1978\)\.Meaning and the moral sciences\. Routledge\.

Reichenbach, H\. \(1938\)\.Experience and prediction: An analysis of the foundation and the structure of knowledge\. University of Chicago Press\.

Schulte, O\. \(1999\)\. Means\-ends epistemology\.The British Journal for the Philosophy of Science, 50\(1\), 1–31\.

Shalev\-Shwartz, S\., & Ben\-David, S\. \(2014\)\.Understanding machine learning: From theory to algorithms\. Cambridge University Press\.

Spirtes, P\., Glymour, C\., & Scheines, R\. \(2000\)\.Causation, prediction, and search\(2nd ed\.\)\. MIT Press\.

Worrall, J\. \(1994\)\. How to remain \(reasonably\) optimistic: Scientific realism and the “luminiferous ether”\.PSA: Proceedings of the Biennial Meeting of the Philosophy of Science Association, 1994\(1\), 334–342\.

Wray, K\. B\. \(2015\)\. Pessimistic inductions: Four varieties\.International Studies in the Philosophy of Science, 29\(1\), 61–73\.

Similar Articles

Solution of the Hempel's statistical ambiguity problem and Causal AI

arXiv cs.AI

This paper presents a solution to Carl Hempel's statistical ambiguity problem in inductive-statistical inference by introducing maximally specific causal relationships (MSCRs) and proving their predictions are consistent, with implications for Causal AI and machine learning.