Solution of the Hempel's statistical ambiguity problem and Causal AI

arXiv cs.AI Papers

Summary

This paper presents a solution to Carl Hempel's statistical ambiguity problem in inductive-statistical inference by introducing maximally specific causal relationships (MSCRs) and proving their predictions are consistent, with implications for Causal AI and machine learning.

arXiv:2607.12826v1 Announce Type: new Abstract: This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity (RMS) for the statistical laws used in the inference. An analysis of the RMS refinements made by Wesley Salmon, Alberto Coffa, and James Fetzer led to the following definition of maximally specific statistical laws: "the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum-phenomenon." However, there was no proof of a solution to the statistical ambiguity problem based on this definition. We use Nancy Cartwright's definition of causes that raise probabilities across background contexts, and then introduce the concept of Causal Rules. Then we define a special semantic probabilistic inference procedure that incrementally refines these causal rules by incorporating all statistically relevant information. This procedure yields Maximally Specific Causal Relationships (MSCRs), for which we prove (Theorem 1) that predictions derived from them are consistent. This resolves the statistical ambiguity problem. The semantic probabilistic inference procedure provides a probabilistic causal learning system, which may be used in such new areas as Causal AI and Causal Machine Learning. They fundamentally explore causal inference as a tool for understanding cause-and-effect relationships within complex systems. Properties similar to RMS remain under discussion. Several notions related to RMS are considered: invariant feature learning, invariant causal prediction, and spurious association.
Original Article
View Cached Full Text

Cached at: 07/15/26, 04:20 AM

# Solution of the Hempel’s statistical ambiguity problem and Causal AI
Source: [https://arxiv.org/html/2607.12826](https://arxiv.org/html/2607.12826)
\[1\]\\fnmEvgenii\\surVityaev

\[1\]\\orgnameSobolev Institute of Mathematics of the SB RAS,\\orgaddress\\streetKoptuga 4,\\cityNovosibirsk,\\postcode630090,\\countryRussia

###### Abstract

This paper addresses Carl Hempel’s longstanding problem of statistical ambiguity in inductive\-statistical inference, in which contradictory predictions are derived from statistical laws\. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity \(RMS\) for the statistical laws used in the inference\. An analysis of the RMS refinements made by Wesley Salmon, Alberto Coffa, and James Fetzer led to the following definition of maximally specific statistical laws: ”the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum\-phenomenon\.” However, there was no proof of a solution to the statistical ambiguity problem based on this definition\. We use Nancy Cartwright’s definition of causes that raise probabilities across background contexts, and then introduce the concept of Causal Rules\. Then we define a special semantic probabilistic inference procedure that incrementally refines these causal rules by incorporating all statistically relevant information\. This procedure yields Maximally Specific Causal Relationships \(MSCRs\), for which we prove \(Theorem 1\) that predictions derived from them are consistent\. This resolves the statistical ambiguity problem\. The semantic probabilistic inference procedure provides a probabilistic causal learning system, which may be used in such new areas as Causal AI and Causal Machine Learning\. They fundamentally explore causal inference as a tool for understanding cause\-and\-effect relationships within complex systems\. Properties similar to RMS remain under discussion\. Several notions related to RMS are considered: invariant feature learning, invariant causal prediction, and spurious association\.

###### keywords:

explanation, inductive\-statistical inference, statistical ambiguity, causal inference, consistency

## 1Introduction

In recent years, within the fields of Causal AI and Causal Machine Learning, causal inference has become a critical tool for understanding cause\-and\-effect relationships\. Integrating these relationships into machine learning models enables the creation of causal models for a deeper understanding of real\-world systems\. This paradigm provides more robust, interpretable, and actionable insights in areas such as medicine, finance, and autonomous systems\.

To develop the probabilistic causal learning system that infers effects without contradictions we need to solve Hempel’sstatistical ambiguity problem\. It concerns the problem that contradictory predictions derived from inductive\-statistical inference\[Hempel65,Hempel68\]\. To avoid such contradictions, Carl Hempel introduced the Requirement of Maximal Specificity \(RMS\) for statistical laws, which means that maximally specific statistical laws must incorporate all information relevant to prediction\. Subsequent analysis of RMS revealed numerous problems and disputes \(see next section on Historical background\)\. As a result, a solution to the ”statistical ambiguity” problem was not reached, and the consistency of predictions for maximally specific statistical laws was not proven, and the problem remained unsolved\.

This article resumes the consideration of the ”statistical ambiguity problem” and RMS\. We define RMS based on Hempel’s definition and define the notion of maximally specific cause\-effect relationships \(MSCR\)\. Then we prove a theorem \(Theorem 1\) that predictions based on MSCRs are consistent\. This solves Hempel’s ”statistical ambiguity problem” for maximally specific cause\-effect relationships\. For the works in Causal AI and Causal Machine Learning it is rather important\.

Based on Cartwright’s definition of causes that raise probabilities of their effects in various background contexts\[Cartwright,Stanford\]we define a stronger notion of Causal Rules \(CR\)\.

- Cartwright: C causes E if and only if P​\(E∣C&B\)\>P​\(E∣¬C&B\)P\(E\\mid C\\&B\)\>P\(E\\mid\\neg\{C\}\\&B\) for every background context B\.
- The background in the Causal Rules consists of all other conditions of the rule\.
- Causal Rule \(see Definition 6\): The ruleA1&⋯&Ak⇒A0A\_\{1\}\\&\\dots\\&A\_\{k\}\\Rightarrow A\_\{0\}, is acausal ruleiff every literalA1,…,AkA\_\{1\},\\dots,A\_\{k\}is a cause ofA0A\_\{0\} P​\(A0∣A1&⋯&Ak\)\>P​\(A0∣¬Ai&B\)P\(A\_\{0\}\\mid A\_\{1\}\\&\\dots\\&A\_\{k\}\)\>P\(A\_\{0\}\\mid\\neg\{A\_\{i\}\}\\&B\),i=1,…,ki=1,\\dots,k, relative to backgroundB=&\{\{A1,…,Ak\}∖Ai\}B=\\&\\\{\\\{A\_\{1\},\\dots,A\_\{k\}\\\}\\setminus A\_\{i\}\\\}, containing all literals exceptAiA\_\{i\}\.

We introduce a refinement procedure for causal rules that incrementally strengthens the conditional probabilities of causal rules by adding all relevant information until the causal rules can no longer be refined, thus producing Maximally Specific Causal Relationships\. It provides the learning method \(Definition 8\) for MSCR discovery\. MSCRs may be considered as causal models, as they identify the key variables that underpin reliable cause\-and\-effect relationships\. This method thus produces a causal learning system\.

Let𝖬𝖲𝖱\\sf MSRis the set of all MSCRs\. It is proved in Theorem 1 that I\-S inference of predictions based on any subset of𝖬𝖲𝖱\\sf MSRis consistent\. The set𝖬𝖲𝖱\\sf MSRthus provides a consistent set of probabilistic knowledge\. The set𝖬𝖲𝖱\\sf MSRmay be considered as aconsistent probabilistic theoryin the same sense as a logical theory that is consistent and infers consistent conclusions\.

The problem of statistical ambiguity for statistical laws of the formA1&⋯&Ak⇒A0A\_\{1\}\\&\\dots\\&A\_\{k\}\\Rightarrow A\_\{0\}, whereA1,…,Ak,A0A\_\{1\},\\dots,A\_\{k\},A\_\{0\}are literals was considered in\[Author1,Author12\], and for the formϕ⇒ψ\\phi\\Rightarrow\\psi, whereϕ\\phiandψ\\psiare propositional formulas was considered in\[Author13\]\.

Our paper is organized as follows\. Section 2 \(RMS history\) is devoted to the historical analysis of RMS and reveals discussions around it\. Section 3 considers relation of RMS to Causal AI and Causal Machine Learning\. Section 4 is devoted to definition of probability on propositional formulas\. Section 5 considers rules, their probabilities and the definition of MSCR\. In Section 6, we define I\-S inference as a prediction operator for𝖬𝖲𝖱\\sf MSRand prove that applying it to a consistent set of MSCRs produces a consistent set\. Finally, Section 7 is conclusion\.

## 2RMS history

The problem of statistical ambiguity in inductive\-statistical \(I\-S\) explanation was first identified by Hempel\. It generated a substantial literature attempting to articulate and refine the conditions under which probabilistic explanations can avoid contradictory conclusions\. This section traces the evolution of the Requirement of Maximal Specificity \(RMS\) from its origins in Hempel’s work through its subsequent critiques and reformulations, culminating in the identification of causal factors as the essential element for resolving the ambiguity\.

The Principle of Total Evidence and Its Limitations\. Before addressing the Requirement of Maximal Specificity directly, it is essential to distinguish it from another fundamental principle of probabilistic inference: the requirement of total evidence\. As articulated by Carnap and others, the requirement of total evidence states that in any rational application of probabilistic inference, the probability assigned to a hypothesis must be determined by reference to all available evidence\. In Hempel’s formulation, this principle directs that when assessing the credibility of a statement, one must consider the entire body of knowledge K: “What degree of belief, or what probability, is it rational to assign to the statement ’Gi’ in a given knowledge situation?”\[Hempel68\]\.

Initially, Hempel had conflated the requirement of maximal specificity with the requirement of total evidence, suggesting that RMS served as a ”rough substitute” for the latter\. However, in his 1968 reappraisal, he explicitly retracted this view, recognizing that the two principles address fundamentally different questions\. The requirement of total evidence concerns the rational credibility of a statement based on all available information\. The Requirement of Maximal Specificity, by contrast, addresses the explanatory status of an argument: it specifies conditions under which two sentences inKK— a statistical lawp​\(G,F\)=rp\(G,F\)=rand a singular sentenceFiF\_\{i\}— can serve to explain, relative toKK, whyiiisGG\. As Hempel emphasizes: The point of an explanation is not to provide evidence for the occurrence of the explanandum phenomenon, but to exhibit it as nomically expectable\. And the probability attached to anI−SI\-Sexplanation is the probability of the conclusion relative to the explanatory premises, not relative to the total classKK\. Thus, the requirement of total evidence simply does not apply to the determination of the probability associated with anI−SI\-Sexplanation, and the requirement of maximal specificity is not ”a rough substitute for the requirement of total evidence\.”\[Hempel68\]

This clarification established RMS as an independent principle with its own distinct rationale, specifically designed to address the problem of explanatory ambiguity in statistical contexts\.

The Problem of Statistical Ambiguity and the Initial Formulation of RMS\. The problem that RMS was designed to solve emerges from the nature of inductive\-statistical explanation itself\. As\[Hempel65\]demonstrated, anI−SI\-Sargument has the form:

p​\(G;F\)=rp\(G;F\)=rF​\(a\)F\(a\)G​\(a\)G\(a\)
where line indicates an inductive relationship, and r represents the inductive probability of the explanandum given the explanans\.

Statistical ambiguity arises when two arguments with true premises yield contradictory conclusions\. It can be illustrates with a classic example \(cited Salmon\)\.

Suppose that we have the following statements\.

- •L1 Almost all cases of streptococcus infection clear up quickly after the administration of penicillin\.
- •L2 Almost no cases of penicillin resistant streptococcus infection clear up quickly after the administration of penicillin\.
- •C1 Jane Jones had streptococcus infection\.
- •C2 Jane Jones received treatment with penicillin\.
- •C3 Jane Jones had a penicillin resistant streptococcus infection\.

Following the above pattern of I\-S explanation it is possible to construct two contradictory arguments based on these statements\. On the base ofL1L\_\{1\}andC​1∧C​2C1\\land C2one can explain why Jane Jones recovered quickly \(E\)\. The second argument with premisesL2L\_\{2\}andC​2∧C​3C2\\land C3explains why Jane Jones did not \(¬E\\lnot E\)\. The premises of both arguments are consistent with each other, they could all be true\. The probability of both argument may be close to 1\.

To block such conflicting explanations, Hempel introduced the Requirement of Maximal Specificity\[Hempel65\]\.

> RMS: For anI−SI\-Sargument to be acceptable relative to a knowledge stateKK, for any predicateHHsuch thatKKcontains both∀x​\(H​\(x\)⇒F​\(x\)\)\\forall x\(H\(x\)\\Rightarrow F\(x\)\)andH​\(a\)H\(a\), there must exist a statistical lawp​\(G;H\)=rp\(G;H\)=rinKKwith the same probabilityrr\.

The basic idea is that ifHHprovides more specific information about the object thanFF, then the law based onHHshould be preferred — and if that law has a different probability, the original argument is invalid\.

Hempel hoped to solve this problem by forcing all statistical laws in an argument to be maximally specific\. That is, they should contain all relevant information with respect to the domain in question\. In our example, then, the premiseC​3C3invalidates the first argument, since it is not maximally specific with respect to all information about Jane Jones\. So, we can only explain¬\\negE, but not E\.

The Problem of Epistemic Relativity\. Coffa criticizes Hempel’s insistence on relativizingI−SI\-Sexplanation to a knowledge state K\[Coffa\]\. Coffa begins by distinguishing between epistemic and non\-epistemic concepts\. A concept is epistemic if its meaning cannot be given without reference to knowledge; non\-epistemic concepts—such as “table,” “chair,” or “truth” — can be characterized independently of any knower\. Hempel’s thesis, Coffa explains, is that inductive explanation is not merely epistemic in the sense that it depends on what we know about the evidence \(as in the case of “well\-confirmed D\-N explanation”\), but rather that it is a “non\-conformational epistemic concept”\. This means, as Hempel himself acknowledges, that “there is no concept that stands to his epistemically relativized notion of inductive explanation as the concept of true D\-N explanations stands to that of well\-confirmed D\-N explanation”\[Coffa\]\. In other words: According to the thesis of epistemic relativity there is no meaningful notion of true inductive explanation\. Hence, we could not possibly have reasons to believe that anything is a “true inductive explanation”\.\[Coffa\]\.

Coffa’s fundamental objection is that this conclusion is not merely surprising but fatal to the project of inductive explanation\. If there is no notion of true inductive explanation, thenI−SI\-Sexplanations relativized toKKcannot be understood as arguments that we have reason to believe are explanations in the same sense thatD−ND\-Nexplanations are\.

More importantly, Coffa argues that Hempel has misidentified the nature of the problem\. The real difficulty is not epistemic but ontological — it is the old problem of the reference class in a new guise\.

We would like to suggest that when Hempel turned his attention to the theory of inductive explanation what he stumbled upon was the fact that the problem of defining a model of inductive explanation for single events was the other side of the coin of the single case problem\. He stumbled, that is, upon the reference\-class problem\[Coffa\]\.

The reference\-class problem arises because a single event can belong to multiple reference classes with different probabilities for the outcome of interest\. Jones belongs to the class of persons with streptococcus infection, the class of persons treated with penicillin, and the class of persons with penicillin\-resistant infections\. Each yields a different probability for recovery\. The frequentist tradition, Coffa notes, held that it is ”strictly meaningless to assign a probability to a single event”\[Coffa\]because all reference classes are, in principle, equally legitimate\.

Hempel’s epistemic relativization attempts to solve this problem by appeal to knowledge: the appropriate reference class is the most specific class to which the individual is known to belong\. But Coffa observes that this strategy has the ironic consequence that “it is ignorance, rather than knowledge, that makes the maximal specificity principle look like a workable demand”\[Coffa\]\. As knowledge increases, the principle becomes increasingly difficult to satisfy; an omniscient being would find no inductive explanations at all\.

Coffa’s proposed alternative points toward an ontic rather than epistemic formulation: the requirement should refer to “all relevant aspects of the explanandum”\[Coffa\], where relevance is understood not as statistical correlation but as nomic connection — “a predicate being nomically relevant to another when a law of nature determines that changes in the first one generate changes in the second one”\[Coffa\]\. This suggestion anticipates later developments in causal approaches to explanation\.

Salmon’s Statistical Relevance Approach and shift to Ontic Homogeneity\. A different line of critique emerged from Wesley Salmon, who questioned not merely the formulation of RMS but its fundamental orientation\. Salmon argued that Hempel’s requirement of high probability is misguided; what matters for explanation is not the magnitude of probability but the presence of statistical relevance relations\.

The following example illustrates this point\[Salmon\]:

John Jones was almost certain to recover from his cold within a week, because he took vitamin C, and almost all colds clear up within a week after administration of vitamin C\.

The difficulty with this example is that colds tend to clear up within a week regardless of the medication administered, and controlled tests indicate that the percentage of recoveries is unaffected by the use of vitamin C\.

Thus, even though the argument satisfies Hempel’s requirements — a high probability, a statistical law, and true premises — it fails to be explanatory because the putative explanatory factor \(taking vitamin C\) is irrelevant to the outcome\. What is needed, Salmon argues, is not high probability but relevance — the fact that taking vitamin C makes a difference to the probability of recovery\.

This led Salmon to develop the Statistical\-Relevance \(S\-R\) model, in which an explanation consists of partitioning a reference class into cells that are homogeneous with respect to the explanandum property, together with information about which cell contains the individual in question\. Crucially, Salmon’s homogeneity requirement is objective rather than epistemic: a class is homogeneous if no further partition can be made that is relevant to the occurrence of the explanandum property, regardless of whether such partitions are known\. This represents a fundamental shift from Hempel’s epistemic relativization to an ontic conception of explanation\. As Salmon later put it, “the identification of the appropriate explanans is fully objective” once the explanandum has been unambiguously specified \(Salmon, 1989\)\.

Fetzer’s Requirement of Strict Maximal Specificity and the Turn to Causation\. James Fetzer \(\[Fetzer81,Fetzer93\]\) carries the ontic turn further, arguing that neither Hempel’s epistemic RMS nor Salmon’s statistical relevance conditions are sufficient\. What is required is not merely statistical homogeneity but causal relevance\. Fetzer introduces the Requirement of Strict Maximal Specificity \(RSMS\), which demands “that the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum\-phenomenon”\[Fetzer93\]\.

Fetzer’s analysis reveals that the problem of statistical ambiguity cannot be resolved at the level of statistical relations alone\. The fundamental question is not which reference class yields the highest probability or even which partition yields homogeneous cells, but rather: which factors are genuinely responsible for the outcome? This is an ontological question about causal structure, not an epistemological question about the choice of reference class\.

Conclusion\. The evolution of the requirements for maximum specificity reveals the depth of the Hempel statistical ambiguity problem\. Salmon and Fetzer came close enough to solving the problem, but did not solve it\. Their research is taken into account in our definition of ”causal rule”\. Salmon’s statistical significance is taken into account in the definition of ”causal rules” as Cartwright’s definition of cause\. James Fetzer’s requirement of strict maximal specificity is taken into account by determining the statistical significance of each condition in the premise of the rule\. Fetzer’s requirement that ”adequate explanations apply to all and only those properties whose presence influenced the explanation” is taken into account in the procedure for clarifying causal rules, which gradually strengthens the conditional probabilities of causal rules by adding all relevant information until the causal rules can no longer be clarified\. This procedure allows us to obtain the Maximally Specific Causal Relationships that solve the problem of statistical ambiguity \(Theorem 1\)\.

## 3Causal AI and Causal Machine Learning

In recent years, causal inference has emerged as a critical tool for understanding cause\-and\-effect relationships within complex systems\. By incorporating causal reasoning into machine learning, models can move to deeper understanding of the real\-world systems\.

The main theoretical achievement of MSCRs is the formal proof of consistency for predictions\. This approach shares the ambition of contemporary causal approaches to identify genuinely operative causes\.

Properties similar to RMS remain under discussion\[Kaddour\]\. There are some notions related to RMS in the Causal Machine Learning\[Kaddour\]:

- •Invariant Feature Learning\. The task of identifying features of our data X that are predictive of Y across a range of environments E\. This definition is very similar to the N\. Cartwright definition of causes that raise probabilities of their effects in various background contexts\. It is included in our definition of causal rules\.
- •Invariant Causal Prediction– an algorithm to find the causal feature set, the minimal set of features which are causal predictors of a target variable\[Kaddour\]\. In our definition of causal rule this property is also fulfilled\.
- •Spurious association\. In the training dataset, pictures of cows typically exhibit alpine pasture backgrounds, a spurious association caused by the cow’s natural habitat\. Definitions of cause by N\. Cartwright and a causal rule exclude these spurious associations\.

Thus, the introduced notion of “causal rule” remains an orientation for other Causal AI and Causal Machine Learning methods to extract precise knowledge from data predicting without contradictions\.

## 4Logic and Probability background

LetF​\(A​t\)F\(At\)be a set of well\-formed formulas constructed from a set of atomsA​tAtusing connectives&\\&,∨\\vee,→\\rightarrow,¬\\neg\. Equivalence↔\\leftrightarrowis an abbreviation,φ↔ψ=\(φ→ψ\)&\(ψ→φ\)\\varphi\\leftrightarrow\\psi=\(\\varphi\\rightarrow\\psi\)\\&\(\\psi\\rightarrow\\varphi\)\. We define⊤\\topasφ∨¬φ\\varphi\\vee\\neg\\varphi, whereφ\\varphiis some fixed formula\.

The conjunction of a finite set of formulasTTis denoted by⋀T\\bigwedge T\.V​\(φ\)V\(\\varphi\)denotes the set of atoms occurring in formulaφ\\varphi\. The algebra of formulasℱ​\(A​t\)\\mathcal\{F\}\(At\)is a setF​\(A​t\)F\(At\)with naturally interpreted connectives\.

###### Definition 1\.

The models are mappings fromA​tAtto truth values\{0,1\}\\\{0,1\\\}\. Such mappings we call \(A​tAt\-\)valuations\. Every valuationv​a​l:A​t→\{0,1\}val:At\\rightarrow\\\{0,1\\\}extends in a standard way to the setF​\(A​t\)F\(At\)asv​a​l:F​\(A​t\)→\{0,1\}val:F\(At\)\\rightarrow\\\{0,1\\\}\.

Let𝔊\\mathfrak\{G\}be a set of valations\. A formulaφ\\varphiis said tosatisfiable in𝔊\\mathfrak\{G\}ifv​a​l​\(φ\)=1val\(\\varphi\)=1for somev​a​l∈𝔊val\\in\\mathfrak\{G\}\. We say thatφ\\varphiholds on𝔊\\mathfrak\{G\}\(and write𝔊⊧φ\\mathfrak\{G\}\\models\\varphi\) ifv​a​l​\(φ\)=1val\(\\varphi\)=1for allv​a​l∈𝔊val\\in\\mathfrak\{G\}\. Finally, a setT⊆F​\(A​t\)T\\subseteq F\(At\)is said to be𝔊\\mathfrak\{G\}\-consistentif there isv​a​l∈𝔊val\\in\\mathfrak\{G\}such thatv​a​l​\(φ\)=1val\(\\varphi\)=1for allφ∈T\\varphi\\in T\. The set of all valuations we denote𝔄\\mathfrak\{A\}\. By a logical tautology we mean a formula that holds on𝔄\\mathfrak\{A\}\.

For a set𝔊\\mathfrak\{G\}of valuations, the relationφ≡𝔊ψ\\varphi\\equiv\_\{\\mathfrak\{G\}\}\\psiis defined byφ↔ψ\\varphi\\leftrightarrow\\psiholding on𝔊\\mathfrak\{G\}\. Then≡𝔊\\equiv\_\{\\mathfrak\{G\}\}is a congruence onℱ​\(A​t\)\\mathcal\{F\}\(At\), the respective quotient is denoted asℬ𝔊​\(A​t\)\\mathcal\{B\}^\{\\mathfrak\{G\}\}\(At\)\. The coset ofφ\\varphiw\.r\.t\.≡𝔊\\equiv\_\{\\mathfrak\{G\}\}is denoted as\[φ\]𝔊\[\\varphi\]\_\{\\mathfrak\{G\}\}and the universe ofℬ𝔊​\(A​t\)\\mathcal\{B\}^\{\\mathfrak\{G\}\}\(At\)equals\{\[φ\]𝔊∣φ∈F​\(A​t\)\}\\\{\[\\varphi\]\_\{\\mathfrak\{G\}\}\\mid\\varphi\\in F\(At\)\\\}\. The operations ofℬ𝔊​\(A​t\)\\mathcal\{B\}^\{\\mathfrak\{G\}\}\(At\)are denoted as&𝔊\\&\_\{\\mathfrak\{G\}\},∨𝔊\\vee\_\{\\mathfrak\{G\}\},→𝔊\\rightarrow\_\{\\mathfrak\{G\}\},¬𝔊\\neg\_\{\\mathfrak\{G\}\}and the lattice order as⊑𝔊\\sqsubseteq\_\{\\mathfrak\{G\}\}\. Recall that\[φ\]𝔊⊑𝔊\[ψ\]𝔊\[\\varphi\]\_\{\\mathfrak\{G\}\}\\sqsubseteq\_\{\\mathfrak\{G\}\}\[\\psi\]\_\{\\mathfrak\{G\}\}iff\[φ\]𝔊=\[φ&ψ\]𝔊=\[φ\]𝔊&𝔊\[ψ\]𝔊\[\\varphi\]\_\{\\mathfrak\{G\}\}=\[\\varphi\\&\\psi\]\_\{\\mathfrak\{G\}\}=\[\\varphi\]\_\{\\mathfrak\{G\}\}\\&\_\{\\mathfrak\{G\}\}\[\\psi\]\_\{\\mathfrak\{G\}\}\.

###### Definition 2\.

Finitely additive measureμ\\muis defined on𝔊\\mathfrak\{G\}asμ:2𝔊→\[0,1\]\\mu:2^\{\\mathfrak\{G\}\}\\rightarrow\[0,1\]:

1. 1\.μ​\(𝔊\)=1\\mu\(\\mathfrak\{G\}\)=1,μ​\(∅\)=0\\mu\(\\varnothing\)=0;
2. 2\.μ​\(A1∪…∪An\)=μ​\(A1\)\+…\+μ​\(An\)\\mu\(A\_\{1\}\\cup\\ldots\\cup A\_\{n\}\)=\\mu\(A\_\{1\}\)\+\\ldots\+\\mu\(A\_\{n\}\)for pairwise disjoint subsetsA1A\_\{1\}, …,An⊆𝔊A\_\{n\}\\subseteq\\mathfrak\{G\}\.
3. 3\.μ​\(A\)=0\\mu\(A\)=0iffA=∅A=\\varnothing

Elements of𝔊\\mathfrak\{G\}may be interpreted as outcomes of experiments\. Further, we assume that only essential experiments are included in𝔊\\mathfrak\{G\}, which explains whyμ​\(\{v​a​l\}\)≠0\\mu\(\\\{val\\\}\)\\neq 0for allv​a​l∈𝔊val\\in\\mathfrak\{G\}\.

For everyφ∈F​\(A​t\)\\varphi\\in F\(At\), we defineφ𝔊:=\{v​a​l∈𝔊∣v​a​l​\(φ\)=1\}\\varphi^\{\\mathfrak\{G\}\}:=\\\{val\\in\\mathfrak\{G\}\\mid val\(\\varphi\)=1\\\}and functionν:F​\(A​t\)→\{0,1\}\\nu:F\(At\)\\rightarrow\\\{0,1\\\}by the ruleν​\(φ\)=μ​\(φ𝔊\)\\nu\(\\varphi\)=\\mu\(\\varphi^\{\\mathfrak\{G\}\}\)\. Then the functionν\\nusatisfies the following properties\.

###### Proposition 1\.

1. 1\.ν​\(φ\)=1\\nu\(\\varphi\)=1iffφ\\varphiholds on𝔊\\mathfrak\{G\}\.
2. 2\.ν​\(φ\)=0\\nu\(\\varphi\)=0iffφ\\varphiis not satisfiable on𝔊\\mathfrak\{G\}\.
3. 3\.ν​\(φ∨ψ\)=ν​\(φ\)\+ν​\(ψ\)\\nu\(\\varphi\\vee\\psi\)=\\nu\(\\varphi\)\+\\nu\(\\psi\)iffφ&ψ\\varphi\\&\\psiis not satisfiable in𝔊\\mathfrak\{G\}\.

Thus we defined the probability on the set of propositional formulas in the sense of\[Fagin\]\.

## 5Method\. Causal Rules and Semantic Probabilistic Inference

By arulewe mean a syntactic object of the form

r=φ⇒ψ,r=\\varphi\\Rightarrow\\psi,whereφ,ψ∈F​\(A​t\)\\varphi,\\psi\\in F\(At\)\. Formulaφ\\varphiis called abodyof the rule, whereasψ\\psiis aheadof the rule:φ=B​\(r\)\\varphi=B\(r\)andψ=H​\(r\)\\psi=H\(r\)\. A rulerrcannot be identified with the implicationφ→ψ\\varphi\\rightarrow\\psibecause the probability ofrrwill be defined in a different way\. Namely, for a ruler=φ⇒ψr=\\varphi\\Rightarrow\\psi, whose body is satisfiable on𝔊\\mathfrak\{G\}, i\.e\.,ν​\(φ\)≠0\\nu\(\\varphi\)\\neq 0, we put

ν​\(r\):=ν​\(ψ\|φ\)=ν​\(ψ&φ\)ν​\(φ\)\.\\nu\(r\):=\\nu\(\\psi\|\\varphi\)=\\frac\{\\nu\(\\psi\\&\\varphi\)\}\{\\nu\(\\varphi\)\}\.In caseB​\(r\)B\(r\)is not satisfiable on𝔊\\mathfrak\{G\}, the valueν​\(r\)\\nu\(r\)remains undefined\.

Notice that the valueν​\(r\)\\nu\(r\)was defined so that it is smaller than the probability of implication\.

###### Proposition 2\.

For every rulerrwithB​\(r\)𝔊≠𝔊B\(r\)^\{\\mathfrak\{G\}\}\\neq\\mathfrak\{G\}, we have

ν​\(r\)≤ν​\(B​\(r\)→H​\(r\)\)\.\\nu\(r\)\\leq\\nu\(B\(r\)\\rightarrow H\(r\)\)\.Moreover, the equalityν​\(r\)=ν​\(B​\(r\)→H​\(r\)\)\\nu\(r\)=\\nu\(B\(r\)\\rightarrow H\(r\)\)is equivalent toB​\(r\)𝔊⊆H​\(r\)𝔊B\(r\)^\{\\mathfrak\{G\}\}\\subseteq H\(r\)^\{\\mathfrak\{G\}\}, i\.e\., to the fact that the implicationB​\(r\)→H​\(r\)B\(r\)\\rightarrow H\(r\)holds on𝔊\\mathfrak\{G\}\.

This statement can be proved analogously to Theorem 2 from\[Author13\]

###### Definition 3\.

Letr1r\_\{1\}andr2r\_\{2\}be two rules with the same head,H​\(r1\)=H​\(r2\)H\(r\_\{1\}\)=H\(r\_\{2\}\)\. We callr1r\_\{1\}a specification ofr2r\_\{2\}, symbolicallyr1≼r2r\_\{1\}\\preccurlyeq r\_\{2\}, ifB​\(r1\)𝔊⊆B​\(r2\)𝔊B\(r\_\{1\}\)^\{\\mathfrak\{G\}\}\\subseteq B\(r\_\{2\}\)^\{\\mathfrak\{G\}\}; ruler1r\_\{1\}is a proper specification ofr2r\_\{2\},r1≺r2r\_\{1\}\\prec r\_\{2\}, ifB​\(r1\)𝔊⫋B​\(r2\)𝔊B\(r\_\{1\}\)^\{\\mathfrak\{G\}\}\\subsetneqq B\(r\_\{2\}\)^\{\\mathfrak\{G\}\}\. We say in this case thatr2r\_\{2\}is a \(proper\) generalization ofr1r\_\{1\}\.

In other words, one of the two rules with the same head is a proper generalization of the other if its body is weaker from the logical point of view\.

###### Definition 4\.

Let ruler1r\_\{1\}be a specification ofr2r\_\{2\}\. We say thatr1r\_\{1\}is a refinement ofr2r\_\{2\}, symbolicallyr1\>r2r\_\{1\}\>r\_\{2\}, ifν​\(r1\)\>ν​\(r2\)\\nu\(r\_\{1\}\)\>\\nu\(r\_\{2\}\)\.

Evidently, the relationr1\>r2r\_\{1\}\>r\_\{2\}implies thatr1r\_\{1\}is a proper specification ofr2r\_\{2\}\.

###### Definition 5\.

A setℛ\\mathcal\{R\}of rules is said to be rich if for everyr∈ℛr\\in\\mathcal\{R\}and an arbitrarysssuch thats\>rs\>r, there isr′∈ℛr^\{\\prime\}\\in\\mathcal\{R\}withr′≼sr^\{\\prime\}\\preccurlyeq sandν​\(r′\)\>ν​\(r\)\\nu\(r^\{\\prime\}\)\>\\nu\(r\)\.

###### Definition 6\.

Letℛ\\mathcal\{R\}be a rich set of rules\. Rulerris a causal rule relative toℛ\\mathcal\{R\}, ifr∈ℛr\\in\\mathcal\{R\}andrris a refinement of all its proper generalizations fromℛ\\mathcal\{R\}\. We assume that ifr∈ℛr\\in\\mathcal\{R\}and\(⊤⇒H\(r\)\)≻r\(\\top\\Rightarrow H\(r\)\)\\succ rthenν\(r\)\>ν\(⊤⇒H\(r\)\)\\nu\(r\)\>\\nu\(\\top\\Rightarrow H\(r\)\)\.

###### Proposition 3\.

Letℛ\\mathcal\{R\}be a rich set of rules over a finite setA​tAtof atoms\. For every ruler∈ℛr\\in\\mathcal\{R\}, there exists its generalizationr′r^\{\\prime\}such thatr′r^\{\\prime\}is a causal rule relative toℛ\\mathcal\{R\}andν​\(r′\)≥ν​\(r\)\\nu\(r^\{\\prime\}\)\\geq\\nu\(r\)\.

###### Proof\.

Letr=φ⇒ψ∈ℛr=\\varphi\\Rightarrow\\psi\\in\\mathcal\{R\}\. Consider the set

Δ=\{\[α\]𝔊∣ν​\(α⇒ψ\)≥ν​\(r\),\[α\]𝔊⊑𝔊\[φ\]𝔊\}\.\\Delta=\\\{\[\\alpha\]\_\{\\mathfrak\{G\}\}\\mid\\nu\(\\alpha\\Rightarrow\\psi\)\\geq\\nu\(r\),\\ \[\\alpha\]\_\{\\mathfrak\{G\}\}\\sqsubseteq\_\{\\mathfrak\{G\}\}\[\\varphi\]\_\{\\mathfrak\{G\}\}\\\}\.Recall that the condition\[α\]𝔊⊑𝔊\[φ\]𝔊\[\\alpha\]\_\{\\mathfrak\{G\}\}\\sqsubseteq\_\{\\mathfrak\{G\}\}\[\\varphi\]\_\{\\mathfrak\{G\}\}means exactly thatα⇒ψ≼r\\alpha\\Rightarrow\\psi\\preccurlyeq r\. SinceΔ\\Deltais finite as a subset ofℬ𝔊​\(A​t\)\\mathcal\{B\}^\{\\mathfrak\{G\}\}\(At\), we can choose an element\[β\]𝔊∈Δ\[\\beta\]\_\{\\mathfrak\{G\}\}\\in\\Deltaminimal w\.r\.t\.⊑𝔊\\sqsubseteq\_\{\\mathfrak\{G\}\}\. Sinceℛ\\mathcal\{R\}is reach andβ⇒ψ≼r\\beta\\Rightarrow\\psi\\preccurlyeq r, there isr′∈ℛr^\{\\prime\}\\in\\mathcal\{R\}withr′≼β⇒ψr^\{\\prime\}\\preccurlyeq\\beta\\Rightarrow\\psi\. Assume thatr′r^\{\\prime\}is not a causal rule, then there is a proper generalizationγ⇒ψ\\gamma\\Rightarrow\\psiofr′r^\{\\prime\}such thatν​\(γ⇒ψ\)≥ν​\(r′\)\\nu\(\\gamma\\Rightarrow\\psi\)\\geq\\nu\(r^\{\\prime\}\)\. Naturally, in this case\[α\]𝔊∈Δ\[\\alpha\]\_\{\\mathfrak\{G\}\}\\in\\Deltaand since a proper generalizationγ⇒ψ\\gamma\\Rightarrow\\psiis a proper generalization ofr′r^\{\\prime\}, we have\[γ\]𝔊⊏𝔊\[β\]𝔊\[\\gamma\]\_\{\\mathfrak\{G\}\}\\sqsubset\_\{\\mathfrak\{G\}\}\[\\beta\]\_\{\\mathfrak\{G\}\}, which contradicts the minimality of\[β\]𝔊\[\\beta\]\_\{\\mathfrak\{G\}\}\. Thus,r′r^\{\\prime\}is the required causal rule\. ∎

###### Definition 7\.

A causal rulerris strong relative toℛ\\mathcal\{R\}, if there is no causal ruler′r^\{\\prime\}such thatr′r^\{\\prime\}is a refinement ofrr\.

###### Definition 8\.

Semantic Probabilistic Inference \(SPI\) ofψ\\psiis a sequencer1,…,rkr\_\{1\},\\dots,r\_\{k\}of causal rulesri∈ℛr\_\{i\}\\in\\mathcal\{R\}with the headψ\\psisuch that:

- ν\(r1\)\>ν\(⊤⇒ψ\);\\nu\(r\_\{1\}\)\>\\nu\(\\top\\Rightarrow\\psi\);
- ri\+1\>rir\_\{i\+1\}\>r\_\{i\}, i=1,…,k\-1;
- rkr\_\{k\}– strong relative toℛ\\mathcal\{R\}\.

Finally, among all strong causal rules with a given head we distinguish rules with maximal probability\.

###### Definition 9\.

Letψ∈F​\(A​t\)\\psi\\in F\(At\)\. A strong causal rulerrwith headψ\\psiis called a maximal specific causal rule forψ\\psirelative toℛ\\mathcal\{R\}, if its probabilityν​\(r\)\\nu\(r\)is greatest among all strong causal rules with headψ\\psi\.

We will say thatrris a maximal specific causal rule relative toℛ\\mathcal\{R\}ifrris a maximal specific causal rule relative toℛ\\mathcal\{R\}forH​\(r\)H\(r\)\. The set of all maximal specific causal rules we denote asMSCR\. Rules fromMSCRwe consider as satisfying the Requirement of Maximal Specificity, because a specification of such rules does not lead to an increase of probability, which means that their bodies contains all statistically relevant information for the prediction of its head\.

###### Proposition 4\.

LetA​tAtbe finite\. For every rulerrsuch thatH​\(r\)=ψH\(r\)=\\psiand the valueν​\(r\)\\nu\(r\)is defined, there exists a maximal specific causal ruler′r^\{\\prime\}forψ\\psisuch thatν​\(r′\)≥ν​\(r\)\\nu\(r^\{\\prime\}\)\\geq\\nu\(r\)\.

###### Proof\.

According to Proposition[3](https://arxiv.org/html/2607.12826#Thmproposition3)the set

Δ=\{\[α\]𝔊∣α⇒ψ​is a probabilistic causal rule\}\\Delta=\\\{\[\\alpha\]\_\{\\mathfrak\{G\}\}\\mid\\alpha\\Rightarrow\\psi\\ \\mbox\{is a probabilistic causal rule\}\\\}is non\-empty\. It follows immediately from Definition[7](https://arxiv.org/html/2607.12826#Thmdefinition7)thatα⇒ψ\\alpha\\Rightarrow\\psiis a strong causal rule iff\[α\]𝔊\[\\alpha\]\_\{\\mathfrak\{G\}\}is a minimal element ofΔ\\Deltaw\.r\.t\.⊑𝔊\\sqsubseteq\_\{\\mathfrak\{G\}\}\. SinceΔ\\Deltais finite, the set of its⊑𝔊\\sqsubseteq\_\{\\mathfrak\{G\}\}\-minimal elements is a finite non\-empty set\. So we can choose in this set an element\[β\]𝔊\[\\beta\]\_\{\\mathfrak\{G\}\}with the greatest valueν​\(β⇒ψ\)\\nu\(\\beta\\Rightarrow\\psi\)\. This is a required maximal specific causal ruler′r^\{\\prime\}forψ\\psi\. Thatν​\(β⇒ψ\)≥ν​\(r\)\\nu\(\\beta\\Rightarrow\\psi\)\\geq\\nu\(r\)follows again from Proposition[3](https://arxiv.org/html/2607.12826#Thmproposition3)\. ∎

## 6Result

LetA​tAt, a set𝔊\\mathfrak\{G\}of models, and a measureμ\\muon𝔊\\mathfrak\{G\}be fixed\. The set of all maximal specific causal rules relative toℛ\\mathcal\{R\}in that case we denote as𝖬𝖲𝖢𝖱​\(ℛ,𝔊,μ\)\{\\sf MSCR\}\(\\mathcal\{R\},\\mathfrak\{G\},\\mu\)\. For the set of rulesΠ⊆𝖬𝖲𝖢𝖱​\(ℛ,𝔊,μ\)\\Pi\\subseteq\{\\sf MSCR\}\(\\mathcal\{R\},\\mathfrak\{G\},\\mu\)andT⊆F​\(A​t\)T\\subseteq F\(At\)we define an operator ofdirect predictions:

PrΠ\(T\)=T∪\{H\(r\)∣r∈Π,∃φ1,…,φn∈T\(𝔊⊧\(φ1&…&φn\)↔B\(r\)\)\}Pr\_\{\\Pi\}\(T\)=T\\cup\\\{H\(r\)\\mid r\\in\\Pi,\\exists\\varphi\_\{1\},\.\.\.,\\varphi\_\{n\}\\in T\(\\mathfrak\{G\}\\models\(\\varphi\_\{1\}\\&\\ldots\\&\\varphi\_\{n\}\)\\leftrightarrow B\(r\)\)\\\}
Further, we put:

P​rΠ0​\(T\)=T,P​rΠn\+1​\(T\)=P​rΠ​\(P​rΠn​\(T\)\),Pr^\{0\}\_\{\\Pi\}\(T\)=T,\\ Pr^\{n\+1\}\_\{\\Pi\}\(T\)=Pr\_\{\\Pi\}\(Pr^\{n\}\_\{\\Pi\}\(T\)\),P​RΠ​\(T\)=⋃n∈ωP​rΠn​\(T\)\.PR\_\{\\Pi\}\(T\)=\\bigcup\_\{n\\in\\omega\}Pr^\{n\}\_\{\\Pi\}\(T\)\.We callP​RΠPR\_\{\\Pi\}aprediction operator forΠ\\Pi\.

###### Theorem 1\.

LetA​tAtbe a finite set of atoms,Π⊆𝖬𝖲𝖢𝖱​\(ℛ,𝔊,μ\)\\Pi\\subseteq\{\\sf MSCR\}\(\\mathcal\{R\},\\mathfrak\{G\},\\mu\), andT⊆F​\(A​t\)T\\subseteq F\(At\)\. IfTTis𝔊\\mathfrak\{G\}\-consistent, thenP​RΠ​\(T\)PR\_\{\\Pi\}\(T\)is𝔊\\mathfrak\{G\}\-consistent too\.

###### Proof\.

Obviously, it will be enough to check that the operator of direct predictions produces a𝔊\\mathfrak\{G\}\-consistent set of formulas\. First of all we show that the setTTof formulas and the setΠ\\Piof rules can be replaced by finite sets\. Consider the family of cosets\{\[φ\]𝔊∣φ∈T\}\\\{\[\\varphi\]\_\{\\mathfrak\{G\}\}\\mid\\varphi\\in T\\\}, which is finite sinceA​tAtis finite\. For every\[φ\]𝔊\[\\varphi\]\_\{\\mathfrak\{G\}\}from this family choose a single representative and put it intoT′T^\{\\prime\}\. Now we consider the set of pairs:

Θ=\{\(\[φ\]𝔊,\[ψ\]𝔊\)∣φ⇒ψ∈Π\.\}\\Theta=\\\{\(\[\\varphi\]\_\{\\mathfrak\{G\}\},\[\\psi\]\_\{\\mathfrak\{G\}\}\)\\mid\\varphi\\Rightarrow\\psi\\in\\Pi\.\\\}A finite set of rulesΠ′\\Pi^\{\\prime\}is defined as follows\. For every pair of cosets\(\[φ\]𝔊,\[ψ\]𝔊\)∈Θ\(\[\\varphi\]\_\{\\mathfrak\{G\}\},\[\\psi\]\_\{\\mathfrak\{G\}\}\)\\in\\Theta, we choose a single pair of representatives\(φ,ψ\)\(\\varphi,\\psi\)and put the ruleφ⇒ψ\\varphi\\Rightarrow\\psiintoΠ′\\Pi^\{\\prime\}\. It is clear that for everyψ∈P​rΠ​\(T\)\\psi\\in Pr\_\{\\Pi\}\(T\)there is aψ′\\psi^\{\\prime\}inP​rΠ′​\(T′\)Pr\_\{\\Pi^\{\\prime\}\}\(T^\{\\prime\}\)such thatψ≡𝔊ψ′\\psi\\equiv\_\{\\mathfrak\{G\}\}\\psi^\{\\prime\}\. In this way, ifP​rΠ′​\(T′\)Pr\_\{\\Pi^\{\\prime\}\}\(T^\{\\prime\}\)is𝔊\\mathfrak\{G\}\-consistent, thenP​rΠ​\(T\)Pr\_\{\\Pi\}\(T\)is𝔊\\mathfrak\{G\}\-consistent too\. Further, letΠ′′\\Pi^\{\\prime\\prime\}be the set of such rulesrrfromΠ′\\Pi^\{\\prime\}that the equivalence\(φ1&…&φn\)→B​\(r\)\(\\varphi\_\{1\}\\&\\ldots\\&\\varphi\_\{n\}\)\\rightarrow B\(r\)holds on𝔊\\mathfrak\{G\}for someφ1\\varphi\_\{1\},…,φn∈T′\\varphi\_\{n\}\\in T^\{\\prime\}\.

LetΠ′′=\{r1,…,rn\}\\Pi^\{\\prime\\prime\}=\\\{r\_\{1\},\\ldots,r\_\{n\}\\\}\. We putT0=T′T\_\{0\}=T^\{\\prime\},Ti\+1=Ti∪\{H​\(ri\)\}T\_\{i\+1\}=T\_\{i\}\\cup\\\{H\(r\_\{i\}\)\\\}\. Clearly,Tn=P​rΠ′​\(T′\)T\_\{n\}=Pr\_\{\\Pi^\{\\prime\}\}\(T^\{\\prime\}\)\. Using induction oniiwe show that everyTiT\_\{i\}is𝔊\\mathfrak\{G\}\-consistent\.

Assume thatTiT\_\{i\}is𝔊\\mathfrak\{G\}\-consistent, butTi\+1T\_\{i\+1\}is not\. Letri=φ⇒ψr\_\{i\}=\\varphi\\Rightarrow\\psi\. By definition ofΠ′′\\Pi^\{\\prime\\prime\}there areχ1,…,χn∈Ti\\chi\_\{1\},\\ldots,\\chi\_\{n\}\\in T\_\{i\}such that\(χ1&…&χn\)↔φ\(\\chi\_\{1\}\\&\\ldots\\&\\chi\_\{n\}\)\\leftrightarrow\\varphiholds on𝔊\\mathfrak\{G\}\. LetN=Ti∖\{χ1,…​χn\}N=T\_\{i\}\\setminus\\\{\\chi\_\{1\},\\ldots\\chi\_\{n\}\\\}\. Assume that\{φ,¬\(⋀N\)\}\\\{\\varphi,\\neg\(\\bigwedge N\)\\\}is𝔊\\mathfrak\{G\}\-consistent, i\.e\.,ν​\(φ&¬\(⋀N\)\)≠0\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\)\\neq 0\. In this case fors=φ&¬\(⋀N\)⇒ψs=\\varphi\\&\\neg\(\\bigwedge N\)\\Rightarrow\\psiwe have:

ν​\(s\)=ν​\(φ&¬\(⋀N\)&ψ\)ν​\(φ&¬\(⋀N\)\)=ν​\(φ&ψ\)−ν​\(φ&⋀N&ψ\)ν​\(φ\)−ν​\(φ&⋀N\)\.\\nu\(s\)=\\frac\{\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\\&\\psi\)\}\{\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\)\}=\\frac\{\\nu\(\\varphi\\&\\psi\)\-\\nu\(\\varphi\\&\\bigwedge N\\&\\psi\)\}\{\\nu\(\\varphi\)\-\\nu\(\\varphi\\&\\bigwedge N\)\}\.
We have𝔊⊧⋀Ti\+1↔\(φ&⋀N&ψ\)\\mathfrak\{G\}\\models\\bigwedge T\_\{i\+1\}\\leftrightarrow\(\\varphi\\&\\bigwedge N\\&\\psi\)and𝔊⊧⋀Ti↔\(φ&⋀N\)\\mathfrak\{G\}\\models\\bigwedge T\_\{i\}\\leftrightarrow\(\\varphi\\&\\bigwedge N\)by choice ofχ1,…,χn\\chi\_\{1\},\\ldots,\\chi\_\{n\}\. Since by assumptionν​\(⋀Ti\+1\)=0\\nu\(\\bigwedge T\_\{i\+1\}\)=0andν​\(⋀Ti\)≠0\\nu\(\\bigwedge T\_\{i\}\)\\neq 0, we conclude thatν​\(φ&⋀N&ψ\)=0\\nu\(\\varphi\\&\\bigwedge N\\&\\psi\)=0andν​\(φ&⋀N\)≠0\\nu\(\\varphi\\&\\bigwedge N\)\\neq 0\. In this way, we have

ν​\(s\)=ν​\(φ&ψ\)ν​\(φ\)−ν​\(φ&⋀N\)\>ν​\(φ&ψ\)ν​\(φ\)=ν​\(ri\)\.\\nu\(s\)=\\frac\{\\nu\(\\varphi\\&\\psi\)\}\{\\nu\(\\varphi\)\-\\nu\(\\varphi\\&\\bigwedge N\)\}\>\\frac\{\\nu\(\\varphi\\&\\psi\)\}\{\\nu\(\\varphi\)\}=\\nu\(r\_\{i\}\)\.
Sinces≺ris\\prec r\_\{i\}andν​\(s\)\>ν​\(ri\)\\nu\(s\)\>\\nu\(r\_\{i\}\), then there isr′∈ℛr^\{\\prime\}\\in\\mathcal\{R\}such thatr′⪯sr^\{\\prime\}\\preceq sandν​\(r′\)\>ν​\(ri\)\\nu\(r^\{\\prime\}\)\>\\nu\(r\_\{i\}\)\. On the other hand, fromri∈𝖬𝖲𝖢𝖱​\(ℛ,𝔊,μ\)r\_\{i\}\\in\{\\sf MSCR\}\(\\mathcal\{R\},\\mathfrak\{G\},\\mu\)andri≻r′r\_\{i\}\\succ r^\{\\prime\}we obtainν​\(ri\)≥ν​\(r′\)\\nu\(r\_\{i\}\)\\geq\\nu\(r^\{\\prime\}\)\. This contradiction proves that the body ofssis not𝔊\\mathfrak\{G\}\-consistent:ν​\(φ&¬\(⋀N\)\)=0\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\)=0\. As a consequence we obtainν​\(φ&¬\(⋀N\)&ψ\)=0\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\\&\\psi\)=0\. Now we have:

ν​\(φ&ψ\)=ν​\(φ&ψ\)−ν​\(φ&¬\(⋀N\)&ψ\)=ν​\(φ&⋀N&ψ\)=0\.\\nu\(\\varphi\\&\\psi\)=\\nu\(\\varphi\\&\\psi\)\-\\nu\(\\varphi\\&\\neg\(\\bigwedge N\)\\&\\psi\)=\\nu\(\\varphi\\&\\bigwedge N\\&\\psi\)=0\.Thus,ν​\(ri\)=0\\nu\(r\_\{i\}\)=0\. At the same timerir\_\{i\}is a causal rule, which implies0=ν\(ri\)\>ν\(⊤⇒ψ\)≥00=\\nu\(r\_\{i\}\)\>\\nu\(\\top\\Rightarrow\\psi\)\\geq 0\. The obtained contradiction concludes the proof\. ∎

## 7Conclusion

Analysis of the history surrounding the RMS discussions has made it possible to precisely formalize and resolve Carl Hempel’s statistical ambiguity problem\. This result does not only sum up the historical debate about RMS requirements, but also may open new directions in the philosophy of science\. It follows from the result that I\-S inference can discover logically consistent natural\-scientific theories\. It also follows that cyclic causal relations, which reflect the integrity of objects of the perceived world, loop back on themselves and form consistent “causal models” of the external world categories\[Rehder,Rehder2\]\. Based on these causal models a “probabilistic formal concepts” were defined that provide the idealized description of categories\[VitDemPon,Author12\]\. The causal relations between actions and their outcomes describe the goal\-directed activity of humans and animals in accordance with the physiological Theory of Functional Systems\[Nadin\]\.

For a class of rules of the formα1∧…∧αn⇒β\\alpha\_\{1\}\\wedge\\ldots\\wedge\\alpha\_\{n\}\\Rightarrow\\beta, whereαi\\alpha\_\{i\}andβ\\betaare literals \(atoms or their negations\), a program system “Discovery” was developed that discovers a set of MSCRs on a sample data D for prediction of some goal propertyψ\\psi, using the exact Fisher test for statistical estimation of probabilistic inequalities of definition 8\[mind\]\. This system was successfully applied to solve several tasks such as financial forecasting\[KovVit00\], medicine\[KovVitR01\], and bio informatics\[VitOVPK02\]\.

Based on𝖬𝖲𝖢𝖱\{\\sf MSCR\}a more powerful program system may be developed for solving Causal AI tasks\. For example, for the digital twins control systems development, where decisions based on predictions are rather important\.

## References

Similar Articles

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Hugging Face Daily Papers

CausaLab is a scalable environment for evaluating LLM agents on interactive causal discovery, assessing both predictive accuracy and faithful recovery of underlying causal mechanisms. Experiments reveal a gap between prediction and mechanism recovery, highlighting limits in current LLM agents as experimental causal reasoners.