Evaluating Bivariate Causal Statements Based on Mutual Compatibility
Summary
This paper introduces compatibility and incompatibility scores for evaluating collections of bivariate causal statements without relying on faithfulness, and demonstrates their applicability by analyzing causal claims from large language models.
View Cached Full Text
Cached at: 06/02/26, 03:46 PM
# Evaluating Bivariate Causal Statements Based on Mutual Compatibility
Source: [https://arxiv.org/abs/2606.00278](https://arxiv.org/abs/2606.00278)
[View PDF](https://arxiv.org/pdf/2606.00278)
> Abstract:For many real\-world systems, causal ground truth is difficult to obtain, making claims about causal effects hard to assess\. We develop methods for evaluating collections of $\\binom\{n\}\{2\}$ bivariate causal statements over a set of $n$ variables\. In the setting of acyclic linear statements, any such collection can be extended to a unique multivariate causal model, but we argue that this induced model is implausible if it imposes substantial additional confounding to explain observed correlations\. We introduce a compatibility score that quantifies this notion of plausibility, notably without relying on the faithfulness assumption\. Additionally, we define an incompatibility score for purely graphical bivariate causal statements, based on global consistency constraints that are derived from acyclicity and faithfulness assumptions\. We give theoretical and empirical evidence that both scores can successfully distinguish correct from incorrect causal statements in generic settings\. Moreover, we demonstrate the practical applicability of our methods by analyzing causal claims made by large language models\. Our work aims to provide a foundation for assessing the reliability of causal information derived from human experts or artificial intelligence in settings where alternative forms of validation are unavailable\.
## Submission history
From: Dominik Janzing \[[view email](https://arxiv.org/show-email/7bf22f62/2606.00278)\] **\[v1\]**Fri, 29 May 2026 19:15:09 UTC \(2,709 KB\)Similar Articles
Score-Based Causal Discovery of Latent Variable Causal Models
This paper introduces score-based methods for causal discovery in the presence of latent variables, offering theoretical guarantees of consistency and score equivalence, and unifies several constraint-based approaches.
Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States
This paper investigates whether language models can correctly judge how diagnostic evidence supports or challenges different causal claims, introducing paired prompts that vary only the causal target. Linear readouts from the penultimate transformer hidden state on models like Qwen2.5-7B-Instruct show moderate balanced accuracy (0.654-0.659) and recover 18–21 out of 49 pairs, indicating some linear decodability of causal relevance.
General Probabilities of Causation with Causal Knowledge
This paper derives tighter bounds for multivalued probabilities of causation by incorporating causal information from covariates and mediators, extending prior work from binary settings.
Medical Causal Hypothesis Verification with Large Language Models
This paper presents a preliminary study evaluating the accuracy of large language models in verifying causal medical hypotheses, finding that while they exhibit strong recall, they often fail to provide valid scientific evidence or reject unsupported claims.
When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models
This paper investigates how large language models internally represent logical validity, showing that validity information is decodable from hidden states despite poor behavioral performance, suggesting distinct roles for representation, expression, and causal use.