Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Summary
This paper formalises counterfactual policy optimisation for Markov Decision Processes under probabilistic nondeterministic causal models, which separate latent confounding from inherent stochasticity, and proposes a practical optimisation procedure for deriving robust counterfactual policies. The approach is validated on a sepsis treatment simulator with diabetes as an unobserved global confounder.
View Cached Full Text
Cached at: 08/05/26, 07:42 AM
# Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Source: [https://arxiv.org/html/2608.02893](https://arxiv.org/html/2608.02893)
Milad KazemiDurham UniversityNicola PaolettiKing’s College LondonDavid WatsonKing’s College LondonSander BeckersUniversity College London
###### Abstract
Counterfactual inference approaches for sequential decision\-making typically assume deterministic causal models, where all randomness stems from latent variables\. However, Markov Decision Processes \(MDPs\) are inherently stochastic\. We address this by formalising counterfactual policy optimisation under probabilistic nondeterministic causal models, which properly separates latent confounding from irreducible stochasticity, and here propose a first practical optimisation problem for identifying robust counterfactual policies under a sensitivity analysis framework\. We validate our approach on a sepsis treatment simulator, where diabetes status acts as a hidden global confounder\.
## 1Introduction
Counterfactual inference in sequential decision\-making asks: given an observed sequence of states, actions, and outcomes, what would the outcome have been under a different policy? In safety\-critical domains, such as healthcare, rerunning experiments under alternative policies is often infeasible or unethical, making counterfactual reasoning crucial for offline policy evaluation\.
Existing work applying counterfactual inference to MDPs\[Oberst and Sontag,[2019](https://arxiv.org/html/2608.02893#bib.bib12), Tsirtsis et al\.,[2021](https://arxiv.org/html/2608.02893#bib.bib15), Killian et al\.,[2022](https://arxiv.org/html/2608.02893#bib.bib8), Tsirtsis and Rodriguez,[2024](https://arxiv.org/html/2608.02893#bib.bib14), Lally et al\.,[2026](https://arxiv.org/html/2608.02893#bib.bib9)\]assumes all randomness comes from unobserved latent variables, and therefore models MDPs as deterministic structural causal models \(SCMs\)\[Pearl,[2009](https://arxiv.org/html/2608.02893#bib.bib13)\]\. However, in practice, randomness can arise from fundamentally different sources: at one extreme, all randomness may arise from unobserved latent variables \(the deterministic SCM setting\); at the other, all randomness may be irreducible \(e\.g\., from random number generators\)\. Recently, probabilistic nondeterministic causal models \(PNSCMs\)\[Beckers,[2026](https://arxiv.org/html/2608.02893#bib.bib3),[2025b](https://arxiv.org/html/2608.02893#bib.bib2)\]have been developed to allow for such inherent stochasticity\. More commonly, however, both sources of randomness coexist: health outcomes usually depend both on unobserved latent factors, and on random differences in applying the treatment, or random individual\-level genetic variation\. Similarly, in the social sciences, variation in the outcome of policies is both due to latent factors and to individual\-level variation that is represented as noise at the aggregate level\.
In these settings, we propose modelling the environment as a PNSCM where randomness arises from both latent variables and irreducible stochasticity, and bounding the influence of latent variables using a sensitivity\-analysis approach similar toKausik et al\. \[[2024](https://arxiv.org/html/2608.02893#bib.bib7)\]\. In this paper, we formalise counterfactual inference for MDPs under PNSCMs, and propose optimisation procedures for deriving robust counterfactual policies under confounding\. We then evaluate these approaches on the Sepsis MDP\[Oberst and Sontag,[2019](https://arxiv.org/html/2608.02893#bib.bib12)\], where diabetes acts as an unobserved global confounder, demonstrating that the derived policies are robust when evaluated on the true, fully\-specified environment\.
## 2Background
In this section, we introduce the necessary background on Markov decision processes and causal models\.
### 2\.1Markov Decision Processes
Markov decision processes \(MDPs\) model sequential decision\-making under uncertainty, and are defined by a tuple\(𝒮,𝒜,𝒫I,𝒫,ℛ,γ\)\(\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{P\}\_\{I\},\\mathcal\{P\},\\mathcal\{R\},\\gamma\)where𝒮\\mathcal\{S\}is the state space,𝒜\\mathcal\{A\}is the action space,𝒫I\\mathcal\{P\}\_\{I\}is the initial state distribution,𝒫:𝒮×𝒜×𝒮→\[0,1\]\\mathcal\{P\}:\\mathcal\{S\}\\times\\mathcal\{A\}\\times\\mathcal\{S\}\\rightarrow\[0,1\]is the transition kernel,ℛ:𝒮×𝒜→ℝ\\mathcal\{R\}:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathbb\{R\}is a reward function, andγ\\gammais a discount factor\. The goal is to identify an optimal policyπ:𝒮→𝒜\\pi:\\mathcal\{S\}\\rightarrow\\mathcal\{A\}, that optimises the expected total discounted reward\. A trajectoryτ\\tauunder policyπ\\piis a sequenceτ=\(s0,a0,s1,a1,…,s\|τ\|−1,a\|τ\|−1,s\|τ\|\)\\tau=\(s\_\{0\},a\_\{0\},s\_\{1\},a\_\{1\},\\ldots,s\_\{\|\\tau\|\-1\},a\_\{\|\\tau\|\-1\},s\_\{\|\\tau\|\}\)wheres0∼𝒫Is\_\{0\}\\sim\\mathcal\{P\}\_\{I\},at=π\(st\)a\_\{t\}=\\pi\(s\_\{t\}\), andst\+1∼𝒫\(⋅∣st,at\)s\_\{t\+1\}\\sim\\mathcal\{P\}\(\\cdot\\mid s\_\{t\},a\_\{t\}\)\.
#### Latent Variables and Confounding
In many applications, the observed state cannot capture all factors influencing system dynamics, e\.g\., electronic health records may have missing information on patient demographics or underlying conditions\. We consider settings where latent factorsUUaffect transition dynamics, such that the true transitionsP\(s′∣s,a,u\)P\(s^\{\\prime\}\\mid s,a,u\)differ from the observed marginalP\(s′∣s,a\)P\(s^\{\\prime\}\\mid s,a\)\. In this initial paper, we focus onglobal confounders, where the value ofUUis fixed throughout the trajectory, confounding all transitions\[Kausik et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib7)\]\.
SinceUUis unobserved, its effect on transitions cannot be identified from observational data alone\. Sensitivity analysis addresses this problem by bounding the degree to whichUUcan influence the transitions\. Existing work on sensitivity analysis of confounding in RL primarily focuses on*off\-policy evaluation*\[Bruns\-Smith,[2021](https://arxiv.org/html/2608.02893#bib.bib6), Namkoong et al\.,[2020](https://arxiv.org/html/2608.02893#bib.bib11), Kausik et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib7), Bruns\-Smith and Zhou,[2023](https://arxiv.org/html/2608.02893#bib.bib5)\]: estimating the expected value of a target policy, given a dataset of observed trajectories under an unknown behaviour policy\. Here, we adapt this sensitivity framework for*counterfactual policy analysis*\[Oberst and Sontag,[2019](https://arxiv.org/html/2608.02893#bib.bib12), Tsirtsis et al\.,[2021](https://arxiv.org/html/2608.02893#bib.bib15), Tsirtsis and Rodriguez,[2024](https://arxiv.org/html/2608.02893#bib.bib14), Lally et al\.,[2026](https://arxiv.org/html/2608.02893#bib.bib9)\]: given an observed trajectory, what might have happened under a different policy, and what policy would have been optimal?
### 2\.2Structural Causal Models
Structural causal models \(SCMs\)\[Pearl,[2009](https://arxiv.org/html/2608.02893#bib.bib13)\]provide a framework for counterfactual inference\. An SCM𝒞=\(𝐔,𝐕,ℱ,P\(𝐔\)\)\\mathcal\{C\}=\(\\mathbf\{U\},\\mathbf\{V\},\\mathcal\{F\},P\(\\mathbf\{U\}\)\)consists of observed variables𝐕\\mathbf\{V\}, unobserved variables𝐔\\mathbf\{U\}with joint distributionP\(𝐔\)P\(\\mathbf\{U\}\), and structural equations\(X=fX\(𝐩𝐚𝐗\)\)∈ℱ\(X=f\_\{X\}\(\\mathbf\{pa\_\{X\}\}\)\)\\in\\mathcal\{F\}that determine the values of eachX∈𝐕X\\in\\mathbf\{V\}as a function of its direct causes – parents –𝐩𝐚X\\mathbf\{pa\}\_\{X\}\. Counterfactual inference proceeds by estimatingP\(𝐔∣𝐯\)P\(\\mathbf\{U\}\\mid\\mathbf\{v\}\)given an observation𝐕=𝐯\\mathbf\{V\}=\\mathbf\{v\}, performing an intervention that modifies the structural equations, and evaluating the values of observed variables\.
*Probabilistic nondeterministic SCMs*\(PNSCMs\)\[Beckers,[2025b](https://arxiv.org/html/2608.02893#bib.bib2)\]replace deterministic structural equationsfXf\_\{X\}with conditional probability distributions𝒫X\(⋅∣𝐩𝐚X\)\\mathcal\{P\}\_\{X\}\(\\cdot\\mid\\mathbf\{pa\}\_\{X\}\), thereby allowing probabilistic causal mechanisms that are not reducible to deterministic mechanisms over unobserved variables\. A fully specified MDP together with a policy is an instance of a PNSCM\. A key feature of counterfactual inference with PNSCMs isactualised refinement: given the values of the observed and unobserved variables that produce the actual setting\(𝐮,𝐯\)\(\\mathbf\{u\},\\mathbf\{v\}\), actualised refinement replaces each conditional distribution𝒫X\\mathcal\{P\}\_\{X\}by𝒫X\(𝐩𝐚X,x\)\\mathcal\{P\}\_\{X\}^\{\(\\mathbf\{pa\}\_\{X\},x\)\}, which behaves identically to𝒫X\\mathcal\{P\}\_\{X\}for all inputs except that it deterministically returns the observed valuexxwhen given the observed parents𝐩𝐚X\\mathbf\{pa\}\_\{X\}\. This expresses the nondeterministic semantics thatallwe learn from the actual setting is that the actual parents resulted in the actual child, and nothing else \(see\[Beckers,[2025a](https://arxiv.org/html/2608.02893#bib.bib1),[b](https://arxiv.org/html/2608.02893#bib.bib2)\]for more details and motivation\.\)
## 3Methodology
Given an observed trajectoryτ=\(s0,a0,…,s\|τ\|\)\\tau=\(s\_\{0\},a\_\{0\},\\ldots,s\_\{\|\\tau\|\}\), our goal is to identify the optimal policyπ\\pithat maximises the worst\-case counterfactual valueV𝐶𝐹\(0,s0\)V^\{\\it CF\}\(0,s\_\{0\}\)under uncertainty over the unobserved global confounderUU:
V𝐶𝐹\(t,s\)=maxπminPt𝐶𝐹∈𝒫t𝐶𝐹,Δ\\displaystyle V^\{\\mathit\{CF\}\}\(t,s\)=\\textstyle\\max\_\{\\pi\}\\textstyle\\min\_\{\{P\}^\{\\it CF\}\_\{t\}\\in\\mathcal\{P\}^\{\\it CF,\\Delta\}\_\{t\}\}\(1\)𝔼s′∼Pt𝐶𝐹\(⋅∣s,π\(s\)\)\[ℛ\(s,π\(s\)\)\+γ⋅V𝐶𝐹\(t\+1,s′\)\]\\displaystyle\\mathbb\{E\}\_\{s^\{\\prime\}\\sim\{P\}^\{\\it CF\}\_\{t\}\(\\cdot\\mid s,\\pi\(s\)\)\}\\left\[\\mathcal\{R\}\(s,\\pi\(s\)\)\+\\gamma\\cdot V^\{\\mathit\{CF\}\}\(t\+1,s^\{\\prime\}\)\\right\]where𝒫t𝐶𝐹,Δ\\mathcal\{P\}^\{\\it CF,\\Delta\}\_\{t\}is the set of feasible counterfactual transition probabilities at timett, whose constraints are defined below\.
### 3\.1Sensitivity Constraints
Similar to the odds\-ratio model used in\[Bruns\-Smith,[2021](https://arxiv.org/html/2608.02893#bib.bib6), Bruns\-Smith and Zhou,[2023](https://arxiv.org/html/2608.02893#bib.bib5), Kausik et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib7), Bennett et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib4)\], we parametrise the influence ofUUvia a sensitivity parameterΔ≥1\\Delta\\geq 1, which constrains how muchP\(s′∣s,a,u\)P\(s^\{\\prime\}\\mid s,a,u\)can differ from the marginalP\(s′∣s,a\)P\(s^\{\\prime\}\\mid s,a\):
1Δ≤odds\(P\(S′=s′∣S=s,A=a,U=u\)\)odds\(P\(S′=s′∣S=s,A=a\)\)≤Δ\\frac\{1\}\{\\Delta\}\\leq\\frac\{\\text\{odds\}\(P\(S^\{\\prime\}=s^\{\\prime\}\\mid S=s,A=a,U=u\)\)\}\{\\text\{odds\}\(P\(S^\{\\prime\}=s^\{\\prime\}\\mid S=s,A=a\)\)\}\\leq\\Deltawhereodds\(p\)=p/\(1−p\)\\text\{odds\}\(p\)=p/\(1\-p\)\. WhenΔ=1\\Delta=1, this is equivalent to a fully specified PNSCM \(i\.e\., no latentUU\)\. AsΔ→∞\\Delta\\rightarrow\\infty, the bounds onP\(S′∣S,A,U\)P\(S^\{\\prime\}\\mid S,A,U\)approach\[0,1\]\[0,1\], representing maximum confounding\.Δ\\Deltacan be global – i\.e\., one value – or state\-action dependentΔ\(s,a\)\\Delta\(s,a\)\[Bennett et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib4)\]\.
### 3\.2ChoosingΔ\\Delta
A key challenge in sensitivity analysis is choosingΔ\\Delta\. Most existing work either delegates this to a domain expert, based on their understanding of the environment dynamics \(e\.g\.,\[Bruns\-Smith,[2021](https://arxiv.org/html/2608.02893#bib.bib6), Bennett et al\.,[2024](https://arxiv.org/html/2608.02893#bib.bib4)\]\), or evaluates over a range fromΔ=1\\Delta=1\(no confounding\) to some large value \(strong confounding\), assessing how the estimated value function changes\. If the target policy consistently outperforms the observed policy, this demonstrates robustness\. Without domain knowledge, a more principled approach is to calibrateΔ\\Deltafrom the observed state variables\[Bruns\-Smith and Zhou,[2023](https://arxiv.org/html/2608.02893#bib.bib5)\]by hiding each variable in turn and computing the resulting odds ratio between the observed and marginalised transitions\. Under the assumption that no unobserved variable influences transitions more than the strongest observed variable, this provides a justifiable upper bound onΔ\(s,a\)\\Delta\(s,a\)\. Formally, for each variableii, lets−is^\{\-i\}denote the state omitting variableii, and letnin\_\{i\}denote the number of distinct values variableiican take\. The marginalised transition distribution when variableiiis hidden is:
P\(St\+1=s′∣St−i=s−i,At=a\)=∑j=0ni−1wi\(j∣s−i\)⋅\\displaystyle P\(S\_\{t\+1\}=s^\{\\prime\}\\mid S^\{\-i\}\_\{t\}=s^\{\-i\},A\_\{t\}=a\)=\\sum\_\{j=0\}^\{n\_\{i\}\-1\}w\_\{i\}\(j\\mid s^\{\-i\}\)\\cdot\(2\)P\(s′∣s−i,Si=j,a\)\\displaystyle P\(s^\{\\prime\}\\mid s^\{\-i\},S^\{i\}=j,a\)wherewi\(j∣s−i\)=P\(Si=j∣s−i\)w\_\{i\}\(j\\mid s^\{\-i\}\)=P\(S^\{i\}=j\\mid s^\{\-i\}\)is the conditional probability that variableiitakes valuejjgiven the projected states−is^\{\-i\}, estimated from the stationary distribution of the MDP\. The sensitivityΔ\(s,a\)\\Delta\(s,a\)can be computed as:
Δi\(s,a\)=maxs′∈𝒮P\(s′∣s,a\)\>0max\(ri\(s,a,s′\),ri\(s,a,s′\)−1\)\\Delta\_\{i\}\(s,a\)=\\max\_\{\\begin\{subarray\}\{c\}s^\{\\prime\}\\in\\mathcal\{S\}\\\\ P\(s^\{\\prime\}\\mid s,a\)\>0\\end\{subarray\}\}\\max\\left\(r\_\{i\}\(s,a,s^\{\\prime\}\),\\ r\_\{i\}\(s,a,s^\{\\prime\}\)^\{\-1\}\\right\)whereri\(s,a,s′\)=odds\(P\(St\+1=s′∣s,a\)\)odds\(P\(St\+1=s′∣s−i,a\)\)r\_\{i\}\(s,a,s^\{\\prime\}\)=\\dfrac\{\\text\{odds\}\\left\(P\(S\_\{t\+1\}=s^\{\\prime\}\\mid s,a\)\\right\)\}\{\\text\{odds\}\\left\(P\(S\_\{t\+1\}=s^\{\\prime\}\\mid s^\{\-i\},a\)\\right\)\}\.
The overall sensitivity for each state\-action pair isΔ\(s,a\)=maxiΔi\(s,a\)\.\\Delta\(s,a\)=\\max\_\{i\}\\Delta\_\{i\}\(s,a\)\.Additionally, a scaling factorΓ≥0\\Gamma\\geq 0can be applied to calibrate the sensitivity, giving a final sensitivity estimate ofΓ⋅Δ\(s,a\)\\Gamma\\cdot\\Delta\(s,a\)\[McClean et al\.,[2025](https://arxiv.org/html/2608.02893#bib.bib10)\]\. This provides an interpretable parameter to assess how counterfactual estimates and policies vary under stronger\(Γ\>1\)\(\\Gamma\>1\)or weaker\(Γ<1\)\(\\Gamma<1\)confounding than the proxyΔ\(s,a\)\\Delta\(s,a\)suggests\.
### 3\.3Actualised Refinement
Given an observed transition\(st,at,st\+1\)\(s\_\{t\},a\_\{t\},s\_\{t\+1\}\), if we were to take the same actionata\_\{t\}insts\_\{t\}in some counterfactual world, the next state would be the samest\+1s\_\{t\+1\}as well\. The reason is that no matter the unobserved valueU=uU=u, it remains constant across worlds\. Still, given that we do not know the actual valueU=uU=u, the counterfactual transition probability must reflect this posterior uncertainty: with probabilityP\(U∣τ\)P\(U\\mid\\tau\),U=uU=ugenerated the trajectory and the transition tost\+1s\_\{t\+1\}is certain; otherwise, with probability1−P\(U∣τ\)1\-P\(U\\mid\\tau\), the counterfactual probability equals the estimated probabilityP~\(s′∣st,at,u\)\\tilde\{P\}\(s^\{\\prime\}\\mid s\_\{t\},a\_\{t\},u\)as chosen by the optimisation\. As a result, the counterfactual transition probability is:
PtCF\(s′\|s,a,u\)=\{P\(U=u\|τ\)⋅𝟏\[s′=st\+1\]\+\(1−P\(U=u\|τ\)\)⋅P~\(s′\|s,a,u\)if\(s,a\)=\(st,at\)P~\(s′\|s,a,u\)otherwiseP^\{CF\}\_\{t\}\(s^\{\\prime\}\|s,a,u\)=\\begin\{cases\}\\begin\{aligned\} &P\(U=u\|\\tau\)\\cdot\\mathbf\{1\}\[s^\{\\prime\}=s\_\{t\+1\}\]\\\\ &\\quad\+\(1\-P\(U=u\|\\tau\)\)\\cdot\\tilde\{P\}\(s^\{\\prime\}\|s,a,u\)\\end\{aligned\}\\\\ \\qquad\\qquad\\qquad\\qquad\\text\{if \}\(s,a\)=\(s\_\{t\},a\_\{t\}\)\\\\\[4\.0pt\] \\tilde\{P\}\(s^\{\\prime\}\|s,a,u\)\\quad\\text\{otherwise\}\\end\{cases\}\(One exception is whenΔ=1\\Delta=1: this corresponds to a setting with no confounderUU, soPtCF\(st\+1∣st,at\)=1P^\{CF\}\_\{t\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\}\)=1\)\.
Mean Differenceδ1¯=VΔ,Γ𝐶𝐹\(0,s0\)−VΔ∗𝐶𝐹\(0,s0\)\\bar\{\\delta\_\{1\}\}=V^\{\\mathit\{CF\}\}\_\{\\Delta,\\Gamma\}\(0,s\_\{0\}\)\-V^\{\\mathit\{CF\}\}\_\{\\Delta^\{\*\}\}\(0,s\_\{0\}\)Mean Differenceδ2¯=V𝑓𝑢𝑙𝑙𝐶𝐹,πΔ,Γ∗\(0,s0\)−V𝑓𝑢𝑙𝑙𝐶𝐹,πΔ∗∗\(0,s0\)\\bar\{\\delta\_\{2\}\}=V^\{\\mathit\{CF\},\\pi^\{\*\}\_\{\\Delta,\\Gamma\}\}\_\{\\it full\}\(0,s\_\{0\}\)\-V^\{\\mathit\{CF\},\\pi^\{\*\}\_\{\\Delta\*\}\}\_\{\\it full\}\(0,s\_\{0\}\)No scaling \(Γ=1\\Gamma=1\)
00\.40\.40\.80\.81\.21\.21\.61\.622−5\-50551010ScaleΓ\\Gammaδ1¯\\bar\{\\delta\_\{1\}\}\(a\)Diabetic Trajectories00\.40\.40\.80\.81\.21\.21\.61\.622−10\-10−5\-50ScaleΓ\\Gammaδ¯2\\bar\{\\delta\}\_\{2\}\(b\)Nondiabetic Trajectories
Figure 1:Mean difference between counterfactual policy values and the oracle \(under true diabetes sensitivityΔ∗\\Delta^\{\*\}\), averaged across1010randomly sampled suboptimal trajectories\.
### 3\.4Optimisation Procedures
Due to the coupling ofP\(S′∣S,A,U\)P\(S^\{\\prime\}\\mid S,A,U\)between all state\-action pairs\(s,a\)\(s,a\)and all time stepstt, identifying globally optimal solutions is challenging on all but toy examples\. In this work, we therefore consider adecoupled approach: we simplify the optimisation problem by removing the time\-homogeneity assumption and the coupling between state\-action pairs\(s,a\)\(s,a\), treating each\(t,s,a\)\(t,s,a\)triple independently\. For each\(t,s\)\(t,s\)we solve:
V𝐶𝐹\(t,s\)=maxa\[R\(s,a\)\+γ⋅min\{P~\(st\+1∣st,at,u\)\}u\{P~\(s′∣s,a,u\)\}u,s′\\displaystyle V^\{\\it CF\}\(t,s\)=\\max\_\{a\}\\big\[R\(s,a\)\+\\gamma\\cdot\\textstyle\\min\_\{\\begin\{subarray\}\{c\}\\\{\\tilde\{P\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\},u\)\\\}\_\{u\}\\\\ \\\{\\tilde\{P\}\(s^\{\\prime\}\\mid s,a,u\)\\\}\_\{u,s^\{\\prime\}\}\\end\{subarray\}\}∑uP\(U=u\|τ\)∑s′PtCF\(s′\|s,a,u\)⋅V𝐶𝐹\(t\+1,s′\)\]\\displaystyle\\sum\_\{u\}P\(U=u\|\\tau\)\\sum\_\{s^\{\\prime\}\}P^\{CF\}\_\{t\}\(s^\{\\prime\}\|s,a,u\)\\cdot V^\{\\it CF\}\(t\+1,s^\{\\prime\}\)\\big\]where the posteriorP\(U=u∣τ\)P\(U=u\\mid\\tau\)is:
P\(U=u∣τ\)=P\(U=u\)∏tP~\(st\+1∣st,at,u\)∑u′P\(U=u′\)∏tP~\(st\+1∣st,at,u′\)P\(U=u\\mid\\tau\)=\\frac\{P\(U=u\)\\prod\_\{t\}\\tilde\{P\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\},u\)\}\{\\sum\_\{u^\{\\prime\}\}P\(U=u^\{\\prime\}\)\\prod\_\{t\}\\tilde\{P\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\},u^\{\\prime\}\)\}SinceP\(U∣τ\)P\(U\\mid\\tau\)is a continuous function of the adversary’s choice of transition probabilities, we compute its feasible range by setting each observed transitionP~\(st\+1∣st,at,u\)\\tilde\{P\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\},u\)to its extremes and treat it as an optimisation variable\. The adversary therefore jointly optimises overP\(U∣τ\)P\(U\\mid\\tau\)and\{P~\(⋅\|s,a,u\)\}u\\\{\\tilde\{P\}\(\\cdot\|s,a,u\)\\\}\_\{u\}at each\(t,s,a\)\(t,s,a\)block, subject to:
∑uP\(U=u\)⋅P~\(s′\|s,a,u\)=P\(s′\|s,a\),∀s′\\displaystyle\\textstyle\\sum\_\{u\}P\(U=u\)\\cdot\\tilde\{P\}\(s^\{\\prime\}\|s,a,u\)=P\(s^\{\\prime\}\|s,a\),\\quad\\forall s^\{\\prime\}∑s′P~\(s′\|s,a,u\)=1,∀u\\displaystyle\\textstyle\\sum\_\{s^\{\\prime\}\}\\tilde\{P\}\(s^\{\\prime\}\|s,a,u\)=1,\\quad\\forall u1Δ⋅P\(s′\|s,a\)≤P~\(s′\|s,a,u\)≤Δ⋅P\(s′\|s,a\),∀s′,u\\displaystyle\\tfrac\{1\}\{\\Delta\}\\cdot P\(s^\{\\prime\}\|s,a\)\\leq\\tilde\{P\}\(s^\{\\prime\}\|s,a,u\)\\leq\\Delta\\cdot P\(s^\{\\prime\}\|s,a\),\\quad\\forall s^\{\\prime\},uNote that at the observed transition\(s,a\)=\(st,at\)\(s,a\)=\(s\_\{t\},a\_\{t\}\),P\(U=u∣τ\)P\(U=u\\mid\\tau\)appears both as a posterior weight and insidePtCFP^\{CF\}\_\{t\}, creating an adversarial tradeoff that makes the optimisation non\-convex; we use Gurobi to guarantee a globally optimal solution\. The decoupling makes the problem tractable but produces a conservative lower bound onVCF\(0,s0\)V^\{\\mathit\{C\}F\}\(0,s\_\{0\}\): the decoupled adversary has strictly more freedom than the true coupled adversary \(which must use the sameP~\(s′∣s,a,u\)\\tilde\{P\}\(s^\{\\prime\}\\mid s,a,u\)andP\(U∣τ\)P\(U\\mid\\tau\)across all\(t,s,a\)\(t,s,a\)blocks\), and since greater adversarial freedom can only decrease the value, this guarantees a valid lower bound\.
#### Incorporating Time\-homogeneity
To obtain a tighter bound, we could incorporate time\-homogeneity by requiring a singleP\(s′∣s,a,u\)P\(s^\{\\prime\}\\mid s,a,u\)for each\(s,a\)\(s,a\)pair across all timesteps\. This creates a trilinear coupling between the posteriorP\(U∣τ\)P\(U\\mid\\tau\), transition probabilitiesP\(s′∣s,a,u\)P\(s^\{\\prime\}\\mid s,a,u\), and value function, which can be solved using gradient descent\. Since gradient descent finds only local minima, we cannot guarantee we will identify the true pessimistic value function\. However, the decoupled and gradient descent solutions will together bound the true pessimistic value function\.
## 4Sepsis Simulation
We demonstrate our approach on a sepsis treatment simulator\[Oberst and Sontag,[2019](https://arxiv.org/html/2608.02893#bib.bib12)\]\. Each state consists of four vital signs \(heart rate, blood pressure, oxygen concentration, and glucose levels\), categorised as low, normal, or high\. At each step, three treatments can be toggled on or off\. Rewards range from−10\-10\(death\) to1010\(discharge\) based on the number of out\-of\-range vital signs\. The full transition matrix isP\(S′∣S,A,U\)P\(S^\{\\prime\}\\mid S,A,U\), whereU∈\{0,1\}U\\in\\\{0,1\\\}indicates whether the patient is diabetic, which acts as a global latent confounder\. In this population,20%20\\%of patients are diabetic\.
We evaluate our approach on the marginalised transition matrix \(where diabetes is hidden\), given observed trajectories from a suboptimal policy \(which acts optimally with probability0\.50\.5, and randomly otherwise\)\. We use the maximum sensitivity of the observed vital signsΔ\(s,a\)\\Delta\(s,a\)\(see Section[3\.2](https://arxiv.org/html/2608.02893#S3.SS2)\) as a proxy for the unobserved diabetes sensitivity, and an optional scaling parameterΓ\\Gamma\. For each trajectory, we derive the optimal pessimistic counterfactual policyπΔ∗\\pi^\{\*\}\_\{\\Delta\}that maximises the worst\-case counterfactual value functionVΔ,Γ𝐶𝐹\(0,s0\)V^\{\\it CF\}\_\{\\Delta,\\Gamma\}\(0,s\_\{0\}\), and evaluate its performance on the fully\-observed MDP \(where the diabetes state is known\) to obtainV𝑓𝑢𝑙𝑙𝐶𝐹,πΔ∗\(0,s0\)V^\{\\it CF,\\pi^\{\*\}\_\{\\Delta\}\}\_\{\\it full\}\(0,s\_\{0\}\)\. We evaluate our approach along two axes: first, whetherπΔ∗\\pi^\{\*\}\_\{\\Delta\}improves on the observation; and second, how closely our estimatedVΔ,Γ𝐶𝐹\(0,s0\)V^\{\\it CF\}\_\{\\Delta,\\Gamma\}\(0,s\_\{0\}\)and policies match those obtained under the true diabetes sensitivityΔ∗\\Delta^\{\*\}\.
#### Improvement over Observation
Table[1](https://arxiv.org/html/2608.02893#S4.T1)shows that our approach can identify counterfactual policies that robustly improve on the observed suboptimal behaviour when evaluated on the fully\-observed MDP\. The estimated valuesVΔ,Γ𝐶𝐹\(0,s0\)V^\{\\it CF\}\_\{\\Delta,\\Gamma\}\(0,s\_\{0\}\)underestimate this improvement, particularly for nondiabetic trajectories as theΔ\(s,a\)\\Delta\(s,a\)proxy overestimates the confounding effect\. This is expected, as the marginalised transition matrix already closely reflects nondiabetic dynamics \(sinceP\(U=0\)=0\.8\)P\(U=0\)=0\.8\)\.
Table 1:Mean difference between counterfactual policy values and observed discounted returnG0G\_\{0\}\(γ=0\.9\\gamma=0\.9\), across1010suboptimal diabetic and nondiabetic trajectories\.Estimatedreports the mean differenceVΔ,Γ𝐶𝐹\(0,s0\)−G0V^\{\\it CF\}\_\{\\Delta,\\Gamma\}\(0,s\_\{0\}\)\-G\_\{0\};Actualreports the mean differenceV𝑓𝑢𝑙𝑙𝐶𝐹,πΔ∗\(0,s0\)−G0V^\{\\it CF,\\pi^\{\*\}\_\{\\Delta\}\}\_\{\\it full\}\(0,s\_\{0\}\)\-G\_\{0\}\.DiabeticNondiabeticEstimatedActualEstimatedActual3\.4±10\.93\.4\\pm 10\.97\.4±10\.77\.4\\pm 10\.7−0\.50±9\.1\-0\.50\\pm 9\.15\.0±7\.45\.0\\pm 7\.4
#### Comparison to Diabetes Sensitivity
Figures[1](https://arxiv.org/html/2608.02893#S3.F1)and[1](https://arxiv.org/html/2608.02893#S3.F1)assess how well our proxyΔ\(s,a\)\\Delta\(s,a\)approximates the true diabetes sensitivityΔ∗\\Delta^\{\*\}\(the oracle\)\. Essentially,δ1¯\\bar\{\\delta\_\{1\}\}\(the blue curve\) reflects how close our estimated counterfactual values are to the oracle, andδ2¯\\bar\{\\delta\_\{2\}\}\(the red curve\) reflects how close the policies derived from our proxyΔ\(s,a\)\\Delta\(s,a\)perform relative to the oracle in practice\. The oracle represents the most accurate estimates achievable by our robust pessimistic approach, if we knew the true sensitivityΔ∗\\Delta^\{\*\}\. SinceΔ\(s,a\)≥1\\Delta\(s,a\)\\geq 1, we clampΓ⋅Δ\(s,a\)\\Gamma\\cdot\\Delta\(s,a\)to a minimum of11\(soΓ=0\\Gamma=0corresponds to a global sensitivityΔ=1\\Delta=1, i\.e\., no confounding\)\. AtΓ=0\\Gamma=0,δ1¯\>0\\bar\{\\delta\_\{1\}\}\>0\(and is much greater than0for diabetic trajectories\), confirming that ignoring confounding leads to inaccurate estimates\. For diabetic trajectories,Γ=1\.0\\Gamma=1\.0minimises bothδ1¯\\bar\{\\delta\_\{1\}\}andδ2¯\\bar\{\\delta\_\{2\}\}, demonstrating that the unscaled proxy reliably estimates the true diabetes sensitivity\. For nondiabetic trajectories,δ1¯\\bar\{\\delta\_\{1\}\}is strongly negative atΓ=1\\Gamma=1, and approaches0as we reduceΓ\\Gamma, showing our estimates are conservative\. This is expected, as the marginalised transition matrix already closely reflects nondiabetic dynamics, so the true confounding effect for nondiabetics is weaker than the proxy suggests\. However, crucially,δ2¯\\bar\{\\delta\_\{2\}\}remains close to0across0<Γ≤10<\\Gamma\\leq 1, suggesting that despite the conservative value estimates, we recover policies close to the oracle\.
## 5Conclusion
In this paper, we introduced a framework for counterfactual policy optimisation under global confounding and demonstrated its robustness on the Sepsis MDP\. Future work will explore extending this framework to settings with different types of latent variables \(e\.g\., memoryless confounders, and confounders that persist over a subset of the time steps\) and evaluating the gradient descent\-based optimisation method to obtain more accurate counterfactual value estimates\.
## References
- Beckers \[2025a\]Sander Beckers\.Causal counterfactuals reconsidered\.*arXiv preprint arXiv:2512\.12804*, 2025a\.
- Beckers \[2025b\]Sander Beckers\.Nondeterministic causal models\.In Biwei Huang and Mathias Drton, editors,*Proceedings of the Fourth Conference on Causal Learning and Reasoning*, volume 275 of*Proceedings of Machine Learning Research*, pages 1532–1554\. PMLR, 07–09 May 2025b\.URL[https://proceedings\.mlr\.press/v275/beckers25b\.html](https://proceedings.mlr.press/v275/beckers25b.html)\.
- Beckers \[2026\]Sander Beckers\.Large language models as nondeterministic causal models\.In*Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning*, 2026\.
- Bennett et al\. \[2024\]Andrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun, and Kaiwen Wang\.Efficient and sharp off\-policy evaluation in robust markov decision processes\.In A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang, editors,*Advances in Neural Information Processing Systems*, volume 37, pages 112962–113000\. Curran Associates, Inc\., 2024\.[10\.52202/079017\-3590](https://arxiv.org/doi.org/10.52202/079017-3590)\.URL[https://doi\.org/10\.52202/079017\-3590](https://doi.org/10.52202/079017-3590)\.
- Bruns\-Smith and Zhou \[2023\]David Bruns\-Smith and Angela Zhou\.Robust fitted\-q\-evaluation and iteration under sequentially exogenous unobserved confounders, 2023\.URL[https://arxiv\.org/abs/2302\.00662](https://arxiv.org/abs/2302.00662)\.
- Bruns\-Smith \[2021\]David A Bruns\-Smith\.Model\-free and model\-based policy evaluation when causality is uncertain\.In Marina Meila and Tong Zhang, editors,*Proceedings of the 38th International Conference on Machine Learning*, volume 139 of*Proceedings of Machine Learning Research*, pages 1116–1126\. PMLR, 18–24 Jul 2021\.URL[https://proceedings\.mlr\.press/v139/bruns\-smith21a\.html](https://proceedings.mlr.press/v139/bruns-smith21a.html)\.
- Kausik et al\. \[2024\]Chinmaya Kausik, Yangyi Lu, Kevin Tan, Maggie Makar, Yixin Wang, and Ambuj Tewari\.Offline policy evaluation and optimization under confounding\.In Sanjoy Dasgupta, Stephan Mandt, and Yingzhen Li, editors,*Proceedings of The 27th International Conference on Artificial Intelligence and Statistics*, volume 238 of*Proceedings of Machine Learning Research*, pages 1459–1467\. PMLR, 02–04 May 2024\.URL[https://proceedings\.mlr\.press/v238/kausik24a\.html](https://proceedings.mlr.press/v238/kausik24a.html)\.
- Killian et al\. \[2022\]Taylor W Killian, Marzyeh Ghassemi, and Shalmali Joshi\.Counterfactually guided policy transfer in clinical settings\.In*Conference on Health, Inference, and Learning*, pages 5–31\. PMLR, 2022\.
- Lally et al\. \[2026\]Jessica Lally, Milad Kazemi, and Nicola Paoletti\.Robust counterfactual inference in markov decision processes\.AAMAS ’26, page 1527–1535, Richland, SC, 2026\. International Foundation for Autonomous Agents and Multiagent Systems\.ISBN 9798400723179\.[10\.65109/TXUQ4572](https://arxiv.org/doi.org/10.65109/TXUQ4572)\.URL[https://doi\.org/10\.65109/TXUQ4572](https://doi.org/10.65109/TXUQ4572)\.
- McClean et al\. \[2025\]Alec McClean, Zach Branson, and Edward H\. Kennedy\.Calibrated sensitivity models, 2025\.URL[https://arxiv\.org/abs/2405\.08738](https://arxiv.org/abs/2405.08738)\.
- Namkoong et al\. \[2020\]Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, and Emma Brunskill\.Off\-policy policy evaluation for sequential decisions under unobserved confounding\.In H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin, editors,*Advances in Neural Information Processing Systems*, volume 33, pages 18819–18831\. Curran Associates, Inc\., 2020\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2020/file/da21bae82c02d1e2b8168d57cd3fbab7\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2020/file/da21bae82c02d1e2b8168d57cd3fbab7-Paper.pdf)\.
- Oberst and Sontag \[2019\]Michael Oberst and David Sontag\.Counterfactual off\-policy evaluation with gumbel\-max structural causal models\.In*International Conference on Machine Learning*, pages 4881–4890\. PMLR, 2019\.
- Pearl \[2009\]Judea Pearl\.*Causality*\.Cambridge University Press, 2nd\{\}^\{\\textnormal\{nd\}\}edition, 2009\.[10\.1017/CBO9780511803161](https://arxiv.org/doi.org/10.1017/CBO9780511803161)\.
- Tsirtsis and Rodriguez \[2024\]Stratis Tsirtsis and Manuel Rodriguez\.Finding counterfactually optimal action sequences in continuous state spaces\.*Advances in Neural Information Processing Systems*, 36, 2024\.
- Tsirtsis et al\. \[2021\]Stratis Tsirtsis, Abir De, and Manuel Rodriguez\.Counterfactual explanations in sequential decision making under uncertainty\.*Advances in Neural Information Processing Systems*, 34:30127–30139, 2021\.Similar Articles
Safe Bayesian Optimization with Counterfactual Policies
This paper introduces a method for safe Bayesian optimization when safety is defined relative to a counterfactual baseline policy. It uses conformal prediction to estimate counterfactual outcomes and provides safety guarantees with user-specified violation rates.
Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness
This paper presents a partial-identification approach to use prior RCT data to bound the causal effect of a new ML model, leveraging assumptions about counterfactual correctness and subgroup predictive accuracy to yield more informative bounds.
Property-driven Causal Abstractions for Markov Decision Processes
This paper introduces a property-driven causal abstraction technique for factored Markov Decision Processes (MDPs), grouping states based on causal relations over state variable predicates to reduce model size while preserving property-relevant behavior. The approach is evaluated on standard benchmarks, yielding small abstractions that support near-optimal policy computation and often generalize to larger MDPs.
Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes
A HEC Montréal/MILA paper proposes fair policy optimization for major-minor weakly coupled MDPs, replacing the utilitarian objective with monotone concave fairness functions and introducing a count-proportion-based deep RL approach with a priority-based sampler, validated on machine replacement and NYC taxi pricing/relocation tasks.
Utility-Constrained Policy Optimization
This paper introduces a simple yet powerful methodology for Utility-Constrained MDPs (UCMDPs) that enables risk-sensitive constraints without fixing constraint limits in advance, outperforming baselines on Safety Gymnasium benchmarks.