Multichain鲁棒平均奖励Markov决策过程的向量Bellman理论
摘要
本文介绍了用于鲁棒平均奖励Markov决策过程的向量Bellman理论,通过有限模型可解性和近似移位的Halpern规划算法,实现了在转移不确定性下的最优性能。
arXiv:2609.28792v1 Announce Type: new
Abstract: Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history-dependent opponents, simultaneously from all initial states. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent-class gains. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average-optimal after a finite, instance-dependent budget. These results thus connect finite Bellman certificates to undiscounted planning for state-dependent robust average rewards, providing theoretical understandings.
查看缓存全文
缓存时间: 2026/09/25 09:37
# Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes
Source: [https://arxiv.org/html/2609.28792](https://arxiv.org/html/2609.28792)
George AtiaAffiliation:University of Central Florida
###### Abstract
Robust average\-reward Markov decision processes provide a fundamental framework for long\-term performance optimization under uncertainty, and can have optimal long\-run rewards that depend on the initial state\. This state dependence requires a vector Bellman theory that accounts for both recurrent\-class rewards and transition uncertainty\. We develop such a theory for finite models with compact, post\-action\(s,a\)\(s,a\)\-rectangular ambiguity\. A gain\-first, bias\-second optimization principle yields a coupled vector gain\-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history\-dependent opponents, simultaneously from all initial states\. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent\-class gains\. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm\. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average\-optimal after a finite, instance\-dependent budget\. These results thus connect finite Bellman certificates to undiscounted planning for state\-dependent robust average rewards, providing theoretical understandings\.
## 1Introduction
Markov decision processes \(MDPs\)[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32)provide a standard framework for sequential decision\-making under stochastic agent\-environment interactions, in which an agent seeks a policy that maximizes expected reward under a specified performance criterion\. A policy optimized for a single nominal MDP with a fixed transition model, however, can perform poorly when the deployed dynamics differ from that model or are imperfectly known\. Such model mismatches widely exist in practice, due to, e\.g\., inaccurate model estimation, non\-stationary environment, or unexpected perturbations upon deployment\. Robust MDPs are thus proposed to address this mismatch by optimizing worst\-case performance over a prescribed family of transition models[Iyengar \(2005\)](https://arxiv.org/html/2609.28792#bib.bib16);[Nilim and El Ghaoui \(2004\)](https://arxiv.org/html/2609.28792#bib.bib28), which captures the uncertain models, and is named ambiguity set or uncertainty set\. To find such a robust policy, one need to set the performance criterion\. Among different criteria, average reward is particularly important as it measures sustained performance without imposing a discount factor, making it natural for systems operating over long horizons\. Robust average\-reward theory must therefore reconcile uncertainty in the transition dynamics with the chain structure that governs long\-run performance\.
In average\-reward models, the limiting average\-reward function \(i\.e\., the long\-run gain\) is naturally dependent on the initial state\. Even under a fixed stationary policy and fixed transition kernel, different starting states may reach recurrent classes with different class gains and different absorption probabilities: even two absorbing states with unequal rewards can yield a nonconstant gain vector\. Under transition uncertainty, stationary gains can also change abruptly and becomes more challenging\. If a zero\-reward state reaches an absorbing unit\-reward state with probabilityε\\varepsilonper step, its gain is11for everyε\>0\\varepsilon\>0but00atε=0\\varepsilon=0\. Vanishing transition probabilities can therefore change the recurrent structure\. This raises two separate questions: whether nature attains the worst stationary gain and whether transient rewards admit a finite common bias\. These features thus require a state\-dependent gain vector and an all\-state gain\-bias formulation, to tackle the complicated statewise dependence and provide a concrete Bellman\-typed characterization\.
Existing work supplies two complementary foundations for such a theory\. The early line of work[Wang et al\. \(2023e\)](https://arxiv.org/html/2609.28792#bib.bib49);[Wang et al\. \(2024d\)](https://arxiv.org/html/2609.28792#bib.bib52)develop the robust Bellman equation and algorithms for\(s,a\)\(s,a\)\-rectangular models under a uniform unichain condition, which ensures all the average reward \(or the gain\) is independent from initial state\. More recently,[Wang and Si \(2025\)](https://arxiv.org/html/2609.28792#bib.bib66)develop robust Bellman optimality with constant \(optimal\) gain, which covers more general conditions like one\-sided weak communication\. These conditions may permit individual stationary controller\-nature pairs to be multichain, while the Bellman result remains in the state\-independent\-gain regime\. Strategic value theory developed in[Grand\-Clement et al\. \(2023\)](https://arxiv.org/html/2609.28792#bib.bib11)establishes stationary controller optimality for general compact\(s,a\)\(s,a\)\-rectangular ambiguity, while polytopic models admit planning through finite stochastic games[Chatterjee et al\. \(2024\)](https://arxiv.org/html/2609.28792#bib.bib5)\. But no Bellman\-typed characterization is studied\. These studies thus leave an open question:when the gain depends on the initial states, does this vector gain admit a finite gain\-bias Bellman certificate, and how can such a certificate support planning?In this work, we provide concrete answers to the question\. Our main contributions are summarized as follows\.
A gain\-first, bias\-second vector Bellman certificate\.We first formulate a coupled Bellman system that gives continuation gain priority for both players \(nature and controller\) to characterize the robust average reward\. Nature first minimizes continuation gain, and the controller maximizes the resulting minimum\. Reward\-bias optimization then takes place among the choices tied at this first level\. We further show that every finite solution identifies the robust optimal gain and produces an all\-state stationary saddle against history\-dependent opponents\. This verification controls nature’s deviations from gain\-minimizing rows, including rows with arbitrarily small gain gaps\. It also connects the certificate to an asymptotically affine trajectory of the original Bellman operator \(Section[4](https://arxiv.org/html/2609.28792#S4)\)\.
Exact solvability beyond communication assumptions\.We then characterize when the finite certificate exists\. For optimal control, the criterion combines a stationary saddle in the gain\-restricted game with a uniform lower bound on canonical biases over all zero\-gain replies to one securing controller\. This separates stationary gain attainment from the control of transient rewards needed for a finite bias\. We further explore concrete regimes that permit distinct recurrent\-class gains: polytopic row sets, a uniform support gap, and different concrete distributional ambiguity sets, showing they are sufficient conditions for Bellman solvability \(Section[5](https://arxiv.org/html/2609.28792#S5)\)\.
Direct undiscounted planning with a vanishing affine defect\.We adapt approximately shifted Halpern iteration\([Zurek and Chen, 2025](https://arxiv.org/html/2609.28792#bib.bib65), Algorithm 1\)to the robust Bellman operator\. For compact row sets, a finite Bellman certificate can produce an asymptotically affine trajectory whose error remains nonzero at every finite time\. We show that this vanishing defect suffices for convergence of the gain estimates and Bellman displacements\. Directional control of the iterates then makes every extracted greedy controller eventually average\-optimal from all states\. The analysis gives finite\-budget bounds in terms of the affine defect and compact examples with arbitrarily small algebraic convergence exponents for the gain estimator \(Section[6](https://arxiv.org/html/2609.28792#S6)\)\.
Roadmap\.Section[3](https://arxiv.org/html/2609.28792#S3)establishes the robust gain vector and an all\-state optimal stationary controller under the standing assumptions\. Section[4](https://arxiv.org/html/2609.28792#S4)develop the vector Bellman equation system and optimality guarantees of its solutions\. Section[5](https://arxiv.org/html/2609.28792#S5)characterizes when that solution exists and identifies sufficient regimes\. Section[6](https://arxiv.org/html/2609.28792#S6)uses the certificate to analyze direct undiscounted planning\. Thus the gain and optimal controller precede the Bellman system logically; finite solvability is the condition for the certificate and the planning theorem\.
## 2Problem Formulation and Preliminaries
We consider a finite robust MDP\(S,A,𝒰,r\)\(S,A,\\mathcal\{U\},r\)\. The state space isS=\{1,…,n\}S=\\\{1,\\ldots,n\\\}, each action setA\(i\)A\(i\)is finite and nonempty, andria∈ℝr\_\{ia\}\\in\\mathbb\{R\}is the one\-stage reward\. At timet≥0t\\geq 0, the controller observesStS\_\{t\}and choosesAt∈A\(St\)A\_\{t\}\\in A\(S\_\{t\}\)\. Nature observes the chosen action, selects a transition rowpt∈𝒰StAt⊆Δ\(S\)p\_\{t\}\\in\\mathcal\{U\}\_\{S\_\{t\}A\_\{t\}\}\\subseteq\\Delta\(S\), and the next state is sampled fromptp\_\{t\}\. We assume\(s,a\)\(s,a\)\-rectangular ambiguity[Iyengar \(2005\)](https://arxiv.org/html/2609.28792#bib.bib16);[Nilim and El Ghaoui \(2004\)](https://arxiv.org/html/2609.28792#bib.bib28):𝒰=∏i∈S∏a∈A\(i\)𝒰ia\\mathcal\{U\}=\\prod\_\{i\\in S\}\\prod\_\{a\\in A\(i\)\}\\mathcal\{U\}\_\{ia\}, where every row set𝒰ia\\mathcal\{U\}\_\{ia\}is nonempty and compact\. In the following, vector inequalities are componentwise\.
The planning objective is to compute a controller policy that maximizes worst\-case long\-run reward\. WriteΠD⊆ΠS⊆ΠH\\Pi\_\{D\}\\subseteq\\Pi\_\{S\}\\subseteq\\Pi\_\{H\}for deterministic stationary, randomized stationary, and randomized history\-dependent controller strategies\. A stationary policy belongs to∏iΔ\(A\(i\)\)\\prod\_\{i\}\\Delta\(A\(i\)\), while a history\-dependent policy maps the observed history\(S0,A0,…,St\)\(S\_\{0\},A\_\{0\},\\ldots,S\_\{t\}\)to a distribution onA\(St\)A\(S\_\{t\}\)\. Nature’s randomized history\-dependent strategies form𝒬H\\mathcal\{Q\}\_\{H\}, and its full stationary selectors form𝒬S=∏i∏a∈A\(i\)𝒰ia\\mathcal\{Q\}\_\{S\}=\\prod\_\{i\}\\prod\_\{a\\in A\(i\)\}\\mathcal\{U\}\_\{ia\}, is the transition kernels nature selected\. A full selector specifies a row for every state\-action pair, including actions unused by a particular controller\. Both players observe the initial state\.
Payoffs\.Our primary performance criterion is the lower limiting expected average reward\. We also record the corresponding upper limit and the two criteria that take the sample\-path limit before expectation\. These distinctions specify the strength of the strategy guarantees proved below\. For\(σ,τ\)∈ΠH×𝒬H\(\\sigma,\\tau\)\\in\\Pi\_\{H\}\\times\\mathcal\{Q\}\_\{H\}, let𝔼iσ,τ\\mathbb\{E\}\_\{i\}^\{\\sigma,\\tau\}denote expectation under the induced process starting atS0=iS\_\{0\}=i\. For the trajectory\(S0,A0,S1,A1,…\)\(S\_\{0\},A\_\{0\},S\_\{1\},A\_\{1\},\\ldots\), defineXN=N−1∑t=0N−1rStAtX\_\{N\}=N^\{\-1\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}andJN\(i,σ,τ\)=𝔼iσ,τXNJ\_\{N\}\(i;\\sigma,\\tau\)=\\mathbb\{E\}\_\{i\}^\{\\sigma,\\tau\}X\_\{N\}\. We distinguish four average\-payoff conventions[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32):
Ji−\\displaystyle J\_\{i\}^\{\-\}=lim infN𝔼iXN,\\displaystyle=\\liminf\_\{N\}\\mathbb\{E\}\_\{i\}X\_\{N\},Ji\+\\displaystyle J\_\{i\}^\{\+\}=lim supN𝔼iXN,\\displaystyle=\\limsup\_\{N\}\\mathbb\{E\}\_\{i\}X\_\{N\},\(1\)Ii−\\displaystyle I\_\{i\}^\{\-\}=𝔼ilim infNXN,\\displaystyle=\\mathbb\{E\}\_\{i\}\\liminf\_\{N\}X\_\{N\},Ii\+\\displaystyle I\_\{i\}^\{\+\}=𝔼ilim supNXN,\\displaystyle=\\mathbb\{E\}\_\{i\}\\limsup\_\{N\}X\_\{N\},with the strategy pair suppressed\. These criteria can differ for a fixed history\-dependent pair\. For a payoffΨ\\Psiand strategy classes𝒞⊆ΠH\\mathcal\{C\}\\subseteq\\Pi\_\{H\},𝒩⊆𝒬H\\mathcal\{N\}\\subseteq\\mathcal\{Q\}\_\{H\}, define the lower and upper values
v¯iΨ\(𝒞,𝒩\):=supσ∈𝒞infτ∈𝒩Ψi\(σ,τ\),v¯iΨ\(𝒞,𝒩\):=infτ∈𝒩supσ∈𝒞Ψi\(σ,τ\)\.\\underline\{v\}\_\{i\}^\{\\Psi\}\(\\mathcal\{C\},\\mathcal\{N\}\):=\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\Psi\_\{i\}\(\\sigma,\\tau\),\\qquad\\overline\{v\}\_\{i\}^\{\\Psi\}\(\\mathcal\{C\},\\mathcal\{N\}\):=\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\Psi\_\{i\}\(\\sigma,\\tau\)\.\(2\)For a fixed controllerσ\\sigma, its robust performance isinfτ∈𝒬HΨi\(σ,τ\)\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\Psi\_\{i\}\(\\sigma,\\tau\)\. The lower value maximizes this guarantee, while the upper value minimizes the controller’s best response\. A pair\(σ¯,τ¯\)\(\\bar\{\\sigma\},\\bar\{\\tau\}\)is an all\-state saddle forΨ\\Psiatu∈ℝSu\\in\\mathbb\{R\}^\{S\}ifΨi\(σ¯,τ\)≥ui≥Ψi\(σ,τ¯\)\\Psi\_\{i\}\(\\bar\{\\sigma\},\\tau\)\\geq u\_\{i\}\\geq\\Psi\_\{i\}\(\\sigma,\\bar\{\\tau\}\)for everyii,σ∈ΠH\\sigma\\in\\Pi\_\{H\}, andτ∈𝒬H\\tau\\in\\mathcal\{Q\}\_\{H\}\. Theorem[8](https://arxiv.org/html/2609.28792#Thmtheorem8)identifies a common value for the four conventions, and Theorem[10](https://arxiv.org/html/2609.28792#Thmtheorem10)characterizes exact stationary attainment by nature\.
Dynamic programming operators\.The order of play within each stage explains the robust Bellman maps\. Given continuation valuexxat stateii, the controller first choosesaaand nature, having observed that action, choosesp∈𝒰iap\\in\\mathcal\{U\}\_\{ia\}\. The resulting one\-step saddle operators are[Iyengar \(2005\)](https://arxiv.org/html/2609.28792#bib.bib16)
\(Tx\)i=maxa∈A\(i\)\{ria\+minp∈𝒰iap⊤x\},\(Tπx\)i=riπ\+∑aπ\(a∣i\)minp∈𝒰iap⊤x,\(Tx\)\_\{i\}=\\max\_\{a\\in A\(i\)\}\\left\\\{r\_\{ia\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}x\\right\\\},\\qquad\(T^\{\\pi\}x\)\_\{i\}=r\_\{i\}^\{\\pi\}\+\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}x,whereriπ=∑aπ\(a∣i\)riar\_\{i\}^\{\\pi\}=\\sum\_\{a\}\\pi\(a\\mid i\)r\_\{ia\}forπ∈ΠS\\pi\\in\\Pi\_\{S\}\. Nature observes the sampled action and then selects the row, which explains the separate minima inTπT^\{\\pi\}\. For discount factor1−ϵ1\-\{\\epsilon\},0<ϵ<10<\{\\epsilon\}<1, define the value vectors coordinatewise by\(Vϵπ\)i=infτ∈𝒬H𝔼iπ,τ∑t=0∞\(1−ϵ\)trStAt\(V\_\{\\epsilon\}^\{\\pi\}\)\_\{i\}=\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\mathbb\{E\}\_\{i\}^\{\\pi,\\tau\}\\sum\_\{t=0\}^\{\\infty\}\(1\-\{\\epsilon\}\)^\{t\}r\_\{S\_\{t\}A\_\{t\}\}, and\(Vϵ\)i=supσ∈ΠHinfτ∈𝒬H𝔼iσ,τ∑t=0∞\(1−ϵ\)trStAt\.\(V\_\{\\epsilon\}\)\_\{i\}=\\sup\_\{\\sigma\\in\\Pi\_\{H\}\}\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\mathbb\{E\}\_\{i\}^\{\\sigma,\\tau\}\\sum\_\{t=0\}^\{\\infty\}\(1\-\{\\epsilon\}\)^\{t\}r\_\{S\_\{t\}A\_\{t\}\}\.Rectangular discounted dynamic programming[Iyengar \(2005\)](https://arxiv.org/html/2609.28792#bib.bib16);[Nilim and El Ghaoui \(2004\)](https://arxiv.org/html/2609.28792#bib.bib28)givesVϵ=T\(\(1−ϵ\)Vϵ\)V\_\{\\epsilon\}=T\(\(1\-\{\\epsilon\}\)V\_\{\\epsilon\}\)andVϵπ=Tπ\(\(1−ϵ\)Vϵπ\)\.V\_\{\\epsilon\}^\{\\pi\}=T^\{\\pi\}\(\(1\-\{\\epsilon\}\)V\_\{\\epsilon\}^\{\\pi\}\)\.Each discounted fixed point is unique and is attained simultaneously from all states by deterministic stationary discounted\-optimal selectors\. Consequently,\(Vϵ\)i=maxπ∈ΠD\(Vϵπ\)i\(V\_\{\\epsilon\}\)\_\{i\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\(V\_\{\\epsilon\}^\{\\pi\}\)\_\{i\}\. Appendix[C](https://arxiv.org/html/2609.28792#A3)reviews the discounted dynamic\-programming and finite\-chain facts used in the analysis\.
Multichain structure and vector gain\.A stationary pair\(π,q\)\(\\pi,q\)induces the transition matrixPijπ,q=∑aπ\(a\|i\)qia,jP^\{\\pi,q\}\_\{ij\}=\\sum\_\{a\}\\pi\(a\|i\)q\_\{ia,j\}\. Its gain vector isηπ,q:=\(Pπ,q\)∞rπ\\eta^\{\\pi,q\}:=\(P^\{\\pi,q\}\)^\{\\infty\}r^\{\\pi\}, where the Cesàro projectorP∞:=limN→∞1N∑t=0N−1PtP^\{\\infty\}:=\\lim\_\{N\\to\\infty\}\\frac\{1\}\{N\}\\sum\_\{t=0\}^\{N\-1\}P^\{t\}exists for every finite stochastic matrix, including periodic multichain matrices, and each coordinate ofηπ,q\\eta^\{\\pi,q\}is an absorption\-weighted average of recurrent\-class rewards[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32)\. This representation explains how initial states can have different gains\. Unichain and irreducible assumptions in prior analyses make each stationary chain’s gain constant[Wang et al\. \(2023e\)](https://arxiv.org/html/2609.28792#bib.bib49);[Wang et al\. \(2024d\)](https://arxiv.org/html/2609.28792#bib.bib52);[Xu et al\. \(2025a\)](https://arxiv.org/html/2609.28792#bib.bib59);[Roch et al\. \(2025\)](https://arxiv.org/html/2609.28792#bib.bib34)\. Here both evaluation and control concern the complete statewise gain vector\. See Appendix[C](https://arxiv.org/html/2609.28792#A3)for a review of existing results\.
## 3Statewise values and stationary strategies
We first identify the gain vector, the main objective of our studies\. The results connect discounted and finite\-horizon values to robust average reward and supply strategies that secure these values simultaneously from every initial state\. We begin with a fixed stationary controller\.
###### Theorem 1\.
\[Extension of\([Grand\-Clement et al\., 2023](https://arxiv.org/html/2609.28792#bib.bib11), Lemmas 3\.3 and 4\.7\)\] For everyπ∈ΠS\\pi\\in\\Pi\_\{S\}, there isgπ∈ℝSg^\{\\pi\}\\in\\mathbb\{R\}^\{S\}such that, with both limits in supremum norm,
gπ=limϵ↓0ϵVϵπ=limN→∞\(Tπ\)N0N,andgiπ=infq∈𝒬Sηiπ,q,∀i∈S\.g^\{\\pi\}=\\lim\_\{\{\\epsilon\}\\downarrow 0\}\{\\epsilon\}V\_\{\\epsilon\}^\{\\pi\}=\\lim\_\{N\\to\\infty\}\\frac\{\(T^\{\\pi\}\)^\{N\}0\}\{N\},\\text\{ and \}g\_\{i\}^\{\\pi\}=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\eta\_\{i\}^\{\\pi,q\},\\forall i\\in S\.\(3\)For everyΨ∈\{I−,J−,J\+,I\+\}\\Psi\\in\\\{I^\{\-\},J^\{\-\},J^\{\+\},I^\{\+\}\\\},infτ∈𝒬HΨi\(π,τ\)=infq∈𝒬SΨi\(π,q\)=giπ\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\Psi\_\{i\}\(\\pi,\\tau\)=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\Psi\_\{i\}\(\\pi,q\)=g\_\{i\}^\{\\pi\}\. Moreover, for eachν\>0\\nu\>0, there exists someqπ,ν∈𝒬Sq^\{\\pi,\\nu\}\\in\\mathcal\{Q\}\_\{S\}such thatgπ≤ηπ,qπ,ν≤gπ\+ν𝟏g^\{\\pi\}\\leq\\eta^\{\\pi,q^\{\\pi,\\nu\}\}\\leq g^\{\\pi\}\+\\nu\\mathbf\{1\}\.
The vectorgπg^\{\\pi\}is the robust gain ofπ\\pi: it is the common normalized discounted and finite\-horizon limit and the worst\-case value under each payoff convention in equation[1](https://arxiv.org/html/2609.28792#S2.E1)\. A stationary selector attaining this vector need not exist, even though the row sets are compact\. For every positive tolerance, however, one stationary selector approximates the entire vector simultaneously\. This is the all\-state guarantee used in policy evaluation\.
The next result identifies the optimal robust gain and strategies securing it from all initial states\.
###### Theorem 2\.
\[Extension of\([Grand\-Clement et al\., 2023](https://arxiv.org/html/2609.28792#bib.bib11), Theorem 5\.2\)\] There existg⋆∈ℝSg^\{\\star\}\\in\\mathbb\{R\}^\{S\}and a deterministic stationaryπ⋆∈ΠD\\pi^\{\\star\}\\in\\Pi\_\{D\}such that
g⋆=limϵ↓0ϵVϵ=limN→∞TN0N=maxπ∈ΠDgπ=gπ⋆,g^\{\\star\}=\\lim\_\{\{\\epsilon\}\\downarrow 0\}\{\\epsilon\}V\_\{\\epsilon\}=\\lim\_\{N\\to\\infty\}\\frac\{T^\{N\}0\}\{N\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}g^\{\\pi\}=g^\{\\pi^\{\\star\}\},\(4\)where the limits are in supremum norm and the maximum is coordinatewise\. Moreover, for everyδ\>0\\delta\>0, there existqδ∈𝒬Sq\_\{\\delta\}\\in\\mathcal\{Q\}\_\{S\}andNδ<∞N\_\{\\delta\}<\\inftysuch that, for alli∈Si\\in SandN≥NδN\\geq N\_\{\\delta\},
JN\(i,π⋆,τ\)≥gi⋆−δfor everyτ∈𝒬H,andJN\(i,σ,qδ\)≤gi⋆\+δfor everyσ∈ΠH\.J\_\{N\}\(i;\\pi^\{\\star\},\\tau\)\\geq g\_\{i\}^\{\\star\}\-\\delta\\text\{ for every \}\\tau\\in\\mathcal\{Q\}\_\{H\},\\text\{ and \}J\_\{N\}\(i;\\sigma,q\_\{\\delta\}\)\\leq g\_\{i\}^\{\\star\}\+\\delta\\text\{ for every \}\\sigma\\in\\Pi\_\{H\}\.\(5\)
Equation[4](https://arxiv.org/html/2609.28792#S3.E4)identifies the optimal robust gain through discounted values, finite\-horizon values, and policy optimization\. One deterministic stationary controllerπ⋆\\pi^\{\\star\}attains every coordinate\. For each toleranceδ\\delta, equation[5](https://arxiv.org/html/2609.28792#S3.E5)supplies a single stationary nature selector and one horizon threshold that work for every initial state, every later horizon, and every history\-dependent opponent\. The controller remains the same for all tolerances; nature’s approximating selector can depend onδ\\delta\.
The vectorg⋆g^\{\\star\}is also the common lower and upper value under all four payoff conventions\. The same controllerπ⋆\\pi^\{\\star\}is optimal in each case\. The values are preserved for intermediate strategy classes containingΠD\\Pi\_\{D\}and𝒬S\\mathcal\{Q\}\_\{S\}, respectively, including the stationary strategy classes\. See Appendix[F](https://arxiv.org/html/2609.28792#A6)\.
For a full stationary selectorqq, letdi\(q\):=maxπ∈ΠDηiπ,qd\_\{i\}\(q\):=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta\_\{i\}^\{\\pi,q\}be the optimal gain of the nominal MDP withqqfixed\. Nature is exactly optimal from all states against every history\-dependent controller if and only ifd\(q\)=g⋆d\(q\)=g^\{\\star\}\. This criterion checks the controller’s best response over every action, including actions unused byπ⋆\\pi^\{\\star\}\. The equalityηπ⋆,q=g⋆\\eta^\{\\pi^\{\\star\},q\}=g^\{\\star\}checks only the prescribed controller\. Appendix[H](https://arxiv.org/html/2609.28792#A8)develops this characterization directly from stationary gains\.
## 4Vector Bellman equation system
Theorems[1](https://arxiv.org/html/2609.28792#Thmtheorem1)\-[2](https://arxiv.org/html/2609.28792#Thmtheorem2)identify the robust value and an all\-state optimal controller without assuming a finite bias\. In this section, we investigate a stronger certificate: a gain\-bias certificate: it must reconcile long\-run gains and transient rewards for both players using one finite pair of vectors\. We now formulate a local system and establish what any finite solution certifies\.
### 4\.1Robust Bellman optimality system: Gain first, bias second
In the constant\-gain setting[Wang et al\. \(2023e\)](https://arxiv.org/html/2609.28792#bib.bib49);[Wang and Si \(2025\)](https://arxiv.org/html/2609.28792#bib.bib66), the robust Bellman equation isρ𝟏\+h=Th\\rho\\mathbf\{1\}\+h=Th\. Every probability row preserves the same continuation gain becausep⊤\(ρ𝟏\)=ρp^\{\\top\}\(\\rho\\mathbf\{1\}\)=\\rho\. For a vector gain, different rows can lead to different long\-run reward rates\. The equationg\+h=Thg\+h=Thalone does not require the selected transitions to preserve the proposed gain\. The Bellman system must therefore compare continuation gains before comparing finite reward and bias terms\.
To see the required order, consider a continuation vectortg\+htg\+hfor largett\. The one\-step objective istp⊤g\+ria\+p⊤htp^\{\\top\}g\+r\_\{ia\}\+p^\{\\top\}h\. A fixed gain difference dominates the bounded reward\-bias term asttgrows\. Nature therefore first minimizesp⊤gp^\{\\top\}g, and the controller maximizes this minimum\. Among the choices tied in gain, both players optimize reward and continuation bias\. Our system below imposes this ordering on both players, extending the nominal multichain gain\-bias separation[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32);[Zurek and Chen \(2025\)](https://arxiv.org/html/2609.28792#bib.bib65)\.
Gain first, bias second\.Forg∈ℝng\\in\\mathbb\{R\}^\{n\}, define the worst continuation gain of each action and the controller’s optimal continuation gain by
mia\(g\):=minp∈𝒰iap⊤g,\(T^g\)i:=maxa∈A\(i\)mia\(g\)\.m\_\{ia\}\(g\):=\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g,\\qquad\(\\widehat\{T\}g\)\_\{i\}:=\\max\_\{a\\in A\(i\)\}m\_\{ia\}\(g\)\.\(6\)For anyg∈ℝng\\in\\mathbb\{R\}^\{n\}, define
Ag\(i\):=\{a∈A\(i\):mia\(g\)=gi\},Fia\(g\):=argminp∈𝒰iap⊤g,A\_\{g\}\(i\):=\\\{a\\in A\(i\):m\_\{ia\}\(g\)=g\_\{i\}\\\},\\qquad F\_\{ia\}\(g\):=\\argmin\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g,\(7\)and we callFia\(g\)F\_\{ia\}\(g\)a gain face even when the row set is nonconvex\. Compactness makes everyFia\(g\)F\_\{ia\}\(g\)nonempty\. Wheng=T^gg=\\widehat\{T\}g, eachAg\(i\)A\_\{g\}\(i\)is also nonempty, and we define the optimal\-control bias operator on these active actions\. Moreover, define the bias operator
\(Lgh\)i:=maxa∈Ag\(i\)minp∈Fia\(g\)\(ria\+p⊤h\),Kgh:=Lgh−g\.\(L\_\{g\}h\)\_\{i\}:=\\max\_\{a\\in A\_\{g\}\(i\)\}\\min\_\{p\\in F\_\{ia\}\(g\)\}\(r\_\{ia\}\+p^\{\\top\}h\),\\qquad K\_\{g\}h:=L\_\{g\}h\-g\.\(8\)Our robust vector Bellman optimality system is
g=T^g,g\+h=Lgh\.g=\\widehat\{T\}g,\\qquad g\+h=L\_\{g\}h\.\(9\)The two equations perform distinct tasks\. The first enforces consistency of continuation gain, while the second determines the reward\-bias balance among gain\-optimal choices\. They must be solved together: the first equation is reward\-free and accepts every constant vector\. The restrictions toAg\(i\)A\_\{g\}\(i\)andFia\(g\)F\_\{ia\}\(g\)preserve gain priority for the controller and nature, respectively\. Specifically, Examples[1](https://arxiv.org/html/2609.28792#Thmexample1)\-[2](https://arxiv.org/html/2609.28792#Thmexample2)show that omitting either restriction can certify an incorrect gain\.
For singleton row sets, the system reduces to the classical multichain MDP optimality equations[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32)\. Ifg=ρ𝟏g=\\rho\\mathbf\{1\}, all actions and rows are gain\-active, soLg=TL\_\{g\}=Tand equation[9](https://arxiv.org/html/2609.28792#S4.E9)becomes the constant\-gain robust Bellman equation[Wang et al\. \(2023e\)](https://arxiv.org/html/2609.28792#bib.bib49);[Wang and Si \(2025\)](https://arxiv.org/html/2609.28792#bib.bib66)\.
The following theorem verifies the system against history\-dependent opponents\.
###### Theorem 3\(Bellman optimality\)\.
Every finite solution\(g,h\)\(g,h\)of equation[9](https://arxiv.org/html/2609.28792#S4.E9)satisfiesg=g⋆g=g^\{\\star\}\. Choose a deterministic stationary policyπ\(i\)∈argmaxa∈Ag\(i\)minp∈Fia\(g\)\{ria\+p⊤h\}\\pi\(i\)\\in\\mathop\{\\rm argmax\}\_\{a\\in A\_\{g\}\(i\)\}\\min\_\{p\\in F\_\{ia\}\(g\)\}\\\{r\_\{ia\}\+p^\{\\top\}h\\\}, and for every active state\-action pair chooseqia∈argminp∈Fia\(g\)p⊤hq\_\{ia\}\\in\\argmin\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h; at inactive actions choose anyqia∈Fia\(g\)q\_\{ia\}\\in F\_\{ia\}\(g\)\. Then\(π,q\)\(\\pi,q\)forms an all\-state saddle for every payoff in equation[1](https://arxiv.org/html/2609.28792#S2.E1):
lim infNJN\(i,π,τ\)≥gi∀τ∈𝒬H,lim supNJN\(i,σ,q\)≤gi∀σ∈ΠH,\\liminf\_\{N\}J\_\{N\}\(i;\\pi,\\tau\)\\geq g\_\{i\}\\quad\\forall\\tau\\in\\mathcal\{Q\}\_\{H\},\\qquad\\limsup\_\{N\}J\_\{N\}\(i;\\sigma,q\)\\leq g\_\{i\}\\quad\\forall\\sigma\\in\\Pi\_\{H\},\(10\)and this pair achieves the optimal robust average reward:limN→∞JN\(i,π,q\)=gi⋆\\lim\_\{N\\to\\infty\}J\_\{N\}\(i;\\pi,q\)=g\_\{i\}^\{\\star\}for everyi∈Si\\in S\.
Fix the selected controller and writed\(p\)=p⊤g−gi≥0d\(p\)=p^\{\\top\}g\-g\_\{i\}\\geq 0ande\(p\)=ria\+p⊤h−gi−hie\(p\)=r\_\{ia\}\+p^\{\\top\}h\-g\_\{i\}\-h\_\{i\}\. On a gain\-minimizing row, the bias equation givese\(p\)≥0e\(p\)\\geq 0\. Compactness implies that for eachε\>0\\varepsilon\>0there is a finiteCεC\_\{\\varepsilon\}withe\(p\)≥−ε−Cεd\(p\)e\(p\)\\geq\-\\varepsilon\-C\_\{\\varepsilon\}d\(p\)on every feasible row\. Along any history\-dependent nature strategy, the cumulative expected gain increase is bounded bysp\(g\)\\operatorname\{sp\}\(g\); telescoping the bias gives the controller’s lower guarantee\. For the selected fullnature plan, active controller actions satisfy the reverse bias inequality, while inactive actions have fixed negative continuation\-gain gaps\. Since there are finitely many controller actions, their bias discrepancies can be charged to those gaps\. The same telescoping argument gives nature’s upper guarantee\.
Asymptotically affine Bellman trajectory\.The certificate also describes the behavior of the original Bellman operatorTT, which optimizes over all actions and all feasible rows\. Our Lemma[6](https://arxiv.org/html/2609.28792#Thmlemma6)proves
g=T^g,g\+h=Lgh⟺T\(h\+tg\)=h\+\(t\+1\)g\+o\(1\),t→∞\.g=\\widehat\{T\}g,\\qquad g\+h=L\_\{g\}h\\quad\\Longleftrightarrow\\quad T\(h\+tg\)=h\+\(t\+1\)g\+o\(1\),\\qquad t\\to\\infty\.This relation connects the gain\-restricted system to the original operator used by the planner\. Asttgrows, actions with a fixed continuation\-gain disadvantage cease to compete, and minimizing rows approach their gain faces\. An optimal row forh\+tgh\+tgcan nevertheless lie outside its gain face at every finitett\. Section[6](https://arxiv.org/html/2609.28792#S6)therefore tracks a vanishing affine defect to design the planning algorithm\.
### 4\.2Fixed\-policy equation
For a fixed stationary controller, the Bellman system certifies its robust gain and a stationary nature selector that is worst from all initial states\. The action maximum is replaced by the policy average, while nature continues to minimize separately after each realized action\.
###### Propsition 1\(Fixed\-policy Bellman equations\)\.
Fix any stationary randomized policyπ∈ΠS\\pi\\in\\Pi\_\{S\}\. Suppose finite vectors\(g,h\)\(g,h\)satisfy the fixed\-policy Bellman system: for every stateii,
gi=∑aπ\(a∣i\)minp∈𝒰iap⊤g,gi\+hi=riπ\+∑aπ\(a∣i\)minp∈Fia\(g\)p⊤h,\\displaystyle g\_\{i\}=\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g,\\quad g\_\{i\}\+h\_\{i\}=r\_\{i\}^\{\\pi\}\+\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h,\(11\)whereriπ:=∑aπ\(a∣i\)riar\_\{i\}^\{\\pi\}:=\\sum\_\{a\}\\pi\(a\\mid i\)r\_\{ia\}, andFia\(g\):=argminp∈𝒰iap⊤gF\_\{ia\}\(g\):=\\arg\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g\. Theng=gπg=g^\{\\pi\}\. Moreover, for every state\-action pair withπ\(a∣i\)\>0\\pi\(a\\mid i\)\>0, chooseqia⋆∈argminp∈Fia\(g\)p⊤hq\_\{ia\}^\{\\star\}\\in\\arg\\min\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h, and choose arbitrary feasible rows for zero\-probability actions\. Then for every initial stateii,
infτ∈𝒬Hlim infN→∞JN\(i,π,τ\)=infτ∈𝒬Hlim supN→∞JN\(i,π,τ\)=giπ,\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\liminf\_\{N\\to\\infty\}J\_\{N\}\(i;\\pi,\\tau\)=\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}\\limsup\_\{N\\to\\infty\}J\_\{N\}\(i;\\pi,\\tau\)=g\_\{i\}^\{\\pi\},\(12\)and the single stationary full selectorq⋆q^\{\\star\}attains both infima simultaneously at every initial state\.
The same finite pair certifies the complete robust gain vector and one stationary selector attaining it\. We will later characterize exactly when such a pair exists\.
## 5Exact stationary conditions for solvability
Section[4](https://arxiv.org/html/2609.28792#S4)establishes that a finite Bellman solution guarantees the optimal robust gain and optimal controller\. We now determine when such a solution exists\.
A finite bias requires both attainment of the stationary gain and uniform control of the associated transient reward corrections\. We first make these requirements precise for policy evaluation\. We then characterize optimal\-control solvability through a stationary saddle in the gain\-restricted game and a one\-sided bound on its canonical biases\. For a finite stochastic matrixPP, letP∞=limN→∞N−1∑t=0N−1PtP^\{\\infty\}=\\lim\_\{N\\to\\infty\}N^\{\-1\}\\sum\_\{t=0\}^\{N\-1\}P^\{t\}andZP=\(I−P\+P∞\)−1Z\_\{P\}=\(I\-P\+P^\{\\infty\}\)^\{\-1\}\. Both exist without irreducibility or aperiodicity[Puterman \(2014\)](https://arxiv.org/html/2609.28792#bib.bib32)\. Fixπ∈ΠS\\pi\\in\\Pi\_\{S\}and define its effective row set by𝒰iπ=\{∑a∈A\(i\)π\(a∣i\)pa:pa∈𝒰iafor everya∈A\(i\)\}\\mathcal\{U\}\_\{i\}^\{\\pi\}=\\\{\\sum\_\{a\\in A\(i\)\}\\pi\(a\\mid i\)p\_\{a\}:p\_\{a\}\\in\\mathcal\{U\}\_\{ia\}\\text\{ for every \}a\\in A\(i\)\\\}\. Independent actionwise minimization gives\(Tπv\)i=riπ\+minp∈𝒰iπp⊤v\(T^\{\\pi\}v\)\_\{i\}=r\_\{i\}^\{\\pi\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{i\}^\{\\pi\}\}p^\{\\top\}v\. Thus, nature’s fixed\-policy problem is a compact\-action MDP with rewardrπr^\{\\pi\}\. For an effective selectorq∈∏i𝒰iπq\\in\\prod\_\{i\}\\mathcal\{U\}\_\{i\}^\{\\pi\}, write\(Pq\)i⋅=qi⊤\(P\_\{q\}\)\_\{i\\cdot\}=q\_\{i\}^\{\\top\}\.
###### Theorem 4\(Fixed\-policy solvability\)\.
Fixπ∈ΠS\\pi\\in\\Pi\_\{S\}\. For effective selectorsq∈∏i𝒰iπq\\in\\prod\_\{i\}\\mathcal\{U\}\_\{i\}^\{\\pi\}, defineη\(q\):=Pq∞rπ,\\eta\(q\):=P\_\{q\}^\{\\infty\}r^\{\\pi\},γiπ:=infqηi\(q\),\\gamma\_\{i\}^\{\\pi\}:=\\inf\_\{q\}\\eta\_\{i\}\(q\),𝒬∗π:=\{q:η\(q\)=γπ\},\\mathcal\{Q\}\_\{\*\}^\{\\pi\}:=\\\{q:\\eta\(q\)=\\gamma^\{\\pi\}\\\},andw\(q\):=ZPq\(rπ−η\(q\)\)\.w\(q\):=Z\_\{P\_\{q\}\}\(r^\{\\pi\}\-\\eta\(q\)\)\.Then equation[11](https://arxiv.org/html/2609.28792#S4.E11)has a finite solution if and only if
𝒬∗π≠∅and∃B<∞such thatw\(q\)≥−B𝟏for anyq∈𝒬∗π\.\\mathcal\{Q\}\_\{\*\}^\{\\pi\}\\neq\\varnothing\\quad\\text\{and\}\\quad\\exists B<\\infty\\text\{ such that \}w\(q\)\\geq\-B\\mathbf\{1\}\\ \\text\{ for any \}q\\in\\mathcal\{Q\}\_\{\*\}^\{\\pi\}\.\(13\)Every finite solution\(g,h\)\(g,h\)satisfiesg=gπ=γπg=g^\{\\pi\}=\\gamma^\{\\pi\}\.
The two conditions identify separate requirements\. The setQ∗πQ\_\{\*\}^\{\\pi\}contains the stationary kernels that attain the robust gain simultaneously from every state\. For such a kernel,w\(q\)w\(q\)is its canonical transient reward correction, with normalizationPq∞w\(q\)=0P\_\{q\}^\{\\infty\}w\(q\)=0\. The uniform lower bound prevents these normalized corrections from becoming arbitrarily negative across gain\-attaining kernels\. Average\-gain attainment and finite transient corrections are therefore distinct parts of Bellman solvability\.
For optimal control, the candidate gain supplies the appropriate reward centering\. Every gain\-active controller\-nature pair satisfiesPπ,qg=gP^\{\\pi,q\}g=g, so replacingriar\_\{ia\}bycia=ria−gic\_\{ia\}=r\_\{ia\}\-g\_\{i\}subtractsggfrom its original gain\. The bias equation is therefore a zero\-gain problem on the gain\-restricted action and row sets\. A stationary saddle secures this zero gain, and a bound on transient corrections determines whether the saddle value has a finite Bellman representation\. The active controller policies and nature selectors areΠg:=∏iAg\(i\)\\Pi\_\{g\}:=\\prod\_\{i\}A\_\{g\}\(i\),𝒬g:=∏i∏a∈Ag\(i\)Fia\(g\)\\mathcal\{Q\}\_\{g\}:=\\prod\_\{i\}\\prod\_\{a\\in A\_\{g\}\(i\)\}F\_\{ia\}\(g\), and𝒬gπ:=∏iFi,π\(i\)\(g\)\\mathcal\{Q\}\_\{g\}^\{\\pi\}:=\\prod\_\{i\}F\_\{i,\\pi\(i\)\}\(g\)\. A full planq∈𝒬gq\\in\\mathcal\{Q\}\_\{g\}specifies rows for every active action, while a replyq∈𝒬gπq\\in\\mathcal\{Q\}\_\{g\}^\{\\pi\}specifies only the rows used byπ\\pi\. Accordingly,Pπ,qP\_\{\\pi,q\}has rowqi,π\(i\)⊤q\_\{i,\\pi\(i\)\}^\{\\top\}for a full plan and rowqi⊤q\_\{i\}^\{\\top\}for a restricted reply\. Setciπ=ci,π\(i\)c\_\{i\}^\{\\pi\}=c\_\{i,\\pi\(i\)\},ηπ,q=Pπ,q∞cπ\\eta^\{\\pi,q\}=P\_\{\\pi,q\}^\{\\infty\}c^\{\\pi\}, andwπ,q=ZPπ,qcπw^\{\\pi,q\}=Z\_\{P\_\{\\pi,q\}\}c^\{\\pi\}whenηπ,q=0\\eta^\{\\pi,q\}=0\. Hereηπ,q\\eta^\{\\pi,q\}uses the centered rewardscc, whereasη\(q\)\\eta\(q\)in Theorem[4](https://arxiv.org/html/2609.28792#Thmtheorem4)uses the original rewardsrπr^\{\\pi\}\.
###### Theorem 5\(Optimal\-control solvability\)\.
Fixg∈ℝSg\\in\\mathbb\{R\}^\{S\}withg=T^gg=\\widehat\{T\}g\. There is a finitehhwithKgh=hK\_\{g\}h=hif and only if there existπ¯∈Πg\\bar\{\\pi\}\\in\\Pi\_\{g\}, a full planq¯∈𝒬g\\bar\{q\}\\in\\mathcal\{Q\}\_\{g\}, andB<∞B<\\inftysuch that
\(i\)ηπ¯,q≥0,for everyq∈𝒬gπ¯;\(ii\)ηπ,q¯≤0,for everyπ∈Πg;\\displaystyle\(i\)\\eta^\{\\bar\{\\pi\},q\}\\geq 0,\\penalty\\ \\text\{for every \}q\\in\\mathcal\{Q\}\_\{g\}^\{\\bar\{\\pi\}\};\\quad\(ii\)\\eta^\{\\pi,\\bar\{q\}\}\\leq 0,\\penalty\\ \\text\{for every \}\\pi\\in\\Pi\_\{g\};\(iii\)wπ¯,q≥−B𝟏,for everyq∈𝒬gπ¯withηπ¯,q=0\.\\displaystyle\(iii\)w^\{\\bar\{\\pi\},q\}\\geq\-B\\mathbf\{1\},\\penalty\\ \\text\{for every \}q\\in\\mathcal\{Q\}\_\{g\}^\{\\bar\{\\pi\}\}\\text\{ with \}\\eta^\{\\bar\{\\pi\},q\}=0\.Equivalently, finite solvability is equivalent tosupN≥0‖KgN0‖∞<∞\\sup\_\{N\\geq 0\}\\\|K\_\{g\}^\{N\}0\\\|\_\{\\infty\}<\\infty\. Another equivalent condition is the existence of finiteℓ,u\\ell,uwithℓ≤Kgℓ\\ell\\leq K\_\{g\}\\ellandKgu≤uK\_\{g\}u\\leq u\.
Conditions\(i\)\(i\)and\(ii\)\(ii\)give a zero\-gain saddle within the gain\-restricted game: one controller secures nonnegative gain against every restricted nature reply, and one full nature plan holds every active controller to nonpositive gain\. Condition\(iii\)\(iii\)supplies the remaining transient control\. Each fixed zero\-gain reply has a finite canonical bias, but the biases can lack a common lower bound as transition probabilities vanish and recurrent classes change\. The bound is one\-sided because, after fixing nature’s full plan, finitely many deterministic controller policies remain\. The theorem thus locates the gap between stationary gain optimality and a common finite Bellman bias\. Appendix[J](https://arxiv.org/html/2609.28792#A10)further treats the case in which the gain is not prescribed\.
A prescribed average\-optimal controller or stationary saddle can fail to share a Bellman bias even when another pair supports a finite certificate \(Example[3](https://arxiv.org/html/2609.28792#Thmexample3)\)\. Appendix[K](https://arxiv.org/html/2609.28792#A11)characterizes which active controller\-nature pairs share one bias satisfying both players’ Bellman inequalities\. The characterization determines the exact minimum span for each compatible pair and, after optimization over pairs, for the full Bellman system\. It also describes the remaining bias freedom through recurrent\-class offsets, which can persist after fixing one reference state\.
Concrete sufficient regimes for solvability\.The following conditions ensure finite Bellman solvability while allowing different recurrent classes to have different gains \(beyond constant gains\)\. They provide concrete model classes covered by the preceding criterion and the planning result later\.
###### Propsition 2\(Sufficient regimes for Bellman solvability\)\.
Under the standing assumptions, each of the following conditions ensures a finite solution\(g⋆,h\)\(g^\{\\star\},h\)to equation[9](https://arxiv.org/html/2609.28792#S4.E9):
\(A\)Polytopic ambiguity:Every row set𝒰ia\\mathcal\{U\}\_\{ia\}is a polytope\.
\(B\)Uniform support gap:There existsδ\>0\\delta\>0such that every feasible row satisfiespj=0p\_\{j\}=0orpj≥δp\_\{j\}\\geq\\delta, for alli,ai,a,p∈𝒰iap\\in\\mathcal\{U\}\_\{ia\}, andj∈Sj\\in S\.
\(C\)Continuous stationary projections:For everyπ∈ΠD\\pi\\in\\Pi\_\{D\}, the mapP↦P∞P\\mapsto P^\{\\infty\}is continuous on the compact induced\-kernel family𝒫π=\{Pπ,q:q∈𝒬S\}\\mathcal\{P\}^\{\\pi\}=\\\{P^\{\\pi,q\}:q\\in\\mathcal\{Q\}\_\{S\}\\\}\.
Polytopic rows makeTTpiecewise affine, so Kohlberg’s invariant half\-line theorem gives a finite pair[Kohlberg \(1980\)](https://arxiv.org/html/2609.28792#bib.bib69)\. Under \(C\), stationary gains vary continuously with the kernel: compactness upgrades simultaneous stationary approximations to an exact full nature plan, while continuity of the Cesàro projections uniformly bounds the canonical corrections of zero\-gain replies\. These are the requirements in Theorem[5](https://arxiv.org/html/2609.28792#Thmtheorem5); condition \(B\) implies \(C\), as proved in Appendix[M](https://arxiv.org/html/2609.28792#A13)\. These regimes accommodate a nonconstant optimal gain\. In that case,ϵsp\(Vϵ\)→sp\(g⋆\)\>0\\epsilon\\operatorname\{sp\}\(V\_\{\\epsilon\}\)\\to\\operatorname\{sp\}\(g^\{\\star\}\)\>0, so the uncentered discounted span grows as the discount vanishes\. The existence analysis in Appendix[M\.1](https://arxiv.org/html/2609.28792#A13.SS1)controls the finite remainderVϵ−g⋆/ϵV\_\{\\epsilon\}\-g^\{\\star\}/\\epsilonin these regimes\. This vector centering isolates transient rewards while preserving the distinct long\-run gains of recurrent classes\. It extends the role played by bounded discounted span in constant\-gain robust theory\([Wang and Si, 2025](https://arxiv.org/html/2609.28792#bib.bib66), Theorem 4\)\.
Further structure of Bellman certificates\.Appendix[K](https://arxiv.org/html/2609.28792#A11)refines the existence result by characterizing which active controller\-nature pairs share one bias satisfying both players’ Bellman inequalities\. An average\-optimal pair can fail this compatibility test even when a different pair supports a finite certificate\. The characterization determines the exact minimum span for each compatible pair and, after optimization over pairs, for the full Bellman system\. It also describes the remaining bias freedom through recurrent\-class offsets, which can persist after fixing one reference state\.
## 6Planning from Finite Bellman Certificates
A finite Bellman certificate supplies asymptotic comparison points for undiscounted planning with a state\-dependent gain\. Algorithm[1](https://arxiv.org/html/2609.28792#alg1)is inspired by the approximately shifted Halpern update of\([Zurek and Chen, 2025](https://arxiv.org/html/2609.28792#bib.bib65), Algorithm 1\), with the robust operatorTT\. Phase I computesxN=TN0x\_\{N\}=T^\{N\}0and uses it both to estimate the gain asg^N=xN/N\\widehat\{g\}\_\{N\}=x\_\{N\}/Nand to initialize Phase II\. The second phase anchors atxNx\_\{N\}and shifts Bellman updates by this estimate\. We assumeTTcan be exactly computed and applied\.
Algorithm 1Approximately Shifted Robust Halpern Iteration1:Input:
N≥1N\\geq 1and an exact robust Bellman oracle
TT\.
2:Initialize:
x0←0x\_\{0\}\\leftarrow 0\.
3:Phase I: Gain estimation and warm start\.
4:for
j=0,1,…,N−1j=0,1,\\ldots,N\-1do
5:
xj\+1←Txjx\_\{j\+1\}\\leftarrow Tx\_\{j\}\.
6:endfor
7:
g^N←xN/N\\widehat\{g\}\_\{N\}\\leftarrow x\_\{N\}/Nand
z0←xNz\_\{0\}\\leftarrow x\_\{N\}\.
8:Phase II: Approximately shifted Halpern iteration\.
9:for
t=0,1,…,N−1t=0,1,\\ldots,N\-1do
10:
zt\+1←2t\+3z0\+t\+1t\+3\(Tzt−g^N\)\\displaystyle z\_\{t\+1\}\\leftarrow\\frac\{2\}\{t\+3\}z\_\{0\}\+\\frac\{t\+1\}\{t\+3\}\\bigl\(Tz\_\{t\}\-\\widehat\{g\}\_\{N\}\\bigr\)\.
11:endfor
12:
ZN←zNZ\_\{N\}\\leftarrow z\_\{N\}\.
13:Evaluate the action scores at
ZNZ\_\{N\}to obtain
TZNTZ\_\{N\}and choose
πN\(i\)∈argmaxa∈A\(i\)\{ria\+minp∈𝒰iap⊤ZN\}\\pi\_\{N\}\(i\)\\in\\arg\\max\_\{a\\in A\(i\)\}\\\{r\_\{ia\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}Z\_\{N\}\\\}for every
i∈Si\\in S\.
14:Output:
g^N\\widehat\{g\}\_\{N\},
ZNZ\_\{N\},
TZN−ZNTZ\_\{N\}\-Z\_\{N\}, and
πN\\pi\_\{N\}\.
###### Theorem 6\.
Under the standing finite\-state, finite\-action, compact post\-action\(s,a\)\(s,a\)\-rectangular model, suppose equation[9](https://arxiv.org/html/2609.28792#S4.E9)has a finite solution\. Then Algorithm[1](https://arxiv.org/html/2609.28792#alg1)satisfies, in supremum norm,
g^N⟶g⋆,TZN−ZN⟶g⋆,ZN/N⟶g⋆\.\\widehat\{g\}\_\{N\}\\longrightarrow g^\{\\star\},\\qquad TZ\_\{N\}\-Z\_\{N\}\\longrightarrow g^\{\\star\},\\qquad Z\_\{N\}/N\\longrightarrow g^\{\\star\}\.\(14\)Moreover, there is a finiteN0N\_\{0\}such that every outputπN\\pi\_\{N\}at anyN≥N0N\\geq N\_\{0\}satisfiesgπN=g⋆g^\{\\pi\_\{N\}\}=g^\{\\star\}, and each such controller is optimal from every initial state against history\-dependent nature\.
For a finite solution\(g,h\)\(g,h\), Theorem[3](https://arxiv.org/html/2609.28792#Thmtheorem3)and Lemma[6](https://arxiv.org/html/2609.28792#Thmlemma6)imply the affine defectωh\(t\):=‖T\(h\+tg\)−h−\(t\+1\)g‖∞⟶0\.\\omega\_\{h\}\(t\):=\\\|T\(h\+tg\)\-h\-\(t\+1\)g\\\|\_\{\\infty\}\\longrightarrow 0\.This relation, however, need not become exact at any finitett: under general compact ambiguity, rows outside the gain face can remain preferable when their bias advantage offsets a small loss in continuation gain \(Appendix[L](https://arxiv.org/html/2609.28792#A12)gives an example withωh\(t\)\>0\\omega\_\{h\}\(t\)\>0for every sufficiently large finitett\)\. We therefore retain the vanishing defectωh\(t\)\\omega\_\{h\}\(t\)in the Halpern comparison and establish bothTZN−ZN→g⋆TZ\_\{N\}\-Z\_\{N\}\\to g^\{\\star\}andZN/N→g⋆Z\_\{N\}/N\\to g^\{\\star\}\. The first limit controls one\-step reward balance; the second eventually excludes each controller action withmia\(g⋆\)<gi⋆m\_\{ia\}\(g^\{\\star\}\)<g\_\{i\}^\{\\star\}from the greedy rule\. At every remaining action, every feasible nature row satisfiesp⊤g⋆≥gi⋆p^\{\\top\}g^\{\\star\}\\geq g\_\{i\}^\{\\star\}, so the displacement error bounds the controller’s statewise robust gain loss\. There are finitely many deterministic controllers; hence vanishing loss makes every sufficiently late greedy controller exactly optimal from all states\. This gain\-active identification is the additional step required when the optimal gain is a vector\. The affine defect also affects the gain estimator’s rate, which is discussed in Appendix[L](https://arxiv.org/html/2609.28792#A12)\.
Numerical verification\.We further numerically evaluate Algorithm[1](https://arxiv.org/html/2609.28792#alg1)against a nominal counterpart using the same anchored updates[Zurek and Chen \(2025\)](https://arxiv.org/html/2609.28792#bib.bib65)on three multichain models: boundary leakage, periodic recurrent classes, and a safe\-risky decision\. Figure[1](https://arxiv.org/html/2609.28792#S6.F1)supports the theoretical convergence of Algorithm[1](https://arxiv.org/html/2609.28792#alg1)and shows that its extracted controllers attain the optimal robust gain in all three models\. The nominal method converges for its reference model but selects controllers with strictly smaller worst\-case gains\. Appendix[B](https://arxiv.org/html/2609.28792#A2)provides all details and further empirical analysis\.
Figure 1:Robust and nominal planning against iterations\. Left logarithmic axis: robust\-target joint errorEN=max\{‖g^N−g⋆‖∞,‖TvN−vN−g⋆‖∞\}E\_\{N\}=\\max\\\{\\\|\\widehat\{g\}\_\{N\}\-g^\{\\star\}\\\|\_\{\\infty\},\\\|Tv\_\{N\}\-v\_\{N\}\-g^\{\\star\}\\\|\_\{\\infty\}\\\}\(lower is better\)\. Right linear axis: exact robust gain for a specific initial stategxπNg^\{\\pi\_\{N\}\}\_\{x\}of the extracted controller \(higher is better\); the black dashed line marks the optimum obtained by solving Bellman equations\.
## 7Conclusion
We developed a vector Bellman certificate for multichain robust average\-reward MDPs\. Our coupled gain\-bias system preserves gain priority for both players and certifies stationary saddle strategies against history\-dependent opponents from every initial state\. Its solvability criterion connects stationary gain attainment to uniform control of canonical transient corrections, with concrete sufficient conditions permitting distinct recurrent\-class gains\. This certificate also provides the asymptotic structure needed for direct undiscounted planning\. Under finite solvability, our approximately shifted Halpern iteration recovers the vector gain through its estimates and Bellman displacements, and its greedy controllers are eventually average\-optimal\. Our studies thus provided comprehensive and systemic understandings of robust average\-reward MDPs beyond constant gains\.
## References
- Abounadiet al\.\(2001\)J\. Abounadi, D\. P\. Bertsekas, and V\. S\. BorkarLearning algorithms for markov decision processes with average cost\.SIAM Journal on Control and Optimization40\(3\),pp\. 681–698\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Akianet al\.\(2012\)M\. Akian, J\. Cochet\-Terrasson, S\. Detournay, and S\. GaubertPolicy iteration algorithm for zero\-sum multichain stochastic games with mean payoff and perfect information\.arXiv preprint arXiv:1208\.0446\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p2.1),[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p1.1),[§M\.1](https://arxiv.org/html/2609.28792#A13.SS1.p6.1)\.
- Akian and Gaubert \(2003\)M\. Akian and S\. GaubertSpectral theorem for convex monotone homogeneous maps, and ergodic control\.Nonlinear Analysis: Theory, Methods & Applications52\(2\),pp\. 637–679\.Cited by:[§A\.3](https://arxiv.org/html/2609.28792#A1.SS3.p2.1),[§K\.5](https://arxiv.org/html/2609.28792#A11.SS5.p3.1)\.
- Badrinath and Kalathil \(2021\)K\. P\. Badrinath and D\. KalathilRobust reinforcement learning using least squares policy iteration with provable performance guarantees\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 511–520\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Basuet al\.\(2013\)A\. Basu, K\. Martin, and C\. T\. RyanProjection: a unified approach to semi\-infinite linear programs and duality in convex programming\.arxiv preprint arXiv:1304\.3030v2\.Cited by:[§A\.3](https://arxiv.org/html/2609.28792#A1.SS3.p3.1),[§K\.1](https://arxiv.org/html/2609.28792#A11.SS1.p4.3)\.
- Bewley and Kohlberg \(1976\)T\. Bewley and E\. KohlbergThe asymptotic theory of stochastic games\.Mathematics of Operations Research1\(3\),pp\. 197–208\.External Links:[Document](https://dx.doi.org/10.1287/moor.1.3.197),[Link](https://pubsonline.informs.org/doi/10.1287/moor.1.3.197)Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p2.1)\.
- Blackwell \(1962\)D\. BlackwellDiscrete dynamic programming\.The Annals of Mathematical Statistics,pp\. 719–726\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p1.1)\.
- Blanchetet al\.\(2023\)J\. Blanchet, M\. Lu, T\. Zhang, and H\. ZhongDouble pessimism is provably efficient for distributionally robust offline reinforcement learning: generic algorithm and robust partial coverage\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.36\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Bolteet al\.\(2015\)J\. Bolte, S\. Gaubert, and G\. VigeralDefinable zero\-sum stochastic games\.Mathematics of Operations Research40\(1\),pp\. 171–191\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p4.1)\.
- Chatterjeeet al\.\(2024\)K\. Chatterjee, E\. K\. Goharshady, M\. Karrabi, P\. Novotnỳ, and Đ\. ŽikelićSolving long\-run average reward robust mdps via stochastic games\.InProc\. International Joint Conferences on Artificial Intelligence \(IJCAI\),Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p3.1),[§1](https://arxiv.org/html/2609.28792#S1.p3.1)\.
- Chenet al\.\(2025\)Z\. Chen, S\. Wang, and N\. SiSample complexity of distributionally robust average\-reward reinforcement learning\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.38,pp\. 85402–85463\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1)\.
- Denardo and Fox \(1968\)E\. V\. Denardo and B\. L\. FoxMultichain Markov renewal programs\.SIAM Journal on Applied Mathematics16\(3\),pp\. 468–487\.External Links:[Document](https://dx.doi.org/10.1137/0116038),[Link](https://epubs.siam.org/doi/10.1137/0116038)Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p1.1)\.
- Dermanet al\.\(2021\)E\. Derman, M\. Geist, and S\. MannorTwice regularized MDPs and the equivalence between robustness and regularization\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Epstein and Schneider \(2003\)L\. G\. Epstein and M\. SchneiderRecursive multiple\-priors\.Journal of Economic Theory113\(1\),pp\. 1–31\.External Links:[Document](https://dx.doi.org/10.1016/S0022-0531%2803%2900097-8),[Link](https://people.bu.edu/lepstein/files-research/RMP-JET2003.pdf)Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p2.1)\.
- Gaubert and Gunawardena \(2004\)S\. Gaubert and J\. GunawardenaThe perron\-frobenius theorem for homogeneous, monotone functions\.Transactions of the American Mathematical Society356\(12\),pp\. 4931–4950\.Cited by:[§A\.3](https://arxiv.org/html/2609.28792#A1.SS3.p2.1),[§J\.2](https://arxiv.org/html/2609.28792#A10.SS2.p2.1.1)\.
- Ghoshet al\.\(2026a\)D\. Ghosh, G\. K\. Atia, and Y\. WangOnline robust reinforcement learning with general function approximation\.InProc\. International Conference on Machine Learning \(ICML\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Ghoshet al\.\(2026b\)D\. Ghosh, G\. K\. Atia, and Y\. WangORVIT: near\-optimal online distributionally robust reinforcement learning\.InAnnual AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 21278–21286\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i25.39273)Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Goyal and Grand\-Clement \(2023\)V\. Goyal and J\. Grand\-ClementRobust markov decision processes: beyond rectangularity\.Mathematics of Operations Research48\(1\),pp\. 203–226\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p3.1)\.
- Grand\-Clementet al\.\(2023\)J\. Grand\-Clement, M\. Petrik, and N\. VieilleBeyond discounted returns: robust markov decision processes with average and blackwell optimality\.arXiv preprint arXiv:2312\.03618v3\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p6.1),[§C\.1](https://arxiv.org/html/2609.28792#A3.SS1.p1.1.1),[§C\.3](https://arxiv.org/html/2609.28792#A3.SS3.p1.3),[Appendix C](https://arxiv.org/html/2609.28792#A3.p2.1),[Appendix C](https://arxiv.org/html/2609.28792#A3.p3.1),[Appendix D](https://arxiv.org/html/2609.28792#A4.p1.1),[Appendix D](https://arxiv.org/html/2609.28792#A4.p2.1.1),[Appendix E](https://arxiv.org/html/2609.28792#A5.p1.1.2),[Appendix E](https://arxiv.org/html/2609.28792#A5.p3.1.1),[Appendix F](https://arxiv.org/html/2609.28792#A6.p1.1),[Appendix G](https://arxiv.org/html/2609.28792#A7.p1.1),[Appendix G](https://arxiv.org/html/2609.28792#A7.p3.1.1),[Appendix H](https://arxiv.org/html/2609.28792#A8.p2.1),[§1](https://arxiv.org/html/2609.28792#S1.p3.1),[Example 6](https://arxiv.org/html/2609.28792#Thmexample6.p1.2.1),[Remark 1](https://arxiv.org/html/2609.28792#Thmremark1.p1.1),[Remark 2](https://arxiv.org/html/2609.28792#Thmremark2.p1.1.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1),[Theorem 1](https://arxiv.org/html/2609.28792#Thmtheorem1.p1.1.1),[Theorem 2](https://arxiv.org/html/2609.28792#Thmtheorem2.p1.1.1)\.
- Grand\-Clément and Petrik \(2023\)J\. Grand\-Clément and M\. PetrikReducing blackwell and average optimality to discounted mdps via the blackwell discount factor\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.36,pp\. 52628–52647\.Cited by:[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Halpern \(1967\)B\. HalpernFixed points of nonexpanding maps\.Bulletin of the American Mathematical Society73\(6\),pp\. 957–961\.Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p2.1)\.
- Heet al\.\(2025\)Y\. He, Z\. Liu, W\. Wang, and P\. XuSample complexity of distributionally robust off\-dynamics reinforcement learning with online interaction\.InProc\. International Conference on Machine Learning \(ICML\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Hoet al\.\(2018\)C\. P\. Ho, M\. Petrik, and W\. WiesemannFast Bellman updates for robust MDPs\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 1979–1988\.Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p3.1)\.
- Hoet al\.\(2021\)C\. P\. Ho, M\. Petrik, and W\. WiesemannPartial policy iteration for l1\-robust Markov decision processes\.Journal of Machine Learning Research22\(275\),pp\. 1–46\.Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p3.1)\.
- Hordijk and Kallenberg \(1979\)A\. Hordijk and L\. C\. M\. KallenbergLinear programming and Markov decision chains\.Management Science25\(4\),pp\. 352–362\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.25.4.352),[Link](https://pubsonline.informs.org/doi/10.1287/mnsc.25.4.352)Cited by:[§A\.3](https://arxiv.org/html/2609.28792#A1.SS3.p3.1)\.
- Iyengar \(2005\)G\. N\. IyengarRobust dynamic programming\.Mathematics of Operations Research30\(2\),pp\. 257–280\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.28792#S1.p1.1),[§2](https://arxiv.org/html/2609.28792#S2.p1.1),[§2](https://arxiv.org/html/2609.28792#S2.p4.1),[§2](https://arxiv.org/html/2609.28792#S2.p4.2)\.
- Jakschet al\.\(2010\)T\. Jaksch, R\. Ortner, and P\. AuerNear\-optimal regret bounds for reinforcement learning\.Journal of Machine Learning Research11,pp\. 1563–1600\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Kohlberg \(1980\)E\. KohlbergInvariant half\-lines of nonexpansive piecewise\-linear transformations\.Mathematics of Operations Research5\(3\),pp\. 366–372\.Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p1.1),[§M\.1](https://arxiv.org/html/2609.28792#A13.SS1.p4.1.1),[§5](https://arxiv.org/html/2609.28792#S5.p8.1)\.
- Kumaret al\.\(2023a\)N\. Kumar, E\. Derman, M\. Geist, K\. Y\. Levy, and S\. MannorPolicy gradient for rectangular robust markov decision processes\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.36,pp\. 59477–59501\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Kumaret al\.\(2023b\)N\. Kumar, K\. Levy, K\. Wang, and S\. MannorAn efficient solution to s\-rectangular robust markov decision processes\.arXiv preprint arXiv:2301\.13642\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Lee and Ryu \(2025\)J\. Lee and E\. RyuOptimal non\-asymptotic rates of value iteration for average\-reward markov decision processes\.InProc\. International Conference on Learning Representations \(ICLR\),Vol\.2025,pp\. 32823–32865\.Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p2.1)\.
- Li and Shapiro \(2025\)Y\. Li and A\. ShapiroRectangularity and duality of distributionally robust Markov decision processes\.Mathematical Programming\.External Links:[Document](https://dx.doi.org/10.1007/s10107-025-02297-y),[Link](https://link.springer.com/article/10.1007/s10107-025-02297-y)Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p3.1)\.
- Lianget al\.\(2023\)Z\. Liang, X\. Ma, J\. Blanchet, J\. Zhang, and Z\. ZhouSingle\-trajectory distributionally robust reinforcement learning\.arXiv preprint arXiv:2301\.11721\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Lieder \(2021\)F\. LiederOn the convergence rate of the Halpern\-iteration\.Optimization Letters15,pp\. 405–418\.External Links:[Document](https://dx.doi.org/10.1007/s11590-020-01617-9),[Link](https://link.springer.com/article/10.1007/s11590-020-01617-9)Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p2.1)\.
- Liggett and Lippman \(1969\)T\. M\. Liggett and S\. A\. LippmanStochastic games with perfect information and time average payoff\.SIAM Review11\(4\),pp\. 604–607\.External Links:[Document](https://dx.doi.org/10.1137/1011093),[Link](https://epubs.siam.org/doi/10.1137/1011093)Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p2.1)\.
- Limet al\.\(2013\)S\. H\. Lim, H\. Xu, and S\. MannorReinforcement learning in robust Markov decision processes\.InProc\. Advances in Neural Information Processing Systems \(NIPS\),pp\. 701–709\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Liu and Xu \(2024\)Z\. Liu and P\. XuMinimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning\.arXiv preprint arXiv:2403\.09621\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Liuet al\.\(2022\)Z\. Liu, Q\. Bai, J\. Blanchet, P\. Dong, W\. Xu, Z\. Zhou, and Z\. ZhouDistributionally robustQQ\-learning\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 13623–13643\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Luet al\.\(2024\)M\. Lu, H\. Zhong, T\. Zhang, and J\. BlanchetDistributionally robust reinforcement learning with interactive data collection: fundamental hardness and near\-optimal algorithms\.arXiv preprint arXiv:2404\.03578\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Maet al\.\(2021\)X\. Ma, X\. Tang, L\. Xia, J\. Yang, and Q\. ZhaoAverage\-reward reinforcement learning with trust region methods\.arXiv preprint arXiv:2106\.03442\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Mahadevan \(1996a\)S\. MahadevanAn Average\-Reward Reinforcement Learning Algorithm for Computing Bias\-Optimal Policies\.InProceedings of the Thirteenth National Conference on Artificial Intelligence \- Volume 1,AAAI’96,pp\. 875–880\.Note:event\-place: Portland, OregonExternal Links:ISBN 0\-262\-51091\-XCited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Mahadevan \(1996b\)S\. MahadevanAverage reward reinforcement learning: foundations, algorithms, and empirical results\.Machine learning22\(1\),pp\. 159–195\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Mertens and Neyman \(1981\)J\. Mertens and A\. NeymanStochastic games\.International Journal of Game Theory10,pp\. 53–66\.External Links:[Document](https://dx.doi.org/10.1007/BF01769259),[Link](https://link.springer.com/article/10.1007/BF01769259)Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p2.1)\.
- Nilim and El Ghaoui \(2004\)A\. Nilim and L\. El GhaouiRobustness in Markov decision problems with uncertain transition matrices\.InProc\. Advances in Neural Information Processing Systems \(NIPS\),pp\. 839–846\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.28792#S1.p1.1),[§2](https://arxiv.org/html/2609.28792#S2.p1.1),[§2](https://arxiv.org/html/2609.28792#S2.p4.2)\.
- Panaganti and Kalathil \(2022\)K\. Panaganti and D\. KalathilSample complexity of robust reinforcement learning with a generative model\.InProc\. International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 9582–9602\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Panagantiet al\.\(2024\)K\. Panaganti, A\. Wierman, and E\. MazumdarModel\-free robustϕ\\phi\-divergence reinforcement learning using both offline and online data\.arXiv preprint arXiv:2405\.05468\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Panagantiet al\.\(2022\)K\. Panaganti, Z\. Xu, D\. Kalathil, and M\. GhavamzadehRobust reinforcement learning using offline data\.arXiv preprint arXiv:2208\.05129\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Puterman \(2014\)M\. L\. PutermanMarkov decision processes: discrete stochastic dynamic programming\.1 edition,Wiley Series in Probability and Statistics,John Wiley & Sons\(en\)\.External Links:ISBN 978\-0\-471\-61977\-2 978\-0\-470\-31688\-7Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p1.1),[§J\.3](https://arxiv.org/html/2609.28792#A10.SS3.p5.1.1.1),[§C\.2](https://arxiv.org/html/2609.28792#A3.SS2.p3.1.1),[§1](https://arxiv.org/html/2609.28792#S1.p1.1),[§2](https://arxiv.org/html/2609.28792#S2.p3.2),[§2](https://arxiv.org/html/2609.28792#S2.p5.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p4.1),[§5](https://arxiv.org/html/2609.28792#S5.p2.1),[Remark 5](https://arxiv.org/html/2609.28792#Thmremark5.p1.1.1)\.
- Rameshet al\.\(2024\)S\. S\. Ramesh, P\. G\. Sessa, Y\. Hu, A\. Krause, and I\. BogunovicDistributionally robust model\-based reinforcement learning with large state spaces\.InProc\. International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 100–108\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Rochet al\.\(2025\)Z\. Roch, G\. Atia, and Y\. WangA reduction framework for distributionally robust reinforcement learning under average reward\.InProc\. International Conference on Machine Learning \(ICML\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[§2](https://arxiv.org/html/2609.28792#S2.p5.1),[Remark 4](https://arxiv.org/html/2609.28792#Thmremark4.p1.1.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Rochet al\.\(2026\)Z\. Roch, G\. Atia, and Y\. WangModel\-free robust average\-reward reinforcement learning with sample complexity analysis\.InProc\. International Conference on Machine Learning \(ICML\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[Remark 4](https://arxiv.org/html/2609.28792#Thmremark4.p1.1.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Ruszczyński \(2010\)A\. RuszczyńskiRisk\-averse dynamic programming for Markov decision processes\.Mathematical Programming125,pp\. 235–261\.Note:See also the erratum, Mathematical Programming 145:601–604 \(2014\), doi:10\.1007/s10107\-014\-0783\-zExternal Links:[Document](https://dx.doi.org/10.1007/s10107-010-0393-3),[Link](https://link.springer.com/article/10.1007/s10107-010-0393-3)Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p2.1)\.
- Schweitzer and Federgruen \(1978\)P\. J\. Schweitzer and A\. FedergruenThe functional equations of undiscounted markov renewal programming\.Mathematics of Operations Research3\(4\),pp\. 308–321\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p1.1)\.
- Schweitzer \(1985\)P\. J\. SchweitzerOn undiscounted markovian decision processes with compact action spaces\.RAIRO\-Operations Research\-Recherche Opérationnelle19\(1\),pp\. 71–86\.Cited by:[§A\.3](https://arxiv.org/html/2609.28792#A1.SS3.p1.1),[§J\.1](https://arxiv.org/html/2609.28792#A10.SS1.p2.1),[§J\.1](https://arxiv.org/html/2609.28792#A10.SS1.p3.1.1),[§K\.5](https://arxiv.org/html/2609.28792#A11.SS5.p3.1),[§M\.1](https://arxiv.org/html/2609.28792#A13.SS1.p18.1),[§M\.1](https://arxiv.org/html/2609.28792#A13.SS1.p7.2.1),[§M\.2](https://arxiv.org/html/2609.28792#A13.SS2.p1.1),[§C\.2](https://arxiv.org/html/2609.28792#A3.SS2.p2.1.1),[Remark 5](https://arxiv.org/html/2609.28792#Thmremark5.p1.1.1)\.
- Shi and Chi \(2022\)L\. Shi and Y\. ChiDistributionally robust model\-based offline reinforcement learning with near\-optimal sample complexity\.arXiv preprint arXiv:2208\.05767\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Shiet al\.\(2023\)L\. Shi, G\. Li, Y\. Wei, Y\. Chen, M\. Geist, and Y\. ChiThe curious price of distributional robustness in reinforcement learning with a generative model\.arXiv preprint arXiv:2305\.16589\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Sunet al\.\(2024\)Z\. Sun, S\. He, F\. Miao, and S\. ZouPolicy optimization for robust average reward mdps\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.37,pp\. 17348–17372\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1)\.
- Tamaret al\.\(2014\)A\. Tamar, S\. Mannor, and H\. XuScaling up robust MDPs using function approximation\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 181–189\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wanet al\.\(2021\)Y\. Wan, A\. Naik, and R\. S\. SuttonLearning and planning in average\-reward markov decision processes\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 10653–10662\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Wan and Sutton \(2022\)Y\. Wan and R\. S\. SuttonOn convergence of average\-reward off\-policy control algorithms in weakly communicating mdps\.arXiv preprint arXiv:2209\.15141\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Wanet al\.\(2024\)Y\. Wan, H\. Yu, and R\. S\. SuttonOn convergence of average\-reward q\-learning in weakly communicating markov decision processes\.arXiv preprint arXiv:2408\.16262\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Wanget al\.\(2024a\)H\. Wang, L\. Shi, and Y\. ChiSample complexity of offline distributionally robust linear markov decision processes\.arXiv preprint arXiv:2403\.12946\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1),[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Wanget al\.\(2023a\)K\. Wang, U\. Gadot, N\. Kumar, K\. Levy, and S\. MannorBring your own \(non\-robust\) algorithm to solve robust mdps by estimating the worst kernel\.arXiv preprint arxiv: 2306\.05859,pp\. arXiv–2306\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wanget al\.\(2025\)Q\. Wang, Y\. Zha, C\. P\. Ho, and M\. PetrikProvable policy gradient for robust average\-reward mdps beyond rectangularity\.InForty\-second International Conference on Machine Learning,Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1)\.
- Wanget al\.\(2023b\)S\. Wang, N\. Si, J\. Blanchet, and Z\. ZhouA finite sample complexity bound for distributionally robust q\-learning\.InProc\. International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 3370–3398\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wanget al\.\(2023c\)S\. Wang, N\. Si, J\. Blanchet, and Z\. ZhouSample complexity of variance\-reduced distributionally robust q\-learning\.arXiv preprint arXiv:2305\.18420\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wang and Si \(2025\)S\. Wang and N\. SiBellman optimality of average\-reward robust markov decision processes with a constant gain\.arXiv preprint arXiv:2509\.14203v3\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p5.1),[§1](https://arxiv.org/html/2609.28792#S1.p3.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p4.1),[§5](https://arxiv.org/html/2609.28792#S5.p8.1)\.
- Wang and Si \(2026\)S\. Wang and N\. SiNon\-rectangular average\-reward robust mdps: optimal policies and their transient values\.arXiv preprint arXiv:2603\.00945\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p3.1)\.
- Wanget al\.\(2024b\)Y\. Wang, S\. Zou, and Y\. WangModel\-free robust reinforcement learning with sample complexity analysis\.InProc\. International Conference on Uncertainty in Artificial Intelligence \(UAI\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wanget al\.\(2024c\)Y\. Wang, Z\. Sun, and S\. ZouA unified principle of pessimism for offline reinforcement learning under model mismatch\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Wanget al\.\(2023d\)Y\. Wang, A\. Velasquez, G\. K\. Atia, A\. Prater\-Bennette, and S\. ZouModel\-free robust average\-reward reinforcement learning\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 36431–36469\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Wanget al\.\(2023e\)Y\. Wang, A\. Velasquez, G\. Atia, A\. Prater\-Bennette, and S\. ZouRobust average\-reward markov decision processes\.InAnnual AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 15215–15223\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p5.1),[§1](https://arxiv.org/html/2609.28792#S1.p3.1),[§2](https://arxiv.org/html/2609.28792#S2.p5.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p4.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Wanget al\.\(2024d\)Y\. Wang, A\. Velasquez, G\. Atia, A\. Prater\-Bennette, and S\. ZouRobust average\-reward reinforcement learning\.Journal of Artificial Intelligence Research80,pp\. 719–803\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p5.1),[§1](https://arxiv.org/html/2609.28792#S1.p3.1),[§2](https://arxiv.org/html/2609.28792#S2.p5.1)\.
- Wang and Zou \(2021\)Y\. Wang and S\. ZouOnline robust reinforcement learning with model uncertainty\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Vol\.34,pp\. 7193–7206\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wang and Zou \(2022\)Y\. Wang and S\. ZouPolicy gradient method for robust reinforcement learning\.InProc\. International Conference on Machine Learning \(ICML\),Vol\.162,pp\. 23484–23526\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Wiesemannet al\.\(2013\)W\. Wiesemann, D\. Kuhn, and B\. RustemRobust Markov decision processes\.Mathematics of Operations Research38\(1\),pp\. 153–183\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p1.1)\.
- Xu and Mannor \(2010\)H\. Xu and S\. MannorDistributionally robust Markov decision processes\.InProc\. Advances in Neural Information Processing Systems \(NIPS\),pp\. 2505–2513\.Cited by:[§A\.1](https://arxiv.org/html/2609.28792#A1.SS1.p1.1)\.
- Xuet al\.\(2025a\)Y\. Xu, S\. Ganesh, and V\. AggarwalEfficientQQ\-learning and actor\-critic methods for robust average reward reinforcement learning\.arXiv preprint arXiv:2506\.07040\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[§2](https://arxiv.org/html/2609.28792#S2.p5.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Xuet al\.\(2025b\)Y\. Xu, W\. U\. Mondal, and V\. AggarwalFinite\-sample analysis of policy evaluation for robust average reward reinforcement learning\.arXiv preprint arXiv:2502\.16816\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Xuet al\.\(2023\)Z\. Xu, K\. Panaganti, and D\. KalathilImproved sample complexity bounds for distributionally robust reinforcement learning\.InProc\. International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 9728–9754\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Yanget al\.\(2022\)W\. Yang, L\. Zhang, and Z\. ZhangToward theoretical understandings of robust markov decision processes: sample complexity and asymptotics\.The Annals of Statistics50\(6\),pp\. 3223–3248\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1)\.
- Yanget al\.\(2026\)Y\. Yang, Y\. Chen, and Y\. ChiRobust average\-reward markov decision processes: minimax\-optimal learning via plug\-in reductions\.arXiv preprint arXiv:2608\.06545\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p2.1),[Remark 6](https://arxiv.org/html/2609.28792#Thmremark6.p1.1.1)\.
- Zhanget al\.\(2021\)S\. Zhang, Y\. Wan, R\. S\. Sutton, and S\. WhitesonAverage\-reward off\-policy policy evaluation with function approximation\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 12578–12588\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Zhouet al\.\(2021\)Z\. Zhou, Z\. Zhou, Q\. Bai, L\. Qiu, J\. Blanchet, and P\. GlynnFinite\-sample regret bound for distributionally robust offline tabular reinforcement learning\.InProc\. International Conference on Artificial Intelligence and Statistics \(AISTATS\),pp\. 3331–3339\.Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p3.1),[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p4.1)\.
- Ziliotto \(2016\)B\. ZiliottoA tauberian theorem for nonexpansive operators and applications to zero\-sum stochastic games\.Mathematics of Operations Research41\(4\),pp\. 1522–1534\.Cited by:[§A\.2](https://arxiv.org/html/2609.28792#A1.SS2.p4.1),[Appendix D](https://arxiv.org/html/2609.28792#A4.p3.2)\.
- Zurek and Chen \(2024\)M\. Zurek and Y\. ChenSpan\-based optimal sample complexity for weakly communicating and general average reward mdps\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§A\.5](https://arxiv.org/html/2609.28792#A1.SS5.p1.1)\.
- Zurek and Chen \(2025\)M\. Zurek and Y\. ChenFaster fixed\-point methods for multichain mdps\.InProc\. Advances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§A\.4](https://arxiv.org/html/2609.28792#A1.SS4.p2.1),[§K\.5](https://arxiv.org/html/2609.28792#A11.SS5.p3.1),[Appendix L](https://arxiv.org/html/2609.28792#A12.p2.1),[§B\.1](https://arxiv.org/html/2609.28792#A2.SS1.p1.1),[§1](https://arxiv.org/html/2609.28792#S1.p6.1),[§4\.1](https://arxiv.org/html/2609.28792#S4.SS1.p2.1),[§6](https://arxiv.org/html/2609.28792#S6.p1.1),[§6](https://arxiv.org/html/2609.28792#S6.p3.1)\.
## Appendix ARelated work
The paper connects robust dynamic programming, multichain average\-reward theory, and the analysis of nonexpansive operators\. The central distinction is between the existence of a long\-run strategic value and its representation by a finite gain\-bias pair\. We organize the literature around this distinction and its consequences for planning\.
### A\.1Transition ambiguity, rectangularity, and dynamic consistency
Robust and distributionally robust MDPs\.Robust MDPs optimize a policy against a family of plausible transition models\. Classical robust dynamic programming identifies rectangularity assumptions under which local worst\-case transition choices yield a recursive description of the value\[[26](https://arxiv.org/html/2609.28792#bib.bib16),[44](https://arxiv.org/html/2609.28792#bib.bib28),[76](https://arxiv.org/html/2609.28792#bib.bib56)\]\. Distributionally robust formulations also model uncertainty through distributions over model parameters and allow statistical information to enter the ambiguity description\[[77](https://arxiv.org/html/2609.28792#bib.bib57)\]\. The precise uncertainty object and the information available to nature matter: an uncertainty set over transition rows, a distribution over kernels, and a single unknown kernel chosen at the outset need not define the same control problem\. Here nature selects a transition row after observing the current state and action, with independent admissibility constraints across state\-action pairs\. This post\-action\(s,a\)\(s,a\)\-rectangular structure determines the order of optimization in our Bellman operator\.
Dynamic consistency and risk\-averse control\.Rectangularity has a broader interpretation in sequential decision theory\.\[[14](https://arxiv.org/html/2609.28792#bib.bib76)\]connect rectangular sets of priors to recursive multiple\-priors preferences and dynamic consistency\. In Markov control,\[[52](https://arxiv.org/html/2609.28792#bib.bib77)\]develop dynamic programming with Markov risk measures for finite\-horizon and discounted problems\. The connection to our operator is visible at the one\-step level: the lower expectationminp∈𝒰sap⊤v\\min\_\{p\\in\\mathcal\{U\}\_\{sa\}\}p^\{\\top\}vequals the negative of the upper expectationmaxp∈𝒰sap⊤\(−v\)\\max\_\{p\\in\\mathcal\{U\}\_\{sa\}\}p^\{\\top\}\(\-v\)\. Thus worst\-case reward evaluation has a natural risk\-averse interpretation after reversing signs\. These connections explain the recursive structure of robust evaluation; the existence of a finite bias for an undiscounted, state\-dependent gain requires additional long\-run analysis\.
Coupled uncertainty and the timing of nature’s choices\.Rectangularity can also be imposed on a representation of uncertainty rather than directly on individual transition rows\.\[[18](https://arxiv.org/html/2609.28792#bib.bib78)\]study factor\-matrix uncertainty that couples transitions across states while retaining tractability under rectangularity in the factor representation\.\[[32](https://arxiv.org/html/2609.28792#bib.bib79)\]examine the relationship between static and game formulations of distributionally robust MDPs and the role of rectangularity in their equivalence and duality\. For average reward,\[[68](https://arxiv.org/html/2609.28792#bib.bib31)\]study nonrectangular uncertainty with a stationary kernel chosen by nature and history\-dependent controller policies\. These models address different forms of dependence and information\. Our results concern the stagewise post\-action model; this specification is essential to the gain\-restricted row sets and the controller\-nature comparisons used below\.
### A\.2Multichain average reward and stochastic\-game values
Classical multichain optimality equations\.Average\-reward MDPs model continuing decisions without an exogenous discount factor\[[48](https://arxiv.org/html/2609.28792#bib.bib32)\]\. The distinction between long\-run gain and transient bias is classical\.\[[7](https://arxiv.org/html/2609.28792#bib.bib3)\]establish the connection between discounted optimization near discount factor one and undiscounted optimality in finite models\.\[[12](https://arxiv.org/html/2609.28792#bib.bib80)\]develop multichain Markov renewal programming, while\[[53](https://arxiv.org/html/2609.28792#bib.bib71)\]study the solution structure of undiscounted functional equations, including the degrees of freedom associated with optimal recurrent behavior\. In general multichain models the gain can depend on the initial state\. Gain\-bias equations then compare continuation gains first and rewards and biases among gain\-optimal actions second\. Under a fixed policy and kernel, recurrent\-class rewards and absorption probabilities determine the gain vector\. Transition ambiguity adds a second optimization that can change those classes and probabilities\. Our vector Bellman system uses the same gain\-first principle while requiring compatible comparisons for both players\.
Finite stochastic games and mean payoff\.The long\-run value problem also belongs to the theory of zero\-sum stochastic games\. For finite state and action spaces,\[[6](https://arxiv.org/html/2609.28792#bib.bib81)\]establish a common asymptotic limit of normalized finite\-horizon and discounted values, and\[[43](https://arxiv.org/html/2609.28792#bib.bib82)\]prove existence of the uniform value\. The latter is a strategic guarantee across sufficiently long horizons; it does not in general imply that both players have stationary optimal strategies\. Perfect\-information games have additional structure, with classical stationary\-strategy results for time\-average payoff\[[35](https://arxiv.org/html/2609.28792#bib.bib83)\]\. Their multichain gain\-bias structure and policy iteration are developed further by\[[2](https://arxiv.org/html/2609.28792#bib.bib70)\]\. These are direct precedents for robust models with finitely many effective nature actions\.
For polytopic\(s,a\)\(s,a\)\-rectangular ambiguity, minimizing a linear continuation value can be reduced to the finitely many extreme rows\.\[[10](https://arxiv.org/html/2609.28792#bib.bib5)\]exploit the resulting connection to finite turn\-based stochastic games to obtain long\-run robust planning and complexity results\. Consequently, direct average\-reward planning beyond scalar\-gain assumptions already has precedents in the polytopic case\. Our analysis also permits compact curved row sets, where a finite reduction need not be available and finite Bellman solvability must be examined separately\.
Compact\-action games and asymptotic values\.Finiteness of the state space alone does not replace assumptions on the action sets in general stochastic\-game value theory\.\[[9](https://arxiv.org/html/2609.28792#bib.bib84)\]use definability and additional structural conditions to establish uniform values for classes of compact\-action games, including definable perfect\-information games\. Definability includes semialgebraic examples and supplies regularity beyond compactness\. A complementary operator approach relates convergence of normalized discounted and finite\-horizon values through Tauberian theorems\[[85](https://arxiv.org/html/2609.28792#bib.bib67)\]\. We use the latter connection after establishing discounted convergence for the present robust model\. The relevant conclusion is convergence of normalized values for arbitrary compact row sets, without imposing definability\. It should be distinguished from both a finite gain\-bias representation and stronger uniform\-strategy conclusions for general compact\-action games\.
Robust average\-reward theory\.Robust average\-reward Bellman equations and algorithms have been developed under unichain assumptions on the policy\-kernel family\[[72](https://arxiv.org/html/2609.28792#bib.bib49),[73](https://arxiv.org/html/2609.28792#bib.bib52)\]\. More recent theory gives conditions for scalar robust Bellman solvability under broader communication and information structures, including one\-sided weak communication\[[67](https://arxiv.org/html/2609.28792#bib.bib66)\]\. Such assumptions can allow multiple recurrent classes for some choices while still producing an optimal gain independent of the initial state\. Thus the relevant distinction for this paper is state dependence of the optimal gain, rather than simply whether any multichain transition matrix is admissible\.
For compact\(s,a\)\(s,a\)\-rectangular sets,\[[19](https://arxiv.org/html/2609.28792#bib.bib11)\]establish deterministic stationary controller optimality, strong duality, equivalence of the principal average\-payoff conventions, and normalized discounted convergence without a unichain assumption\. They also show that nature’s stationary worst case need not be attained\. These results provide the strategic foundation for our study\. We formulate the all\-state guarantees needed by the Bellman analysis and obtain normalized finite\-horizon convergence using the Tauberian connection\. The subsequent questions are whether the value admits a finite vector gain\-bias certificate, which stationary choices such a certificate supports, and how to recover the gain and a controller policy by direct iteration\.
### A\.3Finite biases, nonlinear operators, and feasibility geometry
Compact\-action Bellman solvability\.The fixed\-policy criterion comes from classical compact\-action MDP theory\. After reversing the reward sign,\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Theorem 1\]characterizes finite one\-player gain\-bias solvability through stationary gain attainment and a uniform lower bound on canonically normalized biases\. The bound controls transient corrections as transition kernels vary; pointwise finiteness for each kernel is insufficient\. We apply this criterion to the nature problem induced by a fixed controller policy\. For robust optimal control, the proof combines a lower barrier obtained from this one\-player result with an upper barrier supplied by a full nature plan\. The two\-player formulation makes explicit which policies, replies, and common bias bounds must be compatible\.
Nonlinear spectral theory and recurrent classes\.Undiscounted Bellman operators are monotone, additively homogeneous with respect to scalar constants, and nonexpansive in the sup norm\. Nonlinear Perron\-Frobenius theory provides bounded\-orbit criteria for additive eigenvectors and fixed points\[[15](https://arxiv.org/html/2609.28792#bib.bib74)\]\. For convex monotone homogeneous maps, critical classes describe the structure and degrees of freedom of eigenspaces\[[3](https://arxiv.org/html/2609.28792#bib.bib72)\]\. These results connect recurrent behavior to Bellman solvability\. A robust max\-min operator, however, need not be convex, so the convex spectral theorem does not apply to it directly\. Our analysis instead fixes controller\-nature selector pairs, uses the classical Poisson representation for their induced chains, and imposes both players’ deviation inequalities on the remaining recurrent\-class offsets\. This identifies when one finite bias supports all required comparisons\.
Linear programming and semi\-infinite certificates\.Linear programming is another classical route to average\-reward control\.\[[25](https://arxiv.org/html/2609.28792#bib.bib85)\]formulate finite MDP average optimization through a single linear program and relate its feasible solutions to stationary policies\. In the present compact\-row model, requiring a bias inequality for every admissible deviation produces a semi\-infinite feasibility problem: there are finitely many bias coordinates but potentially infinitely many constraints\. Projection and duality methods for such systems are developed by\[[5](https://arxiv.org/html/2609.28792#bib.bib73)\]\. Our mixed constraint cone combines controller and nature deviations\. Its closure records limiting inconsistencies that can arise even when no finite combination gives an exact contradiction, and a reward\-to\-flow ratio determines the minimum compatible bias span\. The separation principles are standard; their role here is to characterize a common two\-player Bellman certificate and quantify its size\.
### A\.4Direct planning and anchored iterations
Invariant half\-lines and approximate affine behavior\.For finite\-action perfect\-information games and polytopic robust MDPs, the Bellman operator is piecewise affine\.\[[28](https://arxiv.org/html/2609.28792#bib.bib69)\]show that a nonexpansive piecewise\-linear map admits an invariant half\-line, which describes an eventual affine trajectory with a fixed growth direction\. This structure underlies multichain game algorithms\[[2](https://arxiv.org/html/2609.28792#bib.bib70)\]\. For general compact row sets, a finite Bellman solution can instead yield an asymptotically affine trajectory: its one\-step defect tends to zero, but the trajectory need not become exactly invariant after a finite threshold\. The distinction matters for planning because an argument based on eventual exact equality does not automatically cover curved ambiguity sets\. Our convergence analysis tracks this vanishing defect explicitly\.
Halpern iteration and multichain planning\.Anchored fixed\-point methods originate in the iteration of\[[21](https://arxiv.org/html/2609.28792#bib.bib86)\]\. Quantitative analyses include sharp residual bounds for nonexpansive maps in Hilbert spaces\[[34](https://arxiv.org/html/2609.28792#bib.bib87)\]\. Those results explain the general anchoring mechanism, while Bellman planning requires estimates in the operator’s relevant norm and may involve a nonzero growth direction rather than an ordinary fixed point\. Recent nominal MDP work develops anchored and shifted methods that address these issues\[[31](https://arxiv.org/html/2609.28792#bib.bib75),[87](https://arxiv.org/html/2609.28792#bib.bib65)\]\. In particular, our update follows the approximately shifted iteration of\[[87](https://arxiv.org/html/2609.28792#bib.bib65)\]\. The robust analysis controls the additional affine defect and connects gain and displacement estimates to controller\-policy extraction\. The iteration’s origin, its robust convergence argument, and the structural conditions ensuring a finite bias are therefore separate parts of the comparison\.
Computing robust Bellman updates\.Iteration complexity and the cost of each inner minimization are complementary questions\.\[[23](https://arxiv.org/html/2609.28792#bib.bib14)\]develop efficient exact Bellman updates forℓ1\\ell\_\{1\}ambiguity, and\[[24](https://arxiv.org/html/2609.28792#bib.bib15)\]combine efficient updates with partial policy iteration for discounted robust MDPs\. Such methods can serve as computational components when the row sets in our model have the corresponding structure\. They do not by themselves provide an undiscounted multichain convergence argument\. Conversely, an operator\-level convergence result for arbitrary compact row sets does not imply a uniformly efficient implementation of every inner optimization problem; its computational use depends on how the ambiguity sets are represented\.
### A\.5Average\-reward and robust reinforcement learning
Nominal average\-reward learning\.Average\-reward reinforcement learning includes differential value estimation, temporal\-difference methods, andQQ\-learning\[[42](https://arxiv.org/html/2609.28792#bib.bib26),[41](https://arxiv.org/html/2609.28792#bib.bib27),[1](https://arxiv.org/html/2609.28792#bib.bib1),[83](https://arxiv.org/html/2609.28792#bib.bib62),[40](https://arxiv.org/html/2609.28792#bib.bib25)\]\. Regret\-based work such as\[[27](https://arxiv.org/html/2609.28792#bib.bib17)\]studies exploration in unknown communicating MDPs using a diameter parameter\. Other analyses develop convergence under weak communication and finite\-sample guarantees governed by bias span or related structural quantities\[[59](https://arxiv.org/html/2609.28792#bib.bib41),[60](https://arxiv.org/html/2609.28792#bib.bib42),[61](https://arxiv.org/html/2609.28792#bib.bib43),[86](https://arxiv.org/html/2609.28792#bib.bib64)\]\. These results show why communication, recurrent structure, and bias size are central to both learning and planning\. Our setting isolates the deterministic robust planning and solvability questions with access to the Bellman operator; a statistical learning guarantee would additionally need to control how transition\-estimation errors affect those quantities\.
Robust average\-reward learning\.Existing methods include relative\-value TD andQQ\-learning, policy optimization, discounted reductions, anchored procedures, and stochastic approximation\[[71](https://arxiv.org/html/2609.28792#bib.bib48),[57](https://arxiv.org/html/2609.28792#bib.bib39),[50](https://arxiv.org/html/2609.28792#bib.bib34),[11](https://arxiv.org/html/2609.28792#bib.bib6),[79](https://arxiv.org/html/2609.28792#bib.bib60),[78](https://arxiv.org/html/2609.28792#bib.bib59),[51](https://arxiv.org/html/2609.28792#bib.bib35),[82](https://arxiv.org/html/2609.28792#bib.bib7),[64](https://arxiv.org/html/2609.28792#bib.bib55)\]\. Their assumptions vary, with many guarantees using unichain, irreducibility, or uniform ergodicity conditions that yield a state\-independent robust gain\. These conditions provide ways to control long\-run sensitivity and transient behavior\. Our finite\-bias characterization addresses the structural question that arises when the gain can vary across initial states: which robust models still admit a finite certificate on which a direct planning analysis can be based?
Discounted and finite\-horizon robust learning\.A large literature studies statistical estimation of worst\-case values under discounted or finite\-horizon criteria\. Early sample\-based approaches include robust temporal\-difference learning, approximate dynamic programming, and linear policy evaluation\[[36](https://arxiv.org/html/2609.28792#bib.bib21),[58](https://arxiv.org/html/2609.28792#bib.bib40),[4](https://arxiv.org/html/2609.28792#bib.bib2),[74](https://arxiv.org/html/2609.28792#bib.bib44)\]\. Subsequent tabular analyses cover model\-based estimation and plug\-in planning, as well as model\-free procedures, for several ambiguity families and data\-access models\[[81](https://arxiv.org/html/2609.28792#bib.bib61),[45](https://arxiv.org/html/2609.28792#bib.bib30),[80](https://arxiv.org/html/2609.28792#bib.bib58),[56](https://arxiv.org/html/2609.28792#bib.bib38),[84](https://arxiv.org/html/2609.28792#bib.bib63),[65](https://arxiv.org/html/2609.28792#bib.bib47),[33](https://arxiv.org/html/2609.28792#bib.bib20),[38](https://arxiv.org/html/2609.28792#bib.bib22),[66](https://arxiv.org/html/2609.28792#bib.bib50),[69](https://arxiv.org/html/2609.28792#bib.bib51),[63](https://arxiv.org/html/2609.28792#bib.bib46),[30](https://arxiv.org/html/2609.28792#bib.bib18),[13](https://arxiv.org/html/2609.28792#bib.bib8)\]\. Online robust learning additionally treats exploration\[[74](https://arxiv.org/html/2609.28792#bib.bib44),[39](https://arxiv.org/html/2609.28792#bib.bib24),[17](https://arxiv.org/html/2609.28792#bib.bib9),[22](https://arxiv.org/html/2609.28792#bib.bib13)\]\. Related model\-free methods use robustQQ\-learning, multilevel Monte Carlo, or variance reduction\[[38](https://arxiv.org/html/2609.28792#bib.bib22),[65](https://arxiv.org/html/2609.28792#bib.bib47),[69](https://arxiv.org/html/2609.28792#bib.bib51),[62](https://arxiv.org/html/2609.28792#bib.bib53),[16](https://arxiv.org/html/2609.28792#bib.bib10)\], while policy\-based approaches analyze robust policy\-gradient and actor\-critic methods\[[75](https://arxiv.org/html/2609.28792#bib.bib45),[29](https://arxiv.org/html/2609.28792#bib.bib19)\]\.
Offline learning and function approximation\.Robust offline methods combine ambiguity\-aware pessimism with tabular or fitted value iteration and structured function approximation\[[84](https://arxiv.org/html/2609.28792#bib.bib63),[47](https://arxiv.org/html/2609.28792#bib.bib29),[55](https://arxiv.org/html/2609.28792#bib.bib37),[8](https://arxiv.org/html/2609.28792#bib.bib4),[37](https://arxiv.org/html/2609.28792#bib.bib23),[46](https://arxiv.org/html/2609.28792#bib.bib36),[62](https://arxiv.org/html/2609.28792#bib.bib53),[70](https://arxiv.org/html/2609.28792#bib.bib54)\]\. Other work considers distributional robustness with function approximation or simulator access in large or continuous state spaces\[[49](https://arxiv.org/html/2609.28792#bib.bib33)\]\. These literatures address estimation, coverage, and approximation errors\. Their discounted contraction or finite\-horizon recursion controls the propagation of those errors\. In the undiscounted multichain problem studied here, finite\-bias solvability and state\-dependent growth must first be understood to obtain an analogous foundation for algorithmic analysis\.
## Appendix BNumerical experiments
We examine convergence of the computed gain and Bellman displacement, and robust average\-reward performance of the extracted controller\. The three models are designed stress tests in which midpoint nominal parameters favor an action with a smaller worst\-case gain\. The first two have curved, nonpolytopic ambiguity and isolate boundary leakage and periodicity, respectively\. The third is a polytopic safe\-risky decision\. Each model has a finite gain\-bias certificate, specified below, so the solvability hypothesis of Theorem[6](https://arxiv.org/html/2609.28792#Thmtheorem6)is satisfied\.
### B\.1Methods, budgets, and evaluation
Methods and initialization\.LetTTdenote the robust Bellman operator and letT0T\_\{0\}replace each uncertain row set by the nominal reference row specified below\. We compare Algorithm[1](https://arxiv.org/html/2609.28792#alg1), usingTT, with a nominal ablation usingT0T\_\{0\}\. The latter is the fixed\-reference multichain method of\[[87](https://arxiv.org/html/2609.28792#bib.bib65)\]\. For either planning operator𝒯∈\{T,T0\}\\mathcal\{T\}\\in\\\{T,T\_\{0\}\\\}, the run starts at zero, formsxN=𝒯N0x\_\{N\}=\\mathcal\{T\}^\{N\}0, setsg^N=xN/N\\widehat\{g\}\_\{N\}=x\_\{N\}/Nandz0=xNz\_\{0\}=x\_\{N\}, and performs
zt\+1=2z0\+\(t\+1\)\(𝒯zt−g^N\)t\+3,t=0,…,N−1\.z\_\{t\+1\}=\\frac\{2z\_\{0\}\+\(t\+1\)\(\\mathcal\{T\}z\_\{t\}\-\\widehat\{g\}\_\{N\}\)\}\{t\+3\},\\qquad t=0,\\ldots,N\-1\.The output vector isvN=zNv\_\{N\}=z\_\{N\}\. One final evaluation of𝒯vN\\mathcal\{T\}v\_\{N\}extracts a controller greedy for that method’s own operator\. Both methods therefore use2N\+12N\+1planning evaluations\. Tied action values are resolved in favor ofccin Models I and II and the risky action in Model III\. The nominal controller is evaluated under the robust model with its selected action fixed\.
Budget convention\.We evaluate every integerN=1,…,16384N=1,\\ldots,16384, including both parities\. Each point represents a complete run from zero with that budget; the anchor and estimated gain depend onNN\. Thus2N\+12N\+1measures the planning budget of a run\. The implementation shares common first\-phase iterates and batches the independent second\-phase calculations, producing the same outputs as individually initialized runs\. Batched execution time is recorded separately from the per\-run oracle count\. The nominal method uses one additional robust operator evaluation to measure its robust displacement\. This external diagnostic is excluded from its nominal planning budget and is never fed into its updates or policy selection\.
Convergence errors\.All methods are evaluated against the exact robust gaing⋆g^\{\\star\}\. Define
Eg\(N\)=‖g^N−g⋆‖∞,Ed\(N\)=‖TvN−vN−g⋆‖∞,EN=max\{Eg\(N\),Ed\(N\)\}\.E\_\{g\}\(N\)=\\\|\\widehat\{g\}\_\{N\}\-g^\{\\star\}\\\|\_\{\\infty\},\\qquad E\_\{d\}\(N\)=\\\|Tv\_\{N\}\-v\_\{N\}\-g^\{\\star\}\\\|\_\{\\infty\},\\qquad E\_\{N\}=\\max\\\{E\_\{g\}\(N\),E\_\{d\}\(N\)\\\}\.\(15\)The main figure usesENE\_\{N\}to show both convergence targets in one panel per model\. The top and middle rows of Figure[2](https://arxiv.org/html/2609.28792#A2.F2)display the two components separately\. For Algorithm[1](https://arxiv.org/html/2609.28792#alg1), Theorem[6](https://arxiv.org/html/2609.28792#Thmtheorem6)givesEN→0E\_\{N\}\\to 0under finite Bellman solvability\. This is an asymptotic statement and does not require the errors to decrease at every finite budget\. For the nominal method, this robust\-target error includes disagreement between the nominal and robust objectives\. We separately measure each solver’s own\-objective error:
ENown=max\{‖g^N−g𝒯⋆‖∞,‖𝒯vN−vN−g𝒯⋆‖∞\},E\_\{N\}^\{\\mathrm\{own\}\}=\\max\\bigl\\\{\\\|\\widehat\{g\}\_\{N\}\-g\_\{\\mathcal\{T\}\}^\{\\star\}\\\|\_\{\\infty\},\\\|\\mathcal\{T\}v\_\{N\}\-v\_\{N\}\-g\_\{\\mathcal\{T\}\}^\{\\star\}\\\|\_\{\\infty\}\\bigr\\\},whereg𝒯⋆=g⋆g\_\{\\mathcal\{T\}\}^\{\\star\}=g^\{\\star\}for𝒯=T\\mathcal\{T\}=Tandg𝒯⋆=g0⋆g\_\{\\mathcal\{T\}\}^\{\\star\}=g\_\{0\}^\{\\star\}for𝒯=T0\\mathcal\{T\}=T\_\{0\}, withg0⋆g\_\{0\}^\{\\star\}the optimal gain of the nominal reference model\. The bottom row of Figure[2](https://arxiv.org/html/2609.28792#A2.F2)reports this diagnostic\. The data retain full output vectors and statewise robust policy gains\.
Robust evaluation of the output controller\.Each model has one decision statexx\. For the controllerπN\\pi\_\{N\}selected by each method, we plot its actual worst\-case average reward
GN=gxπN=infq∈QSηxπN,q,G⋆=gx⋆\.G\_\{N\}=g\_\{x\}^\{\\pi\_\{N\}\}=\\inf\_\{q\\in Q\_\{S\}\}\\eta\_\{x\}^\{\\pi\_\{N\},q\},\\qquad G^\{\\star\}=g\_\{x\}^\{\\star\}\.\(16\)This is evaluated analytically from the selected controller, rather than estimated by its value iterate or by a finite simulated rollout\. In the first two models, stateyyhas the same gain asxx, and all remaining states have policy\-independent gains\. The latter property also holds in the third model\. Consequently, in every experiment,
‖g⋆−gπN‖∞=G⋆−GN\.\\\|g^\{\\star\}\-g^\{\\pi\_\{N\}\}\\\|\_\{\\infty\}=G^\{\\star\}\-G\_\{N\}\.Reaching the horizontal referenceG⋆G^\{\\star\}therefore verifies optimality from every initial state for these models\. Policy gains are plotted on linear axes, including negative values\.
Exact row minimization and implementation\.Linear minimization over the convex hull of a curve has the same value as minimization over the generating curve\. In Model I, the row objective isvx\+u\(vy−vx\)\+u3\(vz−vx\)v\_\{x\}\+u\(v\_\{y\}\-v\_\{x\}\)\+u^\{3\}\(v\_\{z\}\-v\_\{x\}\); we compare both endpoints and any real stationary point in\[0,1/2\]\[0,1/2\]\. In Model II it is quadratic, so both endpoints and any interior minimizing vertex suffice\. The alternative\-action row objectives and the Model III objective are affine, so their minima occur at interval endpoints\. All calculations use these analytic minimizers in double precision, with no discretization of the ambiguity set and no Monte Carlo policy evaluation\. The computations are deterministic\. The nominal reference rows are specified as part of each model\.
### B\.2Model I: boundary leakage
The ordered state space is\(x,y,z,w,H\)\(x,y,z,w,H\)\. Statexxhas actionscc\(curved\) andtt\(risky\); every other state has one action\. The rewards arer\(x,c\)=0r\(x,c\)=0,r\(x,t\)=10r\(x,t\)=10,r\(y\)=−1r\(y\)=\-1,r\(z\)=1r\(z\)=1,r\(w\)=−2r\(w\)=\-2, andr\(H\)=5r\(H\)=5\. The uncertain rows are
U\(x,c\)\\displaystyle U\(x,c\)=co\{\(1−u−u3,u,u3,0,0\):0≤u≤1/2\},\\displaystyle=\\operatorname\{co\}\\\{\(1\-u\-u^\{3\},u,u^\{3\},0,0\):0\\leq u\\leq 1/2\\\},U\(x,t\)\\displaystyle U\(x,t\)=\{\(0,0,0,1−ρ,ρ\):0≤ρ≤1\}\.\\displaystyle=\\\{\(0,0,0,1\-\\rho,\\rho\):0\\leq\\rho\\leq 1\\\}\.Stateyyreturns toxx, andz,w,Hz,w,Hare absorbing\. These remaining rows are known\. The nominal parameters are the interval midpointsu0=1/4u\_\{0\}=1/4andρ0=1/2\\rho\_\{0\}=1/2\.
The exact robust gain and a gain\-face bias are
g⋆=\(0,0,1,−2,5\),h=\(0,−1,0,0,0\)\.g^\{\\star\}=\(0,0,1,\-2,5\),\\qquad h=\(0,\-1,0,0,0\)\.The curved continuation gain isu3u^\{3\}, minimized atu=0u=0, whereas the risky continuation gain is−2\+7ρ\-2\+7\\rho, minimized atρ=0\\rho=0\. Thus onlyccis gain\-active, and its minimizing gain face is the self\-loop\. The stated bias satisfies the restricted equation there and the deterministic equations at the other states\. The curved controller has robust gaing⋆g^\{\\star\}, while the risky controller has gain\(−2,−2,1,−2,5\)\(\-2,\-2,1,\-2,5\)\. HenceGN∈\{0,−2\}G\_\{N\}\\in\\\{0,\-2\\\}\.
Under the nominal model, actioncceventually reacheszzand has gain11atx,yx,y\. Actiontthas nominal gain\(1−ρ0\)\(−2\)\+5ρ0=3/2\(1\-\\rho\_\{0\}\)\(\-2\)\+5\\rho\_\{0\}=3/2, so it is strictly preferable nominally\. The nominal optimal gain isg0⋆=\(3/2,3/2,1,−2,5\)g\_\{0\}^\{\\star\}=\(3/2,3/2,1,\-2,5\), and its optimal controller has robust loss22\. The immediate reward1010affects transient decisions but contributes zero to the long\-run average after absorption\.
The curved rows create long transients nearu=0u=0: every fixedu\>0u\>0eventually leads tozz, whileu=0u=0leavesxxrecurrent\. This change in recurrent structure makes the model a test of planning near the boundary of the uncertainty set\.
### B\.3Model II: periodic basins
The ordered states are\(x,y,h0,h1,m,ℓ\)\(x,y,h\_\{0\},h\_\{1\},m,\\ell\)\. Statexxhas actionscc\(curved\) andff\(risky\), both with reward zero\. Every other state has one action\. The rewards at\(y,h0,h1,m,ℓ\)\(y,h\_\{0\},h\_\{1\},m,\\ell\)are\(−0\.3,1,−1,−1,1\)\(\-0\.3,1,\-1,\-1,1\)\. For the curved action,
U\(x,c\)=co\{p\(u\):0≤u≤1\},U\(x,c\)=\\operatorname\{co\}\\\{p\(u\):0\\leq u\\leq 1\\\},where, in the stated coordinate order,
p\(u\)=\(0,0,0\.55\+0\.10u,0,0\.25−0\.05u2,0\.20−0\.10u\+0\.05u2\)\.p\(u\)=\(0,0,0\.55\+0\.10u,0,0\.25\-0\.05u^\{2\},0\.20\-0\.10u\+0\.05u^\{2\}\)\.The risky row set is
U\(x,f\)=\{\(0,0,0,0,1−ρ,ρ\):0≤ρ≤1\}\.U\(x,f\)=\\\{\(0,0,0,0,1\-\\rho,\\rho\):0\\leq\\rho\\leq 1\\\}\.Stateyygoes toxx, andh0,h1h\_\{0\},h\_\{1\}alternate deterministically\. Statesm,ℓm,\\ellare absorbing\. The nominal parameters areu0=ρ0=1/2u\_\{0\}=\\rho\_\{0\}=1/2\.
The exact robust certificate is
g⋆=\(−3/40,−3/40,0,0,−1,1\),h=\(27/40,9/20,1,0,0,0\)\.g^\{\\star\}=\(\-3/40,\-3/40,0,0,\-1,1\),\\qquad h=\(27/40,9/20,1,0,0,0\)\.Indeed,p\(u\)⊤g⋆=−3/40\+\(u−1/2\)2/10p\(u\)^\{\\top\}g^\{\\star\}=\-3/40\+\(u\-1/2\)^\{2\}/10, uniquely minimized atu=1/2u=1/2, andp\(1/2\)⊤h=3/5=gx⋆\+hxp\(1/2\)^\{\\top\}h=3/5=g\_\{x\}^\{\\star\}\+h\_\{x\}\. The risky continuation gain is2ρ−12\\rho\-1, whose worst value is−1\-1\. Thusccis strictly gain\-preferred, and the remaining bias equations follow from the deterministic transitions\. The curved and risky controllers have robust gainsg⋆g^\{\\star\}and\(−1,−1,0,0,−1,1\)\(\-1,\-1,0,0,\-1,1\), respectively\. Consequently,GN∈\{−3/40,−1\}G\_\{N\}\\in\\\{\-3/40,\-1\\\}\.
Nominally, the risky action has gain2ρ0−1=02\\rho\_\{0\}\-1=0, which exceeds the curved action’s−3/40\-3/40\. The nominal optimal gain isg0⋆=\(0,0,0,0,−1,1\)g\_\{0\}^\{\\star\}=\(0,0,0,0,\-1,1\), and the nominal optimal controller has robust loss37/4037/40\. The deterministic two\-cycle has alternating rewards and zero average gain\. This model tests convergence with periodic recurrent dynamics and different gains across recurrent classes\.
### B\.4Model III: a safe\-risky decision
This model gives a direct safe\-risky comparison\. There are three states\(x,L,H\)\(x,L,H\)\. StatesL,HL,Hare absorbing with rewards0,10,1, respectively\. Atxx, both available actions have reward zero\. The safe action has known row\(0,0\.6,0\.4\)\(0,0\.6,0\.4\), while the risky action has row set
U\(x,risky\)=\{\(0,1−p,p\):0\.1≤p≤0\.9\}\.U\(x,\\mathrm\{risky\}\)=\\\{\(0,1\-p,p\):0\.1\\leq p\\leq 0\.9\\\}\.The nominal risky row usesp0=0\.5p\_\{0\}=0\.5; the safe row is unchanged\. The robust gain isg⋆=\(0\.4,0,1\)g^\{\\star\}=\(0\.4,0,1\), achieved by the safe action, with biash=\(−0\.4,0,0\)h=\(\-0\.4,0,0\)\. These vectors directly satisfy the gain\-first system; the row sets are also polytopes\. The risky controller has robust gain\(0\.1,0,1\)\(0\.1,0,1\), so its robust loss is0\.30\.3\. In contrast, nominal planning prefers the risky action and has optimal nominal gaing0⋆=\(0\.5,0,1\)g\_\{0\}^\{\\star\}=\(0\.5,0,1\)\. Thus the nominal objective changes the selected controller in this model\. This polytopic example complements the two nonpolytopic constructions by making the robustness distinction explicit\.
### B\.5Results and interpretation
Figure[1](https://arxiv.org/html/2609.28792#S6.F1)pairs each model’s joint error with robust controller gain, and the top and middle rows of Figure[2](https://arxiv.org/html/2609.28792#A2.F2)separate the gain and displacement errors\. Algorithm[1](https://arxiv.org/html/2609.28792#alg1)’s errors are consistent with the predicted convergence to zero, and its controllers attain the optimal robust gain in all three models\. In Model I it selectsttatN=1,2N=1,2andccat every testedN≥3N\\geq 3\. It selects an optimal controller at every tested budget in Models II and III\.
The nominal baseline selects the risky controller at every tested budget in Models I and III\. In Model II it selectsccatN=1,2,3N=1,2,3andffat every testedN≥4N\\geq 4\. Its final robust gains are therefore−2\-2,−1\-1, and0\.10\.1, compared with the optimal gains00,−3/40\-3/40, and0\.40\.4\. These gaps follow from the different objectives: the nominal controller optimizes its reference model, whose favorable outcomes are less reliable under worst\-case transition evaluation\.
AtN=16384N=16384, Algorithm[1](https://arxiv.org/html/2609.28792#alg1)has joint errors approximately5\.51×10−35\.51\\times 10^\{\-3\},6\.10×10−56\.10\\times 10^\{\-5\}, and2\.44×10−52\.44\\times 10^\{\-5\}in Models I\-III, respectively\. The bottom row of Figure[2](https://arxiv.org/html/2609.28792#A2.F2)shows both methods approaching zero error for their own objectives\. The nominal method’s robust displacement can nevertheless grow because its output follows the nominal gain direction\. Its robust policy losses are22,37/4037/40, and0\.30\.3, respectively\. Thus convergence for the reference model and robust controller performance are distinct properties\.
Figure 2:Additional convergence diagnostics at every integer budget\. Columns correspond to Models I\-III\. Top: robust gain\-estimation errorEg\(N\)E\_\{g\}\(N\)\. Middle: robust displacement errorEd\(N\)E\_\{d\}\(N\)\. Bottom: joint errorENownE\_\{N\}^\{\\mathrm\{own\}\}against each method’s own operator and optimal gain\. Both methods converge for their own planning objectives, while the nominal controllers retain the robust policy losses shown in Figure[1](https://arxiv.org/html/2609.28792#S6.F1)\.
## Appendix CPreliminaries and applicability of prior results
We use the model and notation of Section[2](https://arxiv.org/html/2609.28792#S2), withn=\|S\|n=\|S\|,R=maxi,a\|ria\|R=\\max\_\{i,a\}\|r\_\{ia\}\|,∥⋅∥∞\\\|\\cdot\\\|\_\{\\infty\}the supremum norm, andsp\(x\)=maxixi−minixi\\operatorname\{sp\}\(x\)=\\max\_\{i\}x\_\{i\}\-\\min\_\{i\}x\_\{i\}\. Vector inequalities and extrema are understood coordinatewise\. A coordinatewise infimum need not be attained by one selector\. Whenever one selector works for every state, we establish this separately\.
The argument has three inputs: rectangular discounted dynamic programming, compact\-action one\-player average\-payoff results, and a nonexpansive Tauberian theorem\. We cite the standard results and verify their applicability to the present post\-action model\. References to theorem numbers in\[[19](https://arxiv.org/html/2609.28792#bib.bib11)\]use arXiv:2312\.03618v3, dated January 14, 2025\.
Nonconvex row sets\.The compactness hypotheses of\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorems 3\.4\-3\.5\]permit nonconvex sets\. Their finite\-restriction argument should then retain the selected rows themselves\. To check this point, fix an initial lawμ\\muand toleranceξ\>0\\xi\>0\. For each of the finitely manyπ∈ΠD\\pi\\in\\Pi\_\{D\}, chooseqπ∈𝒬Sq^\{\\pi\}\\in\\mathcal\{Q\}\_\{S\}such that
μ⊤ηπ,qπ≤infq∈𝒬Sμ⊤ηπ,q\+ξ\.\\mu^\{\\top\}\\eta^\{\\pi,q^\{\\pi\}\}\\leq\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}\+\\xi\.SetEia=\{qiaπ:π∈ΠD\}E\_\{ia\}=\\\{q^\{\\pi\}\_\{ia\}:\\pi\\in\\Pi\_\{D\}\\\}and𝒬E=∏i,aEia\\mathcal\{Q\}\_\{E\}=\\prod\_\{i,a\}E\_\{ia\}\. Then𝒬E⊆𝒬S\\mathcal\{Q\}\_\{E\}\\subseteq\\mathcal\{Q\}\_\{S\}, and it contains each selectedqπq^\{\\pi\}\. The restricted model is a finite perfect\-information stochastic game, so the finite\-game stationary duality used in their proof gives
infq∈𝒬Smaxπ∈ΠDμ⊤ηπ,q\\displaystyle\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}≤minq∈𝒬Emaxπ∈ΠDμ⊤ηπ,q\\displaystyle\\leq\\min\_\{q\\in\\mathcal\{Q\}\_\{E\}\}\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}=maxπ∈ΠDminq∈𝒬Eμ⊤ηπ,q\\displaystyle=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\min\_\{q\\in\\mathcal\{Q\}\_\{E\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}≤maxinfq∈𝒬Sπ∈ΠDμ⊤ηπ,q\+ξ\.\\displaystyle\\leq\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}\+\\xi\.Weak duality andξ↓0\\xi\\downarrow 0prove the stationary duality needed below for the original row sets\. Convexification is unnecessary for this average\-reward argument\.
### C\.1Rectangular dynamic programming and the fixed\-policy reduction
###### Lemma 1\.
The operatorsTTandTπT^\{\\pi\}are order preserving, additively homogeneous, and nonexpansive in∥⋅∥∞\\\|\\cdot\\\|\_\{\\infty\}\. ForF∈\{T,Tπ\}F\\in\\\{T,T^\{\\pi\}\\\}, the mapx↦F\(\(1−ϵ\)x\)x\\mapsto F\(\(1\-\{\\epsilon\}\)x\)is a\(1−ϵ\)\(1\-\{\\epsilon\}\)\-contraction\. Its unique fixed point is, respectively,VϵV\_\{\\epsilon\}orVϵπV\_\{\\epsilon\}^\{\\pi\}, with norm at mostR/ϵR/\{\\epsilon\}\. These vectors are discounted values against history\-dependent opponents\. In the control problem, both players have deterministic stationary discounted\-optimal selectors that work simultaneously from every state\. For fixedπ∈ΠS\\pi\\in\\Pi\_\{S\}, nature has such a selector\. The finite\-horizon total values areTN0T^\{N\}0and\(Tπ\)N0\(T^\{\\pi\}\)^\{N\}0\.
###### Applicability of standard dynamic programming\.
These are the rectangular dynamic\-programming results summarized in\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Section 2\.1, equations \(2\.3\)\-\(2\.5\), and Proposition 2\.2\], including its Appendix B for history\-dependent nature\. The next\-state\-dependent reward in that reference is specialized here toriaj=riar\_\{iaj\}=r\_\{ia\}\. For the convexity assumption in Proposition 2\.2, each row set may first be replaced by its compact convex hull\. Linear minimization, and hence both Bellman maps, is unchanged\. Compactness then permits every minimizing Bellman row to be chosen in the original𝒰ia\\mathcal\{U\}\_\{ia\}, for every action, while finiteness attains the controller’s maximum\. These selectors satisfy the discounted Bellman inequalities against every admissible original\-model opponent\. The fixed\-policy reduction below gives the same conclusion for randomizedπ\\pi\. The finite\-horizon statement uses the same recursion with terminal value zero\. We use the standard contraction and fixed\-point results without reproving them\. ∎
Forπ∈ΠS\\pi\\in\\Pi\_\{S\}, let nature’s effective action at stateiibe a tupleb=\(pa\)a∈Bi:=∏a∈A\(i\)𝒰iab=\(p\_\{a\}\)\_\{a\}\\in B\_\{i\}:=\\prod\_\{a\\in A\(i\)\}\\mathcal\{U\}\_\{ia\}, with rewardriπr\_\{i\}^\{\\pi\}and transitionPi\(b\)=∑aπ\(a∣i\)paP\_\{i\}\(b\)=\\sum\_\{a\}\\pi\(a\\mid i\)p\_\{a\}\. Thus
𝒰iπ=\{∑aπ\(a∣i\)pa:pa∈𝒰ia\},\(Tπx\)i=riπ\+minb∈BiPi\(b\)⊤x\.\\mathcal\{U\}\_\{i\}^\{\\pi\}=\\left\\\{\\sum\_\{a\}\\pi\(a\\mid i\)p\_\{a\}:p\_\{a\}\\in\\mathcal\{U\}\_\{ia\}\\right\\\},\\qquad\(T^\{\\pi\}x\)\_\{i\}=r\_\{i\}^\{\\pi\}\+\\min\_\{b\\in B\_\{i\}\}P\_\{i\}\(b\)^\{\\top\}x\.\(17\)The finite productBiB\_\{i\}is compact, andb↦Pi\(b\)b\\mapsto P\_\{i\}\(b\)is continuous\. The rewardriπr\_\{i\}^\{\\pi\}is constant inbb\. Rectangularity gives the equality because each positive\-weight summand can be minimized independently\. A deterministic stationary tuple policy specifies a full selector in the original row sets, including arbitrary feasible rows at zero\-weight actions\.
This reduction also respects history\-dependent randomization\. Given a pre\-action historyHHand nature’s conditional row lawsκH,a\\kappa\_\{H,a\}, sample a tuple from⨂aκH,a\\bigotimes\_\{a\}\\kappa\_\{H,a\}, draw the current action fromπ\(⋅∣St\)\\pi\(\\cdot\\mid S\_\{t\}\), and use the corresponding component\. Conditional onHHand the tupleb=\(pa\)ab=\(p\_\{a\}\)\_\{a\}, the next\-state law isp¯\(b\)=∑aπ\(a∣St\)pa\\bar\{p\}\(b\)=\\sum\_\{a\}\\pi\(a\\mid S\_\{t\}\)p\_\{a\}\. The original action history can be retained as auxiliary randomization with its correct conditional law: after observing a next statejjwithp¯j\(b\)\>0\\bar\{p\}\_\{j\}\(b\)\>0, the conditional probability of its action labelaaisπ\(a∣St\)pa,j/p¯j\(b\)\\pi\(a\\mid S\_\{t\}\)p\_\{a,j\}/\\bar\{p\}\_\{j\}\(b\)\. Labels on zero\-probability events can be chosen arbitrarily\. Marginalizing these auxiliary labels conditional on the tuple and state history gives an admissible history\-dependent randomized policy in the compact\-action MDP with transition lawp¯\(b\)\\bar\{p\}\(b\)\. Conversely, a tuple is implemented by using its component after the sampled action is observed\. The state\-process law is preserved\. For the pre\-action filtrationℱt\\mathscr\{F\}\_\{t\}, stationarity of the controller gives𝔼\[rStAt−rStπ∣ℱt\]=0\\mathbb\{E\}\[r\_\{S\_\{t\}A\_\{t\}\}\-r^\{\\pi\}\_\{S\_\{t\}\}\\mid\\mathscr\{F\}\_\{t\}\]=0and\|rStAt−rStπ\|≤2R\|r\_\{S\_\{t\}A\_\{t\}\}\-r^\{\\pi\}\_\{S\_\{t\}\}\|\\leq 2R\. The martingale strong law therefore makes its sample average converge to zero almost surely\. This justifies applying the one\-player results to the tuple model\.
### C\.2Finite\-chain facts used by the Bellman arguments
For a finite stochastic matrixPPand reward vectorcc, define
P∞=limN→∞1N∑t=0N−1Pt,ZP=\(I−P\+P∞\)−1,η=P∞c,w=ZP\(c−η\)\.P^\{\\infty\}=\\lim\_\{N\\to\\infty\}\\frac\{1\}\{N\}\\sum\_\{t=0\}^\{N\-1\}P^\{t\},\\quad Z\_\{P\}=\(I\-P\+P^\{\\infty\}\)^\{\-1\},\\quad\\eta=P^\{\\infty\}c,\\quad w=Z\_\{P\}\(c\-\\eta\)\.\(18\)
###### Lemma 2\.
The projectorP∞P^\{\\infty\}is stochastic,PP∞=P∞P=\(P∞\)2=P∞PP^\{\\infty\}=P^\{\\infty\}P=\(P^\{\\infty\}\)^\{2\}=P^\{\\infty\}, andZPZ\_\{P\}exists\. The canonical bias is the unique solution of
\(I−P\)w=c−P∞c,P∞w=0\.\(I\-P\)w=c\-P^\{\\infty\}c,\\qquad P^\{\\infty\}w=0\.\(19\)Ifd≥0d\\geq 0andP∞d=0P^\{\\infty\}d=0, thenddvanishes on recurrent states and
ZPd=∑t=0∞Ptd≥0\.Z\_\{P\}d=\\sum\_\{t=0\}^\{\\infty\}P^\{t\}d\\geq 0\.\(20\)
###### Proof\.
The projection and fundamental\-matrix identities are the finite\-chain specialization of\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Section 2, equations \(2\.2\)\-\(2\.9\)\], with unit holding times\. They apply without irreducibility or aperiodicity\. For the last assertion, the stationary distribution of each recurrent class is strictly positive on that class\. Its mean of the nonnegative vectorddis zero, soddis zero there\. IfQQis the transient block anddtrd\_\{\\rm tr\}is the restriction ofddto that block, thenρ\(Q\)<1\\rho\(Q\)<1andv=∑t≥0Ptdv=\\sum\_\{t\\geq 0\}P^\{t\}dequals\(I−Q\)−1dtr\(I\-Q\)^\{\-1\}d\_\{\\rm tr\}on the transient states and zero elsewhere\. Hencev≥0v\\geq 0,\(I−P\)v=d\(I\-P\)v=d, andP∞v=0P^\{\\infty\}v=0\. Uniqueness in equation[19](https://arxiv.org/html/2609.28792#A3.E19)givesv=ZPdv=Z\_\{P\}d\. ∎
###### Lemma 3\.
For each fixed finite chain,ϵ\(I−\(1−ϵ\)P\)−1c→P∞c\{\\epsilon\}\(I\-\(1\-\{\\epsilon\}\)P\)^\{\-1\}c\\to P^\{\\infty\}c\. Under a stationary pair\(π,q\)\(\\pi,q\),XNX\_\{N\}converges almost surely to the invariant mean reward of the recurrent class eventually entered\. Its expected limit isηiπ,q\\eta\_\{i\}^\{\\pi,q\}; all four payoffs in equation[1](https://arxiv.org/html/2609.28792#S2.E1)equal this number\.
###### Proof\.
The finite\-chain Cesàro limit gives its Abel limit\. The recurrent\-class ergodic theorem identifies the almost\-sure average ofrStπr^\{\\pi\}\_\{S\_\{t\}\}\. For sampled actions, the differencesrStAt−rStπr\_\{S\_\{t\}A\_\{t\}\}\-r^\{\\pi\}\_\{S\_\{t\}\}are bounded martingale differences, so their averages converge to zero almost surely\. Finally\|XN\|≤R\|X\_\{N\}\|\\leq Rpermits bounded convergence\. These are finite\-chain statements and allow periodic recurrent classes; see\[[48](https://arxiv.org/html/2609.28792#bib.bib32), Chapters 8\-9\]\. ∎
### C\.3The pathwise one\-player input
For bounded rewards, Fatou’s inequalities give
Ii−≤Ji−≤Ji\+≤Ii\+\.I\_\{i\}^\{\-\}\\leq J\_\{i\}^\{\-\}\\leq J\_\{i\}^\{\+\}\\leq I\_\{i\}^\{\+\}\.\(21\)We use the compact one\-player payoff result recorded in\[[19](https://arxiv.org/html/2609.28792#bib.bib11)\]Lemma 3\.3, Appendix E, and the proof of Corollary 3\.7 in Appendix G\. In the two applications needed here it reads
infτ∈𝒬HIi−\(π,τ\)\\displaystyle\\inf\_\{\\tau\\in\\mathcal\{Q\}\_\{H\}\}I\_\{i\}^\{\-\}\(\\pi,\\tau\)=infq∈𝒬Sηiπ,q\\displaystyle=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\eta\_\{i\}^\{\\pi,q\}\(π∈ΠS\),\\displaystyle\(\\pi\\in\\Pi\_\{S\}\),\(22\)supσ∈ΠHIi\+\(σ,q\)\\displaystyle\\sup\_\{\\sigma\\in\\Pi\_\{H\}\}I\_\{i\}^\{\+\}\(\\sigma,q\)=maxπ∈ΠDηiπ,q=:di\(q\)\\displaystyle=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta\_\{i\}^\{\\pi,q\}=:d\_\{i\}\(q\)\(q∈𝒬S\)\.\\displaystyle\(q\\in\\mathcal\{Q\}\_\{S\}\)\.Apply the cited one\-player theorem with initial laweie\_\{i\}\. In the first line, nature controls the compact\-action tuple MDP in equation[17](https://arxiv.org/html/2609.28792#A3.E17)\. Its hypotheses are finite state space, compact actions, and continuous rewards and transitions\. The expected\-limit\-inferior version is explicitly covered in the cited Appendix G\. The reward martingale argument above preserves this pathwise criterion for randomizedπ\\pi\. In the second line, fixingqqleaves a finite nominal MDP\. Sign reversal changes minimizing expected limit inferior into maximizing expected limit superior, and finiteness attains the stationary maximum\. Thus equation[22](https://arxiv.org/html/2609.28792#A3.E22)supplies the two pathwise endpoints needed below, independently of convergence of expected finite\-horizon averages\.
## Appendix DProof of Theorem[1](https://arxiv.org/html/2609.28792#Thmtheorem1): policy evaluation
###### Lemma 4\.
For everyπ∈ΠS\\pi\\in\\Pi\_\{S\},
ϵVϵπ⟶gπ,giπ=infq∈𝒬Sηiπ,q\.\{\\epsilon\}V\_\{\\epsilon\}^\{\\pi\}\\longrightarrow g^\{\\pi\},\\qquad g\_\{i\}^\{\\pi\}=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\eta\_\{i\}^\{\\pi,q\}\.\(23\)For everyν\>0\\nu\>0, one full selector satisfies
gπ≤ηπ,qπ,ν≤gπ\+ν𝟏\.g^\{\\pi\}\\leq\\eta^\{\\pi,q^\{\\pi,\\nu\}\}\\leq g^\{\\pi\}\+\\nu\\mathbf\{1\}\.\(24\)
###### Proof\.
\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Lemma 4\.7\]applies to the fixed stationary randomized policy and compact row sets\. Applying it with initial laweie\_\{i\}gives equation[23](https://arxiv.org/html/2609.28792#A4.E23)coordinatewise, hence in supremum norm becauseSSis finite\. The tuple reduction ensures that nature’s selectors belong to the original row sets\.
For one selector that works from all states, apply\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorem 4\.3\]to the tuple\-action minimizing MDP, the uniform initial lawμi=1/n\\mu\_\{i\}=1/n, and scalar toleranceν/n\\nu/n\. Its compactness and continuity assumptions were checked in Appendix[C](https://arxiv.org/html/2609.28792#A3)\. Discounted stationary optimality givesinfqμ⊤Vϵπ,q=μ⊤Vϵπ\\inf\_\{q\}\\mu^\{\\top\}V\_\{\\epsilon\}^\{\\pi,q\}=\\mu^\{\\top\}V\_\{\\epsilon\}^\{\\pi\}, so the theorem provides oneqπ,νq^\{\\pi,\\nu\}such that, for all sufficiently smallϵ\{\\epsilon\},
0≤ϵμ⊤\(Vϵπ,qπ,ν−Vϵπ\)≤ν/n\.0\\leq\{\\epsilon\}\\mu^\{\\top\}\(V\_\{\\epsilon\}^\{\\pi,q^\{\\pi,\\nu\}\}\-V\_\{\\epsilon\}^\{\\pi\}\)\\leq\\nu/n\.Each discounted coordinate gap is nonnegative\. Sinceμi=1/n\\mu\_\{i\}=1/n, each coordinate is at mostnntimes theirμ\\mu\-weighted mean\. Therefore
0≤ϵ\(Vϵπ,qπ,ν−Vϵπ\)≤ν𝟏\.0\\leq\{\\epsilon\}\(V\_\{\\epsilon\}^\{\\pi,q^\{\\pi,\\nu\}\}\-V\_\{\\epsilon\}^\{\\pi\}\)\\leq\\nu\\mathbf\{1\}\.\(25\)Keep the selector fixed and pass to the established discounted and finite\-chain Abel limits\. This yields equation[24](https://arxiv.org/html/2609.28792#A4.E24)\. ∎
Tauberian applicability under compactness\.For eitherF=TF=TorF=TπF=T^\{\\pi\}andλ,μ∈\(0,1\]\\lambda,\\mu\\in\(0,1\],
‖λF\(x/λ\)−μF\(x/μ\)‖∞≤R\|λ−μ\|\.\\\|\\lambda F\(x/\\lambda\)\-\\mu F\(x/\\mu\)\\\|\_\{\\infty\}\\leq R\|\\lambda\-\\mu\|\.\(26\)Indeed,λTi\(x/λ\)=maxa\{λria\+minpp⊤x\}\\lambda T\_\{i\}\(x/\\lambda\)=\\max\_\{a\}\\\{\\lambda r\_\{ia\}\+\\min\_\{p\}p^\{\\top\}x\\\}; changingλ\\lambdachanges each expression by at mostR\|λ−μ\|R\|\\lambda\-\\mu\|\. ForTπT^\{\\pi\}the difference is exactly\(λ−μ\)rπ\(\\lambda\-\\mu\)r^\{\\pi\}\. Together with nonexpansiveness, this verifies Assumption 1 of\[[85](https://arxiv.org/html/2609.28792#bib.bib67)\]on the Banach space\(ℝS,∥⋅∥∞\)\(\\mathbb\{R\}^\{S\},\\\|\\cdot\\\|\_\{\\infty\}\)\. IfR=0R=0, any positive constant also satisfies that assumption\. The normalized fixed point in Theorem 1\.2 of that reference is
vϵ=ϵF\(\(1−ϵ\)vϵ/ϵ\),v\_\{\\epsilon\}=\{\\epsilon\}F\\bigl\(\(1\-\{\\epsilon\}\)v\_\{\\epsilon\}/\{\\epsilon\}\\bigr\),namelyϵVϵ\{\\epsilon\}V\_\{\\epsilon\}orϵVϵπ\{\\epsilon\}V\_\{\\epsilon\}^\{\\pi\}\. Whenever this normalized discounted vector has a supremum\-norm limit, the theorem gives convergence ofFN0/NF^\{N\}0/Nto the same vector\. No definability assumption enters equation[26](https://arxiv.org/html/2609.28792#A4.E26)\.
###### Propsition 3\.
For eachπ∈ΠS\\pi\\in\\Pi\_\{S\},\(Tπ\)N0/N→gπ\(T^\{\\pi\}\)^\{N\}0/N\\to g^\{\\pi\}\. Consequently, for everyδ\>0\\delta\>0, all states and all sufficiently largeNNsatisfyJN\(i,π,τ\)≥giπ−δJ\_\{N\}\(i;\\pi,\\tau\)\\geq g\_\{i\}^\{\\pi\}\-\\deltafor everyτ∈𝒬H\\tau\\in\\mathcal\{Q\}\_\{H\}\. The stationary and history\-dependent infima of all four average payoffs equalgiπg\_\{i\}^\{\\pi\}\.
###### Proof\.
The first conclusion follows from equation[23](https://arxiv.org/html/2609.28792#A4.E23)and the Tauberian application\. Finite\-horizon dynamic programming gives
JN\(i,π,τ\)≥\[\(Tπ\)N0\]iN\(τ∈𝒬H\),J\_\{N\}\(i;\\pi,\\tau\)\\geq\\frac\{\[\(T^\{\\pi\}\)^\{N\}0\]\_\{i\}\}\{N\}\\quad\(\\tau\\in\\mathcal\{Q\}\_\{H\}\),which proves the uniform lower bound\. For the four payoffs, use equation[22](https://arxiv.org/html/2609.28792#A3.E22), equation[21](https://arxiv.org/html/2609.28792#A3.E21), and stationary\-pair convergence to obtain
giπ≤infτΨi\(π,τ\)≤infqΨi\(π,q\)≤ηiπ,qπ,ν≤giπ\+ν\.g\_\{i\}^\{\\pi\}\\leq\\inf\_\{\\tau\}\\Psi\_\{i\}\(\\pi,\\tau\)\\leq\\inf\_\{q\}\\Psi\_\{i\}\(\\pi,q\)\\leq\\eta\_\{i\}^\{\\pi,q^\{\\pi,\\nu\}\}\\leq g\_\{i\}^\{\\pi\}\+\\nu\.Letν↓0\\nu\\downarrow 0\. This completes Theorem[1](https://arxiv.org/html/2609.28792#Thmtheorem1)\. ∎
## Appendix EProof of Theorem[2](https://arxiv.org/html/2609.28792#Thmtheorem2): uniform control
###### Theorem 7\.
The limits and common controller in equation[4](https://arxiv.org/html/2609.28792#S3.E4)exist, and they satisfy the uniform bounds equation[5](https://arxiv.org/html/2609.28792#S3.E5)\.
###### Proof\.
*The value and one common controller\.*\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Lemma 4\.8\], applied to eachp0=eip\_\{0\}=e\_\{i\}, givesϵVϵ→g⋆=maxπ∈ΠDgπ\{\\epsilon\}V\_\{\\epsilon\}\\to g^\{\\star\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}g^\{\\pi\}in supremum norm\. The fixed\-policy identity equation[23](https://arxiv.org/html/2609.28792#A4.E23)identifies its scalar limit ateie\_\{i\}withmaxπ∈ΠDgiπ\\max\_\{\\pi\\in\\Pi\_\{D\}\}g\_\{i\}^\{\\pi\}\. The normalized discounted value is the same Bellman vector in every application\. To select one all\-state optimizer, choose deterministic stationary discounted\-optimal policies alongϵk↓0\{\\epsilon\}\_\{k\}\\downarrow 0\. SinceΠD\\Pi\_\{D\}is finite, a policyπ⋆\\pi^\{\\star\}occurs on an infinite subsequence\. Reindexing that subsequence,
gπ⋆=limkϵkVϵkπ⋆=limkϵkVϵk=g⋆\.g^\{\\pi^\{\\star\}\}=\\lim\_\{k\}\{\\epsilon\}\_\{k\}V\_\{\{\\epsilon\}\_\{k\}\}^\{\\pi^\{\\star\}\}=\\lim\_\{k\}\{\\epsilon\}\_\{k\}V\_\{\{\\epsilon\}\_\{k\}\}=g^\{\\star\}\.The Tauberian application equation[26](https://arxiv.org/html/2609.28792#A4.E26)givesTN0/N→g⋆T^\{N\}0/N\\to g^\{\\star\}\. Applying Proposition[3](https://arxiv.org/html/2609.28792#Thmproposition3)toπ⋆\\pi^\{\\star\}proves the lower inequality in equation[5](https://arxiv.org/html/2609.28792#S3.E5)\.
*One common selector for nature\.*It remains to construct a stationary upper strategy that works from all states\. Forq∈𝒬Sq\\in\\mathcal\{Q\}\_\{S\}, setdi\(q\)=maxπ∈ΠDηiπ,qd\_\{i\}\(q\)=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta\_\{i\}^\{\\pi,q\}\. Fixingqqleaves a finite nominal MDP\. Its discounted values are the coordinatewise maximum of finitely many policy values; their normalized limit isd\(q\)d\(q\)by the finite\-chain Abel limit\. A constant\-policy subsequence of its discounted optimizers therefore yields one policy attainingd\(q\)d\(q\)at every state\. In particular,maxπμ⊤ηπ,q=μ⊤d\(q\)\\max\_\{\\pi\}\\mu^\{\\top\}\\eta^\{\\pi,q\}=\\mu^\{\\top\}d\(q\)for everyμ\\mu\. Also
d\(q\)≥ηπ⋆,q≥gπ⋆=g⋆\.d\(q\)\\geq\\eta^\{\\pi^\{\\star\},q\}\\geq g^\{\\pi^\{\\star\}\}=g^\{\\star\}\.
Take the uniform initial lawμi=1/n\\mu\_\{i\}=1/n\. The common fixed\-policy approximations in equation[24](https://arxiv.org/html/2609.28792#A4.E24)implyinfqμ⊤ηπ,q=μ⊤gπ\\inf\_\{q\}\\mu^\{\\top\}\\eta^\{\\pi,q\}=\\mu^\{\\top\}g^\{\\pi\}\. The common controller givesmaxπ∈ΠDμ⊤gπ=μ⊤g⋆\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\mu^\{\\top\}g^\{\\pi\}=\\mu^\{\\top\}g^\{\\star\}, and the common nominal optimizer above givesmaxπ∈ΠDμ⊤ηπ,q=μ⊤d\(q\)\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}=\\mu^\{\\top\}d\(q\)\. Thus\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorem 3\.5, equation \(3\.6\)\], with the nonconvex applicability check above, gives
μ⊤g⋆=maxinfq∈𝒬Sπ∈ΠDμ⊤ηπ,q=infq∈𝒬Sμ⊤d\(q\)\.\\mu^\{\\top\}g^\{\\star\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\mu^\{\\top\}\\eta^\{\\pi,q\}=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\mu^\{\\top\}d\(q\)\.\(27\)For anyα\>0\\alpha\>0, choose an approximate minimizer withμ⊤\(d\(qα\)−g⋆\)≤α/n\\mu^\{\\top\}\(d\(q\_\{\\alpha\}\)\-g^\{\\star\}\)\\leq\\alpha/n\. Every coordinate gap is nonnegative, so the full\-support averaging argument yields
g⋆≤d\(qα\)≤g⋆\+α𝟏\.g^\{\\star\}\\leq d\(q\_\{\\alpha\}\)\\leq g^\{\\star\}\+\\alpha\\mathbf\{1\}\.\(28\)
*Uniform finite\-horizon guarantees\.*Let\(Hqx\)i=maxa\{ria\+qia⊤x\}\(H\_\{q\}x\)\_\{i\}=\\max\_\{a\}\\\{r\_\{ia\}\+q\_\{ia\}^\{\\top\}x\\\}\. This operator also satisfies equation[26](https://arxiv.org/html/2609.28792#A4.E26); its discounted limit just identified therefore givesHqN0/N→d\(q\)H\_\{q\}^\{N\}0/N\\to d\(q\)\. Finite\-horizon dynamic programming yields
JN\(i,σ,q\)≤\[HqN0\]i/N\(σ∈ΠH\)\.J\_\{N\}\(i;\\sigma,q\)\\leq\[H\_\{q\}^\{N\}0\]\_\{i\}/N\\quad\(\\sigma\\in\\Pi\_\{H\}\)\.Useqδ/2q\_\{\\delta/2\}and takeNNlarge enough that‖Hqδ/2N0/N−d\(qδ/2\)‖∞≤δ/2\\\|H\_\{q\_\{\\delta/2\}\}^\{N\}0/N\-d\(q\_\{\\delta/2\}\)\\\|\_\{\\infty\}\\leq\\delta/2\. This proves the upper bound\. Taking the larger of the lower and upper horizon thresholds makes both guarantees simultaneous\. ∎
The constant\-policy subsequence produces one optimal controller, and the full\-support initial law produces one approximate minimizing selector\. These are the two steps that strengthen pointwise value identities to simultaneous all\-state guarantees\.
## Appendix FPayoff conventions and strategy\-class duality
###### Theorem 8\.
LetΠD⊆𝒞⊆ΠH\\Pi\_\{D\}\\subseteq\\mathcal\{C\}\\subseteq\\Pi\_\{H\}and𝒬S⊆𝒩⊆𝒬H\\mathcal\{Q\}\_\{S\}\\subseteq\\mathcal\{N\}\\subseteq\\mathcal\{Q\}\_\{H\}\. For everyΨ∈\{I−,J−,J\+,I\+\}\\Psi\\in\\\{I^\{\-\},J^\{\-\},J^\{\+\},I^\{\+\}\\\}andi∈Si\\in S,
supσ∈𝒞infτ∈𝒩Ψi\(σ,τ\)=infτ∈𝒩supσ∈𝒞Ψi\(σ,τ\)=gi⋆\.\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\Psi\_\{i\}\(\\sigma,\\tau\)=\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\Psi\_\{i\}\(\\sigma,\\tau\)=g\_\{i\}^\{\\star\}\.\(29\)For an initial distributionμ\\mu, the value isμ⊤g⋆\\mu^\{\\top\}g^\{\\star\}\. The controllerπ⋆\\pi^\{\\star\}from Theorem[2](https://arxiv.org/html/2609.28792#Thmtheorem2)is optimal for every stated criterion, state, and initial distribution\.
The proof makes the simultaneous statewise guarantees explicit in the setting of\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorems 3\.5\-3\.6 and Corollary 3\.7\]\. It combines the one\-player pathwise bounds with the common stationary strategies constructed above\.
###### Proof of Theorem[8](https://arxiv.org/html/2609.28792#Thmtheorem8)\.
Fix a toleranceδ\>0\\delta\>0and choose a common selector from equation[28](https://arxiv.org/html/2609.28792#A5.E28)\. The one\-player endpoints equation[22](https://arxiv.org/html/2609.28792#A3.E22)give
Ii−\(π⋆,τ\)≥gi⋆\(τ∈𝒬H\),Ii\+\(σ,qδ\)≤di\(qδ\)≤gi⋆\+δ\(σ∈ΠH\)\.I\_\{i\}^\{\-\}\(\\pi^\{\\star\},\\tau\)\\geq g\_\{i\}^\{\\star\}\\quad\(\\tau\\in\\mathcal\{Q\}\_\{H\}\),\\qquad I\_\{i\}^\{\+\}\(\\sigma,q\_\{\\delta\}\)\\leq d\_\{i\}\(q\_\{\\delta\}\)\\leq g\_\{i\}^\{\\star\}\+\\delta\\quad\(\\sigma\\in\\Pi\_\{H\}\)\.The payoff ordering equation[21](https://arxiv.org/html/2609.28792#A3.E21)makes these lower and upper bounds valid for eachΨ∈\{I−,J−,J\+,I\+\}\\Psi\\in\\\{I^\{\-\},J^\{\-\},J^\{\+\},I^\{\+\}\\\}\. Sinceπ⋆∈𝒞\\pi^\{\\star\}\\in\\mathcal\{C\}andqδ∈𝒩q\_\{\\delta\}\\in\\mathcal\{N\}, weak duality yields
gi⋆≤supσ∈𝒞infτ∈𝒩Ψi≤infτ∈𝒩supσ∈𝒞Ψi≤gi⋆\+δ\.g\_\{i\}^\{\\star\}\\leq\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\Psi\_\{i\}\\leq\\inf\_\{\\tau\\in\\mathcal\{N\}\}\\sup\_\{\\sigma\\in\\mathcal\{C\}\}\\Psi\_\{i\}\\leq g\_\{i\}^\{\\star\}\+\\delta\.Letδ↓0\\delta\\downarrow 0\. The lower guarantee proves that the sameπ⋆\\pi^\{\\star\}attains every outer supremum in the lower value\.
For an initial lawμ\\mu, condition the pathwise endpoints onS0=iS\_\{0\}=iand sum their bounds with weightsμi\\mu\_\{i\}\. This gives lower and upper guaranteesμ⊤g⋆\\mu^\{\\top\}g^\{\\star\}andμ⊤g⋆\+δ\\mu^\{\\top\}g^\{\\star\}\+\\delta\. The payoff ordering holds under this initial law as well, so the same squeeze proves the claim for all four criteria\. Only the pathwise endpoint expectations are decomposed in this argument\. No equality between a limit inferior and a weighted sum of limit inferiors is needed\. ∎
## Appendix GDiscounted characterizations of stationary policies
###### Theorem 9\.
For everyπ∈ΠS\\pi\\in\\Pi\_\{S\},
limϵ↓0ϵ‖Vϵ−Vϵπ‖∞=‖g⋆−gπ‖∞\.\\lim\_\{\{\\epsilon\}\\downarrow 0\}\{\\epsilon\}\\\|V\_\{\\epsilon\}\-V\_\{\\epsilon\}^\{\\pi\}\\\|\_\{\\infty\}=\\\|g^\{\\star\}\-g^\{\\pi\}\\\|\_\{\\infty\}\.\(30\)Consequently,π\\piis average optimal from all states if and only if this limit is zero\. Moreover, there existsϵ0\>0\{\\epsilon\}\_\{0\}\>0such that
π∈ΠD,0<ϵ<ϵ0,Vϵπ=Vϵ⟹gπ=g⋆\.\\pi\\in\\Pi\_\{D\},\\qquad 0<\{\\epsilon\}<\{\\epsilon\}\_\{0\},\\qquad V\_\{\\epsilon\}^\{\\pi\}=V\_\{\\epsilon\}\\quad\\Longrightarrow\\quad g^\{\\pi\}=g^\{\\star\}\.\(31\)
The gap identity is a direct consequence of the two discounted limits in Theorems[1](https://arxiv.org/html/2609.28792#Thmtheorem1)\-[2](https://arxiv.org/html/2609.28792#Thmtheorem2)\. It records the exact limiting all\-state loss, including stationary randomized policies\. The corresponding deterministic\-policy connections are\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorems 4\.6 and 4\.10\]\. The threshold in equation[31](https://arxiv.org/html/2609.28792#A7.E31)is existential and supplies no computable stopping rule\.
###### Proof of Theorem[9](https://arxiv.org/html/2609.28792#Thmtheorem9)\.
The two discounted limits imply
ϵ\(Vϵ−Vϵπ\)⟶g⋆−gπin∥⋅∥∞\.\{\\epsilon\}\(V\_\{\\epsilon\}\-V\_\{\\epsilon\}^\{\\pi\}\)\\longrightarrow g^\{\\star\}\-g^\{\\pi\}\\quad\\text\{in \}\\\|\\cdot\\\|\_\{\\infty\}\.Continuity of the norm gives equation[30](https://arxiv.org/html/2609.28792#A7.E30)\. Theorems[1](https://arxiv.org/html/2609.28792#Thmtheorem1)and[8](https://arxiv.org/html/2609.28792#Thmtheorem8)identifygπ=g⋆g^\{\\pi\}=g^\{\\star\}with all\-state average optimality, including randomized stationaryπ\\pi\.
For equation[31](https://arxiv.org/html/2609.28792#A7.E31), apply\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorem 4\.10\]with a full\-support initial lawμ\\mu\. A vector\-discount\-optimal policy is discount optimal for this law; their theorem makes it average optimal for that law wheneverϵ\{\\epsilon\}is small enough\. By equation[24](https://arxiv.org/html/2609.28792#A4.E24), its scalar robust average reward isμ⊤gπ\\mu^\{\\top\}g^\{\\pi\}, while Theorem[8](https://arxiv.org/html/2609.28792#Thmtheorem8)identifies the scalar optimal value withμ⊤g⋆\\mu^\{\\top\}g^\{\\star\}\. Henceμ⊤\(g⋆−gπ\)=0\\mu^\{\\top\}\(g^\{\\star\}\-g^\{\\pi\}\)=0\. Discounted domination impliesg⋆−gπ≥0g^\{\\star\}\-g^\{\\pi\}\\geq 0, and full support impliesgπ=g⋆g^\{\\pi\}=g^\{\\star\}\. The cited theorem supplies one threshold for all deterministic stationary discounted optimizers\. ∎
## Appendix HExact nature attainment and stationary saddles
Forq∈𝒬Sq\\in\\mathcal\{Q\}\_\{S\}, letdi\(q\)=maxπ∈ΠDηiπ,qd\_\{i\}\(q\)=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta\_\{i\}^\{\\pi,q\}denote the optimal gain of the nominal MDP obtained by fixing the full selectorqq\.
###### Theorem 10\.
The stationary performance vectors satisfy
gi⋆\\displaystyle g\_\{i\}^\{\\star\}=maxπ∈ΠDinfq∈𝒬Sηiπ,q=infq∈𝒬Sdi\(q\),i∈S,\\displaystyle=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}\\eta\_\{i\}^\{\\pi,q\}=\\inf\_\{q\\in\\mathcal\{Q\}\_\{S\}\}d\_\{i\}\(q\),\\qquad i\\in S,\(32\)d\(q\)\\displaystyle d\(q\)≥g⋆,q∈𝒬S\.\\displaystyle\\geq g^\{\\star\},\\qquad q\\in\\mathcal\{Q\}\_\{S\}\.For everyδ\>0\\delta\>0, oneq∈𝒬Sq\\in\\mathcal\{Q\}\_\{S\}satisfiesd\(q\)≤g⋆\+δ𝟏d\(q\)\\leq g^\{\\star\}\+\\delta\\mathbf\{1\}\.
A stationary selectorqqis exactly optimal for nature against every history\-dependent controller, from every state and for every payoff in equation[1](https://arxiv.org/html/2609.28792#S2.E1), if and only if
d\(q\)=g⋆\.d\(q\)=g^\{\\star\}\.\(33\)
Forπ¯∈ΠS\\bar\{\\pi\}\\in\\Pi\_\{S\},q¯∈𝒬S\\bar\{q\}\\in\\mathcal\{Q\}\_\{S\}, andu∈ℝSu\\in\\mathbb\{R\}^\{S\}, the stationary gain inequalities
ηπ¯,q≥u∀q∈𝒬S,ηπ,q¯≤u∀π∈ΠD\\eta^\{\\bar\{\\pi\},q\}\\geq u\\quad\\forall q\\in\\mathcal\{Q\}\_\{S\},\\qquad\\eta^\{\\pi,\\bar\{q\}\}\\leq u\\quad\\forall\\pi\\in\\Pi\_\{D\}\(34\)hold if and only ifu=g⋆u=g^\{\\star\}and\(π¯,q¯\)\(\\bar\{\\pi\},\\bar\{q\}\)is an all\-state stationary saddle against history\-dependent opponents for all four payoffs\.
The dual representation builds on\[[19](https://arxiv.org/html/2609.28792#bib.bib11), Theorem 3\.5\]\. The formulation throughd\(q\)d\(q\)isolates simultaneous exact attainment, while equation[34](https://arxiv.org/html/2609.28792#A8.E34)expresses the full saddle property using only stationary\-chain gains\.
###### Lemma 5\.
For eachq∈𝒬Sq\\in\\mathcal\{Q\}\_\{S\}, the nominal MDP with operatorHqH\_\{q\}has all\-state valued\(q\)d\(q\), attained by one policy inΠD\\Pi\_\{D\}\. Moreover, for every payoffΨ\\Psiand every state,
supσ∈ΠHΨi\(σ,q\)=di\(q\),HqN0/N⟶d\(q\)\.\\sup\_\{\\sigma\\in\\Pi\_\{H\}\}\\Psi\_\{i\}\(\\sigma,q\)=d\_\{i\}\(q\),\\qquad H\_\{q\}^\{N\}0/N\\longrightarrow d\(q\)\.For everyδ\>0\\delta\>0, all sufficiently largeNNsatisfyJN\(i,σ,q\)≤di\(q\)\+δJ\_\{N\}\(i;\\sigma,q\)\\leq d\_\{i\}\(q\)\+\\deltafor alliiandσ∈ΠH\\sigma\\in\\Pi\_\{H\}\.
###### Proof\.
The all\-state optimizer and finite\-horizon limit were established in the proof of Theorem[7](https://arxiv.org/html/2609.28792#Thmtheorem7)by the finite\-policy Abel limit and Tauberian argument\. The upper pathwise endpoint is equation[22](https://arxiv.org/html/2609.28792#A3.E22); the payoff ordering and the common stationary optimizer identify all four values\. Finite\-horizon dynamic programming gives the last assertion\. ∎
###### Proof of Theorem[10](https://arxiv.org/html/2609.28792#Thmtheorem10)\.
The primal representation follows from Theorem[1](https://arxiv.org/html/2609.28792#Thmtheorem1)andgπ⋆=g⋆g^\{\\pi^\{\\star\}\}=g^\{\\star\}\. The inequalityd\(q\)≥g⋆d\(q\)\\geq g^\{\\star\}and the common approximation equation[28](https://arxiv.org/html/2609.28792#A5.E28)implyinfqdi\(q\)=gi⋆\\inf\_\{q\}d\_\{i\}\(q\)=g\_\{i\}^\{\\star\}, proving equation[32](https://arxiv.org/html/2609.28792#A8.E32)\.
Ifd\(q\)=g⋆d\(q\)=g^\{\\star\}, Lemma[5](https://arxiv.org/html/2609.28792#Thmlemma5)bounds every controller’s payoff byg⋆g^\{\\star\}from every state\. Conversely, an upper guarantee for even one of the four payoff conventions bounds in particular the stationary gainsηπ,q\\eta^\{\\pi,q\}for allπ∈ΠD\\pi\\in\\Pi\_\{D\}\. Taking their maximum givesd\(q\)≤g⋆d\(q\)\\leq g^\{\\star\}; the reverse inequality always holds\. This proves equation[33](https://arxiv.org/html/2609.28792#A8.E33)for every payoff convention\.
For the gain\-only saddle test, the two inequalities imply
u≤infqηπ¯,q=gπ¯≤g⋆≤d\(q¯\)=maxπ∈ΠDηπ,q¯≤u\.u\\leq\\inf\_\{q\}\\eta^\{\\bar\{\\pi\},q\}=g^\{\\bar\{\\pi\}\}\\leq g^\{\\star\}\\leq d\(\\bar\{q\}\)=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta^\{\\pi,\\bar\{q\}\}\\leq u\.All inequalities are therefore equalities\. The fixed\-policy and fixed\-nature evaluations extend the two stationary guarantees to all history\-dependent opponents, for all four payoffs\. At\(π¯,q¯\)\(\\bar\{\\pi\},\\bar\{q\}\)the payoff isu=g⋆u=g^\{\\star\}, so this is an exact saddle\. Conversely, restricting any such saddle guarantees to stationary opponents and using Lemma[3](https://arxiv.org/html/2609.28792#Thmlemma3)gives equation[34](https://arxiv.org/html/2609.28792#A8.E34)\. ∎
A worst reply can leave a profitable deviation\.Consider statess,L,Hs,L,H, withL,HL,Habsorbing and rewards0,0,10,0,1, respectively\. Atss, actionaagoes toLLand actionbbhas row set\{eL,eH\}\\\{e\_\{L\},e\_\{H\}\\\}\. The two deterministic policies satisfygπa=gπb=g⋆=\(0,0,1\)g^\{\\pi\_\{a\}\}=g^\{\\pi\_\{b\}\}=g^\{\\star\}=\(0,0,1\)\. Choose the full selector withqsb=eHq\_\{sb\}=e\_\{H\}\. Thenηπa,q=\(0,0,1\)\\eta^\{\\pi\_\{a\},q\}=\(0,0,1\)butηπb,q=\(1,0,1\)\\eta^\{\\pi\_\{b\},q\}=\(1,0,1\)\. Thusqqis exactly worst against the optimal controllerπa\\pi\_\{a\}, yetd\(q\)=\(1,0,1\)≠g⋆d\(q\)=\(1,0,1\)\\neq g^\{\\star\}: its unused action permits a deviation\. The same example works with row setco\{eL,eH\}\\operatorname\{co\}\\\{e\_\{L\},e\_\{H\}\\\}\.
## Appendix IProof of Theorem[3](https://arxiv.org/html/2609.28792#Thmtheorem3): vector Bellman verification
The proof has two ingredients\. First, the large\-ttexpansion ofT\(tg\+h\)T\(tg\+h\)identifies the gain\-restricted bias operator\. Second, a compactness estimate controls the bias loss when nature leaves a gain\-minimizing face\. Combining this estimate with a finite budget for cumulative gain drift proves verification against arbitrary history\-dependent opponents\. Fixed\-policy specializations and two counterexamples follow\.
The leading gain and the finite bias require two successive optimizations\. Define the recession map by
\(T^g\)i=maxa∈A\(i\)minp∈𝒰iap⊤g\.\(\\widehat\{T\}g\)\_\{i\}=\\max\_\{a\\in A\(i\)\}\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g\.\(35\)For every vectorg∈ℝSg\\in\\mathbb\{R\}^\{S\}, define the row minima and their minimizing sets by
mia\(g\)=minp∈𝒰iap⊤g,Fia\(g\)=\{p∈𝒰ia:p⊤g=mia\(g\)\}\.m\_\{ia\}\(g\)=\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g,\\qquad F\_\{ia\}\(g\)=\\\{p\\in\\mathcal\{U\}\_\{ia\}:p^\{\\top\}g=m\_\{ia\}\(g\)\\\}\.Ifg=T^gg=\\widehat\{T\}g, additionally put
Ag\(i\)=\{a∈A\(i\):mia\(g\)=gi\},\(Lgh\)i=maxa∈Ag\(i\)minp∈Fia\(g\)\{ria\+p⊤h\}\.A\_\{g\}\(i\)=\\\{a\\in A\(i\):m\_\{ia\}\(g\)=g\_\{i\}\\\},\\qquad\(L\_\{g\}h\)\_\{i\}=\\max\_\{a\\in A\_\{g\}\(i\)\}\\min\_\{p\\in F\_\{ia\}\(g\)\}\\\{r\_\{ia\}\+p^\{\\top\}h\\\}\.\(36\)Compactness of each row set attains its linear minimum and makesFia\(g\)F\_\{ia\}\(g\)nonempty and compact\. Finiteness ofA\(i\)A\(i\)andgi=maxamia\(g\)g\_\{i\}=\\max\_\{a\}m\_\{ia\}\(g\)makeAg\(i\)A\_\{g\}\(i\)nonempty\. We calla∈Ag\(i\)a\\in A\_\{g\}\(i\)a gain\-active action and use “gain face” forFia\(g\)F\_\{ia\}\(g\)even when𝒰ia\\mathcal\{U\}\_\{ia\}is nonconvex\. Definingmia\(g\)m\_\{ia\}\(g\)andFia\(g\)F\_\{ia\}\(g\)for arbitraryggalso permits their use in fixed\-policy evaluation, whereggneed not solve the optimal\-control gain equation\.
A finite vector gain\-bias certificate is a pair\(g,h\)∈ℝS×ℝS\(g,h\)\\in\\mathbb\{R\}^\{S\}\\times\\mathbb\{R\}^\{S\}satisfying
g=T^g,g\+h=Lgh\.g=\\widehat\{T\}g,\\qquad g\+h=L\_\{g\}h\.\(37\)The first equation compares the leading gain\. The second compares the one\-step reward and continuation bias after both players’ choices have been restricted to the first equation’s optimizers\. In particular,Fia\(g\)F\_\{ia\}\(g\)is defined relative tomia\(g\)m\_\{ia\}\(g\)for every action, including inactive actions for whichmia\(g\)<gim\_\{ia\}\(g\)<g\_\{i\}\.
###### Lemma 6\.
Ifg=T^gg=\\widehat\{T\}g, then, for every finitehh,
T\(tg\+h\)−tg⟶Lghast⟶∞\.T\(tg\+h\)\-tg\\longrightarrow L\_\{g\}h\\quad\\text\{as \}t\\longrightarrow\\infty\.\(38\)The convergence is uniform whenhhranges over any fixed compact subset ofℝS\\mathbb\{R\}^\{S\}\. Consequently, equation[37](https://arxiv.org/html/2609.28792#A9.E37)is equivalent to
T\(tg\+h\)=\(t\+1\)g\+h\+o\(1\)\.T\(tg\+h\)=\(t\+1\)g\+h\+o\(1\)\.\(39\)If every ambiguity set is a polytope, the error in equation[38](https://arxiv.org/html/2609.28792#A9.E38)is zero for all sufficiently largett\.
###### Proof\.
The row limit\.Fix\(i,a\)\(i,a\)and abbreviatem=mia\(g\)m=m\_\{ia\}\(g\),F=Fia\(g\)F=F\_\{ia\}\(g\), andb=minp∈F\(ria\+p⊤h\)b=\\min\_\{p\\in F\}\(r\_\{ia\}\+p^\{\\top\}h\)\. Write
bt=minp∈𝒰ia\{t\(p⊤g−m\)\+ria\+p⊤h\}\.b\_\{t\}=\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}\\\{t\(p^\{\\top\}g\-m\)\+r\_\{ia\}\+p^\{\\top\}h\\\}\.The gain gap is nonnegative, and testing a bias\-minimizing row inFFgivesria−‖h‖∞≤bt≤br\_\{ia\}\-\\\|h\\\|\_\{\\infty\}\\leq b\_\{t\}\\leq b\. For any minimizerptp\_\{t\}, it follows that
0≤t\(pt⊤g−m\)=bt−ria−pt⊤h≤2‖h‖∞\.0\\leq t\(p\_\{t\}^\{\\top\}g\-m\)=b\_\{t\}\-r\_\{ia\}\-p\_\{t\}^\{\\top\}h\\leq 2\\\|h\\\|\_\{\\infty\}\.Thus every cluster point of minimizing rows ast→∞t\\to\\inftylies inFF\. To identify the value, take a sequence along whichbtb\_\{t\}tends to its limit inferior and a further subsequence withpt→p0∈Fp\_\{t\}\\to p\_\{0\}\\in F\. Discarding the nonnegative gain gap yields
lim inft→∞bt≥ria\+p0⊤h≥b\.\\liminf\_\{t\\to\\infty\}b\_\{t\}\\geq r\_\{ia\}\+p\_\{0\}^\{\\top\}h\\geq b\.Together withbt≤bb\_\{t\}\\leq b, this provesbt→bb\_\{t\}\\to b\.
The action maximum and compact uniformity\.Restoring row indices gives
\[T\(tg\+h\)−tg\]i=maxa∈A\(i\)\{t\(mia\(g\)−gi\)\+bia,t\(h\)\}\.\[T\(tg\+h\)\-tg\]\_\{i\}=\\max\_\{a\\in A\(i\)\}\\\{t\(m\_\{ia\}\(g\)\-g\_\{i\}\)\+b\_\{ia,t\}\(h\)\\\}\.For an active action the expression tends tobia\(h\)=minp∈Fia\(g\)\(ria\+p⊤h\)b\_\{ia\}\(h\)=\\min\_\{p\\in F\_\{ia\}\(g\)\}\(r\_\{ia\}\+p^\{\\top\}h\)\. For an inactive action,mia\(g\)−gi<0m\_\{ia\}\(g\)\-g\_\{i\}<0, so it tends to−∞\-\\infty\. There are finitely many actions and states\. Taking their maxima proves equation[38](https://arxiv.org/html/2609.28792#A9.E38)in the supremum norm\.
Both mapsh↦T\(tg\+h\)−tgh\\mapsto T\(tg\+h\)\-tgandh↦Lghh\\mapsto L\_\{g\}hare11\-Lipschitz in that norm, by the stochastic\-row estimate in Lemma[1](https://arxiv.org/html/2609.28792#Thmlemma1)\. Ifh1,…,hMh^\{1\},\\ldots,h^\{M\}form a finiteδ\\delta\-net of a compact setKK, then
suph∈K‖T\(tg\+h\)−tg−Lgh‖∞≤2δ\+maxj≤M‖T\(tg\+hj\)−tg−Lghj‖∞\.\\sup\_\{h\\in K\}\\\|T\(tg\+h\)\-tg\-L\_\{g\}h\\\|\_\{\\infty\}\\leq 2\\delta\+\\max\_\{j\\leq M\}\\\|T\(tg\+h^\{j\}\)\-tg\-L\_\{g\}h^\{j\}\\\|\_\{\\infty\}\.First lett→∞t\\to\\inftyfor this finite net and thenδ↓0\\delta\\downarrow 0\. This proves compact uniformity\.
The equivalence and polytopic exactness\.The two Bellman equations imply equation[39](https://arxiv.org/html/2609.28792#A9.E39)by the limit just proved\. Conversely, that asymptotic identity impliesT\(tg\+h\)/t→gT\(tg\+h\)/t\\to g, while
‖T\(tg\+h\)t−T^g‖∞≤R\+‖h‖∞t\.\\left\\\|\\frac\{T\(tg\+h\)\}\{t\}\-\\widehat\{T\}g\\right\\\|\_\{\\infty\}\\leq\\frac\{R\+\\\|h\\\|\_\{\\infty\}\}\{t\}\.The latter bound follows by deleting the uniformly bounded reward\-bias perturbation inside every row optimization\. Henceg=T^gg=\\widehat\{T\}g, and the tangent limit then givesLgh=g\+hL\_\{g\}h=g\+h\.
If each row set is a polytope, minimize over its finitely many vertices\. A vertexppoutsideFia\(g\)F\_\{ia\}\(g\)has positive gapdp=p⊤g−mia\(g\)d\_\{p\}=p^\{\\top\}g\-m\_\{ia\}\(g\)\. Its centered objective exceeds the face minimum as soon as
t\>\[bia\(h\)−ria−p⊤h\]\+dp\.t\>\\frac\{\[b\_\{ia\}\(h\)\-r\_\{ia\}\-p^\{\\top\}h\]\_\{\+\}\}\{d\_\{p\}\}\.Take a common threshold over the finitely many outside vertices\. Above it, every row minimum equals its face minimum exactly\. A further finite threshold excludes all inactive actions, since their gain gapsgi−mia\(g\)g\_\{i\}\-m\_\{ia\}\(g\)are strictly positive\. This proves eventual exactness for fixedg,hg,h\. ∎
### I\.1A compactness estimate for rows close to a gain face
###### Lemma 7\.
LetXXbe compact, letd,e:X→ℝd,e:X\\to\\mathbb\{R\}be continuous, and supposed≥0d\\geq 0onXXande≥0e\\geq 0on\{x:d\(x\)=0\}\\\{x:d\(x\)=0\\\}\. For everyη\>0\\eta\>0there isCη<∞C\_\{\\eta\}<\\inftysuch that
e\(x\)≥−η−Cηd\(x\)\(x∈X\)\.e\(x\)\\geq\-\\eta\-C\_\{\\eta\}d\(x\)\\qquad\(x\\in X\)\.\(40\)IfXXis a polytope andd,ed,eare affine, there isC0<∞C\_\{0\}<\\inftyfor which the same inequality holds withη=0\\eta=0\.
###### Proof\.
For fixedη\>0\\eta\>0, letBη=\{x∈X:e\(x\)≤−η\}B\_\{\\eta\}=\\\{x\\in X:e\(x\)\\leq\-\\eta\\\}andM=maxx∈X\[−e\(x\)\]\+M=\\max\_\{x\\in X\}\[\-e\(x\)\]\_\{\+\}, treating emptyXXseparately as vacuous\. IfBηB\_\{\\eta\}is empty, takeCη=0C\_\{\\eta\}=0\. OtherwiseBηB\_\{\\eta\}is compact and contains no zero ofdd, becausee≥0e\\geq 0whereverd=0d=0\. Thereforeδη=minx∈Bηd\(x\)\>0\\delta\_\{\\eta\}=\\min\_\{x\\in B\_\{\\eta\}\}d\(x\)\>0\. WithCη=M/δηC\_\{\\eta\}=M/\\delta\_\{\\eta\}, onBηB\_\{\\eta\}we havee≥−M≥−Cηde\\geq\-M\\geq\-C\_\{\\eta\}d, and outsideBηB\_\{\\eta\}we havee\>−η≥−η−Cηde\>\-\\eta\\geq\-\\eta\-C\_\{\\eta\}d\. This proves the bound\.
For a polytope with vertex set𝒱\\mathcal\{V\}, take
C0=maxv∈𝒱:d\(v\)\>0\[−e\(v\)\]\+d\(v\),C\_\{0\}=\\max\_\{v\\in\\mathcal\{V\}:\\,d\(v\)\>0\}\\frac\{\[\-e\(v\)\]\_\{\+\}\}\{d\(v\)\},with an empty maximum equal to zero\. This finite constant givese\(v\)\+C0d\(v\)≥0e\(v\)\+C\_\{0\}d\(v\)\\geq 0at vertices withd\(v\)\>0d\(v\)\>0, and the hypothesis gives the same inequality at vertices withd\(v\)=0d\(v\)=0\. Affineness extends it to every convex combination of the vertices\. ∎
The additiveη\\etacannot generally be removed: onX=\[0,1\]X=\[0,1\],d\(x\)=x2d\(x\)=x^\{2\}ande\(x\)=−xe\(x\)=\-xsatisfy the hypotheses, but no finiteC0C\_\{0\}can satisfy−x≥−C0x2\-x\\geq\-C\_\{0\}x^\{2\}for allx\>0x\>0\. For curved row sets this distinction leads to ano\(N\)o\(N\), rather than necessarily bounded, finite\-horizon error\.
### I\.2Verification against history\-dependent opponents
Writesp\(x\)=maxixi−minixi\\operatorname\{sp\}\(x\)=\\max\_\{i\}x\_\{i\}\-\\min\_\{i\}x\_\{i\}\. Nature’s stationary strategy must specify a row for every state and every controller action\. Specifying rows only for the actions used by the selected controller would not define a strategy against a different controller\.
###### Theorem 11\.
Suppose equation[37](https://arxiv.org/html/2609.28792#A9.E37)has a finite solution\. At every state choose
π∗\(i\)\\displaystyle\\pi^\{\*\}\(i\)∈argmaxa∈Ag\(i\)minp∈Fia\(g\)\{ria\+p⊤h\},\\displaystyle\\in\\mathop\{\\rm argmax\}\_\{a\\in A\_\{g\}\(i\)\}\\min\_\{p\\in F\_\{ia\}\(g\)\}\\\{r\_\{ia\}\+p^\{\\top\}h\\\},\(41\)qia∗\\displaystyle q^\{\*\}\_\{ia\}∈\{argminp∈Fia\(g\)p⊤h,a∈Ag\(i\),Fia\(g\),a∉Ag\(i\),\\displaystyle\\in\\begin\{cases\}\\mathop\{\\rm argmin\}\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h,&a\\in A\_\{g\}\(i\),\\\\ F\_\{ia\}\(g\),&a\\notin A\_\{g\}\(i\),\\end\{cases\}\(42\)which specify a deterministic stationary policy and a stationary kernel\. Then, there exist constantsC<∞C<\\inftyand, for everyη\>0\\eta\>0,Cη<∞C\_\{\\eta\}<\\infty, such that for every initial stateii, horizonN≥1N\\geq 1, and randomized history\-dependent strategiesσ,τ\\sigma,\\tau,
𝔼iπ∗,τ∑t=0N−1rStAt\\displaystyle\\mathbb\{E\}\_\{i\}^\{\\pi^\{\*\},\\tau\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}≥N\(gi−η\)−sp\(h\)−Cηsp\(g\),\\displaystyle\\geq N\(g\_\{i\}\-\\eta\)\-\\operatorname\{sp\}\(h\)\-C\_\{\\eta\}\\operatorname\{sp\}\(g\),\(43\)𝔼iσ,q∗∑t=0N−1rStAt\\displaystyle\\mathbb\{E\}\_\{i\}^\{\\sigma,q^\{\*\}\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}≤Ngi\+sp\(h\)\+Csp\(g\)\.\\displaystyle\\leq Ng\_\{i\}\+\\operatorname\{sp\}\(h\)\+C\\operatorname\{sp\}\(g\)\.\(44\)The constants can be chosen uniformly over all selectors satisfying equation[41](https://arxiv.org/html/2609.28792#A9.E41)\-equation[42](https://arxiv.org/html/2609.28792#A9.E42), with dependence
C=C\(r,𝒰,g,h\)<∞,Cη=Cη\(r,𝒰,g,h,η\)<∞\.C=C\(r,\\mathcal\{U\},g,h\)<\\infty,\\qquad C\_\{\\eta\}=C\_\{\\eta\}\(r,\\mathcal\{U\},g,h,\\eta\)<\\infty\.\(45\)They are independent of the initial state, horizon, and opposing strategies\. In generalCηC\_\{\\eta\}need not remain bounded asη↓0\\eta\\downarrow 0\. In particular, it holds that
lim infN→∞JN\(i,π∗,τ\)≥gi,lim supN→∞JN\(i,σ,q∗\)≤gi\.\\liminf\_\{N\\to\\infty\}J\_\{N\}\(i;\\pi^\{\*\},\\tau\)\\geq g\_\{i\},\\qquad\\limsup\_\{N\\to\\infty\}J\_\{N\}\(i;\\sigma,q^\{\*\}\)\\leq g\_\{i\}\.\(46\)For eachΨ∈\{I−,J−,J\+,I\+\}\\Psi\\in\\\{I^\{\-\},J^\{\-\},J^\{\+\},I^\{\+\}\\\}, both the max\-min and min\-max values equalgig\_\{i\}, and\(π∗,q∗\)\(\\pi^\{\*\},q^\{\*\}\)is an all\-state stationary saddle against history\-dependent opponents\. The equalities remain valid when either or both strategy classes are restricted to stationary strategies\. Moreover,
𝔼iπ∗,q∗∑t=0N−1rStAt=Ngi\+hi−𝔼iπ∗,q∗h\(SN\),\\mathbb\{E\}\_\{i\}^\{\\pi^\{\*\},q^\{\*\}\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}=Ng\_\{i\}\+h\_\{i\}\-\\mathbb\{E\}\_\{i\}^\{\\pi^\{\*\},q^\{\*\}\}h\(S\_\{N\}\),\(47\)and
TN0N⟶g\.\\frac\{T^\{N\}0\}\{N\}\\longrightarrow g\.\(48\)Thus the gain component is unique across finite certificates\. If the ambiguity sets for the actionsπ∗\(i\)\\pi^\{\*\}\(i\)are polytopes, then equation[43](https://arxiv.org/html/2609.28792#A9.E43)also holds withη=0\\eta=0and a finiteC0C\_\{0\}\.
###### Proof\.
A gain\-face inequality controls only gain\-minimizing rows\. The first part of the proof extends it to all feasible rows, paying for departures with their nonnegative gain drift\. The cumulative drift is bounded becauseg\(St\)g\(S\_\{t\}\)remains in the finite interval\[minjgj,maxjgj\]\[\\min\_\{j\}g\_\{j\},\\max\_\{j\}g\_\{j\}\]\.
Letℱt\\mathcal\{F\}\_\{t\}denote the full process history beforeAtA\_\{t\}is drawn, and let𝒢t\\mathcal\{G\}\_\{t\}additionally contain the realized actionAtA\_\{t\}and nature’s selected rowptp\_\{t\}, but notSt\+1S\_\{t\+1\}\. Thusℱt⊆𝒢t⊆ℱt\+1\\mathcal\{F\}\_\{t\}\\subseteq\\mathcal\{G\}\_\{t\}\\subseteq\\mathcal\{F\}\_\{t\+1\}and
𝔼\[f\(St\+1\)∣𝒢t\]=pt⊤f\(f∈ℝS\)\.\\mathbb\{E\}\[f\(S\_\{t\+1\}\)\\mid\\mathcal\{G\}\_\{t\}\]=p\_\{t\}^\{\\top\}f\\qquad\(f\\in\\mathbb\{R\}^\{S\}\)\.Using a full filtration for the analysis does not enlarge either player’s admissible information\. It simply includes all already realized randomizations in the joint process\.
1\. Controller inequalities on all feasible rows\.LetAg,h⋆\(i\)A\_\{g,h\}^\{\\star\}\(i\)be the maximizer set in equation[41](https://arxiv.org/html/2609.28792#A9.E41)\. Fora∈Ag,h⋆\(i\)a\\in A\_\{g,h\}^\{\\star\}\(i\), put
dia\(p\)=p⊤g−gi,eia\(p\)=ria\+p⊤h−gi−hi\.d\_\{ia\}\(p\)=p^\{\\top\}g\-g\_\{i\},\\qquad e\_\{ia\}\(p\)=r\_\{ia\}\+p^\{\\top\}h\-g\_\{i\}\-h\_\{i\}\.Gain activity givesdia≥0d\_\{ia\}\\geq 0on𝒰ia\\mathcal\{U\}\_\{ia\}, with zero setFia\(g\)F\_\{ia\}\(g\)\. Bias optimality givesminp∈Fia\(g\)eia\(p\)=0\\min\_\{p\\in F\_\{ia\}\(g\)\}e\_\{ia\}\(p\)=0\. Lemma[7](https://arxiv.org/html/2609.28792#Thmlemma7)therefore applies\. Taking the maximum of its constants over the finitely many pairs\(i,a\)\(i,a\)witha∈Ag,h⋆\(i\)a\\in A\_\{g,h\}^\{\\star\}\(i\)gives, for every allowed selector,
ei,π∗\(i\)\(p\)≥−η−Cηdi,π∗\(i\)\(p\)\(p∈𝒰i,π∗\(i\)\)\.e\_\{i,\\pi^\{\*\}\(i\)\}\(p\)\\geq\-\\eta\-C\_\{\\eta\}d\_\{i,\\pi^\{\*\}\(i\)\}\(p\)\\quad\(p\\in\\mathcal\{U\}\_\{i,\\pi^\{\*\}\(i\)\}\)\.\(49\)This choice ofCηC\_\{\\eta\}is uniform over controller tie\-breaking\.
2\. The controller’s finite\-horizon guarantee\.Fix any nature strategyτ\\tauand useπ∗\\pi^\{\*\}\. Writedt=dStAt\(pt\)d\_\{t\}=d\_\{S\_\{t\}A\_\{t\}\}\(p\_\{t\}\)andet=eStAt\(pt\)e\_\{t\}=e\_\{S\_\{t\}A\_\{t\}\}\(p\_\{t\}\)\. Thendt≥0d\_\{t\}\\geq 0, and the transition rule gives
𝔼\[g\(St\+1\)−g\(St\)∣𝒢t\]=dt\.\\mathbb\{E\}\[g\(S\_\{t\+1\}\)\-g\(S\_\{t\}\)\\mid\\mathcal\{G\}\_\{t\}\]=d\_\{t\}\.The tower property makesg\(St\)g\(S\_\{t\}\)a bounded submartingale, so𝔼ig\(St\)≥gi\\mathbb\{E\}\_\{i\}g\(S\_\{t\}\)\\geq g\_\{i\}\. Taking expectations and summing gives
∑t=0N−1𝔼idt=𝔼ig\(SN\)−gi≤sp\(g\)\.\\sum\_\{t=0\}^\{N\-1\}\\mathbb\{E\}\_\{i\}d\_\{t\}=\\mathbb\{E\}\_\{i\}g\(S\_\{N\}\)\-g\_\{i\}\\leq\\operatorname\{sp\}\(g\)\.\(50\)Similarly, the definition ofete\_\{t\}yields the exact expected reward identity
𝔼i∑t=0N−1rStAt=∑t=0N−1𝔼ig\(St\)\+hi−𝔼ih\(SN\)\+∑t=0N−1𝔼iet\.\\mathbb\{E\}\_\{i\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}=\\sum\_\{t=0\}^\{N\-1\}\\mathbb\{E\}\_\{i\}g\(S\_\{t\}\)\+h\_\{i\}\-\\mathbb\{E\}\_\{i\}h\(S\_\{N\}\)\+\\sum\_\{t=0\}^\{N\-1\}\\mathbb\{E\}\_\{i\}e\_\{t\}\.Here the bias terms telescope, without requiring a limit of the state process\. Substituting𝔼ig\(St\)≥gi\\mathbb\{E\}\_\{i\}g\(S\_\{t\}\)\\geq g\_\{i\},hi−𝔼ih\(SN\)≥−sp\(h\)h\_\{i\}\-\\mathbb\{E\}\_\{i\}h\(S\_\{N\}\)\\geq\-\\operatorname\{sp\}\(h\), and equation[49](https://arxiv.org/html/2609.28792#A9.E49)\-equation[50](https://arxiv.org/html/2609.28792#A9.E50)proves equation[43](https://arxiv.org/html/2609.28792#A9.E43)\. Dividing byNN, taking the limit inferior at fixedη\\eta, and then lettingη↓0\\eta\\downarrow 0proves the controller half of equation[46](https://arxiv.org/html/2609.28792#A9.E46)\. Neither the selector nor the opponent changes withη\\eta\.
3\. Nature’s full selector and upper guarantee\.For every action, including those unused byπ∗\\pi^\{\*\}, define
d¯ia=\(qia∗\)⊤g−gi,e¯ia=ria\+\(qia∗\)⊤h−gi−hi\.\\bar\{d\}\_\{ia\}=\(q^\{\*\}\_\{ia\}\)^\{\\top\}g\-g\_\{i\},\\qquad\\bar\{e\}\_\{ia\}=r\_\{ia\}\+\(q^\{\*\}\_\{ia\}\)^\{\\top\}h\-g\_\{i\}\-h\_\{i\}\.For active actions, the gain\-face and bias\-minimizing choices gived¯ia=0\\bar\{d\}\_\{ia\}=0and
e¯ia=minp∈Fia\(g\)\(ria\+p⊤h\)−\(Lgh\)i≤0\.\\bar\{e\}\_\{ia\}=\\min\_\{p\\in F\_\{ia\}\(g\)\}\(r\_\{ia\}\+p^\{\\top\}h\)\-\(L\_\{g\}h\)\_\{i\}\\leq 0\.For inactive actions,d¯ia=mia\(g\)−gi<0\\bar\{d\}\_\{ia\}=m\_\{ia\}\(g\)\-g\_\{i\}<0\. Their possibly positive bias residual can be charged to this strictly negative gain drift\. Specifically, set
C=max\(i,a\):a∉Ag\(i\)maxp∈Fia\(g\)\[ria\+p⊤h−gi−hi\]\+gi−mia\(g\),C=\\max\_\{\(i,a\):\\,a\\notin A\_\{g\}\(i\)\}\\frac\{\\max\_\{p\\in F\_\{ia\}\(g\)\}\[r\_\{ia\}\+p^\{\\top\}h\-g\_\{i\}\-h\_\{i\}\]\_\{\+\}\}\{g\_\{i\}\-m\_\{ia\}\(g\)\},with an empty maximum equal to zero\. Compactness bounds the numerators, and there are finitely many positive denominators\. ThusCCis finite, independent of nature’s tie\-breaking, and
d¯ia≤0,e¯ia≤C\(−d¯ia\)for every\(i,a\)\.\\bar\{d\}\_\{ia\}\\leq 0,\\qquad\\bar\{e\}\_\{ia\}\\leq C\(\-\\bar\{d\}\_\{ia\}\)\\quad\\text\{for every \}\(i,a\)\.\(51\)
Against any randomized history\-dependent controllerσ\\sigma, putd¯t=d¯StAt\\bar\{d\}\_\{t\}=\\bar\{d\}\_\{S\_\{t\}A\_\{t\}\}ande¯t=e¯StAt\\bar\{e\}\_\{t\}=\\bar\{e\}\_\{S\_\{t\}A\_\{t\}\}\. Conditioning on its realized action makes the preceding inequalities applicable\. Consequentlyg\(St\)g\(S\_\{t\}\)is a bounded supermartingale, with
𝔼ig\(St\)≤gi,∑t<N𝔼i\(−d¯t\)=gi−𝔼ig\(SN\)≤sp\(g\)\.\\mathbb\{E\}\_\{i\}g\(S\_\{t\}\)\\leq g\_\{i\},\\qquad\\sum\_\{t<N\}\\mathbb\{E\}\_\{i\}\(\-\\bar\{d\}\_\{t\}\)=g\_\{i\}\-\\mathbb\{E\}\_\{i\}g\(S\_\{N\}\)\\leq\\operatorname\{sp\}\(g\)\.The same reward identity as in part 2 now gives
𝔼iσ,q∗∑t<NrStAt≤Ngi\+sp\(h\)\+Csp\(g\),\\mathbb\{E\}\_\{i\}^\{\\sigma,q^\{\*\}\}\\sum\_\{t<N\}r\_\{S\_\{t\}A\_\{t\}\}\\leq Ng\_\{i\}\+\\operatorname\{sp\}\(h\)\+C\\operatorname\{sp\}\(g\),proving equation[44](https://arxiv.org/html/2609.28792#A9.E44)and the nature half of equation[46](https://arxiv.org/html/2609.28792#A9.E46)\.
4\. The two pathwise payoff conventions\.The expected bounds above alone do not imply bounds on𝔼lim inf\\mathbb\{E\}\\liminfor𝔼lim sup\\mathbb\{E\}\\limsup\. We establish those directly\. For either one\-sided strategy pair define
ξt\+1=h\(St\+1\)−pt⊤h,MN=∑t=0N−1ξt\+1\.\\xi\_\{t\+1\}=h\(S\_\{t\+1\}\)\-p\_\{t\}^\{\\top\}h,\\qquad M\_\{N\}=\\sum\_\{t=0\}^\{N\-1\}\\xi\_\{t\+1\}\.The transition rule and tower property give𝔼\[ξt\+1∣ℱt\]=0\\mathbb\{E\}\[\\xi\_\{t\+1\}\\mid\\mathcal\{F\}\_\{t\}\]=0, and\|ξt\+1\|≤sp\(h\)\|\\xi\_\{t\+1\}\|\\leq\\operatorname\{sp\}\(h\)\. HenceMN/N→0M\_\{N\}/N\\to 0almost surely, by the bounded martingale\-difference strong law \(equivalently, Azuma\-Hoeffding and Borel\-Cantelli\)\. The exact sample\-path identity is
∑t<NrStAt=∑t<Ng\(St\)\+hi−h\(SN\)\+∑t<Net\+MN,\\sum\_\{t<N\}r\_\{S\_\{t\}A\_\{t\}\}=\\sum\_\{t<N\}g\(S\_\{t\}\)\+h\_\{i\}\-h\(S\_\{N\}\)\+\\sum\_\{t<N\}e\_\{t\}\+M\_\{N\},usinge¯t\\bar\{e\}\_\{t\}for the nature pair\.
Under\(π∗,τ\)\(\\pi^\{\*\},\\tau\), bounded\-submartingale convergence givesg\(St\)→G∞g\(S\_\{t\}\)\\to G\_\{\\infty\}almost surely and inL1L^\{1\}, with𝔼iG∞≥gi\\mathbb\{E\}\_\{i\}G\_\{\\infty\}\\geq g\_\{i\}\. Monotone convergence applied to equation[50](https://arxiv.org/html/2609.28792#A9.E50)gives𝔼i∑t≥0dt≤sp\(g\)\\mathbb\{E\}\_\{i\}\\sum\_\{t\\geq 0\}d\_\{t\}\\leq\\operatorname\{sp\}\(g\), hence∑t≥0dt<∞\\sum\_\{t\\geq 0\}d\_\{t\}<\\inftyalmost surely\. For eachη\>0\\eta\>0, the row inequality thus implies
lim infN1N∑t<NrStAt≥G∞−ηalmost surely\.\\liminf\_\{N\}\\frac\{1\}\{N\}\\sum\_\{t<N\}r\_\{S\_\{t\}A\_\{t\}\}\\geq G\_\{\\infty\}\-\\eta\\quad\\text\{almost surely\}\.Taking a countable sequenceη↓0\\eta\\downarrow 0proves the same bound withη=0\\eta=0\. After expectations, this isIi−\(π∗,τ\)≥giI\_\{i\}^\{\-\}\(\\pi^\{\*\},\\tau\)\\geq g\_\{i\}\.
Under\(σ,q∗\)\(\\sigma,q^\{\*\}\), bounded\-supermartingale convergence instead givesg\(St\)→G¯∞g\(S\_\{t\}\)\\to\\bar\{G\}\_\{\\infty\}with𝔼iG¯∞≤gi\\mathbb\{E\}\_\{i\}\\bar\{G\}\_\{\\infty\}\\leq g\_\{i\}, and∑t\(−d¯t\)<∞\\sum\_\{t\}\(\-\\bar\{d\}\_\{t\}\)<\\inftyalmost surely\. From equation[51](https://arxiv.org/html/2609.28792#A9.E51)and the same path identity,
lim supN1N∑t<NrStAt≤G¯∞almost surely,\\limsup\_\{N\}\\frac\{1\}\{N\}\\sum\_\{t<N\}r\_\{S\_\{t\}A\_\{t\}\}\\leq\\bar\{G\}\_\{\\infty\}\\quad\\text\{almost surely\},soIi\+\(σ,q∗\)≤giI\_\{i\}^\{\+\}\(\\sigma,q^\{\*\}\)\\leq g\_\{i\}\. For bounded rewards, Fatou’s lemma and its reverse giveI−≤J−≤J\+≤I\+I^\{\-\}\\leq J^\{\-\}\\leq J^\{\+\}\\leq I^\{\+\}\. Thus these two guarantees bracket every payoffΨ\\Psiin the theorem\. Weak duality then gives
gi≤supσinfτΨi\(σ,τ\)≤infτsupσΨi\(σ,τ\)≤gi\.g\_\{i\}\\leq\\sup\_\{\\sigma\}\\inf\_\{\\tau\}\\Psi\_\{i\}\(\\sigma,\\tau\)\\leq\\inf\_\{\\tau\}\\sup\_\{\\sigma\}\\Psi\_\{i\}\(\\sigma,\\tau\)\\leq g\_\{i\}\.The same argument holds for any restricted strategy classes containingπ∗\\pi^\{\*\}andq∗q^\{\*\}, including the stated stationary classes\.
5\. The selected pair, finite\-horizon limit, and uniqueness\.When both selected strategies are used, gain and bias residuals vanish:
Pπ∗q∗g=g,rπ∗\+Pπ∗q∗h=g\+h\.P^\{\\pi^\{\*\}q^\{\*\}\}g=g,\\qquad r^\{\\pi^\{\*\}\}\+P^\{\\pi^\{\*\}q^\{\*\}\}h=g\+h\.The reward identity from part 2 is therefore exactly equation[47](https://arxiv.org/html/2609.28792#A9.E47)\. The finite\-horizon dynamic\-programming value isTN0T^\{N\}0\(Lemma[1](https://arxiv.org/html/2609.28792#Thmlemma1)\)\. The two uniform guarantees yield
gi−η−sp\(h\)\+Cηsp\(g\)N≤\(TN0\)iN≤gi\+sp\(h\)\+Csp\(g\)N\.g\_\{i\}\-\\eta\-\\frac\{\\operatorname\{sp\}\(h\)\+C\_\{\\eta\}\\operatorname\{sp\}\(g\)\}\{N\}\\leq\\frac\{\(T^\{N\}0\)\_\{i\}\}\{N\}\\leq g\_\{i\}\+\\frac\{\\operatorname\{sp\}\(h\)\+C\\operatorname\{sp\}\(g\)\}\{N\}\.First sendN→∞N\\to\\inftyat fixedη\\etaand thenη↓0\\eta\\downarrow 0\. There are finitely many states, so this proves equation[48](https://arxiv.org/html/2609.28792#A9.E48)in norm and uniqueness of the gain in any finite certificate\. It also identifies that gain with the robust valueg⋆g^\{\\star\}\. Finally, if the selected\-action row sets are polytopes, the affine part of Lemma[7](https://arxiv.org/html/2609.28792#Thmlemma7)suppliesC0<∞C\_\{0\}<\\infty\. Repeating part 2 withη=0\\eta=0proves the final assertion\. ∎
###### Corollary 1\.
Under\(π∗,q∗\)\(\\pi^\{\*\},q^\{\*\}\), there is a bounded random variableG∞G\_\{\\infty\}, the recurrent classwise gain, such that
g\(St\)⟶G∞almost surely,1N∑t=0N−1rStAt⟶G∞almost surely,𝔼iG∞=gi\.g\(S\_\{t\}\)\\longrightarrow G\_\{\\infty\}\\quad\\text\{almost surely\},\\qquad\\frac\{1\}\{N\}\\sum\_\{t=0\}^\{N\-1\}r\_\{S\_\{t\}A\_\{t\}\}\\longrightarrow G\_\{\\infty\}\\quad\\text\{almost surely\},\\qquad\\mathbb\{E\}\_\{i\}G\_\{\\infty\}=g\_\{i\}\.\(52\)In every recurrent class of the selected finite Markov chain,ggis constant and equals that class’s average reward\. From a transient state,gig\_\{i\}is the absorption\-probability weighted average of these class gains\.
###### Proof\.
WriteP∗=Pπ∗q∗P^\{\*\}=P^\{\\pi^\{\*\}q^\{\*\}\}andri∗=ri,π∗\(i\)r\_\{i\}^\{\*\}=r\_\{i,\\pi^\{\*\}\(i\)\}\. The selections give
P∗g=g,\(I−P∗\)h=r∗−g\.P^\{\*\}g=g,\\qquad\(I\-P^\{\*\}\)h=r^\{\*\}\-g\.\(53\)Let\(P∗\)∞\(P^\{\*\}\)^\{\\infty\}be its Cesàro projector from Lemma[2](https://arxiv.org/html/2609.28792#Thmlemma2)\. The first equality implies\(P∗\)∞g=g\(P^\{\*\}\)^\{\\infty\}g=g\. Applying the projector to the second, and using\(P∗\)∞\(I−P∗\)=0\(P^\{\*\}\)^\{\\infty\}\(I\-P^\{\*\}\)=0, gives\(P∗\)∞r∗=g\(P^\{\*\}\)^\{\\infty\}r^\{\*\}=g\. Thusggis the ordinary stationary gain of the selected chain\. The finite\-chain limit and absorption formula in Lemma[3](https://arxiv.org/html/2609.28792#Thmlemma3)now prove both almost\-sure limits and𝔼iG∞=gi\\mathbb\{E\}\_\{i\}G\_\{\\infty\}=g\_\{i\}\. Periodic recurrent classes require no extra assumption, since the reward averages are Cesàro averages\. ∎
### I\.3Fixed\-policy specializations of the verification theorem
For a deterministic stationary controllerπ\\pi, put
\(Tπx\)i=ri,π\(i\)\+minp∈𝒰i,π\(i\)p⊤x,\(T^πg\)i=minp∈𝒰i,π\(i\)p⊤g\.\(T^\{\\pi\}x\)\_\{i\}=r\_\{i,\\pi\(i\)\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{i,\\pi\(i\)\}\}p^\{\\top\}x,\\qquad\(\\widehat\{T\}^\{\\pi\}g\)\_\{i\}=\\min\_\{p\\in\\mathcal\{U\}\_\{i,\\pi\(i\)\}\}p^\{\\top\}g\.Giveng=T^πgg=\\widehat\{T\}^\{\\pi\}g, defineFiπ\(g\)=\{p∈𝒰i,π\(i\):p⊤g=gi\}F\_\{i\}^\{\\pi\}\(g\)=\\\{p\\in\\mathcal\{U\}\_\{i,\\pi\(i\)\}:p^\{\\top\}g=g\_\{i\}\\\}\. The two evaluation equations are
g=T^πg,gi\+hi=ri,π\(i\)\+minp∈Fiπ\(g\)p⊤h\(i∈S\)\.g=\\widehat\{T\}^\{\\pi\}g,\\qquad g\_\{i\}\+h\_\{i\}=r\_\{i,\\pi\(i\)\}\+\\min\_\{p\\in F\_\{i\}^\{\\pi\}\(g\)\}p^\{\\top\}h\\quad\(i\\in S\)\.\(54\)Both equations are necessary for the certificate\. In particular, merely substituting a vector for the scalar gain ing\+h=Tπhg\+h=T^\{\\pi\}hdoes not impose the first\-level row restriction\.
###### Propsition 4\.
If\(g,h\)\(g,h\)solves equation[54](https://arxiv.org/html/2609.28792#A9.E54), choose
qi∗∈argminp∈Fiπ\(g\)p⊤h\.q\_\{i\}^\{\*\}\\in\\mathop\{\\rm argmin\}\_\{p\\in F\_\{i\}^\{\\pi\}\(g\)\}p^\{\\top\}h\.Then, for every initial state,
infτlim infNJN\(i,π,τ\)=infτlim supNJN\(i,π,τ\)=gi,\(Tπ\)N0N⟶g\.\\inf\_\{\\tau\}\\liminf\_\{N\}J\_\{N\}\(i;\\pi,\\tau\)=\\inf\_\{\\tau\}\\limsup\_\{N\}J\_\{N\}\(i;\\pi,\\tau\)=g\_\{i\},\\qquad\\frac\{\(T^\{\\pi\}\)^\{N\}0\}\{N\}\\longrightarrow g\.\(55\)The infima may be taken over all history\-dependent randomized nature strategies or only stationary strategies\. The one stationary selectorq∗q^\{\*\}attains both infima simultaneously at every initial state\. Under\(π,q∗\)\(\\pi,q^\{\*\}\), the identity equation[47](https://arxiv.org/html/2609.28792#A9.E47)and the pathwise interpretation in Corollary[1](https://arxiv.org/html/2609.28792#Thmcorollary1)hold\.
###### Proof\.
Restrict the action set at stateiito\{π\(i\)\}\\\{\\pi\(i\)\\\}\. The gain and bias equations of this one\-action model are precisely equation[54](https://arxiv.org/html/2609.28792#A9.E54), so Theorem[11](https://arxiv.org/html/2609.28792#Thmtheorem11)applies\. For everyη\>0\\eta\>0and every history\-dependentτ\\tau,
JN\(i,π,τ\)≥gi−η−sp\(h\)\+Cηsp\(g\)N,JN\(i,π,q∗\)=gi\+hi−𝔼iπ,q∗h\(SN\)N\.J\_\{N\}\(i;\\pi,\\tau\)\\geq g\_\{i\}\-\\eta\-\\frac\{\\operatorname\{sp\}\(h\)\+C\_\{\\eta\}\\operatorname\{sp\}\(g\)\}\{N\},\\qquad J\_\{N\}\(i;\\pi,q^\{\*\}\)=g\_\{i\}\+\\frac\{h\_\{i\}\-\\mathbb\{E\}\_\{i\}^\{\\pi,q^\{\*\}\}h\(S\_\{N\}\)\}\{N\}\.The second numerator is bounded bysp\(h\)\\operatorname\{sp\}\(h\)in absolute value\. Taking limits proves both infimum identities and simultaneous stationary attainment\. The finite\-horizon and pathwise conclusions are the corresponding conclusions of the same verification theorem and Corollary[1](https://arxiv.org/html/2609.28792#Thmcorollary1)\. ∎
#### I\.3\.1Randomized stationary policies and action\-contingent nature
Letπ\(a∣i\)\\pi\(a\\mid i\)be a fixed stationary randomized policy\. Nature observes the realized action\. Define the expected one\-step reward and the set of effective transition rows by
riπ=∑a∈A\(i\)π\(a∣i\)ria,𝒰iπ=\{∑a∈A\(i\)π\(a∣i\)pa:pa∈𝒰iafor alla\}\.r\_\{i\}^\{\\pi\}=\\sum\_\{a\\in A\(i\)\}\\pi\(a\\mid i\)r\_\{ia\},\\qquad\\mathcal\{U\}\_\{i\}^\{\\pi\}=\\left\\\{\\sum\_\{a\\in A\(i\)\}\\pi\(a\\mid i\)p\_\{a\}:p\_\{a\}\\in\\mathcal\{U\}\_\{ia\}\\text\{ for all \}a\\right\\\}\.\(56\)The set𝒰iπ\\mathcal\{U\}\_\{i\}^\{\\pi\}is nonempty and compact as the continuous image of a finite product of compact sets\. Zero\-probability actions can be omitted from this product without changing𝒰iπ\\mathcal\{U\}\_\{i\}^\{\\pi\}\.
###### Propsition 5\.
For the post\-action nature model, the fixed\-policy Bellman operator is
\(Tπx\)i=∑aπ\(a∣i\)\(ria\+minp∈𝒰iap⊤x\)=riπ\+minp¯∈𝒰iπp¯⊤x\.\(T^\{\\pi\}x\)\_\{i\}=\\sum\_\{a\}\\pi\(a\\mid i\)\\left\(r\_\{ia\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}x\\right\)=r\_\{i\}^\{\\pi\}\+\\min\_\{\\bar\{p\}\\in\\mathcal\{U\}\_\{i\}^\{\\pi\}\}\\bar\{p\}^\{\\top\}x\.\(57\)Its finite vector evaluation certificate is
gi\\displaystyle g\_\{i\}=∑aπ\(a∣i\)mia\(g\),\\displaystyle=\\sum\_\{a\}\\pi\(a\\mid i\)m\_\{ia\}\(g\),\(58\)gi\+hi\\displaystyle g\_\{i\}\+h\_\{i\}=∑aπ\(a∣i\)\(ria\+minp∈Fia\(g\)p⊤h\)\.\\displaystyle=\\sum\_\{a\}\\pi\(a\\mid i\)\\left\(r\_\{ia\}\+\\min\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h\\right\)\.\(59\)Every finite solution has all the expected\-value conclusions in Proposition[4](https://arxiv.org/html/2609.28792#Thmproposition4)\. A worst stationary nature strategy is obtained by choosing, separately for every positive\-probability action,
qia∗∈argminp∈Fia\(g\)p⊤h\.q^\{\*\}\_\{ia\}\\in\\mathop\{\\rm argmin\}\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h\.\(60\)In this statementggmay be nonconstant, and the individual quantitiesmia\(g\)m\_\{ia\}\(g\)need not equalgig\_\{i\}\.
###### Proof\.
The effective\-row reduction in equation[17](https://arxiv.org/html/2609.28792#A3.E17)applies because nature sees the realized action\. We spell out the gain\-face calculation, which is the additional point needed for vector gains\.
For each vectorxx, independent minimization of positive\-weight summands gives
minp¯∈𝒰iπp¯⊤x=∑aπ\(a∣i\)minp∈𝒰iap⊤x\.\\min\_\{\\bar\{p\}\\in\\mathcal\{U\}\_\{i\}^\{\\pi\}\}\\bar\{p\}^\{\\top\}x=\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}x\.Indeed, any tuple gives at least the right\-hand side, and compactness allows each component minimum to be attained simultaneously\. This proves equation[57](https://arxiv.org/html/2609.28792#A9.E57)and its recession equation equation[58](https://arxiv.org/html/2609.28792#A9.E58)\.
Assume that gain equation\. For every tuple representingp¯\\bar\{p\},
p¯⊤g−gi=∑aπ\(a∣i\)\(pa⊤g−mia\(g\)\)\.\\bar\{p\}^\{\\top\}g\-g\_\{i\}=\\sum\_\{a\}\\pi\(a\\mid i\)\\bigl\(p\_\{a\}^\{\\top\}g\-m\_\{ia\}\(g\)\\bigr\)\.Every summand is nonnegative\. The sum vanishes exactly whenpa∈Fia\(g\)p\_\{a\}\\in F\_\{ia\}\(g\)for every positive\-weight action\. This statement holds for every representation ofp¯\\bar\{p\}\. Consequently the effective gain face is the set of weighted sums of these component faces, and its bias minimum is
minp¯∈𝒰iπ:p¯⊤g=gip¯⊤h=∑aπ\(a∣i\)minp∈Fia\(g\)p⊤h\.\\min\_\{\\bar\{p\}\\in\\mathcal\{U\}\_\{i\}^\{\\pi\}:\\,\\bar\{p\}^\{\\top\}g=g\_\{i\}\}\\bar\{p\}^\{\\top\}h=\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h\.This identifies equation[59](https://arxiv.org/html/2609.28792#A9.E59)as the one\-action model’s bias equation and proves feasibility of the selector equation[60](https://arxiv.org/html/2609.28792#A9.E60)\. For actions of zero probability, fill the unused selector entries with arbitrary feasible rows\.
For completeness, the row inequalities also survive history\-dependent post\-action randomization directly\. Conditional on the pre\-action historyHHandSt=iS\_\{t\}=i, letκH,a\\kappa\_\{H,a\}be nature’s conditional law on𝒰ia\\mathcal\{U\}\_\{ia\}after actionaa\. Integrate the effective\-row inequalities over the product law⨂aκH,a\\bigotimes\_\{a\}\\kappa\_\{H,a\}\. For every vectorff, the resulting continuation term is
∑aπ\(a∣i\)∫pa⊤fκH,a\(dpa\)=𝔼\[f\(St\+1\)∣H\],riπ=𝔼\[rStAt∣H\]\.\\sum\_\{a\}\\pi\(a\\mid i\)\\int p\_\{a\}^\{\\top\}f\\,\\kappa\_\{H,a\}\(dp\_\{a\}\)=\\mathbb\{E\}\[f\(S\_\{t\+1\}\)\\mid H\],\\qquad r\_\{i\}^\{\\pi\}=\\mathbb\{E\}\[r\_\{S\_\{t\}A\_\{t\}\}\\mid H\]\.The controller lower\-bound proof of Theorem[11](https://arxiv.org/html/2609.28792#Thmtheorem11)therefore applies after this pre\-action conditioning\. This argument integrates inequalities valid for every tuple\. It does not require a mean row to belong to a nonconvex row set\.
Conversely, the chosen effective row is implemented by its selected componentqia∗q^\{\*\}\_\{ia\}after the sampled action\. The resulting stationary pair satisfiesPπ,q∗g=gP^\{\\pi,q^\{\*\}\}g=gandrπ\+Pπ,q∗h=g\+hr^\{\\pi\}\+P^\{\\pi,q^\{\*\}\}h=g\+h, yielding the expected Poisson identity and attainment\. Its finite\-horizon operator is equation[57](https://arxiv.org/html/2609.28792#A9.E57), so the normalized finite\-horizon limit also follows\. These are all the expected\-value conclusions claimed\. ∎
### I\.4Why both gain restrictions are necessary
The following two counterexamples show separately that nature’s minimizing face and the controller’s active action set are necessary\. They use the same transition geometry, so the controller example can reuse the absorption calculation from the nature example\.
###### Example 1\(Nature’s gain restriction cannot be omitted\)\.
There are states\(s,m,z\)\(s,m,z\), rewards\(0,−1,0\)\(0,\-1,0\), and absorbing statesm,zm,z\. Atss, let
𝒰s=\{pλ=\(1−λ/2,λ/4,λ/4\):0≤λ≤1\}\.\\mathcal\{U\}\_\{s\}=\\\{p^\{\\lambda\}=\(1\-\\lambda/2,\\lambda/4,\\lambda/4\):0\\leq\\lambda\\leq 1\\\}\.The robust gain isg=\(−1/2,−1,0\)g=\(\-1/2,\-1,0\)andh=\(0,−2,0\)h=\(0,\-2,0\)is a Bellman bias\. However, the unrestricted equationsT^g~=g~\\widehat\{T\}\\widetilde\{g\}=\\widetilde\{g\}andTh~=g~\+h~T\\widetilde\{h\}=\\widetilde\{g\}\+\\widetilde\{h\}also accept the incorrect pairg~=\(−3/4,−1,0\)\\widetilde\{g\}=\(\-3/4,\-1,0\),h~=\(0,−3,0\)\\widetilde\{h\}=\(0,\-3,0\)\.
###### Proof\.
Letτm,τz\\tau\_\{m\},\\tau\_\{z\}be the hitting times of the absorbing states\. For any adaptive choiceλt\\lambda\_\{t\}, the two first\-entry probabilities at timet\+1t\+1are equal:
ℙs\(τm=t\+1\)=𝔼s\[𝟏\{St=s\}λt/4\]=ℙs\(τz=t\+1\)\.\\mathbb\{P\}\_\{s\}\(\\tau\_\{m\}=t\+1\)=\\mathbb\{E\}\_\{s\}\[\\mathbf\{1\}\_\{\\\{S\_\{t\}=s\\\}\}\\lambda\_\{t\}/4\]=\\mathbb\{P\}\_\{s\}\(\\tau\_\{z\}=t\+1\)\.Summing overttshows that each eventual absorption probability is at most1/21/2\. The sample average converges to−𝟏\{τm<∞\}\-\\mathbf\{1\}\_\{\\\{\\tau\_\{m\}<\\infty\\\}\}; bounded convergence therefore gives expected average at least−1/2\-1/2\. Takingλt=1\\lambda\_\{t\}=1at every visit tossattains−1/2\-1/2, because the survival probability aftertttransitions is2−t2^\{\-t\}\. This proves the stated gain\. Moreover,
\(pλ\)⊤g=−1/2,\(pλ\)⊤h=−λ/2,\(p^\{\\lambda\}\)^\{\\top\}g=\-1/2,\\qquad\(p^\{\\lambda\}\)^\{\\top\}h=\-\\lambda/2,so every row is gain\-minimizing and the gain\-face bias equation holds\. The absorbing\-state equations are identities\.
For the proposed incorrect pair, direct substitution gives
\(pλ\)⊤g~=−3/4\+λ/8,\(pλ\)⊤h~=−3λ/4\.\(p^\{\\lambda\}\)^\{\\top\}\\widetilde\{g\}=\-3/4\+\\lambda/8,\\qquad\(p^\{\\lambda\}\)^\{\\top\}\\widetilde\{h\}=\-3\\lambda/4\.Thus the gain minimum is−3/4\-3/4, attained only atλ=0\\lambda=0, whereas the unrestricted bias minimum is−3/4\-3/4, attained atλ=1\\lambda=1\. The unrestricted equations combine these incompatible choices\. On the actual gain\-minimizing face\{p0\}\\\{p^\{0\}\\\}, the bias equation would instead require−3/4\+0=0\-3/4\+0=0, which is impossible\. ∎
###### Example 2\(The controller’s gain restriction cannot be omitted\)\.
Use states\(s,m,z\)\(s,m,z\)with rewards\(0,1,0\)\(0,1,0\), and makem,zm,zabsorbing\. Atss, the controller chooses between the nominal rowsp0=\(1,0,0\)p^\{0\}=\(1,0,0\)andp1=\(1/2,1/4,1/4\)p^\{1\}=\(1/2,1/4,1/4\)\. The true gain is\(1/2,1,0\)\(1/2,1,0\), but the unrestricted equations acceptg~=\(3/4,1,0\)\\widetilde\{g\}=\(3/4,1,0\)andh~=\(0,3,0\)\\widetilde\{h\}=\(0,3,0\)\.
###### Proof\.
Under either action, the probabilities of enteringmmandzzon the next step are equal\. The first\-entry calculation in Example[1](https://arxiv.org/html/2609.28792#Thmexample1)therefore givesℙs\(τm<∞\)≤1/2\\mathbb\{P\}\_\{s\}\(\\tau\_\{m\}<\\infty\)\\leq 1/2under every controller strategy\. The sample average converges to𝟏\{τm<∞\}\\mathbf\{1\}\_\{\\\{\\tau\_\{m\}<\\infty\\\}\}\. Always usingp1p^\{1\}attains absorption probability1/21/2, which proves the true gain, including against history\-dependent control\. For the incorrect pair,
\(p0\)⊤g~=3/4,\(p1\)⊤g~=5/8,\(p0\)⊤h~=0,\(p1\)⊤h~=3/4\.\(p^\{0\}\)^\{\\top\}\\widetilde\{g\}=3/4,\\qquad\(p^\{1\}\)^\{\\top\}\\widetilde\{g\}=5/8,\\qquad\(p^\{0\}\)^\{\\top\}\\widetilde\{h\}=0,\\qquad\(p^\{1\}\)^\{\\top\}\\widetilde\{h\}=3/4\.The gain maximum is attained only by actionp0p^\{0\}, whereas the unrestricted bias maximum usesp1p^\{1\}\. Hence the unrestricted equations hold, but the gain\-active bias equation requires3/4\+0=03/4\+0=0atss\. This contradiction proves that the controller restriction is necessary independently of nature’s restriction\. ∎
## Appendix JProofs of Theorems[4](https://arxiv.org/html/2609.28792#Thmtheorem4)and[5](https://arxiv.org/html/2609.28792#Thmtheorem5): finite Bellman solvability
Theorem[12](https://arxiv.org/html/2609.28792#Thmtheorem12)and Corollary[2](https://arxiv.org/html/2609.28792#Thmcorollary2)prove Theorem[4](https://arxiv.org/html/2609.28792#Thmtheorem4)\. Theorems[13](https://arxiv.org/html/2609.28792#Thmtheorem13)and[14](https://arxiv.org/html/2609.28792#Thmtheorem14)prove the two equivalent forms of Theorem[5](https://arxiv.org/html/2609.28792#Thmtheorem5)\. Structural sufficient conditions and the two failure mechanisms are collected separately in Appendix[M](https://arxiv.org/html/2609.28792#A13)\.
This section separates a prescribed gain from an unspecified gain\. For a prescribed recession fixed pointgg, the question is whether the equationg\+h=Lghg\+h=L\_\{g\}hhas a finite solution\. When the gain is unspecified, an additive eigenvalue of the tangent operator can be absorbed into a scalar shift ofgg\. We first specialize the classical compact\-action criterion, then use it in the paper\-specific optimal\-control argument\.
### J\.1The classical criterion and fixed\-policy specialization
We state the one\-player result in a slightly broader form so that it applies both to fixed\-policy evaluation and to the tangent\-game argument below\. At stateii, letBiB\_\{i\}be a nonempty compact metric action space\. An actionb∈Bib\\in B\_\{i\}has a continuous rewardci\(b\)c\_\{i\}\(b\)and continuous transition rowpi\(b\)p\_\{i\}\(b\)\. The minimizing player chooses a stationary deterministic policyq∈𝒬:=∏iBiq\\in\\mathcal\{Q\}:=\\prod\_\{i\}B\_\{i\}\. WritePqP\_\{q\}for its transition matrix andciq=ci\(qi\)c^\{q\}\_\{i\}=c\_\{i\}\(q\_\{i\}\)\. Define componentwise
γi=infq∈𝒬\(Pq∞cq\)i,𝒬∗=\{q∈𝒬:Pq∞cq=γ\},w\(q\)=ZPq\(cq−Pq∞cq\)\.\\gamma\_\{i\}=\\inf\_\{q\\in\\mathcal\{Q\}\}\(P\_\{q\}^\{\\infty\}c^\{q\}\)\_\{i\},\\quad\\mathcal\{Q\}\_\{\*\}=\\\{q\\in\\mathcal\{Q\}:P\_\{q\}^\{\\infty\}c^\{q\}=\\gamma\\\},\\quad w\(q\)=Z\_\{P\_\{q\}\}\(c^\{q\}\-P\_\{q\}^\{\\infty\}c^\{q\}\)\.\(61\)The infimum defines a finite vector because rewards are bounded\. It need not be attained by one policy simultaneously at all states\.
###### Theorem 12\.
The coupled equations
zi\\displaystyle z\_\{i\}=minb∈Bipi\(b\)𝖳z,\\displaystyle=\\min\_\{b\\in B\_\{i\}\}p\_\{i\}\(b\)^\{\\mathsf\{T\}\}z,\(62\)zi\+hi\\displaystyle z\_\{i\}\+h\_\{i\}=minb:pi\(b\)𝖳z=zi\{ci\(b\)\+pi\(b\)𝖳h\}\\displaystyle=\\min\_\{b:\\,p\_\{i\}\(b\)^\{\\mathsf\{T\}\}z=z\_\{i\}\}\\\{c\_\{i\}\(b\)\+p\_\{i\}\(b\)^\{\\mathsf\{T\}\}h\\\}\(63\)have a finite solution if and only if
𝒬∗≠∅,∃B<∞w\(q\)≥−B𝟏for everyq∈𝒬∗\.\\mathcal\{Q\}\_\{\*\}\\neq\\varnothing,\\qquad\\exists B<\\infty\\quad w\(q\)\\geq\-B\\mathbf\{1\}\\quad\\text\{for every \}q\\in\\mathcal\{Q\}\_\{\*\}\.\(64\)Every solution hasz=γz=\\gamma\. The bound is one\-sided and ranges over*all*simultaneously gain\-optimal policies\.
This is the discrete\-time minimization specialization of\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Theorem 1\]\. We check the change of convention and normalization below, and use that published theorem for existence\.
###### Proof\.
This is\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Theorem 1\]after reversing rewards\. We check its hypotheses and the two conventions that affect the statement\. Use unit holding times, so the source’s holding\-time matrix isHq=PqH\_\{q\}=P\_\{q\}, and rewardc~=−c\\widetilde\{c\}=\-c\. The finite state space, compact metric action spaces, and continuous data satisfy its assumptions\. Forη\(q\)=Pq∞cq\\eta\(q\)=P\_\{q\}^\{\\infty\}c^\{q\}, stochasticity andPqPq∞=Pq∞P\_\{q\}P\_\{q\}^\{\\infty\}=P\_\{q\}^\{\\infty\}give
η~\(q\)=−η\(q\),w~\(q\)=ZPq\(−cq\+Pqη\(q\)\)=−w\(q\)\.\\widetilde\{\\eta\}\(q\)=\-\\eta\(q\),\\qquad\\widetilde\{w\}\(q\)=Z\_\{P\_\{q\}\}\(\-c^\{q\}\+P\_\{q\}\\eta\(q\)\)=\-w\(q\)\.The transformed maximal\-gain vector is−γ\-\\gamma, and its simultaneously optimal policies are precisely𝒬∗\\mathcal\{Q\}\_\{\*\}\. Thus the source’s uniform upper bound on their canonical biases is the lower bound in equation[64](https://arxiv.org/html/2609.28792#A10.E64)\. The normalization is unchanged:Pq∞w\(q\)=0P\_\{q\}^\{\\infty\}w\(q\)=0\.
For the equations, substitutez~=−z\\widetilde\{z\}=\-zandh~=−h\\widetilde\{h\}=\-hin the maximizing system\. Its first equation becomes equation[62](https://arxiv.org/html/2609.28792#A10.E62)\. Its maximizing actions are exactly\{b:pi\(b\)⊤z=zi\}\\\{b:p\_\{i\}\(b\)^\{\\top\}z=z\_\{i\}\\\}, and its second equation becomes
−hi=maxb:pi\(b\)⊤z=zi\{−ci\(b\)\+pi\(b\)⊤z−pi\(b\)⊤h\}=zi−minb:pi\(b\)⊤z=zi\{ci\(b\)\+pi\(b\)⊤h\}\.\-h\_\{i\}=\\max\_\{b:\\,p\_\{i\}\(b\)^\{\\top\}z=z\_\{i\}\}\\\{\-c\_\{i\}\(b\)\+p\_\{i\}\(b\)^\{\\top\}z\-p\_\{i\}\(b\)^\{\\top\}h\\\}=z\_\{i\}\-\\min\_\{b:\\,p\_\{i\}\(b\)^\{\\top\}z=z\_\{i\}\}\\\{c\_\{i\}\(b\)\+p\_\{i\}\(b\)^\{\\top\}h\\\}\.This is equation[63](https://arxiv.org/html/2609.28792#A10.E63)\. The cited theorem therefore supplies both directions of the existence criterion and identifies every solution’s gain asγ\\gamma\. ∎
###### Corollary 2\.
For a deterministic stationary controllerπ\\pi, apply Theorem[12](https://arxiv.org/html/2609.28792#Thmtheorem12)withBi=𝒰i,π\(i\)B\_\{i\}=\\mathcal\{U\}\_\{i,\\pi\(i\)\},pi\(b\)=bp\_\{i\}\(b\)=b, andci\(b\)=ri,π\(i\)c\_\{i\}\(b\)=r\_\{i,\\pi\(i\)\}\. Its equations are exactly
giπ=minp∈𝒰i,π\(i\)p𝖳gπ,giπ\+hiπ=ri,π\(i\)\+minp∈Fi,π\(i\)\(gπ\)p𝖳hπ\.g\_\{i\}^\{\\pi\}=\\min\_\{p\\in\\mathcal\{U\}\_\{i,\\pi\(i\)\}\}p^\{\\mathsf\{T\}\}g^\{\\pi\},\\qquad g\_\{i\}^\{\\pi\}\+h\_\{i\}^\{\\pi\}=r\_\{i,\\pi\(i\)\}\+\\min\_\{p\\in F\_\{i,\\pi\(i\)\}\(g^\{\\pi\}\)\}p^\{\\mathsf\{T\}\}h^\{\\pi\}\.They are solvable precisely when a stationary nature kernel attains the complete worst\-gain vector and all such kernels have a common lower bound on their canonical biases\.
For a fixed stationary randomized policy, takeBi=∏a:π\(a∣i\)\>0𝒰iaB\_\{i\}=\\prod\_\{a:\\pi\(a\\mid i\)\>0\}\\mathcal\{U\}\_\{ia\}and setpi\(b\)=∑aπ\(a∣i\)bap\_\{i\}\(b\)=\\sum\_\{a\}\\pi\(a\\mid i\)b\_\{a\}andci\(b\)=∑aπ\(a∣i\)riac\_\{i\}\(b\)=\\sum\_\{a\}\\pi\(a\\mid i\)r\_\{ia\}\. This is again a compact continuous one\-player model\. Its gain and bias equations are the correspondingπ\\pi\-weighted sums of the actionwise gain and gain\-face minima\.
###### Proof\.
For deterministicπ\\pi, the substitution in the statement preserves the stationary matrices, rewards, gains, and canonical biases\. Theorem[12](https://arxiv.org/html/2609.28792#Thmtheorem12)therefore gives the asserted criterion directly\.
For randomizedπ\\pi, writep¯\(b\)=∑aπ\(a∣i\)ba\\bar\{p\}\(b\)=\\sum\_\{a\}\\pi\(a\\mid i\)b\_\{a\}on the compact product of its positive\-weight action row sets\. Rectangularity gives
minbp¯\(b\)⊤g=∑aπ\(a∣i\)mia\(g\)\.\\min\_\{b\}\\bar\{p\}\(b\)^\{\\top\}g=\\sum\_\{a\}\\pi\(a\\mid i\)m\_\{ia\}\(g\)\.The excess of a feasible tuple over this minimum is∑aπ\(a∣i\)\[ba⊤g−mia\(g\)\]\\sum\_\{a\}\\pi\(a\\mid i\)\[b\_\{a\}^\{\\top\}g\-m\_\{ia\}\(g\)\]\. Every summand is nonnegative, so the excess vanishes exactly whenba∈Fia\(g\)b\_\{a\}\\in F\_\{ia\}\(g\)for each positive\-weight action\. Its restricted bias minimum is consequently∑aπ\(a∣i\)minp∈Fia\(g\)p⊤h\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in F\_\{ia\}\(g\)\}p^\{\\top\}h\. This establishes both evaluation equations and the same canonical\-bias criterion\. Proposition[5](https://arxiv.org/html/2609.28792#Thmproposition5)identifies the tuple model with post\-action policy evaluation\. ∎
### J\.2Fixed points, bounded orbits, and barriers
Forz=T^zz=\\widehat\{T\}z, defineKz\(x\)=Lzx−zK\_\{z\}\(x\)=L\_\{z\}x\-z\. These maps are monotone, additively homogeneous, and nonexpansive in the supremum norm\. Writesp\(x\)=maxixi−minixi\\operatorname\{sp\}\(x\)=\\max\_\{i\}x\_\{i\}\-\\min\_\{i\}x\_\{i\}\.
###### Lemma 8\.
LetF:ℝn→ℝnF:\\mathbb\{R\}^\{n\}\\to\\mathbb\{R\}^\{n\}be monotone and satisfyF\(x\+c𝟏\)=F\(x\)\+c𝟏F\(x\+c\\mathbf\{1\}\)=F\(x\)\+c\\mathbf\{1\}\. The following are equivalent:
1. 1\.Fh=h\+λ𝟏Fh=h\+\\lambda\\mathbf\{1\}for somehhandλ\\lambda;
2. 2\.supN≥0sp\(FN0\)<∞\\sup\_\{N\\geq 0\}\\operatorname\{sp\}\(F^\{N\}0\)<\\infty;
3. 3\.there areℓ,u,λ\\ell,u,\\lambdawithℓ\+λ𝟏≤Fℓ\\ell\+\\lambda\\mathbf\{1\}\\leq F\\ellandFu≤u\+λ𝟏Fu\\leq u\+\\lambda\\mathbf\{1\}\.
Furthermore,FFhas a fixed point if and only if one of its orbits is bounded in the supremum norm\.
###### Proof\.
A monotone, additively homogeneous map is called topical\. It is nonexpansive in both the supremum norm and the span seminorm: applyFFtoy\+mini\(xi−yi\)𝟏≤x≤y\+maxi\(xi−yi\)𝟏y\+\\min\_\{i\}\(x\_\{i\}\-y\_\{i\}\)\\mathbf\{1\}\\leq x\\leq y\+\\max\_\{i\}\(x\_\{i\}\-y\_\{i\}\)\\mathbf\{1\}\. In particular,FFis continuous\. The equivalence \(1\)⇔\\Leftrightarrow\(2\) is the additive bounded\-orbit theorem of\[[15](https://arxiv.org/html/2609.28792#bib.bib74), Theorem 9\]\. Their Lemma 3, with scalar growth rate zero, gives the final fixed\-point assertion\. These results require precisely monotonicity and scalar additive homogeneity on finite\-dimensionalℝn\\mathbb\{R\}^\{n\}\.
We retain the short barrier argument because it is used below\. If \(1\) holds, chooseℓ=u=h\\ell=u=hin \(3\)\. Conversely, under \(3\), setG=F−λ𝟏G=F\-\\lambda\\mathbf\{1\}\. A scalar shift ofuupreservesGu≤uGu\\leq uand makesℓ≤u\\ell\\leq u\. Starting fromh0=ℓh\_\{0\}=\\ell, monotonicity gives
ℓ=h0≤h1:=Gh0≤h2:=Gh1≤⋯≤u\.\\ell=h\_\{0\}\\leq h\_\{1\}:=Gh\_\{0\}\\leq h\_\{2\}:=Gh\_\{1\}\\leq\\cdots\\leq u\.Indeed, the lower inequality propagates fromℓ≤Gℓ\\ell\\leq G\\ell, andhk≤uh\_\{k\}\\leq uimplieshk\+1≤Gu≤uh\_\{k\+1\}\\leq Gu\\leq u\. Coordinatewise convergence and continuity giveGh=hGh=h, hence \(1\)\. A bounded orbit from any starting vector is equivalent to one from zero because‖FNx−FN0‖∞≤‖x‖∞\\\|F^\{N\}x\-F^\{N\}0\\\|\_\{\\infty\}\\leq\\\|x\\\|\_\{\\infty\}\. ∎
Bounded span permits a scalar drift, whereas bounded supremum norm forces that drift to be zero\. This distinction is essential when the gain is prescribed\.
###### Theorem 13\.
For a prescribedg=T^gg=\\widehat\{T\}g, the following are equivalent:
1. 1\.there is a finitehhwithg\+h=Lghg\+h=L\_\{g\}h;
2. 2\.supN≥0‖KgN0‖∞<∞\\sup\_\{N\\geq 0\}\\\|K\_\{g\}^\{N\}0\\\|\_\{\\infty\}<\\infty;
3. 3\.there areℓ,u\\ell,uwithℓ≤Kgℓ\\ell\\leq K\_\{g\}\\ellandKgu≤uK\_\{g\}u\\leq u\.
With the gain unspecified, a finite gain\-bias pair exists if and only if somez=T^zz=\\widehat\{T\}zhas a span\-boundedKzK\_\{z\}\-orbit\. Equivalently, some suchKzK\_\{z\}has an additive eigenpairKzh=h\+λ𝟏K\_\{z\}h=h\+\\lambda\\mathbf\{1\}\. In that case the solution is\(g,h\)\(g,h\)withg=z\+λ𝟏g=z\+\\lambda\\mathbf\{1\}\.
###### Proof\.
The prescribed\-gain equation isKgh=hK\_\{g\}h=h\. The mapKgK\_\{g\}is monotone and additively homogeneous, so Lemma[8](https://arxiv.org/html/2609.28792#Thmlemma8)gives the equivalence with a bounded orbit\. Its barrier argument atλ=0\\lambda=0gives the third equivalent condition\.
For an unspecified gain, supposez=T^zz=\\widehat\{T\}zandKzh=h\+λ𝟏K\_\{z\}h=h\+\\lambda\\mathbf\{1\}\. Probability rows satisfyp⊤𝟏=1p^\{\\top\}\\mathbf\{1\}=1, so scalar translations preserve the active sets:
T^\(z\+λ𝟏\)=z\+λ𝟏,Az\+λ𝟏\(i\)=Az\(i\),Fia\(z\+λ𝟏\)=Fia\(z\)\.\\widehat\{T\}\(z\+\\lambda\\mathbf\{1\}\)=z\+\\lambda\\mathbf\{1\},\\quad A\_\{z\+\\lambda\\mathbf\{1\}\}\(i\)=A\_\{z\}\(i\),\\quad F\_\{ia\}\(z\+\\lambda\\mathbf\{1\}\)=F\_\{ia\}\(z\)\.ThusLz\+λ𝟏=LzL\_\{z\+\\lambda\\mathbf\{1\}\}=L\_\{z\}\. Forg=z\+λ𝟏g=z\+\\lambda\\mathbf\{1\},Kgh=Kzh−λ𝟏=hK\_\{g\}h=K\_\{z\}h\-\\lambda\\mathbf\{1\}=h, giving a finite Bellman pair\. Conversely a finite pair supplies the eigenpairKgh=hK\_\{g\}h=hwith scalar eigenvalue zero\. Lemma[8](https://arxiv.org/html/2609.28792#Thmlemma8)identifies additive eigenpairs with span\-bounded orbits\. Only scalar translations are used here; vector centering need not preserve the gain faces or commute with iteration\. ∎
### J\.3Stationary tangent saddle and canonical\-bias envelope
Fixg=T^gg=\\widehat\{T\}g, putcia=ria−gic\_\{ia\}=r\_\{ia\}\-g\_\{i\}, and define
Πg=∏iAg\(i\),𝒬g=∏i∏a∈Ag\(i\)Fia\(g\),𝒬gπ=∏iFi,π\(i\)\(g\)\.\\Pi\_\{g\}=\\prod\_\{i\}A\_\{g\}\(i\),\\qquad\\mathcal\{Q\}\_\{g\}=\\prod\_\{i\}\\prod\_\{a\\in A\_\{g\}\(i\)\}F\_\{ia\}\(g\),\\qquad\\mathcal\{Q\}\_\{g\}^\{\\pi\}=\\prod\_\{i\}F\_\{i,\\pi\(i\)\}\(g\)\.A full planq∈𝒬gq\\in\\mathcal\{Q\}\_\{g\}specifies a row for every active action\. Forπ∈Πg\\pi\\in\\Pi\_\{g\}, the matrixPπ,qP\_\{\\pi,q\}has rowqi,π\(i\)q\_\{i,\\pi\(i\)\}whenqqis a full plan and rowqiq\_\{i\}whenq∈𝒬gπq\\in\\mathcal\{Q\}\_\{g\}^\{\\pi\}is a restricted reply\. Write
ηπ,q=Pπ,q∞cπ,wπ,q=ZPπ,qcπwhenηπ,q=0\.\\eta^\{\\pi,q\}=P\_\{\\pi,q\}^\{\\infty\}c^\{\\pi\},\\qquad w^\{\\pi,q\}=Z\_\{P\_\{\\pi,q\}\}c^\{\\pi\}\\quad\\text\{when \}\\eta^\{\\pi,q\}=0\.These gains use the centered rewardscc, and all vector inequalities below are componentwise\.
###### Theorem 14\.
The equationKgh=hK\_\{g\}h=hhas a finite solution if and only if there existπ¯∈Πg\\bar\{\\pi\}\\in\\Pi\_\{g\}, a full planq¯∈𝒬g\\bar\{q\}\\in\\mathcal\{Q\}\_\{g\}, andB<∞B<\\inftysuch that
ηπ¯,q\\displaystyle\\eta^\{\\bar\{\\pi\},q\}≥0\\displaystyle\\geq 0for everyq∈𝒬gπ¯,\\displaystyle\\text\{for every \}q\\in\\mathcal\{Q\}\_\{g\}^\{\\bar\{\\pi\}\},\(65\)ηπ,q¯\\displaystyle\\eta^\{\\pi,\\bar\{q\}\}≤0\\displaystyle\\leq 0for everyπ∈Πg,\\displaystyle\\text\{for every \}\\pi\\in\\Pi\_\{g\},\(66\)wπ¯,q\\displaystyle w^\{\\bar\{\\pi\},q\}≥−B𝟏\\displaystyle\\geq\-B\\mathbf\{1\}for everyq∈𝒬gπ¯withηπ¯,q=0\.\\displaystyle\\text\{for every \}q\\in\\mathcal\{Q\}\_\{g\}^\{\\bar\{\\pi\}\}\\text\{ with \}\\eta^\{\\bar\{\\pi\},q\}=0\.\(67\)The first and third conditions fix the controllerπ¯\\bar\{\\pi\}and vary nature’s replyqq\. The second fixes the full nature planq¯\\bar\{q\}and varies the controllerπ\\pi\. Only one securing controller is required, and its canonical\-bias bound must cover all zero\-gain replies\.
###### Proof\.
Necessity: select strategies from a common bias\.SupposeKgh=hK\_\{g\}h=h\. Chooseπ¯\(i\)\\bar\{\\pi\}\(i\)attaining its action maximum, and chooseq¯ia\\bar\{q\}\_\{ia\}attaining its row minimum for every active action\. Finiteness and compactness ensure these selections exist\. For every restricted replyq∈𝒬gπ¯q\\in\\mathcal\{Q\}\_\{g\}^\{\\bar\{\\pi\}\}and everyπ∈Πg\\pi\\in\\Pi\_\{g\}, respectively,
cπ¯\+Pπ¯,qh−h≥0,cπ\+Pπ,q¯h−h≤0\.c^\{\\bar\{\\pi\}\}\+P\_\{\\bar\{\\pi\},q\}h\-h\\geq 0,\\qquad c^\{\\pi\}\+P\_\{\\pi,\\bar\{q\}\}h\-h\\leq 0\.Indeed, every row of the selected controller action is at least that action’s minimizing valuehih\_\{i\}\. For the second inequality, the selected row at each active action realizes an action minimum no greater than the maximumhih\_\{i\}\. Multiplication by the corresponding nonnegative Cesàro projector eliminates the terms\(P−I\)h\(P\-I\)hand proves equation[65](https://arxiv.org/html/2609.28792#A10.E65)\-equation[66](https://arxiv.org/html/2609.28792#A10.E66)\.
Fix any replyqqwithηπ¯,q=0\\eta^\{\\bar\{\\pi\},q\}=0\. SetP=Pπ¯,qP=P\_\{\\bar\{\\pi\},q\},Γ=P∞\\Gamma=P^\{\\infty\}, andd=cπ¯\+Ph−hd=c^\{\\bar\{\\pi\}\}\+Ph\-h\. Thend≥0d\\geq 0andΓd=0\\Gamma d=0\. UsingZP\(I−P\)=I−ΓZ\_\{P\}\(I\-P\)=I\-\\Gammagives
wπ¯,q=ZP\(\(I−P\)h\+d\)=\(I−Γ\)h\+ZPd\.w^\{\\bar\{\\pi\},q\}=Z\_\{P\}\\bigl\(\(I\-P\)h\+d\\bigr\)=\(I\-\\Gamma\)h\+Z\_\{P\}d\.Lemma[2](https://arxiv.org/html/2609.28792#Thmlemma2)givesZPd≥0Z\_\{P\}d\\geq 0\. SinceΓ\\Gammais stochastic,hi−\(Γh\)i≥−sp\(h\)h\_\{i\}\-\(\\Gamma h\)\_\{i\}\\geq\-\\operatorname\{sp\}\(h\)for every state\. Thus equation[67](https://arxiv.org/html/2609.28792#A10.E67)holds withB=sp\(h\)B=\\operatorname\{sp\}\(h\), uniformly over all zero\-gain replies with their canonical normalizations\.
Sufficiency: construct a lower barrier\.Assume the three stationary conditions\. Restrictq¯\\bar\{q\}to the actions ofπ¯\\bar\{\\pi\}\. Applying both gain inequalities to this pair givesηπ¯,q¯=0\\eta^\{\\bar\{\\pi\},\\bar\{q\}\}=0\. Withπ¯\\bar\{\\pi\}fixed, nature’s compact\-action MDP has rewardscπ¯c^\{\\bar\{\\pi\}\}and row setsFi,π¯\(i\)\(g\)F\_\{i,\\bar\{\\pi\}\(i\)\}\(g\)\. All its stationary gains are nonnegative, and the restrictedq¯\\bar\{q\}attains zero at every state\. Its simultaneously optimal stationary policies are therefore exactly the replies withηπ¯,q=0\\eta^\{\\bar\{\\pi\},q\}=0\. Condition equation[67](https://arxiv.org/html/2609.28792#A10.E67)is the canonical lower bound required by Theorem[12](https://arxiv.org/html/2609.28792#Thmtheorem12)\. That theorem gives a finite vectorℓ\\ellsatisfying
ℓi=minp∈Fi,π¯\(i\)\(g\)\{ci,π¯\(i\)\+p⊤ℓ\}\.\\ell\_\{i\}=\\min\_\{p\\in F\_\{i,\\bar\{\\pi\}\(i\)\}\(g\)\}\\\{c\_\{i,\\bar\{\\pi\}\(i\)\}\+p^\{\\top\}\\ell\\\}\.Here the one\-player gain is zero, so its gain restriction retains every available row\. The outer maximum inKgK\_\{g\}can chooseπ¯\(i\)\\bar\{\\pi\}\(i\), and henceℓ≤Kgℓ\\ell\\leq K\_\{g\}\\ell\.
Construct an upper barrier\.With the full planq¯\\bar\{q\}fixed, the controller has finite action setsAg\(i\)A\_\{g\}\(i\)and rowsq¯ia\\bar\{q\}\_\{ia\}\. Every stationary policy has centered gain at most zero, whileπ¯\\bar\{\\pi\}attains zero from all states\. The finite\-action multichain optimality equations therefore admit a finite bias\[[48](https://arxiv.org/html/2609.28792#bib.bib32), Chapters 8\-9\]\. Equivalently, apply Theorem[12](https://arxiv.org/html/2609.28792#Thmtheorem12)to rewards−c\-c: the optimal gain is zero and the canonical\-bias bound is automatic because there are finitely many deterministic policies\. Reversing the resulting bias gives
ui=maxa∈Ag\(i\)\{cia\+q¯ia⊤u\},\(Kgu\)i≤maxa∈Ag\(i\)\{cia\+q¯ia⊤u\}=ui\.u\_\{i\}=\\max\_\{a\\in A\_\{g\}\(i\)\}\\\{c\_\{ia\}\+\\bar\{q\}\_\{ia\}^\{\\top\}u\\\},\\qquad\(K\_\{g\}u\)\_\{i\}\\leq\\max\_\{a\\in A\_\{g\}\(i\)\}\\\{c\_\{ia\}\+\\bar\{q\}\_\{ia\}^\{\\top\}u\\\}=u\_\{i\}\.The inequality uses feasibility ofq¯ia\\bar\{q\}\_\{ia\}in each row minimum\. It requires a full plan covering every active controller action\.
Construct a common fixed point\.Lets=maxi\(ℓi−ui\)s=\\max\_\{i\}\(\\ell\_\{i\}\-u\_\{i\}\)andu~=u\+s𝟏\\widetilde\{u\}=u\+s\\mathbf\{1\}\. Scalar additive homogeneity givesKgu~≤u~K\_\{g\}\\widetilde\{u\}\\leq\\widetilde\{u\}, andℓ≤u~\\ell\\leq\\widetilde\{u\}\. Starting fromh0=ℓh\_\{0\}=\\elland iteratinghk\+1=Kghkh\_\{k\+1\}=K\_\{g\}h\_\{k\}, monotonicity gives
ℓ=h0≤h1≤⋯≤hk≤u~\.\\ell=h\_\{0\}\\leq h\_\{1\}\\leq\\cdots\\leq h\_\{k\}\\leq\\widetilde\{u\}\.Each coordinate converges to a finite limit\. Continuity ofKgK\_\{g\}therefore gives a finite fixed pointhh\. The securing pair\(π¯,q¯\)\(\\bar\{\\pi\},\\bar\{q\}\)need not be compatible with this same bias\. The common\-bias selectors are obtained by taking maximizing actions and minimizing rows at the resultinghh\. ∎
## Appendix KGeometry and structure of Bellman certificates
Section[5](https://arxiv.org/html/2609.28792#S5)characterizes finite Bellman solvability\. Here we study the certificates themselves: which controller\-nature pairs share a bias, how small its span can be, and how all compatible biases can be parameterized\. The same two families of linear inequalities answer all three questions\. We first establish the geometric characterization and its minimum\-span formula, then describe the remaining freedom through recurrent\-class offsets\. We finish with a sufficient condition for finite span and a finite linear\-program formulation for listed polyhedral gain faces\.
### K\.1Common biases and the mixed\-flow characterization
Throughout this section, fix a recession\-fixed vectorg=T^gg=\\widehat\{T\}gand putcia=ria−gic\_\{ia\}=r\_\{ia\}\-g\_\{i\}\. The state set hasn≥1n\\geq 1elements, the action sets are finite and nonempty, and each ambiguity row set is nonempty and compact\. RecallΠg=∏iAg\(i\)\\Pi\_\{g\}=\\prod\_\{i\}A\_\{g\}\(i\)and𝒬g=∏i∏a∈Ag\(i\)Fia\(g\)\\mathcal\{Q\}\_\{g\}=\\prod\_\{i\}\\prod\_\{a\\in A\_\{g\}\(i\)\}F\_\{ia\}\(g\)\. The gain\-active sets are nonempty, everyFia\(g\)F\_\{ia\}\(g\)is compact, andp⊤g=gip^\{\\top\}g=g\_\{i\}fora∈Ag\(i\)a\\in A\_\{g\}\(i\)andp∈Fia\(g\)p\\in F\_\{ia\}\(g\)\. A pair\(π,q\)∈Πg×𝒬g\(\\pi,q\)\\in\\Pi\_\{g\}\\times\\mathcal\{Q\}\_\{g\}chooses one active controller action per state and one nature row for*every*active state\-action pair\. Specifying onlyqi,π\(i\)q\_\{i,\\pi\(i\)\}would leave controller deviations uncontrolled\. All infima over selector pairs below range over this product\.
Defineℋπq\\mathcal\{H\}\_\{\\pi q\}as the set ofh∈ℝnh\\in\\mathbb\{R\}^\{n\}satisfying
hi\\displaystyle h\_\{i\}≤ci,π\(i\)\+p⊤h\\displaystyle\\leq c\_\{i,\\pi\(i\)\}\+p^\{\\top\}hfor everyi,p∈Fi,π\(i\)\(g\),\\displaystyle\\text\{for every \}i,\\ p\\in F\_\{i,\\pi\(i\)\}\(g\),cia\+qia⊤h\\displaystyle c\_\{ia\}\+q\_\{ia\}^\{\\top\}h≤hi\\displaystyle\\leq h\_\{i\}for everyi,a∈Ag\(i\)\.\\displaystyle\\text\{for every \}i,\\ a\\in A\_\{g\}\(i\)\.\(68\)The first family secures the controller’s lower Bellman bound against every row of its chosen action\. The second secures nature’s upper bound against every active action\. A common bias satisfies both families with the same vector\. At the selected action and row they forcehi=ci,π\(i\)\+qi,π\(i\)⊤hh\_\{i\}=c\_\{i,\\pi\(i\)\}\+q\_\{i,\\pi\(i\)\}^\{\\top\}h\. Every common bias solvesKgh=hK\_\{g\}h=h, and every Bellman bias admits a compatible pair, as established below\. Verification then identifiesg=g⋆g=g^\{\\star\}and supplies the stationary payoff guarantees\.
Rewriting equation[68](https://arxiv.org/html/2609.28792#A11.E68)as linear inequalities gives the compact generator set, witheie\_\{i\}theiith coordinate vector,
Gπq=\\displaystyle G\_\{\\pi q\}=\{\}\{\(ei−p,ci,π\(i\)\):i=1,…,n,p∈Fi,π\(i\)\(g\)\}\\displaystyle\\\{\(e\_\{i\}\-p,c\_\{i,\\pi\(i\)\}\):i=1,\\ldots,n,\\ p\\in F\_\{i,\\pi\(i\)\}\(g\)\\\}∪\{\(qia−ei,−cia\):i=1,…,n,a∈Ag\(i\)\},\\displaystyle\\quad\\cup\\\{\(q\_\{ia\}\-e\_\{i\},\-c\_\{ia\}\):i=1,\\ldots,n,\\ a\\in A\_\{g\}\(i\)\\\},\(69\)andCπq=cone\(Gπq\)C\_\{\\pi q\}=\\operatorname\{cone\}\(G\_\{\\pi q\}\), consisting of finite nonnegative combinations, including zero\. A point\(f,b\)\(f,b\)in this cone combines the original constraints intof⊤h≤bf^\{\\top\}h\\leq b\. The vectorffmeasures their remaining statewise imbalance and satisfies𝟏⊤f=g⊤f=0\\mathbf\{1\}^\{\\top\}f=g^\{\\top\}f=0\. The scalarbbis the corresponding signed combination of centered rewards\. These combinations use both players’ inequalities\.
Define
β\(π,q\)=sup\(f,b\)∈Cπq\[−b\]\+‖f‖1\.\\beta\(\\pi,q\)=\\sup\_\{\(f,b\)\\in C\_\{\\pi q\}\}\\frac\{\[\-b\]\_\{\+\}\}\{\\\|f\\\|\_\{1\}\}\.\(70\)Here\[x\]\+=max\{x,0\}\[x\]\_\{\+\}=\\max\\\{x,0\\\}\. A zero denominator gives\+∞\+\\inftywhenb<0b<0, and zero whenb≥0b\\geq 0\. Equivalently,
ℋπq=\{h∈ℝn:f⊤h≤bfor all\(f,b\)∈Cπq\}\.\\mathcal\{H\}\_\{\\pi q\}=\\\{h\\in\\mathbb\{R\}^\{n\}:f^\{\\top\}h\\leq b\\text\{ for all \}\(f,b\)\\in C\_\{\\pi q\}\\\}\.\(71\)For a fixed pair this is a semi\-infinite linear feasibility problem\. Related feasibility criteria based on projected inequalities appear in\[[5](https://arxiv.org/html/2609.28792#bib.bib73), Theorem 2\.14\]\. Here the constraints come from the two players’ Bellman comparisons, and the ratio in equation[70](https://arxiv.org/html/2609.28792#A11.E70)determines the exact minimum compatible bias span\. The general separation principle is standard convex geometry\. We give its short form below because the cone need not be closed\. The Bellman\-specific content is the use of*both*deviation families and the exact compatible\-span formula\.
###### Theorem 15\.
For fixed\(π,q\)\(\\pi,q\)the following are equivalent:
ℋπq≠∅⟺\(0,−1\)∉cl\(Cπq\)⟺β\(π,q\)<∞\.\\mathcal\{H\}\_\{\\pi q\}\\neq\\varnothing\\quad\\Longleftrightarrow\\quad\(0,\-1\)\\notin\\operatorname\{cl\}\(C\_\{\\pi q\}\)\\quad\\Longleftrightarrow\\quad\\beta\(\\pi,q\)<\\infty\.\(72\)If these conditions hold, the minimum span is attained and equals
minh∈ℋπqsp\(h\)=2β\(π,q\)\.\\min\_\{h\\in\\mathcal\{H\}\_\{\\pi q\}\}\\operatorname\{sp\}\(h\)=2\\beta\(\\pi,q\)\.\(73\)With the convention that the infimum of an empty set is\+∞\+\\infty, the full Bellman problem satisfies
Bg:=inf\{sp\(h\):Kgh=h\}=2infπ,qβ\(π,q\)\.B\_\{g\}:=\\inf\\\{\\operatorname\{sp\}\(h\):K\_\{g\}h=h\\\}=2\\inf\_\{\\pi,q\}\\beta\(\\pi,q\)\.\(74\)IfBg<\+∞B\_\{g\}<\+\\infty, both infima in equation[74](https://arxiv.org/html/2609.28792#A11.E74)are minima\. In particular, a finite Bellman bias exists if and only if one pair of active selectors has finite mixed\-flow ratio\. Infeasibility for a fixed pair has the sparse limiting witnesses of Proposition[6](https://arxiv.org/html/2609.28792#Thmproposition6)\.
Interpretation and relation to solvability\.A point\(0,−1\)\(0,\-1\)in the cone expresses an inconsistent combination0≤−10\\leq\-1\. Its presence only in the closure is equally obstructive: constraintsfk⊤h≤−1f\_\{k\}^\{\\top\}h\\leq\-1withfk→0f\_\{k\}\\to 0cannot hold for a finitehh\. Compact generators can have a nonclosed cone, so the closure retains limiting obstructions created by vanishing transition probabilities\. Excluding them gives a common finite bias\. The exact factor two in the span formula follows from constant\-shift invariance: midpoint centering givesinft‖h−t𝟏‖∞=sp\(h\)/2\\inf\_\{t\}\\\|h\-t\\mathbf\{1\}\\\|\_\{\\infty\}=\\operatorname\{sp\}\(h\)/2\.
For a prescribed pair the theorem characterizes compatibility with a common bias\. Taking the union over pairs characterizes the same fixed\-point existence as Theorem[5](https://arxiv.org/html/2609.28792#Thmtheorem5), while also minimizing the required span\. An incompatible pair can coexist with another pair supporting a Bellman solution, even when both are average optimal\. Example[3](https://arxiv.org/html/2609.28792#Thmexample3)exhibits this distinction\. Full solvability withggunknown requires the condition for someg=T^gg=\\widehat\{T\}g; verification identifies every successful gain withg⋆g^\{\\star\}\.
### K\.2Proof of the mixed\-flow theorem
We first prove the elementary separation statement used below\.
###### Lemma 9\.
LetC⊂ℝn×ℝC\\subset\\mathbb\{R\}^\{n\}\\times\\mathbb\{R\}be a convex cone containing zero\. Then\(0,−1\)∉cl\(C\)\(0,\-1\)\\notin\\operatorname\{cl\}\(C\)if and only if there existsh∈ℝnh\\in\\mathbb\{R\}^\{n\}such thata⊤h≤ba^\{\\top\}h\\leq bfor every\(a,b\)∈C\(a,b\)\\in C\.
###### Proof\.
If a potential exists, its closed halfspace\{\(a,b\):b−a⊤h≥0\}\\\{\(a,b\):b\-a^\{\\top\}h\\geq 0\\\}containscl\(C\)\\operatorname\{cl\}\(C\)and excludes\(0,−1\)\(0,\-1\)\. Conversely, strictly separate\(0,−1\)\(0,\-1\)from the closed convex conecl\(C\)\\operatorname\{cl\}\(C\)\. Since the set is a cone containing zero, the separating functional can be writtenu⊤a\+λb≥0u^\{\\top\}a\+\\lambda b\\geq 0on the cone and−λ<0\-\\lambda<0at\(0,−1\)\(0,\-1\)\. Henceλ\>0\\lambda\>0, andh=−u/λh=\-u/\\lambdasatisfies all the required inequalities\. ∎
###### Lemma 10\.
LetC⊂ℝn×ℝC\\subset\\mathbb\{R\}^\{n\}\\times\\mathbb\{R\}be a convex cone containing zero and letB≥0B\\geq 0\. There existshhwith‖h‖∞≤B\\\|h\\\|\_\{\\infty\}\\leq Banda⊤h≤ba^\{\\top\}h\\leq bonCCif and only if
b≥−B‖a‖1for every\(a,b\)∈C\.b\\geq\-B\\\|a\\\|\_\{1\}\\qquad\\text\{for every \}\(a,b\)\\in C\.\(75\)
###### Proof\.
Necessity follows fromb≥a⊤h≥−B‖a‖1b\\geq a^\{\\top\}h\\geq\-B\\\|a\\\|\_\{1\}\. For sufficiency, letEB=\{\(a,b\):b≥B‖a‖1\}E\_\{B\}=\\\{\(a,b\):b\\geq B\\\|a\\\|\_\{1\}\\\}andD=C\+EBD=C\+E\_\{B\}\. The coneEBE\_\{B\}consists precisely of the inequalities valid throughout the box\[−B,B\]n\[\-B,B\]^\{n\}\. Thus addingEBE\_\{B\}will force the separating potential into that box\. For\(a,b\)=\(a1,b1\)\+\(a2,b2\)∈D\(a,b\)=\(a\_\{1\},b\_\{1\}\)\+\(a\_\{2\},b\_\{2\}\)\\in D, the assumed bound and the triangle inequality give
b≥−B‖a1‖1\+B‖a2‖1≥−B‖a1\+a2‖1=−B‖a‖1\.b\\geq\-B\\\|a\_\{1\}\\\|\_\{1\}\+B\\\|a\_\{2\}\\\|\_\{1\}\\geq\-B\\\|a\_\{1\}\+a\_\{2\}\\\|\_\{1\}=\-B\\\|a\\\|\_\{1\}\.This bound extends tocl\(D\)\\operatorname\{cl\}\(D\)by continuity and excludes\(0,−1\)\(0,\-1\)\. Lemma[9](https://arxiv.org/html/2609.28792#Thmlemma9)supplies a potential valid onDD\. BecauseC⊆DC\\subseteq D, it satisfies the original constraints\. Applying it to\(ei,B\),\(−ei,B\)∈EB⊆D\(e\_\{i\},B\),\(\-e\_\{i\},B\)\\in E\_\{B\}\\subseteq Dgives\|hi\|≤B\|h\_\{i\}\|\\leq Bfor every coordinate\. The argument includesB=0B=0\. ∎
###### Proof of Theorem[15](https://arxiv.org/html/2609.28792#Thmtheorem15)\.
We first identify the fixed points represented by the mixed inequalities\. Separation then gives fixed\-pair feasibility and the exact span\. Finally, a compactness argument attains the optimum over selectors\.
Step 1: turn the mixed inequalities into the full Bellman equation\.Fix active selectors\(π,q\)\(\\pi,q\)\. Ifh∈ℋπqh\\in\\mathcal\{H\}\_\{\\pi q\}, the lower generator indexed by stateiiand rowp∈Fi,π\(i\)\(g\)p\\in F\_\{i,\\pi\(i\)\}\(g\)gives
\(ei−p\)⊤h≤ci,π\(i\)⟺hi≤ci,π\(i\)\+p⊤h\.\(e\_\{i\}\-p\)^\{\\top\}h\\leq c\_\{i,\\pi\(i\)\}\\quad\\Longleftrightarrow\\quad h\_\{i\}\\leq c\_\{i,\\pi\(i\)\}\+p^\{\\top\}h\.It holds for every row of the chosen action, so it holds for their minimum\. Allowing the controller to maximize over all active actions then gives
hi≤minp∈Fi,π\(i\)\(g\)\(ci,π\(i\)\+p⊤h\)≤\(Kgh\)i\.h\_\{i\}\\leq\\min\_\{p\\in F\_\{i,\\pi\(i\)\}\(g\)\}\(c\_\{i,\\pi\(i\)\}\+p^\{\\top\}h\)\\leq\(K\_\{g\}h\)\_\{i\}\.For each active action, the corresponding upper generator gives
\(qia−ei\)⊤h≤−cia⟺cia\+qia⊤h≤hi\.\(q\_\{ia\}\-e\_\{i\}\)^\{\\top\}h\\leq\-c\_\{ia\}\\quad\\Longleftrightarrow\\quad c\_\{ia\}\+q\_\{ia\}^\{\\top\}h\\leq h\_\{i\}\.The face minimum is no greater than its value at the feasibleqiaq\_\{ia\}\. Thus
\(Kgh\)i=maxa∈Ag\(i\)minp∈Fia\(g\)\(cia\+p⊤h\)≤maxa∈Ag\(i\)\(cia\+qia⊤h\)≤hi\.\(K\_\{g\}h\)\_\{i\}=\\max\_\{a\\in A\_\{g\}\(i\)\}\\min\_\{p\\in F\_\{ia\}\(g\)\}\(c\_\{ia\}\+p^\{\\top\}h\)\\leq\\max\_\{a\\in A\_\{g\}\(i\)\}\(c\_\{ia\}\+q\_\{ia\}^\{\\top\}h\)\\leq h\_\{i\}\.Combining the lower and upper inequalities provesKgh=hK\_\{g\}h=h\.
Conversely, supposeKgh=hK\_\{g\}h=h\. At each state choose a maximizingπ\(i\)∈Ag\(i\)\\pi\(i\)\\in A\_\{g\}\(i\); for every active action choose a minimizingqia∈Fia\(g\)q\_\{ia\}\\in F\_\{ia\}\(g\)forp⊤hp^\{\\top\}h\. Finiteness and compactness ensure these choices exist\. The chosen action has minimumhih\_\{i\}, so every row of that action satisfies the lower inequality\. Each action minimum is at most the maximumhih\_\{i\}, so its selected minimizing row satisfies the upper inequality\. Hence
Fix\(Kg\)=⋃π,qℋπq\.\\operatorname\{Fix\}\(K\_\{g\}\)=\\bigcup\_\{\\pi,q\}\\mathcal\{H\}\_\{\\pi q\}\.\(76\)This set equality requires both families of inequalities to use the same vectorhh\.
Step 2: apply separation to the mixed cone\.An inequalitya⊤h≤ba^\{\\top\}h\\leq bholding on the generators holds on every finite nonnegative combination: if\(a,b\)=∑j=1mλj\(aj,bj\)\(a,b\)=\\sum\_\{j=1\}^\{m\}\\lambda\_\{j\}\(a\_\{j\},b\_\{j\}\)withλj≥0\\lambda\_\{j\}\\geq 0, then
a⊤h=∑jλjaj⊤h≤∑jλjbj=b\.a^\{\\top\}h=\\sum\_\{j\}\\lambda\_\{j\}a\_\{j\}^\{\\top\}h\\leq\\sum\_\{j\}\\lambda\_\{j\}b\_\{j\}=b\.Conversely every generator belongs to the cone\. Thereforeℋπq\\mathcal\{H\}\_\{\\pi q\}is exactly the potential set forCπqC\_\{\\pi q\}\. Lemma[9](https://arxiv.org/html/2609.28792#Thmlemma9)yields
ℋπq≠∅⟺\(0,−1\)∉cl\(Cπq\)\.\\mathcal\{H\}\_\{\\pi q\}\\neq\\varnothing\\quad\\Longleftrightarrow\\quad\(0,\-1\)\\notin\\operatorname\{cl\}\(C\_\{\\pi q\}\)\.The closure is necessary because a continuous potential inequality also holds at limits of cone points\.
Step 3: derive the sharp lower bound on every feasible span\.Takeh∈ℋπqh\\in\\mathcal\{H\}\_\{\\pi q\}, letMh=maxihiM\_\{h\}=\\max\_\{i\}h\_\{i\},mh=minihim\_\{h\}=\\min\_\{i\}h\_\{i\}, and seth′=h−\(Mh\+mh\)𝟏/2h^\{\\prime\}=h\-\(M\_\{h\}\+m\_\{h\}\)\\mathbf\{1\}/2\. Every generator flow has zero sum and so does every conic combination\. Consequentlya⊤h′=a⊤h≤ba^\{\\top\}h^\{\\prime\}=a^\{\\top\}h\\leq bon the cone\. The largest and smallest entries ofh′h^\{\\prime\}are\(Mh−mh\)/2\(M\_\{h\}\-m\_\{h\}\)/2and−\(Mh−mh\)/2\-\(M\_\{h\}\-m\_\{h\}\)/2, respectively\. Hence
‖h′‖∞=sp\(h\)2\.\\\|h^\{\\prime\}\\\|\_\{\\infty\}=\\frac\{\\operatorname\{sp\}\(h\)\}\{2\}\.The norm bound in Lemma[10](https://arxiv.org/html/2609.28792#Thmlemma10), or directly Hölder’s inequality, givesb≥−\(sp\(h\)/2\)‖a‖1b\\geq\-\(\\operatorname\{sp\}\(h\)/2\)\\\|a\\\|\_\{1\}for every cone point\. In particular, ifa=0a=0, feasibility forcesb≥0b\\geq 0\. Ifa≠0a\\neq 0andb<0b<0, division by‖a‖1\>0\\\|a\\\|\_\{1\}\>0gives\(−b\)/‖a‖1≤sp\(h\)/2\(\-b\)/\\\|a\\\|\_\{1\}\\leq\\operatorname\{sp\}\(h\)/2\. Points withb≥0b\\geq 0contribute zero to the ratio\. Taking the supremum yields
2β\(π,q\)≤sp\(h\)\.2\\beta\(\\pi,q\)\\leq\\operatorname\{sp\}\(h\)\.Thus every feasible potential impliesβ\(π,q\)<∞\\beta\(\\pi,q\)<\\infty\.
Step 4: attain the lower span bound when the ratio is finite\.Supposeβ=β\(π,q\)<∞\\beta=\\beta\(\\pi,q\)<\\infty\. For a cone point witha≠0a\\neq 0andb<0b<0, the definition givesb≥−β‖a‖1b\\geq\-\\beta\\\|a\\\|\_\{1\}\. Forb≥0b\\geq 0the same inequality holds since its right side is nonpositive\. Fora=0a=0, finiteness ofβ\\betarules outb<0b<0by the stated convention, so again the inequality holds\. Therefore
b≥−β‖a‖1\(\(a,b\)∈Cπq\)\.b\\geq\-\\beta\\\|a\\\|\_\{1\}\\qquad\(\(a,b\)\\in C\_\{\\pi q\}\)\.Lemma[10](https://arxiv.org/html/2609.28792#Thmlemma10)supplies a feasiblehhwith‖h‖∞≤β\\\|h\\\|\_\{\\infty\}\\leq\\beta\. Combine this with Step 3:
2β≤sp\(h\)≤2‖h‖∞≤2β\.2\\beta\\leq\\operatorname\{sp\}\(h\)\\leq 2\\\|h\\\|\_\{\\infty\}\\leq 2\\beta\.Equality holds throughout\. This proves feasibility, the exact minimum span, and attainment, including the caseβ=0\\beta=0\. Together with Step 2 it proves the three\-way alternative\.
Step 5: optimize over selectors without assuming continuity of their ratios\.The union identity equation[76](https://arxiv.org/html/2609.28792#A11.E76)and the fixed\-pair span formula give
Bg\\displaystyle B\_\{g\}=infKgh=hsp\(h\)=infπ,qinfh∈ℋπqsp\(h\)=2infπ,qβ\(π,q\)\.\\displaystyle=\\inf\_\{K\_\{g\}h=h\}\\operatorname\{sp\}\(h\)=\\inf\_\{\\pi,q\}\\inf\_\{h\\in\\mathcal\{H\}\_\{\\pi q\}\}\\operatorname\{sp\}\(h\)=2\\inf\_\{\\pi,q\}\\beta\(\\pi,q\)\.The formula includes\+∞\+\\inftybecause an empty fixed\-pair region has infinite ratio and contributes an infinite infimum\. SupposeBg<∞B\_\{g\}<\\infty\. Choose fixed pointshkh^\{k\}withBg≤sp\(hk\)≤Bg\+1/kB\_\{g\}\\leq\\operatorname\{sp\}\(h^\{k\}\)\\leq B\_\{g\}\+1/k\. Replace each byhk−minihik𝟏h^\{k\}\-\\min\_\{i\}h\_\{i\}^\{k\}\\mathbf\{1\}\. Additive homogeneity ofKgK\_\{g\}preserves its fixed\-point equation, and now
0≤hik≤Bg\+1for everyi,k\.0\\leq h\_\{i\}^\{k\}\\leq B\_\{g\}\+1\\qquad\\text\{for every \}i,k\.A subsequence converges to a finiteh∗h^\{\*\}\. Continuity givesKgh∗=h∗K\_\{g\}h^\{\*\}=h^\{\*\}, and continuity of the finite maximum and minimum givessp\(h∗\)=Bg\\operatorname\{sp\}\(h^\{\*\}\)=B\_\{g\}\. Choose\(π∗,q∗\)\(\\pi^\{\*\},q^\{\*\}\)from this fixed point as in Step 1\. Then
Bg≤minh∈ℋπ∗q∗sp\(h\)=2β\(π∗,q∗\)≤sp\(h∗\)=Bg\.B\_\{g\}\\leq\\min\_\{h\\in\\mathcal\{H\}\_\{\\pi^\{\*\}q^\{\*\}\}\}\\operatorname\{sp\}\(h\)=2\\beta\(\\pi^\{\*\},q^\{\*\}\)\\leq\\operatorname\{sp\}\(h^\{\*\}\)=B\_\{g\}\.The first inequality holds because this region is a subset of the full fixed\-point set, and the second becauseh∗h^\{\*\}belongs to the region\. Equality throughout proves attainment over selectors\. No continuity ofβ\(π,q\)\\beta\(\\pi,q\)as the selectors vary is required\. ∎
### K\.3Sparse obstructions
###### Propsition 6\.
Letrrbe the dimension of the linear span of the generator flows\. Each point ofCπqC\_\{\\pi q\}is a nonnegative combination of at mostr\+1≤nr\+1\\leq ngenerators\. Ifggis nonconstant,r\+1≤n−1r\+1\\leq n\-1\. Ifℋπq=∅\\mathcal\{H\}\_\{\\pi q\}=\\varnothing, there is a sequence
\(ak,−1\)=∑j=1mkλkjzkj,mk≤r\+1,λkj≥0,zkj∈Gπq,‖ak‖1⟶0\.\(a\_\{k\},\-1\)=\\sum\_\{j=1\}^\{m\_\{k\}\}\\lambda\_\{kj\}z\_\{kj\},\\qquad m\_\{k\}\\leq r\+1,\\quad\\lambda\_\{kj\}\\geq 0,\\quad z\_\{kj\}\\in G\_\{\\pi q\},\\quad\\\|a\_\{k\}\\\|\_\{1\}\\longrightarrow 0\.If a negative exactly balanced combination exists, the sequence can be constant\. Otherwise failure is witnessed by increasingly balanced combinations with at mostr\+1r\+1generators at each index\.
###### Proof\.
LetLLbe the span of the generator flows\. Each flowaasatisfies𝟏⊤a=g⊤a=0\\mathbf\{1\}^\{\\top\}a=g^\{\\top\}a=0, sor≤n−1r\\leq n\-1, andr≤n−2r\\leq n\-2whenggis nonconstant\. Every full generator belongs toL×ℝL\\times\\mathbb\{R\}, whose dimension isr\+1r\+1\. The conic Carathéodory theorem therefore represents every nonzero cone point with at mostr\+1r\+1generators\. Its usual linear\-dependence argument applies without any closedness assumption on the cone: from a representation with too many positive coefficients, subtract a suitable multiple of a linear dependence until one coefficient becomes zero\. Repeating removes the excess terms\. Zero has the empty representation\.
Suppose now thatℋπq=∅\\mathcal\{H\}\_\{\\pi q\}=\\varnothing\. Theorem[15](https://arxiv.org/html/2609.28792#Thmtheorem15)gives\(0,−1\)∈cl\(Cπq\)\(0,\-1\)\\in\\operatorname\{cl\}\(C\_\{\\pi q\}\)\. Choose\(a~k,b~k\)∈Cπq\(\\widetilde\{a\}\_\{k\},\\widetilde\{b\}\_\{k\}\)\\in C\_\{\\pi q\}converging to\(0,−1\)\(0,\-1\)\. After discarding finitely many terms,b~k<0\\widetilde\{b\}\_\{k\}<0, so positive rescaling gives
\(ak,−1\)=\(a~k,b~k\)−b~k∈Cπq,‖ak‖1=‖a~k‖1−b~k⟶0\.\(a\_\{k\},\-1\)=\\frac\{\(\\widetilde\{a\}\_\{k\},\\widetilde\{b\}\_\{k\}\)\}\{\-\\widetilde\{b\}\_\{k\}\}\\in C\_\{\\pi q\},\\qquad\\\|a\_\{k\}\\\|\_\{1\}=\\frac\{\\\|\\widetilde\{a\}\_\{k\}\\\|\_\{1\}\}\{\-\\widetilde\{b\}\_\{k\}\}\\longrightarrow 0\.Sparsify each point using the first textbf\. If\(0,b\)∈Cπq\(0,b\)\\in C\_\{\\pi q\}for someb<0b<0, use its rescaling to\(0,−1\)\(0,\-1\)at every index\. Otherwise no such constant exactly balanced witness exists, and the limiting sequence is necessary\. ∎
Interpretation\.Sparsity bounds the number of generators in each witness\. The coefficients can diverge and the rows can vary withkk\. Thus the proposition retains limiting obstructions without replacing compact ambiguity by one finite row list\.
### K\.4Average optimality and common\-bias compatibility
###### Example 3\(Separate barriers do not certify the same selectors\)\.
There are states\(x,y,z,w\)\(x,y,z,w\), one controller action, and rewards\(0,−1,0,1\)\(0,\-1,0,1\)\. Stateyymoves toxx; statesz,wz,ware absorbing\. Let
𝒰x=co\{p0,p1\},p0=\(0,0,1,0\),p1=\(0,1/2,1/2,0\)\.\\mathcal\{U\}\_\{x\}=\\operatorname\{co\}\\\{p^\{0\},p^\{1\}\\\},\\qquad p^\{0\}=\(0,0,1,0\),\\quad p^\{1\}=\(0,1/2,1/2,0\)\.For the selectorqx=p0q\_\{x\}=p^\{0\}, separate lower and upper potentials exist, but no common potential exists\. Forqx′=p1q^\{\\prime\}\_\{x\}=p^\{1\}, the minimum common bias span is two andβ\(π,q′\)=1\\beta\(\\pi,q^\{\\prime\}\)=1\.
###### Proof\.
At every visit toxx, absorption atzzhas probability at least1/21/2, and stateyyreturns immediately toxx\. This bound holds conditionally on every history\. Consequently absorption occurs almost surely, and the expected number of visits toyyis finite under every nature strategy\. The total negative reward before absorption therefore has finite expectation\. All four average\-payoff conventions giveg=\(0,0,0,1\)g=\(0,0,0,1\)for every stationary selector, and these selectors are average optimal\. All rows are gain\-active andc=\(0,−1,0,0\)c=\(0,\-1,0,0\)\. The lower potentialℓ=\(−1,−2,0,0\)\\ell=\(\-1,\-2,0,0\)satisfies
\(p0\)⊤ℓ=0≥ℓx,\(p1\)⊤ℓ=−1=ℓx,−1\+ℓx=ℓy\.\(p^\{0\}\)^\{\\top\}\\ell=0\\geq\\ell\_\{x\},\\qquad\(p^\{1\}\)^\{\\top\}\\ell=\-1=\\ell\_\{x\},\\qquad\-1\+\\ell\_\{x\}=\\ell\_\{y\}\.The absorbing\-state lower inequalities are equalities; affineness extends the endpoint checks to every row in𝒰x\\mathcal\{U\}\_\{x\}\. WritingPqP\_\{q\}for the selected transition matrix, the upper potentialu=\(0,−1,0,0\)u=\(0,\-1,0,0\)satisfiesu=c\+Pquu=c\+P\_\{q\}u\.
A common potential forqqwould havehx=hzh\_\{x\}=h\_\{z\}andhy=hx−1h\_\{y\}=h\_\{x\}\-1, because its selected\-row lower and upper inequalities must both hold\. The lower inequality forp1p^\{1\}would then require
hx≤\(hy\+hz\)/2=hx−1/2,h\_\{x\}\\leq\(h\_\{y\}\+h\_\{z\}\)/2=h\_\{x\}\-1/2,which is impossible\. Equivalently, its mixed cone contains
\(ex−12ey−12ez,0\)\+12\(ey−ex,−1\)\+12\(ez−ex,0\)=\(0,−1/2\)\.\(e\_\{x\}\-\\tfrac\{1\}\{2\}e\_\{y\}\-\\tfrac\{1\}\{2\}e\_\{z\},0\)\+\\tfrac\{1\}\{2\}\(e\_\{y\}\-e\_\{x\},\-1\)\+\\tfrac\{1\}\{2\}\(e\_\{z\}\-e\_\{x\},0\)=\(0,\-1/2\)\.The first two terms are lower generators and the last is an upper generator, soβ\(π,q\)=\+∞\\beta\(\\pi,q\)=\+\\infty\.
Forq′q^\{\\prime\}, the vectorh=ℓh=\\ellsatisfies both families of constraints\. Every common bias obeys
hy=hx−1,hx=\(hy\+hz\)/2,hencehz=hx\+1\.h\_\{y\}=h\_\{x\}\-1,\\qquad h\_\{x\}=\(h\_\{y\}\+h\_\{z\}\)/2,\\qquad\\text\{hence \}h\_\{z\}=h\_\{x\}\+1\.Its span is at leasthz−hy=2h\_\{z\}\-h\_\{y\}=2, andℓ\\ellattains this bound\. Theorem[15](https://arxiv.org/html/2609.28792#Thmtheorem15)therefore givesβ\(π,q′\)=2/2=1\\beta\(\\pi,q^\{\\prime\}\)=2/2=1\. ∎
Discussion\.The example has polytopic ambiguity, a nonconstant gain, and stationary gain attainment for every nature selector\. Separate barriers guarantee that some Bellman fixed point exists, but they need not use the prescribed stationary selector in a common bias certificate\. Mixing the constraints is essential both for a fixed selector characterization and for the exact minimum\-span formula\.
### K\.5Compatible recurrent\-class offsets
For a fixed pair, its selected Poisson equation leaves one free constant per recurrent class\. The mixed inequalities determine which choices of these constants are compatible with all deviations\.
Writes=\(π,q\)s=\(\\pi,q\),ℋs=ℋπq\\mathcal\{H\}\_\{s\}=\\mathcal\{H\}\_\{\\pi q\},\(Ps\)i⋅=qi,π\(i\)⊤\(P\_\{s\}\)\_\{i\\cdot\}=q\_\{i,\\pi\(i\)\}^\{\\top\}, andciπ=ri,π\(i\)−gic\_\{i\}^\{\\pi\}=r\_\{i,\\pi\(i\)\}\-g\_\{i\}\. The Cesàro projectorPs∞P\_\{s\}^\{\\infty\}and the fundamental matrix are defined in equation[18](https://arxiv.org/html/2609.28792#A3.E18)\. A compatible pair must satisfyPs∞cπ=0P\_\{s\}^\{\\infty\}c^\{\\pi\}=0\. For such a pair, define
hs0=\(I−Ps\+Ps∞\)−1cπ\.h\_\{s\}^\{0\}=\(I\-P\_\{s\}\+P\_\{s\}^\{\\infty\}\)^\{\-1\}c^\{\\pi\}\.LetCs,1,…,Cs,msC\_\{s,1\},\\ldots,C\_\{s,m\_\{s\}\}be its recurrent classes\. Define\(Hs\)ij\(H\_\{s\}\)\_\{ij\}as the probability of eventually enteringCs,jC\_\{s,j\}from stateii, and let rowjjofNsN\_\{s\}be the invariant distribution of that class, extended by zero to the remaining states\. ThusHs∈ℝn×msH\_\{s\}\\in\\mathbb\{R\}^\{n\\times m\_\{s\}\}andNs∈ℝms×nN\_\{s\}\\in\\mathbb\{R\}^\{m\_\{s\}\\times n\}\. Finally, put𝒵s=\{z∈ℝms:hs0\+Hsz∈ℋs\}\\mathcal\{Z\}\_\{s\}=\\\{z\\in\\mathbb\{R\}^\{m\_\{s\}\}:h\_\{s\}^\{0\}\+H\_\{s\}z\\in\\mathcal\{H\}\_\{s\}\\\}\.
###### Propsition 7\.
Forg=T^gg=\\widehat\{T\}g,
\{h:Kgh=h\}=⋃s:Ps∞cπ=0\{hs0\+Hsz:z∈𝒵s\}\.\\\{h:K\_\{g\}h=h\\\}=\\bigcup\_\{s:\\,P\_\{s\}^\{\\infty\}c^\{\\pi\}=0\}\\\{h\_\{s\}^\{0\}\+H\_\{s\}z:z\\in\\mathcal\{Z\}\_\{s\}\\\}\.\(77\)For each fixed pair,𝒵s\\mathcal\{Z\}\_\{s\}is closed and convex and the mapz↦hs0\+Hszz\\mapsto h\_\{s\}^\{0\}\+H\_\{s\}zis one\-to\-one\. Regions from different pairs may overlap\. For polytopic gain faces, finitely many vertex\-selector pairs suffice and their regions are polyhedral\.
The Poisson representation leaves one constant per recurrent class; the mixed inequalities determine which constants work against*both*players’ deviations\. Thus the canonical choicez=0z=0can fail even when the same selector pair has a compatible repair\. Solving forz∈𝒵sz\\in\\mathcal\{Z\}\_\{s\}repairs precisely this failure\. The linear\-chain representation is classical\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Section 2, equations \(2\.6\)\-\(2\.9\)\]; the additional restriction here is the common two\-player certificateℋs\\mathcal\{H\}\_\{s\}\. The proof below applies the finite\-chain facts already collected in Lemma[2](https://arxiv.org/html/2609.28792#Thmlemma2)\. Example[4](https://arxiv.org/html/2609.28792#Thmexample4)illustrates a repair using only class offsets\. Class\-based descriptions of Bellman solutions have substantial precedents:\[[3](https://arxiv.org/html/2609.28792#bib.bib72), Theorem 1\.1\]characterize eigenspaces of convex monotone homogeneous maps using critical classes, and\[[87](https://arxiv.org/html/2609.28792#bib.bib65), Lemma D\.3 and Theorem D\.4\]study gain\-direction shifts that enforce additional Bellman inequalities and can substantially increase bias span in nominal multichain MDPs\. Here the possibly nonconvex max\-min operator is handled pair by pair: the mixed inequalities impose compatibility with both players\.
###### Proof of Proposition[7](https://arxiv.org/html/2609.28792#Thmproposition7)\.
The two mixed inequalities at the selected action and row force\(I−Ps\)h=cπ\(I\-P\_\{s\}\)h=c^\{\\pi\}\. Multiplying byPs∞P\_\{s\}^\{\\infty\}shows why pairs withPs∞cπ≠0P\_\{s\}^\{\\infty\}c^\{\\pi\}\\neq 0must be discarded\. For every other pair, Lemma[2](https://arxiv.org/html/2609.28792#Thmlemma2)gives the particular solutionhs0h\_\{s\}^\{0\}withPs∞hs0=0P\_\{s\}^\{\\infty\}h\_\{s\}^\{0\}=0\.
We now identify all solutions of this Poisson equation\. The absorption representation givesPs∞=HsNsP\_\{s\}^\{\\infty\}=H\_\{s\}N\_\{s\}, whileNsHs=ImsN\_\{s\}H\_\{s\}=I\_\{m\_\{s\}\}because starting in a recurrent class leads to that same class with probability one\. Alsoker\(I−Ps\)=imPs∞\\ker\(I\-P\_\{s\}\)=\\operatorname\{im\}P\_\{s\}^\{\\infty\}: a harmonic vector is unchanged by every Cesàro average, and every vector in the image ofPs∞P\_\{s\}^\{\\infty\}is harmonic\. Hence
\(I−Ps\)h=cπ⟺h=hs0\+Hszfor a uniquez∈ℝms\.\(I\-P\_\{s\}\)h=c^\{\\pi\}\\quad\\Longleftrightarrow\\quad h=h\_\{s\}^\{0\}\+H\_\{s\}z\\quad\\text\{for a unique \}z\\in\\mathbb\{R\}^\{m\_\{s\}\}\.IndeedNshs0=0N\_\{s\}h\_\{s\}^\{0\}=0, sinceHsNshs0=0H\_\{s\}N\_\{s\}h\_\{s\}^\{0\}=0andHsH\_\{s\}has full column rank, so the unique coordinates arez=Nshz=N\_\{s\}h\. Each coordinate is the invariant average ofhhon its recurrent class\.
Substituting this representation into the mixed inequalities shows explicitly which offsets are admissible:
\(ei−p\)⊤Hsz\\displaystyle\(e\_\{i\}\-p\)^\{\\top\}H\_\{s\}z≤ci,π\(i\)−\(ei−p\)⊤hs0\\displaystyle\\leq c\_\{i,\\pi\(i\)\}\-\(e\_\{i\}\-p\)^\{\\top\}h\_\{s\}^\{0\}\(p∈Fi,π\(i\)\(g\)\),\\displaystyle\(p\\in F\_\{i,\\pi\(i\)\}\(g\)\),\(qia−ei\)⊤Hsz\\displaystyle\(q\_\{ia\}\-e\_\{i\}\)^\{\\top\}H\_\{s\}z≤−cia−\(qia−ei\)⊤hs0\\displaystyle\\leq\-c\_\{ia\}\-\(q\_\{ia\}\-e\_\{i\}\)^\{\\top\}h\_\{s\}^\{0\}\(a∈Ag\(i\)\)\.\\displaystyle\(a\\in A\_\{g\}\(i\)\)\.Their solution set is exactly𝒵s\\mathcal\{Z\}\_\{s\}, a closed convex intersection of affine halfspaces\. The union identity equation[76](https://arxiv.org/html/2609.28792#A11.E76)now proves equation[77](https://arxiv.org/html/2609.28792#A11.E77)\. For polytopic gain faces, the lower inequalities need only be checked at vertices and all minimizing upper rows can be chosen at vertices\. There are finitely many such selector pairs, and each corresponding offset region is polyhedral\. ∎
Because every active row satisfiesp⊤𝟏=1p^\{\\top\}\\mathbf\{1\}=1andp⊤g=gip^\{\\top\}g=g\_\{i\}, the bias set is invariant under addition ofu𝟏\+tgu\\mathbf\{1\}\+tgfor anyu,t∈ℝu,t\\in\\mathbb\{R\}\. IndeedKg\(h\+u𝟏\+tg\)=Kgh\+u𝟏\+tgK\_\{g\}\(h\+u\\mathbf\{1\}\+tg\)=K\_\{g\}h\+u\\mathbf\{1\}\+tg\. These two directions need not exhaust the allowable recurrent\-class offsets\.
###### Example 4\(A canonical bias can be repaired without changing the pair\)\.
There are three nominal states\. At state11, the controller may stay with reward zero or move to state22with reward−1\-1\. State22moves to state33with reward22, and state33is absorbing with reward zero\. Every stationary controller has gaing=0g=0\. Choose the controller that stays at state11\. Its recurrent classes are\{1\}\\\{1\\\}and\{3\}\\\{3\\\}, and
h0=\(0,2,0\)⊤,H=\(100101\),h=h0\+Hz=\(z1,2\+z3,z3\)⊤\.h^\{0\}=\(0,2,0\)^\{\\top\},\\qquad H=\\begin\{pmatrix\}1&0\\\\ 0&1\\\\ 0&1\\end\{pmatrix\},\\qquad h=h^\{0\}\+Hz=\(z\_\{1\},2\+z\_\{3\},z\_\{3\}\)^\{\\top\}\.The selected Poisson equation holds for everyzz\. The only additional Bellman inequality comes from moving at state11and is−1\+h2≤h1\-1\+h\_\{2\}\\leq h\_\{1\}, orz1−z3≥1z\_\{1\}\-z\_\{3\}\\geq 1\. Thusz=0z=0fails, but the same pair admitsz=\(1,0\)⊤z=\(1,0\)^\{\\top\}and the exact biash=\(1,2,0\)⊤h=\(1,2,0\)^\{\\top\}\.
###### Proof\.
Under staying, state11and state33have zero reward forever, while state22receives reward22once before absorption\. Moving at state11adds only one reward−1\-1, so all policies have zero average gain\. The displayedh0h^\{0\}is the selected Poisson solution with zero recurrent\-class averages\. The Bellman equation ish1=max\{h1,−1\+h2\}h\_\{1\}=\\max\\\{h\_\{1\},\-1\+h\_\{2\}\\\},h2=2\+h3h\_\{2\}=2\+h\_\{3\}, andh3=h3h\_\{3\}=h\_\{3\}, giving the stated offset condition\. ∎
### K\.6A quantitative geometric sufficient condition
LetDπq=conv\(Gπq\)D\_\{\\pi q\}=\\operatorname\{conv\}\(G\_\{\\pi q\}\)and letAπqA\_\{\\pi q\}be its projection onto the flow coordinate\. Both are compact, by compactness ofGπqG\_\{\\pi q\}and the finite\-dimensional convex\-hull theorem\. PutLπq=span\(Aπq\)L\_\{\\pi q\}=\\operatorname\{span\}\(A\_\{\\pi q\}\)\. For any state, the two generators for the selected row satisfy
12\(ei−qiπ\(i\),ciπ\(i\)\)\+12\(qiπ\(i\)−ei,−ciπ\(i\)\)=\(0,0\)\.\\tfrac\{1\}\{2\}\(e\_\{i\}\-q\_\{i\\pi\(i\)\},c\_\{i\\pi\(i\)\}\)\+\\tfrac\{1\}\{2\}\(q\_\{i\\pi\(i\)\}\-e\_\{i\},\-c\_\{i\\pi\(i\)\}\)=\(0,0\)\.Hence0∈Dπq0\\in D\_\{\\pi q\}and its flow projection contains zero\.
###### Theorem 16\.
Fix selectors\(π,q\)\(\\pi,q\)\. Assume
1. 1\.every\(0,b\)∈Dπq\(0,b\)\\in D\_\{\\pi q\}satisfiesb≥0b\\geq 0;
2. 2\.for someρ\>0\\rho\>0, \{a∈Lπq:‖a‖1≤ρ\}⊂Aπq\.\\\{a\\in L\_\{\\pi q\}:\\\|a\\\|\_\{1\}\\leq\\rho\\\}\\subset A\_\{\\pi q\}\.
WithR=max\{\[b\]\+:\(a,b\)∈Dπq\}R=\\max\\\{\[b\]\_\{\+\}:\(a,b\)\\in D\_\{\\pi q\}\\\}, there existsh∈ℋπqh\\in\\mathcal\{H\}\_\{\\pi q\}such that
sp\(h\)=2β\(π,q\)≤2Rρ\.\\operatorname\{sp\}\(h\)=2\\beta\(\\pi,q\)\\leq\\frac\{2R\}\{\\rho\}\.\(78\)The second assumption is equivalent to00belonging to the relative interior ofAπqA\_\{\\pi q\}\. It permits multiple recurrent classes and does not require a lower bound on positive probabilities\.
###### Proof\.
We show that any negative centered reward can be bounded by the size of its flow, using an opposite flow to make an exactly balanced mixture\. WriteD=DπqD=D\_\{\\pi q\},A=AπqA=A\_\{\\pi q\}, andL=LπqL=L\_\{\\pi q\}\. IfL=\{0\}L=\\\{0\\\}, every point ofDDhas zero flow, so the first assumption gives nonnegative reward everywhere inDDand its cone\. Thusβ=0\\beta=0, and Theorem[15](https://arxiv.org/html/2609.28792#Thmtheorem15)proves the conclusion\.
SupposeL≠\{0\}L\\neq\\\{0\\\}and take\(a,b\)∈D\(a,b\)\\in D\. Ata=0a=0the first assumption already givesb≥0b\\geq 0\. Ata≠0a\\neq 0, the vectora′=−ρa/∥a∥1a^\{\\prime\}=\-\\rho a/\\\|a\\\|\_\{1\}belongs toLLand has‖a′‖1=ρ\\\|a^\{\\prime\}\\\|\_\{1\}=\\rho\. The second assumption impliesa′∈Aa^\{\\prime\}\\in A, so some realb′b^\{\\prime\}has\(a′,b′\)∈D\(a^\{\\prime\},b^\{\\prime\}\)\\in D\. Define the positive weights
α=ρρ\+‖a‖1,1−α=‖a‖1ρ\+‖a‖1\.\\alpha=\\frac\{\\rho\}\{\\rho\+\\\|a\\\|\_\{1\}\},\\qquad 1\-\\alpha=\\frac\{\\\|a\\\|\_\{1\}\}\{\\rho\+\\\|a\\\|\_\{1\}\}\.They sum to one and cancel the flow:
αa\+\(1−α\)a′=ρaρ\+‖a‖1−‖a‖1ρ\+‖a‖1ρa‖a‖1=0\.\\alpha a\+\(1\-\\alpha\)a^\{\\prime\}=\\frac\{\\rho a\}\{\\rho\+\\\|a\\\|\_\{1\}\}\-\\frac\{\\\|a\\\|\_\{1\}\}\{\\rho\+\\\|a\\\|\_\{1\}\}\\frac\{\\rho a\}\{\\\|a\\\|\_\{1\}\}=0\.SinceDDis convex, the mixture belongs toDD\. Its second coordinate must be nonnegative by the first assumption:
ρb\+‖a‖1b′ρ\+‖a‖1≥0\.\\frac\{\\rho b\+\\\|a\\\|\_\{1\}b^\{\\prime\}\}\{\\rho\+\\\|a\\\|\_\{1\}\}\\geq 0\.Multiplication by the positive denominator, followed by division byρ\>0\\rho\>0, givesb≥−\(b′/ρ\)‖a‖1b\\geq\-\(b^\{\\prime\}/\\rho\)\\\|a\\\|\_\{1\}\. By definition ofRR,b′≤\[b′\]\+≤Rb^\{\\prime\}\\leq\[b^\{\\prime\}\]\_\{\+\}\\leq R; multiplying this upper bound by−∥a∥1/ρ≤0\-\\\|a\\\|\_\{1\}/\\rho\\leq 0reverses it\. Therefore
b≥−b′ρ‖a‖1≥−Rρ‖a‖1\.b\\geq\-\\frac\{b^\{\\prime\}\}\{\\rho\}\\\|a\\\|\_\{1\}\\geq\-\\frac\{R\}\{\\rho\}\\\|a\\\|\_\{1\}\.This also includes the previously treated casea=0a=0\.
To transfer the bound fromDDto the conic hull, write any nonzero conic combination as
∑jλjzj=Λ∑jλjΛzj,Λ=∑jλj\>0\.\\sum\_\{j\}\\lambda\_\{j\}z\_\{j\}=\\Lambda\\sum\_\{j\}\\frac\{\\lambda\_\{j\}\}\{\\Lambda\}z\_\{j\},\\qquad\\Lambda=\\sum\_\{j\}\\lambda\_\{j\}\>0\.The normalized sum lies inDD\. Multiplying its inequality byΛ\\Lambdapreserves the sign and uses‖Λa‖1=Λ‖a‖1\\\|\\Lambda a\\\|\_\{1\}=\\Lambda\\\|a\\\|\_\{1\}\. Thus the same bound holds onCπqC\_\{\\pi q\}, including its zero point\. It follows thatβ\(π,q\)≤R/ρ\\beta\(\\pi,q\)\\leq R/\\rho\. Theorem[15](https://arxiv.org/html/2609.28792#Thmtheorem15)supplies a feasible bias with span exactly2β\(π,q\)2\\beta\(\\pi,q\)and hence at most2R/ρ2R/\\rho\.
Finally,0∈A0\\in Aimpliesaff\(A\)=span\(A\)=L\\operatorname\{aff\}\(A\)=\\operatorname\{span\}\(A\)=L\. By definition,0∈ri\(A\)0\\in\\operatorname\{ri\}\(A\)means thatAAcontains an open neighborhood of zero inLL\. In finite dimension, this is equivalent to containing anℓ1\\ell\_\{1\}ball of sufficiently small positive radius\. Shrinking the radius if needed makes that ball closed\. Conversely the displayed closed ball contains a relative open neighborhood\. This proves the stated relative\-interior equivalence\. ∎
Discussion\.The two assumptions have distinct roles\. The first rules out a negative exactly balanced mixture\. The second ensures that a small imbalance can be canceled using a proportionately small added mixture\. Their combination converts an exact\-cycle condition into the linear leakage bound required for finite bias\. These assumptions are sufficient; they are not claimed necessary\. In particular, polyhedral models may be solvable even when the flow projection has the origin on its relative boundary\.
One useful way to verify the flow\-interior condition uses only a finite set of reference flows\. Define the stochastic matrixQi⋅=qi,π\(i\)⊤Q\_\{i\\cdot\}=q\_\{i,\\pi\(i\)\}^\{\\top\}and letL0L\_\{0\}be the row space ofI−QI\-Q\. If every flow inGπqG\_\{\\pi q\}belongs toL0L\_\{0\}, thenLπq=L0L\_\{\\pi q\}=L\_\{0\}and the flow\-interior condition holds\. Indeed, the flow projection contains±\(ei−Qi⋅⊤\)\\pm\(e\_\{i\}\-Q\_\{i\\cdot\}^\{\\top\}\)for every state\. Choose a basisb1,…,brb\_\{1\},\\ldots,b\_\{r\}from these flows and writeℒα=∑j=1rαjbj\\mathcal\{L\}\\alpha=\\sum\_\{j=1\}^\{r\}\\alpha\_\{j\}b\_\{j\}\. Ifr\>0r\>0, the inverse coordinate map has a finite normC=‖ℒ−1‖1→1\>0C=\\\|\\mathcal\{L\}^\{\-1\}\\\|\_\{1\\to 1\}\>0\. Fora∈L0a\\in L\_\{0\}with‖a‖1≤1/C\\\|a\\\|\_\{1\}\\leq 1/C,
‖ℒ−1a‖1≤C‖a‖1≤1⟹a∈conv\{±b1,…,±br\}⊆Aπq\.\\\|\\mathcal\{L\}^\{\-1\}a\\\|\_\{1\}\\leq C\\\|a\\\|\_\{1\}\\leq 1\\quad\\Longrightarrow\\quad a\\in\\operatorname\{conv\}\\\{\\pm b\_\{1\},\\ldots,\\pm b\_\{r\}\\\}\\subseteq A\_\{\\pi q\}\.The implication follows by weighting the signedbjb\_\{j\}with the absolute values of their coordinates and allocating any unused weight to zero\. IfL0=\{0\}L\_\{0\}=\\\{0\\\}, the relative\-neighborhood condition is immediate\. All other flows lie inL0L\_\{0\}by assumption\. Equivalently, each candidate flow annihilates every harmonic vectorddsatisfyingQd=dQd=d, because the orthogonal complement of the row space ofI−QI\-Qis its nullspace\. This condition can preserve several classwise harmonic coordinates\.
### K\.7Finite linear programs for polyhedral gain faces
###### Propsition 8\.
Suppose every gain face is the convex hull of finitely many listed rows\. In computingBgB\_\{g\}, it suffices to enumerate active controllers and full selectors taking one listed row per active state\-action pair\. For each enumerated pair, write its finitely many generator inequalities asMh≤dMh\\leq d\. Its minimum bias span is the linear\-program value
minh,t\{t:Mh≤d,0≤hi≤t\(i=1,…,n\),t≥0\}\.\\min\_\{h,t\}\\\{t:Mh\\leq d,\\quad 0\\leq h\_\{i\}\\leq t\\ \(i=1,\\ldots,n\),\\quad t\\geq 0\\\}\.\(79\)When feasible, the same value is
2maxλ≥0\{−d⊤λ:‖M⊤λ‖1≤1\}\.2\\max\_\{\\lambda\\geq 0\}\\\{\-d^\{\\top\}\\lambda:\\\|M^\{\\top\}\\lambda\\\|\_\{1\}\\leq 1\\\}\.\(80\)If the primal is infeasible, the maximization in equation[80](https://arxiv.org/html/2609.28792#A11.E80)is unbounded\. The smallest enumerated primal value equalsBgB\_\{g\}; if all are infeasible,Bg=\+∞B\_\{g\}=\+\\infty\.
###### Proof\.
An affine inequality holds throughout the convex hull of a finite list exactly when it holds at every listed row\. A linear minimum over that hull is attained at a listed row\. Thus every Bellman fixed point admits a listed selector pair, and each such pair has finitely many inequalitiesMh≤dMh\\leq d\.
Fix one pair\. SinceM𝟏=0M\\mathbf\{1\}=0, translating any feasiblehhby−minihi𝟏\-\\min\_\{i\}h\_\{i\}\\mathbf\{1\}gives0≤hi≤sp\(h\)0\\leq h\_\{i\}\\leq\\operatorname\{sp\}\(h\)without altering its constraints\. Conversely, a feasible point of equation[79](https://arxiv.org/html/2609.28792#A11.E79)satisfiessp\(h\)≤t\\operatorname\{sp\}\(h\)\\leq t\. This proves the primal span formula\.
For the dual formula, midpoint centering shows that half this span value equals the finite linear\-program value
minh,B\{B:Mh≤d,−B𝟏≤h≤B𝟏,B≥0\}\.\\min\_\{h,B\}\\\{B:Mh\\leq d,\\ \-B\\mathbf\{1\}\\leq h\\leq B\\mathbf\{1\},\\ B\\geq 0\\\}\.Associate nonnegative multipliersλ\\lambdawithMh≤dMh\\leq d\. For a fixedBB, minimizing the Lagrangian over the box gives
min‖h‖∞≤B\{B\+λ⊤\(Mh−d\)\}=−d⊤λ\+B\(1−‖M⊤λ‖1\)\.\\min\_\{\\\|h\\\|\_\{\\infty\}\\leq B\}\\\{B\+\\lambda^\{\\top\}\(Mh\-d\)\\\}=\-d^\{\\top\}\\lambda\+B\(1\-\\\|M^\{\\top\}\\lambda\\\|\_\{1\}\)\.Minimizing further overB≥0B\\geq 0yields−d⊤λ\-d^\{\\top\}\\lambdawhen‖M⊤λ‖1≤1\\\|M^\{\\top\}\\lambda\\\|\_\{1\}\\leq 1, and−∞\-\\inftyotherwise\. Finite linear\-program duality therefore gives equation[80](https://arxiv.org/html/2609.28792#A11.E80)whenever the primal is feasible\. Its value is finite and attained: a feasible bias supplies a finite upper bound, whileB≥0B\\geq 0supplies a lower bound\. The norm constraint in the dual is itself polyhedral, for example by introducingz≥0z\\geq 0with−z≤M⊤λ≤z\-z\\leq M^\{\\top\}\\lambda\\leq zand𝟏⊤z≤1\\mathbf\{1\}^\{\\top\}z\\leq 1\.
IfMh≤dMh\\leq dis infeasible, Farkas’ lemma suppliesλ≥0\\lambda\\geq 0withM⊤λ=0M^\{\\top\}\\lambda=0andd⊤λ<0d^\{\\top\}\\lambda<0\. Every positive multiple remains dual feasible and its objective diverges to\+∞\+\\infty, proving unboundedness\. Finally, the finite selector reduction and equation[74](https://arxiv.org/html/2609.28792#A11.E74)identify the smallest enumerated value withBgB\_\{g\}, including the case in which every pair is infeasible\. ∎
Scope\.This finite optimization computes the minimum bias span for a prescribed recession\-fixed gain and listed polyhedral gain faces\. Selector enumeration can be exponential\. Its role is to evaluate the geometric characterization explicitly in this special case\. The unknown\-gain planner of Section[6](https://arxiv.org/html/2609.28792#S6)instead uses ordinary robust Bellman updates under compact ambiguity and finite Bellman solvability\.
## Appendix LProof of unknown\-gain anchored planning
This section proves Theorem[6](https://arxiv.org/html/2609.28792#Thmtheorem6)\. All vector norms are sup norms unless stated otherwise\. We use two previously established facts:TTis nonexpansive by Lemma[1](https://arxiv.org/html/2609.28792#Thmlemma1), and every finite Bellman solution\(g,h\)\(g,h\)hasg=g⋆g=g^\{\\star\}and affine defectωh\(t\)=‖T\(h\+tg\)−h−\(t\+1\)g‖∞→0\\omega\_\{h\}\(t\)=\\\|T\(h\+tg\)\-h\-\(t\+1\)g\\\|\_\{\\infty\}\\to 0by Theorem[3](https://arxiv.org/html/2609.28792#Thmtheorem3)and Lemma[6](https://arxiv.org/html/2609.28792#Thmlemma6)\. The proof has three parts\. First, we estimate Halpern iteration around an approximate fixed point\. Second, we apply that estimate at the budget\-dependent pointh\+Ngh\+Ng\. Third, we show why convergence of both displacement and direction is sufficient for robust policy extraction\.
The updates are the robust\-operator specialization of approximately shifted Halpern iteration in\[[87](https://arxiv.org/html/2609.28792#bib.bib65)\]\. The additional issue here is that a finite robust Bellman solution need only generate an asymptotically affine trajectory, so its finite\-time defect must be retained throughout the analysis\.
### L\.1Halpern iteration near an approximate fixed point
###### Lemma 11\.
LetS:X→XS:X\\to Xbe nonexpansive on a normed vector space\. Fix an anchorz0z\_\{0\}and a comparison pointz∗z\_\{\*\}, and setD=‖z0−z∗‖D=\\\|z\_\{0\}\-z\_\{\*\}\\\|ande=‖Sz∗−z∗‖e=\\\|Sz\_\{\*\}\-z\_\{\*\}\\\|\. For
zt\+1=2t\+3z0\+t\+1t\+3Szt,t≥0,z\_\{t\+1\}=\\frac\{2\}\{t\+3\}z\_\{0\}\+\\frac\{t\+1\}\{t\+3\}Sz\_\{t\},\\qquad t\\geq 0,we have, for everyt,N≥0t,N\\geq 0,
‖zt−z∗‖≤D\+t3e,‖SzN−zN‖≤8DN\+3\+43e\.\\\|z\_\{t\}\-z\_\{\*\}\\\|\\leq D\+\\frac\{t\}\{3\}e,\\qquad\\\|Sz\_\{N\}\-z\_\{N\}\\\|\\leq\\frac\{8D\}\{N\+3\}\+\\frac\{4\}\{3\}e\.\(81\)No exact fixed point ofSSis required\.
###### Proof\.
We first control distance from the comparison point\. Nonexpansiveness gives
‖zt\+1−z∗‖≤2Dt\+3\+t\+1t\+3\(‖zt−z∗‖\+e\)\.\\\|z\_\{t\+1\}\-z\_\{\*\}\\\|\\leq\\frac\{2D\}\{t\+3\}\+\\frac\{t\+1\}\{t\+3\}\\bigl\(\\\|z\_\{t\}\-z\_\{\*\}\\\|\+e\\bigr\)\.The claimed bound is an equality att=0t=0\. Substituting the bound at timettinto this recursion givesD\+\(t\+1\)e/3D\+\(t\+1\)e/3, sincet\+1t\+3\(1\+t/3\)=\(t\+1\)/3\\frac\{t\+1\}\{t\+3\}\(1\+t/3\)=\(t\+1\)/3\. This proves the first assertion by induction\.
For the residual, fixNNand defineM=2D\+\(N/3\+1\)eM=2D\+\(N/3\+1\)e\. For0≤t≤N0\\leq t\\leq N, the distance estimate gives
‖Szt−z0‖≤‖Szt−Sz∗‖\+‖Sz∗−z∗‖\+‖z∗−z0‖≤M\.\\\|Sz\_\{t\}\-z\_\{0\}\\\|\\leq\\\|Sz\_\{t\}\-Sz\_\{\*\}\\\|\+\\\|Sz\_\{\*\}\-z\_\{\*\}\\\|\+\\\|z\_\{\*\}\-z\_\{0\}\\\|\\leq M\.Letdt=‖zt−zt−1‖d\_\{t\}=\\\|z\_\{t\}\-z\_\{t\-1\}\\\|fort≥1t\\geq 1\. The first update givesd1≤M/3d\_\{1\}\\leq M/3\. For1≤t≤N1\\leq t\\leq N, subtracting consecutive updates and using nonexpansiveness yields
dt\+1≤t\+1t\+3dt\+2M\(t\+2\)\(t\+3\)\.d\_\{t\+1\}\\leq\\frac\{t\+1\}\{t\+3\}d\_\{t\}\+\\frac\{2M\}\{\(t\+2\)\(t\+3\)\}\.The coefficient ofMMis the change in the anchor weight\. Starting fromd1≤2M/3d\_\{1\}\\leq 2M/3, induction givesdt≤2M/\(t\+2\)d\_\{t\}\\leq 2M/\(t\+2\)for1≤t≤N\+11\\leq t\\leq N\+1: indeed, substituting this estimate in the preceding display gives2M/\(t\+3\)2M/\(t\+3\)\. Finally, the update at timeNNimplies
‖SzN−zN‖≤‖SzN−zN\+1‖\+dN\+1≤2MN\+3\+2MN\+3=8DN\+3\+43e\.\\\|Sz\_\{N\}\-z\_\{N\}\\\|\\leq\\\|Sz\_\{N\}\-z\_\{N\+1\}\\\|\+d\_\{N\+1\}\\leq\\frac\{2M\}\{N\+3\}\+\\frac\{2M\}\{N\+3\}=\\frac\{8D\}\{N\+3\}\+\\frac\{4\}\{3\}e\.The additional iteratezN\+1z\_\{N\+1\}is used only in the proof\. EvaluatingSzNSz\_\{N\}suffices to compute the residual\. ∎
### L\.2Gain, displacement, and direction estimates
Fix a finite Bellman solution\(g,h\)\(g,h\)and define
AN=‖h‖∞\+∑j=0N−1ωh\(j\),εN=AN\+‖h‖∞N,N≥1\.A\_\{N\}=\\\|h\\\|\_\{\\infty\}\+\\sum\_\{j=0\}^\{N\-1\}\\omega\_\{h\}\(j\),\\qquad\\varepsilon\_\{N\}=\\frac\{A\_\{N\}\+\\\|h\\\|\_\{\\infty\}\}\{N\},\\qquad N\\geq 1\.\(82\)Becauseωh\(j\)→0\\omega\_\{h\}\(j\)\\to 0, its Cesàro average tends to zero\. HenceAN/N→0A\_\{N\}/N\\to 0andεN→0\\varepsilon\_\{N\}\\to 0\. These quantities analyze the algorithm and are not required as inputs\.
###### Propsition 9\.
The outputs of Algorithm[1](https://arxiv.org/html/2609.28792#alg1)satisfy
‖g^N−g‖∞\\displaystyle\\\|\\widehat\{g\}\_\{N\}\-g\\\|\_\{\\infty\}≤εN,\\displaystyle\\leq\\varepsilon\_\{N\},\(83\)ηN:=‖TZN−ZN−g‖∞\\displaystyle\\eta\_\{N\}:=\\\|TZ\_\{N\}\-Z\_\{N\}\-g\\\|\_\{\\infty\}≤8ANN\+3\+43ωh\(N\)\+73εN,\\displaystyle\\leq\\frac\{8A\_\{N\}\}\{N\+3\}\+\\frac\{4\}\{3\}\\omega\_\{h\}\(N\)\+\\frac\{7\}\{3\}\\varepsilon\_\{N\},\(84\)‖ZN−h−Ng‖∞\\displaystyle\\\|Z\_\{N\}\-h\-Ng\\\|\_\{\\infty\}≤AN\+N3\(ωh\(N\)\+εN\)\.\\displaystyle\\leq A\_\{N\}\+\\frac\{N\}\{3\}\\bigl\(\\omega\_\{h\}\(N\)\+\\varepsilon\_\{N\}\\bigr\)\.\(85\)In particular, all three limits in equation[14](https://arxiv.org/html/2609.28792#S6.E14)hold\.
###### Proof\.
Gain estimation\.Nonexpansiveness and the definition of the affine defect give
‖xj\+1−h−\(j\+1\)g‖∞≤‖xj−h−jg‖∞\+ωh\(j\)\.\\\|x\_\{j\+1\}\-h\-\(j\+1\)g\\\|\_\{\\infty\}\\leq\\\|x\_\{j\}\-h\-jg\\\|\_\{\\infty\}\+\\omega\_\{h\}\(j\)\.Starting fromx0=0x\_\{0\}=0and summing overj=0,…,N−1j=0,\\ldots,N\-1yields‖xN−h−Ng‖∞≤AN\\\|x\_\{N\}\-h\-Ng\\\|\_\{\\infty\}\\leq A\_\{N\}\. Consequently,‖xN/N−g‖∞≤\(AN\+‖h‖∞\)/N\\\|x\_\{N\}/N\-g\\\|\_\{\\infty\}\\leq\(A\_\{N\}\+\\\|h\\\|\_\{\\infty\}\)/N, which proves equation[83](https://arxiv.org/html/2609.28792#A12.E83)\.
Displacement and direction\.For the second phase, use the nonexpansive mapSN\(v\)=Tv−g^NS\_\{N\}\(v\)=Tv\-\\widehat\{g\}\_\{N\}and comparison pointz∗=h\+Ngz\_\{\*\}=h\+Ng\. Its anchor isz0=xNz\_\{0\}=x\_\{N\}, and the first\-phase estimate gives
‖z0−z∗‖∞≤AN,‖SNz∗−z∗‖∞≤ωh\(N\)\+εN\.\\\|z\_\{0\}\-z\_\{\*\}\\\|\_\{\\infty\}\\leq A\_\{N\},\\qquad\\\|S\_\{N\}z\_\{\*\}\-z\_\{\*\}\\\|\_\{\\infty\}\\leq\\omega\_\{h\}\(N\)\+\\varepsilon\_\{N\}\.This comparison point therefore need not be a fixed point, but its defect vanishes\. Applying Lemma[11](https://arxiv.org/html/2609.28792#Thmlemma11)at timeNNgives
‖TZN−ZN−g^N‖∞≤8ANN\+3\+43\(ωh\(N\)\+εN\)\.\\\|TZ\_\{N\}\-Z\_\{N\}\-\\widehat\{g\}\_\{N\}\\\|\_\{\\infty\}\\leq\\frac\{8A\_\{N\}\}\{N\+3\}\+\\frac\{4\}\{3\}\\bigl\(\\omega\_\{h\}\(N\)\+\\varepsilon\_\{N\}\\bigr\)\.Adding the gain\-estimation error proves equation[84](https://arxiv.org/html/2609.28792#A12.E84)\. The distance estimate in the same lemma proves equation[85](https://arxiv.org/html/2609.28792#A12.E85)\. Dividing the latter byNN, and accounting forh/N→0h/N\\to 0, provesZN/N→gZ\_\{N\}/N\\to g\. Every term on the right of equation[84](https://arxiv.org/html/2609.28792#A12.E84)also tends to zero\. ∎
### L\.3Gain\-active action identification and average optimality
###### Completion of the proof of Theorem[6](https://arxiv.org/html/2609.28792#Thmtheorem6)\.
The preceding proposition proves the three vector limits\. We now convert them into a policy guarantee\. This requires identifying gain\-active actions before telescoping the displacement inequality\.
Gain\-active identification\.PutR=maxi,a\|ria\|R=\\max\_\{i,a\}\|r\_\{ia\}\|andaN=‖ZN/N−g‖∞a\_\{N\}=\\\|Z\_\{N\}/N\-g\\\|\_\{\\infty\}\. The row Lipschitz bound gives, uniformly ini,ai,a,
\|ria\+minp∈𝒰iap⊤ZNN−mia\(g\)\|≤R/N\+aN\.\\left\|\\frac\{r\_\{ia\}\+\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}Z\_\{N\}\}\{N\}\-m\_\{ia\}\(g\)\\right\|\\leq R/N\+a\_\{N\}\.An action maximizing the first expression has gain score within twice this error of the largest gain score\. Sincemaxamia\(g\)=gi\\max\_\{a\}m\_\{ia\}\(g\)=g\_\{i\}, every permitted greedy selector satisfies
mi,πN\(i\)\(g\)≥gi−2\(R/N\+aN\)\.m\_\{i,\\pi\_\{N\}\(i\)\}\(g\)\\geq g\_\{i\}\-2\(R/N\+a\_\{N\}\)\.\(86\)If an inactive action exists, define
ΔA=mini,a:mia\(g\)<gi\{gi−mia\(g\)\}\>0\.\\Delta\_\{A\}=\\min\_\{i,a:\\,m\_\{ia\}\(g\)<g\_\{i\}\}\\\{g\_\{i\}\-m\_\{ia\}\(g\)\\\}\>0\.The minimum is positive because the state and controller\-action sets are finite\. For all sufficiently largeNN,2\(R/N\+aN\)<ΔA2\(R/N\+a\_\{N\}\)<\\Delta\_\{A\}, so every greedy action belongs toAg\(i\)A\_\{g\}\(i\)\. If there are no inactive actions, this conclusion holds for every budget\. No finiteness assumption on nature’s row sets is used\.
Uniform performance after identification\.Fix such a budgetNN\. For each selected action and every feasible row, gain activity impliesp⊤g≥gip^\{\\top\}g\\geq g\_\{i\}\. Greediness and the definition ofηN\\eta\_\{N\}imply
ri,πN\(i\)\+p⊤ZN−ZN,i≥gi−ηN\(p∈𝒰i,πN\(i\)\)\.r\_\{i,\\pi\_\{N\}\(i\)\}\+p^\{\\top\}Z\_\{N\}\-Z\_\{N,i\}\\geq g\_\{i\}\-\\eta\_\{N\}\\qquad\(p\\in\\mathcal\{U\}\_\{i,\\pi\_\{N\}\(i\)\}\)\.\(87\)Consider any randomized history\-dependent nature strategyτ\\tau\. The first inequality makesg\(St\)g\(S\_\{t\}\)a bounded submartingale, so𝔼iπN,τg\(St\)≥gi\\mathbb\{E\}\_\{i\}^\{\\pi\_\{N\},\\tau\}g\(S\_\{t\}\)\\geq g\_\{i\}\. Taking conditional expectations in equation[87](https://arxiv.org/html/2609.28792#A12.E87)and summing over a horizonHHgives
𝔼iπN,τ∑t=0H−1rSt,πN\(St\)\\displaystyle\\mathbb\{E\}\_\{i\}^\{\\pi\_\{N\},\\tau\}\\sum\_\{t=0\}^\{H\-1\}r\_\{S\_\{t\},\\pi\_\{N\}\(S\_\{t\}\)\}≥∑t=0H−1𝔼iπN,τg\(St\)−HηN\+ZN,i−𝔼iπN,τZN\(SH\)\\displaystyle\\geq\\sum\_\{t=0\}^\{H\-1\}\\mathbb\{E\}\_\{i\}^\{\\pi\_\{N\},\\tau\}g\(S\_\{t\}\)\-H\\eta\_\{N\}\+Z\_\{N,i\}\-\\mathbb\{E\}\_\{i\}^\{\\pi\_\{N\},\\tau\}Z\_\{N\}\(S\_\{H\}\)≥H\(gi−ηN\)−sp\(ZN\)\.\\displaystyle\\geq H\(g\_\{i\}\-\\eta\_\{N\}\)\-\\operatorname\{sp\}\(Z\_\{N\}\)\.\(88\)Here the budgetNNis fixed whileH→∞H\\to\\infty, so the potential term divided byHHvanishes even thoughsp\(ZN\)\\operatorname\{sp\}\(Z\_\{N\}\)may grow withNN\. The statewise value characterization then yields
0≤g−gπN≤ηN𝟏\.0\\leq g\-g^\{\\pi\_\{N\}\}\\leq\\eta\_\{N\}\\mathbf\{1\}\.\(89\)The left inequality follows from optimality ofg=g⋆g=g^\{\\star\}\. The fixed\-policy payoff equivalences in Theorem[1](https://arxiv.org/html/2609.28792#Thmtheorem1)transfer this guarantee to all payoff conventions used in the paper\.
Eventual exact optimality\.There are finitely many deterministic stationary controllers\. If any are suboptimal, their positive errors have a positive minimum
ΔΠ=minπ∈ΠD:gπ≠g∥g−gπ∥∞\>0\.\\Delta\_\{\\Pi\}=\\min\_\{\\pi\\in\\Pi\_\{D\}:\\,g^\{\\pi\}\\neq g\}\\\|g\-g^\{\\pi\}\\\|\_\{\\infty\}\>0\.SinceηN→0\\eta\_\{N\}\\to 0, for every sufficiently largeNNinequality equation[89](https://arxiv.org/html/2609.28792#A12.E89)excludes every suboptimal deterministic controller\. Combining this threshold with the gain\-active threshold gives a singleN0N\_\{0\}valid for all permitted greedy ties\. If all deterministic controllers are optimal, no policy\-gap argument is needed\. Each selected controller consequently attainsggfrom every initial state against arbitrary history\-dependent nature\. ∎
Rates when a defect modulus is available\.If a finite Bellman solution satisfiesωh\(t\)≤C\(1\+t\)−α\\omega\_\{h\}\(t\)\\leq C\(1\+t\)^\{\-\\alpha\}, summing this bound in equation[82](https://arxiv.org/html/2609.28792#A12.E82)gives
‖g^N−g‖∞\+ηN=\{O\(N−α\),0<α<1,O\(log\(N\+1\)/N\),α=1,O\(N−1\),α\>1\.\\\|\\widehat\{g\}\_\{N\}\-g\\\|\_\{\\infty\}\+\\eta\_\{N\}=\\begin\{cases\}O\(N^\{\-\\alpha\}\),&0<\\alpha<1,\\\\ O\(\\log\(N\+1\)/N\),&\\alpha=1,\\\\ O\(N^\{\-1\}\),&\\alpha\>1\.\\end\{cases\}\(90\)After gain\-active identification, equation[89](https://arxiv.org/html/2609.28792#A12.E89)gives the corresponding controller\-loss bound\. If the affine defect is eventually zero, its sum is finite and the same argument givesO\(1/N\)O\(1/N\)\. The constants depend on the chosen solution and its defect modulus\. Under compactness alone, the proof uses onlyωh\(t\)→0\\omega\_\{h\}\(t\)\\to 0\. The existence ofN0N\_\{0\}is therefore an eventual\-optimality statement, not a computable stopping rule from the observable residual alone\.
### L\.4Finite solvability does not imply an inverse\-budget rate
###### Propsition 10\.
For every integerk≥3k\\geq 3, there is a three\-state robust MDP with one controller action, compact convex semialgebraic row uncertainty, and a finite vector Bellman solution such that Algorithm[1](https://arxiv.org/html/2609.28792#alg1)satisfies
∥g^N−g∥∞=Θ\(N−1/\(k−1\)\)\.\\\|\\widehat\{g\}\_\{N\}\-g\\\|\_\{\\infty\}=\\Theta\\bigl\(N^\{\-1/\(k\-1\)\}\\bigr\)\.Consequently, finite Bellman solvability does not imply anO\(N−1\)O\(N^\{\-1\}\)gain\-estimation rate, or any fixed positive algebraic exponent throughout this class\.
###### Proof\.
Model and Bellman certificate\.Use states\(x,y,z\)\(x,y,z\), one action per state, and rewards\(0,−1,1\)\(0,\-1,1\)\. Stateyymoves deterministically toxx, andzzis absorbing\. Atxx, set
pk\(u\)=\(1−u−uk,u,uk\),𝒰x=co\{pk\(u\):0≤u≤1/2\}\.p\_\{k\}\(u\)=\(1\-u\-u^\{k\},u,u^\{k\}\),\\qquad\\mathcal\{U\}\_\{x\}=\\operatorname\{co\}\\\{p\_\{k\}\(u\):0\\leq u\\leq 1/2\\\}\.The entries are nonnegative and sum to one\. The row set is compact and convex\. It is semialgebraic by Carathéodory’s theorem and the Tarski\-Seidenberg projection theorem: at most four curve points suffice, and their convex combinations admit a finite polynomial description with the curve parameters as auxiliary variables\.
Considerg=\(0,0,1\)g=\(0,0,1\)andh=\(0,−1,0\)h=\(0,\-1,0\)\. Atxx,pk\(u\)⊤g=ukp\_\{k\}\(u\)^\{\\top\}g=u^\{k\}, so the unique gain\-minimizing row isex=pk\(0\)e\_\{x\}=p\_\{k\}\(0\)and the gain\-face bias equation is0=00=0\. Atyy, that equation is0−1=−1\+00\-1=\-1\+0, and atzzit is1\+0=1\+01\+0=1\+0\. Thus\(g,h\)\(g,h\)solves the vector Bellman system, and verification identifiesggwith the robust gain\.
Upper bound from the affine defect\.The defect vanishes aty,zy,z\. Atxx, for all sufficiently largett, the objective−u\+tuk\-u\+tu^\{k\}has interior minimizeru=\(kt\)−1/\(k−1\)u=\(kt\)^\{\-1/\(k\-1\)\}\. Differentiating gives−1\+ktuk−1=0\-1\+ktu^\{k\-1\}=0, and hencetuk=u/ktu^\{k\}=u/kat this minimizer\. Therefore
ωh\(t\)=−min0≤u≤1/2\{−u\+tuk\}=k−1k\(kt\)−1/\(k−1\)\.\\omega\_\{h\}\(t\)=\-\\min\_\{0\\leq u\\leq 1/2\}\\\{\-u\+tu^\{k\}\\\}=\\frac\{k\-1\}\{k\}\(kt\)^\{\-1/\(k\-1\)\}\.The finitely many smallerttcontribute a bounded amount to the accumulated defect\. Applying equation[83](https://arxiv.org/html/2609.28792#A12.E83)gives the claimedO\(N−1/\(k−1\)\)O\(N^\{\-1/\(k\-1\)\}\)upper bound\.
Matching lower bound from a feasible nature strategy\.FixN≥2N\\geq 2and let nature use the stationary rowpk\(u\)p\_\{k\}\(u\)withu=14N−1/\(k−1\)u=\\frac\{1\}\{4\}N^\{\-1/\(k\-1\)\}\. Starting atxx, writeat,bt,cta\_\{t\},b\_\{t\},c\_\{t\}for the probabilities of being atx,y,zx,y,z\. The transition rules imply
bt=uat−1\(t≥1\),ct=uk∑s=0t−1as≤tuk\.b\_\{t\}=ua\_\{t\-1\}\\quad\(t\\geq 1\),\\qquad c\_\{t\}=u^\{k\}\\sum\_\{s=0\}^\{t\-1\}a\_\{s\}\\leq tu^\{k\}\.Fort≤Nt\\leq N, we havebt≤ub\_\{t\}\\leq uandct≤Nukc\_\{t\}\\leq Nu^\{k\}\. Sinceu≤1/4u\\leq 1/4andNuk−1=4−\(k−1\)≤1/16Nu^\{k\-1\}=4^\{\-\(k\-1\)\}\\leq 1/16, it follows thatat=1−bt−ct≥1−u−Nuk≥1/2a\_\{t\}=1\-b\_\{t\}\-c\_\{t\}\\geq 1\-u\-Nu^\{k\}\\geq 1/2\. Consequently,
𝔼x∑t=0N−1r\(St\)\\displaystyle\\mathbb\{E\}\_\{x\}\\sum\_\{t=0\}^\{N\-1\}r\(S\_\{t\}\)=−∑t=1N−1bt\+∑t=1N−1ct\\displaystyle=\-\\sum\_\{t=1\}^\{N\-1\}b\_\{t\}\+\\sum\_\{t=1\}^\{N\-1\}c\_\{t\}≤−u\(N−1\)2\+ukN\(N−1\)2\\displaystyle\\leq\-\\frac\{u\(N\-1\)\}\{2\}\+\\frac\{u^\{k\}N\(N\-1\)\}\{2\}=−u\(N−1\)2\(1−4−\(k−1\)\)\.\\displaystyle=\-\\frac\{u\(N\-1\)\}\{2\}\\bigl\(1\-4^\{\-\(k\-1\)\}\\bigr\)\.The finite\-horizon robust value is no larger than the payoff under this feasible nature strategy\. Usinggx=0g\_\{x\}=0,g^N=TN0/N\\widehat\{g\}\_\{N\}=T^\{N\}0/N, and\(N−1\)/N≥1/2\(N\-1\)/N\\geq 1/2, we obtain
∥g^N−g∥∞≥−\(TN0\)xN≥1−4−\(k−1\)16N−1/\(k−1\)\.\\\|\\widehat\{g\}\_\{N\}\-g\\\|\_\{\\infty\}\\geq\-\\frac\{\(T^\{N\}0\)\_\{x\}\}\{N\}\\geq\\frac\{1\-4^\{\-\(k\-1\)\}\}\{16\}\\,N^\{\-1/\(k\-1\)\}\.This proves the matching order\. For any prescribedα\>0\\alpha\>0, choosingkkwith1/\(k−1\)<α1/\(k\-1\)<\\alpharules out anO\(N−α\)O\(N^\{\-\\alpha\}\)bound for this instance, even with an instance\-dependent constant\. ∎
The obstruction concerns the finite\-horizon gain estimator used by Algorithm[1](https://arxiv.org/html/2609.28792#alg1), not every possible planning algorithm\. Rows with small positive leakage can produce long negative transients even though the gain\-minimizing limiting row stays atxx\. This is the behavior measured by the affine defect\.
## Appendix MSupporting results and solvability obstructions
These results support the main\-text discussions without interrupting the proof sequence for the main theorems\. The first subsection proves structural sufficient conditions for finite Bellman solvability; the second verifies its two distinct failure mechanisms; the third records representation and reward invariances and the necessary recession equation\.
### M\.1Structural multichain existence regimes
We now give sufficient assumptions on the original ambiguity sets\. None requires a unique recurrent class\. We useVεV\_\{\\varepsilon\}for the unnormalized discounted value satisfyingVε=T\(\(1−ε\)Vε\)V\_\{\\varepsilon\}=T\(\(1\-\\varepsilon\)V\_\{\\varepsilon\}\)\.
###### Lemma 12\.
Supposeεk↓0\\varepsilon\_\{k\}\\downarrow 0andhk=Vεk−g/εkh\_\{k\}=V\_\{\\varepsilon\_\{k\}\}\-g/\\varepsilon\_\{k\}is bounded for somegg\. Every convergent subsequencehk→hh\_\{k\}\\to hgives a solutiong=T^gg=\\widehat\{T\}g,g\+h=Lghg\+h=L\_\{g\}h\.
###### Proof\.
Pass to the stated subsequence, sohk→hh\_\{k\}\\to h, and putvk=ϵkVϵk=g\+ϵkhkv\_\{k\}=\{\\epsilon\}\_\{k\}V\_\{\{\\epsilon\}\_\{k\}\}=g\+\{\\epsilon\}\_\{k\}h\_\{k\}\. The normalized discounted equation and bounded rewards give
‖vk−T^g‖∞≤ϵkR\+\(1−ϵk\)‖vk−g‖∞\+ϵk‖g‖∞⟶0\.\\\|v\_\{k\}\-\\widehat\{T\}g\\\|\_\{\\infty\}\\leq\{\\epsilon\}\_\{k\}R\+\(1\-\{\\epsilon\}\_\{k\}\)\\\|v\_\{k\}\-g\\\|\_\{\\infty\}\+\{\\epsilon\}\_\{k\}\\\|g\\\|\_\{\\infty\}\\longrightarrow 0\.Sincevk→gv\_\{k\}\\to g, this provesg=T^gg=\\widehat\{T\}g\.
For the next order, settk=ϵk−1−1t\_\{k\}=\{\\epsilon\}\_\{k\}^\{\-1\}\-1andzk=\(1−ϵk\)hkz\_\{k\}=\(1\-\{\\epsilon\}\_\{k\}\)h\_\{k\}\. The discounted equation becomes
g\+hk=T\(tkg\+zk\)−tkg\.g\+h\_\{k\}=T\(t\_\{k\}g\+z\_\{k\}\)\-t\_\{k\}g\.Heretk→∞t\_\{k\}\\to\\inftyandzk→hz\_\{k\}\\to h\. Nonexpansiveness and Lemma[6](https://arxiv.org/html/2609.28792#Thmlemma6)imply
‖T\(tkg\+zk\)−tkg−Lgh‖∞≤‖zk−h‖∞\+‖T\(tkg\+h\)−tkg−Lgh‖∞⟶0\.\\\|T\(t\_\{k\}g\+z\_\{k\}\)\-t\_\{k\}g\-L\_\{g\}h\\\|\_\{\\infty\}\\leq\\\|z\_\{k\}\-h\\\|\_\{\\infty\}\+\\\|T\(t\_\{k\}g\+h\)\-t\_\{k\}g\-L\_\{g\}h\\\|\_\{\\infty\}\\longrightarrow 0\.Taking limits givesg\+h=Lghg\+h=L\_\{g\}h\. The conclusion requires a bounded subsequence, not convergence of the entire centered discounted family\. ∎
###### Theorem 17\.
If every𝒰ia\\mathcal\{U\}\_\{ia\}is a polytope, a finite vector gain\-bias pair exists\. Moreover there areg,h,t0g,h,t\_\{0\}such that
T\(tg\+h\)=\(t\+1\)g\+hfor allt≥t0\.T\(tg\+h\)=\(t\+1\)g\+h\\qquad\\text\{for all \}t\\geq t\_\{0\}\.\(91\)
###### Proof\.
IfEiaE\_\{ia\}is the finite vertex set of𝒰ia\\mathcal\{U\}\_\{ia\}, linear minimization gives\(Tx\)i=maxaminp∈Eia\(ria\+p⊤x\)\(Tx\)\_\{i\}=\\max\_\{a\}\\min\_\{p\\in E\_\{ia\}\}\(r\_\{ia\}\+p^\{\\top\}x\)\. HenceTTis piecewise affine and nonexpansive in the supremum norm\. Kohlberg’s invariant\-half\-line theorem\[[28](https://arxiv.org/html/2609.28792#bib.bib69)\]gives equation[91](https://arxiv.org/html/2609.28792#A13.E91), without assumptions on recurrent classes\. Dividing that identity byttand using‖T\(tg\+h\)/t−T^g‖∞≤\(R\+‖h‖∞\)/t\\\|T\(tg\+h\)/t\-\\widehat\{T\}g\\\|\_\{\\infty\}\\leq\(R\+\\\|h\\\|\_\{\\infty\}\)/tgivesT^g=g\\widehat\{T\}g=g\. Subtractingtgtgand taking the tangent limit givesLgh=g\+hL\_\{g\}h=g\+h\.
The invariant half\-line also gives the bounded discounted centering used in the main text\. For small enoughϵ\>0\{\\epsilon\}\>0, sett=ϵ−1−1≥t0t=\{\\epsilon\}^\{\-1\}\-1\\geq t\_\{0\}andWϵ=g/ϵ\+hW\_\{\\epsilon\}=g/\{\\epsilon\}\+h\. Nonexpansiveness and equation[91](https://arxiv.org/html/2609.28792#A13.E91)yield
‖T\(\(1−ϵ\)Wϵ\)−Wϵ‖∞=‖T\(tg\+\(1−ϵ\)h\)−T\(tg\+h\)‖∞≤ϵ‖h‖∞\.\\\|T\(\(1\-\{\\epsilon\}\)W\_\{\\epsilon\}\)\-W\_\{\\epsilon\}\\\|\_\{\\infty\}=\\\|T\(tg\+\(1\-\{\\epsilon\}\)h\)\-T\(tg\+h\)\\\|\_\{\\infty\}\\leq\{\\epsilon\}\\\|h\\\|\_\{\\infty\}\.The discounted operator has contraction factor1−ϵ1\-\{\\epsilon\}, so its fixed\-point residual bound gives‖Vϵ−Wϵ‖∞≤‖h‖∞\\\|V\_\{\\epsilon\}\-W\_\{\\epsilon\}\\\|\_\{\\infty\}\\leq\\\|h\\\|\_\{\\infty\}\. Thus‖Vϵ−g/ϵ‖∞≤2‖h‖∞\\\|V\_\{\\epsilon\}\-g/\{\\epsilon\}\\\|\_\{\\infty\}\\leq 2\\\|h\\\|\_\{\\infty\}for all sufficiently smallϵ\{\\epsilon\}\. Verification identifiesg=g⋆g=g^\{\\star\}\. ∎
The polytope reduction checks the hypotheses of Kohlberg’s theorem; the final limit calculation identifies its invariant half\-line with the present gain\-bias equations\. The finite perfect\-information treatment is also developed in\[[2](https://arxiv.org/html/2609.28792#bib.bib70)\]\.
###### Lemma 13\.
Let𝒫\\mathcal\{P\}be a compact family of finite stochastic matrices such that, for someδ\>0\\delta\>0, every entry of everyP∈𝒫P\\in\\mathcal\{P\}is either zero or at leastδ\\delta\. ThenP↦P∞P\\mapsto P^\{\\infty\}is continuous on𝒫\\mathcal\{P\}, and
supP∈𝒫,0≤β<1‖\(I−βP\)−1−P∞1−β‖∞<∞\.\\sup\_\{P\\in\\mathcal\{P\},\\,0\\leq\\beta<1\}\\left\\\|\(I\-\\beta P\)^\{\-1\}\-\\frac\{P^\{\\infty\}\}\{1\-\\beta\}\\right\\\|\_\{\\infty\}<\\infty\.\(92\)
###### Proof\.
WriteΓ\(P\)=P∞\\Gamma\(P\)=P^\{\\infty\}\. The support gap makes every convergent sequence in𝒫\\mathcal\{P\}eventually have the support of its limit\. On a fixed support, the transient set and recurrent classes are fixed\. IfQQis the transient block,BjB\_\{j\}the block leading to recurrent classCjC\_\{j\}, andνj\\nu\_\{j\}that class’s invariant distribution, the standard finite\-chain decomposition gives
Γ\(P\)D,Cj=\(I−Q\)−1Bj𝟏νj⊤,Γ\(P\)Cj,Cj=𝟏νj⊤\.\\Gamma\(P\)\_\{D,C\_\{j\}\}=\(I\-Q\)^\{\-1\}B\_\{j\}\\mathbf\{1\}\\,\\nu\_\{j\}^\{\\top\},\\qquad\\Gamma\(P\)\_\{C\_\{j\},C\_\{j\}\}=\\mathbf\{1\}\\nu\_\{j\}^\{\\top\}\.The remaining blocks are zero\. Inversion ofI−QI\-Qis continuous sinceρ\(Q\)<1\\rho\(Q\)<1, andνj\\nu\_\{j\}is continuous as the uniquely normalized solution of a finite irreducible stationary system\. ThusΓ\\Gammais continuous on every fixed\-support stratum and hence on𝒫\\mathcal\{P\}\. This is also the support\-gap implication in\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Theorem 2\]\.
To bound the discounted remainder, use the invariant decompositionℝn=imΓ⊕kerΓ\\mathbb\{R\}^\{n\}=\\operatorname\{im\}\\Gamma\\oplus\\ker\\Gamma, on whichPPacts as the identity and as its restriction tokerΓ\\ker\\Gamma, respectively\. It gives
\(I−βP\)−1−Γ\(P\)1−β=\[I−β\(P−Γ\(P\)\)\]−1\(I−Γ\(P\)\)\.\(I\-\\beta P\)^\{\-1\}\-\\frac\{\\Gamma\(P\)\}\{1\-\\beta\}=\[I\-\\beta\(P\-\\Gamma\(P\)\)\]^\{\-1\}\(I\-\\Gamma\(P\)\)\.\(93\)Forβ<1\\beta<1the inverse exists by the discounted resolvent\. Atβ=1\\beta=1its matrix isI−P\+Γ\(P\)I\-P\+\\Gamma\(P\), invertible by Lemma[2](https://arxiv.org/html/2609.28792#Thmlemma2)\. The right side is therefore continuous on the compact set𝒫×\[0,1\]\\mathcal\{P\}\\times\[0,1\], so it is uniformly bounded\. Periodicity does not affect this argument: only the eigenvalue one is removed\. ∎
###### Theorem 18\.
Suppose there isδ\>0\\delta\>0such that
pj=0orpj≥δfor everyi,a,p∈𝒰ia,j\.p\_\{j\}=0\\quad\\text\{or\}\\quad p\_\{j\}\\geq\\delta\\qquad\\text\{for every \}i,a,p\\in\\mathcal\{U\}\_\{ia\},j\.\(94\)Then a finite gain\-bias pair exists\. More strongly, for some vectorggand some finite constantCC,
‖Vε−gε‖∞≤C\(0<ε<1\)\.\\left\\\|V\_\{\\varepsilon\}\-\\frac\{g\}\{\\varepsilon\}\\right\\\|\_\{\\infty\}\\leq C\\qquad\(0<\\varepsilon<1\)\.\(95\)The same assertion holds for every fixed stationary policy\. For a fixed randomized policy, the constant may depend on its positive action probabilities\.
###### Proof\.
We first obtain a uniform bound for stationary chains, then pass through the two optimizations\. For every deterministic controllerπ\\pi, its induced\-kernel family is compact and inherits the support gap\. Lemma[13](https://arxiv.org/html/2609.28792#Thmlemma13), the bound‖rπ‖∞≤R\\\|r^\{\\pi\}\\\|\_\{\\infty\}\\leq R, and finiteness ofΠD\\Pi\_\{D\}give a commonC<∞C<\\inftysuch that
‖\(I−\(1−ϵ\)Pπq\)−1rπ−ηπqϵ‖∞≤C,ηπq=\(Pπq\)∞rπ,\\left\\\|\(I\-\(1\-\{\\epsilon\}\)P^\{\\pi q\}\)^\{\-1\}r^\{\\pi\}\-\\frac\{\\eta^\{\\pi q\}\}\{\{\\epsilon\}\}\\right\\\|\_\{\\infty\}\\leq C,\\qquad\\eta^\{\\pi q\}=\(P^\{\\pi q\}\)^\{\\infty\}r^\{\\pi\},\(96\)for every stationary pair and0<ϵ<10<\{\\epsilon\}<1\. The same constant works for every pair because the reduced resolvent bound is uniform and there are finitely many controller policies\.
For fixedπ\\pi, discounted dynamic programming identifiesVϵ,iπV\_\{\{\\epsilon\},i\}^\{\\pi\}with the infimum of the displayed stationary\-chain value overqq\. Setγiπ=infqηiπq\\gamma\_\{i\}^\{\\pi\}=\\inf\_\{q\}\\eta\_\{i\}^\{\\pi q\}\. Taking coordinatewise infima in equation[96](https://arxiv.org/html/2609.28792#A13.E96)gives
‖Vϵπ−γπϵ‖∞≤C\.\\left\\\|V\_\{\\epsilon\}^\{\\pi\}\-\\frac\{\\gamma^\{\\pi\}\}\{\{\\epsilon\}\}\\right\\\|\_\{\\infty\}\\leq C\.\(97\)No stationary attainment is needed for this step\. Multiplication byϵ\{\\epsilon\}and Lemma[4](https://arxiv.org/html/2609.28792#Thmlemma4)identifyγπ=gπ\\gamma^\{\\pi\}=g^\{\\pi\}\.
Discounted optimality givesVϵ,i=maxπ∈ΠDVϵ,iπV\_\{\{\\epsilon\},i\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}V\_\{\{\\epsilon\},i\}^\{\\pi\}\. Taking this finite maximum in the preceding coordinate bounds yields
‖Vϵ−g¯ϵ‖∞≤C,g¯i=maxπ∈ΠDgiπ\.\\left\\\|V\_\{\\epsilon\}\-\\frac\{\\bar\{g\}\}\{\{\\epsilon\}\}\\right\\\|\_\{\\infty\}\\leq C,\\qquad\\bar\{g\}\_\{i\}=\\max\_\{\\pi\\in\\Pi\_\{D\}\}g\_\{i\}^\{\\pi\}\.Theorem[7](https://arxiv.org/html/2609.28792#Thmtheorem7)identifiesg¯=g⋆\\bar\{g\}=g^\{\\star\}\. In finite dimension, the bounded centered family has a convergent subsequence asϵ↓0\{\\epsilon\}\\downarrow 0\. Lemma[12](https://arxiv.org/html/2609.28792#Thmlemma12)gives a finite optimal Bellman pair\. The same argument withTπT^\{\\pi\}gives the fixed deterministic\-policy assertion\.
For fixed stationary randomizedπ\\pi, letαπ=min\{π\(a∣i\):π\(a∣i\)\>0\}\>0\\alpha\_\{\\pi\}=\\min\\\{\\pi\(a\\mid i\):\\pi\(a\\mid i\)\>0\\\}\>0\. If thejjth coordinate of an effective rowp¯=∑aπ\(a∣i\)pa\\bar\{p\}=\\sum\_\{a\}\\pi\(a\\mid i\)p\_\{a\}is positive, one of its nonnegative summands is at leastαπδ\\alpha\_\{\\pi\}\\delta\. Hence every effective row has support gapαπδ\\alpha\_\{\\pi\}\\delta\. Its compact effective row sets satisfy the same fixed\-policy argument\. The resulting bound may depend onπ\\pi\. ∎
###### Corollary 3\.
For every deterministic stationary controllerπ\\pi, let𝒫π=\{Pπq:q∈∏i𝒰i,π\(i\)\}\\mathcal\{P\}^\{\\pi\}=\\\{P\_\{\\pi q\}:q\\in\\prod\_\{i\}\\mathcal\{U\}\_\{i,\\pi\(i\)\}\\\}\. SupposeP↦P∞P\\mapsto P^\{\\infty\}is continuous on each compact family𝒫π\\mathcal\{P\}^\{\\pi\}\. Then the optimal discounted values have bounded vector centering as in equation[95](https://arxiv.org/html/2609.28792#A13.E95), and the full vector Bellman system has a finite solution\. The same conclusion holds for every fixed deterministic policy\.
###### Proof\.
The proof of Lemma[13](https://arxiv.org/html/2609.28792#Thmlemma13)uses the support gap only to establish continuity ofP∞P^\{\\infty\}\. Under the present hypothesis, equation[93](https://arxiv.org/html/2609.28792#A13.E93)is continuous on each compact family𝒫π×\[0,1\]\\mathcal\{P\}^\{\\pi\}\\times\[0,1\]directly\. Consequently equation[96](https://arxiv.org/html/2609.28792#A13.E96)holds with a uniform constant after maximizing over finitely manyπ\\pi\. Taking infima overqqand maxima overπ\\pias above proves bounded vector centering and then finite Bellman solvability\.
We also make explicit the connection with the stationary criterion of Theorem[5](https://arxiv.org/html/2609.28792#Thmtheorem5)\. On the compact full\-selector space𝒬S\\mathcal\{Q\}\_\{S\}, every vectorηπ,q=Pπ,q∞rπ\\eta^\{\\pi,q\}=P\_\{\\pi,q\}^\{\\infty\}r^\{\\pi\}is continuous inqq\. Hencedi\(q\)=maxπ∈ΠDηiπ,qd\_\{i\}\(q\)=\\max\_\{\\pi\\in\\Pi\_\{D\}\}\\eta\_\{i\}^\{\\pi,q\}is continuous\. Choose the common stationary approximations from equation[28](https://arxiv.org/html/2609.28792#A5.E28)with errors tending to zero and take a convergent subsequence\. Its limitq⋆q^\{\\star\}satisfiesd\(q⋆\)=g⋆d\(q^\{\\star\}\)=g^\{\\star\}\. Together with the common optimal controllerπ⋆\\pi^\{\\star\}, this gives the stationary gain guarantees of Theorem[10](https://arxiv.org/html/2609.28792#Thmtheorem10)\.
To restrict this pair to the tangent game, one must check activity, including nature’s rows at unused active actions\. Fixed\-policy discounted optimality, after normalization and passage to the limit, givesmi,π⋆\(i\)\(g⋆\)=gi⋆m\_\{i,\\pi^\{\\star\}\(i\)\}\(g^\{\\star\}\)=g\_\{i\}^\{\\star\}, soπ⋆∈Πg⋆\\pi^\{\\star\}\\in\\Pi\_\{g^\{\\star\}\}\. For the nominal MDP obtained by fixingq⋆q^\{\\star\}, the same normalized Bellman limit gives
maxa∈A\(i\)\(qia⋆\)⊤g⋆=di\(q⋆\)=gi⋆\.\\max\_\{a\\in A\(i\)\}\(q^\{\\star\}\_\{ia\}\)^\{\\top\}g^\{\\star\}=d\_\{i\}\(q^\{\\star\}\)=g\_\{i\}^\{\\star\}\.Thus\(qia⋆\)⊤g⋆≤gi⋆\(q^\{\\star\}\_\{ia\}\)^\{\\top\}g^\{\\star\}\\leq g\_\{i\}^\{\\star\}for every action\. Ifa∈Ag⋆\(i\)a\\in A\_\{g^\{\\star\}\}\(i\), feasibility also gives\(qia⋆\)⊤g⋆≥mia\(g⋆\)=gi⋆\(q^\{\\star\}\_\{ia\}\)^\{\\top\}g^\{\\star\}\\geq m\_\{ia\}\(g^\{\\star\}\)=g\_\{i\}^\{\\star\}\. Equality holds, so every such row belongs toFia\(g⋆\)F\_\{ia\}\(g^\{\\star\}\)\. The restrictionq¯\\bar\{q\}is a full tangent plan, and we may takeπ¯=π⋆\\bar\{\\pi\}=\\pi^\{\\star\}\.
Every tangent pair preservesg⋆g^\{\\star\}:Pπ,qg⋆=g⋆P\_\{\\pi,q\}g^\{\\star\}=g^\{\\star\}, hencePπ,q∞g⋆=g⋆P\_\{\\pi,q\}^\{\\infty\}g^\{\\star\}=g^\{\\star\}\. Its centered gain is therefore its original gain minusg⋆g^\{\\star\}\. The original stationary guarantees give equation[65](https://arxiv.org/html/2609.28792#A10.E65)\-equation[66](https://arxiv.org/html/2609.28792#A10.E66)\. Finally, continuity ofP∞P^\{\\infty\}makesZP=\(I−P\+P∞\)−1Z\_\{P\}=\(I\-P\+P^\{\\infty\}\)^\{\-1\}continuous and uniformly bounded on the compactπ¯\\bar\{\\pi\}\-kernel family\. Withci=ri,π¯\(i\)−gi⋆c\_\{i\}=r\_\{i,\\bar\{\\pi\}\(i\)\}\-g\_\{i\}^\{\\star\}, this boundsZPcZ\_\{P\}cuniformly over all zero\-centered\-gain replies and proves equation[67](https://arxiv.org/html/2609.28792#A10.E67)\. Thus all three stationary conditions are verified directly\. ∎
###### Corollary 4\.
Suppose each compact row family𝒰ia\\mathcal\{U\}\_\{ia\}has a fixed support: for each coordinatejj, eitherpj=0p\_\{j\}=0for every row in that family orpj\>0p\_\{j\}\>0for every row\. Then Theorem[18](https://arxiv.org/html/2609.28792#Thmtheorem18)applies\.
###### Proof\.
For each\(i,a,j\)\(i,a,j\)whose coordinate is positive throughout𝒰ia\\mathcal\{U\}\_\{ia\}, compactness and continuity giveδiaj=minp∈𝒰iapj\>0\\delta\_\{iaj\}=\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p\_\{j\}\>0\. There are finitely many such coordinates and at least one in each stochastic row family\. Their minimumδ\>0\\delta\>0satisfies
pj=0orpj≥δiaj≥δ\(i,a,p∈𝒰ia,j\)\.p\_\{j\}=0\\quad\\text\{or\}\\quad p\_\{j\}\\geq\\delta\_\{iaj\}\\geq\\delta\\qquad\(i,a,p\\in\\mathcal\{U\}\_\{ia\},j\)\.This is the hypothesis of Theorem[18](https://arxiv.org/html/2609.28792#Thmtheorem18)\. ∎
The support\-gap condition allows several support patterns and any number of recurrent classes\. Its one\-player antecedent appears in Schweitzer’s Theorem 2\[[54](https://arxiv.org/html/2609.28792#bib.bib68)\]\. Here the uniform resolvent estimate and the finite maximization over controller policies establish the robust optimality version\. Compactness, convexity, and semialgebraicity alone do not imply the gap and do not imply finite\-bias solvability\. The curved examples later in the paper exhibit the corresponding failure modes\.
#### M\.1\.1Divergence balls and the support\-loss boundary
For a nominal rowp0∈Δ\(S\)p^\{0\}\\in\\Delta\(S\), writeJ\(p0\)=\{j:pj0\>0\}J\(p^\{0\}\)=\\\{j:p^\{0\}\_\{j\}\>0\\\}andΔ\(J\)=\{p∈Δ\(S\):pj=0forj∉J\}\\Delta\(J\)=\\\{p\\in\\Delta\(S\):p\_\{j\}=0\\text\{ for \}j\\notin J\\\}\. We useDKL\(p∥p0\)=∑j:pj\>0pjlog\(pj/pj0\)D\_\{\\mathrm\{KL\}\}\(p\\\|p^\{0\}\)=\\sum\_\{j:p\_\{j\}\>0\}p\_\{j\}\\log\(p\_\{j\}/p^\{0\}\_\{j\}\), with value\+∞\+\\inftyifpj\>0=pj0p\_\{j\}\>0=p^\{0\}\_\{j\}for somejj\. The two orders of KL divergence give different support conditions\.
###### Corollary 5\.
Suppose each row family𝒰ia\\mathcal\{U\}\_\{ia\}is either a polytope or is compact and satisfiespj=0p\_\{j\}=0orpj≥δiap\_\{j\}\\geq\\delta\_\{ia\}for all its rows and coordinates, whereδia\>0\\delta\_\{ia\}\>0\. Then the full vector Bellman system has a finite solution\(g⋆,h\)\(g^\{\\star\},h\), andsup0<ϵ<1‖Vϵ−g⋆/ϵ‖∞<∞\\sup\_\{0<\{\\epsilon\}<1\}\\\|V\_\{\\epsilon\}\-g^\{\\star\}/\{\\epsilon\}\\\|\_\{\\infty\}<\\infty\. The fixed\-policy conclusions of Theorem[18](https://arxiv.org/html/2609.28792#Thmtheorem18)also hold\.
###### Proof\.
For every polytopic row, replace𝒰ia\\mathcal\{U\}\_\{ia\}by its finite vertex setEiaE\_\{ia\}\. The minimum ofp⊤xp^\{\\top\}xover either set is the same for everyx∈ℝSx\\in\\mathbb\{R\}^\{S\}\. Thus the replacement leaves the discounted Bellman operators and their fixed pointsVϵV\_\{\\epsilon\}unchanged\. The reduced row families are compact and have a common support gap: take the minimum of the finitely manyδia\\delta\_\{ia\}and all positive coordinates of all the finitely many vertices\.
For each deterministic controller policy, Lemma[13](https://arxiv.org/html/2609.28792#Thmlemma13)therefore gives a uniform reduced\-resolvent bound on its reduced stationary kernel family\. Discounted dynamic programming for the reduced compact, rectangular row families, followed by coordinatewise infima over nature and maxima over the finitely many deterministic policies, gives the bound onVϵ−g⋆/ϵV\_\{\\epsilon\}\-g^\{\\star\}/\{\\epsilon\}exactly as in Steps 1–3 of Theorem[18](https://arxiv.org/html/2609.28792#Thmtheorem18)\. This argument uses compactness, not convexity, of the reduced families\. Lemma[12](https://arxiv.org/html/2609.28792#Thmlemma12)applied to the unchanged original discounted operator then gives a finite pair for the original row sets\. For a fixed randomized policy, the same argument uses the effective\-row gapαπδ\\alpha\_\{\\pi\}\\deltafrom Step 4 of Theorem[18](https://arxiv.org/html/2609.28792#Thmtheorem18)\. ∎
###### Corollary 6\.
For each\(i,a\)\(i,a\), letpia0p^\{0\}\_\{ia\}be a nominal row, putJia=J\(pia0\)J\_\{ia\}=J\(p^\{0\}\_\{ia\}\), and letρia≥0\\rho\_\{ia\}\\geq 0\. Assume that every row family is one of the following:
1. 1\.a forward KL ball𝒰ia=\{p∈Δ\(S\):DKL\(p∥pia0\)≤ρia\}\\mathcal\{U\}\_\{ia\}=\\\{p\\in\\Delta\(S\):D\_\{\\mathrm\{KL\}\}\(p\\\|p^\{0\}\_\{ia\}\)\\leq\\rho\_\{ia\}\\\}with ρia<ρiadel:=minj∈Jia−log\(1−pia,j0\),−log0:=\+∞;\\rho\_\{ia\}<\\rho\_\{ia\}^\{\\mathrm\{del\}\}:=\\min\_\{j\\in J\_\{ia\}\}\-\\log\(1\-p^\{0\}\_\{ia,j\}\),\\qquad\-\\log 0:=\+\\infty;\(98\)
2. 2\.a support\-restricted reverse KL ball𝒰ia=\{p∈Δ\(Jia\):DKL\(pia0∥p\)≤ρia\}\\mathcal\{U\}\_\{ia\}=\\\{p\\in\\Delta\(J\_\{ia\}\):D\_\{\\mathrm\{KL\}\}\(p^\{0\}\_\{ia\}\\\|p\)\\leq\\rho\_\{ia\}\\\}withρia<∞\\rho\_\{ia\}<\\infty; or
3. 3\.a polytope \(including a total\-variation ball\)\.
Then the conclusions of Corollary[5](https://arxiv.org/html/2609.28792#Thmcorollary5)hold\. For a full\-support nominal row, the support restriction in item 2 is automatic\.
###### Proof\.
Fix a forward KL row and abbreviate its nominal support byJJ\. Finite forward KL divergence forcesp∈Δ\(J\)p\\in\\Delta\(J\)\. If\|J\|=1\|J\|=1, the ball is the nominal singleton\. Otherwise, forj∈Jj\\in Jand a row withpj=0p\_\{j\}=0, direct substitution and the nonnegativity of KL give
DKL\(p∥p0\)=−log\(1−pj0\)\+DKL\(p∥p0\|J∖\{j\}1−pj0\)≥−log\(1−pj0\)\.D\_\{\\mathrm\{KL\}\}\(p\\\|p^\{0\}\)=\-\\log\(1\-p^\{0\}\_\{j\}\)\+D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\\middle\\\|\\frac\{p^\{0\}\|\_\{J\\setminus\\\{j\\\}\}\}\{1\-p^\{0\}\_\{j\}\}\\right\)\\geq\-\\log\(1\-p^\{0\}\_\{j\}\)\.\(99\)The displayed conditional distribution is extended by zero outsideJ∖\{j\}J\\setminus\\\{j\\\}; it attains equality, so this is the exact cost of deleting coordinatejj\. Under equation[98](https://arxiv.org/html/2609.28792#A13.E98), no coordinate inJJcan vanish\. The ball is compact, hence each such coordinate has a positive minimum over it\.
For a reverse KL row,DKL\(p0∥p\)<∞D\_\{\\mathrm\{KL\}\}\(p^\{0\}\\\|p\)<\\inftyforcespj\>0p\_\{j\}\>0for eachj∈Jj\\in J\. More explicitly, withH\(p0\)=−∑j∈Jpj0logpj0H\(p^\{0\}\)=\-\\sum\_\{j\\in J\}p^\{0\}\_\{j\}\\log p^\{0\}\_\{j\},
DKL\(p0∥p\)=−H\(p0\)−∑j∈Jpj0logpj≥−H\(p0\)−pj0logpj\.D\_\{\\mathrm\{KL\}\}\(p^\{0\}\\\|p\)=\-H\(p^\{0\}\)\-\\sum\_\{j\\in J\}p^\{0\}\_\{j\}\\log p\_\{j\}\\geq\-H\(p^\{0\}\)\-p^\{0\}\_\{j\}\\log p\_\{j\}\.Thuspj≥exp\[−\(ρ\+H\(p0\)\)/pj0\]\>0p\_\{j\}\\geq\\exp\[\-\(\\rho\+H\(p^\{0\}\)\)/p^\{0\}\_\{j\}\]\>0throughout the ball\. The restriction toΔ\(J\)\\Delta\(J\)makes its support exactlyJJ\. There are finitely many state\-action pairs, so all the nonpolytopic rows have a common positive support gap\. Apply Corollary[5](https://arxiv.org/html/2609.28792#Thmcorollary5)\. ∎
The strict forward radius in equation[98](https://arxiv.org/html/2609.28792#A13.E98)is exact for preserving support\. At equality a conditional nominal row in equation[99](https://arxiv.org/html/2609.28792#A13.E99)loses a coordinate; mixtures of this row withp0p^\{0\}have arbitrarily small positive mass there\. It is not a necessary condition for solvability of a particular model\. For example, ifρia≥log\(1/minj∈Jiapia,j0\)\\rho\_\{ia\}\\geq\\log\(1/\\min\_\{j\\in J\_\{ia\}\}p^\{0\}\_\{ia,j\}\), thenDKL\(p∥pia0\)≤log\(1/minj∈Jiapia,j0\)D\_\{\\mathrm\{KL\}\}\(p\\\|p^\{0\}\_\{ia\}\)\\leq\\log\(1/\\min\_\{j\\in J\_\{ia\}\}p^\{0\}\_\{ia,j\}\)for everyp∈Δ\(Jia\)p\\in\\Delta\(J\_\{ia\}\); the ball is the entire simplex face, and item 3 applies\. Forward KL balls on faces of size at most two are also polytopes at every radius\.
The same argument applies to other divergence balls\. For a convex, lower\-semicontinuousf:\[0,∞\)→ℝ∪\{\+∞\}f:\[0,\\infty\)\\to\\mathbb\{R\}\\cup\\\{\+\\infty\\\}withf\(1\)=0f\(1\)=0, defineDf\(p∥p0\)=∑j∈Jpj0f\(pj/pj0\)D\_\{f\}\(p\\\|p^\{0\}\)=\\sum\_\{j\\in J\}p^\{0\}\_\{j\}f\(p\_\{j\}/p^\{0\}\_\{j\}\)onΔ\(J\)\\Delta\(J\)\. Ifpj=0p\_\{j\}=0, Jensen’s inequality on the remaining coordinates yields
Df\(p∥p0\)≥βf\(pj0\),βf\(t\):=tf\(0\)\+\(1−t\)f\(11−t\)\.D\_\{f\}\(p\\\|p^\{0\}\)\\geq\\beta\_\{f\}\(p^\{0\}\_\{j\}\),\\qquad\\beta\_\{f\}\(t\):=tf\(0\)\+\(1\-t\)f\\\!\\left\(\\frac\{1\}\{1\-t\}\\right\)\.\(100\)Equality holds forp=p0\(⋅∣J∖\{j\}\)p=p^\{0\}\(\\cdot\\mid J\\setminus\\\{j\\\}\); setβf\(1\)=\+∞\\beta\_\{f\}\(1\)=\+\\infty\. Therefore, if every such row hasρia<minj∈Jiaβf\(pia,j0\)\\rho\_\{ia\}<\\min\_\{j\\in J\_\{ia\}\}\\beta\_\{f\}\(p^\{0\}\_\{ia,j\}\), it has fixed support and can replace either KL type in Corollary[6](https://arxiv.org/html/2609.28792#Thmcorollary6)\. In particular, the face costs are−log\(1−t\)\-\\log\(1\-t\)for forward KL,t/\(1−t\)t/\(1\-t\)for Pearsonχ2\(p∥p0\)=∑j\(pj−pj0\)2/pj0\\chi^\{2\}\(p\\\|p^\{0\}\)=\\sum\_\{j\}\(p\_\{j\}\-p^\{0\}\_\{j\}\)^\{2\}/p^\{0\}\_\{j\}, and1−1−t1\-\\sqrt\{1\-t\}forH2\(p,p0\)=1−∑jpjpj0H^\{2\}\(p,p^\{0\}\)=1\-\\sum\_\{j\}\\sqrt\{p\_\{j\}p^\{0\}\_\{j\}\}\. For Hellinger balls with sparse nominal rows, the stated support restriction must be imposed explicitly\. Total\-variation balls need no radius restriction because they are polytopes\.
###### Example 5\(Failure at the first forward\-KL support loss\)\.
There are three states\(x,y,z\)\(x,y,z\)and one action per state\. Stateyyreturns deterministically toxx, statezzis absorbing, and the rewards are\(rx,ry,rz\)=\(1/2,−1,0\)\(r\_\{x\},r\_\{y\},r\_\{z\}\)=\(1/2,\-1,0\)\. Atxx, take the KL ball
𝒰x=\{p∈Δ\(\{x,y,z\}\):DKL\(p∥\(1/3,1/3,1/3\)\)≤log\(3/2\)\}\.\\mathcal\{U\}\_\{x\}=\\left\\\{p\\in\\Delta\(\\\{x,y,z\\\}\):D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\\middle\\\|\(1/3,1/3,1/3\)\\right\)\\leq\\log\(3/2\)\\right\\\}\.Its radius equals equation[98](https://arxiv.org/html/2609.28792#A13.E98)\. The robust gain isg⋆=0g^\{\\star\}=0, but the vector Bellman system has no finite solution\.
###### Proof\.
Writep=\(1−u−v,u,v\)p=\(1\-u\-v,u,v\)\. On the facev=0v=0,
DKL\(p∥\(1/3,1/3,1/3\)\)=log\(3/2\)\+dBer\(u∥1/2\),D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\\middle\\\|\(1/3,1/3,1/3\)\\right\)=\\log\(3/2\)\+d\_\{\\mathrm\{Ber\}\}\(u\\\|1/2\),wheredBer\(u∥q\)=ulog\(u/q\)\+\(1−u\)log\(\(1−u\)/\(1−q\)\)d\_\{\\mathrm\{Ber\}\}\(u\\\|q\)=u\\log\(u/q\)\+\(1\-u\)\\log\(\(1\-u\)/\(1\-q\)\)\. Since binary KL vanishes only whenu=1/2u=1/2, the sole feasible zero\-leak row isp∗=\(1/2,1/2,0\)p^\{\*\}=\(1/2,1/2,0\)\. It has recurrent class\{x,y\}\\\{x,y\\\}with average reward\(rx\+\(1/2\)ry\)/\(1\+1/2\)=0\(r\_\{x\}\+\(1/2\)r\_\{y\}\)/\(1\+1/2\)=0\. Under every feasible row withv\>0v\>0, the process visitsxxrepeatedly until it reaches absorbingzz, so its stationary gain is also zero\. Lemma[4](https://arxiv.org/html/2609.28792#Thmlemma4)identifies the fixed\-policy robust gain as the coordinatewise infimum of these stationary gains\. There is only one controller policy; henceg⋆=0g^\{\\star\}=0\.
For sufficiently smallt\>0t\>0, letpt=\(1/2−t−t2,1/2\+t,t2\)p\_\{t\}=\(1/2\-t\-t^\{2\},\\,1/2\+t,\\,t^\{2\}\), which is a stochastic row\. Expandings↦slog\(3s\)s\\mapsto s\\log\(3s\)ats=1/2s=1/2in its first two coordinates gives
DKL\(pt∥\(1/3,1/3,1/3\)\)−log\(3/2\)=t2\[1\+log\(2t2\)\]\+O\(t3\)<0\.D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\_\{t\}\\middle\\\|\(1/3,1/3,1/3\)\\right\)\-\\log\(3/2\)=t^\{2\}\\bigl\[1\+\\log\(2t^\{2\}\)\\bigr\]\+O\(t^\{3\}\)<0\.Thuspt∈𝒰xp\_\{t\}\\in\\mathcal\{U\}\_\{x\}for all sufficiently smallt\>0t\>0\. If a finite Bellman pair existed, verification \(Theorem[11](https://arxiv.org/html/2609.28792#Thmtheorem11)\) would give gaing⋆=0g^\{\\star\}=0\. The equation atyywould givehy=hx−1h\_\{y\}=h\_\{x\}\-1, and the equation atxxwould require
0=min\(1−u−v,u,v\)∈𝒰x\{1/2−u\+v\(hz−hx\)\}\.0=\\min\_\{\(1\-u\-v,u,v\)\\in\\mathcal\{U\}\_\{x\}\}\\bigl\\\{1/2\-u\+v\(h\_\{z\}\-h\_\{x\}\)\\bigr\\\}\.The feasible rowptp\_\{t\}makes the expression−t\+t2\(hz−hx\)<0\-t\+t^\{2\}\(h\_\{z\}\-h\_\{x\}\)<0for sufficiently smalltt, contradicting this equality\. Equivalently, with canonical normalizationwz=0w\_\{z\}=0, the Poisson equation forptp\_\{t\}giveswx=\(1/2−\(1/2\+t\)\)/t2=−1/tw\_\{x\}=\(1/2\-\(1/2\+t\)\)/t^\{2\}=\-1/t; the one\-sided canonical\-bias envelope fails\. ∎
### M\.2Two distinct obstructions to finite Bellman solvability
The two obstruction mechanisms have classical one\-player antecedents in\[[54](https://arxiv.org/html/2609.28792#bib.bib68), Examples 1\-2\]\. The models below realize them with state\-dependent rewards and compact convex ambiguity, and identify their gain\-face equations explicitly\.
The exact fixed\-policy criterion separates two obstructions that can occur even with compact convex semialgebraic ambiguity\. In the first example, no single stationary nature kernel attains the complete worst\-gain vector\. In the second, stationary worst\-gain kernels exist, but their canonically normalized biases have no common componentwise lower bound\.
###### Example 6\(Compact convex ambiguity without stationary attainment\)\.
There are states\(s,m,z\)\(s,m,z\)with rewards\(0,−1,0\)\(0,\-1,0\); statesm,zm,zare absorbing\. Atss, let
𝒰s=\{\(1−a,β,a−β\):0≤a≤1,0≤β≤a\(1−a\)\}\.\\mathcal\{U\}\_\{s\}=\\\{\(1\-a,\\beta,a\-\\beta\):0\\leq a\\leq 1,\\ 0\\leq\\beta\\leq a\(1\-a\)\\\}\.\(101\)This compact convex semialgebraic model has robust gain\(−1,−1,0\)\(\-1,\-1,0\), but no stationary nature selector attains it fromssand no finite Bellman pair exists\. This is the state\-reward form of the nonattainment mechanism in\[[19](https://arxiv.org/html/2609.28792#bib.bib11)\]\.
###### Proof\.
The parameter set is closed and bounded\. It is convex becauseβ≤a\(1−a\)\\beta\\leq a\(1\-a\)is the hypograph of a concave function and its other constraints are affine\. Its affine image is a compact convex set of probability rows, and the displayed polynomial constraints make it semialgebraic\.
For a stationary row witha\>0a\>0, the probability of eventual absorption inmmand the gain fromssare
ℙs\(τm<∞\)=∑t≥0\(1−a\)tβ=β/a,ηs=−β/a≥−\(1−a\)\>−1\.\\mathbb\{P\}\_\{s\}\(\\tau\_\{m\}<\\infty\)=\\sum\_\{t\\geq 0\}\(1\-a\)^\{t\}\\beta=\\beta/a,\\qquad\\eta\_\{s\}=\-\\beta/a\\geq\-\(1\-a\)\>\-1\.Ata=0a=0, the row is the self\-loop atssand has gain zero\. The boundary rowsβ=a\(1−a\)\\beta=a\(1\-a\)approach gain−1\-1asa↓0a\\downarrow 0\. Every reward is at least−1\-1, so adaptive nature cannot obtain a smaller average under any of the four conventions\. Thus the robust gain is\(−1,−1,0\)\(\-1,\-1,0\), and its coordinate atssis not attained by a stationary row\.
For completeness, the bias obstruction is also explicit\. By Theorem[11](https://arxiv.org/html/2609.28792#Thmtheorem11), every finite Bellman pair must have this true gain\. At that gain,
\(1−a,β,a−β\)⊤g=−1\+a−β≥−1\+a2\.\(1\-a,\\beta,a\-\\beta\)^\{\\top\}g=\-1\+a\-\\beta\\geq\-1\+a^\{2\}\.Equality withgs=−1g\_\{s\}=\-1is possible only ata=β=0a=\\beta=0\. The gain face is therefore the self\-loop, and its bias equation is−1\+hs=hs\-1\+h\_\{s\}=h\_\{s\}, a contradiction\. ∎
###### Example 7\(Stationary attainment with an unbounded lower bias envelope\)\.
Let the states be\(x,y,z\)\(x,y,z\), with rewards\(0,−1,0\)\(0,\-1,0\)\. Stateyyreturns toxx, statezzis absorbing, and
𝒰x=co\{\(1−k−k2,k,k2\):0≤k≤1/2\}\.\\mathcal\{U\}\_\{x\}=\\operatorname\{co\}\\\{\(1\-k\-k^\{2\},k,k^\{2\}\):0\\leq k\\leq 1/2\\\}\.The row set is compact, convex, and semialgebraic\. Every stationary kernel has gain zero and attains the robust value, but their canonical biases have no common lower bound\. No finite Bellman pair exists\.
###### Proof\.
Write a row as\(1−u−v,u,v\)\(1\-u\-v,u,v\)\. A convex combination with weightsλj\\lambda\_\{j\}and parameterskjk\_\{j\}satisfies
u=∑jλjkj,v=∑jλjkj2≥u2,v=0⟺u=0\.u=\\sum\_\{j\}\\lambda\_\{j\}k\_\{j\},\\qquad v=\\sum\_\{j\}\\lambda\_\{j\}k\_\{j\}^\{2\}\\geq u^\{2\},\\qquad v=0\\Longleftrightarrow u=0\.Compactness follows from the compact generating curve and Carathéodory’s theorem\. The parameter region is exactly0≤u≤1/20\\leq u\\leq 1/2,u2≤v≤u/2u^\{2\}\\leq v\\leq u/2: necessity follows from the preceding inequality andk2≤k/2k^\{2\}\\leq k/2; conversely, at a givenuu, the lower endpoint is generated byk=uk=uand the upper endpoint by mixingk=0,1/2k=0,1/2\. Their mixtures fill the interval\. This also verifies semialgebraicity\.
Ifu\>0u\>0, thenv\>0v\>0\. Every nonabsorbed path visitsxxinfinitely often, and its chance of avoidingzzthroughjjvisits is\(1−v\)j\(1\-v\)^\{j\}\. Thuszzis the unique recurrent class\. Ifu=v=0u=v=0, the recurrent classes are\{x\}\\\{x\\\}and\{z\}\\\{z\\\}\. In either case every stationary gain is zero\. Lemma[4](https://arxiv.org/html/2609.28792#Thmlemma4)identifies their infimum with the fixed\-policy gain, and Theorem[7](https://arxiv.org/html/2609.28792#Thmtheorem7)extends that value to history\-dependent nature\. Hence the robust gain is zero\.
Foru\>0u\>0, canonical normalization iswz=0w\_\{z\}=0\. The Poisson equations give
wy=wx−1,wx=\(1−u−v\)wx\+uwy=\(1−v\)wx−u,w\_\{y\}=w\_\{x\}\-1,\\qquad w\_\{x\}=\(1\-u\-v\)w\_\{x\}\+uw\_\{y\}=\(1\-v\)w\_\{x\}\-u,sowx=−u/vw\_\{x\}=\-u/vandwy=−1−u/vw\_\{y\}=\-1\-u/v\. Along the generating curve,
w\(k\)=\(−1/k,−1−1/k,0\)\(k\>0\),w\(0\)=\(0,−1,0\)\.w\(k\)=\(\-1/k,\-1\-1/k,0\)\\quad\(k\>0\),\\qquad w\(0\)=\(0,\-1,0\)\.All these kernels attain gain zero, but their first two bias coordinates tend to−∞\-\\inftyask↓0k\\downarrow 0\.
Finally, a finite bias at gain zero would satisfyhy=hx−1h\_\{y\}=h\_\{x\}\-1and
0=min0≤k≤1/2\{−k\+k2\(hz−hx\)\}\.0=\\min\_\{0\\leq k\\leq 1/2\}\\\{\-k\+k^\{2\}\(h\_\{z\}\-h\_\{x\}\)\\\}\.ForH=hz−hx≤0H=h\_\{z\}\-h\_\{x\}\\leq 0, every positivekkmakes this expression negative\. ForH\>0H\>0, any0<k<min\{1/2,1/H\}0<k<\\min\\\{1/2,1/H\\\}does so because−k\+k2H=k\(−1\+kH\)<0\-k\+k^\{2\}H=k\(\-1\+kH\)<0\. Thus no such bias exists\. Theorem[11](https://arxiv.org/html/2609.28792#Thmtheorem11)excludes a pair with a different gain\. ∎
Corollary[2](https://arxiv.org/html/2609.28792#Thmcorollary2)diagnoses the examples clause by clause: the first violates simultaneous worst\-gain attainment, while the second violates the one\-sided canonical\-bias envelope\. Yet their statewise average values remain well defined by Appendix[E](https://arxiv.org/html/2609.28792#A5), which establishes value existence without assuming a finite bias\.
### M\.3Representation invariance and elementary stability
###### Propsition 11\.
The following statements hold without restrictions on recurrent classes\.
1. \(i\)Replacing each𝒰ia\\mathcal\{U\}\_\{ia\}byco\(𝒰ia\)\\operatorname\{co\}\(\\mathcal\{U\}\_\{ia\}\)leaves the discounted, finite\-horizon, and robust average values unchanged, both for optimal control and for every fixed stationary randomized controller\. It also leaves the set of average\-optimal stationary controllers unchanged\.
2. \(ii\)With the ambiguity sets fixed, writeg⋆\(r\)g^\{\\star\}\(r\)for the optimal gain under reward arrayrr\. Ifδ=maxi,a\|ria−r~ia\|\\delta=\\max\_\{i,a\}\|r\_\{ia\}\-\\widetilde\{r\}\_\{ia\}\|, then ‖g⋆\(r\)−g⋆\(r~\)‖∞≤δ\.\\left\\\|g^\{\\star\}\(r\)\-g^\{\\star\}\(\\widetilde\{r\}\)\\right\\\|\_\{\\infty\}\\leq\\delta\.The same bound holds for every fixed\-policy gain\. The gain is monotone in rewards, satisfiesg⋆\(r\+c\)=g⋆\(r\)\+c𝟏g^\{\\star\}\(r\+c\)=g^\{\\star\}\(r\)\+c\\mathbf\{1\}for any scalar reward shiftcc, and satisfiesg⋆\(αr\)=αg⋆\(r\)g^\{\\star\}\(\\alpha r\)=\\alpha g^\{\\star\}\(r\)forα≥0\\alpha\\geq 0\.
3. \(iii\)If𝒰ia⊆𝒰~ia\\mathcal\{U\}\_\{ia\}\\subseteq\\widetilde\{\\mathcal\{U\}\}\_\{ia\}at every pair, theng⋆\(r,𝒰~\)≤g⋆\(r,𝒰\)g^\{\\star\}\(r,\\widetilde\{\\mathcal\{U\}\}\)\\leq g^\{\\star\}\(r,\\mathcal\{U\}\); the same order holds for every fixed\-policy gain\.
###### Proof\.
Linear minimization over a set and over its convex hull gives the same value\. Finite\-dimensional compactness makesco\(𝒰ia\)\\operatorname\{co\}\(\\mathcal\{U\}\_\{ia\}\)compact, so convexification leavesTT, everyTπT^\{\\pi\}, and their discounted and finite\-horizon values unchanged\. Their normalized limits identify the same gains\. Since a stationary policy is all\-state average optimal exactly whengπ=g⋆g^\{\\pi\}=g^\{\\star\}, the set of such policies is unchanged as well\.
For rewards at distanceδ\\deltain supremum norm,Tr~x−δ𝟏≤Trx≤Tr~x\+δ𝟏T\_\{\\widetilde\{r\}\}x\-\\delta\\mathbf\{1\}\\leq T\_\{r\}x\\leq T\_\{\\widetilde\{r\}\}x\+\\delta\\mathbf\{1\}\. Monotonicity and scalar additive homogeneity give, by induction,
Tr~N0−Nδ𝟏≤TrN0≤Tr~N0\+Nδ𝟏\.T\_\{\\widetilde\{r\}\}^\{N\}0\-N\\delta\\mathbf\{1\}\\leq T\_\{r\}^\{N\}0\\leq T\_\{\\widetilde\{r\}\}^\{N\}0\+N\\delta\\mathbf\{1\}\.Dividing byNNand taking the finite\-horizon gain limits proves the Lipschitz bound\. The same comparison without an error term proves reward monotonicity\. The identitiesTr\+cN0=TrN0\+Nc𝟏T\_\{r\+c\}^\{N\}0=T\_\{r\}^\{N\}0\+Nc\\mathbf\{1\}andTαrN0=αTrN0T\_\{\\alpha r\}^\{N\}0=\\alpha T\_\{r\}^\{N\}0forα≥0\\alpha\\geq 0prove scalar shifts and positive homogeneity upon normalization\.
Larger ambiguity sets decrease every row minimum\. Monotonicity propagates this operator order through all finite\-horizon iterates, and normalized limits give the gain order\. Each argument applies unchanged toTπT^\{\\pi\}, whose outer weights are nonnegative and sum to one\. ∎
###### Propsition 12\.
The optimal gain satisfiesg⋆=T^g⋆g^\{\\star\}=\\widehat\{T\}g^\{\\star\}\. For every fixed stationary randomized controller,
giπ=∑aπ\(a∣i\)minp∈𝒰iap⊤gπ\.g\_\{i\}^\{\\pi\}=\\sum\_\{a\}\\pi\(a\\mid i\)\\min\_\{p\\in\\mathcal\{U\}\_\{ia\}\}p^\{\\top\}g^\{\\pi\}\.These gain\-only equations do not determine the average reward\.
###### Proof\.
Setvϵ=ϵVϵ→g⋆v\_\{\\epsilon\}=\{\\epsilon\}V\_\{\\epsilon\}\\to g^\{\\star\}\. Its discounted equation implies
‖vϵ−T^g⋆‖∞≤ϵR\+\(1−ϵ\)‖vϵ−g⋆‖∞\+ϵ‖g⋆‖∞⟶0\.\\\|v\_\{\\epsilon\}\-\\widehat\{T\}g^\{\\star\}\\\|\_\{\\infty\}\\leq\{\\epsilon\}R\+\(1\-\{\\epsilon\}\)\\\|v\_\{\\epsilon\}\-g^\{\\star\}\\\|\_\{\\infty\}\+\{\\epsilon\}\\\|g^\{\\star\}\\\|\_\{\\infty\}\\longrightarrow 0\.This uniform bound follows from\|p⊤x\|≤‖x‖∞\|p^\{\\top\}x\|\\leq\\\|x\\\|\_\{\\infty\}and is preserved by row minima and action maxima\. It identifies the limit asg⋆=T^g⋆g^\{\\star\}=\\widehat\{T\}g^\{\\star\}\. For fixedπ\\pi, the weighted action sum preserves the same error bound and gives the displayed evaluation equation\.
Every constant vectorc𝟏c\\mathbf\{1\}satisfies either recession equation, independently of rewards\. In a one\-state model with rewardrr, however, the average gain isrr\. Thus the recession equation alone cannot identify the reward\-dependent gain\. ∎相似文章
鲁棒平均奖励马尔可夫决策过程:通过插入式归约实现极小极大最优学习
本文研究了鲁棒平均奖励马尔可夫决策过程的样本复杂度,在总变差不确定集下通过插入式归约导出了极小极大最优学习率。
面向状态依赖可行动作集的马尔可夫决策过程的Bellman-Taylor Score Decoding
本文介绍了Bellman-Taylor Score Decoding,一种用于处理马尔可夫决策过程中状态依赖可行动作集的方法,解决了将深度强化学习应用于运筹学问题的一个关键挑战。
面向多目标强化学习的确定性帕累托最优策略综合
本文引入了一种基于切比雪夫标量化的新颖偏好条件贝尔曼算子,用于计算多目标马尔可夫决策过程中的确定性帕累托最优策略,并证明了该算子的收敛性及其在捕获完整帕累托前沿方面的有效性。
具有重尾奖励和信息不对称的鲁棒多智能体多臂老虎机
本文研究了三种信息不对称机制下具有重尾奖励的多智能体多臂老虎机问题,提出了鲁棒的分布式算法,其遗憾保证几乎匹配集中式算法的速率,并在帕累托分布奖励环境中进行了验证。
面向安全强化学习的鲁棒防护
提出了一种新颖的防护框架,用于鲁棒马尔可夫决策过程(RMDP),该框架在不确定的转移动态下正式保证安全性,并证明了其正确性和最优性。该方法结合了学习模型的PAC保证,使得在未知环境中实现安全强化学习成为可能。