Restless bandits with imperfect binary feedback: PCL-indexability analysis and computation

arXiv cs.LG Papers

Summary

This paper studies restless bandits with binary latent states and imperfect binary feedback, developing a partial conservation laws (PCL)-based framework for establishing indexability and computing the Whittle index, with applications to opportunistic spectrum access.

arXiv:2606.11192v1 Announce Type: new Abstract: We study restless bandits with binary latent states and imperfect binary feedback, motivated by opportunistic spectrum access with sensing errors. For the associated belief-state model, we develop a partial conservation laws (PCL)-based analytical and computational framework for establishing indexability and evaluating the Whittle index, building on a verification theorem for real-state discounted restless bandits. The framework analyzes the stochastic dynamics via an associated deterministic skeleton, renewal decompositions, and combinatorics on words. It yields tractable expressions for discounted reward and resource metrics in several threshold regimes, enabling full verification of the PCL-indexability conditions there. For the remaining regime, where a complete analytic verification is not achieved in this paper, we derive efficient numerical schemes for computing the relevant marginal metrics and the marginal productivity (MP) index, which equals the Whittle index when those conditions hold. Extensive computational experiments provide strong evidence that these conditions also hold in that regime across broad parameter ranges and without the stringent parameter restrictions imposed in prior work. The experiments further show that theMP index policy typically outperforms standard benchmark policies, often by a substantial margin.
Original Article
View Cached Full Text

Cached at: 06/11/26, 01:44 PM

# Restless bandits with imperfect binary feedback: PCL-indexability analysis and computation
Source: [https://arxiv.org/html/2606.11192](https://arxiv.org/html/2606.11192)
José Niño\-Mora Departamento de Estadística Universidad Carlos III de Madrid 28903 Getafe \(Madrid\), Spain [jose\.nino@uc3m\.es](https://arxiv.org/html/2606.11192v1/mailto:[email protected]) [https://alum\.mit\.edu/www/jnimora](https://alum.mit.edu/www/jnimora) [http://orcid\.org/0000\-0002\-2172\-3983](http://orcid.org/0000-0002-2172-3983)The author’s work was supported in part by Universidad Carlos III de Madrid \(UC3M\) through an internal research program grant and a grant for the acquisition of research tools\.

\(Submitted 27/3/2026\)

###### Abstract

We study restless bandits with binary latent states and imperfect binary feedback, motivated by opportunistic spectrum access with sensing errors\. For the associated belief\-state model, we develop a partial conservation laws \(PCL\)\-based analytical and computational framework for establishing indexability and evaluating the Whittle index, building on a verification theorem for real\-state discounted restless bandits\. The framework analyzes the stochastic dynamics via an associated deterministic skeleton, renewal decompositions, and combinatorics on words\. It yields tractable expressions for discounted reward and resource metrics in several threshold regimes, enabling full verification of the PCL\-indexability conditions there\. For the remaining regime, where a complete analytic verification is not achieved in this paper, we derive efficient numerical schemes for computing the relevant marginal metrics and the marginal productivity \(MP\) index, which equals the Whittle index when those conditions hold\. Extensive computational experiments provide strong evidence that these conditions also hold in that regime across broad parameter ranges and without the stringent parameter restrictions imposed in prior work\. The experiments further show that theMP index policy typically outperforms standard benchmark policies, often by a substantial margin\. Keywords:restless multi\-armed bandits; partial observability; sensing errors; Whittle index; indexability; conservation laws; opportunistic spectrum access

## 1Introduction

### 1\.1Background and motivation

Many sequential resource\-allocation problems involve a collection of stochastically evolving entities, here called*projects*, whose latent states are only partially observed\. In each discrete\-time period, the controller may activate only a limited number of projects, thereby obtaining information and potentially earning reward\. We consider a practically relevant setting in which each project’s latent state is binary \(bad/good\), and activation yields imperfect,*one\-sided*binary ACK/NACK feedback: an ACK certifies a good state and yields immediate reward, whereas a NACK \(no observed success\) yields no reward and is inherently ambiguous, since it may occur even when the latent state is good\. Because projects continue to evolve when not selected, the controller must balance immediate exploitation against information gathering and future opportunities\. This exploration–exploitation tradeoff leads naturally to a*restless multi\-armed bandit problem*\(RMABP\), a widely studied modeling framework introduced byWhittle\[[36](https://arxiv.org/html/2606.11192#bib.bib36)\]; see, e\.g\.,\[[17](https://arxiv.org/html/2606.11192#bib.bib17);[22](https://arxiv.org/html/2606.11192#bib.bib22);[23](https://arxiv.org/html/2606.11192#bib.bib23);[15](https://arxiv.org/html/2606.11192#bib.bib15);[19](https://arxiv.org/html/2606.11192#bib.bib19);[5](https://arxiv.org/html/2606.11192#bib.bib5)\]for representative work on partially observed restless bandits, and\[[32](https://arxiv.org/html/2606.11192#bib.bib32)\]for a broader review\.

We study a partially observed RMABP in whichNNprojects, labeled byn∈𝒩≜\{1,…,N\}n\\in\\mathscr\{N\}\\triangleq\\\{1,\\ldots,N\\\}, have latent binary states that evolve exogenously as Markov chains, independently of the controller’s actions\. In each period, the controller chooses which projects to activate, subject to an activation\-capacity constraint; nonselected projects are passive\. Passive actions yield neither feedback nor reward, and the corresponding belief states are updated predictively using the latent\-state dynamics\. By contrast, when projectnnis activated, it generates an ACK with probability0<κn<10<\\kappa\_\{n\}<1conditional on the latent state being good, and the ACK yields rewardrn\>0r\_\{n\}\>0\. Upon observing ACK/NACK, the controller updates its posterior belief via Bayes’ rule\. LetXn​\(t\)∈\[0,1\]X\_\{n\}\(t\)\\in\[0,1\]denote the posterior probability that projectnnis in the good state at the start of periodtt\. The vector𝐗​\(t\)=\(Xn​\(t\)\)n∈𝒩\\mathbf\{X\}\(t\)=\(X\_\{n\}\(t\)\)\_\{n\\in\\mathscr\{N\}\}is a sufficient statistic, so the overall control problem becomes a high\-dimensional*Markov decision process*\(MDP\) on the belief\-state space, induced by partial observability\[[16](https://arxiv.org/html/2606.11192#bib.bib16)\]\.

Because belief\-state RMABPs are generally intractable to solve exactly, the main goal of this paper is to exploit the special structure of the model—action\-independent latent\-state transitions, one\-sided ACK/NACK feedback, and ACK\-dependent rewards—to apply the*partial conservation laws*\(PCL\)\-based verification theorem ofNiño\-Mora\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]to this setting\. That theorem provides sufficient conditions for indexability and Whittle index evaluation in discounted real\-state restless bandits\. In several threshold regimes, we obtain full analytical verification of the required*PCL\-indexability conditions*\. For the remaining regime, in which a complete analysis remains elusive, we derive efficient computational schemes and use them to provide broad numerical evidence that those conditions also hold\. We further compute the associated*marginal productivity*\(MP\) index, which equals the Whittle index under the PCL conditions\.

This problem structure is motivated by several modern engineering applications\.*Opportunistic spectrum access*\(OSA\) in cognitive radio networks is a natural example\. Each project corresponds to a primary\-user \(PU\) channel whose availability evolves as a two\-state Markov chain, and the controller senses up toMMchannels in each period\. Under sensing errors and an exogenous collision\-tolerance constraint, secondary users \(SUs\) use the sensing outcomes to decide whether to transmit\. By the separation principle ofChen et al\.\[[10](https://arxiv.org/html/2606.11192#bib.bib10)\], the access decision is myopically optimal conditional on the sensing outcome, so the effect of sensing errors on each channel can be summarized in our reduced model by a single parameterκn\\kappa\_\{n\}, the maximum feasible transmission probability when channelnnis available\. This sensing\-access mechanism induces the one\-sided feedback structure in our model: an ACK certifies that the channel is available, whereasκn\\kappa\_\{n\}determines the effective ACK\-given\-available probability\. See Section[2\.2](https://arxiv.org/html/2606.11192#S2.SS2)for details\.

Related models also arise in link\- or server\-probing problems, where a controller sequentially tests or selects time\-varying resources whose availability evolves exogenously and earns reward only upon successful transmission or service; see, e\.g\.,\[[17](https://arxiv.org/html/2606.11192#bib.bib17);[24](https://arxiv.org/html/2606.11192#bib.bib24)\]\. Another example is remote diagnostics in distributed devices, where successful heartbeat checks confirm liveness whereas failed checks remain ambiguous yet informative; see, e\.g\.,\[[9](https://arxiv.org/html/2606.11192#bib.bib9);[2](https://arxiv.org/html/2606.11192#bib.bib2)\]\.

### 1\.2Related work on partially observed restless bandits, Whittle indices, and OSA

Related work most relevant to this paper falls into three strands: \(i\)*partially observed MDP*\(POMDP\) formulations and structural results for OSA with binary Markov channels; \(ii\) Whittle\-index approaches for sensing and scheduling problems with continuous belief states; and \(iii\) results on threshold optimality and indexability under imperfect binary feedback, including work closest to the one\-sided ACK/NACK setting studied here\.

#### 1\.2\.1OSA as a belief\-state POMDP

Control of partially observed binary Markov systems is naturally formulated as a belief\-state MDP, i\.e\., a POMDP whose sufficient statistic is the posterior probability of the good state\. In the OSA literature, early work studied optimal transmission on a single channel—often under the Gilbert–Elliott model—and established structural properties such as optimality of threshold policies\[[14](https://arxiv.org/html/2606.11192#bib.bib14)\]\. Subsequent work considered sensing errors and times, and delay or energy costs through constrained POMDP formulations and heuristic approaches\[[13](https://arxiv.org/html/2606.11192#bib.bib13)\], while multichannel OSA with sensing errors and collision constraints was formulated as a POMDP in\[[37](https://arxiv.org/html/2606.11192#bib.bib37)\]\. A key later development is the*separation principle*ofChen et al\.\[[10](https://arxiv.org/html/2606.11192#bib.bib10)\], which decouples sensing and access decisions in a broad class of imperfect\-sensing OSA models\.

#### 1\.2\.2From POMDPs to RMABPs and Whittle indices

Because such POMDPs are generally intractable to solve exactly, a substantial literature exploits their structure as RMABPs with continuous belief states\. In the classic*rested*case, the Gittins index yields an optimal index policy\[[12](https://arxiv.org/html/2606.11192#bib.bib12)\], but OSA and many monitoring and scheduling problems are*restless*\. This motivates using the*Whittle index*\[[36](https://arxiv.org/html/2606.11192#bib.bib36)\], which provides a scalable heuristic policy and is often accompanied by computable performance bounds\. Within OSA,Ahmad et al\.\[[4](https://arxiv.org/html/2606.11192#bib.bib4)\]established conditions under which*myopic sensing*is optimal, whileLiu and Zhao\[[17](https://arxiv.org/html/2606.11192#bib.bib17)\]established*indexability*and derived closed\-form Whittle\-index expressions in the perfect\-sensing case\. Related imperfect\-detection models were analyzed in\[[18](https://arxiv.org/html/2606.11192#bib.bib18)\], although indexability was not addressed there\.

More broadly, restless bandits with imperfect, limited, or constrained observations have been studied in frameworks encompassing OSA\-type models, including hidden\-Markov and constrained\-feedback formulations\[[22](https://arxiv.org/html/2606.11192#bib.bib22);[23](https://arxiv.org/html/2606.11192#bib.bib23);[15](https://arxiv.org/html/2606.11192#bib.bib15)\], as well as approaches aimed at low\-complexity Whittle\-index computation under imperfect feedback\[[19](https://arxiv.org/html/2606.11192#bib.bib19)\]\. Recent work has also considered data\-driven and risk\-aware RMAB extensions\[[20](https://arxiv.org/html/2606.11192#bib.bib20);[21](https://arxiv.org/html/2606.11192#bib.bib21)\]; see also\[[1](https://arxiv.org/html/2606.11192#bib.bib1);[5](https://arxiv.org/html/2606.11192#bib.bib5)\]\.

#### 1\.2\.3Closest work under imperfect binary feedback

The work most closely related to our setting addresses threshold optimality and/or Whittle indexability for belief\-state RMABPs with imperfect binary feedback, typically under parameter restrictions\.Meshram et al\.\[[23](https://arxiv.org/html/2606.11192#bib.bib23)\]study a broader model allowing action\-dependent transitions and introduce*approximate indexability*\. Their threshold\-optimality and indexability results hold only under stringent parameter restrictions; when specialized to our setting, these reduce to low autocorrelation\(0≤ρ≤1/5\)\(0\\leq\\rho\\leq 1/5\)and strong discounting\(β<1/3\)\(\\beta<1/3\)\.

Wang et al\.\[[35](https://arxiv.org/html/2606.11192#bib.bib35)\]study the OSA special case of our model and propose a fixed\-point/region\-partition analysis of the induced nonlinear belief dynamics, together with a piecewise\-linearization argument, to support indexability and closed\-form index expressions\. Their treatment highlights the role of piecewise threshold dynamics, but several continuity and monotonicity issues across regime boundaries are handled only implicitly\. The follow\-up workWang and Chen\[[34](https://arxiv.org/html/2606.11192#bib.bib34), Ch\. 3\]provides additional detail\. Their analysis focuses on selected iterates of the deterministic belief\-update mapsϕ0\\phi^\{0\}andϕ1\\phi^\{1\}\(see Section[4\.1](https://arxiv.org/html/2606.11192#S4.SS1)\)\. By contrast, threshold policies in our setting generate a broader family of mixed compositions of these maps, together with ACK\-induced resets in the one\-sided model, and this richer structure drives the piecewise and discontinuous behavior of the resulting metrics\. Their sufficient sensing\-error condition, in terms of the false\-alarm probabilityϵ\\epsilon\(whereκ=1−ϵ\\kappa=1\-\\epsilon\), specializes in our notation to

ϵ≤p01​p10\(1−p01\)​\(1−p10\)\.\\epsilon\\leq\\frac\{p\_\{01\}p\_\{10\}\}\{\(1\-p\_\{01\}\)\(1\-p\_\{10\}\)\}\.This can be quite restrictive whenρ\>0\\rho\>0, since the right\-hand side may be arbitrarily small\.

Kaza et al\.\[[15](https://arxiv.org/html/2606.11192#bib.bib15)\]study a hidden\-state RMABP with ACK/NACK feedback and sparse observations\. They establish threshold optimality for single\-project subproblems under strong conditions \(Theorem 1\), requiring low autocorrelation\(0<ρ<b/5\)\(0<\\rho<b/5\)and strong discounting\(β<b/5\)\(\\beta<b/5\), whereb=min⁡\{1,r\}b=\\min\\\{1,r\\\}in our model\. Their indexability claim \(Theorem 2\) is stated under the conditionβ<1/3\\beta<1/3, but the preceding development still relies on threshold optimality\.

Liu et al\.\[[19](https://arxiv.org/html/2606.11192#bib.bib19)\]also consider the OSA case of our model, assuming positive channel autocorrelation\. They give a sufficient condition for threshold optimality,

β≤1\(3−ϵ\)​ρ=1\(2\+κ\)​ρ,\\beta\\leq\\frac\{1\}\{\(3\-\\epsilon\)\\rho\}=\\frac\{1\}\{\(2\+\\kappa\)\\rho\},derive sufficient indexability conditions, and show that the model is indexable forβ≤1/2\\beta\\leq 1/2\. They also propose a low\-complexity algorithm for*approximate*Whittle\-index computation under the condition

β≤min⁡\{1\(2\+κ\)​ρ,12\},\\beta\\leq\\min\\Bigl\\\{\\frac\{1\}\{\(2\+\\kappa\)\\rho\},\\,\\frac\{1\}\{2\}\\Bigr\\\},which ensures both threshold optimality and indexability\. This restriction can still be stringent, since1/\(\(2\+κ\)​ρ\)1/\(\(2\+\\kappa\)\\rho\)can be close to1/31/3asκ↗1\\kappa\\nearrow 1andρ↗1\\rho\\nearrow 1\.

Finally, PCL methods have previously been applied to OSA\-type models with binary success feedback\. The perfect\-feedback caseκ=1\\kappa=1was treated via PCLs in\[[29](https://arxiv.org/html/2606.11192#bib.bib29)\]and later given a complete analysis in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), Section 12\.2\]\. The imperfect\-feedback case0<κ<10<\\kappa<1was outlined in\[[30](https://arxiv.org/html/2606.11192#bib.bib30)\], including a regime\-wise decomposition of threshold dynamics, but without proofs of the PCL verification steps\. The present paper builds on that line of work, develops the missing computational machinery, and uses it as a basis for a broad study of indexability and of the performance of the resulting MP index policy\.

Beyond the OSA literature,Dance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]gave the first proof of both Whittle indexability and threshold\-policy optimality for Kalman\-filter restless bandits, using the real\-state PCL verification theorem of\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]\. As further context, see\[[33](https://arxiv.org/html/2606.11192#bib.bib33)\]for a PCL\-based complete indexability analysis of a belief\-state adherence RMAB model\.

The present paper extends this PCL program for real\-state bandits to a one\-sided imperfect\-observation RMAB setting\. Existing analytical results establish indexability only under fairly restrictive parameter conditions, which could suggest that such restrictions are intrinsic to the model\. By contrast, our numerical results provide strong evidence that indexability persists over a much broader parameter range than is currently supported analytically\.

### 1\.3Goals, approach, and contributions

Our main goal is to develop an analytical and computational framework for studying indexability and evaluating the MP index in belief\-state restless bandits with one\-sided ACK/NACK feedback and ACK\-dependent rewards\. Relative to much of the OSA literature reviewed above, our work differs both in emphasis and in methodology\. Although we also exploit fixed points and regime\-wise decompositions of the belief dynamics, our analysis is organized around the PCL\-based verification theorem ofNiño\-Mora\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\], which reduces Whittle indexability and threshold\-policy optimality to the verification of three conditions for certain project performance metrics under threshold policies\.

A common route in the RMAB literature is first to establish optimality of threshold policies for single\-project Lagrangian subproblems, and then to deduce indexability and evaluate the Whittle index from the associated optimal\-threshold map and its monotonicity\. By contrast, the PCL framework in\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]yields the MP index directly as a ratio of marginal reward and marginal work metrics under threshold policies, thereby bypassing the need to invert a threshold map\. In the present model, full analytical verification of \(PCLI1–PCLI3\) remains elusive over part of the threshold range, because thresholding and ACK\-induced resets generate intricate belief dynamics\. Nevertheless, we show that the PCL route leads to a tractable and robust computational methodology: substantial parts of \(PCLI1–PCLI3\) can be verified analytically, while the remaining parts can be investigated systematically by numerical means, yielding strong evidence in support of indexability over a broader parameter range than is currently supported analytically\.

A terminological distinction is worth making explicit\. The*Whittle index*is defined implicitly, as the subsidy for passivity that makes active and passive actions equally desirable in the corresponding single\-project Lagrangian problem\. By contrast, the*MP index*is defined explicitly as a ratio of marginal reward and marginal work metrics under threshold policies\. Under \(PCLI1–PCLI3\) over the full threshold range, the verification theorem in\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]implies that the MP index coincides with the Whittle index\. Since full analytical verification of \(PCLI1–PCLI3\) remains open in some regimes, our numerical computations are formally for the MP index\. Thus, throughout the paper we refer to the computed ratio as the MP index, and to the induced policy as the MP index policy; if the PCL conditions were verified over the full threshold range, this would ensure that the MP index is indeed the Whittle index\.

This PCL program was initiated for the present model in our conference paper\[[30](https://arxiv.org/html/2606.11192#bib.bib30)\]\. The present paper develops that line of work into a coherent analytical and computational framework\. Our main contributions are as follows:

1. \(C1\)We develop a regime\-wise decomposition of the threshold\-induced belief dynamics and derive closed\-form expressions and renewal\-type identities for the key reward and work metrics, tailored to the one\-sided success\-feedback structure\. This yields explicit formulas in tractable regimes and efficient evaluation schemes elsewhere\.
2. \(C2\)We obtain a direct PCL\-based procedure for evaluating the MP index, and identify threshold regimes in which the sufficient conditions \(PCLI1–PCLI3\) can be fully verified analytically\.
3. \(C3\)We provide broad numerical evidence for the PCL conditions, and hence for indexability, well beyond the stringent discount\-factor and autocorrelation restrictions imposed in prior work: we carry out numerical tests for detecting violations of the PCL conditions, and do not observe any on the parameter grids explored\.
4. \(C4\)We carry out extensive computational experiments showing that the resulting MP index policy is readily computable and typically outperforms the benchmark policies considered, often by a substantial margin\. At the same time, the observed gaps to a Lagrangian dual upper bound can remain substantial, and experiments with two\-type instances do not provide numerical evidence of asymptotic optimality\.

Finally, to help clarify a possible route to a complete indexability proof in the remaining threshold regime, we connect the threshold\-induced belief dynamics with tools from*combinatorics on words*\(see, e\.g\.,\[[7](https://arxiv.org/html/2606.11192#bib.bib7)\]\), in particular Christoffel–Sturmian structure\. In the spirit ofDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\], this points to a structural program for analyzing the ordering of action itineraries induced by threshold policies\. In the present paper, however, the actual computation of the reward and work metrics relies primarily on renewal\-type decompositions rather than on a full characterization of that word structure\. Further details on this symbolic\-dynamics viewpoint are collected in Appendix[C](https://arxiv.org/html/2606.11192#A3)\.

### 1\.4Organization of the paper

The remainder of the paper is organized as follows\. Section[2](https://arxiv.org/html/2606.11192#S2)introduces the binary\-success\-feedback restless bandit model and shows how OSA with sensing errors and collision constraints fits within that framework\. Section[3](https://arxiv.org/html/2606.11192#S3)reviews restless\-bandit indexation via Lagrangian relaxation and presents the PCL\-indexability framework and verification theorem underlying our approach\. Section[4](https://arxiv.org/html/2606.11192#S4)develops the computational machinery for evaluating threshold\-policy performance metrics, based on no\-ACK skeleton dynamics and renewal decompositions, and identifies tractable threshold regimes together with explicit formulas for the associated marginal metrics and MP index\. Section[5](https://arxiv.org/html/2606.11192#S5)reports the computational experiments, including broad numerical evidence on the PCL\-indexability conditions and benchmarking of the resulting MP index policy against standard baseline policies and a Lagrangian upper bound\. Section[6](https://arxiv.org/html/2606.11192#S6)concludes and discusses directions for future work\.

Appendix[A](https://arxiv.org/html/2606.11192#A1)collects supplementary proofs for the computational machinery developed in Section[4](https://arxiv.org/html/2606.11192#S4)\. Appendix[B](https://arxiv.org/html/2606.11192#A2)presents an illustrative parameter instance together with diagnostic plots of the key metrics and the MP index\. Appendix[C](https://arxiv.org/html/2606.11192#A3)records the symbolic organization of intermediate\-regime threshold itineraries via Christoffel and Sturmian words\. Appendix[D](https://arxiv.org/html/2606.11192#A4)briefly outlines the corresponding time\-average criterion and its connection with the discounted analysis\. Appendix[E](https://arxiv.org/html/2606.11192#A5)collects supplementary parameter\-dependence plots and implementation details\.

## 2Binary\-success\-feedback restless bandit model and an OSA instantiation

### 2\.1A generic binary\-success\-feedback restless bandit model

We consider a collection ofNNprojects, labeled byn∈𝒩≜\{1,…,N\}n\\in\\mathscr\{N\}\\triangleq\\\{1,\\ldots,N\\\}, that evolve over discrete periodst=0,1,…t=0,1,\\ldots\. Projectnnhas a latent binary stateSn​\(t\)∈\{0,1\}S\_\{n\}\(t\)\\in\\\{0,1\\\}, where11represents a*good*state and0a*bad*state\. The latent state evolves as a two\-state Markov chain with transition probabilities

pn,i​j=ℙ​\{Sn​\(t\+1\)=j∣Sn​\(t\)=i\}\>0,i,j∈\{0,1\}\.p\_\{n,ij\}=\\mathbb\{P\}\\\{S\_\{n\}\(t\+1\)=j\\mid S\_\{n\}\(t\)=i\\\}\>0,\\qquad i,j\\in\\\{0,1\\\}\.We restrict attention to the positively correlated case

ρn≜1−pn,10−pn,01=pn,11−pn,01\>0\.\\rho\_\{n\}\\triangleq 1\-p\_\{n,10\}\-p\_\{n,01\}=p\_\{n,11\}\-p\_\{n,01\}\>0\.This is natural in settings where latent conditions are persistent: a channel, server, or device that is currently good \(respectively bad\) is typically more likely to remain so in the next period than to switch immediately\. In communication settings, this is the standard persistence captured by Gilbert–Elliott and related finite\-state Markov channel models for correlated fading and bursty error processes\[[14](https://arxiv.org/html/2606.11192#bib.bib14);[17](https://arxiv.org/html/2606.11192#bib.bib17)\]\. The latent\-state processes are independent across projects\.

At the start of each periodtt, the controller selects an actionAn​\(t\)∈\{0,1\}A\_\{n\}\(t\)\\in\\\{0,1\\\}for each projectnn, whereAn​\(t\)=1A\_\{n\}\(t\)=1\(*active*\) means that the project is selected or probed, andAn​\(t\)=0A\_\{n\}\(t\)=0\(*passive*\) means that it is not selected\. At mostMMprojects can be active in any period, so

∑n∈𝒩An​\(t\)≤M,t=0,1,2,…\.\\sum\_\{n\\in\\mathscr\{N\}\}A\_\{n\}\(t\)\\leq M,\\qquad t=0,1,2,\\ldots\.\(1\)
IfAn​\(t\)=1A\_\{n\}\(t\)=1, the controller observes a one\-sided binary feedback signalKn​\(t\)∈\{0,1\}K\_\{n\}\(t\)\\in\\\{0,1\\\}, interpreted as success/failure \(ACK/NACK\), whereKn​\(t\)=1K\_\{n\}\(t\)=1\(ACK\) certifies that the project is in the good stateSn​\(t\)=1S\_\{n\}\(t\)=1, whereasKn​\(t\)=0K\_\{n\}\(t\)=0\(NACK\) may occur in either latent state\. Specifically, for some0<κn<10<\\kappa\_\{n\}<1,

ℙ​\{Kn​\(t\)=1∣Sn​\(t\)=1,An​\(t\)=1\}=κn,ℙ​\{Kn​\(t\)=1∣Sn​\(t\)=0,An​\(t\)=1\}=0\.\\mathbb\{P\}\\\{K\_\{n\}\(t\)=1\\mid S\_\{n\}\(t\)=1,\\,A\_\{n\}\(t\)=1\\\}=\\kappa\_\{n\},\\qquad\\mathbb\{P\}\\\{K\_\{n\}\(t\)=1\\mid S\_\{n\}\(t\)=0,\\,A\_\{n\}\(t\)=1\\\}=0\.IfAn​\(t\)=0A\_\{n\}\(t\)=0, no feedback is received; for notational convenience, we setKn​\(t\)=0K\_\{n\}\(t\)=0in that case\. Upon an ACK, projectnnearns rewardrn\>0r\_\{n\}\>0, so the realized reward in periodttisrn​Kn​\(t\)r\_\{n\}K\_\{n\}\(t\)\.

Letℋt\\mathscr\{H\}\_\{t\}denote the global history available at the start of periodtt, consisting of the initial prior together with all past actions and observed feedback up to timett\. Since the latent states are unobserved, decisions are based on belief states\. For projectnn, the belief state is

Xn​\(t\)≜ℙ​\{Sn​\(t\)=1∣ℋt\}∈𝒳≜\[0,1\],X\_\{n\}\(t\)\\triangleq\\mathbb\{P\}\\\{S\_\{n\}\(t\)=1\\mid\\mathscr\{H\}\_\{t\}\\\}\\in\\mathscr\{X\}\\triangleq\[0,1\],and the vector𝐗​\(t\)=\(Xn​\(t\)\)n∈𝒩\\mathbf\{X\}\(t\)=\(X\_\{n\}\(t\)\)\_\{n\\in\\mathscr\{N\}\}is a sufficient statistic for control\. We consider scheduling policiesπ∈Π\\pi\\in\\Pithat select a feasible action vector𝐀​\(t\)\\mathbf\{A\}\(t\)as a function of𝐗​\(t\)\\mathbf\{X\}\(t\)\.

GivenXn​\(t\)=xnX\_\{n\}\(t\)=x\_\{n\}andAn​\(t\)=1A\_\{n\}\(t\)=1, Bayes’ rule yields

ℙ​\{Sn​\(t\)=1∣Kn​\(t\)=1,An​\(t\)=1\}=1,ℙ​\{Sn​\(t\)=1∣Kn​\(t\)=0,An​\(t\)=1\}=ψn​\(xn\)≜\(1−κn\)​xn1−κn​xn\.\\mathbb\{P\}\\\{S\_\{n\}\(t\)=1\\mid K\_\{n\}\(t\)=1,\\,A\_\{n\}\(t\)=1\\\}=1,\\qquad\\mathbb\{P\}\\\{S\_\{n\}\(t\)=1\\mid K\_\{n\}\(t\)=0,\\,A\_\{n\}\(t\)=1\\\}=\\psi\_\{n\}\(x\_\{n\}\)\\triangleq\\frac\{\(1\-\\kappa\_\{n\}\)x\_\{n\}\}\{1\-\\kappa\_\{n\}x\_\{n\}\}\.Thus, an ACK confirms a good latent state, whereas a NACK lowers the belief\.

The belief dynamics follow by combining this Bayesian update with the latent Markov transition\. Under the passive actionAn​\(t\)=0A\_\{n\}\(t\)=0, the update is deterministic:

Xn​\(t\+1\)=ϕn0​\(Xn​\(t\)\)≜pn,01\+ρn​Xn​\(t\)\.X\_\{n\}\(t\+1\)=\\phi\_\{n\}^\{0\}\(X\_\{n\}\(t\)\)\\triangleq p\_\{n,01\}\+\\rho\_\{n\}X\_\{n\}\(t\)\.\(2\)Under the active actionAn​\(t\)=1A\_\{n\}\(t\)=1, an ACK occurs at the end of periodttwith probability

ℙ​\{Kn​\(t\)=1∣Xn​\(t\)=xn,An​\(t\)=1\}=κn​xn\.\\mathbb\{P\}\\\{K\_\{n\}\(t\)=1\\mid X\_\{n\}\(t\)=x\_\{n\},\\,A\_\{n\}\(t\)=1\\\}=\\kappa\_\{n\}x\_\{n\}\.Hence the active belief update is

Xn​\(t\+1\)=\{pn,01\+ρn=pn,11,w\.p\.​κn​Xn​\(t\),ϕn1​\(Xn​\(t\)\)≜pn,01\+ρn​ψn​\(Xn​\(t\)\),w\.p\.​1−κn​Xn​\(t\)\.X\_\{n\}\(t\+1\)=\\begin\{cases\}p\_\{n,01\}\+\\rho\_\{n\}=p\_\{n,11\},&\\textup\{w\.p\. \}\\kappa\_\{n\}X\_\{n\}\(t\),\\\\\[5\.69054pt\] \\phi\_\{n\}^\{1\}\(X\_\{n\}\(t\)\)\\triangleq p\_\{n,01\}\+\\rho\_\{n\}\\,\\psi\_\{n\}\(X\_\{n\}\(t\)\),&\\textup\{w\.p\. \}1\-\\kappa\_\{n\}X\_\{n\}\(t\)\.\\end\{cases\}\(3\)
We focus primarily on the discounted\-reward criterion\. Given an initial belief vector𝐱\\mathbf\{x\}, the objective is

maxπ∈Π⁡𝔼𝐱π​\[∑t=0∞∑n∈𝒩βt​rn​Kn​\(t\)\],0<β<1,\\max\_\{\\pi\\in\\Pi\}\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\sum\_\{n\\in\\mathscr\{N\}\}\\beta^\{t\}\\,r\_\{n\}K\_\{n\}\(t\)\\right\],\\qquad 0<\\beta<1,\(4\)where𝔼𝐱π\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}is expectation under policyπ\\pistarting from𝐗​\(0\)=𝐱\\mathbf\{X\}\(0\)=\\mathbf\{x\}\. Equivalently, the expected one\-step reward is

Rn​\(xn,an\)≜𝔼​\[rn​Kn​\(t\)∣Xn​\(t\)=xn,An​\(t\)=an\]=rn​κn​xn​an,R\_\{n\}\(x\_\{n\},a\_\{n\}\)\\triangleq\\mathbb\{E\}\[r\_\{n\}K\_\{n\}\(t\)\\mid X\_\{n\}\(t\)=x\_\{n\},\\,A\_\{n\}\(t\)=a\_\{n\}\]=r\_\{n\}\\kappa\_\{n\}x\_\{n\}a\_\{n\},so \([4](https://arxiv.org/html/2606.11192#S2.E4)\) can be written as

maxπ∈Π⁡𝔼𝐱π​\[∑t=0∞∑n∈𝒩βt​Rn​\(Xn​\(t\),An​\(t\)\)\]\.\\max\_\{\\pi\\in\\Pi\}\\;\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\sum\_\{n\\in\\mathscr\{N\}\}\\beta^\{t\}\\,R\_\{n\}\(X\_\{n\}\(t\),A\_\{n\}\(t\)\)\\right\]\.\(5\)
For later reference, we also consider the average\-reward criterion

maxπ∈Π​lim infT→∞1T\+1​𝔼𝐱π​\[∑t=0T∑n∈𝒩Rn​\(Xn​\(t\),An​\(t\)\)\]\.\\max\_\{\\pi\\in\\Pi\}\\;\\liminf\_\{T\\to\\infty\}\\frac\{1\}\{T\+1\}\\,\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{T\}\\sum\_\{n\\in\\mathscr\{N\}\}R\_\{n\}\(X\_\{n\}\(t\),A\_\{n\}\(t\)\)\\right\]\.\(6\)Its connection with the discounted formulation is outlined in Appendix[D](https://arxiv.org/html/2606.11192#A4)\.

Problems \([4](https://arxiv.org/html/2606.11192#S2.E4)\) and \([6](https://arxiv.org/html/2606.11192#S2.E6)\) are partially observed RMABPs with continuous belief states and are generally intractable\. When the model is*indexable*, the*Whittle index policy*provides a scalable and well\-grounded heuristic: at each time, it activates up toMMprojects with the largest nonnegative Whittle indices\. In the PCL framework outlined later, we work instead with the MP index, which coincides with the Whittle index when conditions \(PCLI1–PCLI3\) hold\.

### 2\.2Illustrative example: OSA with sensing errors

We next show how the binary\-success\-feedback model of Section[2\.1](https://arxiv.org/html/2606.11192#S2.SS1)arises in opportunistic spectrum access \(OSA\) with sensing errors and collision constraints\. Consider a cognitive radio system withNNlicensed*primary\-user*\(PU\) channels, of which at mostMMcan be sensed by the controller in each slot\. Channelnnoffers throughputrnr\_\{n\}Mb/slot and has PU occupancy stateSn​\(t\)∈\{0,1\}S\_\{n\}\(t\)\\in\\\{0,1\\\}, whereSn​\(t\)=1S\_\{n\}\(t\)=1means that the channel is available to secondary users \(SUs\), andSn​\(t\)=0S\_\{n\}\(t\)=0means that it is unavailable \(busy\)\. If channelnnis sensed at timett, a binary sensor outcomeOn​\(t\)∈\{0,1\}O\_\{n\}\(t\)\\in\\\{0,1\\\}is observed, whereOn​\(t\)=1O\_\{n\}\(t\)=1andOn​\(t\)=0O\_\{n\}\(t\)=0indicate “sensed available” and “sensed busy,” respectively\. Sensing errors are modeled by positive miss\-detection and false\-alarm probabilitiesδn\>0\\delta\_\{n\}\>0andϵn\>0\\epsilon\_\{n\}\>0, defined by

δn≜ℙ​\{On​\(t\)=1∣Sn​\(t\)=0\},ϵn≜ℙ​\{On​\(t\)=0∣Sn​\(t\)=1\}\.\\delta\_\{n\}\\triangleq\\mathbb\{P\}\\\{O\_\{n\}\(t\)=1\\mid S\_\{n\}\(t\)=0\\\},\\qquad\\epsilon\_\{n\}\\triangleq\\mathbb\{P\}\\\{O\_\{n\}\(t\)=0\\mid S\_\{n\}\(t\)=1\\\}\.We assume that the sensor is informative\[[10](https://arxiv.org/html/2606.11192#bib.bib10)\], so

δn\+ϵn<1\.\\delta\_\{n\}\+\\epsilon\_\{n\}<1\.\(7\)
An access attempt whenSn​\(t\)=0S\_\{n\}\(t\)=0causes a collision and therefore an unsuccessful transmission\. Each channelnnimposes a collision\-tolerance requirement0<ζn<10<\\zeta\_\{n\}<1, namely,

ℙ​\{access at​t∣Sn​\(t\)=0,An​\(t\)=1\}≤ζn\.\\mathbb\{P\}\\\{\\text\{access at \}t\\mid S\_\{n\}\(t\)=0,\\,A\_\{n\}\(t\)=1\\\}\\leq\\zeta\_\{n\}\.
To avoid SU contention, the controller selects at mostMMchannels to be sensed in each slot, via actionsAn​\(t\)∈\{0,1\}A\_\{n\}\(t\)\\in\\\{0,1\\\}satisfying \([1](https://arxiv.org/html/2606.11192#S2.E1)\)\. Conditional on sensing, an*access rule*is specified by probabilitiesyn​\(o\)∈\[0,1\]y\_\{n\}\(o\)\\in\[0,1\],o∈\{0,1\}o\\in\\\{0,1\\\}, whereyn​\(o\)y\_\{n\}\(o\)is the probability of attempting access given sensor outcomeoo\. Then

ℙ​\{access at​t∣Sn​\(t\)=1,An​\(t\)=1\}\\displaystyle\\mathbb\{P\}\\\{\\text\{access at \}t\\mid S\_\{n\}\(t\)=1,\\,A\_\{n\}\(t\)=1\\\}=ϵn​yn​\(0\)\+\(1−ϵn\)​yn​\(1\),\\displaystyle=\\epsilon\_\{n\}y\_\{n\}\(0\)\+\(1\-\\epsilon\_\{n\}\)y\_\{n\}\(1\),ℙ​\{access at​t∣Sn​\(t\)=0,An​\(t\)=1\}\\displaystyle\\mathbb\{P\}\\\{\\text\{access at \}t\\mid S\_\{n\}\(t\)=0,\\,A\_\{n\}\(t\)=1\\\}=\(1−δn\)​yn​\(0\)\+δn​yn​\(1\)\.\\displaystyle=\(1\-\\delta\_\{n\}\)y\_\{n\}\(0\)\+\\delta\_\{n\}y\_\{n\}\(1\)\.
Following the*separation principle*ofChen et al\.\[[10](https://arxiv.org/html/2606.11192#bib.bib10)\], we choose the access probabilitiesyn​\(o\)y\_\{n\}\(o\)myopically so as to maximize the access probability when the channel is truly available, subject to the collision constraint:

κn≜max0≤yn​\(0\),yn​\(1\)≤1⁡\{ϵn​yn​\(0\)\+\(1−ϵn\)​yn​\(1\):\(1−δn\)​yn​\(0\)\+δn​yn​\(1\)≤ζn\}\.\\kappa\_\{n\}\\triangleq\\max\_\{0\\leq y\_\{n\}\(0\),\\,y\_\{n\}\(1\)\\leq 1\}\\Bigl\\\{\\epsilon\_\{n\}y\_\{n\}\(0\)\+\(1\-\\epsilon\_\{n\}\)y\_\{n\}\(1\):\(1\-\\delta\_\{n\}\)y\_\{n\}\(0\)\+\\delta\_\{n\}y\_\{n\}\(1\)\\leq\\zeta\_\{n\}\\Bigr\\\}\.\(8\)Thus,κn\\kappa\_\{n\}is the maximal probability of attempting access on channelnnwhen it is in fact available, while respecting the collision\-tolerance requirement\. Under \([7](https://arxiv.org/html/2606.11192#S2.E7)\), the optimizer admits the closed form \(see\[[10](https://arxiv.org/html/2606.11192#bib.bib10), Proposition 2\]\)

yn∗​\(1\)=min⁡\{1,ζnδn\},yn∗​\(0\)=\(ζn−δn1−δn\)\+,κn=ϵn​yn∗​\(0\)\+\(1−ϵn\)​yn∗​\(1\)\.y\_\{n\}^\{\*\}\(1\)=\\min\\\!\\left\\\{1,\\frac\{\\zeta\_\{n\}\}\{\\delta\_\{n\}\}\\right\\\},\\qquad y\_\{n\}^\{\*\}\(0\)=\\left\(\\frac\{\\zeta\_\{n\}\-\\delta\_\{n\}\}\{1\-\\delta\_\{n\}\}\\right\)^\{\+\},\\qquad\\kappa\_\{n\}=\\epsilon\_\{n\}y\_\{n\}^\{\*\}\(0\)\+\(1\-\\epsilon\_\{n\}\)y\_\{n\}^\{\*\}\(1\)\.\(9\)Ifδn=ζn\\delta\_\{n\}=\\zeta\_\{n\}\(the “trust\-the\-sensor” point inChen et al\.\[[10](https://arxiv.org/html/2606.11192#bib.bib10), Theorem 2\]\), then\(yn∗​\(0\),yn∗​\(1\)\)=\(0,1\)\(y\_\{n\}^\{\*\}\(0\),y\_\{n\}^\{\*\}\(1\)\)=\(0,1\)andκn=1−ϵn\\kappa\_\{n\}=1\-\\epsilon\_\{n\}\.

Under the collision\-tolerance condition0<ζn<10<\\zeta\_\{n\}<1, we haveyn∗​\(1\)=min⁡\{1,ζn/δn\}\>0y\_\{n\}^\{\*\}\(1\)=\\min\\\{1,\\zeta\_\{n\}/\\delta\_\{n\}\\\}\>0, and henceκn\>0\\kappa\_\{n\}\>0\. Further, the collision constraint rules out\(yn∗​\(0\),yn∗​\(1\)\)=\(1,1\)\(y\_\{n\}^\{\*\}\(0\),y\_\{n\}^\{\*\}\(1\)\)=\(1,1\), and since0<ϵn<10<\\epsilon\_\{n\}<1, this impliesκn<1\\kappa\_\{n\}<1\.

LetKn​\(t\)∈\{0,1\}K\_\{n\}\(t\)\\in\\\{0,1\\\}denote the ACK/NACK indicator, whereKn​\(t\)=1K\_\{n\}\(t\)=1if an SU attempts access on channelnnat timettand the channel is available, i\.e\.,Sn​\(t\)=1S\_\{n\}\(t\)=1\. Under the optimal access rule, conditioning onXn​\(t\)=xnX\_\{n\}\(t\)=x\_\{n\}gives

ℙ​\{Kn​\(t\)=1∣Xn​\(t\)=xn,An​\(t\)=1\}=κn​xn\.\\mathbb\{P\}\\\{K\_\{n\}\(t\)=1\\mid X\_\{n\}\(t\)=x\_\{n\},\\,A\_\{n\}\(t\)=1\\\}=\\kappa\_\{n\}x\_\{n\}\.Moreover, a successful SU transmission on channelnnyields throughputrnr\_\{n\}Mb/slot, so the one\-slot expected reward is

Rn​\(xn,an\)≜rn​κn​xn​an\.R\_\{n\}\(x\_\{n\},a\_\{n\}\)\\triangleq r\_\{n\}\\kappa\_\{n\}x\_\{n\}a\_\{n\}\.Thus, in the OSA interpretation, the good latent state corresponds to channel availability, the ACK event corresponds to successful transmission, andκn\\kappa\_\{n\}summarizes the optimal sensing\-conditioned access rule under the collision constraint\.

## 3Restless bandit indexation

### 3\.1Indexability, dual bound, and Whittle index policy

Problem \([4](https://arxiv.org/html/2606.11192#S2.E4)\) is a discounted\-reward RMABP with the per\-period activation\-capacity constraint \([1](https://arxiv.org/html/2606.11192#S2.E1)\)\. Since computing an optimal policy is generally intractable, we adopt Whittle’s index approach\[[36](https://arxiv.org/html/2606.11192#bib.bib36)\], based on a Lagrangian relaxation and decomposition into single\-project subproblems\.

LetΠ^\\widehat\{\\Pi\}denote the class of stationary randomized policies that, at each timett, select an action vector𝐀​\(t\)∈\{0,1\}N\\mathbf\{A\}\(t\)\\in\\\{0,1\\\}^\{N\}as a possibly randomized function of the current belief vector𝐗​\(t\)\\mathbf\{X\}\(t\), without being required to satisfy the per\-period constraint \([1](https://arxiv.org/html/2606.11192#S2.E1)\)\. We relax \([1](https://arxiv.org/html/2606.11192#S2.E1)\) by imposing only that the expected total discounted activation effort not exceedM/\(1−β\)M/\(1\-\\beta\):

𝔼𝐱π​\[∑t=0∞∑n∈𝒩βt​An​\(t\)\]⩽M1−β\.\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\sum\_\{n\\in\\mathscr\{N\}\}\\beta^\{t\}A\_\{n\}\(t\)\\right\]\\;\\leqslant\\;\\frac\{M\}\{1\-\\beta\}\.\(10\)The resulting relaxed problem has optimal value

V^∗​\(𝐱\)≜maxπ∈Π^⁡\{𝔼𝐱π​\[∑t=0∞∑n∈𝒩Rn​\(Xn​\(t\),An​\(t\)\)​βt\]:\([10](https://arxiv.org/html/2606.11192#S3.E10)\)\}\.\\widehat\{V\}^\{\*\}\(\\mathbf\{x\}\)\\;\\triangleq\\;\\max\_\{\\pi\\in\\widehat\{\\Pi\}\}\\left\\\{\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\sum\_\{n\\in\\mathscr\{N\}\}R\_\{n\}\\bigl\(X\_\{n\}\(t\),A\_\{n\}\(t\)\\bigr\)\\,\\beta^\{t\}\\right\]\\;:\\;\\eqref\{eq:avversac\}\\right\\\}\.\(11\)SinceΠ⊆Π^\\Pi\\subseteq\\widehat\{\\Pi\}and \([1](https://arxiv.org/html/2606.11192#S2.E1)\) implies \([10](https://arxiv.org/html/2606.11192#S3.E10)\), it follows thatV∗​\(𝐱\)⩽V^∗​\(𝐱\)\.V^\{\*\}\(\\mathbf\{x\}\)\\leqslant\\widehat\{V\}^\{\*\}\(\\mathbf\{x\}\)\.

Attaching a multiplierλ\\lambdato \([10](https://arxiv.org/html/2606.11192#S3.E10)\), define

L​\(𝐱;λ\)≜maxπ∈Π^⁡𝔼𝐱π​\[∑t=0∞∑n∈𝒩\(Rn​\(Xn​\(t\),An​\(t\)\)−λ​An​\(t\)\)​βt\]\+M1−β​λ\.L\(\\mathbf\{x\};\\lambda\)\\;\\triangleq\\;\\max\_\{\\pi\\in\\widehat\{\\Pi\}\}\\mathbb\{E\}\_\{\\mathbf\{x\}\}^\{\\pi\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\sum\_\{n\\in\\mathscr\{N\}\}\\bigl\(R\_\{n\}\\bigl\(X\_\{n\}\(t\),A\_\{n\}\(t\)\\bigr\)\-\\lambda\\,A\_\{n\}\(t\)\\bigr\)\\,\\beta^\{t\}\\right\]\\;\+\\;\\frac\{M\}\{1\-\\beta\}\\,\\lambda\.\(12\)Forλ⩾0\\lambda\\geqslant 0, this is a Lagrangian relaxation of \([11](https://arxiv.org/html/2606.11192#S3.E11)\), and weak duality yieldsV^∗​\(𝐱\)⩽L​\(𝐱;λ\)\.\\widehat\{V\}^\{\*\}\(\\mathbf\{x\}\)\\leqslant L\(\\mathbf\{x\};\\lambda\)\.Minimizing overλ⩾0\\lambda\\geqslant 0gives the following Lagrangian dual bound, which upper\-bounds bothV^∗​\(𝐱\)\\widehat\{V\}^\{\*\}\(\\mathbf\{x\}\)andV∗​\(𝐱\)V^\{\*\}\(\\mathbf\{x\}\):

D​\(𝐱\)≜minλ⩾0⁡L​\(𝐱;λ\),D\(\\mathbf\{x\}\)\\;\\triangleq\\;\\min\_\{\\lambda\\geqslant 0\}L\(\\mathbf\{x\};\\lambda\),\(13\)
The relaxation \([12](https://arxiv.org/html/2606.11192#S3.E12)\) decouples across projects into single\-project subproblems\. For eachn∈𝒩n\\in\\mathscr\{N\}, letΠn\\Pi\_\{n\}denote the class of stationary policies for projectnnin isolation\. Define the single\-project value under activity chargeλ\\lambdaby

Ln​\(xn;λ\)≜maxπn∈Πn⁡𝔼xnπn​\[∑t=0∞\(Rn​\(Xn​\(t\),An​\(t\)\)−λ​An​\(t\)\)​βt\]\.L\_\{n\}\(x\_\{n\};\\lambda\)\\;\\triangleq\\;\\max\_\{\\pi\_\{n\}\\in\\Pi\_\{n\}\}\\mathbb\{E\}\_\{x\_\{n\}\}^\{\\pi\_\{n\}\}\\\!\\left\[\\sum\_\{t=0\}^\{\\infty\}\\bigl\(R\_\{n\}\\bigl\(X\_\{n\}\(t\),A\_\{n\}\(t\)\\bigr\)\-\\lambda\\,A\_\{n\}\(t\)\\bigr\)\\,\\beta^\{t\}\\right\]\.\(14\)Then

L​\(𝐱;λ\)=∑n∈𝒩Ln​\(xn;λ\)\+M1−β​λ\.L\(\\mathbf\{x\};\\lambda\)\\;=\\;\\sum\_\{n\\in\\mathscr\{N\}\}L\_\{n\}\(x\_\{n\};\\lambda\)\\;\+\\;\\frac\{M\}\{1\-\\beta\}\\,\\lambda\.\(15\)
We say that projectnnis*indexable*if there exists a functionλn∗:𝒳→ℝ\\lambda\_\{n\}^\{\*\}:\\mathscr\{X\}\\to\\mathbb\{R\}such that, for everyxn∈𝒳x\_\{n\}\\in\\mathscr\{X\}andλ∈ℝ\\lambda\\in\\mathbb\{R\}, the active actionan=1a\_\{n\}=1is optimal in \([14](https://arxiv.org/html/2606.11192#S3.E14)\) at belief statexnx\_\{n\}if and only ifλn∗​\(xn\)≥λ\\lambda\_\{n\}^\{\*\}\(x\_\{n\}\)\\geq\\lambda, while the passive actionan=0a\_\{n\}=0is optimal atxnx\_\{n\}if and only ifλn∗​\(xn\)≤λ\\lambda\_\{n\}^\{\*\}\(x\_\{n\}\)\\leq\\lambda\. Thus both actions are optimal atxnx\_\{n\}if and only ifλn∗​\(xn\)=λ\\lambda\_\{n\}^\{\*\}\(x\_\{n\}\)=\\lambda\. This statewise formulation is the one adopted in our PCL verification framework\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]\. It is equivalent to Whittle’s original set\-expansion definition when the indifference charge is unique at each state\.

Under indexability, Whittle’s policy activates at each periodttup toMMprojects with the largest nonnegative indicesλn∗​\(Xn​\(t\)\)\\lambda\_\{n\}^\{\*\}\\\!\\bigl\(X\_\{n\}\(t\)\\bigr\)\. Later, within the PCL framework, we will introduce the MP index as an explicit candidate index; when the PCL conditions hold, it coincides with the Whittle index\.

### 3\.2PCL\-based verification theorem for threshold\-indexability

To apply Whittle’s index approach\[[36](https://arxiv.org/html/2606.11192#bib.bib36)\], indexability must first be established for the single\-projectλ\\lambda\-charge subproblems\. Rather than following the conventional route—first proving optimality of threshold policies for each subproblem \([14](https://arxiv.org/html/2606.11192#S3.E14)\), and then establishing monotonicity of the optimal threshold as a function ofλ\\lambda—we adopt the alternative PCL\-based approach\. In the discrete\-state setting, PCLs yield general sufficient indexability conditions and an index algorithm; see\[[25](https://arxiv.org/html/2606.11192#bib.bib25);[26](https://arxiv.org/html/2606.11192#bib.bib26);[28](https://arxiv.org/html/2606.11192#bib.bib28);[27](https://arxiv.org/html/2606.11192#bib.bib27)\]\. Here we use the PCL\-indexability framework for real\-state projects in\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\]\.

We henceforth focus on a single project and suppress its label\. LetΠ\\Pidenote the class of stationary policies for the single\-project belief process\. For anyπ∈Π\\pi\\in\\Piand initial beliefx∈𝒳x\\in\\mathscr\{X\}, define the discounted*reward*and*work*metrics

F​\(x,π\)≜𝔼xπ​\[∑t=0∞R​\(X​\(t\),A​\(t\)\)​βt\],G​\(x,π\)≜𝔼xπ​\[∑t=0∞A​\(t\)​βt\],F\(x,\\pi\)\\triangleq\\mathbb\{E\}^\{\\pi\}\_\{x\}\\\!\\Bigg\[\\sum\_\{t=0\}^\{\\infty\}R\\bigl\(X\(t\),A\(t\)\\bigr\)\\,\\beta^\{t\}\\Bigg\],\\qquad G\(x,\\pi\)\\triangleq\\mathbb\{E\}^\{\\pi\}\_\{x\}\\\!\\Bigg\[\\sum\_\{t=0\}^\{\\infty\}A\(t\)\\,\\beta^\{t\}\\Bigg\],\(16\)whereR​\(x,a\)≜a​r​κ​x,a∈\{0,1\}\.R\(x,a\)\\triangleq a\\,r\\,\\kappa\\,x,\\kern 5\.0pta\\in\\\{0,1\\\}\.Given an activation chargeλ∈ℝ\\lambda\\in\\mathbb\{R\}, define the correspondingλ\\lambda\-charge performance by

ℒ​\(x,π;λ\)≜F​\(x,π\)−λ​G​\(x,π\),\\mathscr\{L\}\(x,\\pi;\\lambda\)\\triangleq F\(x,\\pi\)\-\\lambda\\,G\(x,\\pi\),so that the single\-projectλ\\lambda\-charge problem is

L​\(x;λ\)≜maxπ∈Π⁡ℒ​\(x,π;λ\)\.L\(x;\\lambda\)\\triangleq\\max\_\{\\pi\\in\\Pi\}\\,\\mathscr\{L\}\(x,\\pi;\\lambda\)\.\(17\)
Threshold policies play a central role\. For anyz∈ℝ¯≜ℝ∪\{−∞,∞\}z\\in\\overline\{\\mathbb\{R\}\}\\triangleq\\mathbb\{R\}\\cup\\\{\-\\infty,\\infty\\\}, the*zz\-threshold policy*activates whenx\>zx\>zand remains passive otherwise\. The extended thresholds have the natural limiting interpretation:z=−∞z=\-\\inftyyields the always\-active policy, whereasz=\+∞z=\+\\inftyyields the always\-passive policy \(though in the present setting finite thresholds suffice\)\. We writeF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)for the reward and work metrics under thezz\-threshold policy, and define

ℒ​\(x,z;λ\)≜F​\(x,z\)−λ​G​\(x,z\)\.\\mathscr\{L\}\(x,z;\\lambda\)\\triangleq F\(x,z\)\-\\lambda G\(x,z\)\.
Marginal metrics are defined as follows\. Fora∈\{0,1\}a\\in\\\{0,1\\\}and thresholdzz, let⟨a,z⟩\\langle a,z\\rangledenote the policy that takes actionaaat time0and thereafter follows thezz\-threshold policy\. Define the*marginal reward*and*marginal work*metrics by

f​\(x,z\)≜F​\(x,⟨1,z⟩\)−F​\(x,⟨0,z⟩\),g​\(x,z\)≜G​\(x,⟨1,z⟩\)−G​\(x,⟨0,z⟩\)\.f\(x,z\)\\triangleq F\\bigl\(x,\\langle 1,z\\rangle\\bigr\)\-F\\bigl\(x,\\langle 0,z\\rangle\\bigr\),\\qquad g\(x,z\)\\triangleq G\\bigl\(x,\\langle 1,z\\rangle\\bigr\)\-G\\bigl\(x,\\langle 0,z\\rangle\\bigr\)\.Ifg​\(x,z\)\>0g\(x,z\)\>0for allxxandzz, define the*marginal productivity*\(MP\) metric and the associated*MP index*by

m​\(x,z\)≜f​\(x,z\)g​\(x,z\),m​\(x\)≜m​\(x,x\),x∈𝒳\.m\(x,z\)\\triangleq\\frac\{f\(x,z\)\}\{g\(x,z\)\},\\qquad m\(x\)\\triangleq m\(x,x\),\\qquad x\\in\\mathscr\{X\}\.
We will verify the following real\-state*PCL\-indexability conditions*:

1. \(PCL1\)g​\(x,z\)\>0g\(x,z\)\>0for allx∈𝒳x\\in\\mathscr\{X\}andz∈ℝz\\in\\mathbb\{R\}\.
2. \(PCL2\)The MP indexm​\(⋅\)m\(\\cdot\)is nondecreasing and continuous on𝒳\\mathscr\{X\}\.
3. \(PCL3\)For eachx∈𝒳x\\in\\mathscr\{X\}and finite thresholdsz1<z2z\_\{1\}<z\_\{2\}, F​\(x,z2\)−F​\(x,z1\)=∫\(z1,z2\]m​\(z\)​G​\(x,d​z\)\.F\(x,z\_\{2\}\)\-F\(x,z\_\{1\}\)=\\int\_\{\(z\_\{1\},z\_\{2\}\]\}m\(z\)\\,G\(x,dz\)\.That is, as a function of the threshold,F​\(x,⋅\)F\(x,\\cdot\)is the indefinite*Lebesgue–Stieltjes integral*\[[8](https://arxiv.org/html/2606.11192#bib.bib8)\]ofm​\(⋅\)m\(\\cdot\)with respect to the signed Stieltjes measureG​\(x,d​z\)G\(x,dz\)induced by the right\-continuous functionz↦G​\(x,z\)z\\mapsto G\(x,z\)\. Under \(PCLI1\), the latter is nonincreasing and of bounded variation; see\[[31](https://arxiv.org/html/2606.11192#bib.bib31), Lemmas 10, 11, and 17\]\.

The single\-projectλ\\lambda\-charge problem \([17](https://arxiv.org/html/2606.11192#S3.E17)\) is called*threshold\-indexable*if the project is indexable in the statewise Whittle sense of Section[3\.1](https://arxiv.org/html/2606.11192#S3.SS1)and, for eachλ∈ℝ\\lambda\\in\\mathbb\{R\}, there exists an optimalz∗​\(λ\)z^\{\*\}\(\\lambda\)\-threshold policy\. A*generalized inverse*of a nondecreasing functionh:𝒳→ℝh:\\mathscr\{X\}\\to\\mathbb\{R\}is any functionz∗:ℝ→ℝz^\{\*\}:\\mathbb\{R\}\\to\\mathbb\{R\}such that, for allλ∈ℝ\\lambda\\in\\mathbb\{R\},

inf\{x∈𝒳:h​\(x\)≥λ\}≤z∗​\(λ\)≤sup\{x∈𝒳:h​\(x\)≤λ\},\\inf\\\{x\\in\\mathscr\{X\}:h\(x\)\\geq\\lambda\\\}\\leq z^\{\*\}\(\\lambda\)\\leq\\sup\\\{x\\in\\mathscr\{X\}:h\(x\)\\leq\\lambda\\\},with the usual conventionsinf∅=\+∞\\inf\\varnothing=\+\\inftyandsup∅=−∞\\sup\\varnothing=\-\\infty\.

With these definitions in place, the following verification theorem, specialized from\[[31](https://arxiv.org/html/2606.11192#bib.bib31)\], shows that the PCLI conditions imply both threshold\-indexability and an explicit index formula\.

###### Theorem 3\.2\(PCL verification theorem for threshold\-indexability\)\.

Under\(PCLI1–PCLI3\), the single\-project discountedλ\\lambda\-charge problem \([17](https://arxiv.org/html/2606.11192#S3.E17)\) is threshold\-indexable\. Moreover, the Whittle index exists and equals the MP indexm​\(x\)m\(x\), and any optimal threshold mapz∗​\(⋅\)z^\{\*\}\(\\cdot\)is a generalized inverse ofm​\(⋅\)m\(\\cdot\)\.

The remainder of the paper is devoted to verifying these conditions and deriving tractable expressions forFF,GG, andm​\(⋅\)m\(\\cdot\)\. Where possible, we establish the required properties analytically; elsewhere, especially in threshold regimes where the belief dynamics become more intricate, we provide systematic numerical evidence\.

## 4Computing threshold\-policy performance metrics via the no\-ACK skeleton

This section develops the computational machinery for evaluating the reward and work metricsF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)underzz\-threshold policies, together with the associated marginal metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)and the MP indexm​\(x\)=f​\(x,x\)/g​\(x,x\)m\(x\)=f\(x,x\)/g\(x,x\)\. The key idea is to decompose threshold\-driven belief evolution into two components: \(i\) deterministic*pre\-ACK*dynamics, described by the*no\-ACK skeleton*and its associated survival probabilities; and \(ii\) random ACK events that reset the belief and restart the same deterministic evolution\. This decomposition yields renewal\-type identities forFF,GG,ff, andggthat are numerically tractable and admit closed forms in several threshold regimes\. These identities also support practical verification of the PCL\-indexability conditions and guide the numerical experiments reported later\. Longer proofs are deferred to Appendix[A](https://arxiv.org/html/2606.11192#A1)\.

Throughout this section we focus on a single project and suppress its label\. We assume

0<β<1,0<p01<1,0<p10<1,ρ≜1−p10−p01=p11−p01\>0,0<κ<1,r\>0,0<\\beta<1,\\quad 0<p\_\{01\}<1,\\quad 0<p\_\{10\}<1,\\quad\\rho\\triangleq 1\-p\_\{10\}\-p\_\{01\}=p\_\{11\}\-p\_\{01\}\>0,\\quad 0<\\kappa<1,\\quad r\>0,and recall that𝒳≜\[0,1\]\\mathscr\{X\}\\triangleq\[0,1\]\. Throughout, “increasing” and “decreasing” are interpreted in the strict sense\.

### 4\.1Belief\-update maps and their properties

We begin by collecting the structural properties of the one\-step belief\-update maps that underpin the subsequent computations under threshold policies\. We introduce the*passive update*ϕ0\\phi^\{0\}and the*active no\-ACK update*ϕ1\\phi^\{1\}, identify a forward\-invariant belief interval, and characterize the fixed points and closed\-form iterates of both maps\. These properties underpin the no\-ACK skeleton and the itinerary/word machinery used later, and also prepare the contractiveness argument obtained after an increasing change of variables to log\-odds\.

Given beliefxxand activation, the posterior belief after observing a NACK is

ψ​\(x\)≜ℙ​\{S​\(t\)=1∣K​\(t\)=0,A​\(t\)=1,X​\(t\)=x\}=\(1−κ\)​x1−κ​x,x∈𝒳\.\\psi\(x\)\\triangleq\\mathbb\{P\}\\\{S\(t\)=1\\mid K\(t\)=0,\\,A\(t\)=1,\\,X\(t\)=x\\\}=\\frac\{\(1\-\\kappa\)x\}\{1\-\\kappa x\},\\qquad x\\in\\mathscr\{X\}\.\(18\)Letϕ0\\phi^\{0\}andϕ1\\phi^\{1\}denote the passive and active no\-ACK belief\-update maps, respectively:

ϕ0​\(x\)≜p01\+ρ​x,ϕ1​\(x\)≜p01\+ρ​ψ​\(x\)=ϕ0​\(x\)−ρ​κ​x​\(1−x\)1−κ​x,x∈𝒳\.\\phi^\{0\}\(x\)\\triangleq p\_\{01\}\+\\rho x,\\qquad\\phi^\{1\}\(x\)\\triangleq p\_\{01\}\+\\rho\\,\\psi\(x\)=\\phi^\{0\}\(x\)\-\\rho\\,\\frac\{\\kappa x\(1\-x\)\}\{1\-\\kappa x\},\\qquad x\\in\\mathscr\{X\}\.\(19\)
A key relation between the belief\-update mapsϕ0\\phi^\{0\}andϕ1\\phi^\{1\}is the following*one\-step posterior\-mean identity*:

𝔼​\[X​\(t\+1\)∣X​\(t\)=x,A​\(t\)=0\]=ϕ0​\(x\)=κ​x​p11\+\(1−κ​x\)​ϕ1​\(x\)=𝔼​\[X​\(t\+1\)∣X​\(t\)=x,A​\(t\)=1\]\.\\mathbb\{E\}\\\!\\left\[X\(t\+1\)\\mid X\(t\)=x,\\ A\(t\)=0\\right\]=\\phi^\{0\}\(x\)=\\kappa x\\,p\_\{11\}\+\(1\-\\kappa x\)\\,\\phi^\{1\}\(x\)=\\mathbb\{E\}\\\!\\left\[X\(t\+1\)\\mid X\(t\)=x,\\ A\(t\)=1\\right\]\.\(20\)Thus, although the action changes the*distribution*of the next beliefX​\(t\+1\)X\(t\+1\), it leaves its*conditional mean*unchanged: under either action, the one\-step mean drift equalsϕ0​\(x\)\\phi^\{0\}\(x\)\. This one\-step mean identity propagates over time and yields the following lemma\. Letϕt0\\phi\_\{t\}^\{0\}denote thett\-fold iterate ofϕ0\\phi^\{0\}, withϕ00​\(x\)≡x\\phi\_\{0\}^\{0\}\(x\)\\equiv x\.

###### Lemma 4\.2\(Policy\-invariant mean belief path\)\.

Fix an initial beliefX​\(0\)=xX\(0\)=xand letπ∈Π\\pi\\in\\Pibe any admissible policy\. Then the mean belief path is*policy\-invariant*: for everyt≥0t\\geq 0,

𝔼xπ​\[X​\(t\)\]=ϕt0​\(x\)\.\\mathbb\{E\}\_\{x\}^\{\\pi\}\[X\(t\)\]=\\phi\_\{t\}^\{0\}\(x\)\.\(21\)

The proof is deferred to Appendix[A\.1](https://arxiv.org/html/2606.11192#A1.SS1)\.

Our analysis of no\-ACK skeleton action itineraries draws on the regularity assumptions used in\[[11](https://arxiv.org/html/2606.11192#bib.bib11), Assumption 2\]\. To state them independently of our model, letha:𝒳→𝒳h^\{a\}:\\mathscr\{X\}\\to\\mathscr\{X\}denote a generic belief\-update map under actiona∈\{0,1\}a\\in\\\{0,1\\\}\.

###### Assumption 4\.3\(Regularity of belief\-update maps\[[11](https://arxiv.org/html/2606.11192#bib.bib11), Assumption 2\]\)\.

1. \(i\)\(Monotonicity\)h0h^\{0\}andh1h^\{1\}are increasing on𝒳\\mathscr\{X\}\.
2. \(ii\)\(Unique fixed points and ordering\)Eachhah^\{a\}admits a unique fixed pointxax^\{a\}in𝒳\\mathscr\{X\}, andx1<x0x^\{1\}<x^\{0\}\.
3. \(iii\)\(Contractiveness up to increasing conjugacy\)There exists an increasing bijectionϑ:\(0,1\)→ℝ\\vartheta:\(0,1\)\\to\\mathbb\{R\}such that the conjugated maps h^a≜ϑ∘ha∘ϑ−1:ℝ→ℝ,a∈\{0,1\},\\hat\{h\}^\{a\}\\triangleq\\vartheta\\circ h^\{a\}\\circ\\vartheta^\{\-1\}:\\mathbb\{R\}\\to\\mathbb\{R\},\\qquad a\\in\\\{0,1\\\},are*contractive*in the strict sense that \|h^a​\(u\)−h^a​\(v\)\|<\|u−v\|for all distinct​u,v∈ℝ,a∈\{0,1\}\.\|\\hat\{h\}^\{a\}\(u\)\-\\hat\{h\}^\{a\}\(v\)\|<\|u\-v\|\\qquad\\text\{for all distinct \}u,v\\in\\mathbb\{R\},\\ \\ a\\in\\\{0,1\\\}\.

In our model, we verify \(i\)–\(ii\) directly for\(h0,h1\)=\(ϕ0,ϕ1\)\(h^\{0\},h^\{1\}\)=\(\\phi^\{0\},\\phi^\{1\}\), and we verify \(iii\) withϑ\\varthetataken to be the logit map, as established in Lemma[4\.7](https://arxiv.org/html/2606.11192#S4.Thmtheorem7)\. The fact thatϑ\\varthetais defined on\(0,1\)\(0,1\)is immaterial here, since the forward\-invariant interval𝒳~=\[p01,p11\]⊂\(0,1\)\\widetilde\{\\mathscr\{X\}\}=\[p\_\{01\},p\_\{11\}\]\\subset\(0,1\)contains all no\-ACK skeleton iterates from time11onward; see Remark[4\.1](https://arxiv.org/html/2606.11192#S4.Thmtheorem1)\(iii\)\.

We next verify the relevant parts of Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)for our model\. We show thatϕ0\\phi^\{0\}andϕ1\\phi^\{1\}are increasing on𝒳\\mathscr\{X\}, that each has a unique fixed point in𝒳\\mathscr\{X\}withx1<x0x^\{1\}<x^\{0\}, and we derive closed\-form expressions for their iterates, thereby establishing Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(i\)–\(ii\)\. The mapϕ0\\phi^\{0\}is immediately seen to be contractive on𝒳\\mathscr\{X\}\. By contrast,ϕ1\\phi^\{1\}need not be contractive in the belief variable, but it becomes a strict contraction after the standard increasing change of variables to log\-odds\. This odds\-contractiveness is sufficient for the itinerary machinery adapted fromDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]\.

Fora∈\{0,1\}a\\in\\\{0,1\\\}, letϕta​\(x\)\\phi\_\{t\}^\{a\}\(x\)denote thett\-fold iterate ofϕa\\phi^\{a\}, withϕ0a​\(x\)≜x\\phi\_\{0\}^\{a\}\(x\)\\triangleq x\. We refer to\(ϕt0​\(x\)\)t≥0\\bigl\(\\phi\_\{t\}^\{0\}\(x\)\\bigr\)\_\{t\\geq 0\}as the*passive iterates*and to\(ϕt1​\(x\)\)t≥0\\bigl\(\\phi\_\{t\}^\{1\}\(x\)\\bigr\)\_\{t\\geq 0\}as the*active no\-ACK iterates*\.

We first consider passive iterates and threshold up\-crossing times\. The mapϕ0​\(x\)=p01\+ρ​x\\phi^\{0\}\(x\)=p\_\{01\}\+\\rho xis contractive and has the unique fixed point

x0=p011−ρ∈\(p01,p11\)\.x^\{0\}=\\frac\{p\_\{01\}\}\{1\-\\rho\}\\in\(p\_\{01\},p\_\{11\}\)\.Its iterates admit the closed form

ϕt0​\(x\)=x0\+\(x−x0\)​ρt,t≥0\.\\phi\_\{t\}^\{0\}\(x\)=x^\{0\}\+\(x\-x^\{0\}\)\\rho^\{t\},\\qquad t\\geq 0\.\(22\)In particular,ϕ0\\phi^\{0\}is increasing on𝒳\\mathscr\{X\}, verifying Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(i\) forh0=ϕ0h^\{0\}=\\phi^\{0\}\. Ifx<x0x<x^\{0\}, then\(ϕt0​\(x\)\)t≥0\\bigl\(\\phi\_\{t\}^\{0\}\(x\)\\bigr\)\_\{t\\geq 0\}increases tox0x^\{0\}, whereas ifx\>x0x\>x^\{0\}, it decreases tox0x^\{0\}\.

For a thresholdz<x0z<x^\{0\}, define the*passive threshold up\-crossing time*

τ↑0​\(x,z\)≜min⁡\{t≥0:ϕt0​\(x\)\>z\}=\{1\+⌊ln⁡\(\(x0−z\)/\(x0−x\)\)ln⁡ρ⌋,x≤z,0,x\>z,\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\\triangleq\\min\\\{t\\geq 0:\\phi\_\{t\}^\{0\}\(x\)\>z\\\}=\\begin\{cases\}1\+\\left\\lfloor\\dfrac\{\\ln\\\!\\bigl\(\(x^\{0\}\-z\)/\(x^\{0\}\-x\)\\bigr\)\}\{\\ln\\rho\}\\right\\rfloor,&x\\leq z,\\\\\[4\.30554pt\] 0,&x\>z,\\end\{cases\}\(23\)where we have used thatϕt0​\(x\)\>z\\phi\_\{t\}^\{0\}\(x\)\>zis equivalent toρt<\(x0−z\)/\(x0−x\)\\rho^\{t\}<\(x^\{0\}\-z\)/\(x^\{0\}\-x\)whenx≤z<x0x\\leq z<x^\{0\}\.

We next turn to the active no\-ACK iterates\. The fixed points ofϕ1\\phi^\{1\}satisfy the quadratic

κ​x2−\(1−ρ\+κ​p11\)​x\+p01=0\.\\kappa x^\{2\}\-\\bigl\(1\-\\rho\+\\kappa p\_\{11\}\\bigr\)x\+p\_\{01\}=0\.\(24\)
###### Lemma 4\.4\.

Equation \([24](https://arxiv.org/html/2606.11192#S4.E24)\) has two distinct real rootsxlo1<xhi1x\_\{\\mathrm\{lo\}\}^\{1\}<x\_\{\\mathrm\{hi\}\}^\{1\}, namely

xlo1≜1−ρ\+κ​p11−Δ​\(κ\)2​κ,xhi1≜1−ρ\+κ​p11\+Δ​\(κ\)2​κ,x\_\{\\mathrm\{lo\}\}^\{1\}\\triangleq\\frac\{1\-\\rho\+\\kappa p\_\{11\}\-\\sqrt\{\\Delta\(\\kappa\)\}\}\{2\\kappa\},\\qquad x\_\{\\mathrm\{hi\}\}^\{1\}\\triangleq\\frac\{1\-\\rho\+\\kappa p\_\{11\}\+\\sqrt\{\\Delta\(\\kappa\)\}\}\{2\\kappa\},\(25\)where

Δ​\(κ\)≜\(1−ρ\+κ​p11\)2−4​κ​p01=\(\(p10\+p01\)−κ​p11\)2\+4​κ​p10​ρ\>0\.\\Delta\(\\kappa\)\\triangleq\\bigl\(1\-\\rho\+\\kappa p\_\{11\}\\bigr\)^\{2\}\-4\\kappa p\_\{01\}=\\bigl\(\(p\_\{10\}\+p\_\{01\}\)\-\\kappa p\_\{11\}\\bigr\)^\{2\}\+4\\kappa p\_\{10\}\\rho\>0\.\(26\)Moreover,xlo1∈\(p01,p11\)x\_\{\\mathrm\{lo\}\}^\{1\}\\in\(p\_\{01\},p\_\{11\}\)andxhi1∈\(1,1/κ\)x\_\{\\mathrm\{hi\}\}^\{1\}\\in\(1,1/\\kappa\)\. Thus,ϕ1\\phi^\{1\}has a unique fixed point in𝒳\\mathscr\{X\}, namelyx1≜xlo1x^\{1\}\\triangleq x\_\{\\mathrm\{lo\}\}^\{1\}\.

The proof is deferred to Appendix[A\.1](https://arxiv.org/html/2606.11192#A1.SS1)\.

A useful identity involving these fixed points follows from Vieta’s formulas:

xlo1\+xhi1=1−ρ\+κ​p11κ,xlo1​xhi1=p01κ\.x\_\{\\mathrm\{lo\}\}^\{1\}\+x\_\{\\mathrm\{hi\}\}^\{1\}=\\frac\{1\-\\rho\+\\kappa p\_\{11\}\}\{\\kappa\},\\qquad x\_\{\\mathrm\{lo\}\}^\{1\}x\_\{\\mathrm\{hi\}\}^\{1\}=\\frac\{p\_\{01\}\}\{\\kappa\}\.Therefore,

\(1−κ​xlo1\)​\(1−κ​xhi1\)=ρ​\(1−κ\),\(1\-\\kappa x\_\{\\mathrm\{lo\}\}^\{1\}\)\(1\-\\kappa x\_\{\\mathrm\{hi\}\}^\{1\}\)=\\rho\(1\-\\kappa\),\(27\)and, lettingx1≜xlo1x^\{1\}\\triangleq x\_\{\\mathrm\{lo\}\}^\{1\},

\(ϕ1\)′​\(x1\)=ρ​\(1−κ\)\(1−κ​x1\)2=1−κ​xhi11−κ​x1\.\(\\phi^\{1\}\)^\{\\prime\}\(x^\{1\}\)=\\frac\{\\rho\(1\-\\kappa\)\}\{\(1\-\\kappa x^\{1\}\)^\{2\}\}=\\frac\{1\-\\kappa x\_\{\\mathrm\{hi\}\}^\{1\}\}\{1\-\\kappa x^\{1\}\}\.\(28\)
###### Lemma 4\.5\(Ordering of the active and passive fixed points\)\.

We havex1<x0x^\{1\}<x^\{0\}\.

The proof is deferred to Appendix[A\.1](https://arxiv.org/html/2606.11192#A1.SS1)\.

Lemmas[4\.4](https://arxiv.org/html/2606.11192#S4.Thmtheorem4)and[4\.5](https://arxiv.org/html/2606.11192#S4.Thmtheorem5)verify Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(ii\) forh0=ϕ0h^\{0\}=\\phi^\{0\}andh1=ϕ1h^\{1\}=\\phi^\{1\}\.

We next record the Möbius structure of the active no\-ACK update\. The mapϕ1\\phi^\{1\}is a nondegenerate Möbius \(fractional\-linear\) map, i\.e\.,ϕ1​\(x\)=\(a​x\+b\)/\(c​x\+d\)\\phi^\{1\}\(x\)=\(ax\+b\)/\(cx\+d\)witha​d−b​c≠0ad\-bc\\neq 0; see, e\.g\.,\[[3](https://arxiv.org/html/2606.11192#bib.bib3);[6](https://arxiv.org/html/2606.11192#bib.bib6)\]\. In particular,

ϕ1​\(x\)=p01\+ρ​\(1−κ\)​x1−κ​x=\(ρ​\(1−κ\)−κ​p01\)​x\+p011−κ​x,x∈𝒳\.\\phi^\{1\}\(x\)=p\_\{01\}\+\\rho\\,\\frac\{\(1\-\\kappa\)x\}\{1\-\\kappa x\}=\\frac\{\\bigl\(\\rho\(1\-\\kappa\)\-\\kappa p\_\{01\}\\bigr\)x\+p\_\{01\}\}\{1\-\\kappa x\},\\qquad x\\in\\mathscr\{X\}\.\(29\)We use the natural extension ofϕ1\\phi^\{1\}to\[0,1/κ\)\[0,1/\\kappa\)to obtain closed forms for the iteratesϕt1\\phi\_\{t\}^\{1\}\.

The following standard conjugacy identity for Möbius maps will be useful\. LetΦ​\(x\)≜\(a​x\+b\)/\(c​x\+d\)\\Phi\(x\)\\triangleq\(ax\+b\)/\(cx\+d\)be a nondegenerate Möbius map with two distinct fixed pointsx−<x\+x\_\{\-\}<x\_\{\+\}, and define

Ψ​\(x\)≜x−x−x−x\+,x≠x\+\.\\Psi\(x\)\\triangleq\\frac\{x\-x\_\{\-\}\}\{x\-x\_\{\+\}\},\\qquad x\\neq x\_\{\+\}\.\(30\)Then, withμ≜Φ′​\(x−\)≠0\\mu\\triangleq\\Phi^\{\\prime\}\(x\_\{\-\}\)\\neq 0,

Ψ​\(Φ​\(x\)\)=μ​Ψ​\(x\)\(x≠x\+\)\.\\Psi\(\\Phi\(x\)\)=\\mu\\,\\Psi\(x\)\\qquad\(x\\neq x\_\{\+\}\)\.\(31\)DefiningΦ0​\(x\)≜x\\Phi\_\{0\}\(x\)\\triangleq xandΦt≜Φ∘Φt−1\\Phi\_\{t\}\\triangleq\\Phi\\circ\\Phi\_\{t\-1\}fort≥1t\\geq 1, we obtain

Ψ​\(Φt​\(x\)\)=μt​Ψ​\(x\)\(t≥0,x≠x\+\)\.\\Psi\(\\Phi\_\{t\}\(x\)\)=\\mu^\{t\}\\Psi\(x\)\\qquad\(t\\geq 0,\\ x\\neq x\_\{\+\}\)\.Moreover,Ψ′​\(x\)=\(x−−x\+\)/\(x−x\+\)2<0\\Psi^\{\\prime\}\(x\)=\(x\_\{\-\}\-x\_\{\+\}\)/\(x\-x\_\{\+\}\)^\{2\}<0, soΨ\\Psiis decreasing onℝ∖\{x\+\}\\mathbb\{R\}\\setminus\\\{x\_\{\+\}\\\}and hence invertible there, with

Ψ−1​\(y\)=y​x\+−x−y−1\(y≠1\)\.\\Psi^\{\-1\}\(y\)=\\frac\{y\\,x\_\{\+\}\-x\_\{\-\}\}\{y\-1\}\\qquad\(y\\neq 1\)\.Therefore,

Φt​\(x\)=Ψ−1​\(μt​Ψ​\(x\)\)=μt​Ψ​\(x\)​x\+−x−μt​Ψ​\(x\)−1,t≥0,x≠x\+\.\\Phi\_\{t\}\(x\)=\\Psi^\{\-1\}\\\!\\bigl\(\\mu^\{t\}\\Psi\(x\)\\bigr\)=\\frac\{\\mu^\{t\}\\Psi\(x\)\\,x\_\{\+\}\-x\_\{\-\}\}\{\\mu^\{t\}\\Psi\(x\)\-1\},\\qquad t\\geq 0,\\ x\\neq x\_\{\+\}\.\(32\)
Specializing now toΦ=ϕ1\\Phi=\\phi^\{1\}, the fixed points arex−=x1x\_\{\-\}=x^\{1\}andx\+=xhi1x\_\{\+\}=x\_\{\\mathrm\{hi\}\}^\{1\}by Lemma[4\.4](https://arxiv.org/html/2606.11192#S4.Thmtheorem4)\. Thus,

Ψ​\(x\)≜x−x1x−xhi1,x≠xhi1\.\\Psi\(x\)\\triangleq\\frac\{x\-x^\{1\}\}\{x\-x\_\{\\mathrm\{hi\}\}^\{1\}\},\\qquad x\\neq x\_\{\\mathrm\{hi\}\}^\{1\}\.\(33\)Sincexhi1\>1x\_\{\\mathrm\{hi\}\}^\{1\}\>1,Ψ\\Psiis well defined on𝒳\\mathscr\{X\}\.

###### Lemma 4\.6\(Closed form and monotone behavior ofϕt1​\(x\)\\phi\_\{t\}^\{1\}\(x\)\)\.

The following holds:

1. \(a\)For everyt≥1t\\geq 1, the mapx↦ϕt1​\(x\)x\\mapsto\\phi\_\{t\}^\{1\}\(x\)is increasing\.
2. \(b\)For allx∈𝒳x\\in\\mathscr\{X\}, ϕ1​\(x\)−x\\displaystyle\\phi^\{1\}\(x\)\-x=−κ​\(x−x1\)​\(xhi1−x\)1−κ​x,\\displaystyle=\-\\frac\{\\kappa\\,\(x\-x^\{1\}\)\(x\_\{\\mathrm\{hi\}\}^\{1\}\-x\)\}\{1\-\\kappa x\},\(34\)ϕ1​\(x\)−x1\\displaystyle\\phi^\{1\}\(x\)\-x^\{1\}=ρ​\(1−κ\)1−κ​x1​x−x11−κ​x\.\\displaystyle=\\frac\{\\rho\(1\-\\kappa\)\}\{1\-\\kappa x^\{1\}\}\\,\\frac\{x\-x^\{1\}\}\{1\-\\kappa x\}\.\(35\)Consequently, x<x1⟺x<ϕ1\(x\)<x1,x\>x1⟺x1<ϕ1\(x\)<x,x<x^\{1\}\\ \\Longleftrightarrow\\ x<\\phi^\{1\}\(x\)<x^\{1\},\\qquad x\>x^\{1\}\\ \\Longleftrightarrow\\ x^\{1\}<\\phi^\{1\}\(x\)<x,and\(ϕt1​\(x\)\)t≥0\\bigl\(\\phi\_\{t\}^\{1\}\(x\)\\bigr\)\_\{t\\geq 0\}is strictly monotone inttand converges tox1x^\{1\}forx≠x1x\\neq x^\{1\}\.
3. \(c\)Let μ≜\(ϕ1\)′​\(x1\)=ρ​\(1−κ\)\(1−κ​x1\)2=1−κ​xhi11−κ​x1∈\(0,1\)\.\\mu\\triangleq\(\\phi^\{1\}\)^\{\\prime\}\(x^\{1\}\)=\\frac\{\\rho\(1\-\\kappa\)\}\{\(1\-\\kappa x^\{1\}\)^\{2\}\}=\\frac\{1\-\\kappa x\_\{\\mathrm\{hi\}\}^\{1\}\}\{1\-\\kappa x^\{1\}\}\\in\(0,1\)\.\(36\)Then for allt≥0t\\geq 0andx∈𝒳x\\in\\mathscr\{X\}, ϕt1​\(x\)=x1​\(x−xhi1\)−xhi1​μt​\(x−x1\)\(x−xhi1\)−μt​\(x−x1\)=x1\+μt​\(xhi1−x1\)​\(x−x1\)\(xhi1−x1\)\+\(μt−1\)​\(x−x1\),\\phi\_\{t\}^\{1\}\(x\)=\\frac\{x^\{1\}\(x\-x\_\{\\mathrm\{hi\}\}^\{1\}\)\-x\_\{\\mathrm\{hi\}\}^\{1\}\\,\\mu^\{t\}\(x\-x^\{1\}\)\}\{\(x\-x\_\{\\mathrm\{hi\}\}^\{1\}\)\-\\mu^\{t\}\(x\-x^\{1\}\)\}=x^\{1\}\+\\frac\{\\mu^\{t\}\(x\_\{\\mathrm\{hi\}\}^\{1\}\-x^\{1\}\)\\,\(x\-x^\{1\}\)\}\{\(x\_\{\\mathrm\{hi\}\}^\{1\}\-x^\{1\}\)\+\(\\mu^\{t\}\-1\)\\,\(x\-x^\{1\}\)\},\(37\)andϕt1​\(x\)→x1\\phi\_\{t\}^\{1\}\(x\)\\to x^\{1\}ast→∞t\\to\\inftywith geometric rateμ\\mu\.

The proof is deferred to Appendix[A\.1](https://arxiv.org/html/2606.11192#A1.SS1)\.

Finally, we address contractiveness\. The passive mapϕ0\\phi^\{0\}is contractive on𝒳\\mathscr\{X\}:

\|ϕ0​\(x\)−ϕ0​\(y\)\|=ρ​\|x−y\|<\|x−y\|\(x≠y\),\|\\phi^\{0\}\(x\)\-\\phi^\{0\}\(y\)\|=\\rho\\,\|x\-y\|<\|x\-y\|\\qquad\(x\\neq y\),verifying Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(iii\) forh0=ϕ0h^\{0\}=\\phi^\{0\}\. The active no\-ACK mapϕ1\\phi^\{1\}need not be contractive on𝒳\\mathscr\{X\}in the belief variable whenκ\\kappais close to11\(indeed,supx∈𝒳\(ϕ1\)′​\(x\)\\sup\_\{x\\in\\mathscr\{X\}\}\(\\phi^\{1\}\)^\{\\prime\}\(x\)may exceed11\)\. We therefore use an increasing change of variables to log\-odds to ensure contractiveness\.

#### 4\.1\.1Conjugacy to increasing contractions

The next lemma shows that our belief\-update maps satisfy the contractiveness part of Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(iii\) after an increasing change of variables, in the sense used inDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]\.

###### Lemma 4\.7\(Logit conjugacy yields increasing contractions\)\.

Letϑ:\(0,1\)→ℝ\\vartheta:\(0,1\)\\to\\mathbb\{R\}be the logit map

ϑ​\(x\)≜log⁡x1−x,ϑ−1​\(t\)=et1\+et\.\\vartheta\(x\)\\triangleq\\log\\frac\{x\}\{1\-x\},\\qquad\\vartheta^\{\-1\}\(t\)=\\frac\{e^\{t\}\}\{1\+e^\{t\}\}\.Fora∈\{0,1\}a\\in\\\{0,1\\\}, define the conjugated mapsϕ^a≜ϑ∘ϕa∘ϑ−1\\hat\{\\phi\}^\{a\}\\triangleq\\vartheta\\circ\\phi^\{a\}\\circ\\vartheta^\{\-1\}onℝ\\mathbb\{R\}\. Then the following hold:

1. \(a\)ϕ^0\\hat\{\\phi\}^\{0\}andϕ^1\\hat\{\\phi\}^\{1\}are increasing and contractive onℝ\\mathbb\{R\};
2. \(b\)eachϕ^a\\hat\{\\phi\}^\{a\}has the unique fixed pointx^a=ϑ​\(xa\)\\hat\{x\}^\{a\}=\\vartheta\(x^\{a\}\);
3. \(c\)x^1<x^0\\hat\{x\}^\{1\}<\\hat\{x\}^\{0\}\.

The proof is deferred to Appendix[A\.1](https://arxiv.org/html/2606.11192#A1.SS1)\.

### 4\.2No\-ACK skeleton dynamics and survival probabilities

A central ingredient in our PCL\-indexability analysis is an explicit description of the belief trajectories induced by threshold policies, since the performance metricsFF,GG,ff, andggcan be expressed in terms of deterministic pre\-ACK evolution together with ACK\-triggered restarts\. Fix a thresholdz∈ℝz\\in\\mathbb\{R\}and an initial beliefx∈𝒳x\\in\\mathscr\{X\}\. Under thezz\-threshold policy, the action at timettisA​\(t\)=𝟙\{X​\(t\)\>z\}A\(t\)=\\mathbbm\{1\}\_\{\\\{X\(t\)\>z\\\}\}\. The resulting belief and action sequences,

orbit​\(x,z\)≜\(X​\(t\)\)t≥0,σ​\(x,z\)≜\(A​\(t\)\)t≥0,\\mathrm\{orbit\}\(x,z\)\\triangleq\(X\(t\)\)\_\{t\\geq 0\},\\qquad\\sigma\(x,z\)\\triangleq\(A\(t\)\)\_\{t\\geq 0\},are called, respectively, the*orbit*and*itinerary*\. Their evolution combines deterministic updates viaϕ0\\phi^\{0\}andϕ1\\phi^\{1\}between ACK times with random resets top11p\_\{11\}upon an ACK\. We next isolate the associated deterministic no\-ACK skeleton, generated by a*map\-with\-a\-gap*, together with the corresponding no\-ACK survival probabilitiesΓt​\(x,z\)\\Gamma\_\{t\}\(x,z\)\.

#### 4\.2\.1The no\-ACK skeleton as an iterated map\-with\-a\-gap

To separate deterministic evolution from random ACK resets, define the*no\-ACK skeleton*mapφ​\(⋅,z\):𝒳→𝒳\\varphi\(\\cdot,z\):\\mathscr\{X\}\\to\\mathscr\{X\}under thezz\-threshold policy by

φ​\(x,z\)≜\{ϕ1​\(x\),x\>z,ϕ0​\(x\),x≤z,\\varphi\(x,z\)\\triangleq\\begin\{cases\}\\phi^\{1\}\(x\),&x\>z,\\\\\[2\.84526pt\] \\phi^\{0\}\(x\),&x\\leq z,\\end\{cases\}which is generally discontinuous atzz; see Remark[4\.1](https://arxiv.org/html/2606.11192#S4.Thmtheorem1)\(i\)\. For each fixedzz, define its iterates by

φ0​\(x,z\)≜x,φt​\(x,z\)≜φ​\(φt−1​\(x,z\),z\),t≥1\.\\varphi\_\{0\}\(x,z\)\\triangleq x,\\qquad\\varphi\_\{t\}\(x,z\)\\triangleq\\varphi\\\!\\bigl\(\\varphi\_\{t\-1\}\(x,z\),z\\bigr\),\\qquad t\\geq 1\.\(38\)The associated deterministic skeleton orbit and itinerary are

orbit~​\(x,z\)≜\(X~t​\(x,z\)\)t≥0,X~t​\(x,z\)≜φt​\(x,z\),σ~​\(x,z\)≜\(A~t​\(x,z\)\)t≥0,A~t​\(x,z\)≜𝟙\{X~t​\(x,z\)\>z\}\.\\widetilde\{\\mathrm\{orbit\}\}\(x,z\)\\triangleq\(\\widetilde\{X\}\_\{t\}\(x,z\)\)\_\{t\\geq 0\},\\ \\ \\widetilde\{X\}\_\{t\}\(x,z\)\\triangleq\\varphi\_\{t\}\(x,z\),\\qquad\\widetilde\{\\sigma\}\(x,z\)\\triangleq\(\\widetilde\{A\}\_\{t\}\(x,z\)\)\_\{t\\geq 0\},\\ \\ \\widetilde\{A\}\_\{t\}\(x,z\)\\triangleq\\mathbbm\{1\}\_\{\\\{\\widetilde\{X\}\_\{t\}\(x,z\)\>z\\\}\}\.\(39\)Thus\(X~t,A~t\)\(\\widetilde\{X\}\_\{t\},\\widetilde\{A\}\_\{t\}\)is the trajectory obtained by iterating the map\-with\-a\-gap under the convention that no ACK is ever observed\. The actual stochastic trajectory agrees with this skeleton up to the first ACK time; when an ACK occurs, the belief resets top11p\_\{11\}and the same deterministic rule restarts from that new initial state\.

We writeφt​\(x,z\)\\varphi\_\{t\}\(x,z\)when emphasizing the iterated map, andX~t​\(x,z\)\\widetilde\{X\}\_\{t\}\(x,z\)andA~t​\(x,z\)\\widetilde\{A\}\_\{t\}\(x,z\)when emphasizing the resulting deterministic skeleton state trajectory and itinerary\.

#### 4\.2\.2Pre\-ACK representation and survival probabilities

Fix\(x,z\)\(x,z\)and let

τack≜inf\{t≥0:an ACK is observed at time​t\},\\tau^\{\\mathrm\{ack\}\}\\triangleq\\inf\\\{t\\geq 0:\\ \\text\{an ACK is observed at time \}t\\\},with the conventionτack=∞\\tau^\{\\mathrm\{ack\}\}=\\inftyif no ACK is ever observed\. On the event\{τack≥t\}\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}, that is, when no ACK occurs in periods0,…,t−10,\\ldots,t\-1, the stochastic trajectory agrees with the skeleton up to and including timett:

\{τack≥t\}⊆\{\(X​\(j\),A​\(j\)\)=\(X~j​\(x,z\),A~j​\(x,z\)\),j=0,…,t\}\.\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}\\subseteq\\bigl\\\{\(X\(j\),A\(j\)\)=\(\\widetilde\{X\}\_\{j\}\(x,z\),\\widetilde\{A\}\_\{j\}\(x,z\)\),\\ j=0,\\ldots,t\\bigr\\\}\.\(40\)Define the*no\-ACK survival probabilities*by

Γt​\(x,z\)≜ℙxz​\{τack≥t\},t≥0\.\\Gamma\_\{t\}\(x,z\)\\triangleq\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\},\\qquad t\\geq 0\.\(41\)Since, conditional on\{τack≥t−1\}\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\-1\\\}, an ACK occurs at timet−1t\-1with probability

κ​X~t−1​\(x,z\)​A~t−1​\(x,z\),\\kappa\\,\\widetilde\{X\}\_\{t\-1\}\(x,z\)\\,\\widetilde\{A\}\_\{t\-1\}\(x,z\),we obtain the recursion

Γ0​\(x,z\)=1,Γt​\(x,z\)=Γt−1​\(x,z\)​\(1−κ​X~t−1​\(x,z\)​A~t−1​\(x,z\)\),t≥1,\\Gamma\_\{0\}\(x,z\)=1,\\qquad\\Gamma\_\{t\}\(x,z\)=\\Gamma\_\{t\-1\}\(x,z\)\\Bigl\(1\-\\kappa\\,\\widetilde\{X\}\_\{t\-1\}\(x,z\)\\,\\widetilde\{A\}\_\{t\-1\}\(x,z\)\\Bigr\),\\qquad t\\geq 1,and hence the product representation

Γt​\(x,z\)=∏j=0t−1\(1−κ​X~j​\(x,z\)​A~j​\(x,z\)\)=∏0≤j<t:X~j​\(x,z\)\>z\(1−κ​X~j​\(x,z\)\),t≥1\.\\Gamma\_\{t\}\(x,z\)=\\prod\_\{j=0\}^\{t\-1\}\\Bigl\(1\-\\kappa\\,\\widetilde\{X\}\_\{j\}\(x,z\)\\,\\widetilde\{A\}\_\{j\}\(x,z\)\\Bigr\)=\\prod\_\{0\\leq j<t:\\ \\widetilde\{X\}\_\{j\}\(x,z\)\>z\}\\bigl\(1\-\\kappa\\,\\widetilde\{X\}\_\{j\}\(x,z\)\\bigr\),\\qquad t\\geq 1\.\(42\)Finally, letΓ∞​\(x,z\)≜limt→∞Γt​\(x,z\)=ℙxz​\{τack=∞\},\\Gamma\_\{\\infty\}\(x,z\)\\triangleq\\lim\_\{t\\to\\infty\}\\Gamma\_\{t\}\(x,z\)=\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=\\infty\\\},where the limit exists because\(Γt​\(x,z\)\)t≥0\(\\Gamma\_\{t\}\(x,z\)\)\_\{t\\geq 0\}is nonincreasing and bounded below by0\. These deterministic quantities will be used below to derive renewal formulas forF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)in the intermediate threshold regimes\.

Thezz\-threshold policy activates at timettif and only ifX​\(t\)\>zX\(t\)\>z, and its no\-ACK skeleton is generated byφ​\(⋅,z\)\\varphi\(\\cdot,z\)above\. It is also convenient to consider the*z−z^\{\-\}\-threshold policy*, which activates if and only ifX​\(t\)≥zX\(t\)\\geq z\. Its no\-ACK skeleton is generated by the mapφ​\(⋅,z−\)\\varphi\(\\cdot,z^\{\-\}\), defined by

φ​\(x,z−\)≜\{ϕ1​\(x\),x≥z,ϕ0​\(x\),x<z,x∈𝒳,\\varphi\(x,z^\{\-\}\)\\triangleq\\begin\{cases\}\\phi^\{1\}\(x\),&x\\geq z,\\\\\[2\.84526pt\] \\phi^\{0\}\(x\),&x<z,\\end\{cases\}\\qquad x\\in\\mathscr\{X\},and iterated as in \([38](https://arxiv.org/html/2606.11192#S4.E38)\)\. The notationφ​\(x,z−\)\\varphi\(x,z^\{\-\}\)reflects the fact that, for each fixedx∈𝒳x\\in\\mathscr\{X\},φ​\(x,z−\)=limζ↗zφ​\(x,ζ\),\\varphi\(x,z^\{\-\}\)=\\lim\_\{\\zeta\\nearrow z\}\\varphi\(x,\\zeta\),that is, it is the left limit of the mapζ↦φ​\(x,ζ\)\\zeta\\mapsto\\varphi\(x,\\zeta\)at thresholdzz\. We writeσ~​\(x,z−\)\\widetilde\{\\sigma\}\(x,z^\{\-\}\)for the resulting skeleton itinerary\.

### 4\.3Renewal decompositions forF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)

We next derive renewal\-type identities expressing the threshold\-policy reward and work metricsF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)in terms of deterministic*pre\-ACK*quantities computed along the no\-ACK skeleton and the restart statep11p\_\{11\}induced by an ACK\. The key observation is that, under a threshold policy, the stochastic trajectory coincides with the skeleton up to the first ACK time, at which point the belief resets top11p\_\{11\}and the same evolution restarts from that state\. This reduces the computation ofFFandGGto two ingredients: discounted sums along the no\-ACK skeleton, weighted by no\-ACK survival probabilities, and a continuation term evaluated at the restart state\. The resulting pre\-ACK metrics admit convergent series representations that can be evaluated numerically by truncation with explicit a priori tail bounds, yielding controllable approximation errors and stable computation\.

For any thresholdz∈ℝz\\in\\mathbb\{R\}, the reward and work metrics under thezz\-threshold policy \(see \([16](https://arxiv.org/html/2606.11192#S3.E16)\)\) are the unique bounded solutions of

F​\(x,z\)=\{r​κ​x\+β​κ​x​F​\(p11,z\)\+β​\(1−κ​x\)​F​\(ϕ1​\(x\),z\),x\>z,β​F​\(ϕ0​\(x\),z\),x≤z,F\(x,z\)=\\begin\{cases\}r\\kappa x\+\\beta\\kappa x\\,F\(p\_\{11\},z\)\+\\beta\(1\-\\kappa x\)\\,F\(\\phi^\{1\}\(x\),z\),&x\>z,\\\\\[2\.84526pt\] \\beta\\,F\(\\phi^\{0\}\(x\),z\),&x\\leq z,\\end\{cases\}\(43\)and

G​\(x,z\)=\{1\+β​κ​x​G​\(p11,z\)\+β​\(1−κ​x\)​G​\(ϕ1​\(x\),z\),x\>z,β​G​\(ϕ0​\(x\),z\),x≤z\.G\(x,z\)=\\begin\{cases\}1\+\\beta\\kappa x\\,G\(p\_\{11\},z\)\+\\beta\(1\-\\kappa x\)\\,G\(\\phi^\{1\}\(x\),z\),&x\>z,\\\\\[2\.84526pt\] \\beta\\,G\(\\phi^\{0\}\(x\),z\),&x\\leq z\.\\end\{cases\}\(44\)
To isolate the pre\-ACK contribution, define the discounted*pre\-ACK reward*and*pre\-ACK work*metrics by

F~​\(x,z\)≜r​κ​∑t=0∞βt​Γt​\(x,z\)​X~t​\(x,z\)​A~t​\(x,z\),G~​\(x,z\)≜∑t=0∞βt​Γt​\(x,z\)​A~t​\(x,z\)\.\\widetilde\{F\}\(x,z\)\\triangleq r\\kappa\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\),\\qquad\\widetilde\{G\}\(x,z\)\\triangleq\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\)\.\(45\)These quantities are well defined, since0≤Γt​\(x,z\)≤10\\leq\\Gamma\_\{t\}\(x,z\)\\leq 1andA~t​\(x,z\)∈\{0,1\}\\widetilde\{A\}\_\{t\}\(x,z\)\\in\\\{0,1\\\}, and they depend only on the no\-ACK skeleton\(X~t,A~t\)\(\\widetilde\{X\}\_\{t\},\\widetilde\{A\}\_\{t\}\)and the associated survival probabilitiesΓt​\(x,z\)\\Gamma\_\{t\}\(x,z\)\.

ForT∈ℤ\+T\\in\\mathbb\{Z\}\_\{\+\}, define the truncated approximations

F~T​\(x,z\)≜r​κ​∑t=0T−1βt​Γt​\(x,z\)​X~t​\(x,z\)​A~t​\(x,z\),G~T​\(x,z\)≜∑t=0T−1βt​Γt​\(x,z\)​A~t​\(x,z\)\.\\widetilde\{F\}\_\{T\}\(x,z\)\\triangleq r\\kappa\\sum\_\{t=0\}^\{T\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\),\\qquad\\widetilde\{G\}\_\{T\}\(x,z\)\\triangleq\\sum\_\{t=0\}^\{T\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\)\.Since0≤Γt≤10\\leq\\Gamma\_\{t\}\\leq 1,0≤X~t≤10\\leq\\widetilde\{X\}\_\{t\}\\leq 1, andA~t∈\{0,1\}\\widetilde\{A\}\_\{t\}\\in\\\{0,1\\\}, the tails satisfy the uniform bounds

0≤G~​\(x,z\)−G~T​\(x,z\)≤∑t=T∞βt=βT1−β,0≤F~​\(x,z\)−F~T​\(x,z\)≤r​κ​∑t=T∞βt=r​κ​βT1−β\.0\\leq\\widetilde\{G\}\(x,z\)\-\\widetilde\{G\}\_\{T\}\(x,z\)\\leq\\sum\_\{t=T\}^\{\\infty\}\\beta^\{t\}=\\frac\{\\beta^\{T\}\}\{1\-\\beta\},\\qquad 0\\leq\\widetilde\{F\}\(x,z\)\-\\widetilde\{F\}\_\{T\}\(x,z\)\\leq r\\kappa\\sum\_\{t=T\}^\{\\infty\}\\beta^\{t\}=\\frac\{r\\kappa\\,\\beta^\{T\}\}\{1\-\\beta\}\.\(46\)Thus, given a toleranceε\>0\\varepsilon\>0, choosingTTso thatβT/\(1−β\)≤ε\\beta^\{T\}/\(1\-\\beta\)\\leq\\varepsilonyields an a prioriε\\varepsilon\-accurate evaluation ofG~\\widetilde\{G\}, while choosingTTso thatr​κ​βT/\(1−β\)≤εr\\kappa\\,\\beta^\{T\}/\(1\-\\beta\)\\leq\\varepsilonyields an a prioriε\\varepsilon\-accurate evaluation ofF~\\widetilde\{F\}\.

For later use, define the discounted transform of the first ACK timeτack\\tau^\{\\mathrm\{ack\}\}by

Θ~​\(x,z\)≜𝔼xz​\[βτack\]=∑t=0∞βt​ℙxz​\{τack=t\},\\widetilde\{\\Theta\}\(x,z\)\\triangleq\\mathbb\{E\}\_\{x\}^\{z\}\[\\beta^\{\\tau^\{\\mathrm\{ack\}\}\}\]=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\},\(47\)with the conventionβ∞≜0\\beta^\{\\infty\}\\triangleq 0\. Since\{τack=t\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\}means that no ACK occurs before periodttand an ACK occurs in periodtt, we have

ℙxz​\{τack=t\}=Γt​\(x,z\)​κ​X~t​\(x,z\)​A~t​\(x,z\),t≥0\.\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\}=\\Gamma\_\{t\}\(x,z\)\\,\\kappa\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\),\\qquad t\\geq 0\.A key simplification, established in Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(a\), is

F~​\(x,z\)=r​Θ~​\(x,z\)\.\\widetilde\{F\}\(x,z\)=r\\,\\widetilde\{\\Theta\}\(x,z\)\.\(48\)Hence the tail bound forF~\\widetilde\{F\}in \([46](https://arxiv.org/html/2606.11192#S4.E46)\) immediately yields an a priori tail bound forΘ~\\widetilde\{\\Theta\}via the identityΘ~​\(x,z\)=F~​\(x,z\)/r\\widetilde\{\\Theta\}\(x,z\)=\\widetilde\{F\}\(x,z\)/r\.

We now decomposeF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)at the first ACK time\. Letτack\\tau^\{\\mathrm\{ack\}\}be the first ACK time under thezz\-threshold policy, withτack=∞\\tau^\{\\mathrm\{ack\}\}=\\inftyif no ACK ever occurs\. An ACK regenerates the process by resetting the belief top11p\_\{11\}at the next period, soF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)split into a pre\-ACK contribution plus a discounted continuation value from the post\-ACK statep11p\_\{11\}\. The next lemma makes this precise\.

###### Lemma 4\.8\(Renewal decomposition at the first ACK time\)\.

Fixz∈ℝz\\in\\mathbb\{R\}andx∈𝒳x\\in\\mathscr\{X\}\. Then:

1. \(a\)Pre\-ACK representations via the no\-ACK skeleton\. 𝔼xz​\[∑t=0τackβt​r​κ​X​\(t\)​A​\(t\)\]=F~​\(x,z\),𝔼xz​\[∑t=0τackβt​A​\(t\)\]=G~​\(x,z\),Θ~​\(x,z\)=1r​F~​\(x,z\)\.\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,r\\kappa\\,X\(t\)A\(t\)\\bigg\]=\\widetilde\{F\}\(x,z\),\\qquad\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,A\(t\)\\bigg\]=\\widetilde\{G\}\(x,z\),\\qquad\\widetilde\{\\Theta\}\(x,z\)=\\frac\{1\}\{r\}\\,\\widetilde\{F\}\(x,z\)\.
2. \(b\)Renewal decomposition\. F​\(x,z\)=F~​\(x,z\)\+β​Θ~​\(x,z\)​F​\(p11,z\),G​\(x,z\)=G~​\(x,z\)\+β​Θ~​\(x,z\)​G​\(p11,z\)\.F\(x,z\)=\\widetilde\{F\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,F\(p\_\{11\},z\),\\qquad G\(x,z\)=\\widetilde\{G\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,G\(p\_\{11\},z\)\.
3. \(c\)Post\-ACK fixed\-point evaluation\. We have0≤β​Θ~​\(p11,z\)<10\\leq\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)<1, and F​\(p11,z\)=F~​\(p11,z\)1−β​Θ~​\(p11,z\)=r​Θ~​\(p11,z\)1−β​Θ~​\(p11,z\),G​\(p11,z\)=G~​\(p11,z\)1−β​Θ~​\(p11,z\)\.F\(p\_\{11\},z\)=\\frac\{\\widetilde\{F\}\(p\_\{11\},z\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\}=\\frac\{r\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\},\\qquad G\(p\_\{11\},z\)=\\frac\{\\widetilde\{G\}\(p\_\{11\},z\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\}\.\(49\)

The proof is deferred to Appendix[A\.2](https://arxiv.org/html/2606.11192#A1.SS2)\.

Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)expressesF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)in terms of the pre\-ACK quantitiesF~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), and the discounted transformΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)of the first ACK timeτack\\tau^\{\\mathrm\{ack\}\}\. In several threshold regimes these objects simplify substantially—for example, whenτack\\tau^\{\\mathrm\{ack\}\}is a\.s\. finite, whenτack=∞\\tau^\{\\mathrm\{ack\}\}=\\inftya\.s\., or when the no\-ACK skeleton permits only finitely many activations—yielding especially transparent closed forms\. We exploit these simplifications next\.

### 4\.4Tractable threshold regimes and regime\-wise formulas

The no\-ACK skeleton and the renewal identities provide a general route to evaluatingF​\(x,z\)F\(x,z\),G​\(x,z\)G\(x,z\), and the marginal metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)\. Several threshold regimes, however, admit substantial simplifications\. In this subsection we classify the corresponding skeleton itineraries, determine the finiteness behavior of the first ACK timeτack\\tau^\{\\mathrm\{ack\}\}, and record the resulting regime\-wise formulas forF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)\.

We use the standard word notation for binary itineraries: fora∈\{0,1\}a\\in\\\{0,1\\\}andm∈ℤ\+m\\in\\mathbb\{Z\}\_\{\+\},ama^\{m\}denotes a block ofmmrepeated symbolsaa, anda∞a^\{\\infty\}denotes the infinite constant sequence\(a,a,…\)\(a,a,\\ldots\)\. Concatenation is written by juxtaposition, so01∞01^\{\\infty\}denotes\(0,1,1,…\)\(0,1,1,\\ldots\), and more generally0t0​1∞0^\{t\_\{0\}\}1^\{\\infty\}denotest0t\_\{0\}consecutive zeros followed by ones forever\.

###### Lemma 4\.9\(Skeleton itineraries by threshold regime\)\.

Fixx∈𝒳x\\in\\mathscr\{X\}andz∈ℝz\\in\\mathbb\{R\}\. The no\-ACK skeleton itineraryσ~​\(x,z\)\\widetilde\{\\sigma\}\(x,z\)behaves as follows:

1. \(a\)Ifz<p01z<p\_\{01\}, thenσ~​\(x,z\)=\{1∞,x\>z,01∞,x≤z\.\\widetilde\{\\sigma\}\(x,z\)=\\begin\{cases\}1^\{\\infty\},&x\>z,\\\\ 01^\{\\infty\},&x\\leq z\.\\end\{cases\}
2. \(b\)Ifp01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}, then, witht0≜τ↑0​\(x,z\)t\_\{0\}\\triangleq\\tau\_\{\\uparrow\}^\{0\}\(x,z\),σ~​\(x,z\)=0t0​1∞\.\\widetilde\{\\sigma\}\(x,z\)=0^\{t\_\{0\}\}1^\{\\infty\}\.
3. \(c\)Ifx1<z<x0x^\{1\}<z<x^\{0\}, thenσ~​\(x,z\)\\widetilde\{\\sigma\}\(x,z\)alternates indefinitely between finite blocks of11’s and finite blocks of0’s\.
4. \(d\)Ifx0≤z<p11x^\{0\}\\leq z<p\_\{11\}, thenσ~​\(x,z\)=\{0∞,x≤z,1t1​0∞,x\>z,t1≜τ↓1​\(x,z\)​when​x\>z\.\\widetilde\{\\sigma\}\(x,z\)=\\begin\{cases\}0^\{\\infty\},&x\\leq z,\\\\ 1^\{t\_\{1\}\}0^\{\\infty\},&x\>z,\\end\{cases\}\\qquad t\_\{1\}\\triangleq\\tau\_\{\\downarrow\}^\{1\}\(x,z\)\\ \\text\{when \}x\>z\.
5. \(e\)Ifz≥p11z\\geq p\_\{11\}, thenσ~​\(x,z\)=\{0∞,x≤z,10∞,x\>z\.\\widetilde\{\\sigma\}\(x,z\)=\\begin\{cases\}0^\{\\infty\},&x\\leq z,\\\\ 10^\{\\infty\},&x\>z\.\\end\{cases\}

The proof is deferred to Appendix[A\.3](https://arxiv.org/html/2606.11192#A1.SS3)\.

The itinerary classification above determines whether the first ACK timeτack\\tau^\{\\mathrm\{ack\}\}is almost surely finite, almost surely infinite, or defective, and hence governs the form of the renewal identities forF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)\.

###### Lemma 4\.10\(Finiteness regimes ofτack\\tau^\{\\mathrm\{ack\}\}\)\.

Fixz∈ℝz\\in\\mathbb\{R\}andx∈𝒳x\\in\\mathscr\{X\}, and letτack\\tau^\{\\mathrm\{ack\}\}be the first ACK time under thezz\-threshold policy\. Then:

1. \(a\)Ifz<x0z<x^\{0\}, thenℙxz​\{τack<∞\}=1\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}=1\.
2. \(b\)Ifx0≤z<p11x^\{0\}\\leq z<p\_\{11\}, then the passive region\[0,z\]\[0,z\]is trapping: - •ifx≤zx\\leq z, thenτack=∞\\tau^\{\\mathrm\{ack\}\}=\\inftya\.s\.; - •ifx\>zx\>z, then an ACK can occur only during the finite time window0≤t<τ↓1​\(x,z\)0\\leq t<\\tau\_\{\\downarrow\}^\{1\}\(x,z\), and ℙxz​\{τack=∞\}=∏j=0τ↓1​\(x,z\)−1\(1−κ​ϕj1​\(x\)\)∈\(0,1\)\.\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=\\infty\\\}=\\prod\_\{j=0\}^\{\\tau\_\{\\downarrow\}^\{1\}\(x,z\)\-1\}\\bigl\(1\-\\kappa\\phi\_\{j\}^\{1\}\(x\)\\bigr\)\\in\(0,1\)\.\(50\)
3. \(c\)Ifz≥p11z\\geq p\_\{11\}, then activation is possible only at timet=0t=0, soℙxz​\{τack<∞\}=κ​x​1\{x\>z\}\.\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}=\\kappa x\\,\\mathbbm\{1\}\_\{\\\{x\>z\\\}\}\.

The proof is deferred to Appendix[A\.3](https://arxiv.org/html/2606.11192#A1.SS3)\.

###### Proposition 4\.11\(Piecewise constancy of the pre\-ACK metrics in the tractable regimes\)\.

Fixx∈𝒳x\\in\\mathscr\{X\}\. The maps

z⟼F~​\(x,z\),z⟼G~​\(x,z\),z⟼Θ~​\(x,z\)z\\longmapsto\\widetilde\{F\}\(x,z\),\\qquad z\\longmapsto\\widetilde\{G\}\(x,z\),\\qquad z\\longmapsto\\widetilde\{\\Theta\}\(x,z\)are piecewise constant on the tractable threshold regimes\(−∞,x1\]∪\[x0,∞\)\.\(\-\\infty,x^\{1\}\]\\cup\[x^\{0\},\\infty\)\.More precisely:

1. \(a\)Regimez<p01z<p\_\{01\}\.The three metrics are constant on each of the intervals\(−∞,min⁡\{x,p01\}\)\(\-\\infty,\\min\\\{x,p\_\{01\}\\\}\)and\[x,p01\)∩\(−∞,p01\)\[x,p\_\{01\}\)\\cap\(\-\\infty,p\_\{01\}\)\. On the first interval the skeleton itinerary is1∞1^\{\\infty\}, and on the second interval \(when nonempty\) it is01∞01^\{\\infty\}\.
2. \(b\)Regimep01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}\.For eachn∈ℤ\+n\\in\\mathbb\{Z\}\_\{\+\}, letIn↑​\(x\)≜\{z∈\[p01,x1\]:τ↑0​\(x,z\)=n\}\.I\_\{n\}^\{\\uparrow\}\(x\)\\triangleq\\\{z\\in\[p\_\{01\},x^\{1\}\]:\\ \\tau\_\{\\uparrow\}^\{0\}\(x,z\)=n\\\}\.ThenIn↑​\(x\)I\_\{n\}^\{\\uparrow\}\(x\)is an interval \(possibly empty\), and on each nonempty intervalIn↑​\(x\)I\_\{n\}^\{\\uparrow\}\(x\)one hasσ~​\(x,z\)=0n​1∞\.\\widetilde\{\\sigma\}\(x,z\)=0^\{n\}1^\{\\infty\}\.Consequently,F~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), andΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)are constant on each nonempty intervalIn↑​\(x\)I\_\{n\}^\{\\uparrow\}\(x\)\.
3. \(c\)Regimex0≤z<p11x^\{0\}\\leq z<p\_\{11\}\.Ifx≤zx\\leq z, thenF~​\(x,z\)=G~​\(x,z\)=Θ~​\(x,z\)=0\.\\widetilde\{F\}\(x,z\)=\\widetilde\{G\}\(x,z\)=\\widetilde\{\\Theta\}\(x,z\)=0\. Ifx\>zx\>z, then for eachn∈ℕn\\in\\mathbb\{N\}, letIn↓​\(x\)≜\{z∈\[x0,min⁡\{x,p11\}\):τ↓1​\(x,z\)=n\}\.I\_\{n\}^\{\\downarrow\}\(x\)\\triangleq\\\{z\\in\[x^\{0\},\\min\\\{x,p\_\{11\}\\\}\):\\ \\tau\_\{\\downarrow\}^\{1\}\(x,z\)=n\\\}\.ThenIn↓​\(x\)I\_\{n\}^\{\\downarrow\}\(x\)is an interval \(possibly empty\), and on each nonempty intervalIn↓​\(x\)I\_\{n\}^\{\\downarrow\}\(x\)one hasσ~​\(x,z\)=1n​0∞\.\\widetilde\{\\sigma\}\(x,z\)=1^\{n\}0^\{\\infty\}\.Consequently,F~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), andΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)are constant on each nonempty intervalIn↓​\(x\)I\_\{n\}^\{\\downarrow\}\(x\)\.
4. \(d\)Regimez≥p11z\\geq p\_\{11\}\.The three metrics are constant on each of the intervals\[p11,x\)∩\[p11,∞\)\[p\_\{11\},x\)\\cap\[p\_\{11\},\\infty\)and\[max⁡\{x,p11\},∞\)\[\\max\\\{x,p\_\{11\}\\\},\\infty\)\. On the first interval \(when nonempty\) the skeleton itinerary is10∞10^\{\\infty\}, and on the second interval it is0∞0^\{\\infty\}\.

The proof is deferred to Appendix[A\.3](https://arxiv.org/html/2606.11192#A1.SS3)\.

Building on the regime classification of the no\-ACK skeleton and the renewal identities at the first ACK time, we next record the corresponding formulas forF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)\.

###### Proposition 4\.12\(Reward/work metrics under azz\-threshold policy: regime\-wise formulas\)\.

Fixx∈𝒳x\\in\\mathscr\{X\}\. On the tractable threshold regimes\(−∞,x1\]∪\[x0,∞\),\(\-\\infty,x^\{1\}\]\\cup\[x^\{0\},\\infty\),the functionsF​\(x,⋅\)F\(x,\\cdot\)andG​\(x,⋅\)G\(x,\\cdot\)are right\-continuous step functions ofzz, possibly with infinitely many jumps, and any discontinuity can occur only at a thresholdzzthat coincides with a belief value visited along a no\-ACK skeleton segment started either fromxxor from the post\-ACK reset beliefp11p\_\{11\}\.

Moreover,F​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)admit the following regime\-wise expressions:

1. \(a\)Regimez<p01z<p\_\{01\}\. F​\(x,z\)=\{r​κ​\[x01−β−x0−x1−β​ρ\],x\>z,β​r​κ​\[x01−β−x0−ϕ0​\(x\)1−β​ρ\],x≤z,G​\(x,z\)=\{11−β,x\>z,β1−β,x≤z\.F\(x,z\)=\\begin\{cases\}r\\kappa\\left\[\\dfrac\{x^\{0\}\}\{1\-\\beta\}\-\\dfrac\{x^\{0\}\-x\}\{1\-\\beta\\rho\}\\right\],&x\>z,\\\\\[5\.69054pt\] \\beta\\,r\\kappa\\left\[\\dfrac\{x^\{0\}\}\{1\-\\beta\}\-\\dfrac\{x^\{0\}\-\\phi^\{0\}\(x\)\}\{1\-\\beta\\rho\}\\right\],&x\\leq z,\\end\{cases\}\\qquad G\(x,z\)=\\begin\{cases\}\\dfrac\{1\}\{1\-\\beta\},&x\>z,\\\\\[2\.84526pt\] \\dfrac\{\\beta\}\{1\-\\beta\},&x\\leq z\.\\end\{cases\}
2. \(b\)Regimep01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}\. F​\(x,z\)=r​κ​βτ↑0​\(x,z\)​\[x01−β−x0−ϕτ↑0​\(x,z\)0​\(x\)1−β​ρ\],G​\(x,z\)=βτ↑0​\(x,z\)1−β\.F\(x,z\)=r\\kappa\\,\\beta^\{\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\}\\left\[\\frac\{x^\{0\}\}\{1\-\\beta\}\-\\frac\{x^\{0\}\-\\phi^\{0\}\_\{\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\}\(x\)\}\{1\-\\beta\\rho\}\\right\],\\qquad G\(x,z\)=\\frac\{\\beta^\{\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\}\}\{1\-\\beta\}\.
3. \(c\)Regimex1<z<x0x^\{1\}<z<x^\{0\}\.By Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8), F​\(x,z\)\\displaystyle F\(x,z\)=F~​\(x,z\)\+β​Θ~​\(x,z\)​F​\(p11,z\),G​\(x,z\)=G~​\(x,z\)\+β​Θ~​\(x,z\)​G​\(p11,z\),\\displaystyle=\\widetilde\{F\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,F\(p\_\{11\},z\),\\qquad G\(x,z\)=\\widetilde\{G\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,G\(p\_\{11\},z\),withF~,G~\\widetilde\{F\},\\widetilde\{G\}defined by \([45](https://arxiv.org/html/2606.11192#S4.E45)\)\. Moreover, F​\(p11,z\)=F~​\(p11,z\)1−β​Θ~​\(p11,z\),G​\(p11,z\)=G~​\(p11,z\)1−β​Θ~​\(p11,z\)\.F\(p\_\{11\},z\)=\\frac\{\\widetilde\{F\}\(p\_\{11\},z\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\},\\qquad G\(p\_\{11\},z\)=\\frac\{\\widetilde\{G\}\(p\_\{11\},z\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\}\.
4. \(d\)Regimex0≤z<p11x^\{0\}\\leq z<p\_\{11\}\.For fixedxx, the functionsF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)are piecewise constant inzz\. Ifx≤zx\\leq z, thenF​\(x,z\)=G​\(x,z\)=0\.F\(x,z\)=G\(x,z\)=0\. Assumex\>zx\>z\. Forn,m∈ℕn,m\\in\\mathbb\{N\}, define the intervals In,m​\(x\)≜\{z∈\[x0,min⁡\{x,p11\}\):τ↓1​\(x,z\)=n,τ↓1​\(p11,z\)=m\}\.I\_\{n,m\}\(x\)\\triangleq\\Bigl\\\{z\\in\[x^\{0\},\\min\\\{x,p\_\{11\}\\\}\):\\tau\_\{\\downarrow\}^\{1\}\(x,z\)=n,\\ \\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},z\)=m\\Bigr\\\}\.Then eachIn,m​\(x\)I\_\{n,m\}\(x\)is an interval \(possibly empty\), andF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)are constant on each nonemptyIn,m​\(x\)I\_\{n,m\}\(x\)\. More explicitly, ifz∈In,m​\(x\)z\\in I\_\{n,m\}\(x\), define Ax\(n\)\\displaystyle A\_\{x\}^\{\(n\)\}≜∑t=0n−1βt​∏j=0t−1\(1−κ​ϕj1​\(x\)\),\\displaystyle\\triangleq\\sum\_\{t=0\}^\{n\-1\}\\beta^\{t\}\\prod\_\{j=0\}^\{t\-1\}\\bigl\(1\-\\kappa\\,\\phi\_\{j\}^\{1\}\(x\)\\bigr\),Bx\(n\)\\displaystyle B\_\{x\}^\{\(n\)\}≜κ​∑t=0n−1βt​∏j=0t−1\(1−κ​ϕj1​\(x\)\)​ϕt1​\(x\),\\displaystyle\\triangleq\\kappa\\sum\_\{t=0\}^\{n\-1\}\\beta^\{t\}\\prod\_\{j=0\}^\{t\-1\}\\bigl\(1\-\\kappa\\,\\phi\_\{j\}^\{1\}\(x\)\\bigr\)\\,\\phi\_\{t\}^\{1\}\(x\),and similarly A11\(m\)\\displaystyle A\_\{11\}^\{\(m\)\}≜∑t=0m−1βt​∏j=0t−1\(1−κ​ϕj1​\(p11\)\),\\displaystyle\\triangleq\\sum\_\{t=0\}^\{m\-1\}\\beta^\{t\}\\prod\_\{j=0\}^\{t\-1\}\\bigl\(1\-\\kappa\\,\\phi\_\{j\}^\{1\}\(p\_\{11\}\)\\bigr\),B11\(m\)\\displaystyle B\_\{11\}^\{\(m\)\}≜κ​∑t=0m−1βt​∏j=0t−1\(1−κ​ϕj1​\(p11\)\)​ϕt1​\(p11\)\.\\displaystyle\\triangleq\\kappa\\sum\_\{t=0\}^\{m\-1\}\\beta^\{t\}\\prod\_\{j=0\}^\{t\-1\}\\bigl\(1\-\\kappa\\,\\phi\_\{j\}^\{1\}\(p\_\{11\}\)\\bigr\)\\,\\phi\_\{t\}^\{1\}\(p\_\{11\}\)\.ThenG~​\(x,z\)=Ax\(n\),Θ~​\(x,z\)=Bx\(n\),F~​\(x,z\)=r​Bx\(n\),\\widetilde\{G\}\(x,z\)=A\_\{x\}^\{\(n\)\},\\qquad\\widetilde\{\\Theta\}\(x,z\)=B\_\{x\}^\{\(n\)\},\\qquad\\widetilde\{F\}\(x,z\)=r\\,B\_\{x\}^\{\(n\)\},and F​\(p11,z\)=r​B11\(m\)1−β​B11\(m\),G​\(p11,z\)=A11\(m\)1−β​B11\(m\)\.F\(p\_\{11\},z\)=\\frac\{r\\,B\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\},\\qquad G\(p\_\{11\},z\)=\\frac\{A\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\}\.Consequently, F​\(x,z\)=r​Bx\(n\)1−β​B11\(m\),G​\(x,z\)=Ax\(n\)\+β​Bx\(n\)​A11\(m\)1−β​B11\(m\)=Ax\(n\)​\(1−β​B11\(m\)\)\+β​Bx\(n\)​A11\(m\)1−β​B11\(m\)\.F\(x,z\)=\\frac\{r\\,B\_\{x\}^\{\(n\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\},\\quad G\(x,z\)=A\_\{x\}^\{\(n\)\}\+\\frac\{\\beta\\,B\_\{x\}^\{\(n\)\}\\,A\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\}=\\frac\{A\_\{x\}^\{\(n\)\}\\bigl\(1\-\\beta\\,B\_\{11\}^\{\(m\)\}\\bigr\)\+\\beta\\,B\_\{x\}^\{\(n\)\}\\,A\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\}\.
5. \(e\)Regimez≥p11z\\geq p\_\{11\}\.F​\(x,z\)=r​κ​x​1\{x\>z\},G​\(x,z\)=𝟙\{x\>z\}\.F\(x,z\)=r\\kappa x\\,\\mathbbm\{1\}\_\{\\\{x\>z\\\}\},\\kern 5\.0ptG\(x,z\)=\\mathbbm\{1\}\_\{\\\{x\>z\\\}\}\.

The proof is deferred to Appendix[A\.3](https://arxiv.org/html/2606.11192#A1.SS3)\.

###### Lemma 4\.13\(Monotonicity of the work metric in the tractable regimes\)\.

Fixz∈\(−∞,x1\]∪\[x0,∞\)z\\in\(\-\\infty,x^\{1\}\]\\cup\[x^\{0\},\\infty\)\. Then the functionx↦G​\(x,z\)x\\mapsto G\(x,z\)is nondecreasing on𝒳\\mathscr\{X\}\.

The proof is deferred to Appendix[A\.3](https://arxiv.org/html/2606.11192#A1.SS3)\.

Together, the preceding lemmas and propositions provide the regime\-wise formulas used later for numerical evaluation ofFF,GG, and the MP index, and for verifying the PCL conditions\.

### 4\.5Computation of marginal metrics

We next turn to the marginal reward and work metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)associated with a one\-step deviation from azz\-threshold policy\. These marginal quantities are the ingredients of the MP index \(m​\(x,z\)≜f​\(x,z\)/g​\(x,z\),m​\(x\)≜m​\(x,x\),m\(x,z\)\\triangleq\{f\(x,z\)\}/\{g\(x,z\)\},m\(x\)\\triangleq m\(x,x\),\) which equals the Whittle index if the PCL conditions hold\. Using the renewal decompositions and pre\-ACK series developed above, we expressffandggin terms of pre\-ACK contributions and continuation values from the post\-ACK reset state, yielding formulas that are efficient to evaluate and admit controlled truncation error\.

Fora∈\{0,1\}a\\in\\\{0,1\\\}and thresholdzz, let⟨a,z⟩\\langle a,z\\rangledenote the policy that takes actionaaat time0and follows thezz\-threshold policy from time11onward\. Letτack\\tau^\{\\mathrm\{ack\}\}denote the corresponding first ACK time under that policy\. Define the associated pre\-ACK reward, work, and transform metrics by \(with the conventionβ∞≜0\\beta^\{\\infty\}\\triangleq 0\)

F~​\(x,⟨a,z⟩\)\\displaystyle\\widetilde\{F\}\(x,\\langle a,z\\rangle\)≜𝔼x⟨a,z⟩​\[∑t=0τackβt​r​κ​X​\(t\)​A​\(t\)\],G~​\(x,⟨a,z⟩\)≜𝔼x⟨a,z⟩​\[∑t=0τackβt​A​\(t\)\],Θ~​\(x,⟨a,z⟩\)≜𝔼x⟨a,z⟩​\[βτack\]\.\\displaystyle\\triangleq\\mathbb\{E\}\_\{x\}^\{\\langle a,z\\rangle\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,r\\kappa\\,X\(t\)A\(t\)\\bigg\],\\quad\\widetilde\{G\}\(x,\\langle a,z\\rangle\)\\triangleq\\mathbb\{E\}\_\{x\}^\{\\langle a,z\\rangle\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,A\(t\)\\bigg\],\\quad\\widetilde\{\\Theta\}\(x,\\langle a,z\\rangle\)\\triangleq\\mathbb\{E\}\_\{x\}^\{\\langle a,z\\rangle\}\\\!\\big\[\\beta^\{\\tau^\{\\mathrm\{ack\}\}\}\\big\]\.
We then define the*pre\-ACK marginal*metrics as the differences between active and passive initializations:

f~​\(x,z\)\\displaystyle\\tilde\{f\}\(x,z\)≜F~​\(x,⟨1,z⟩\)−F~​\(x,⟨0,z⟩\)=r​κ​x\+β​\(\(1−κ​x\)​F~​\(ϕ1​\(x\),z\)−F~​\(ϕ0​\(x\),z\)\),\\displaystyle\\triangleq\\widetilde\{F\}\(x,\\langle 1,z\\rangle\)\-\\widetilde\{F\}\(x,\\langle 0,z\\rangle\)=r\\kappa x\+\\beta\\Big\(\(1\-\\kappa x\)\\,\\widetilde\{F\}\(\\phi^\{1\}\(x\),z\)\-\\widetilde\{F\}\(\\phi^\{0\}\(x\),z\)\\Big\),\(51\)g~​\(x,z\)\\displaystyle\\tilde\{g\}\(x,z\)≜G~​\(x,⟨1,z⟩\)−G~​\(x,⟨0,z⟩\)=1\+β​\(\(1−κ​x\)​G~​\(ϕ1​\(x\),z\)−G~​\(ϕ0​\(x\),z\)\),\\displaystyle\\triangleq\\widetilde\{G\}\(x,\\langle 1,z\\rangle\)\-\\widetilde\{G\}\(x,\\langle 0,z\\rangle\)=1\+\\beta\\Big\(\(1\-\\kappa x\)\\,\\widetilde\{G\}\(\\phi^\{1\}\(x\),z\)\-\\widetilde\{G\}\(\\phi^\{0\}\(x\),z\)\\Big\),\(52\)θ~​\(x,z\)\\displaystyle\\tilde\{\\theta\}\(x,z\)≜Θ~​\(x,⟨1,z⟩\)−Θ~​\(x,⟨0,z⟩\)=κ​x\+β​\(\(1−κ​x\)​Θ~​\(ϕ1​\(x\),z\)−Θ~​\(ϕ0​\(x\),z\)\),\\displaystyle\\triangleq\\widetilde\{\\Theta\}\(x,\\langle 1,z\\rangle\)\-\\widetilde\{\\Theta\}\(x,\\langle 0,z\\rangle\)=\\kappa x\+\\beta\\Big\(\(1\-\\kappa x\)\\,\\widetilde\{\\Theta\}\(\\phi^\{1\}\(x\),z\)\-\\widetilde\{\\Theta\}\(\\phi^\{0\}\(x\),z\)\\Big\),\(53\)where the identities on the right follow from one\-step conditioning\.

###### Lemma 4\.14\(Decomposition of marginal metrics via pre\-ACK metrics\)\.

For everyx∈𝒳x\\in\\mathscr\{X\}andz∈ℝz\\in\\mathbb\{R\}, the marginal reward and work metrics satisfy

f​\(x,z\)\\displaystyle f\(x,z\)=f~​\(x,z\)\+β​θ~​\(x,z\)​F​\(p11,z\),\\displaystyle=\\tilde\{f\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,F\(p\_\{11\},z\),\(54\)g​\(x,z\)\\displaystyle g\(x,z\)=g~​\(x,z\)\+β​θ~​\(x,z\)​G​\(p11,z\)\.\\displaystyle=\\tilde\{g\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,G\(p\_\{11\},z\)\.\(55\)

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

###### Lemma 4\.15\(Relation betweenf~\\tilde\{f\}andθ~\\tilde\{\\theta\}\)\.

Assumer\>0r\>0\. Then for everyx∈𝒳x\\in\\mathscr\{X\}andz∈ℝz\\in\\mathbb\{R\},

f~​\(x,z\)=r​θ~​\(x,z\)\.\\tilde\{f\}\(x,z\)=r\\,\\tilde\{\\theta\}\(x,z\)\.\(56\)

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

###### Proposition 4\.16\(Marginal reward and work metrics: one\-step and regime\-wise forms\)\.

For everyx∈𝒳x\\in\\mathscr\{X\}andz∈ℝz\\in\\mathbb\{R\}, we have the one\-step identities

f​\(x,z\)\\displaystyle f\(x,z\)=r​κ​x\+β​\(κ​x​F​\(p11,z\)\+\(1−κ​x\)​F​\(ϕ1​\(x\),z\)−F​\(ϕ0​\(x\),z\)\),\\displaystyle=r\\kappa x\+\\beta\\Big\(\\kappa x\\,F\(p\_\{11\},z\)\+\(1\-\\kappa x\)\\,F\(\\phi^\{1\}\(x\),z\)\-F\(\\phi^\{0\}\(x\),z\)\\Big\),\(57\)g​\(x,z\)\\displaystyle g\(x,z\)=1\+β​\(κ​x​G​\(p11,z\)\+\(1−κ​x\)​G​\(ϕ1​\(x\),z\)−G​\(ϕ0​\(x\),z\)\),\\displaystyle=1\+\\beta\\Big\(\\kappa x\\,G\(p\_\{11\},z\)\+\(1\-\\kappa x\)\\,G\(\\phi^\{1\}\(x\),z\)\-G\(\\phi^\{0\}\(x\),z\)\\Big\),\(58\)whereF​\(⋅,z\)F\(\\cdot,z\)andG​\(⋅,z\)G\(\\cdot,z\)are the reward and work metrics under thezz\-threshold policy from Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\.

Substituting the regime\-wise forms ofF​\(⋅,z\)F\(\\cdot,z\)andG​\(⋅,z\)G\(\\cdot,z\)yields the following case split inzz\.

1. \(a\)Regimez<p01z<p\_\{01\}\.We havef​\(x,z\)=r​κ​x,g​\(x,z\)=1\.f\(x,z\)=r\\kappa x,\\qquad g\(x,z\)=1\.
2. \(b\)Regimep01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}\.Letτ↑0​\(⋅,z\)\\tau\_\{\\uparrow\}^\{0\}\(\\cdot,z\)be the passive up\-crossing time, and set τ0≜τ↑0​\(ϕ0​\(x\),z\),τ1≜τ↑0​\(ϕ1​\(x\),z\),ϕ^a≜ϕτa0​\(ϕa​\(x\)\),a∈\{0,1\}\.\\tau\_\{0\}\\triangleq\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{0\}\(x\),z\),\\qquad\\tau\_\{1\}\\triangleq\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{1\}\(x\),z\),\\qquad\\widehat\{\\phi\}\_\{a\}\\triangleq\\phi^\{0\}\_\{\\tau\_\{a\}\}\(\\phi^\{a\}\(x\)\),\\quad a\\in\\\{0,1\\\}\.DefineF\+​\(u\)≜r​κ​\[x01−β−x0−u1−β​ρ\]\.F^\{\+\}\(u\)\\triangleq r\\kappa\\left\[\\frac\{x^\{0\}\}\{1\-\\beta\}\-\\frac\{x^\{0\}\-u\}\{1\-\\beta\\rho\}\\right\]\.Thenf​\(x,z\)=r​κ​x\+β​\(κ​x​F\+​\(p11\)\+\(1−κ​x\)​βτ1​F\+​\(ϕ^1\)−βτ0​F\+​\(ϕ^0\)\),f\(x,z\)=r\\kappa x\+\\beta\\Big\(\\kappa x\\,F^\{\+\}\(p\_\{11\}\)\+\(1\-\\kappa x\)\\,\\beta^\{\\tau\_\{1\}\}F^\{\+\}\(\\widehat\{\\phi\}\_\{1\}\)\-\\beta^\{\\tau\_\{0\}\}F^\{\+\}\(\\widehat\{\\phi\}\_\{0\}\)\\Big\),andg​\(x,z\)=1\+β1−β​\(κ​x\+\(1−κ​x\)​βτ1−βτ0\)\.g\(x,z\)=1\+\\frac\{\\beta\}\{1\-\\beta\}\\Big\(\\kappa x\+\(1\-\\kappa x\)\\beta^\{\\tau\_\{1\}\}\-\\beta^\{\\tau\_\{0\}\}\\Big\)\.
3. \(c\)Regimex1<z<x0x^\{1\}<z<x^\{0\}\.LetD​\(z\)≜1−β​Θ~​\(p11,z\)\.D\(z\)\\triangleq 1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\.Then f​\(x,z\)\\displaystyle f\(x,z\)=r​θ~​\(x,z\)D​\(z\),\\displaystyle=\\frac\{r\\,\\tilde\{\\theta\}\(x,z\)\}\{D\(z\)\},\(59\)g​\(x,z\)\\displaystyle g\(x,z\)=g~​\(x,z\)\+β​θ~​\(x,z\)D​\(z\)​G~​\(p11,z\)\.\\displaystyle=\\tilde\{g\}\(x,z\)\+\\beta\\,\\frac\{\\tilde\{\\theta\}\(x,z\)\}\{D\(z\)\}\\,\\widetilde\{G\}\(p\_\{11\},z\)\.\(60\)
4. \(d\)Regimex0≤z<p11x^\{0\}\\leq z<p\_\{11\}\.The formulas in part\(c\)remain valid; moreover,F~​\(⋅,z\)\\widetilde\{F\}\(\\cdot,z\),G~​\(⋅,z\)\\widetilde\{G\}\(\\cdot,z\), andΘ~​\(⋅,z\)\\widetilde\{\\Theta\}\(\\cdot,z\)reduce to finite sums, truncating as in Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\(d\)\. In particular,F~​\(y,z\)=G~​\(y,z\)=Θ~​\(y,z\)=0for​y≤z\.\\widetilde\{F\}\(y,z\)=\\widetilde\{G\}\(y,z\)=\\widetilde\{\\Theta\}\(y,z\)=0\\qquad\\text\{for \}y\\leq z\.
5. \(e\)Regimez≥p11z\\geq p\_\{11\}\.We havef​\(x,z\)=r​κ​x,g​\(x,z\)=1\.f\(x,z\)=r\\kappa x,\\qquad g\(x,z\)=1\.

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

###### Lemma 4\.17\(Marginal\-work positivity in the tractable regimes\)\.

For everyx∈𝒳x\\in\\mathscr\{X\}:

1. \(a\)Ifz<p01z<p\_\{01\}orz≥p11z\\geq p\_\{11\}, theng​\(x,z\)=1g\(x,z\)=1\.
2. \(b\)Ifp01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}, theng​\(x,z\)≥1−βg\(x,z\)\\geq 1\-\\beta\.
3. \(c\)Ifx0≤z<p11x^\{0\}\\leq z<p\_\{11\}, theng​\(x,z\)≥1−βg\(x,z\)\\geq 1\-\\beta\.

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

### 4\.6Partial closed\-form formulas for the MP index

We now specialize the marginal formulas to the MP indexm​\(x\)≜m​\(x,x\)\.m\(x\)\\triangleq m\(x,x\)\.The next lemma gives general identities forf​\(x,x\)f\(x,x\),g​\(x,x\)g\(x,x\), andm​\(x\)m\(x\)in terms of the pre\-ACK metrics\.

###### Lemma 4\.18\(General formulas forf​\(x,x\)f\(x,x\),g​\(x,x\)g\(x,x\), andm​\(x\)m\(x\)\)\.

For everyx∈𝒳x\\in\\mathscr\{X\}, define

D​\(x\)≜1−β​Θ~​\(p11,x\)\.D\(x\)\\triangleq 1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\.\(61\)ThenD​\(x\)\>0D\(x\)\>0, and

f​\(x,x\)\\displaystyle f\(x,x\)=f~​\(x,x\)D​\(x\)=r​θ~​\(x,x\)D​\(x\),\\displaystyle=\\frac\{\\tilde\{f\}\(x,x\)\}\{D\(x\)\}=\\frac\{r\\,\\tilde\{\\theta\}\(x,x\)\}\{D\(x\)\},\(62\)g​\(x,x\)\\displaystyle g\(x,x\)=g~​\(x,x\)\+βr​f~​\(x,x\)D​\(x\)​G~​\(p11,x\)=g~​\(x,x\)\+β​θ~​\(x,x\)D​\(x\)​G~​\(p11,x\)\.\\displaystyle=\\tilde\{g\}\(x,x\)\+\\frac\{\\beta\}\{r\}\\,\\frac\{\\tilde\{f\}\(x,x\)\}\{D\(x\)\}\\,\\widetilde\{G\}\(p\_\{11\},x\)=\\tilde\{g\}\(x,x\)\+\\beta\\,\\frac\{\\tilde\{\\theta\}\(x,x\)\}\{D\(x\)\}\\,\\widetilde\{G\}\(p\_\{11\},x\)\.\(63\)Consequently,

m​\(x\)\\displaystyle m\(x\)=f~​\(x,x\)D​\(x\)​g~​\(x,x\)\+βr​f~​\(x,x\)​G~​\(p11,x\)=r​θ~​\(x,x\)D​\(x\)​g~​\(x,x\)\+β​θ~​\(x,x\)​G~​\(p11,x\)\.\\displaystyle=\\frac\{\\tilde\{f\}\(x,x\)\}\{D\(x\)\\,\\tilde\{g\}\(x,x\)\+\\frac\{\\beta\}\{r\}\\,\\tilde\{f\}\(x,x\)\\,\\widetilde\{G\}\(p\_\{11\},x\)\}=\\frac\{r\\,\\tilde\{\\theta\}\(x,x\)\}\{D\(x\)\\,\\tilde\{g\}\(x,x\)\+\\beta\\,\\tilde\{\\theta\}\(x,x\)\\,\\widetilde\{G\}\(p\_\{11\},x\)\}\.\(64\)

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

The next proposition gives exact formulas for the MP index in the tractable regimes\. In the low\- and high\-belief regions it coincides with the myopic indexr​κ​xr\\kappa x, while on\[x0,p11\]\[x^\{0\},p\_\{11\}\]it admits an exact finite\-sum representation\.

###### Proposition 4\.19\(Exact formulas for the MP index in the tractable regimes\)\.

1. \(a\)Forx∈\[0,x1\]x\\in\[0,x^\{1\}\],m​\(x\)=r​κ​x\.m\(x\)=r\\kappa x\.
2. \(b\)Forx∈\[x0,p11\]x\\in\[x^\{0\},p\_\{11\}\], letut≜ϕt1​\(p11\),t≥0,u\_\{t\}\\triangleq\\phi\_\{t\}^\{1\}\(p\_\{11\}\),\\kern 5\.0ptt\\geq 0,andΓ011≜1,Γt11≜∏j=0t−1\(1−κ​uj\),t≥1\.\\Gamma\_\{0\}^\{11\}\\triangleq 1,\\kern 5\.0pt\\Gamma\_\{t\}^\{11\}\\triangleq\\prod\_\{j=0\}^\{t\-1\}\(1\-\\kappa u\_\{j\}\),\\kern 5\.0ptt\\geq 1\.Then G~​\(p11,x\)\\displaystyle\\widetilde\{G\}\(p\_\{11\},x\)=∑t=0τ↓1​\(p11,x\)−1βt​Γt11,\\displaystyle=\\sum\_\{t=0\}^\{\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x\)\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}^\{11\},\(65\)Θ~​\(p11,x\)\\displaystyle\\widetilde\{\\Theta\}\(p\_\{11\},x\)=κ​∑t=0τ↓1​\(p11,x\)−1βt​Γt11​ut,\\displaystyle=\\kappa\\sum\_\{t=0\}^\{\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x\)\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}^\{11\}\\,u\_\{t\},\(66\)and f​\(x,x\)\\displaystyle f\(x,x\)=r​κ​x1−β​Θ~​\(p11,x\),\\displaystyle=\\frac\{r\\kappa x\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\},\(67\)g​\(x,x\)\\displaystyle g\(x,x\)=1−β​Θ~​\(p11,x\)\+β​κ​x​G~​\(p11,x\)1−β​Θ~​\(p11,x\),\\displaystyle=\\frac\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\+\\beta\\kappa x\\,\\widetilde\{G\}\(p\_\{11\},x\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\},\(68\)m​\(x\)\\displaystyle m\(x\)=r​κ​x1−β​Θ~​\(p11,x\)\+β​κ​x​G~​\(p11,x\)\.\\displaystyle=\\frac\{r\\kappa x\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\+\\beta\\kappa x\\,\\widetilde\{G\}\(p\_\{11\},x\)\}\.\(69\)Equivalently, m​\(x\)=r​κ​x1\+β​κ​∑t=0τ↓1​\(p11,x\)−1βt​Γt11​\(x−ut\)\.m\(x\)=\\frac\{r\\kappa x\}\{1\+\\beta\\kappa\\displaystyle\\sum\_\{t=0\}^\{\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x\)\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}^\{11\}\\,\\bigl\(x\-u\_\{t\}\\bigr\)\}\.\(70\)
3. \(c\)Forx∈\[p11,1\]x\\in\[p\_\{11\},1\],m​\(x\)=r​κ​x\.m\(x\)=r\\kappa x\.

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

###### Proposition 4\.20\(Continuity and monotonicity of the MP index in the tractable regimes\)\.

Letut,u\_\{t\},Γ011≜1,\\Gamma\_\{0\}^\{11\}\\triangleq 1,andΓt11≜∏j=0t−1\(1−κ​uj\)\\Gamma\_\{t\}^\{11\}\\triangleq\\prod\_\{j=0\}^\{t\-1\}\(1\-\\kappa u\_\{j\}\)be as in Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\. Also letN0≜τ↓1​\(p11,x0\),N\_\{0\}\\triangleq\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x^\{0\}\),which is finite sinceut↓x1<x0u\_\{t\}\\downarrow x^\{1\}<x^\{0\}\. Forn=1,…,N0n=1,\\dots,N\_\{0\}, defineJn≜\[max⁡\{x0,un\},un−1\),An≜∑t=0n−1βt​Γt11,Bn≜∑t=0n−1βt​Γt11​ut\.J\_\{n\}\\triangleq\[\\max\\\{x^\{0\},u\_\{n\}\\\},\\,u\_\{n\-1\}\),\\kern 5\.0ptA\_\{n\}\\triangleq\\sum\_\{t=0\}^\{n\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}^\{11\},\\kern 5\.0ptB\_\{n\}\\triangleq\\sum\_\{t=0\}^\{n\-1\}\\beta^\{t\}\\,\\Gamma\_\{t\}^\{11\}u\_\{t\}\.Then the following hold\.

1. \(a\)On\[0,x1\]\[0,x^\{1\}\], we havem​\(x\)=r​κ​xm\(x\)=r\\kappa x\. Hencem​\(⋅\)m\(\\cdot\)is continuous and increasing on\[0,x1\]\[0,x^\{1\}\]\.
2. \(b\)For eachn=1,…,N0n=1,\\dots,N\_\{0\}and everyx∈Jnx\\in J\_\{n\}, m​\(x\)=r​κ​x1\+β​κ​\(An​x−Bn\)\.m\(x\)=\\frac\{r\\kappa x\}\{1\+\\beta\\kappa\(A\_\{n\}x\-B\_\{n\}\)\}\.\(71\)Hencem​\(⋅\)m\(\\cdot\)is continuous and increasing on each intervalJnJ\_\{n\}\.
3. \(c\)The functionm​\(⋅\)m\(\\cdot\)is continuous at every breakpointunu\_\{n\}with1≤n≤N0−11\\leq n\\leq N\_\{0\}\-1, and also atx=p11x=p\_\{11\}\.
4. \(d\)On\[p11,1\]\[p\_\{11\},1\], we havem​\(x\)=r​κ​xm\(x\)=r\\kappa x\. Hencem​\(⋅\)m\(\\cdot\)is continuous and increasing on\[p11,1\]\[p\_\{11\},1\]\.

Consequently,m​\(⋅\)m\(\\cdot\)is continuous and increasing on each tractable regime\[0,x1\],\[x0,p11\],\[p11,1\],\[0,x^\{1\}\],\\kern 5\.0pt\[x^\{0\},p\_\{11\}\],\\kern 5\.0pt\[p\_\{11\},1\],and in particular is continuous atx=p11x=p\_\{11\}\.

The proof is deferred to Appendix[A\.4](https://arxiv.org/html/2606.11192#A1.SS4)\.

An illustrative parameter instance, together with diagnostic plots of the metricsF~,G~,Θ~,F,G,f,g\\widetilde\{F\},\\widetilde\{G\},\\widetilde\{\\Theta\},F,G,f,gand the diagonal MP indexm​\(x\)m\(x\), is presented in Appendix[B](https://arxiv.org/html/2606.11192#A2)\. Those plots provide a concrete visualization of the tractable\-regime formulas developed above and of the staircase structure that arises in the intermediate regime\.

### 4\.7What remains to be proven

The tractable\-regime analysis of Sections[4\.4](https://arxiv.org/html/2606.11192#S4.SS4)–[4\.6](https://arxiv.org/html/2606.11192#S4.SS6)already verifies the required PCL properties outside the intermediate threshold regionx1<z<x0\.x^\{1\}<z<x^\{0\}\.Moreover, because the belief\-update mapsϕ0\\phi^\{0\}andϕ1\\phi^\{1\}satisfy the analogue ofDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]’s Assumption A2 after an increasing change of variables \(Lemma[4\.7](https://arxiv.org/html/2606.11192#S4.Thmtheorem7)\), the symbolic structure of threshold itineraries in this intermediate region can be imported from the maps\-with\-gaps theory ofDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]; see Theorem[C\.3](https://arxiv.org/html/2606.11192#A3.Thmtheorem3), Corollary[C\.4](https://arxiv.org/html/2606.11192#A3.Thmtheorem4), and Proposition[C\.5](https://arxiv.org/html/2606.11192#A3.Thmtheorem5)below\.

Accordingly, the remaining issue is no longer to describe the no\-ACK threshold itineraries themselves, but to convert that symbolic description into the metric statements needed for full verification of \(PCLI1\)–\(PCLI3\) onx1<z<x0x^\{1\}<z<x^\{0\}\. The subsections below push this reduction substantially further: they derive exact Christoffel\-interval formulas for the diagonal marginal metricsf​\(x,x\)f\(x,x\),g​\(x,x\)g\(x,x\), andm​\(x\)m\(x\), and reduce the proof of \(PCLI2\) to an intervalwise monotonicity problem\.

Thus, the main remaining analytic task is to establish \(PCLI1\) and \(PCLI2\) onx1<z<x0x^\{1\}<z<x^\{0\}\. By contrast, \(PCLI3\) is not an independent obstacle: as shown in Subsection[4\.8](https://arxiv.org/html/2606.11192#S4.SS8), once \(PCLI1\) and \(PCLI2\) are available on that region, \(PCLI3\) follows from Proposition 6 of\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]via Assumption 4\.

The symbolic organization of threshold itineraries onx1<z<x0x^\{1\}<z<x^\{0\}, including the Christoffel–Sturmian structure inherited from the maps\-with\-gaps theory ofDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\], is recorded in Appendix[C](https://arxiv.org/html/2606.11192#A3)\. That appendix identifies the itinerary patterns that would need to be translated into metric statements in order to complete the remaining analytic verification\.

### 4\.8Reduction of \(PCLI3\) to \(PCLI1\) and \(PCLI2\)

Condition \(PCLI3\) is not an independent obstacle in the present model\. Indeed, Proposition 6 in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]shows that, under \(PCLI1\) and \(PCLI2\), \(PCLI3\) follows once one verifies either Assumption 3 \(piecewise\-constantG​\(x,⋅\)G\(x,\\cdot\)\) or the more structural Assumption 4, which requires that, for each initial statexx, all states visited under any threshold policy lie almost surely in a countable setD​\(x\)D\(x\)\. For our model, Assumption 4 is natural and easy to verify\.

###### Proposition 4\.22\(Verification of Assumption 4 for the present model\)\.

Fix an initial beliefx∈𝒳x\\in\\mathscr\{X\}, and define

D​\(x\)≜\{ϕw​\(x\):w∈\{0,1\}∗\}∪\{ϕw​\(p11\):w∈\{0,1\}∗\},D\(x\)\\triangleq\\\{\\phi^\{w\}\(x\):\\,w\\in\\\{0,1\\\}^\{\*\}\\\}\\ \\cup\\ \\\{\\phi^\{w\}\(p\_\{11\}\):\\,w\\in\\\{0,1\\\}^\{\*\}\\\},\(72\)where\{0,1\}∗\\\{0,1\\\}^\{\*\}denotes the set of all finite binary words, including the empty word\. ThenD​\(x\)D\(x\)is countable, and for every thresholdz∈ℝz\\in\\mathbb\{R\},

ℙxz​\{\{X​\(t\)\}t≥0⊆D​\(x\)\}=1\.\\mathbb\{P\}\_\{x\}^\{z\}\\\!\\bigl\\\{\\\{X\(t\)\\\}\_\{t\\geq 0\}\\subseteq D\(x\)\\bigr\\\}=1\.\(73\)Hence Assumption 4 in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]holds for the present model\.

###### Proof\.

The set\{0,1\}∗\\\{0,1\\\}^\{\*\}of finite binary words is countable, soD​\(x\)D\(x\)is countable\. Now fix a thresholdzz\. Before the first ACK time, the belief state follows the no\-ACK skeleton under the strict\-threshold policy, and hence every visited state is of the formϕw​\(x\)\\phi^\{w\}\(x\)for some finite wordww\. When an ACK occurs, the belief resets top11p\_\{11\}at the next period\. Thereafter, until the next ACK, the visited states are of the formϕw​\(p11\)\\phi^\{w\}\(p\_\{11\}\)for finite wordsww\. The same argument applies after every subsequent ACK, because each post\-ACK excursion again starts fromp11p\_\{11\}\. Therefore every state visited along the sample path belongs toD​\(x\)D\(x\), proving \([73](https://arxiv.org/html/2606.11192#S4.E73)\)\. This is precisely Assumption 4 in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]\. ∎

The preceding proposition has the following immediate consequence\.

###### Corollary 4\.23\(Reduction of \(PCLI3\) to \(PCLI1\) and \(PCLI2\)\)\.

For the present model, whenever\(PCLI1\)and\(PCLI2\)hold on a given threshold region,\(PCLI3\)also holds on that region\.

###### Proof\.

By Proposition[4\.22](https://arxiv.org/html/2606.11192#S4.Thmtheorem22), Assumption 4 in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]holds for the present model\. Hence Proposition 6 in\[[31](https://arxiv.org/html/2606.11192#bib.bib31), §10\.2\]applies and yields \(PCLI3\) under \(PCLI1\) and \(PCLI2\)\. ∎

Corollary[4\.23](https://arxiv.org/html/2606.11192#S4.Thmtheorem23)shows that \(PCLI3\) is not an independent obstacle herein\. In particular, on the tractable threshold regimes, where \(PCLI1\) and \(PCLI2\) have already been established, \(PCLI3\) follows immediately\. Likewise, on the intermediate regimex1<z<x0x^\{1\}<z<x^\{0\}, once \(PCLI1\) and \(PCLI2\) are proved, \(PCLI3\) follows automatically by the same argument\. Thus the only genuinely unresolved analytical task is the verification of \(PCLI1\) and \(PCLI2\) on that regime\.

## 5Computational experiments

This section complements the analytical results with large\-scale computational experiments\. Our goals are threefold: first, to provide broad numerical evidence for the remaining PCL\-indexability conditions in parameter regimes not covered analytically; second, to assess the behavior of the MP index across representative problem families; and third, to benchmark the resulting MP index policy against standard alternatives and the Lagrangian dual upper bound\. We begin with extensive numerical tests of \(PCLI1\) and \(PCLI2\), and then turn to policy\-performance comparisons on heterogeneous multi\-project instances\. Additional parameter\-dependence plots for the MP index, together with implementation and reproducibility details, are reported in Appendix[E](https://arxiv.org/html/2606.11192#A5)\.

### 5\.1Numerical verification of PCLI conditions

We begin with large\-scale numerical tests of the two key PCL conditions that remain analytically unresolved in the intermediate regime, namely \(PCLI1\) and \(PCLI2\)\. In both experiments, we explore a broad four\-dimensional parameter grid in\(q,ρ,κ,β\)\(q,\\rho,\\kappa,\\beta\), withρ\\rhoparametrized asρ=α​\(1−q\),\\rho=\\alpha\(1\-q\),so that the condition0<ρ<1−q0<\\rho<1\-qis enforced\.

#### 5\.1\.1Testing condition \(PCLI1\)

We first test the stronger conditiong​\(x,z\)≥1−β\.g\(x,z\)\\geq 1\-\\beta\.Since this inequality has already been established analytically in the tractable regimes, the numerical study is restricted to the remaining stripx1≤z≤x0,x^\{1\}\\leq z\\leq x^\{0\},wherex1x^\{1\}andx0x^\{0\}are the fixed points ofϕ1\\phi^\{1\}andϕ0\\phi^\{0\}, respectively\.

For each parameter tuple\(q,ρ,κ,β\)\(q,\\rho,\\kappa,\\beta\), we computed the minimum sampled slack

Δmin​\(q,ρ,κ,β\)≜minx,z⁡\(g​\(x,z\)−\(1−β\)\),\\Delta\_\{\\min\}\(q,\\rho,\\kappa,\\beta\)\\triangleq\\min\_\{x,z\}\\bigl\(g\(x,z\)\-\(1\-\\beta\)\\bigr\),over a finite grid withx∈\[0,1\]x\\in\[0,1\]andz∈\[x1,x0\]z\\in\[x^\{1\},x^\{0\}\]\. The outer parameter grid consisted of1414\-point uniform grids on

q∈\[0\.05,0\.95\],α∈\[0\.10,0\.90\],κ∈\[0\.05,0\.95\],q\\in\[0\.05,0\.95\],\\qquad\\alpha\\in\[0\.10,0\.90\],\\qquad\\kappa\\in\[0\.05,0\.95\],together with the discount\-factor gridβ∈\{0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.7,0\.8,0\.9,0\.95,0\.99\}\.\\beta\\in\\\{0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.7,0\.8,0\.9,0\.95,0\.99\\\}\.This yields14×14×14×11=30,18414\\times 14\\times 14\\times 11=30\{,\}184parameter tuples\.

For each tuple, we first computedx1x^\{1\}andx0x^\{0\}, and then sampledxxandzzusing endpoint\-enriched cosine grids of sizeNx=Nu=121,N\_\{x\}=N\_\{u\}=121,withz=x1\+u​\(x0−x1\),u∈\[0,1\]\.z=x^\{1\}\+u\(x^\{0\}\-x^\{1\}\),\\qquad u\\in\[0,1\]\.Thusg​\(x,z\)g\(x,z\)was evaluated at121×121=14,641121\\times 121=14\{,\}641sampled\(x,z\)\(x,z\)\-pairs per tuple, for a total of30,184×121×121=441,923,94430\{,\}184\\times 121\\times 121=441\{,\}923\{,\}944function evaluations\.

The computation was implemented in Julia and run on88threads; the full experiment completed in about1\.381\.38hours\. No sampled violation ofg​\(x,z\)≥1−βg\(x,z\)\\geq 1\-\\betawas found\. The smallest observed slack over all sampled points wasΔmin=2\.60209967284708×10−4,\\Delta\_\{\\min\}=2\.60209967284708\\times 10^\{\-4\},attained at\(q,ρ,κ,β\)=\(0\.05,0\.153462,0\.95,0\.1\),\(x⋆,z⋆\)=\(0\.00273905,0\.0504062\)\.\(q,\\rho,\\kappa,\\beta\)=\\bigl\(0\.05,\\;0\.153462,\\;0\.95,\\;0\.1\\bigr\),\\kern 5\.0pt\(x^\{\\star\},z^\{\\star\}\)=\\bigl\(0\.00273905,\\;0\.0504062\\bigr\)\.Hence, this experiment provides strong numerical evidence that the stronger inequalityg​\(x,z\)≥1−βg\(x,z\)\\geq 1\-\\betaholds throughout the tested parameter region, including the intermediate stripx1≤z≤x0x^\{1\}\\leq z\\leq x^\{0\}\.

#### 5\.1\.2Testing condition \(PCLI2\)

We next test condition \(PCLI2\), namely that the MP indexm​\(x\)≜f​\(x,x\)/g​\(x,x\)m\(x\)\\triangleq f\(x,x\)/g\(x,x\)is nondecreasing and continuous as a function ofxx\. We use the same outer parameter grid as in Subsection[5\.1\.1](https://arxiv.org/html/2606.11192#S5.SS1.SSS1), and therefore again cover14×14×14×11=30,18414\\times 14\\times 14\\times 11=30\{,\}184parameter tuples\.

For each tuple\(q,ρ,κ,β\)\(q,\\rho,\\kappa,\\beta\), we first computed the fixed pointsx1x^\{1\}andx0x^\{0\}\. We then sampledm​\(x\)m\(x\)on a uniform grid over the core interval\[x1,x0\]\[x^\{1\},x^\{0\}\], usingNmid=2001N\_\{\\mathrm\{mid\}\}=2001equally spaced points, so that both endpoints were included exactly\. To probe monotonicity and possible continuity effects atx1x^\{1\}andx0x^\{0\}, we enlarged the interval to

\[xL,xR\],xL=max⁡\{0,x1−η​\(x0−x1\)\},xR=min⁡\{1,x0\+η​\(x0−x1\)\},\[x\_\{L\},x\_\{R\}\],\\qquad x\_\{L\}=\\max\\\{0,\\;x^\{1\}\-\\eta\(x^\{0\}\-x^\{1\}\)\\\},\\qquad x\_\{R\}=\\min\\\{1,\\;x^\{0\}\+\\eta\(x^\{0\}\-x^\{1\}\)\\\},with padding fractionη=0\.05,\\eta=0\.05,and addedNside=201N\_\{\\mathrm\{side\}\}=201uniformly spaced points on each side, excluding duplicated endpoints\. Thus the nominal number of sampledxx\-values per tuple wasNmid\+2​Nside=2001\+2⋅201=2403,N\_\{\\mathrm\{mid\}\}\+2N\_\{\\mathrm\{side\}\}=2001\+2\\cdot 201=2403,for a total of30,184×2403=72,532,15230\{,\}184\\times 2403=72\{,\}532\{,\}152evaluations ofm​\(x\)m\(x\)\.

To assess monotonicity, we computed the forward differencesΔi≜m​\(xi\+1\)−m​\(xi\),\\Delta\_\{i\}\\triangleq m\(x\_\{i\+1\}\)\-m\(x\_\{i\}\),and recorded both the minimum value over the full padded interval\[xL,xR\]\[x\_\{L\},x\_\{R\}\]and the minimum value restricted to the core interval\[x1,x0\]\[x^\{1\},x^\{0\}\]\. To assess continuity near the endpoints, we also recorded the one\-step absolute differences immediately to the left and right ofx1x^\{1\}andx0x^\{0\}, and used their maximum as an endpoint continuity proxy\.

The computation was implemented in Julia and run on88threads; the full experiment completed in about5\.65\.6minutes\. No sampled violation of monotonicity was found\. Over the padded interval\[xL,xR\]\[x\_\{L\},x\_\{R\}\], the smallest forward difference was1\.350445×10−10,1\.350445\\times 10^\{\-10\},attained for\(q,ρ,κ,β\)=\(0\.95,0\.005,0\.05,0\.99\),\(q,\\rho,\\kappa,\\beta\)=\(0\.95,\\;0\.005,\\;0\.05,\\;0\.99\),at a point nearx≈0\.954774x\\approx 0\.954774\. Restricting attention to the core interval\[x1,x0\]\[x^\{1\},x^\{0\}\], the smallest forward difference was3\.11747×10−10,3\.11747\\times 10^\{\-10\},attained for\(q,ρ,κ,β\)=\(0\.95,0\.005,0\.05,0\.1\),\(q,\\rho,\\kappa,\\beta\)=\(0\.95,\\;0\.005,\\;0\.05,\\;0\.1\),nearx≈0\.954763x\\approx 0\.954763\. Thus all sampled forward differences were positive, both on the core interval and on the enlarged interval used to probe the endpoints\.

The largest observed endpoint continuity proxy was6\.846899×10−4,6\.846899\\times 10^\{\-4\},which is attained for\(q,ρ,κ,β\)=\(0\.188462,0\.730385,0\.95,0\.99\)\.\(q,\\rho,\\kappa,\\beta\)=\(0\.188462,\\;0\.730385,\\;0\.95,\\;0\.99\)\.For this case, the four one\-step absolute differences around the endpoints were

1\.18554×10−4,6\.84690×10−4,3\.94556×10−4,2\.55875×10−5\.1\.18554\\times 10^\{\-4\},\\qquad 6\.84690\\times 10^\{\-4\},\\qquad 3\.94556\\times 10^\{\-4\},\\qquad 2\.55875\\times 10^\{\-5\}\.These values are small, and no visible jump\-type behavior was detected at either endpoint in any sampled case\.

Therefore, while this experiment does not constitute a proof, it provides strong numerical evidence in support of condition \(PCLI2\): throughout the tested parameter region, the sampled MP indexm​\(x\)m\(x\)was nondecreasing, including at and around the boundary pointsx1x^\{1\}andx0x^\{0\}, and no visible discontinuities were observed\.

### 5\.2Policy benchmarking and dual bound comparisons

We next compare the MP index policy with three natural alternatives: the myopic policy, a round\-robin policy, and a random policy\. Here the myopic policy activates theMMprojects with the largest current one\-step expected rewardsrn​κn​xnr\_\{n\}\\kappa\_\{n\}x\_\{n\}; the round\-robin policy ignores beliefs and cycles deterministically through the project labels, activating the nextMMprojects in a fixed cyclic order at each period; and the random policy ignores beliefs and activatesMMprojects chosen uniformly at random at each period\. The purpose of these experiments is twofold\. First, we assess the practical value of the MP index policy in heterogeneous populations, where it need not coincide with the myopic rule\. Second, we compare achieved rewards with the normalized Lagrangian dual upper bound in order to assess how close the best simulated policy comes to that benchmark and to look for evidence of asymptotic optimality as the system size grows\.

We therefore focus on heterogeneous instances with two project types, denotedAAandBB, intended to play the same conceptual role as the “self\-healing” and “fragile” types used in the adherence\-model experiments\. TypeAAprojects are assigned more favorable latent dynamics, while typeBBprojects are assigned more fragile or persistent\-bad dynamics\. To obtain a reasonably rich but still interpretable grid, we considered44variants of typeAAcrossed with33variants of typeBB, yielding1212scenario families in total\. The fourAA\-type parameter sets were

\(p01,ρ,κ,r\)∈\{\(0\.01,0\.90,0\.70,1\),\(0\.03,0\.80,0\.70,1\),\(0\.05,0\.70,0\.70,1\),\(0\.02,0\.85,0\.55,1\)\},\(p\_\{01\},\\rho,\\kappa,r\)\\in\\\{\(0\.01,0\.90,0\.70,1\),\\ \(0\.03,0\.80,0\.70,1\),\\ \(0\.05,0\.70,0\.70,1\),\\ \(0\.02,0\.85,0\.55,1\)\\\},and the threeBB\-type parameter sets were

\(p01,ρ,κ,r\)∈\{\(0\.10,0\.10,0\.95,1\),\(0\.15,0\.05,0\.95,1\),\(0\.08,0\.20,0\.95,1\)\}\.\(p\_\{01\},\\rho,\\kappa,r\)\\in\\\{\(0\.10,0\.10,0\.95,1\),\\ \(0\.15,0\.05,0\.95,1\),\\ \(0\.08,0\.20,0\.95,1\)\\\}\.Thus each scenario consists of oneAA\-type and oneBB\-type parameter vector\.

Within each scenario, we varied the population composition over the nine type\-AAproportionspA∈\{0\.1,…,0\.9\}p\_\{A\}\\in\\\{0\.1,\\ldots,0\.9\\\},pB=1−pA,p\_\{B\}=1\-p\_\{A\},capacity ratiosαM∈\{0\.05,0\.10,0\.15,0\.20,0\.25,0\.30,0\.40,0\.50\},\\alpha\_\{M\}\\in\\\{0\.05,0\.10,0\.15,0\.20,0\.25,0\.30,0\.40,0\.50\\\},and system sizesN∈\{100,200,400,800,1600\}\.N\\in\\\{100,200,400,800,1600\\\}\.All projects were initialized at the common beliefx0=0\.5x\_\{0\}=0\.5, so that performance differences are driven by the latent dynamics and the evolving information state rather than by exogenous heterogeneity in the initial beliefs\. We chose all values ofNNas multiples of2020, so thatαM​N\\alpha\_\{M\}Nis an integer for every tested capacity ratio and the realized capacityMMmatches the target ratio exactly\. Altogether, the grid comprises12×9×8×5=432012\\times 9\\times 8\\times 5=4320problem instances\.

For each instance, we estimated the normalized discounted reward per project,

Jπ≈1−βN​𝔼​\[∑t=0T−1βt​Rπ​\(t\)\],J^\{\\pi\}\\approx\\frac\{1\-\\beta\}\{N\}\\,\\mathbb\{E\}\\\!\\left\[\\sum\_\{t=0\}^\{T\-1\}\\beta^\{t\}R^\{\\pi\}\(t\)\\right\],under each benchmark policyπ\\pi, usingT=300T=300periods,10001000independent Monte Carlo replications, discount factorβ=0\.99\\beta=0\.99, and95%95\\%confidence intervals\. The MP index policy used precomputed type\-specific index lookup tables on a fine belief grid, with linear interpolation between grid points\. The myopic policy ranked projects by the one\-period expected reward proxyrn​κn​xnr\_\{n\}\\kappa\_\{n\}x\_\{n\}\. The round\-robin and random rules served as low\-information baselines\.

For each realized instance\(N,M\)\(N,M\), we also computed the normalized Lagrangian dual upper bound\. A direct implementation of the bound routine proved to be the dominant computational bottleneck\. We therefore replaced it by a grouped convex solver that exploits the fact that only two project types are present\. Writing the dual objective as

Lvec​\(λ\)=M​λ1−β\+∑n=1NLn​\(xn,λ\),L\_\{\\mathrm\{vec\}\}\(\\lambda\)=\\frac\{M\\lambda\}\{1\-\\beta\}\+\\sum\_\{n=1\}^\{N\}L\_\{n\}\(x\_\{n\},\\lambda\),we evaluate the type\-specific thresholdzk⋆​\(λ\)z\_\{k\}^\{\\star\}\(\\lambda\)only once per typekk, rather than once per project, and minimize the resulting continuous convex objective by bisection on a subgradient,

s​\(λ\)=M1−β−∑n=1NGn⋆​\(xn,λ\)\.s\(\\lambda\)=\\frac\{M\}\{1\-\\beta\}\-\\sum\_\{n=1\}^\{N\}G\_\{n\}^\{\\star\}\(x\_\{n\},\\lambda\)\.On representative test instances this yielded the same minimizer and bound value as the previous implementation, while reducing runtime by more than an order of magnitude; the full43204320\-instance experiment then completed in a few minutes rather than hours\. This reduction in cost was essential for making the larger benchmark grid computationally feasible\. Additional implementation details, numerical tolerances, and reproducibility notes are collected in Appendix[E](https://arxiv.org/html/2606.11192#A5)\.

Across the full grid of43204320instances, the MP index policy achieved the largest estimated mean reward in41854185cases, or96\.9%96\.9\\%of the total\. The myopic policy was best in the remaining135135cases \(3\.1%3\.1\\%\), while round\-robin and random were never best\. Hence the MP index policy is the strongest benchmark overall, although, unlike in the homogeneous\-project setting, it is not uniformly dominant\.

The same conclusion holds under a more conservative confidence\-interval comparison\. Using the criterion that the lower endpoint of the MP index95%95\\%confidence interval exceed the upper endpoint of the competing policy’s interval, the MP index policy dominates round\-robin and random in100%100\\%of cases and dominates myopic in91\.8%91\.8\\%of cases\. Thus, even when it is not the best policy in point estimate, the MP index rule remains highly competitive, and clear inferiority to myopic is rare\.

We next compare performance with the normalized Lagrangian dual upper bound through the relative gap\(L¯−Jπ\)/L¯\.\(\\bar\{L\}\-J^\{\\pi\}\)/\\bar\{L\}\.Among the four policies tested, this gap was smallest for the MP index policy throughout the grid\. Over all43204320instances, the MP index relative gap ranged from5\.66%5\.66\\%to39\.02%39\.02\\%, with mean14\.46%14\.46\\%\. For the myopic policy, the corresponding range was6\.58%6\.58\\%to62\.93%62\.93\\%, with mean24\.20%24\.20\\%\. Round\-robin and random performed substantially worse, with relative gaps typically above40%40\\%and often much larger\. Thus the MP index rule is not only the strongest simulated policy overall; it is also the one that remains closest to the dual benchmark across the tested policy class\.

The smallest MP index relative gap occurred in the scenarioA\_harder\_obs\_B\_very\_fragile,\\texttt\{A\\\_harder\\\_obs\\\_B\\\_very\\\_fragile\},with type proportions\(0\.5,0\.5\)\(0\.5,0\.5\), capacity ratioαM=0\.5\\alpha\_\{M\}=0\.5, and\(N,M\)=\(200,100\)\(N,M\)=\(200,100\), where the relative gap was5\.66%5\.66\\%\. The largest MP index relative gap occurred inA\_harder\_obs\_B\_persistent\_bad,\\texttt\{A\\\_harder\\\_obs\\\_B\\\_persistent\\\_bad\},with proportions\(0\.1,0\.9\)\(0\.1,0\.9\),αM=0\.05\\alpha\_\{M\}=0\.05, and\(N,M\)=\(100,5\)\(N,M\)=\(100,5\), where the relative gap was39\.02%39\.02\\%\. These extremes already suggest two clear patterns: low\-capacity regimes are harder, and scenarios involving the persistent\-badBB\-type tend to produce larger gaps\.

To measure the practical value of the MP index rule relative to the strongest simpler benchmark, we also computed the relative improvement over myopic,\(JMP−Jmyopic\)/\(Jmyopic\)\.\(J^\{\\mathrm\{MP\}\}\-J^\{\\mathrm\{myopic\}\}\)/\(J^\{\\mathrm\{myopic\}\}\)\.Over the full grid, this quantity ranged from−0\.91%\-0\.91\\%to69\.18%69\.18\\%, with mean approximately15\.5%15\.5\\%\. Thus the MP index policy often provides a substantial gain over myopic, and even in the rare cases where it underperforms, the loss is very small\.

The largest relative gain of the MP index policy over myopic occurred in the scenarioA\_mild\_selfhealing\_B\_fragile,\\texttt\{A\\\_mild\\\_selfhealing\\\_B\\\_fragile\},with proportions\(0\.3,0\.7\)\(0\.3,0\.7\),αM=0\.05\\alpha\_\{M\}=0\.05, and\(N,M\)=\(1600,80\)\(N,M\)=\(1600,80\), where the gain reached69\.18%69\.18\\%\. The smallest relative improvement occurred inA\_mild\_selfhealing\_B\_persistent\_bad,\\texttt\{A\\\_mild\\\_selfhealing\\\_B\\\_persistent\\\_bad\},with proportions\(0\.7,0\.3\)\(0\.7,0\.3\),αM=0\.05\\alpha\_\{M\}=0\.05, and\(N,M\)=\(100,5\)\(N,M\)=\(100,5\), where the MP index policy fell below myopic by only0\.91%0\.91\\%\. Hence the grid does contain a small number of cases where myopic slightly outperforms MP, but the overall comparison strongly favors the MP index rule\.

Among all design factors, the capacity ratioαM=M/N\\alpha\_\{M\}=M/Nhas the clearest effect on the MP index gap to the dual bound\. Averaging over all scenarios, proportions, and system sizes, the mean MP index relative gap decreases monotonically from about30\.4%30\.4\\%atαM=0\.05\\alpha\_\{M\}=0\.05to about7\.0%7\.0\\%atαM=0\.50\\alpha\_\{M\}=0\.50\. Thus the MP index policy comes much closer to the dual benchmark when capacity is less scarce\.

The dependence on type composition is present but weaker\. Averaged over scenarios, capacities, and system sizes, the mean MP index relative gap is smallest for intermediate mixtures, around\(0\.4,0\.6\)\(0\.4,0\.6\)and\(0\.5,0\.5\)\(0\.5,0\.5\), and larger at the extremes, especially forAA\-heavy populations\. This suggests that strongly unbalanced mixtures are somewhat more difficult, especially when the more persistent type dominates\.

By contrast, there is no convincing evidence here that the MP index relative gap vanishes asNNgrows\. Averaging over scenarios, proportions, and capacities, the mean relative gap is essentially flat:

14\.53%,14\.46%,14\.44%,14\.43%,14\.42%14\.53\\%,\\ 14\.46\\%,\\ 14\.44\\%,\\ 14\.43\\%,\\ 14\.42\\%forN=100,200,400,800,1600N=100,200,400,800,1600, respectively\. Thus, over the tested range of system sizes, the experiments do not support asymptotic optimality of the MP index policy relative to the Lagrangian bound\. What they do support is the more modest but practically relevant conclusion that the MP index policy remains closer to the dual benchmark than the simpler alternatives\.

Overall, the numerical evidence points to three main conclusions\. First, the MP index policy is the strongest benchmark across a large and systematically constructed grid of heterogeneous two\-type instances\. Second, its practical advantage over myopic can be substantial, often dramatic, especially in low\-capacity and strongly heterogeneous regimes\. Third, although the Lagrangian dual bound remains a useful common upper benchmark, it is not sufficiently tight here to certify asymptotic optimality of the MP index policy: the relative MP index gap remains non\-negligible and essentially flat inNNover the tested range\.

Accordingly, the experiments support the MP index policy as a highly effective practical benchmark in heterogeneous populations, but they do not justify any claim of asymptotic optimality with respect to the Lagrangian relaxation\. In particular, the low\-capacity regime with persistent\-badBB\-type projects emerges as the most challenging part of the design space, whereas higher\-capacity regimes with more fragileBB\-types appear substantially easier\.

## 6Conclusions

We have studied a class of belief\-state restless bandits with imperfect binary observations, motivated in particular by opportunistic spectrum access with sensing errors and collision constraints\. The main contribution of the paper is a PCL\-based analytical and computational framework for evaluating threshold\-policy performance metrics and the associated marginal productivity \(MP\) index in this one\-sided ACK/NACK setting\.

At the technical level, the key step is the decomposition of threshold\-policy dynamics into deterministic no\-ACK skeleton evolution and ACK\-triggered regenerative restarts\. This leads to explicit renewal\-type representations for the reward and work metricsFFandGG, their marginal counterpartsffandgg, and hence the MP index\. In the tractable threshold regimes we obtained closed\-form expressions for these quantities, including exact formulas for the diagonal MP index on\[0,x1\]\[0,x^\{1\}\],\[x0,p11\]\[x^\{0\},p\_\{11\}\], and\[p11,1\]\[p\_\{11\},1\]\. In the intermediate regimex1<z<x0x^\{1\}<z<x^\{0\}, where the no\-ACK skeleton alternates indefinitely between active and passive phases, the same renewal structure yields stable numerical evaluation schemes and isolates the remaining analytical gaps\.

A second contribution is the clarification of the symbolic organization of threshold itineraries in the intermediate regime\. Using the fact that the belief\-update maps satisfy the analogue of the maps\-with\-gaps regularity conditions ofDance and Silander\[[11](https://arxiv.org/html/2606.11192#bib.bib11)\]after an increasing change of variables, we showed that the threshold axis is partitioned into Christoffel intervals and Sturmian points, and that these determine the corresponding one\-step\-deviation itineraries\. This identifies a concrete route toward a complete analytical treatment of the unresolved part of the PCL verification\.

The computational experiments provide broad numerical evidence for the PCL\-indexability conditions beyond the parameter restrictions currently available in the literature\. In particular, the large\-scale tests of \(PCLI1\) and \(PCLI2\) revealed no counterexamples on extensive parameter grids, including the intermediate threshold stripx1≤z≤x0x^\{1\}\\leq z\\leq x^\{0\}\. The experiments also showed that the MP index behaves continuously and monotonically in the tractable regimes and exhibits highly structured staircase behavior in the intermediate regime, consistent with the Christoffel–Sturmian organization of threshold itineraries\.

From the policy\-performance viewpoint, the MP\-index policy emerged as the strongest benchmark across a large heterogeneous two\-type test bed\. It consistently outperformed round\-robin and random baselines, and in the vast majority of instances also outperformed the myopic policy\. Moreover, it generally remained closer than the benchmark policies to the normalized Lagrangian dual upper bound\. At the same time, the experiments did*not*provide convincing evidence that the MP\-index policy becomes asymptotically optimal relative to that dual bound as the number of projects grows\. Thus the numerical results support the MP\-index policy as a highly effective practical benchmark, but not as an empirically established asymptotically optimal rule\.

Taken together, the tractable\-regime analysis, the reduction of \(PCLI3\) to \(PCLI1\) and \(PCLI2\), the Christoffel–Sturmian organization of the intermediate regime, and the extensive numerical evidence lead us to the following:

###### Conjecture 6\.1\(Global PCL\-indexability\)\.

For every feasible parameter tuple satisfying

0<p01<1,0<ρ<1−p01,0<κ<1,0<β<1,0<p\_\{01\}<1,\\qquad 0<\\rho<1\-p\_\{01\},\\qquad 0<\\kappa<1,\\qquad 0<\\beta<1,the discounted single\-project problem is PCL\-indexable, that is,\(PCLI1\)–\(PCLI3\)hold on the full threshold range\.

If true, this would imply that the model is Whittle\-indexable throughout that parameter domain, and that the MP index coincides globally with the Whittle index\.

Several directions remain open\. The most immediate is to complete the analytical verification of \(PCLI1\) and \(PCLI2\) on the intermediate regimex1<z<x0x^\{1\}<z<x^\{0\}, thereby resolving the conjecture above and obtaining a full proof of Whittle indexability for the discounted model\. A second direction is to develop the average\-reward counterpart of the present discounted PCL analysis, building on the outline given in Appendix[D](https://arxiv.org/html/2606.11192#A4)\. More broadly, the computational framework developed here should also be useful for other partially observed restless bandit models with real\-valued belief states\.

## 7Declaration of generative AI and AI\-assisted technologies in the manuscript preparation process

During the preparation of this work, the author used ChatGPT \(OpenAI\) in order to assist with editing text, improving readability, and refining the presentation of the manuscript\. After using this tool, the author reviewed and edited the content as needed and takes full responsibility for the content of the manuscript\.

## References

- Aalto et al\. \[2019\]S\. Aalto, P\. Lassila, and I\. Taboada\.Whittle index approach to opportunistic scheduling with partial channel information\.*Performance Eval\.*, 136:102052, 2019\.
- Aguilera et al\. \[1999\]M\. K\. Aguilera, W\. Chen, and S\. Toueg\.Using the heartbeat failure detector for quiescent reliable communication and consensus in partitionable networks\.*Theoret\. Comput\. Sci\.*, 220:3–30, 1999\.
- Ahlfors \[1979\]L\. V\. Ahlfors\.*Complex Analysis: An Introduction to the Theory of Analytic Functions of One Complex Variable*\.McGraw\-Hill, 3 edition, 1979\.
- Ahmad et al\. \[2009\]S\. H\. A\. Ahmad, M\. Y\. Liu, T\. Javidi, and Q\. Zhao\.Optimality of myopic sensing in multichannel opportunistic access\.*IEEE Trans\. Inform\. Theory*, 55:4040–4050, 2009\.
- Akbarzadeh and Mahajan \[2024\]N\. Akbarzadeh and A\. Mahajan\.Two families of indexable partially observable restless bandits and Whittle index computation\.*Performance Eval\.*, 163:102394, 2024\.
- Beardon \[2000\]Alan F\. Beardon\.*Iteration of Rational Functions: Complex Analytic Dynamical Systems*\.Springer, 2000\.
- Berstel et al\. \[2009\]J\. Berstel, A\. Lauve, C\. Reutenauer, and F\. V\. Saliola\.*Combinatorics on Words: Christoffel Words and Repetitions in Words*, volume 27 of*CRM Monograph Series*\.AMS, Providence, RI, 2009\.
- Carter and van Brunt \[2000\]M\. Carter and B\. van Brunt\.*The Lebesgue–Stieltjes Integral: A Practical Introduction*\.Springer, New York, NY, USA, 2000\.
- Chandra and Toueg \[1996\]T\. D\. Chandra and S\. Toueg\.Unreliable failure detectors for reliable distributed systems\.*J\. ACM*, 43:225–267, 1996\.
- Chen et al\. \[2008\]Y\. Chen, Q\. Zhao, and A\. Swami\.Joint design and separation principle for opportunistic spectrum access in the presence of sensing errors\.*IEEE Trans\. Inform\. Theory*, 54:2053–2071, 2008\.
- Dance and Silander \[2019\]C\. R\. Dance and T\. Silander\.Optimal policies for observing time series and related restless bandit problems\.*J\. Mach\. Learn\. Res\.*, 20:35, 2019\.
- Gittins \[1979\]J\. C\. Gittins\.Bandit processes and dynamic allocation indices \(with discussion\)\.*J\. Roy\. Statist\. Soc\. Ser\. B*, 41:148–177, 1979\.
- Hoang et al\. \[2009\]A\. T\. Hoang, Y\. \-C\. Liang, D\. T\. C\. Wong, Y\. Zeng, and R\. Zhang\.Opportunistic spectrum access for energy\-constrained cognitive radios\.*IEEE Trans\. Wireless Commun\.*, 8:1206–1211, 2009\.
- Johnston and Krishnamurthy \[2006\]L\. A\. Johnston and V\. Krishnamurthy\.Opportunistic file transfer over a fading channel: a POMDP search theory formulation with optimal threshold policies\.*IEEE Trans\. Wireless Commun\.*, 5:394–405, 2006\.
- Kaza et al\. \[2019\]K\. Kaza, R\. Meshram, V\. Mehta, and S\. N\. Merchant\.Sequential decision making with limited observation capability: application to wireless networks\.*IEEE Trans\. Cognitive Commun\. Netw\.*, 5:237–251, 2019\.
- Krishnamurthy \[2025\]V\. Krishnamurthy\.*Partially Observed Markov Decision Processes: Filtering, Learning and Controlled Sensing*\.Cambridge University Press, Cambridge, UK, 2nd edition, 2025\.
- Liu and Zhao \[2010\]K\. Q\. Liu and Q\. Zhao\.Indexability of restless bandit problems and optimality of Whittle index for dynamic multichannel access\.*IEEE Trans\. Inform\. Theory*, 56:5547–5567, 2010\.
- Liu et al\. \[2010\]K\. Q\. Liu, Q\. Zhao, and B\. Krishnamachari\.Dynamic multichannel access with imperfect channel state detection\.*IEEE Trans\. Signal Process\.*, 58:2795–2808, 2010\.
- Liu et al\. \[2024\]K\. Q\. Liu, R\. Weber, and C\. Z\. Zhang\.Low\-complexity algorithm for restless bandits with imperfect observations\.*Math\. Meth\. Oper\. Res\.*, 100:467–508, 2024\.
- Mate et al\. \[2020\]A\. Mate, J\. A\. Killian, H\. Xu, A\. Perrault, and M\. Tambe\.Collapsing bandits and their application to public health interventions\.In*Advances in Neural Information Processing Systems 33 \(NeurIPS 2020\)*, volume 33, pages 15639–15650, 2020\.
- Mate et al\. \[2021\]A\. Mate, A\. Perrault, and M\. Tambe\.Risk\-aware interventions in public health: Planning with restless multi\-armed bandits\.In*Proceedings of AAMAS 2021*, pages 880–888, 2021\.
- Mehta et al\. \[2018\]V\. Mehta, R\. Meshram, K\. Kaza, S\. N\. Merchant, and U\. B\. Desai\.Rested and restless bandits with constrained arms and hidden states: Applications in social networks and 5G networks\.*IEEE Access*, 6:56782–56799, 2018\.
- Meshram et al\. \[2018\]R\. Meshram, D\. Manjunath, and A\. Gopalan\.On the Whittle index for restless multiarmed hidden Markov bandits\.*IEEE Trans\. Automat\. Control*, 63:3046–3053, 2018\.
- Murugesan et al\. \[2012\]S\. Murugesan, P\. Schniter, and N\. B\. Shroff\.Multiuser scheduling in a Markov\-modeled downlink using randomly delayed ARQ feedback\.*IEEE Trans\. Inform\. Theory*, 58:1025–1042, 2012\.
- Niño\-Mora \[2001\]J\. Niño\-Mora\.Restless bandits, partial conservation laws and indexability\.*Adv\. Appl\. Probab\.*, 33:76–98, 2001\.
- Niño\-Mora \[2002\]J\. Niño\-Mora\.Dynamic allocation indices for restless projects and queueing admission control: A polyhedral approach\.*Math\. Program\.*, 93:361–413, 2002\.
- Niño\-Mora \[2006\]J\. Niño\-Mora\.Restless bandit marginal productivity indices, diminishing returns and optimal control of make\-to\-order/make\-to\-stock M/G/1 queues\.*Math\. Oper\. Res\.*, 31:50–84, 2006\.
- Niño\-Mora \[2007\]J\. Niño\-Mora\.Dynamic priority allocation via restless bandit marginal productivity indices \(with discussion\)\.*TOP*, 15:161–198, 2007\.
- Niño\-Mora \[2008\]J\. Niño\-Mora\.An index policy for dynamic fading\-channel allocation to heterogeneous mobile users with partial observations\.In*Proceedings of the 4th EuroNGI Conference on Next Generation Internet Networks, Krakow, Poland*, pages 231–238, New York, NY, USA, 2008\. IEEE\.
- Niño\-Mora \[2009\]J\. Niño\-Mora\.A restless bandit marginal productivity index for opportunistic spectrum access with sensing errors\.In*Proceedings of the 3rd Euro\-NF Conference on Network Control and Optimization \(NETCOOP\), Eindhoven, The Netherlands*, volume 5894 of*Lecture Notes in Computer Science*, pages 60–74, Berlin, Germany, 2009\. Springer\.
- Niño\-Mora \[2020\]J\. Niño\-Mora\.A verification theorem for threshold\-indexability of real\-state discounted restless bandits\.*Math\. Oper\. Res\.*, 45:465–496, 2020\.
- Niño\-Mora \[2023\]J\. Niño\-Mora\.Markovian restless bandits and index policies: A review\.*Mathematics*, 11:1639, 2023\.
- Niño\-Mora and Pellitero García \[2026\]J\. Niño\-Mora and Á\. Pellitero García\.A belief\-state restless bandit model for treatment adherence: Whittle indexability via partial conservation laws, 2026\.arXiv:2601\.06976\.
- Wang and Chen \[2021\]K\. H\. Wang and L\. Chen\.*Restless Multi\-Armed Bandit in Opportunistic Scheduling*\.Springer, Cham, Switzerland, 2021\.
- Wang et al\. \[2018\]K\. H\. Wang, L\. Chen, J\. H\. Yu, and M\. Win\.Opportunistic multichannel access with imperfect observation: A fixed point analysis on indexability and index\-based policy\.In*Proceedings of IEEE INFOCOM*, New York, NY, 2018\. IEEE\.
- Whittle \[1988\]P\. Whittle\.Restless bandits: Activity allocation in a changing world\.*J\. Appl\. Probab\.*, 25A:287–298, 1988\.
- Zhao et al\. \[2007\]Q\. Zhao, L\. Tong, A\. Swami, and Y\. Chen\.Decentralized cognitive MAC for opportunistic spectrum access in ad hoc networks: a POMDP framework\.*IEEE J\. Sel\. Areas Commun\.*, 25:589–600, 2007\.

## Appendix ASupplementary proofs for the computational machinery

### A\.1Proof details for Section[4\.1](https://arxiv.org/html/2606.11192#S4.SS1)

We collect here the proofs omitted from Section[4\.1](https://arxiv.org/html/2606.11192#S4.SS1)\.

###### Proof of Lemma[4\.2](https://arxiv.org/html/2606.11192#S4.Thmtheorem2)\.

Letℱt\\mathscr\{F\}\_\{t\}denote the history up to timett\. SinceA​\(t\)A\(t\)isℱt\\mathscr\{F\}\_\{t\}\-measurable and \([20](https://arxiv.org/html/2606.11192#S4.E20)\) implies that𝔼​\[X​\(t\+1\)∣X​\(t\),A​\(t\)\]=ϕ0​\(X​\(t\)\)\\mathbb\{E\}\[X\(t\+1\)\\mid X\(t\),A\(t\)\]=\\phi^\{0\}\(X\(t\)\)a\.s\., we have

𝔼​\[X​\(t\+1\)∣ℱt\]=𝔼​\[𝔼​\[X​\(t\+1\)∣X​\(t\),A​\(t\)\]∣ℱt\]=𝔼​\[ϕ0​\(X​\(t\)\)∣ℱt\]=ϕ0​\(X​\(t\)\)\.\\mathbb\{E\}\\\!\\left\[X\(t\+1\)\\mid\\mathscr\{F\}\_\{t\}\\right\]=\\mathbb\{E\}\\\!\\left\[\\mathbb\{E\}\[X\(t\+1\)\\mid X\(t\),A\(t\)\]\\mid\\mathscr\{F\}\_\{t\}\\right\]=\\mathbb\{E\}\\\!\\left\[\\phi^\{0\}\(X\(t\)\)\\mid\\mathscr\{F\}\_\{t\}\\right\]=\\phi^\{0\}\(X\(t\)\)\.Taking expectations and using the affine form ofϕ0\\phi^\{0\}yields the recursion

𝔼xπ​\[X​\(t\+1\)\]=p01\+ρ​𝔼xπ​\[X​\(t\)\]\.\\mathbb\{E\}\_\{x\}^\{\\pi\}\[X\(t\+1\)\]=p\_\{01\}\+\\rho\\,\\mathbb\{E\}\_\{x\}^\{\\pi\}\[X\(t\)\]\.Iterating it, with𝔼xπ​\[X​\(0\)\]=x\\mathbb\{E\}\_\{x\}^\{\\pi\}\[X\(0\)\]=x, gives𝔼xπ​\[X​\(t\)\]=ϕt0​\(x\)\\mathbb\{E\}\_\{x\}^\{\\pi\}\[X\(t\)\]=\\phi\_\{t\}^\{0\}\(x\)for allt≥0t\\geq 0\. ∎

###### Proof of Lemma[4\.4](https://arxiv.org/html/2606.11192#S4.Thmtheorem4)\.

LetB≜1−ρ\+κ​p11B\\triangleq 1\-\\rho\+\\kappa p\_\{11\}\. The discriminant of \([24](https://arxiv.org/html/2606.11192#S4.E24)\) isΔ​\(κ\)=B2−4​κ​p01\\Delta\(\\kappa\)=B^\{2\}\-4\\kappa p\_\{01\}, which can be rewritten as in \([26](https://arxiv.org/html/2606.11192#S4.E26)\) using1−ρ=p10\+p011\-\\rho=p\_\{10\}\+p\_\{01\}andp11=1−p10p\_\{11\}=1\-p\_\{10\}\. HenceΔ​\(κ\)\>0\\Delta\(\\kappa\)\>0, and \([24](https://arxiv.org/html/2606.11192#S4.E24)\) has the roots in \([25](https://arxiv.org/html/2606.11192#S4.E25)\)\.

To locate the roots, writeH​\(x\)≜ϕ1​\(x\)−xH\(x\)\\triangleq\\phi^\{1\}\(x\)\-xforx∈\[0,1/κ\)x\\in\[0,1/\\kappa\)\. Then

H​\(x\)=P​\(x\)1−κ​x,P​\(x\)≜κ​x2−B​x\+p01\.H\(x\)=\\frac\{P\(x\)\}\{1\-\\kappa x\},\\qquad P\(x\)\\triangleq\\kappa x^\{2\}\-Bx\+p\_\{01\}\.Since1−κ​x\>01\-\\kappa x\>0on\[0,1\]\[0,1\], the sign ofH​\(x\)H\(x\)on𝒳\\mathscr\{X\}is the sign of the quadratic numeratorP​\(x\)P\(x\)\. We have

P​\(p01\)=p01​ρ​\(1−κ\)\>0,P​\(p11\)=p01−\(1−ρ\)​p11=−p10​ρ<0,P\(p\_\{01\}\)=p\_\{01\}\\rho\(1\-\\kappa\)\>0,\\qquad P\(p\_\{11\}\)=p\_\{01\}\-\(1\-\\rho\)p\_\{11\}=\-p\_\{10\}\\rho<0,so by continuityPP\(and henceHH\) has a root in\(p01,p11\)\(p\_\{01\},p\_\{11\}\)\. This is the smaller rootxlo1x\_\{\\mathrm\{lo\}\}^\{1\}, and we setx1≜xlo1∈\(p01,p11\)x^\{1\}\\triangleq x\_\{\\mathrm\{lo\}\}^\{1\}\\in\(p\_\{01\},p\_\{11\}\)\.

Next,P​\(1\)=κ−B\+p01=−p10​\(1−κ\)<0,P\(1\)=\\kappa\-B\+p\_\{01\}=\-p\_\{10\}\(1\-\\kappa\)<0,whileP​\(x\)→\+∞P\(x\)\\to\+\\inftyasx→\+∞x\\to\+\\inftysinceκ\>0\\kappa\>0, so the larger root satisfiesxhi1\>1x\_\{\\mathrm\{hi\}\}^\{1\}\>1\. Finally, Vieta’s identityxlo1​xhi1=p01/κx\_\{\\mathrm\{lo\}\}^\{1\}x\_\{\\mathrm\{hi\}\}^\{1\}=\{p\_\{01\}\}/\{\\kappa\}together withxlo1\>p01x\_\{\\mathrm\{lo\}\}^\{1\}\>p\_\{01\}givesxhi1<1/κx\_\{\\mathrm\{hi\}\}^\{1\}<1/\\kappa\. Thusxhi1∈\(1,1/κ\)x\_\{\\mathrm\{hi\}\}^\{1\}\\in\(1,1/\\kappa\)\.

Sinceϕ1​\(𝒳\)⊆\[p01,p11\]⊆𝒳\\phi^\{1\}\(\\mathscr\{X\}\)\\subseteq\[p\_\{01\},p\_\{11\}\]\\subseteq\\mathscr\{X\}, any fixed point ofϕ1\\phi^\{1\}in𝒳\\mathscr\{X\}must lie in\[p01,p11\]\[p\_\{01\},p\_\{11\}\], hence it must bex1x^\{1\}\. ∎

###### Proof of Lemma[4\.5](https://arxiv.org/html/2606.11192#S4.Thmtheorem5)\.

Setb≜1−ρ\>0b\\triangleq 1\-\\rho\>0andy≜p11∈\(0,1\)y\\triangleq p\_\{11\}\\in\(0,1\)\. Fromϕ1​\(x\)=x\\phi^\{1\}\(x\)=xwe obtain

κ​x2−\(b\+κ​y\)​x\+p01=0,\\kappa x^\{2\}\-\\bigl\(b\+\\kappa y\\bigr\)x\+p\_\{01\}=0,hencep01=x​\(b\+κ​y−κ​x\)p\_\{01\}=x\\bigl\(b\+\\kappa y\-\\kappa x\\bigr\)\. Dividing bybband usingx0=p01/bx^\{0\}=p\_\{01\}/byields

x0−x=p01b−x=κ​x​\(y−x\)b\.x^\{0\}\-x=\\frac\{p\_\{01\}\}\{b\}\-x=\\frac\{\\kappa x\(y\-x\)\}\{b\}\.Takingx=x1x=x^\{1\}and usingx1∈\(p01,y\)x^\{1\}\\in\(p\_\{01\},y\)from Lemma[4\.4](https://arxiv.org/html/2606.11192#S4.Thmtheorem4), we getx1\>0x^\{1\}\>0andy−x1\>0y\-x^\{1\}\>0, hencex0−x1\>0x^\{0\}\-x^\{1\}\>0\. ∎

###### Proof of Lemma[4\.6](https://arxiv.org/html/2606.11192#S4.Thmtheorem6)\.

For part \(a\), from \([29](https://arxiv.org/html/2606.11192#S4.E29)\) we have

\(ϕ1\)′​\(x\)=ρ​\(1−κ\)\(1−κ​x\)2\>0,x∈𝒳,\(\\phi^\{1\}\)^\{\\prime\}\(x\)=\\frac\{\\rho\(1\-\\kappa\)\}\{\(1\-\\kappa x\)^\{2\}\}\>0,\\qquad x\\in\\mathscr\{X\},soϕ1\\phi^\{1\}is increasing; compositions preserve strict monotonicity\.

For part \(b\), writingϕ1​\(x\)=\(A​x\+p01\)/\(1−κ​x\)\\phi^\{1\}\(x\)=\(Ax\+p\_\{01\}\)/\(1\-\\kappa x\)withA=ρ​\(1−κ\)−κ​p01A=\\rho\(1\-\\kappa\)\-\\kappa p\_\{01\}gives

ϕ1​\(x\)−x=κ​x2−\(1−ρ\+κ​p11\)​x\+p011−κ​x\.\\phi^\{1\}\(x\)\-x=\\frac\{\\kappa x^\{2\}\-\\bigl\(1\-\\rho\+\\kappa p\_\{11\}\\bigr\)x\+p\_\{01\}\}\{1\-\\kappa x\}\.The numerator has rootsx1x^\{1\}andxhi1x\_\{\\mathrm\{hi\}\}^\{1\}, hence equalsκ​\(x−x1\)​\(x−xhi1\)\\kappa\(x\-x^\{1\}\)\(x\-x\_\{\\mathrm\{hi\}\}^\{1\}\), yielding \([34](https://arxiv.org/html/2606.11192#S4.E34)\)\. Subtractingx1=ϕ1​\(x1\)x^\{1\}=\\phi^\{1\}\(x^\{1\}\)fromϕ1​\(x\)\\phi^\{1\}\(x\)gives \([35](https://arxiv.org/html/2606.11192#S4.E35)\)\. The stated inequalities and monotonicity inttfollow\.

For part \(c\), apply \([31](https://arxiv.org/html/2606.11192#S4.E31)\)–\([32](https://arxiv.org/html/2606.11192#S4.E32)\) toΦ=ϕ1\\Phi=\\phi^\{1\}withx−=x1x\_\{\-\}=x^\{1\},x\+=xhi1x\_\{\+\}=x\_\{\\mathrm\{hi\}\}^\{1\}, andΨ\\Psias in \([33](https://arxiv.org/html/2606.11192#S4.E33)\), obtaining the first equality in \([37](https://arxiv.org/html/2606.11192#S4.E37)\)\. The second equality follows by algebraic simplification\. Sinceμ∈\(0,1\)\\mu\\in\(0,1\), convergence tox1x^\{1\}is geometric\. ∎

###### Proof of Lemma[4\.7](https://arxiv.org/html/2606.11192#S4.Thmtheorem7)\.

Fixa∈\{0,1\}a\\in\\\{0,1\\\}and writet=ϑ​\(x\)t=\\vartheta\(x\)withx∈\(0,1\)x\\in\(0,1\)\. Sinceϑ′​\(x\)=1/\(x​\(1−x\)\)\\vartheta^\{\\prime\}\(x\)=1/\(x\(1\-x\)\), we have

\(ϕ^a\)′​\(t\)=ϑ′​\(ϕa​\(x\)\)ϑ′​\(x\)​\(ϕa\)′​\(x\)=\(ϕa\)′​\(x\)​x​\(1−x\)ϕa​\(x\)​\(1−ϕa​\(x\)\)\.\(\\hat\{\\phi\}^\{a\}\)^\{\\prime\}\(t\)=\\frac\{\\vartheta^\{\\prime\}\(\\phi^\{a\}\(x\)\)\}\{\\vartheta^\{\\prime\}\(x\)\}\\,\(\\phi^\{a\}\)^\{\\prime\}\(x\)=\\frac\{\(\\phi^\{a\}\)^\{\\prime\}\(x\)\\,x\(1\-x\)\}\{\\phi^\{a\}\(x\)\\bigl\(1\-\\phi^\{a\}\(x\)\\bigr\)\}\.\(74\)In particular,\(ϕ^a\)′​\(t\)\>0\(\\hat\{\\phi\}^\{a\}\)^\{\\prime\}\(t\)\>0for alltt, soϕ^a\\hat\{\\phi\}^\{a\}is increasing\. Moreover, if\(ϕ^a\)′​\(t\)<1\(\\hat\{\\phi\}^\{a\}\)^\{\\prime\}\(t\)<1for alltt, then by the mean value theoremϕ^a\\hat\{\\phi\}^\{a\}is contractive onℝ\\mathbb\{R\}in the sense of Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)\(iii\)\.

Fora=0a=0, we haveϕ0​\(x\)=p01\+ρ​x\\phi^\{0\}\(x\)=p\_\{01\}\+\\rho xand\(ϕ0\)′​\(x\)=ρ\(\\phi^\{0\}\)^\{\\prime\}\(x\)=\\rho, so by \([74](https://arxiv.org/html/2606.11192#A1.E74)\) it suffices to show

ϕ0​\(x\)​\(1−ϕ0​\(x\)\)−ρ​x​\(1−x\)\>0,x∈\(0,1\)\.\\phi^\{0\}\(x\)\\bigl\(1\-\\phi^\{0\}\(x\)\\bigr\)\-\\rho x\(1\-x\)\>0,\\qquad x\\in\(0,1\)\.A direct expansion yields

ϕ0​\(x\)​\(1−ϕ0​\(x\)\)−ρ​x​\(1−x\)=p01​\(1−p01\)−2​p01​ρ​x\+ρ​\(1−ρ\)​x2\.\\phi^\{0\}\(x\)\\bigl\(1\-\\phi^\{0\}\(x\)\\bigr\)\-\\rho x\(1\-x\)=p\_\{01\}\(1\-p\_\{01\}\)\-2p\_\{01\}\\rho\\,x\+\\rho\(1\-\\rho\)\\,x^\{2\}\.The discriminant of this quadratic equals4​p01​ρ​\(p11−1\)=−4​ρ​p01​p10<0,4p\_\{01\}\\rho\\,\(p\_\{11\}\-1\)=\-4\\rho p\_\{01\}p\_\{10\}<0,so the quadratic is positive onℝ\\mathbb\{R\}, and therefore\(ϕ^0\)′​\(t\)∈\(0,1\)\(\\hat\{\\phi\}^\{0\}\)^\{\\prime\}\(t\)\\in\(0,1\)for alltt\.

Fora=1a=1, we have

ϕ1​\(x\)=p01\+ρ​\(1−κ\)​x1−κ​x,\(ϕ1\)′​\(x\)=ρ​\(1−κ\)\(1−κ​x\)2\.\\phi^\{1\}\(x\)=p\_\{01\}\+\\rho\\,\\frac\{\(1\-\\kappa\)x\}\{1\-\\kappa x\},\\qquad\(\\phi^\{1\}\)^\{\\prime\}\(x\)=\\frac\{\\rho\(1\-\\kappa\)\}\{\(1\-\\kappa x\)^\{2\}\}\.By \([74](https://arxiv.org/html/2606.11192#A1.E74)\) it suffices to show

ϕ1​\(x\)​\(1−ϕ1​\(x\)\)−\(ϕ1\)′​\(x\)​x​\(1−x\)\>0,x∈\(0,1\)\.\\phi^\{1\}\(x\)\\bigl\(1\-\\phi^\{1\}\(x\)\\bigr\)\-\(\\phi^\{1\}\)^\{\\prime\}\(x\)\\,x\(1\-x\)\>0,\\qquad x\\in\(0,1\)\.A straightforward calculation shows that the left\-hand side equalsQ​\(x\)/\(1−κ​x\)2Q\(x\)/\(1\-\\kappa x\)^\{2\}, whereQQis a quadratic satisfyingQ​\(0\)=p01​\(1−p01\)\>0Q\(0\)=p\_\{01\}\(1\-p\_\{01\}\)\>0and whose discriminant is4​p01​ρ​\(κ−1\)2​\(p11−1\)=−4​ρ​\(1−κ\)2​p01​p10<0\.4p\_\{01\}\\rho\(\\kappa\-1\)^\{2\}\(p\_\{11\}\-1\)=\-4\\rho\(1\-\\kappa\)^\{2\}p\_\{01\}p\_\{10\}<0\.HenceQ​\(x\)\>0Q\(x\)\>0for allx∈ℝx\\in\\mathbb\{R\}, and therefore\(ϕ^1\)′​\(t\)∈\(0,1\)\(\\hat\{\\phi\}^\{1\}\)^\{\\prime\}\(t\)\\in\(0,1\)for alltt\.

Thusϕ^0\\hat\{\\phi\}^\{0\}andϕ^1\\hat\{\\phi\}^\{1\}are increasing and contractive onℝ\\mathbb\{R\}, proving \(a\)\. Sinceϑ\\varthetais a bijection, fixed points are preserved by conjugacy:xxis a fixed point ofϕa\\phi^\{a\}if and only ifϑ​\(x\)\\vartheta\(x\)is a fixed point ofϕ^a\\hat\{\\phi\}^\{a\}\. Uniqueness of the fixed points then follows from contractiveness, proving \(b\)\. Finally, becauseϑ\\varthetais increasing,x1<x0x^\{1\}<x^\{0\}impliesx^1<x^0\\hat\{x\}^\{1\}<\\hat\{x\}^\{0\}, proving \(c\)\. ∎

### A\.2Proof details for Section[4\.3](https://arxiv.org/html/2606.11192#S4.SS3)

We collect here the proofs omitted from Section[4\.3](https://arxiv.org/html/2606.11192#S4.SS3)\.

###### Proof of Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\.

For part \(a\), on the event\{τack≥t\}\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}, no ACK has occurred in periods0,…,t−10,\\ldots,t\-1, so the trajectory up to timettcoincides with the no\-ACK skeleton\. Hence

𝔼xz​\[𝟙\{τack≥t\}​r​κ​X​\(t\)​A​\(t\)\]=r​κ​Γt​\(x,z\)​X~t​\(x,z\)​A~t​\(x,z\),\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\big\[\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}\}\\,r\\kappa X\(t\)A\(t\)\\big\]=r\\kappa\\,\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\),and

𝔼xz​\[𝟙\{τack≥t\}​A​\(t\)\]=Γt​\(x,z\)​A~t​\(x,z\)\.\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\big\[\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}\}\\,A\(t\)\\big\]=\\Gamma\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\)\.Since

∑t=0τackβt​r​κ​X​\(t\)​A​\(t\)=∑t=0∞βt​1\{τack≥t\}​r​κ​X​\(t\)​A​\(t\),∑t=0τackβt​A​\(t\)=∑t=0∞βt​1\{τack≥t\}​A​\(t\),\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,r\\kappa X\(t\)A\(t\)=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}\}\\,r\\kappa X\(t\)A\(t\),\\qquad\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,A\(t\)=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}\}\\,A\(t\),taking expectations yields the first two identities\.

ForΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\), note that\{τack=t\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\}means that no ACK occurs before periodttand an ACK occurs in periodtt\. Given\{τack≥t\}\\\{\\tau^\{\\mathrm\{ack\}\}\\geq t\\\}, an ACK in periodttoccurs only ifA​\(t\)=1A\(t\)=1, and then with conditional probabilityκ​X​\(t\)=κ​X~t​\(x,z\)\\kappa X\(t\)=\\kappa\\widetilde\{X\}\_\{t\}\(x,z\)\. Thus

ℙxz​\{τack=t\}=Γt​\(x,z\)​κ​X~t​\(x,z\)​A~t​\(x,z\),\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\}=\\Gamma\_\{t\}\(x,z\)\\,\\kappa\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\),and therefore

Θ~​\(x,z\)=∑t=0∞βt​ℙxz​\{τack=t\}=∑t=0∞βt​Γt​\(x,z\)​κ​X~t​\(x,z\)​A~t​\(x,z\)=1r​F~​\(x,z\)\.\\widetilde\{\\Theta\}\(x,z\)=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=t\\\}=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\Gamma\_\{t\}\(x,z\)\\,\\kappa\\,\\widetilde\{X\}\_\{t\}\(x,z\)\\,\\widetilde\{A\}\_\{t\}\(x,z\)=\\frac\{1\}\{r\}\\,\\widetilde\{F\}\(x,z\)\.
For part \(b\), for any nonnegative process\(Y​\(t\)\)t≥0\(Y\(t\)\)\_\{t\\geq 0\},

∑t=0∞βt​Y​\(t\)=∑t=0τackβt​Y​\(t\)\+𝟙\{τack<∞\}​∑t=τack\+1∞βt​Y​\(t\)\.\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}Y\(t\)=\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}Y\(t\)\+\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}\}\\sum\_\{t=\\tau^\{\\mathrm\{ack\}\}\+1\}^\{\\infty\}\\beta^\{t\}Y\(t\)\.TakingY​\(t\)=r​κ​X​\(t\)​A​\(t\)Y\(t\)=r\\kappa X\(t\)A\(t\)and conditioning onℱτack\+1\\mathscr\{F\}\_\{\\tau^\{\\mathrm\{ack\}\}\+1\}, on\{τack<∞\}\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}an ACK is observed at the end of periodτack\\tau^\{\\mathrm\{ack\}\}, soX​\(τack\+1\)=p11X\(\\tau^\{\\mathrm\{ack\}\}\+1\)=p\_\{11\}\. By the strong Markov property at timeτack\+1\\tau^\{\\mathrm\{ack\}\}\+1,

𝔼xz​\[𝟙\{τack<∞\}​∑t=τack\+1∞βt​r​κ​X​\(t\)​A​\(t\)\|ℱτack\+1\]=𝟙\{τack<∞\}​βτack\+1​F​\(p11,z\)\.\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\left\[\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}\}\\sum\_\{t=\\tau^\{\\mathrm\{ack\}\}\+1\}^\{\\infty\}\\beta^\{t\}\\,r\\kappa X\(t\)A\(t\)\\,\\Big\|\\,\\mathscr\{F\}\_\{\\tau^\{\\mathrm\{ack\}\}\+1\}\\right\]=\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}\}\\beta^\{\\tau^\{\\mathrm\{ack\}\}\+1\}F\(p\_\{11\},z\)\.Taking expectations gives

F​\(x,z\)=𝔼xz​\[∑t=0τackβt​r​κ​X​\(t\)​A​\(t\)\]\+𝔼xz​\[βτack\+1​𝟙\{τack<∞\}\]​F​\(p11,z\)\.F\(x,z\)=\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\big\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}\\beta^\{t\}\\,r\\kappa X\(t\)A\(t\)\\big\]\+\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\big\[\\beta^\{\\tau^\{\\mathrm\{ack\}\}\+1\}\\mathbbm\{1\}\_\{\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}\}\\big\]F\(p\_\{11\},z\)\.Using \(a\) and𝔼xz​\[βτack\+1\]=β​Θ~​\(x,z\)\\mathbb\{E\}\_\{x\}^\{z\}\\big\[\\beta^\{\\tau^\{\\mathrm\{ack\}\}\+1\}\\big\]=\\beta\\,\\widetilde\{\\Theta\}\(x,z\)yields the identity forFF\. The argument forGGis identical withY​\(t\)=A​\(t\)Y\(t\)=A\(t\)\.

For part \(c\), apply part \(b\) withx=p11x=p\_\{11\}and rearrange to obtain \([49](https://arxiv.org/html/2606.11192#S4.E49)\)\. Finally, sinceF​\(p11,z\)F\(p\_\{11\},z\)andG​\(p11,z\)G\(p\_\{11\},z\)are finite andF~,G~≥0\\widetilde\{F\},\\widetilde\{G\}\\geq 0, the identities in \([49](https://arxiv.org/html/2606.11192#S4.E49)\) imply1−β​Θ~​\(p11,z\)\>01\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\>0, that is,0≤β​Θ~​\(p11,z\)<10\\leq\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)<1\. ∎

### A\.3Proof details for Section[4\.4](https://arxiv.org/html/2606.11192#S4.SS4)

We collect here the proofs omitted from Section[4\.4](https://arxiv.org/html/2606.11192#S4.SS4)\.

###### Proof of Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\.

Both updates satisfyϕ0​\(𝒳\)⊆\[p01,p11\]\\phi^\{0\}\(\\mathscr\{X\}\)\\subseteq\[p\_\{01\},p\_\{11\}\]andϕ1​\(𝒳\)⊆\[p01,p11\]\\phi^\{1\}\(\\mathscr\{X\}\)\\subseteq\[p\_\{01\},p\_\{11\}\], so after at most one step the skeleton state lies in\[p01,p11\]\[p\_\{01\},p\_\{11\}\]forever\. Moreover, the passive iteratesϕt0​\(⋅\)\\phi\_\{t\}^\{0\}\(\\cdot\)converge monotonically tox0x^\{0\}, while the active no\-ACK iteratesϕt1​\(⋅\)\\phi\_\{t\}^\{1\}\(\\cdot\)converge monotonically tox1x^\{1\}\(Lemma[4\.6](https://arxiv.org/html/2606.11192#S4.Thmtheorem6)\(b\)–\(c\)\)\. Finally, ifz≥x0z\\geq x^\{0\}, thenϕ0​\(z\)≤z\\phi^\{0\}\(z\)\\leq z, and monotonicity ofϕ0\\phi^\{0\}impliesϕ0​\(y\)≤z\\phi^\{0\}\(y\)\\leq zfor ally≤zy\\leq z\.

*\(a\)*Ifz<p01z<p\_\{01\}, thenX~1​\(x,z\)∈\[p01,p11\]⊂\(z,∞\)\\widetilde\{X\}\_\{1\}\(x,z\)\\in\[p\_\{01\},p\_\{11\}\]\\subset\(z,\\infty\)regardless ofxx\. Hence the skeleton is active from time11onward, and from time0onward ifx\>zx\>z\.

*\(b\)*Ifp01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}, then under passivity the skeleton drifts towardx0\>x1≥zx^\{0\}\>x^\{1\}\\geq z, so it up\-crosseszzin finite timet0=τ↑0​\(x,z\)t\_\{0\}=\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\. OnceX~t0​\(x,z\)\>z\\widetilde\{X\}\_\{t\_\{0\}\}\(x,z\)\>z, the active no\-ACK evolution stays abovezz, since it converges monotonically tox1≥zx^\{1\}\\geq z, and in the boundary casez=x1z=x^\{1\}it remains strictly abovex1x^\{1\}after the first up\-crossing\.

*\(c\)*Assumex1<z<x0x^\{1\}<z<x^\{0\}\. Sincez<x0z<x^\{0\}, any passive phase at or belowzzmust up\-crosszzin finite time\. On the other hand, for everyy\>zy\>z, the active no\-ACK iterates satisfy

ϕt1​\(y\)↓x1<z,\\phi\_\{t\}^\{1\}\(y\)\\downarrow x^\{1\}<z,so every active excursion abovezzmust down\-crosszzin finite time\. After each return to\[0,z\]\[0,z\], the skeleton must up\-cross again in finite time under passivity\. Hence finite blocks of11’s and0’s alternate indefinitely\.

*\(d\)*Ifx0≤z<p11x^\{0\}\\leq z<p\_\{11\}, then once the skeleton enters\[0,z\]\[0,z\]it cannot exceedzzagain\. Thus ifx≤zx\\leq zit is passive forever, while ifx\>zx\>zit performs one active excursion until its first down\-crossing ofzzand is passive forever thereafter\.

*\(e\)*Ifz≥p11z\\geq p\_\{11\}, then fort≥1t\\geq 1the skeleton state lies in\[p01,p11\]⊆\(−∞,z\]\[p\_\{01\},p\_\{11\}\]\\subseteq\(\-\\infty,z\], soA~t​\(x,z\)=0\\widetilde\{A\}\_\{t\}\(x,z\)=0, whileA~0​\(x,z\)=𝟙\{x\>z\}\\widetilde\{A\}\_\{0\}\(x,z\)=\\mathbbm\{1\}\_\{\\\{x\>z\\\}\}\. ∎

###### Proof of Lemma[4\.10](https://arxiv.org/html/2606.11192#S4.Thmtheorem10)\.

*\(a\)*Ifz<x0z<x^\{0\}, then by Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(a\)–\(c\) the skeleton itinerary contains infinitely many11’s\.

WritingNtact​\(x,z\)≜∑j=0t−1A~j​\(x,z\),N\_\{t\}^\{\\mathrm\{act\}\}\(x,z\)\\triangleq\\sum\_\{j=0\}^\{t\-1\}\\widetilde\{A\}\_\{j\}\(x,z\),we therefore haveNtact​\(x,z\)→∞N\_\{t\}^\{\\mathrm\{act\}\}\(x,z\)\\to\\infty\. Moreover, on every active periodjj,X~j​\(x,z\)∈\[p01,p11\]\\widetilde\{X\}\_\{j\}\(x,z\)\\in\[p\_\{01\},p\_\{11\}\], so

1−κ​X~j​\(x,z\)≤1−κ​p01<1\.1\-\\kappa\\widetilde\{X\}\_\{j\}\(x,z\)\\leq 1\-\\kappa p\_\{01\}<1\.Using \([42](https://arxiv.org/html/2606.11192#S4.E42)\),

Γt​\(x,z\)=∏0≤j<t:A~j​\(x,z\)=1\(1−κ​X~j​\(x,z\)\)≤\(1−κ​p01\)Ntact​\(x,z\)→0\.\\Gamma\_\{t\}\(x,z\)=\\prod\_\{0\\leq j<t:\\ \\widetilde\{A\}\_\{j\}\(x,z\)=1\}\\bigl\(1\-\\kappa\\,\\widetilde\{X\}\_\{j\}\(x,z\)\\bigr\)\\leq\(1\-\\kappa p\_\{01\}\)^\{N\_\{t\}^\{\\mathrm\{act\}\}\(x,z\)\}\\to 0\.Henceℙxz​\{τack=∞\}=Γ∞​\(x,z\)=0\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}=\\infty\\\}=\\Gamma\_\{\\infty\}\(x,z\)=0\.

*\(b\)*Assumex0≤z<p11x^\{0\}\\leq z<p\_\{11\}\. Ifx≤zx\\leq z, then the policy is passive att=0t=0, and sinceϕ0​\(y\)≤z\\phi^\{0\}\(y\)\\leq zfor ally≤zy\\leq z, it remains passive forever\. Thusτack=∞\\tau^\{\\mathrm\{ack\}\}=\\inftya\.s\. Ifx\>zx\>z, then absent an ACK the skeleton is active until timeτ↓1​\(x,z\)\\tau\_\{\\downarrow\}^\{1\}\(x,z\)and passive forever thereafter \(Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(d\)\)\. Hence ACKs are possible only in periodsj=0,1,…,τ↓1​\(x,z\)−1j=0,1,\\ldots,\\tau\_\{\\downarrow\}^\{1\}\(x,z\)\-1, during whichX~j​\(x,z\)=ϕj1​\(x\)\\widetilde\{X\}\_\{j\}\(x,z\)=\\phi\_\{j\}^\{1\}\(x\)andA~j​\(x,z\)=1\\widetilde\{A\}\_\{j\}\(x,z\)=1\. Multiplying the corresponding no\-ACK probabilities yields \([50](https://arxiv.org/html/2606.11192#S4.E50)\)\.

*\(c\)*Ifz≥p11z\\geq p\_\{11\}, then after time0the belief lies in\[p01,p11\]⊆\(−∞,z\]\[p\_\{01\},p\_\{11\}\]\\subseteq\(\-\\infty,z\]regardless of the outcome at time0, so the policy cannot activate aftert=0t=0\. Thereforeℙxz​\{τack<∞\}=κ​x​1\{x\>z\}\\mathbb\{P\}\_\{x\}^\{z\}\\\{\\tau^\{\\mathrm\{ack\}\}<\\infty\\\}=\\kappa x\\,\\mathbbm\{1\}\_\{\\\{x\>z\\\}\}\. ∎

###### Proof of Proposition[4\.11](https://arxiv.org/html/2606.11192#S4.Thmtheorem11)\.

The pre\-ACK metricsF~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), andΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)are defined in \([45](https://arxiv.org/html/2606.11192#S4.E45)\) and \([47](https://arxiv.org/html/2606.11192#S4.E47)\) as deterministic functionals of the no\-ACK skeleton orbitX~t​\(x,z\)\\widetilde\{X\}\_\{t\}\(x,z\), itineraryA~t​\(x,z\)\\widetilde\{A\}\_\{t\}\(x,z\), and survival probabilitiesΓt​\(x,z\)\\Gamma\_\{t\}\(x,z\)\. It is therefore enough to show that, on each of the statedzz\-intervals, the no\-ACK skeleton orbit and survival probabilities are independent of the particular value ofzz\.

*\(a\)*By Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(a\), ifz<p01z<p\_\{01\}, then the itinerary is1∞1^\{\\infty\}whenx\>zx\>z, and01∞01^\{\\infty\}whenx≤zx\\leq z\. Thus, on each of the two stated intervals, the skeleton itinerary is fixed\. Moreover, the corresponding skeleton orbit is also fixed: it is\(ϕt1​\(x\)\)t≥0\(\\phi\_\{t\}^\{1\}\(x\)\)\_\{t\\geq 0\}in the case1∞1^\{\\infty\}, and, in the case01∞01^\{\\infty\},

x,ϕ0​\(x\),ϕ1​\(ϕ0​\(x\)\),ϕ21​\(ϕ0​\(x\)\),…\.x,\\ \\phi^\{0\}\(x\),\\ \\phi^\{1\}\(\\phi^\{0\}\(x\)\),\\ \\phi\_\{2\}^\{1\}\(\\phi^\{0\}\(x\)\),\\ldots\.
Hence the survival probabilitiesΓt​\(x,z\)\\Gamma\_\{t\}\(x,z\)are also fixed, and the three pre\-ACK metrics are constant\.

*\(b\)*Fixn∈ℤ\+n\\in\\mathbb\{Z\}\_\{\+\}andz∈In↑​\(x\)z\\in I\_\{n\}^\{\\uparrow\}\(x\)\. By definition,τ↑0​\(x,z\)=n\\tau\_\{\\uparrow\}^\{0\}\(x,z\)=n, and by Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(b\), the skeleton itinerary is0n​1∞0^\{n\}1^\{\\infty\}\.

Therefore the skeleton orbit depends only onxxandnn, not on the particularz∈In↑​\(x\)z\\in I\_\{n\}^\{\\uparrow\}\(x\), being

x,ϕ0​\(x\),ϕ20​\(x\),…,ϕn0​\(x\),ϕ1​\(ϕn0​\(x\)\),ϕ21​\(ϕn0​\(x\)\),…\.x,\\ \\phi^\{0\}\(x\),\\ \\phi\_\{2\}^\{0\}\(x\),\\ \\ldots,\\ \\phi\_\{n\}^\{0\}\(x\),\\ \\phi^\{1\}\(\\phi\_\{n\}^\{0\}\(x\)\),\\ \\phi\_\{2\}^\{1\}\(\\phi\_\{n\}^\{0\}\(x\)\),\\ldots\.
The corresponding survival probabilitiesΓt​\(x,z\)\\Gamma\_\{t\}\(x,z\)are therefore also independent ofzzonIn↑​\(x\)I\_\{n\}^\{\\uparrow\}\(x\)\. HenceF~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), andΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)are constant onIn↑​\(x\)I\_\{n\}^\{\\uparrow\}\(x\)\.

*\(c\)*Ifx≤zx\\leq z, then Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(d\) gives the itinerary0∞0^\{\\infty\}, so no activation ever occurs and therefore

F~​\(x,z\)=G~​\(x,z\)=Θ~​\(x,z\)=0\.\\widetilde\{F\}\(x,z\)=\\widetilde\{G\}\(x,z\)=\\widetilde\{\\Theta\}\(x,z\)=0\.
Assume now thatx\>zx\>z, and fixn∈ℕn\\in\\mathbb\{N\}andz∈In↓​\(x\)z\\in I\_\{n\}^\{\\downarrow\}\(x\)\. Thenτ↓1​\(x,z\)=n\\tau\_\{\\downarrow\}^\{1\}\(x,z\)=n, and by Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(d\) the skeleton itinerary is1n​0∞1^\{n\}0^\{\\infty\}\. Hence the skeleton orbit depends only onxxandnn, not on the particularz∈In↓​\(x\)z\\in I\_\{n\}^\{\\downarrow\}\(x\), being

x,ϕ1​\(x\),ϕ21​\(x\),…,ϕn1​\(x\),ϕ0​\(ϕn1​\(x\)\),ϕ20​\(ϕn1​\(x\)\),…\.x,\\ \\phi^\{1\}\(x\),\\ \\phi\_\{2\}^\{1\}\(x\),\\ \\ldots,\\ \\phi\_\{n\}^\{1\}\(x\),\\ \\phi^\{0\}\(\\phi\_\{n\}^\{1\}\(x\)\),\\ \\phi\_\{2\}^\{0\}\(\\phi\_\{n\}^\{1\}\(x\)\),\\ldots\.
Thus the survival probabilities are fixed onIn↓​\(x\)I\_\{n\}^\{\\downarrow\}\(x\), and the three pre\-ACK metrics are constant there\.

*\(d\)*By Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(e\), ifz≥p11z\\geq p\_\{11\}, then the itinerary is10∞10^\{\\infty\}whenx\>zx\>z, and0∞0^\{\\infty\}whenx≤zx\\leq z\.

Hence, on each of the two stated intervals, both the skeleton itinerary and the corresponding skeleton orbit are fixed, so the survival probabilities are fixed as well\. ThereforeF~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\), andΘ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\)are constant on each interval\. ∎

###### Proof of Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\.

The step\-function property follows because, in each regime,F​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)depend onzzonly through threshold comparisons along the no\-ACK skeleton and through integer\-valued crossing times \(τ↑0​\(x,z\)\\tau\_\{\\uparrow\}^\{0\}\(x,z\)andτ↓1​\(x,z\)\\tau\_\{\\downarrow\}^\{1\}\(x,z\)\)\. Hence jumps can occur only whenzzcoincides with a skeleton belief value reached fromxxor fromp11p\_\{11\}\.

\(a\)Ifz<p01z<p\_\{01\}, then\[p01,p11\]⊂\(z,∞\)\[p\_\{01\},p\_\{11\}\]\\subset\(z,\\infty\), so after at most one step the policy is always active\. By Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(a\), the itinerary is1∞1^\{\\infty\}ifx\>zx\>zand01∞01^\{\\infty\}ifx≤zx\\leq z\. Ifx\>zx\>z, thenA​\(t\)≡1A\(t\)\\equiv 1, so

F​\(x,z\)=r​κ​∑t=0∞βt​𝔼xz​\[X​\(t\)\],G​\(x,z\)=∑t=0∞βt=11−β\.F\(x,z\)=r\\kappa\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}\\,\\mathbb\{E\}\_\{x\}^\{z\}\[X\(t\)\],\\qquad G\(x,z\)=\\sum\_\{t=0\}^\{\\infty\}\\beta^\{t\}=\\frac\{1\}\{1\-\\beta\}\.Writingmt≜𝔼xz​\[X​\(t\)\]m\_\{t\}\\triangleq\\mathbb\{E\}\_\{x\}^\{z\}\[X\(t\)\], the one\-step posterior\-mean identity \([20](https://arxiv.org/html/2606.11192#S4.E20)\) yieldsmt\+1=p01\+ρ​mt,m0=x,m\_\{t\+1\}=p\_\{01\}\+\\rho m\_\{t\},\\kern 5\.0ptm\_\{0\}=x,hencemt=x0\+\(x−x0\)​ρt\.m\_\{t\}=x^\{0\}\+\(x\-x^\{0\}\)\\rho^\{t\}\.Substituting and summing the geometric series gives

F​\(x,z\)=r​κ​\[x01−β−x0−x1−β​ρ\]\.F\(x,z\)=r\\kappa\\left\[\\frac\{x^\{0\}\}\{1\-\\beta\}\-\\frac\{x^\{0\}\-x\}\{1\-\\beta\\rho\}\\right\]\.
Ifx≤zx\\leq z, thenA​\(0\)=0A\(0\)=0,X​\(1\)=ϕ0​\(x\)X\(1\)=\\phi^\{0\}\(x\), and thereafter the policy is always active\. Thus

G​\(x,z\)=∑t=1∞βt=β1−β,F​\(x,z\)=β​F​\(ϕ0​\(x\),z\),G\(x,z\)=\\sum\_\{t=1\}^\{\\infty\}\\beta^\{t\}=\\frac\{\\beta\}\{1\-\\beta\},\\qquad F\(x,z\)=\\beta\\,F\(\\phi^\{0\}\(x\),z\),and the preceding formula applied at initial stateϕ0​\(x\)\\phi^\{0\}\(x\)yields the stated expression\.

\(b\)Lett0≜τ↑0​\(x,z\)t\_\{0\}\\triangleq\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\. By Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(b\), the itinerary is0t0​1∞0^\{t\_\{0\}\}1^\{\\infty\}\. Hence

G​\(x,z\)=∑t=t0∞βt=βt01−β\.G\(x,z\)=\\sum\_\{t=t\_\{0\}\}^\{\\infty\}\\beta^\{t\}=\\frac\{\\beta^\{t\_\{0\}\}\}\{1\-\\beta\}\.
Also, up to timet0t\_\{0\}the process is passive, soX​\(t0\)=ϕt00​\(x\)X\(t\_\{0\}\)=\\phi\_\{t\_\{0\}\}^\{0\}\(x\)deterministically, and from timet0t\_\{0\}onward the policy is always active\. Applying part \(a\) from timet0t\_\{0\}onward yields

F​\(x,z\)=r​κ​βt0​\[x01−β−x0−ϕt00​\(x\)1−β​ρ\]\.F\(x,z\)=r\\kappa\\,\\beta^\{t\_\{0\}\}\\left\[\\frac\{x^\{0\}\}\{1\-\\beta\}\-\\frac\{x^\{0\}\-\\phi\_\{t\_\{0\}\}^\{0\}\(x\)\}\{1\-\\beta\\rho\}\\right\]\.
\(c\)This is exactly Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(b\)–\(c\)\.

\(d\)Ifx≤zx\\leq z, then the policy never activates by Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(d\), soF​\(x,z\)=G​\(x,z\)=0F\(x,z\)=G\(x,z\)=0\.

Assume next thatx\>zx\>z\. By Lemma[4\.9](https://arxiv.org/html/2606.11192#S4.Thmtheorem9)\(d\), the no\-ACK skeleton starting fromxxhas the form

σ~​\(x,z\)=1n​0∞,n=τ↓1​\(x,z\),\\widetilde\{\\sigma\}\(x,z\)=1^\{n\}0^\{\\infty\},\\qquad n=\\tau\_\{\\downarrow\}^\{1\}\(x,z\),and the no\-ACK skeleton starting from the post\-ACK reset statep11p\_\{11\}has the form

σ~​\(p11,z\)=1m​0∞,m=τ↓1​\(p11,z\)\.\\widetilde\{\\sigma\}\(p\_\{11\},z\)=1^\{m\}0^\{\\infty\},\\qquad m=\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},z\)\.Hence, by Proposition[4\.11](https://arxiv.org/html/2606.11192#S4.Thmtheorem11)\(c\), on every interval

In,m​\(x\)=\{z∈\[x0,min⁡\{x,p11\}\):τ↓1​\(x,z\)=n,τ↓1​\(p11,z\)=m\}I\_\{n,m\}\(x\)=\\Bigl\\\{z\\in\[x^\{0\},\\min\\\{x,p\_\{11\}\\\}\):\\tau\_\{\\downarrow\}^\{1\}\(x,z\)=n,\\ \\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},z\)=m\\Bigr\\\}the pre\-ACK metrics are constant, and are given by the stated finite sums

G~​\(x,z\)=Ax\(n\),Θ~​\(x,z\)=Bx\(n\),F~​\(x,z\)=r​Bx\(n\),\\widetilde\{G\}\(x,z\)=A\_\{x\}^\{\(n\)\},\\qquad\\widetilde\{\\Theta\}\(x,z\)=B\_\{x\}^\{\(n\)\},\\qquad\\widetilde\{F\}\(x,z\)=r\\,B\_\{x\}^\{\(n\)\},together with

G~​\(p11,z\)=A11\(m\),Θ~​\(p11,z\)=B11\(m\),F~​\(p11,z\)=r​B11\(m\)\.\\widetilde\{G\}\(p\_\{11\},z\)=A\_\{11\}^\{\(m\)\},\\qquad\\widetilde\{\\Theta\}\(p\_\{11\},z\)=B\_\{11\}^\{\(m\)\},\\qquad\\widetilde\{F\}\(p\_\{11\},z\)=r\\,B\_\{11\}^\{\(m\)\}\.
Applying Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(c\) atp11p\_\{11\}yields

F​\(p11,z\)=r​B11\(m\)1−β​B11\(m\),G​\(p11,z\)=A11\(m\)1−β​B11\(m\)\.F\(p\_\{11\},z\)=\\frac\{r\\,B\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\},\\qquad G\(p\_\{11\},z\)=\\frac\{A\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\}\.Then Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(b\) gives

F​\(x,z\)=F~​\(x,z\)\+β​Θ~​\(x,z\)​F​\(p11,z\)=r​Bx\(n\)\+β​Bx\(n\)​r​B11\(m\)1−β​B11\(m\)=r​Bx\(n\)1−β​B11\(m\),F\(x,z\)=\\widetilde\{F\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,F\(p\_\{11\},z\)=r\\,B\_\{x\}^\{\(n\)\}\+\\beta\\,B\_\{x\}^\{\(n\)\}\\frac\{r\\,B\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\}=\\frac\{r\\,B\_\{x\}^\{\(n\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\},G​\(x,z\)=G~​\(x,z\)\+β​Θ~​\(x,z\)​G​\(p11,z\)=Ax\(n\)\+β​Bx\(n\)​A11\(m\)1−β​B11\(m\),G\(x,z\)=\\widetilde\{G\}\(x,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(x,z\)\\,G\(p\_\{11\},z\)=A\_\{x\}^\{\(n\)\}\+\\beta\\,B\_\{x\}^\{\(n\)\}\\frac\{A\_\{11\}^\{\(m\)\}\}\{1\-\\beta\\,B\_\{11\}^\{\(m\)\}\},which is the stated formula\.

Since these expressions depend onzzonly through the pair of integers\(τ↓1​\(x,z\),τ↓1​\(p11,z\)\)\\bigl\(\\tau\_\{\\downarrow\}^\{1\}\(x,z\),\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},z\)\\bigr\), it follows thatF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)are constant on each nonempty intervalIn,m​\(x\)I\_\{n,m\}\(x\), and therefore piecewise constant on\[x0,p11\)\[x^\{0\},p\_\{11\}\)\.

\(e\)Ifz≥p11z\\geq p\_\{11\}, then regardless of the action and outcome at time0, the belief at time11lies in\[p01,p11\]⊆\(−∞,z\]\[p\_\{01\},p\_\{11\}\]\\subseteq\(\-\\infty,z\]\. Hence the policy cannot activate after time0, so

G​\(x,z\)=𝔼xz​\[A​\(0\)\]=𝟙\{x\>z\},F​\(x,z\)=𝔼xz​\[r​κ​X​\(0\)​A​\(0\)\]=r​κ​x​1\{x\>z\}\.G\(x,z\)=\\mathbb\{E\}\_\{x\}^\{z\}\[A\(0\)\]=\\mathbbm\{1\}\_\{\\\{x\>z\\\}\},\\qquad F\(x,z\)=\\mathbb\{E\}\_\{x\}^\{z\}\[r\\kappa X\(0\)A\(0\)\]=r\\kappa x\\,\\mathbbm\{1\}\_\{\\\{x\>z\\\}\}\.∎

###### Proof of Lemma[4\.13](https://arxiv.org/html/2606.11192#S4.Thmtheorem13)\.

We consider the tractable regimes separately\.

*Case 1:z<p01z<p\_\{01\}\.*By Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\(a\),

G​\(x,z\)=\{β1−β,x≤z,11−β,x\>z,G\(x,z\)=\\begin\{cases\}\\dfrac\{\\beta\}\{1\-\\beta\},&x\\leq z,\\\\\[2\.84526pt\] \\dfrac\{1\}\{1\-\\beta\},&x\>z,\\end\{cases\}which is clearly nondecreasing inxx\.

*Case 2:p01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}\.*By Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\(b\),

G​\(x,z\)=βτ↑0​\(x,z\)1−β\.G\(x,z\)=\\frac\{\\beta^\{\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\}\}\{1\-\\beta\}\.For fixedzz, the mapx↦τ↑0​\(x,z\)x\\mapsto\\tau\_\{\\uparrow\}^\{0\}\(x,z\)is nonincreasing, becauseϕt0​\(⋅\)\\phi\_\{t\}^\{0\}\(\\cdot\)is increasing for everyt≥0t\\geq 0\. Since0<β<10<\\beta<1, the mapn↦βnn\\mapsto\\beta^\{n\}is decreasing onℤ\+\\mathbb\{Z\}\_\{\+\}, hencex↦βτ↑0​\(x,z\)x\\mapsto\\beta^\{\\tau\_\{\\uparrow\}^\{0\}\(x,z\)\}is nondecreasing\. Thereforex↦G​\(x,z\)x\\mapsto G\(x,z\)is nondecreasing\.

*Case 3:x0≤z<p11x^\{0\}\\leq z<p\_\{11\}\.*Let𝒯z\\mathscr\{T\}\_\{z\}be the Bellman evaluation operator for the work metric:

\(𝒯z​H\)​\(x\)≜\{β​H​\(ϕ0​\(x\)\),x≤z,1\+β​\(κ​x​H​\(p11\)\+\(1−κ​x\)​H​\(ϕ1​\(x\)\)\),x\>z\.\(\\mathscr\{T\}\_\{z\}H\)\(x\)\\triangleq\\begin\{cases\}\\beta\\,H\(\\phi^\{0\}\(x\)\),&x\\leq z,\\\\\[2\.84526pt\] 1\+\\beta\\Big\(\\kappa x\\,H\(p\_\{11\}\)\+\(1\-\\kappa x\)\\,H\(\\phi^\{1\}\(x\)\)\\Big\),&x\>z\.\\end\{cases\}SetH0≡0H\_\{0\}\\equiv 0andHn\+1≜𝒯z​HnH\_\{n\+1\}\\triangleq\\mathscr\{T\}\_\{z\}H\_\{n\}\. ThenHn​\(x\)↑G​\(x,z\)H\_\{n\}\(x\)\\uparrow G\(x,z\)pointwise asn→∞n\\to\\infty\.

We prove by induction that eachHnH\_\{n\}is nondecreasing\. This is clear forH0H\_\{0\}\. AssumeHnH\_\{n\}is nondecreasing\. Ifx1≤x2≤zx\_\{1\}\\leq x\_\{2\}\\leq z, thenϕ0​\(x1\)≤ϕ0​\(x2\)\\phi^\{0\}\(x\_\{1\}\)\\leq\\phi^\{0\}\(x\_\{2\}\), and hence

Hn\+1​\(x1\)=β​Hn​\(ϕ0​\(x1\)\)≤β​Hn​\(ϕ0​\(x2\)\)=Hn\+1​\(x2\)\.H\_\{n\+1\}\(x\_\{1\}\)=\\beta H\_\{n\}\(\\phi^\{0\}\(x\_\{1\}\)\)\\leq\\beta H\_\{n\}\(\\phi^\{0\}\(x\_\{2\}\)\)=H\_\{n\+1\}\(x\_\{2\}\)\.Ifz<x1≤x2z<x\_\{1\}\\leq x\_\{2\}, then

Hn\+1​\(x2\)−Hn\+1​\(x1\)\\displaystyle H\_\{n\+1\}\(x\_\{2\}\)\-H\_\{n\+1\}\(x\_\{1\}\)=β​κ​\(x2−x1\)​Hn​\(p11\)\+β​\(\(1−κ​x2\)​Hn​\(ϕ1​\(x2\)\)−\(1−κ​x1\)​Hn​\(ϕ1​\(x1\)\)\)\\displaystyle=\\beta\\kappa\(x\_\{2\}\-x\_\{1\}\)H\_\{n\}\(p\_\{11\}\)\+\\beta\\Big\(\(1\-\\kappa x\_\{2\}\)H\_\{n\}\(\\phi^\{1\}\(x\_\{2\}\)\)\-\(1\-\\kappa x\_\{1\}\)H\_\{n\}\(\\phi^\{1\}\(x\_\{1\}\)\)\\Big\)=β​κ​\(x2−x1\)​\(Hn​\(p11\)−Hn​\(ϕ1​\(x2\)\)\)\+β​\(1−κ​x1\)​\(Hn​\(ϕ1​\(x2\)\)−Hn​\(ϕ1​\(x1\)\)\)≥0,\\displaystyle=\\beta\\kappa\(x\_\{2\}\-x\_\{1\}\)\\bigl\(H\_\{n\}\(p\_\{11\}\)\-H\_\{n\}\(\\phi^\{1\}\(x\_\{2\}\)\)\\bigr\)\+\\beta\(1\-\\kappa x\_\{1\}\)\\bigl\(H\_\{n\}\(\\phi^\{1\}\(x\_\{2\}\)\)\-H\_\{n\}\(\\phi^\{1\}\(x\_\{1\}\)\)\\bigr\)\\geq 0,becauseϕ1\\phi^\{1\}is increasing,HnH\_\{n\}is nondecreasing, andϕ1​\(x2\)≤p11\\phi^\{1\}\(x\_\{2\}\)\\leq p\_\{11\}\.

Ifx1≤z<x2x\_\{1\}\\leq z<x\_\{2\}, thenHn\+1​\(x1\)=β​Hn​\(ϕ0​\(x1\)\)H\_\{n\+1\}\(x\_\{1\}\)=\\beta H\_\{n\}\(\\phi^\{0\}\(x\_\{1\}\)\), while

Hn\+1​\(x2\)=1\+β​\(κ​x2​Hn​\(p11\)\+\(1−κ​x2\)​Hn​\(ϕ1​\(x2\)\)\)≥1\.H\_\{n\+1\}\(x\_\{2\}\)=1\+\\beta\\Big\(\\kappa x\_\{2\}H\_\{n\}\(p\_\{11\}\)\+\(1\-\\kappa x\_\{2\}\)H\_\{n\}\(\\phi^\{1\}\(x\_\{2\}\)\)\\Big\)\\geq 1\.Sincez≥x0z\\geq x^\{0\}, the interval\[0,z\]\[0,z\]is trapping underϕ0\\phi^\{0\}, soϕ0​\(y\)≤z\\phi^\{0\}\(y\)\\leq zfor everyy≤zy\\leq z\. Starting fromH0≡0H\_\{0\}\\equiv 0, it follows by induction thatHn​\(y\)=0H\_\{n\}\(y\)=0for everyy≤zy\\leq zand everynn\. In particular,Hn\+1​\(x1\)=0≤Hn\+1​\(x2\)H\_\{n\+1\}\(x\_\{1\}\)=0\\leq H\_\{n\+1\}\(x\_\{2\}\)\.

ThusHn\+1H\_\{n\+1\}is nondecreasing\. Passing to the limit yields thatG​\(⋅,z\)G\(\\cdot,z\)is nondecreasing\.

*Case 4:z≥p11z\\geq p\_\{11\}\.*By Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\(e\),G​\(x,z\)=𝟙\{x\>z\},G\(x,z\)=\\mathbbm\{1\}\_\{\\\{x\>z\\\}\},which is nondecreasing inxx\.

Combining the four cases proves the result\. ∎

### A\.4Proof details for Sections[4\.5](https://arxiv.org/html/2606.11192#S4.SS5)–[4\.6](https://arxiv.org/html/2606.11192#S4.SS6)

We collect here the proofs omitted from Sections[4\.5](https://arxiv.org/html/2606.11192#S4.SS5)–[4\.6](https://arxiv.org/html/2606.11192#S4.SS6)\.

###### Proof of Lemma[4\.14](https://arxiv.org/html/2606.11192#S4.Thmtheorem14)\.

By one\-step conditioning at time0,

f​\(x,z\)\\displaystyle f\(x,z\)=r​κ​x\+β​\(κ​x​F​\(p11,z\)\+\(1−κ​x\)​F​\(ϕ1​\(x\),z\)−F​\(ϕ0​\(x\),z\)\),\\displaystyle=r\\kappa x\+\\beta\\Big\(\\kappa x\\,F\(p\_\{11\},z\)\+\(1\-\\kappa x\)\\,F\(\\phi^\{1\}\(x\),z\)\-F\(\\phi^\{0\}\(x\),z\)\\Big\),g​\(x,z\)\\displaystyle g\(x,z\)=1\+β​\(κ​x​G​\(p11,z\)\+\(1−κ​x\)​G​\(ϕ1​\(x\),z\)−G​\(ϕ0​\(x\),z\)\)\.\\displaystyle=1\+\\beta\\Big\(\\kappa x\\,G\(p\_\{11\},z\)\+\(1\-\\kappa x\)\\,G\(\\phi^\{1\}\(x\),z\)\-G\(\\phi^\{0\}\(x\),z\)\\Big\)\.
Applying Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(b\) toF​\(y,z\)F\(y,z\)andG​\(y,z\)G\(y,z\)fory∈\{ϕ1​\(x\),ϕ0​\(x\)\}y\\in\\\{\\phi^\{1\}\(x\),\\phi^\{0\}\(x\)\\\}gives

F​\(y,z\)=F~​\(y,z\)\+β​Θ~​\(y,z\)​F​\(p11,z\),G​\(y,z\)=G~​\(y,z\)\+β​Θ~​\(y,z\)​G​\(p11,z\)\.F\(y,z\)=\\widetilde\{F\}\(y,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(y,z\)\\,F\(p\_\{11\},z\),\\qquad G\(y,z\)=\\widetilde\{G\}\(y,z\)\+\\beta\\,\\widetilde\{\\Theta\}\(y,z\)\\,G\(p\_\{11\},z\)\.
Substituting these expressions and collecting the constant terms and the coefficients ofF​\(p11,z\)F\(p\_\{11\},z\)andG​\(p11,z\)G\(p\_\{11\},z\)yields \([54](https://arxiv.org/html/2606.11192#S4.E54)\)–\([55](https://arxiv.org/html/2606.11192#S4.E55)\), together with \([51](https://arxiv.org/html/2606.11192#S4.E51)\)–\([53](https://arxiv.org/html/2606.11192#S4.E53)\)\. ∎

###### Proof of Lemma[4\.15](https://arxiv.org/html/2606.11192#S4.Thmtheorem15)\.

By Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(a\), for everyy∈𝒳y\\in\\mathscr\{X\}we haveΘ~​\(y,z\)=F~​\(y,z\)/r\.\\widetilde\{\\Theta\}\(y,z\)=\{\\widetilde\{F\}\(y,z\)\}/\{r\}\.Substituting this into \([53](https://arxiv.org/html/2606.11192#S4.E53)\) and comparing with \([51](https://arxiv.org/html/2606.11192#S4.E51)\) gives

θ~​\(x,z\)=1r​\[r​κ​x\+β​\(\(1−κ​x\)​F~​\(ϕ1​\(x\),z\)−F~​\(ϕ0​\(x\),z\)\)\]=f~​\(x,z\)r\.\\tilde\{\\theta\}\(x,z\)=\\frac\{1\}\{r\}\\Bigl\[r\\kappa x\+\\beta\\bigl\(\(1\-\\kappa x\)\\widetilde\{F\}\(\\phi^\{1\}\(x\),z\)\-\\widetilde\{F\}\(\\phi^\{0\}\(x\),z\)\\bigr\)\\Bigr\]=\\frac\{\\tilde\{f\}\(x,z\)\}\{r\}\.∎

###### Proof of Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\.

The one\-step identities \([57](https://arxiv.org/html/2606.11192#S4.E57)\)–\([58](https://arxiv.org/html/2606.11192#S4.E58)\) follow by conditioning on the time\-0outcome: under action11, an ACK occurs with probabilityκ​x\\kappa x, sending the belief top11p\_\{11\}, while with probability1−κ​x1\-\\kappa xno ACK is observed and the belief updates toϕ1​\(x\)\\phi^\{1\}\(x\); under action0, the belief updates deterministically toϕ0​\(x\)\\phi^\{0\}\(x\)\. In either case, the policy follows thezz\-threshold policy from time11onward\.

*Regimes \(a\), \(b\), and \(e\)\.*These follow by substituting into \([57](https://arxiv.org/html/2606.11192#S4.E57)\)–\([58](https://arxiv.org/html/2606.11192#S4.E58)\) the corresponding formulas forF​\(⋅,z\)F\(\\cdot,z\)andG​\(⋅,z\)G\(\\cdot,z\)from Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\.

*Regime \(c\)\.*Forx1<z<x0x^\{1\}<z<x^\{0\}, letD​\(z\)≜1−β​Θ~​\(p11,z\)\.D\(z\)\\triangleq 1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},z\)\.Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(b\)–\(c\) gives

F​\(p11,z\)=F~​\(p11,z\)D​\(z\),G​\(p11,z\)=G~​\(p11,z\)D​\(z\)\.F\(p\_\{11\},z\)=\\frac\{\\widetilde\{F\}\(p\_\{11\},z\)\}\{D\(z\)\},\\qquad G\(p\_\{11\},z\)=\\frac\{\\widetilde\{G\}\(p\_\{11\},z\)\}\{D\(z\)\}\.
By Lemma[4\.14](https://arxiv.org/html/2606.11192#S4.Thmtheorem14),

f​\(x,z\)=f~​\(x,z\)\+β​θ~​\(x,z\)​F​\(p11,z\),g​\(x,z\)=g~​\(x,z\)\+β​θ~​\(x,z\)​G​\(p11,z\)\.f\(x,z\)=\\tilde\{f\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,F\(p\_\{11\},z\),\\qquad g\(x,z\)=\\tilde\{g\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,G\(p\_\{11\},z\)\.
Using Lemma[4\.15](https://arxiv.org/html/2606.11192#S4.Thmtheorem15),

f​\(x,z\)=r​θ~​\(x,z\)\+β​θ~​\(x,z\)​F~​\(p11,z\)D​\(z\)=θ~​\(x,z\)​\(r\+β​F~​\(p11,z\)D​\(z\)\)=r​θ~​\(x,z\)D​\(z\),f\(x,z\)=r\\,\\tilde\{\\theta\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,\\frac\{\\widetilde\{F\}\(p\_\{11\},z\)\}\{D\(z\)\}=\\tilde\{\\theta\}\(x,z\)\\left\(r\+\\beta\\,\\frac\{\\widetilde\{F\}\(p\_\{11\},z\)\}\{D\(z\)\}\\right\)=\\frac\{r\\,\\tilde\{\\theta\}\(x,z\)\}\{D\(z\)\},which gives \([59](https://arxiv.org/html/2606.11192#S4.E59)\)\. Similarly, we obtain \([60](https://arxiv.org/html/2606.11192#S4.E60)\):

g​\(x,z\)=g~​\(x,z\)\+β​θ~​\(x,z\)​G~​\(p11,z\)D​\(z\)\.g\(x,z\)=\\tilde\{g\}\(x,z\)\+\\beta\\,\\tilde\{\\theta\}\(x,z\)\\,\\frac\{\\widetilde\{G\}\(p\_\{11\},z\)\}\{D\(z\)\}\.
*Regime \(d\)\.*The same substitutions as in regime \(c\) apply\. The fact thatF~​\(⋅,z\)\\widetilde\{F\}\(\\cdot,z\),G~​\(⋅,z\)\\widetilde\{G\}\(\\cdot,z\), andΘ~​\(⋅,z\)\\widetilde\{\\Theta\}\(\\cdot,z\)reduce to finite sums follows from Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)\(d\) and Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(a\)\. ∎

###### Proof of Lemma[4\.17](https://arxiv.org/html/2606.11192#S4.Thmtheorem17)\.

*\(a\)*Ifz<p01z<p\_\{01\}orz≥p11z\\geq p\_\{11\}, then Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(a\) and \(e\) giveg​\(x,z\)=1g\(x,z\)=1\.

*\(b\)*Forp01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}, Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(b\) yields

g​\(x,z\)=1\+β1−β​\(κ​x\+\(1−κ​x\)​βτ1−βτ0\),g\(x,z\)=1\+\\frac\{\\beta\}\{1\-\\beta\}\\Big\(\\kappa x\+\(1\-\\kappa x\)\\beta^\{\\tau\_\{1\}\}\-\\beta^\{\\tau\_\{0\}\}\\Big\),withτ0=τ↑0​\(ϕ0​\(x\),z\),τ1=τ↑0​\(ϕ1​\(x\),z\)\.\\tau\_\{0\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{0\}\(x\),z\),\\kern 5\.0pt\\tau\_\{1\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{1\}\(x\),z\)\.

We first show that

τ1≤τ0\+1\.\\tau\_\{1\}\\leq\\tau\_\{0\}\+1\.\(75\)
Ifx\>x1x\>x^\{1\}, then Lemma[4\.6](https://arxiv.org/html/2606.11192#S4.Thmtheorem6)\(b\) givesϕ1​\(x\)\>x1≥z,\\phi^\{1\}\(x\)\>x^\{1\}\\geq z,soτ1=0\\tau\_\{1\}=0, and \([75](https://arxiv.org/html/2606.11192#A1.E75)\) is immediate\. Assume now thatx≤x1x\\leq x^\{1\}\. Then Lemma[4\.6](https://arxiv.org/html/2606.11192#S4.Thmtheorem6)\(b\) impliesϕ1​\(x\)≥x,\\phi^\{1\}\(x\)\\geq x,hencex0−ϕ1​\(x\)≤x0−x\.x^\{0\}\-\\phi^\{1\}\(x\)\\leq x^\{0\}\-x\.Also,

x0−ϕ0​\(x\)=x0−\(p01\+ρ​x\)=ρ​\(x0−x\),x^\{0\}\-\\phi^\{0\}\(x\)=x^\{0\}\-\(p\_\{01\}\+\\rho x\)=\\rho\(x^\{0\}\-x\),so

x0−ϕ1​\(x\)≤x0−ϕ0​\(x\)ρ\.x^\{0\}\-\\phi^\{1\}\(x\)\\leq\\frac\{x^\{0\}\-\\phi^\{0\}\(x\)\}\{\\rho\}\.Lett=τ0t=\\tau\_\{0\}\. By definition ofτ↑0\\tau\_\{\\uparrow\}^\{0\},ϕt0​\(ϕ0​\(x\)\)\>z\.\\phi\_\{t\}^\{0\}\(\\phi^\{0\}\(x\)\)\>z\.Using the closed formϕt0​\(y\)=x0\+\(y−x0\)​ρt\\phi\_\{t\}^\{0\}\(y\)=x^\{0\}\+\(y\-x^\{0\}\)\\rho^\{t\}, this is equivalent toρt​\(x0−ϕ0​\(x\)\)<x0−z\.\\rho^\{t\}\\bigl\(x^\{0\}\-\\phi^\{0\}\(x\)\\bigr\)<x^\{0\}\-z\.Therefore,

ρt\+1​\(x0−ϕ1​\(x\)\)≤ρt​\(x0−ϕ0​\(x\)\)<x0−z,\\rho^\{t\+1\}\\bigl\(x^\{0\}\-\\phi^\{1\}\(x\)\\bigr\)\\leq\\rho^\{t\}\\bigl\(x^\{0\}\-\\phi^\{0\}\(x\)\\bigr\)<x^\{0\}\-z,which impliesϕt\+10​\(ϕ1​\(x\)\)\>z\.\\phi\_\{t\+1\}^\{0\}\(\\phi^\{1\}\(x\)\)\>z\.Henceτ1≤t\+1=τ0\+1\\tau\_\{1\}\\leq t\+1=\\tau\_\{0\}\+1, proving \([75](https://arxiv.org/html/2606.11192#A1.E75)\)\.

Since0<β<10<\\beta<1, \([75](https://arxiv.org/html/2606.11192#A1.E75)\) impliesβτ1≥βτ0\+1\.\\beta^\{\\tau\_\{1\}\}\\geq\\beta^\{\\tau\_\{0\}\+1\}\.Because1−κ​x\>01\-\\kappa x\>0, substituting this into the formula forg​\(x,z\)g\(x,z\)gives

g​\(x,z\)\\displaystyle g\(x,z\)≥1\+β1−β​\(κ​x\+\(1−κ​x\)​βτ0\+1−βτ0\)=1\+β1−β​\(κ​x​\(1−βτ0\+1\)−\(1−β\)​βτ0\)\\displaystyle\\geq 1\+\\frac\{\\beta\}\{1\-\\beta\}\\Big\(\\kappa x\+\(1\-\\kappa x\)\\beta^\{\\tau\_\{0\}\+1\}\-\\beta^\{\\tau\_\{0\}\}\\Big\)=1\+\\frac\{\\beta\}\{1\-\\beta\}\\Big\(\\kappa x\(1\-\\beta^\{\\tau\_\{0\}\+1\}\)\-\(1\-\\beta\)\\beta^\{\\tau\_\{0\}\}\\Big\)=1−βτ0\+1\+β​κ​x1−β​\(1−βτ0\+1\)=\(1−βτ0\+1\)​\(1\+β​κ​x1−β\)≥1−βτ0\+1≥1−β\.\\displaystyle=1\-\\beta^\{\\tau\_\{0\}\+1\}\+\\frac\{\\beta\\kappa x\}\{1\-\\beta\}\\bigl\(1\-\\beta^\{\\tau\_\{0\}\+1\}\\bigr\)=\(1\-\\beta^\{\\tau\_\{0\}\+1\}\)\\left\(1\+\\frac\{\\beta\\kappa x\}\{1\-\\beta\}\\right\)\\geq 1\-\\beta^\{\\tau\_\{0\}\+1\}\\geq 1\-\\beta\.
*\(c\)*Fixz∈\[x0,p11\)z\\in\[x^\{0\},p\_\{11\}\)\. Ifx≤zx\\leq z, thenϕ1​\(x\)≤ϕ0​\(x\)≤z\\phi^\{1\}\(x\)\\leq\\phi^\{0\}\(x\)\\leq z, where the second inequality follows fromz≥x0z\\geq x^\{0\}and the monotonicity ofϕ0\\phi^\{0\}\. HenceG​\(ϕ1​\(x\),z\)=G​\(ϕ0​\(x\),z\)=0,G\(\\phi^\{1\}\(x\),z\)=G\(\\phi^\{0\}\(x\),z\)=0,and the one\-step identity \([58](https://arxiv.org/html/2606.11192#S4.E58)\) yieldsg​\(x,z\)=1\+β​κ​x​G​\(p11,z\)≥1\.g\(x,z\)=1\+\\beta\\,\\kappa x\\,G\(p\_\{11\},z\)\\geq 1\.Hence assumex\>zx\>z\. Define

Hz​\(y\)≜κ​y​G​\(p11,z\)\+\(1−κ​y\)​G​\(ϕ1​\(y\),z\),y∈\(z,1\]\.H\_\{z\}\(y\)\\triangleq\\kappa y\\,G\(p\_\{11\},z\)\+\(1\-\\kappa y\)\\,G\(\\phi^\{1\}\(y\),z\),\\qquad y\\in\(z,1\]\.By the evaluation equation \([44](https://arxiv.org/html/2606.11192#S4.E44)\), for everyy\>zy\>z,

G​\(y,z\)=1\+β​Hz​\(y\)\.G\(y,z\)=1\+\\beta\\,H\_\{z\}\(y\)\.
We claim thatHz​\(⋅\)H\_\{z\}\(\\cdot\)is nondecreasing on\(z,1\]\(z,1\]\. Indeed, ify1≤y2y\_\{1\}\\leq y\_\{2\}, then

Hz​\(y2\)−Hz​\(y1\)\\displaystyle H\_\{z\}\(y\_\{2\}\)\-H\_\{z\}\(y\_\{1\}\)=κ​\(y2−y1\)​G​\(p11,z\)\+\(\(1−κ​y2\)​G​\(ϕ1​\(y2\),z\)−\(1−κ​y1\)​G​\(ϕ1​\(y1\),z\)\)\\displaystyle=\\kappa\(y\_\{2\}\-y\_\{1\}\)\\,G\(p\_\{11\},z\)\+\\Big\(\(1\-\\kappa y\_\{2\}\)G\(\\phi^\{1\}\(y\_\{2\}\),z\)\-\(1\-\\kappa y\_\{1\}\)G\(\\phi^\{1\}\(y\_\{1\}\),z\)\\Big\)=κ​\(y2−y1\)​\(G​\(p11,z\)−G​\(ϕ1​\(y2\),z\)\)\+\(1−κ​y1\)​\(G​\(ϕ1​\(y2\),z\)−G​\(ϕ1​\(y1\),z\)\)≥0,\\displaystyle=\\kappa\(y\_\{2\}\-y\_\{1\}\)\\Bigl\(G\(p\_\{11\},z\)\-G\(\\phi^\{1\}\(y\_\{2\}\),z\)\\Bigr\)\+\(1\-\\kappa y\_\{1\}\)\\Bigl\(G\(\\phi^\{1\}\(y\_\{2\}\),z\)\-G\(\\phi^\{1\}\(y\_\{1\}\),z\)\\Bigr\)\\geq 0,becauseG​\(⋅,z\)G\(\\cdot,z\)is nondecreasing on𝒳\\mathscr\{X\}by Lemma[4\.13](https://arxiv.org/html/2606.11192#S4.Thmtheorem13),ϕ1\\phi^\{1\}is increasing, andϕ1​\(y2\)≤p11\\phi^\{1\}\(y\_\{2\}\)\\leq p\_\{11\}\.

Now, sincez≥x0z\\geq x^\{0\}andx\>zx\>z, we haveϕ0​\(x\)≤x\\phi^\{0\}\(x\)\\leq x\. Ifϕ0​\(x\)≤z\\phi^\{0\}\(x\)\\leq z, thenG​\(ϕ0​\(x\),z\)=0G\(\\phi^\{0\}\(x\),z\)=0, and the one\-step identity \([58](https://arxiv.org/html/2606.11192#S4.E58)\) yieldsg​\(x,z\)=1\+β​Hz​\(x\)≥1\.g\(x,z\)=1\+\\beta\\,H\_\{z\}\(x\)\\geq 1\.If insteadϕ0​\(x\)\>z\\phi^\{0\}\(x\)\>z, then by \([44](https://arxiv.org/html/2606.11192#S4.E44)\),

G​\(ϕ0​\(x\),z\)=1\+β​Hz​\(ϕ0​\(x\)\)≤1\+β​Hz​\(x\),G\(\\phi^\{0\}\(x\),z\)=1\+\\beta\\,H\_\{z\}\(\\phi^\{0\}\(x\)\)\\leq 1\+\\beta\\,H\_\{z\}\(x\),becauseHzH\_\{z\}is nondecreasing andϕ0​\(x\)≤x\\phi^\{0\}\(x\)\\leq x\. Substituting this bound into \([58](https://arxiv.org/html/2606.11192#S4.E58)\), we obtain

g​\(x,z\)\\displaystyle g\(x,z\)=1\+β​\(Hz​\(x\)−G​\(ϕ0​\(x\),z\)\)≥1\+β​\(Hz​\(x\)−1−β​Hz​\(x\)\)=1−β\+β​\(1−β\)​Hz​\(x\)\.\\displaystyle=1\+\\beta\\Bigl\(H\_\{z\}\(x\)\-G\(\\phi^\{0\}\(x\),z\)\\Bigr\)\\geq 1\+\\beta\\Bigl\(H\_\{z\}\(x\)\-1\-\\beta H\_\{z\}\(x\)\\Bigr\)=1\-\\beta\+\\beta\(1\-\\beta\)H\_\{z\}\(x\)\.SinceG​\(⋅,z\)≥0G\(\\cdot,z\)\\geq 0, we haveHz​\(x\)≥0H\_\{z\}\(x\)\\geq 0, and thereforeg​\(x,z\)≥1−β\.g\(x,z\)\\geq 1\-\\beta\.∎

###### Proof of Lemma[4\.18](https://arxiv.org/html/2606.11192#S4.Thmtheorem18)\.

For everyx∈𝒳x\\in\\mathscr\{X\}, letD​\(x\)≜1−β​Θ~​\(p11,x\)\.D\(x\)\\triangleq 1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\.Applying Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(c\) withz=xz=x, and usingF~​\(p11,x\)=r​Θ~​\(p11,x\)\\widetilde\{F\}\(p\_\{11\},x\)=r\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)from Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(a\), we obtain

F​\(p11,x\)=r​Θ~​\(p11,x\)D​\(x\),G​\(p11,x\)=G~​\(p11,x\)D​\(x\)\.F\(p\_\{11\},x\)=\\frac\{r\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{D\(x\)\},\\qquad G\(p\_\{11\},x\)=\\frac\{\\widetilde\{G\}\(p\_\{11\},x\)\}\{D\(x\)\}\.Moreover, Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(c\) gives0≤β​Θ~​\(p11,x\)<1,0\\leq\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)<1,henceD​\(x\)\>0D\(x\)\>0\.

Now apply Lemma[4\.14](https://arxiv.org/html/2606.11192#S4.Thmtheorem14)withz=xz=x:

f​\(x,x\)=f~​\(x,x\)\+β​θ~​\(x,x\)​F​\(p11,x\),g​\(x,x\)=g~​\(x,x\)\+β​θ~​\(x,x\)​G​\(p11,x\)\.f\(x,x\)=\\tilde\{f\}\(x,x\)\+\\beta\\,\\tilde\{\\theta\}\(x,x\)\\,F\(p\_\{11\},x\),\\qquad g\(x,x\)=\\tilde\{g\}\(x,x\)\+\\beta\\,\\tilde\{\\theta\}\(x,x\)\\,G\(p\_\{11\},x\)\.Substituting the preceding formulas into the first identity gives

f​\(x,x\)=f~​\(x,x\)\+β​θ~​\(x,x\)​r​Θ~​\(p11,x\)D​\(x\)\.f\(x,x\)=\\tilde\{f\}\(x,x\)\+\\beta\\,\\tilde\{\\theta\}\(x,x\)\\,\\frac\{r\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{D\(x\)\}\.Using Lemma[4\.15](https://arxiv.org/html/2606.11192#S4.Thmtheorem15), namelyf~​\(x,x\)=r​θ~​\(x,x\)\\tilde\{f\}\(x,x\)=r\\,\\tilde\{\\theta\}\(x,x\), this becomes

f​\(x,x\)=r​θ~​\(x,x\)​\(1\+β​Θ~​\(p11,x\)D​\(x\)\)\.f\(x,x\)=r\\,\\tilde\{\\theta\}\(x,x\)\\left\(1\+\\frac\{\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{D\(x\)\}\\right\)\.SinceD​\(x\)=1−β​Θ~​\(p11,x\)D\(x\)=1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\), we have

1\+β​Θ~​\(p11,x\)D​\(x\)=D​\(x\)\+β​Θ~​\(p11,x\)D​\(x\)=1D​\(x\),1\+\\frac\{\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{D\(x\)\}=\\frac\{D\(x\)\+\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{D\(x\)\}=\\frac\{1\}\{D\(x\)\},and therefore

f​\(x,x\)=r​θ~​\(x,x\)D​\(x\)=f~​\(x,x\)D​\(x\),f\(x,x\)=\\frac\{r\\,\\tilde\{\\theta\}\(x,x\)\}\{D\(x\)\}=\\frac\{\\tilde\{f\}\(x,x\)\}\{D\(x\)\},which proves \([62](https://arxiv.org/html/2606.11192#S4.E62)\)\. Similarly,

g​\(x,x\)=g~​\(x,x\)\+β​θ~​\(x,x\)​G~​\(p11,x\)D​\(x\)=g~​\(x,x\)\+βr​f~​\(x,x\)D​\(x\)​G~​\(p11,x\),g\(x,x\)=\\tilde\{g\}\(x,x\)\+\\beta\\,\\tilde\{\\theta\}\(x,x\)\\,\\frac\{\\widetilde\{G\}\(p\_\{11\},x\)\}\{D\(x\)\}=\\tilde\{g\}\(x,x\)\+\\frac\{\\beta\}\{r\}\\,\\frac\{\\tilde\{f\}\(x,x\)\}\{D\(x\)\}\\,\\widetilde\{G\}\(p\_\{11\},x\),which proves \([63](https://arxiv.org/html/2606.11192#S4.E63)\)\. Finally, substituting \([62](https://arxiv.org/html/2606.11192#S4.E62)\) and \([63](https://arxiv.org/html/2606.11192#S4.E63)\) intom​\(x\)≜f​\(x,x\)/g​\(x,x\)m\(x\)\\triangleq\{f\(x,x\)\}/\{g\(x,x\)\}yields \([64](https://arxiv.org/html/2606.11192#S4.E64)\)\. ∎

###### Proof of Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\.

*\(a\)*Letz=xz=x\. If0≤x<p010\\leq x<p\_\{01\}, thenz=x<p01z=x<p\_\{01\}, so Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(a\) givesg​\(x,x\)=1,f​\(x,x\)=r​κ​x,g\(x,x\)=1,\\kern 5\.0ptf\(x,x\)=r\\kappa x,hencem​\(x\)=r​κ​xm\(x\)=r\\kappa x\. Ifp01≤x<x1p\_\{01\}\\leq x<x^\{1\}, thenϕ0​\(x\)\>x\\phi^\{0\}\(x\)\>x\(sincex<x0x<x^\{0\}\) andϕ1​\(x\)\>x\\phi^\{1\}\(x\)\>x\(sincex<x1x<x^\{1\}\)\. Thus in Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(b\) withz=xz=xone hasτ0=τ↑0​\(ϕ0​\(x\),x\)=0,τ1=τ↑0​\(ϕ1​\(x\),x\)=0,\\tau\_\{0\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{0\}\(x\),x\)=0,\\kern 5\.0pt\\tau\_\{1\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{1\}\(x\),x\)=0,so

g​\(x,x\)=1\+β1−β​\(κ​x\+\(1−κ​x\)−1\)=1\.g\(x,x\)=1\+\\frac\{\\beta\}\{1\-\\beta\}\\bigl\(\\kappa x\+\(1\-\\kappa x\)\-1\\bigr\)=1\.Moreover, since

F\+​\(u\)≜r​κ​\[x01−β−x0−u1−β​ρ\]F^\{\+\}\(u\)\\triangleq r\\kappa\\left\[\\frac\{x^\{0\}\}\{1\-\\beta\}\-\\frac\{x^\{0\}\-u\}\{1\-\\beta\\rho\}\\right\]is affine inuu, andϕ0​\(x\)=κ​x​p11\+\(1−κ​x\)​ϕ1​\(x\)\\phi^\{0\}\(x\)=\\kappa x\\,p\_\{11\}\+\(1\-\\kappa x\)\\phi^\{1\}\(x\)by the one\-step posterior\-mean identity, the bracketed term in Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(b\) cancels, yieldingf​\(x,x\)=r​κ​x\.f\(x,x\)=r\\kappa x\.Hencem​\(x\)=r​κ​xm\(x\)=r\\kappa x\.

Finally, atx=x1x=x^\{1\}, Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(b\) givesτ0=τ↑0​\(ϕ0​\(x1\),x1\)=0,τ1=τ↑0​\(ϕ1​\(x1\),x1\)=1,\\tau\_\{0\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{0\}\(x^\{1\}\),x^\{1\}\)=0,\\kern 5\.0pt\\tau\_\{1\}=\\tau\_\{\\uparrow\}^\{0\}\(\\phi^\{1\}\(x^\{1\}\),x^\{1\}\)=1,and thereforeg​\(x1,x1\)=1−β\+β​κ​x1\.g\(x^\{1\},x^\{1\}\)=1\-\\beta\+\\beta\\kappa x^\{1\}\.Similarly,

f​\(x,x\)−r​κ​x​g​\(x,x\)=−β​κ​r1−β​ρ​\[κ​x2−\(1−ρ\+κ​p11\)​x\+p01\]\.f\(x,x\)\-r\\kappa x\\,g\(x,x\)=\-\\frac\{\\beta\\kappa r\}\{1\-\\beta\\rho\}\\,\\Bigl\[\\kappa x^\{2\}\-\(1\-\\rho\+\\kappa p\_\{11\}\)x\+p\_\{01\}\\Bigr\]\.Evaluating atx=x1x=x^\{1\}, the bracket vanishes becausex1x^\{1\}is the fixed point ofϕ1\\phi^\{1\}, i\.e\., the root of \([24](https://arxiv.org/html/2606.11192#S4.E24)\)\. Hencef​\(x1,x1\)=r​κ​x1​g​\(x1,x1\)f\(x^\{1\},x^\{1\}\)=r\\kappa x^\{1\}\\,g\(x^\{1\},x^\{1\}\), som​\(x1\)=r​κ​x1m\(x^\{1\}\)=r\\kappa x^\{1\}\. This proves part \(a\)\.

*\(b\)*Fixx∈\[x0,p11\]x\\in\[x^\{0\},p\_\{11\}\]and setz=xz=x\. Sincex≥x0x\\geq x^\{0\}, we haveϕ0​\(x\)≤x\\phi^\{0\}\(x\)\\leq x, and sincex≥x0\>x1x\\geq x^\{0\}\>x^\{1\}, we also haveϕ1​\(x\)<x\\phi^\{1\}\(x\)<x\. Hence, after either continuation stateϕ0​\(x\)\\phi^\{0\}\(x\)orϕ1​\(x\)\\phi^\{1\}\(x\), thexx\-threshold policy is passive forever\. Therefore,

F~​\(ϕ0​\(x\),x\)=F~​\(ϕ1​\(x\),x\)=0,G~​\(ϕ0​\(x\),x\)=G~​\(ϕ1​\(x\),x\)=0,Θ~​\(ϕ0​\(x\),x\)=Θ~​\(ϕ1​\(x\),x\)=0\.\\widetilde\{F\}\(\\phi^\{0\}\(x\),x\)=\\widetilde\{F\}\(\\phi^\{1\}\(x\),x\)=0,\\qquad\\widetilde\{G\}\(\\phi^\{0\}\(x\),x\)=\\widetilde\{G\}\(\\phi^\{1\}\(x\),x\)=0,\\qquad\\widetilde\{\\Theta\}\(\\phi^\{0\}\(x\),x\)=\\widetilde\{\\Theta\}\(\\phi^\{1\}\(x\),x\)=0\.By \([51](https://arxiv.org/html/2606.11192#S4.E51)\)–\([53](https://arxiv.org/html/2606.11192#S4.E53)\), it follows thatf~​\(x,x\)=r​κ​x,g~​\(x,x\)=1,θ~​\(x,x\)=κ​x\.\\tilde\{f\}\(x,x\)=r\\kappa x,\\kern 5\.0pt\\tilde\{g\}\(x,x\)=1,\\kern 5\.0pt\\tilde\{\\theta\}\(x,x\)=\\kappa x\.Hence, by Lemma[4\.14](https://arxiv.org/html/2606.11192#S4.Thmtheorem14),

f​\(x,x\)=r​κ​x\+β​κ​x​F​\(p11,x\),g​\(x,x\)=1\+β​κ​x​G​\(p11,x\)\.f\(x,x\)=r\\kappa x\+\\beta\\kappa x\\,F\(p\_\{11\},x\),\\qquad g\(x,x\)=1\+\\beta\\kappa x\\,G\(p\_\{11\},x\)\.
Under thexx\-threshold policy started fromp11p\_\{11\}, the no\-ACK skeleton remains active up to, but not including, timeτ↓1​\(p11,x\),\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x\),and is passive thereafter\. Therefore \([65](https://arxiv.org/html/2606.11192#S4.E65)\)–\([66](https://arxiv.org/html/2606.11192#S4.E66)\) hold\. Applying Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(c\) gives

F​\(p11,x\)=r​Θ~​\(p11,x\)1−β​Θ~​\(p11,x\),G​\(p11,x\)=G~​\(p11,x\)1−β​Θ~​\(p11,x\)\.F\(p\_\{11\},x\)=\\frac\{r\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\},\\qquad G\(p\_\{11\},x\)=\\frac\{\\widetilde\{G\}\(p\_\{11\},x\)\}\{1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\}\.Substituting these into the preceding identities yields \([67](https://arxiv.org/html/2606.11192#S4.E67)\) and \([68](https://arxiv.org/html/2606.11192#S4.E68)\), and taking their ratio gives \([69](https://arxiv.org/html/2606.11192#S4.E69)\)\. Finally, substituting \([65](https://arxiv.org/html/2606.11192#S4.E65)\)–\([66](https://arxiv.org/html/2606.11192#S4.E66)\) into \([69](https://arxiv.org/html/2606.11192#S4.E69)\) gives \([70](https://arxiv.org/html/2606.11192#S4.E70)\)\.

*\(c\)*Ifx∈\[p11,1\]x\\in\[p\_\{11\},1\], thenz=x≥p11z=x\\geq p\_\{11\}, so by Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\(e\)g​\(x,x\)=1,f​\(x,x\)=r​κ​x\.g\(x,x\)=1,f\(x,x\)=r\\kappa x\.Hencem​\(x\)=r​κ​xm\(x\)=r\\kappa x\. ∎

###### Proof of Proposition[4\.20](https://arxiv.org/html/2606.11192#S4.Thmtheorem20)\.

*\(a\)*This is Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\(a\)\.

*\(b\)*Fixn∈\{1,…,N0\}n\\in\\\{1,\\dots,N\_\{0\}\\\}andx∈Jnx\\in J\_\{n\}\. By construction,un≤x<un−1\.u\_\{n\}\\leq x<u\_\{n\-1\}\.Forn<N0n<N\_\{0\}this is immediate from the definition ofJnJ\_\{n\}\. Forn=N0n=N\_\{0\}, it follows fromJN0=\[x0,uN0−1\)and​uN0≤x0<uN0−1,J\_\{N\_\{0\}\}=\[x^\{0\},u\_\{N\_\{0\}\-1\}\)\\qquad\\text\{and\}\\kern 5\.0ptu\_\{N\_\{0\}\}\\leq x^\{0\}<u\_\{N\_\{0\}\-1\},the latter being equivalent toN0=τ↓1​\(p11,x0\)N\_\{0\}=\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x^\{0\}\)\. Henceτ↓1​\(p11,x\)=n\.\\tau\_\{\\downarrow\}^\{1\}\(p\_\{11\},x\)=n\.Therefore, Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\(b\) givesG~​\(p11,x\)=An,Θ~​\(p11,x\)=κ​Bn\.\\widetilde\{G\}\(p\_\{11\},x\)=A\_\{n\},\\kern 5\.0pt\\widetilde\{\\Theta\}\(p\_\{11\},x\)=\\kappa B\_\{n\}\.Substituting these identities into \([69](https://arxiv.org/html/2606.11192#S4.E69)\) yields \([71](https://arxiv.org/html/2606.11192#S4.E71)\)\.

Now differentiate \([71](https://arxiv.org/html/2606.11192#S4.E71)\) onJnJ\_\{n\}:

m′​\(x\)=r​κ​\(1−β​κ​Bn\)\(1\+β​κ​\(An​x−Bn\)\)2\.m^\{\\prime\}\(x\)=\\frac\{r\\kappa\\bigl\(1\-\\beta\\kappa B\_\{n\}\\bigr\)\}\{\\bigl\(1\+\\beta\\kappa\(A\_\{n\}x\-B\_\{n\}\)\\bigr\)^\{2\}\}\.Since, onJnJ\_\{n\},1−β​κ​Bn=1−β​Θ~​\(p11,x\)\>01\-\\beta\\kappa B\_\{n\}=1\-\\beta\\,\\widetilde\{\\Theta\}\(p\_\{11\},x\)\>0by Lemma[4\.8](https://arxiv.org/html/2606.11192#S4.Thmtheorem8)\(c\), it follows thatm′​\(x\)\>0m^\{\\prime\}\(x\)\>0onJnJ\_\{n\}\. Thusm​\(⋅\)m\(\\cdot\)is continuous and increasing on eachJnJ\_\{n\}\.

*\(c\)*Let1≤n≤N0−11\\leq n\\leq N\_\{0\}\-1\. SinceAn\+1=An\+βn​Γn11,Bn\+1=Bn\+βn​Γn11​un,A\_\{n\+1\}=A\_\{n\}\+\\beta^\{n\}\\Gamma\_\{n\}^\{11\},\\kern 5\.0ptB\_\{n\+1\}=B\_\{n\}\+\\beta^\{n\}\\Gamma\_\{n\}^\{11\}u\_\{n\},we obtain1\+β​κ​\(An\+1​un−Bn\+1\)=1\+β​κ​\(An​un−Bn\)\.1\+\\beta\\kappa\(A\_\{n\+1\}u\_\{n\}\-B\_\{n\+1\}\)=1\+\\beta\\kappa\(A\_\{n\}u\_\{n\}\-B\_\{n\}\)\.Hence the left and right formulas in \([71](https://arxiv.org/html/2606.11192#S4.E71)\) match atx=unx=u\_\{n\}, sommis continuous at eachunu\_\{n\}\.

Atx=p11x=p\_\{11\}, Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\(c\) givesm​\(p11\)=r​κ​p11\.m\(p\_\{11\}\)=r\\kappa p\_\{11\}\.Also, sinceu0=p11u\_\{0\}=p\_\{11\}, the interval immediately belowp11p\_\{11\}isJ1=\[max⁡\{x0,u1\},p11\)J\_\{1\}=\[\\max\\\{x^\{0\},u\_\{1\}\\\},p\_\{11\}\), and onJ1J\_\{1\},A1=1,B1=p11\.A\_\{1\}=1,\\qquad B\_\{1\}=p\_\{11\}\.Thus

m​\(x\)=r​κ​x1\+β​κ​\(x−p11\),x∈J1,m\(x\)=\\frac\{r\\kappa x\}\{1\+\\beta\\kappa\(x\-p\_\{11\}\)\},\\qquad x\\in J\_\{1\},solimx↗p11m​\(x\)=r​κ​p11=m​\(p11\)\.\\lim\_\{x\\nearrow p\_\{11\}\}m\(x\)=r\\kappa p\_\{11\}=m\(p\_\{11\}\)\.Thereforemmis continuous atx=p11x=p\_\{11\}\.

*\(d\)*This is Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)\(c\)\. ∎

## Appendix BIllustrative instance and diagnostic plots

To illustrate the preceding analytical and computational results, we consider the parameter instance

p01=0\.25,ρ=0\.6,κ=0\.8,β=0\.95,r=1\.p\_\{01\}=0\.25,\\qquad\\rho=0\.6,\\qquad\\kappa=0\.8,\\qquad\\beta=0\.95,\\qquad r=1\.For this instance,p11=p01\+ρ=0\.85,x0=p01/1−ρ=0\.625,p\_\{11\}=p\_\{01\}\+\\rho=0\.85,\\kern 5\.0ptx^\{0\}=\{p\_\{01\}\}/\{1\-\\rho\}=0\.625,and the active no\-ACK fixed point isx1≈0\.2967\.x^\{1\}\\approx 0\.2967\.Thus the analytically unresolved threshold region is the intermediate intervalx1<z<x0,x^\{1\}<z<x^\{0\},i\.e\.,z∈\(0\.2967,0\.625\),z\\in\(0\.2967,\\,0\.625\),while the complementary regions are covered by the tractable\-regime analysis above\.

Figures[1](https://arxiv.org/html/2606.11192#A2.F1)–[4](https://arxiv.org/html/2606.11192#A2.F4)display, for a representative fixed beliefxx, the threshold dependence of the pre\-ACK metricsF~​\(x,z\)\\widetilde\{F\}\(x,z\),G~​\(x,z\)\\widetilde\{G\}\(x,z\),Θ~​\(x,z\)\\widetilde\{\\Theta\}\(x,z\), the reward/work metricsF​\(x,z\)F\(x,z\),G​\(x,z\)G\(x,z\), and the marginal metricsf~​\(x,z\)\\tilde\{f\}\(x,z\),g~​\(x,z\)\\tilde\{g\}\(x,z\),f​\(x,z\)f\(x,z\), andg​\(x,z\)g\(x,z\)\. The plots are consistent with the tractable\-regime theory\. In particular, outside the intervalx1<z<x0x^\{1\}<z<x^\{0\}, the curves exhibit the right\-continuous stepwise behavior predicted by Propositions[4\.11](https://arxiv.org/html/2606.11192#S4.Thmtheorem11),[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12), and[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16), with jumps only at threshold values where the underlying skeleton itinerary changes\. Within the intermediate regimex1<z<x0x^\{1\}<z<x^\{0\}, the plots display a much finer staircase structure\. This is consistent with the Christoffel–Sturmian itinerary organization of Theorem[C\.3](https://arxiv.org/html/2606.11192#A3.Thmtheorem3)and Corollary[C\.4](https://arxiv.org/html/2606.11192#A3.Thmtheorem4), even though the corresponding metric consequences are not derived here\.

Becauser=1r=1in this instance, Lemma[4\.15](https://arxiv.org/html/2606.11192#S4.Thmtheorem15)impliesf~​\(x,z\)=θ~​\(x,z\),\\tilde\{f\}\(x,z\)=\\tilde\{\\theta\}\(x,z\),so Figure[3](https://arxiv.org/html/2606.11192#A2.F3)provides a useful consistency check on the pre\-ACK marginal quantities\. Likewise, Figures[4](https://arxiv.org/html/2606.11192#A2.F4)and[5](https://arxiv.org/html/2606.11192#A2.F5)indicate that the marginal work metric remains positive throughout the plotted range, in agreement with the proved lower bounds on the tractable regimes and with the broader numerical evidence reported later\.

![Refer to caption](https://arxiv.org/html/2606.11192v1/x1.png)Figure 1:Pre\-ACK reward and work metricsF~​\(x,z\)\\widetilde\{F\}\(x,z\)andG~​\(x,z\)\\widetilde\{G\}\(x,z\)vs\.zzfor fixed beliefxx\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x2.png)Figure 2:Reward and work metricsF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)vs\.zzfor fixed beliefxx\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x3.png)Figure 3:Marginal metricsf~​\(x,z\)\\tilde\{f\}\(x,z\)andg~​\(x,z\)\\tilde\{g\}\(x,z\)vs\.zzfor fixed beliefxx\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x4.png)Figure 4:Marginal metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)vs\.zzfor fixed beliefxx\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x5.png)Figure 5:Marginal metricsf​\(x,x\)f\(x,x\)andg​\(x,x\)g\(x,x\)vs\.xx\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x6.png)Figure 6:MP indexm​\(x\)m\(x\)vs\.xx\.Figures[5](https://arxiv.org/html/2606.11192#A2.F5)and[6](https://arxiv.org/html/2606.11192#A2.F6)show the diagonal quantitiesf​\(x,x\)f\(x,x\),g​\(x,x\)g\(x,x\), and the MP indexm​\(x\)m\(x\)\. In the low\- and high\-belief regions, the plots agree with the exact formulas in Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19), namelym​\(x\)=r​κ​x=0\.8​xm\(x\)=r\\kappa x=0\.8xforx∈\[0,x1\]∪\[p11,1\]\.x\\in\[0,x^\{1\}\]\\cup\[p\_\{11\},1\]\.On the interval\[x0,p11\]\[x^\{0\},p\_\{11\}\], the plot agrees with the finite\-sum representation established in the same proposition\. In the remaining intervalx1<x<x0x^\{1\}<x<x^\{0\}, where a complete analytical treatment is left for future work, the graph ofm​\(x\)m\(x\)appears continuous and nondecreasing, which is precisely the behavior required for \(PCLI2\)\.

Overall, this illustrative instance highlights the main message of the paper\. The tractable\-regime analysis accurately captures the coarse transitions of the metrics and the MP index, while the intermediate regime exhibits a rich but highly structured staircase behavior that strongly suggests that the same underlying symbolic organization persists at the metric level\. This motivates the broader numerical investigation of Section[5](https://arxiv.org/html/2606.11192#S5)\.

## Appendix CItinerary organization via words

This appendix records the symbolic\-dynamics material underlying the discussion in Section[4\.7](https://arxiv.org/html/2606.11192#S4.SS7)\. It describes the organization of no\-ACK threshold itineraries in the intermediate regime in terms of binary words, and states the corresponding Christoffel–Sturmian structure inherited from the maps\-with\-gaps theory ofDance and Silander\[[2019](https://arxiv.org/html/2606.11192#bib.bib11)\]\.

A key simplification is that, for a fixed thresholdzz, the no\-ACK skeleton is fully determined by the binary decisions “active” and “passive” along the trajectory\. We therefore encode skeleton itineraries by words over the alphabet𝒜≜\{0,1\},\\mathscr\{A\}\\triangleq\\\{0,1\\\},where11denotes an active step \(so the update isϕ1\\phi^\{1\}\) and0denotes a passive step \(so the update isϕ0\\phi^\{0\}\)\. This converts a threshold\-driven piecewise iteration into a composition of the two deterministic mapsϕ0\\phi^\{0\}andϕ1\\phi^\{1\}, and makes it possible to compare trajectories through the combinatorial structure of the associated words\.

A finite word is a stringw=w0​⋯​wℓ−1∈𝒜ℓw=w\_\{0\}\\cdots w\_\{\\ell\-1\}\\in\\mathscr\{A\}^\{\\ell\}of length\|w\|=ℓ<∞\|w\|=\\ell<\\infty, while an infinite word isw=w0​w1​w2​⋯∈𝒜ℕw=w\_\{0\}w\_\{1\}w\_\{2\}\\cdots\\in\\mathscr\{A\}^\{\\mathbb\{N\}\}\. We denote by∅\\emptysetthe empty word\. For a finite wordww, we write\|w\|1\|w\|\_\{1\}and\|w\|0\|w\|\_\{0\}for the number of11’s and0’s occurring inww, respectively\. Ifuuandvvare words, their concatenation is denotedu​vuv\. We say thatuuis a prefix ofwwifw=u​vw=uvfor some \(possibly empty\) wordvv\. Form∈ℤ\+m\\in\\mathbb\{Z\}\_\{\+\}, we denote bywmw^\{m\}themm\-fold concatenation of a finite wordww\(withw0=∅w^\{0\}=\\emptyset\), and letw∞w^\{\\infty\}denote infinite periodic repetition\. For a wordwwand0≤i≤j≤\|w\|−10\\leq i\\leq j\\leq\|w\|\-1, we writewi:j≜wi​⋯​wj\.w\_\{i:j\}\\triangleq w\_\{i\}\\cdots w\_\{j\}\.

For a wordw=w0​⋯​wt−1w=w\_\{0\}\\cdots w\_\{t\-1\}, we denote byϕw\\phi^\{w\}the corresponding composition of the deterministic belief\-update mapsϕ0\\phi^\{0\}andϕ1\\phi^\{1\},ϕw≜ϕwt−1∘⋯∘ϕw0,\\phi^\{w\}\\triangleq\\phi^\{w\_\{t\-1\}\}\\circ\\cdots\\circ\\phi^\{w\_\{0\}\},with the conventionϕ∅​\(x\)≡x\.\\phi^\{\\emptyset\}\(x\)\\equiv x\.Ifw=σ~​\(x,z\)w=\\widetilde\{\\sigma\}\(x,z\)is the infinite skeleton itinerary, then its length\-ttprefix isw0:t−1=A~0​\(x,z\)​⋯​A~t−1​\(x,z\),t≥1,w\_\{0:t\-1\}=\\widetilde\{A\}\_\{0\}\(x,z\)\\cdots\\widetilde\{A\}\_\{t\-1\}\(x,z\),\\kern 5\.0ptt\\geq 1,andX~t​\(x,z\)=φt​\(x,z\)=ϕw0:t−1​\(x\),t≥1\.\\widetilde\{X\}\_\{t\}\(x,z\)=\\varphi\_\{t\}\(x,z\)=\\phi^\{w\_\{0:t\-1\}\}\(x\),\\kern 5\.0ptt\\geq 1\.

Fixx∈𝒳x\\in\\mathscr\{X\}, and letu=u0​⋯​ut−1∈𝒜tu=u\_\{0\}\\cdots u\_\{t\-1\}\\in\\mathscr\{A\}^\{t\}be a finite word\. Define the associated deterministic belief sequence by

y0u​\(x\)≜x,yk\+1u​\(x\)≜ϕuk​\(yku​\(x\)\),k=0,…,t−1,y\_\{0\}^\{u\}\(x\)\\triangleq x,\\qquad y\_\{k\+1\}^\{u\}\(x\)\\triangleq\\phi^\{u\_\{k\}\}\\\!\\bigl\(y\_\{k\}^\{u\}\(x\)\\bigr\),\\qquad k=0,\\ldots,t\-1,so thatytu​\(x\)=ϕu​\(x\)y\_\{t\}^\{u\}\(x\)=\\phi^\{u\}\(x\)\. Set

Lu​\(x\)≜max⁡\{yku​\(x\):uk=0\},Uu​\(x\)≜min⁡\{yku​\(x\):uk=1\},L^\{u\}\(x\)\\triangleq\\max\\\{y\_\{k\}^\{u\}\(x\):\\,u\_\{k\}=0\\\},\\qquad U^\{u\}\(x\)\\triangleq\\min\\\{y\_\{k\}^\{u\}\(x\):\\,u\_\{k\}=1\\\},with the conventionsmax⁡∅=−∞\\max\\emptyset=\-\\inftyandmin⁡∅=\+∞\\min\\emptyset=\+\\infty, and define the word\-realization interval

Iu​\(x\)≜\[Lu​\(x\),Uu​\(x\)\)⊂ℝ\.I^\{u\}\(x\)\\triangleq\[\\,L^\{u\}\(x\),\\,U^\{u\}\(x\)\\,\)\\subset\\mathbb\{R\}\.\(76\)Thenuuis realized as the length\-ttstrict\-threshold skeleton prefix if and only ifz∈Iu​\(x\)z\\in I^\{u\}\(x\), and in that case

φt​\(x,z\)=ϕu​\(x\)=ytu​\(x\)\.\\varphi\_\{t\}\(x,z\)=\\phi^\{u\}\(x\)=y\_\{t\}^\{u\}\(x\)\.\(77\)
The corresponding statement for infinite words is obtained by intersection over all finite prefixes\.

###### Lemma C\.1\.

Fixx∈𝒳x\\in\\mathscr\{X\}\.

1. \(a\)For everyt∈ℤ\+t\\in\\mathbb\{Z\}\_\{\+\}, the mapz↦φt​\(x,z\)z\\mapsto\\varphi\_\{t\}\(x,z\)is piecewise constant and right\-continuous onℝ\\mathbb\{R\}\. Any point of discontinuity must be a self\-consistency point, in the sense thatz=φk​\(x,z\)​for some​k∈\{0,1,…,t−1\}\.z=\\varphi\_\{k\}\(x,z\)\\kern 5\.0pt\\text\{for some \}k\\in\\\{0,1,\\dots,t\-1\\\}\.
2. \(b\)For every infinite wordw∈𝒜ℕw\\in\\mathscr\{A\}^\{\\mathbb\{N\}\}, if the itinerary intervalIw​\(x\)≜⋂t≥1Iw0:t−1​\(x\)I^\{w\}\(x\)\\triangleq\\bigcap\_\{t\\geq 1\}I^\{w\_\{0:t\-1\}\}\(x\)is nonempty, then for everyz∈Iw​\(x\)z\\in I^\{w\}\(x\)the strict\-threshold skeleton itinerary equalsww, and for everyt≥1t\\geq 1, φt​\(x,z\)=ϕw0:t−1​\(x\)\.\\varphi\_\{t\}\(x,z\)=\\phi^\{\\,w\_\{0:t\-1\}\}\(x\)\.\(78\)In particular, eachφt​\(x,⋅\)\\varphi\_\{t\}\(x,\\cdot\)is constant onIw​\(x\)I^\{w\}\(x\)\.

###### Proof\.

Part \(a\) follows from \([77](https://arxiv.org/html/2606.11192#A3.E77)\), since for fixedttthere are finitely many wordsu∈𝒜tu\\in\\mathscr\{A\}^\{t\}\. Right\-continuity follows because each realization interval is of the form\[L,U\)\[L,U\)\. Part \(b\) follows by applying \([77](https://arxiv.org/html/2606.11192#A3.E77)\) to every prefixw0:t−1w\_\{0:t\-1\}\. ∎

The next lemma formalizes the standard fact, used explicitly in\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), §2\.4, remark after Theorem 12\], that threshold itineraries are preserved by an increasing change of variables\.

###### Lemma C\.2\(Itinerary invariance under increasing conjugacy\)\.

Letγ:I→J\\gamma:I\\to Jbe increasing and setϕ^a=γ∘ϕa∘γ−1,a∈\{0,1\}\.\\hat\{\\phi\}^\{a\}=\\gamma\\circ\\phi^\{a\}\\circ\\gamma^\{\-1\},\\kern 5\.0pta\\in\\\{0,1\\\}\.Then for allx,z∈Ix,z\\in I,

σ~​\(x,z∣ϕ0,ϕ1\)=σ~​\(γ​\(x\),γ​\(z\)∣ϕ^0,ϕ^1\),σ~​\(x,z−∣ϕ0,ϕ1\)=σ~​\(γ​\(x\),γ​\(z\)−∣ϕ^0,ϕ^1\)\.\\widetilde\{\\sigma\}\(x,z\\mid\\phi^\{0\},\\phi^\{1\}\)=\\widetilde\{\\sigma\}\(\\gamma\(x\),\\gamma\(z\)\\mid\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\),\\qquad\\widetilde\{\\sigma\}\(x,z^\{\-\}\\mid\\phi^\{0\},\\phi^\{1\}\)=\\widetilde\{\\sigma\}\(\\gamma\(x\),\\gamma\(z\)^\{\-\}\\mid\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\)\.

###### Proof\.

Let\(xj\)\(x\_\{j\}\)be thezz\-threshold orbit fromxxunder\(ϕ0,ϕ1\)\(\\phi^\{0\},\\phi^\{1\}\), and setyj=γ​\(xj\)y\_\{j\}=\\gamma\(x\_\{j\}\)\. Sinceγ\\gammais increasing,xj\>zx\_\{j\}\>ziffyj\>γ​\(z\)y\_\{j\}\>\\gamma\(z\), andxj≥zx\_\{j\}\\geq ziffyj≥γ​\(z\)y\_\{j\}\\geq\\gamma\(z\)\. Moreover,yj\+1=γ​\(ϕ0​\(xj\)\)=ϕ^0​\(yj\)y\_\{j\+1\}=\\gamma\(\\phi^\{0\}\(x\_\{j\}\)\)=\\hat\{\\phi\}^\{0\}\(y\_\{j\}\)whenyj≤γ​\(z\)y\_\{j\}\\leq\\gamma\(z\)\(resp\.<γ​\(z\)<\\gamma\(z\)\), andyj\+1=γ​\(ϕ1​\(xj\)\)=ϕ^1​\(yj\)y\_\{j\+1\}=\\gamma\(\\phi^\{1\}\(x\_\{j\}\)\)=\\hat\{\\phi\}^\{1\}\(y\_\{j\}\)whenyj\>γ​\(z\)y\_\{j\}\>\\gamma\(z\)\(resp\.≥γ​\(z\)\\geq\\gamma\(z\)\)\. Thus\(yj\)\(y\_\{j\}\)is exactly theγ​\(z\)\\gamma\(z\)\-threshold \(resp\.γ​\(z\)−\\gamma\(z\)^\{\-\}\-threshold\) orbit fromγ​\(x\)\\gamma\(x\)under\(ϕ^0,ϕ^1\)\(\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\), and the induced itinerary letters coincide\. ∎

A word is*balanced*if, for everyℓ≥1\\ell\\geq 1, the number of11’s in any two length\-ℓ\\ellfactors differs by at most one\. A source of balanced words is provided by*mechanical words*: forα∈\(0,1\)\\alpha\\in\(0,1\)andη∈ℝ\\eta\\in\\mathbb\{R\}, the*lower mechanical word*is

\(Lα,η\)j≜⌊\(j\+1\)​α\+η⌋−⌊j​α\+η⌋,j≥0\.\(L\_\{\\alpha,\\eta\}\)\_\{j\}\\triangleq\\big\\lfloor\(j\+1\)\\alpha\+\\eta\\big\\rfloor\-\\big\\lfloor j\\alpha\+\\eta\\big\\rfloor,\\qquad j\\geq 0\.Whenα\\alphais rational, the lower mechanical word is periodic, and its primitive period is, up to cyclic shift, a*Christoffel word*\. Whenα\\alphais irrational, one obtains an aperiodic balanced word, i\.e\., a*Sturmian word*\.

Ifα=m/n∈\(0,1\)\\alpha=m/n\\in\(0,1\)is rational in lowest terms, the corresponding lower Christoffel word isCm/n≜\(Lm/n,0\)0:\(n−1\)\.C\_\{m/n\}\\triangleq\(L\_\{m/n,0\}\)\_\{0:\(n\-1\)\}\.It has the standard palindromic factorizationCm/n=0​p​1,C\_\{m/n\}=0p1,whereppis a \(possibly empty\) palindrome\. Ifα\\alphais irrational, the lower mechanical wordLα,0L\_\{\\alpha,0\}is a Sturmian word\. FollowingDance and Silander\[[2019](https://arxiv.org/html/2606.11192#bib.bib11), Definition 6\], we callCm/nC\_\{m/n\}theℳ\\mathscr\{M\}\-word of rateα=m/n\\alpha=m/nwhenα\\alphais rational, andLα,0L\_\{\\alpha,0\}theℳ\\mathscr\{M\}\-word of rateα\\alphawhenα\\alphais irrational\.

For a finite non\-empty wordww, Assumption[4\.3](https://arxiv.org/html/2606.11192#S4.Thmtheorem3)implies thatϕw\\phi^\{w\}has a unique fixed pointxw∈𝒳x^\{w\}\\in\\mathscr\{X\}\. For a Sturmianℳ\\mathscr\{M\}\-word0​s0s, definexsx^\{s\}via the common limit of fixed points along the associated Christoffel approximants, exactly as in\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), §2\.4\]; existence and coincidence of the two limits follow from\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), Lemma 55\]\.

###### Theorem C\.3\(Threshold itineraries as Christoffel and Sturmian words\)\.

Let0​p​10p1be a Christoffel word and let0​s0sbe a Sturmianℳ\\mathscr\{M\}\-word\. Then the fixed pointsx01​px^\{01p\},x10​px^\{10p\}, andxsx^\{s\}exist in\(0,1\)\(0,1\)\. Moreover, the active\-at\-threshold left\-threshold itineraryz↦σ~​\(z,z−\)z\\mapsto\\widetilde\{\\sigma\}\(z,z^\{\-\}\)is lexicographically nonincreasing onz∈\(0,1\)z\\in\(0,1\)and is given by

σ~​\(z,z−\)=\{1∞,if and only if​z≤x1,\(10​p\)∞,if and only if​z∈\[x01​p,x10​p\],10​s,if and only if​z=xs,10∞,if and only if​z≥x0\.\\widetilde\{\\sigma\}\(z,z^\{\-\}\)=\\begin\{cases\}1^\{\\infty\},&\\text\{if and only if \}z\\leq x^\{1\},\\\\ \(10p\)^\{\\infty\},&\\text\{if and only if \}z\\in\[x^\{01p\},x^\{10p\}\],\\\\ 10s,&\\text\{if and only if \}z=x^\{s\},\\\\ 10^\{\\infty\},&\\text\{if and only if \}z\\geq x^\{0\}\.\\end\{cases\}

###### Proof\.

Letϑ\\varthetabe the logit map and let\(ϕ^0,ϕ^1\)\(\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\)be the conjugated maps from Lemma[4\.7](https://arxiv.org/html/2606.11192#S4.Thmtheorem7)\. By that lemma,\(ϕ^0,ϕ^1\)\(\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\)satisfies Assumption 2 ofDance and Silander\[[2019](https://arxiv.org/html/2606.11192#bib.bib11)\]onℝ\\mathbb\{R\}, with ordered fixed pointsx^1<x^0\\hat\{x\}^\{1\}<\\hat\{x\}^\{0\}\. Therefore\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), Theorem 12\]applies toσ~​\(⋅,⋅−\)\\widetilde\{\\sigma\}\(\\cdot,\\cdot^\{\-\}\)for the conjugated maps\.

By Lemma[C\.2](https://arxiv.org/html/2606.11192#A3.Thmtheorem2), for eachz∈\(0,1\)z\\in\(0,1\),σ~​\(z,z−∣ϕ0,ϕ1\)=σ~​\(ϑ​\(z\),ϑ​\(z\)−∣ϕ^0,ϕ^1\),\\widetilde\{\\sigma\}\(z,z^\{\-\}\\mid\\phi^\{0\},\\phi^\{1\}\)=\\widetilde\{\\sigma\}\(\\vartheta\(z\),\\vartheta\(z\)^\{\-\}\\mid\\hat\{\\phi\}^\{0\},\\hat\{\\phi\}^\{1\}\),and fixed points for word compositions correspond underϑ\\vartheta\. Translating the conclusion of\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), Theorem 12\]back throughϑ\\varthetayields the stated characterization\. ∎

###### Corollary C\.4\(One\-step\-deviation itineraries\)\.

Let0​p​10p1be a Christoffel word and let0​s0sbe a Sturmianℳ\\mathscr\{M\}\-word\. Then the pair of itineraries\(σ~​\(ϕ0​\(z\),z\),σ~​\(ϕ1​\(z\),z\)\)\\bigl\(\\widetilde\{\\sigma\}\(\\phi^\{0\}\(z\),z\),\\ \\widetilde\{\\sigma\}\(\\phi^\{1\}\(z\),z\)\\bigr\)satisfies

\(σ~​\(ϕ0​\(z\),z\),σ~​\(ϕ1​\(z\),z\)\)=\{\(1∞,1∞\),z<x1,\(1∞,01∞\),z=x1,\(\(1​p​0\)∞,\(0​p​1\)∞\),z∈\[x01​p,x10​p\),\(\(1​p​0\)∞,0​p​\(01​p\)∞\),z=x10​p,\(1​s,0​s\),z=xs,\(0∞,0∞\),z≥x0\.\\bigl\(\\widetilde\{\\sigma\}\(\\phi^\{0\}\(z\),z\),\\ \\widetilde\{\\sigma\}\(\\phi^\{1\}\(z\),z\)\\bigr\)=\\begin\{cases\}\(1^\{\\infty\},1^\{\\infty\}\),&z<x^\{1\},\\\\ \(1^\{\\infty\},01^\{\\infty\}\),&z=x^\{1\},\\\\ \(\(1p0\)^\{\\infty\},\(0p1\)^\{\\infty\}\),&z\\in\[x^\{01p\},x^\{10p\}\),\\\\ \(\(1p0\)^\{\\infty\},\\,0p\(01p\)^\{\\infty\}\),&z=x^\{10p\},\\\\ \(1s,0s\),&z=x^\{s\},\\\\ \(0^\{\\infty\},0^\{\\infty\}\),&z\\geq x^\{0\}\.\\end\{cases\}

###### Proof\.

Apply Lemmas[4\.7](https://arxiv.org/html/2606.11192#S4.Thmtheorem7)and[C\.2](https://arxiv.org/html/2606.11192#A3.Thmtheorem2)as in the proof of Theorem[C\.3](https://arxiv.org/html/2606.11192#A3.Thmtheorem3), and then invoke\[Dance and Silander,[2019](https://arxiv.org/html/2606.11192#bib.bib11), Corollary 13\]for the conjugated maps\. ∎

###### Proposition C\.5\(Christoffel\-interval partition of the threshold axis\)\.

Let𝒞\\mathscr\{C\}denote the set of Christoffel words0​p​10p1with rational rate in\(0,1\)\(0,1\)\. For each0​p​1∈𝒞0p1\\in\\mathscr\{C\}, define the Christoffel intervalI0​p​1≜\[x01​p,x10​p\)\.I\_\{0p1\}\\triangleq\[\\,x^\{01p\},\\,x^\{10p\}\\,\)\.Define the Sturmian set𝒮≜\{xs:0​s​is a Sturmianℳ\-word\}\.\\mathscr\{S\}\\triangleq\\\{x^\{s\}:\\ 0s\\text\{ is a Sturmian $\\mathscr\{M\}$\-word\}\\\}\.Then:

1. \(i\)the family\{I0​p​1:0​p​1∈𝒞\}\\\{I\_\{0p1\}:0p1\\in\\mathscr\{C\}\\\}is pairwise disjoint and \(x1,x0\)=\(⨆0​p​1∈𝒞I0​p​1\)⊔𝒮;\(x^\{1\},x^\{0\}\)=\\Big\(\\bigsqcup\_\{0p1\\in\\mathscr\{C\}\}I\_\{0p1\}\\Big\)\\sqcup\\mathscr\{S\};
2. \(ii\)on each Christoffel intervalI0​p​1I\_\{0p1\}, the left\-threshold itinerary at the threshold is constant: σ~​\(z,z−\)=\(10​p\)∞,z∈I0​p​1;\\widetilde\{\\sigma\}\(z,z^\{\-\}\)=\(10p\)^\{\\infty\},\\qquad z\\in I\_\{0p1\};
3. \(iii\)for each Sturmian pointxs∈𝒮x^\{s\}\\in\\mathscr\{S\}, one hasσ~​\(xs,\(xs\)−\)=10​s;\\widetilde\{\\sigma\}\(x^\{s\},\(x^\{s\}\)^\{\-\}\)=10s;
4. \(iv\)the extreme regimes are σ~​\(z,z−\)=1∞⟺z≤x1,σ~​\(z,z−\)=10∞⟺z≥x0\.\\widetilde\{\\sigma\}\(z,z^\{\-\}\)=1^\{\\infty\}\\ \\Longleftrightarrow\\ z\\leq x^\{1\},\\qquad\\widetilde\{\\sigma\}\(z,z^\{\-\}\)=10^\{\\infty\}\\ \\Longleftrightarrow\\ z\\geq x^\{0\}\.

###### Proof\.

Theorem[C\.3](https://arxiv.org/html/2606.11192#A3.Thmtheorem3)gives the threshold\-itinerary classification on the closed intervals\[x01​p,x10​p\]\[x^\{01p\},x^\{10p\}\]\. Passing to the half\-open conventionI0​p​1=\[x01​p,x10​p\)I\_\{0p1\}=\[x^\{01p\},x^\{10p\}\)removes the endpoint overlaps between adjacent Christoffel intervals and yields a disjoint partition of\(x1,x0\)\(x^\{1\},x^\{0\}\), with the remaining threshold values given by the Sturmian pointsxsx^\{s\}\. The itinerary statements in \(ii\)–\(iv\) are then exactly those of Theorem[C\.3](https://arxiv.org/html/2606.11192#A3.Thmtheorem3)\. ∎

Proposition[C\.5](https://arxiv.org/html/2606.11192#A3.Thmtheorem5)shows that the threshold axis in the intermediate regime is partitioned into Christoffel intervals and Sturmian points\. On each Christoffel interval, the active\-at\-threshold itinerary is fixed and periodic, while Corollary[C\.4](https://arxiv.org/html/2606.11192#A3.Thmtheorem4)fixes the corresponding one\-step\-deviation itineraries for the passive\-at\-threshold case\. These are precisely the symbolic inputs needed for a future intervalwise analysis of the pre\-ACK quantitiesF~,G~,Θ~\\widetilde\{F\},\\widetilde\{G\},\\widetilde\{\\Theta\}, the marginal metricsf,gf,g, and hence the MP indexmmon the intermediate regime\.

## Appendix DOutline of the time\-average criterion

The discounted analysis developed above naturally suggests a corresponding long\-run average formulation\. Since the present paper is primarily concerned with the discounted criterion, we only sketch here how the no\-ACK skeleton and renewal machinery extend to the time\-average setting\. A full treatment of the average\-criterion PCL\-indexability framework for real\-state projects is left for future work\.

Fix a thresholdz∈ℝz\\in\\mathbb\{R\}\. LetFβ​\(x,z\)F\_\{\\beta\}\(x,z\),Gβ​\(x,z\)G\_\{\\beta\}\(x,z\),fβ​\(x,z\)f\_\{\\beta\}\(x,z\), andgβ​\(x,z\)g\_\{\\beta\}\(x,z\)denote the discounted reward, work, and marginal metrics introduced above, with the discount factor made explicit\. Motivated by Abelian\-limit considerations, define the corresponding time\-average quantities by

F​\(x,z\)≜limβ↗1\(1−β\)​Fβ​\(x,z\),G​\(x,z\)≜limβ↗1\(1−β\)​Gβ​\(x,z\),F\(x,z\)\\triangleq\\lim\_\{\\beta\\nearrow 1\}\(1\-\\beta\)F\_\{\\beta\}\(x,z\),\\qquad G\(x,z\)\\triangleq\\lim\_\{\\beta\\nearrow 1\}\(1\-\\beta\)G\_\{\\beta\}\(x,z\),\(79\)and

f​\(x,z\)≜limβ↗1fβ​\(x,z\),g​\(x,z\)≜limβ↗1gβ​\(x,z\)\.f\(x,z\)\\triangleq\\lim\_\{\\beta\\nearrow 1\}f\_\{\\beta\}\(x,z\),\\qquad g\(x,z\)\\triangleq\\lim\_\{\\beta\\nearrow 1\}g\_\{\\beta\}\(x,z\)\.\(80\)In this outline we take for granted that these limits exist and thatF​\(x,z\)F\(x,z\)andG​\(x,z\)G\(x,z\)are independent of the initial beliefxx\. This is consistent with the regenerative structure of the model and with the numerical evidence discussed below\. When this independence holds, we simply writeF​\(z\)F\(z\)andG​\(z\)G\(z\)\. The associated average\-criterion MP index is then

m​\(x,z\)≜f​\(x,z\)g​\(x,z\),m​\(x\)≜m​\(x,x\),m\(x,z\)\\triangleq\\frac\{f\(x,z\)\}\{g\(x,z\)\},\\qquad m\(x\)\\triangleq m\(x,x\),\(81\)wheneverg​\(x,z\)\>0g\(x,z\)\>0\.

In the tractable regimes, the discounted formulas of Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12)suggest the followingβ↗1\\beta\\nearrow 1limits\. In the eventually\-always\-active regimesz≤x1z\\leq x^\{1\}, namely cases \(a\)–\(b\) of Proposition[4\.12](https://arxiv.org/html/2606.11192#S4.Thmtheorem12), one is led to

F​\(z\)=r​κ​x0,G​\(z\)=1,z≤x1,F\(z\)=r\\kappa x^\{0\},\\qquad G\(z\)=1,\\qquad z\\leq x^\{1\},\(82\)wherex0=p01/\(1−ρ\)x^\{0\}=p\_\{01\}/\(1\-\\rho\)is the passive fixed point\. Thus, once the threshold is at or belowx1x^\{1\}, the policy is eventually always active, the average work rate is11, and the average reward rate isr​κ​x0r\\kappa x^\{0\}\.

At the other extreme, in the eventually\-passive regimesz≥x0z\\geq x^\{0\}, the threshold policy activates only finitely many times \(or not at all\), so the discounted reward and work totals remainO​\(1\)O\(1\)asβ↗1\\beta\\nearrow 1\. This suggests

F​\(z\)=0,G​\(z\)=0,z≥x0\.F\(z\)=0,\\qquad G\(z\)=0,\\qquad z\\geq x^\{0\}\.\(83\)In particular, \([83](https://arxiv.org/html/2606.11192#A4.E83)\) covers bothx0≤z<p11x^\{0\}\\leq z<p\_\{11\}andz≥p11z\\geq p\_\{11\}\.

Therefore, the only genuinely nontrivial average\-criterion regime is the intermediate stripx1<z<x0,x^\{1\}<z<x^\{0\},which is precisely the regime where the no\-ACK skeleton alternates indefinitely between active and passive phases\.

The limiting marginal metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)may be obtained by passing to the limit in the discounted one\-step formulas of Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\. In the extreme regimesz<p01z<p\_\{01\}andz≥p11z\\geq p\_\{11\}, those formulas remain unchanged:

f​\(x,z\)=r​κ​x,g​\(x,z\)=1,m​\(x,z\)=r​κ​x\.f\(x,z\)=r\\kappa x,\\qquad g\(x,z\)=1,\\qquad m\(x,z\)=r\\kappa x\.\(84\)In the remaining tractable subregimesp01≤z≤x1p\_\{01\}\\leq z\\leq x^\{1\}andx0≤z<p11x^\{0\}\\leq z<p\_\{11\}, explicit closed forms forf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)are suggested by takingβ↗1\\beta\\nearrow 1in the corresponding discounted formulas; they involve the same threshold crossing timesτ↑0\\tau\_\{\\uparrow\}^\{0\}andτ↓1\\tau\_\{\\downarrow\}^\{1\}, and, in thex0≤z<p11x^\{0\}\\leq z<p\_\{11\}regime, finite sums over the active no\-ACK skeleton segment\. Since the present section is only meant as an outline, we do not record those expressions here\.

Forx1<z<x0x^\{1\}<z<x^\{0\}, the average\-criterion metrics can be approached in two natural ways\. The first is the simplest computationally: evaluate the discounted metricsFβ​\(x,z\)F\_\{\\beta\}\(x,z\),Gβ​\(x,z\)G\_\{\\beta\}\(x,z\),fβ​\(x,z\)f\_\{\\beta\}\(x,z\), andgβ​\(x,z\)g\_\{\\beta\}\(x,z\)forβ\\betasufficiently close to11, and then use \([79](https://arxiv.org/html/2606.11192#A4.E79)\)–\([80](https://arxiv.org/html/2606.11192#A4.E80)\) as numerical approximations\. The second is to work directly with the regenerative structure at ACK times\. Writingτack\\tau^\{\\mathrm\{ack\}\}for the first ACK time under thezz\-threshold policy, define the undiscounted cycle quantities

R¯​\(x,z\)\\displaystyle\\overline\{R\}\(x,z\)≜𝔼xz​\[∑t=0τackR​\(X​\(t\),A​\(t\)\)\],W¯​\(x,z\)≜𝔼xz​\[∑t=0τackA​\(t\)\],T¯​\(x,z\)≜𝔼xz​\[τack\+1\]\.\\displaystyle\\triangleq\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}R\(X\(t\),A\(t\)\)\\bigg\],\\quad\\overline\{W\}\(x,z\)\\triangleq\\mathbb\{E\}\_\{x\}^\{z\}\\\!\\bigg\[\\sum\_\{t=0\}^\{\\tau^\{\\mathrm\{ack\}\}\}A\(t\)\\bigg\],\\quad\\overline\{T\}\(x,z\)\\triangleq\\mathbb\{E\}\_\{x\}^\{z\}\\big\[\\tau^\{\\mathrm\{ack\}\}\+1\\big\]\.\(85\)Formally, when the threshold policy regenerates at the post\-ACK statep11p\_\{11\}, one expects the average reward and work rates to satisfy

F​\(z\)=R¯​\(p11,z\)T¯​\(p11,z\),G​\(z\)=W¯​\(p11,z\)T¯​\(p11,z\)\.F\(z\)=\\frac\{\\overline\{R\}\(p\_\{11\},z\)\}\{\\overline\{T\}\(p\_\{11\},z\)\},\\qquad G\(z\)=\\frac\{\\overline\{W\}\(p\_\{11\},z\)\}\{\\overline\{T\}\(p\_\{11\},z\)\}\.\(86\)Likewise, the average marginal metricsf​\(x,z\)f\(x,z\)andg​\(x,z\)g\(x,z\)may be obtained either as limits of the discounted marginals or by combining one\-step deviation formulas with the same regenerative continuation values\. Thus, the discounted skeleton and renewal machinery provide a natural computational route to the time\-average criterion as well\.

The tractable\-regime formulas above strongly suggest that the average\-criterion metrics inherit the same structural properties as their discounted counterparts\. In particular, in the tractable regimes one expects the average\-criterion analogues of \(PCLI1\) and \(PCLI2\) to hold, namely positivity ofg​\(x,z\)g\(x,z\)and monotonicity of the diagonal MP indexm​\(x\)m\(x\)\. For the extreme regimesz<p01z<p\_\{01\}andz≥p11z\\geq p\_\{11\}, this is immediate from \([84](https://arxiv.org/html/2606.11192#A4.E84)\)\. The corresponding statements in the remaining tractable regimes are suggested by the limiting closed forms obtained from Proposition[4\.16](https://arxiv.org/html/2606.11192#S4.Thmtheorem16)\.

This leads naturally to the following conjecture\.

###### Conjecture D\.1\.

The average\-criterion limiting metricsf​\(x,z\)f\(x,z\),g​\(x,z\)g\(x,z\), andm​\(x\)=f​\(x,x\)/g​\(x,x\)m\(x\)=f\(x,x\)/g\(x,x\)satisfy the average\-criterion analogues of\(PCLI1\)and\(PCLI2\)on the full threshold range, including the regimex1<z<x0x^\{1\}<z<x^\{0\}\.

If true, this would yield an average\-criterion MP index for the present real\-state project model and provide the time\-average counterpart of the discounted PCL\-indexability analysis developed in this paper\.

The main point of this section is that the no\-ACK skeleton and renewal decomposition are not specific to the discounted criterion\. They also suggest a natural average\-criterion theory, with explicit tractable\-regime formulas, direct computational schemes in the intermediate regime, and a plausible extension of the PCL\-indexability conditions\. This average\-criterion development is beyond the scope of the present paper, but the discounted results obtained here provide its natural starting point\.

## Appendix ESupplementary experimental plots and implementation details

### E\.1Parameter dependence of the MP index

#### E\.1\.1Dependence of the MP index onβ\\beta

We begin with the dependence of the MP index on the discount factorβ\\beta, using the same base parameter instance as in Appendix[B](https://arxiv.org/html/2606.11192#A2), namelyp01=0\.25,ρ=0\.6,κ=0\.8,r=1\.p\_\{01\}=0\.25,\\rho=0\.6,\\kappa=0\.8,r=1\.Figure[7](https://arxiv.org/html/2606.11192#A5.F7)plots the MP indexm​\(x\)m\(x\)forβ∈\{0\.1,0\.3,0\.5,0\.70,0\.90,0\.95,0\.999\}\.\\beta\\in\\\{0\.1,\\,0\.3,\\,0\.5,\\,0\.70,\\,0\.90,\\,0\.95,\\,0\.999\\\}\.

Several features are apparent\. First, over the displayed range, the MP index appears to be increasing inβ\\beta, and the curves appear to converge to a limiting profile asβ↗1\\beta\\nearrow 1, consistent with the time\-average discussion of Appendix[D](https://arxiv.org/html/2606.11192#A4)\. Second, by Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19), the MP index coincides with the myopic indexm​\(x\)=r​κ​xm\(x\)=r\\kappa xon the low\- and high\-belief regionsx∈\[0,x1\]∪\[p11,1\]\.x\\in\[0,x^\{1\}\]\\cup\[p\_\{11\},1\]\.Accordingly, the visible dependence onβ\\betais concentrated in the intermediate belief region\(x1,p11\)\(x^\{1\},p\_\{11\}\), where future information and restart effects matter most\. The sensitivity toβ\\betais strongest in the nontrivial intervalx1<x<x0x^\{1\}<x<x^\{0\}and persists on\[x0,p11\]\[x^\{0\},p\_\{11\}\]\)\.

![Refer to caption](https://arxiv.org/html/2606.11192v1/x7.png)Figure 7:MP indexm​\(x\)m\(x\)versusxxfor several values ofβ\\beta\.
#### E\.1\.2Dependence of the MP index onκ\\kappa

We next examine the dependence of the MP index on the sensing\-success parameterκ\\kappa, using the same base parameter instance as in Appendix[B](https://arxiv.org/html/2606.11192#A2), withp01=0\.25,ρ=0\.6,β=0\.95,r=1p\_\{01\}=0\.25,\\rho=0\.6,\\beta=0\.95,r=1held fixed\. Figure[8](https://arxiv.org/html/2606.11192#A5.F8)plots the MP indexm​\(x\)m\(x\)forκ∈\{0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.70,0\.8,0\.90,0\.95,0\.999\}\.\\kappa\\in\\\{0\.1,\\,0\.2,\\,0\.3,\\,0\.4,\\,0\.5,\\,0\.6,\\,0\.70,\\,0\.8,\\,0\.90,\\,0\.95,\\,0\.999\\\}\.

Unlike theβ\\beta\-family case, the active no\-ACK fixed pointx1x^\{1\}varies withκ\\kappa, whereas the passive fixed pointx0x^\{0\}and the reset statep11p\_\{11\}remain unchanged\. For this reason, the boundary between the low\-belief myopic region and the intermediate region moves withκ\\kappa, while the pointsx0x^\{0\}andp11p\_\{11\}remain fixed on the horizontal axis\. Accordingly, the defaultxx\-axis ticks in Figure[8](https://arxiv.org/html/2606.11192#A5.F8)show only theseκ\\kappa\-invariant reference points\.

The figure shows that the MP index appears to be increasing inκ\\kappaover the displayed range, and to converge to a limiting profile asκ↗1\\kappa\\nearrow 1\. As in the previous subsection, the most significant changes occur in the nontrivial belief region\. In particular, since Proposition[4\.19](https://arxiv.org/html/2606.11192#S4.Thmtheorem19)givesm​\(x\)=r​κ​xm\(x\)=r\\kappa xforx∈\[0,x1\]∪\[p11,1\],x\\in\[0,x^\{1\}\]\\cup\[p\_\{11\},1\],the dependence onκ\\kappais especially informative on the interval\(x1,p11\)\(x^\{1\},p\_\{11\}\), where bothx1x^\{1\}and the magnitude of the index vary withκ\\kappa\.

![Refer to caption](https://arxiv.org/html/2606.11192v1/x8.png)Figure 8:MP indexm​\(x\)m\(x\)versusxxfor several values ofκ\\kappa\.
#### E\.1\.3Dependence of the MP index onρ\\rho

We next examine the dependence of the MP index on the persistence parameterρ\\rho, using the same base parameter instance as in Appendix[B](https://arxiv.org/html/2606.11192#A2), withp01=0\.25,κ=0\.8,β=0\.95,r=1p\_\{01\}=0\.25,\\kappa=0\.8,\\beta=0\.95,r=1held fixed\. Figure[9](https://arxiv.org/html/2606.11192#A5.F9)plots the MP indexm​\(x\)m\(x\)forρ∈\{0\.001,0\.1,0\.2,0\.3,0\.4,0\.5,0\.6,0\.70,0\.749\}\.\\rho\\in\\\{0\.001,\\,0\.1,\\,0\.2,\\,0\.3,\\,0\.4,\\,0\.5,\\,0\.6,\\,0\.70,\\,0\.749\\\}\.Sincep01=0\.25p\_\{01\}=0\.25is held fixed, feasibility requires0<ρ<1−p01=0\.75,0<\\rho<1\-p\_\{01\}=0\.75,so the largest plotted valueρ=0\.749\\rho=0\.749lies just below the upper admissible boundary\.

Unlike theβ\\beta\- andκ\\kappa\-families, the key internal reference pointsx1,x0,p11=p01\+ρx^\{1\},x^\{0\},p\_\{11\}=p\_\{01\}\+\\rhoall depend onρ\\rho, so they no longer occur at common locations on the horizontal axis\. For this reason, the defaultxx\-axis ticks in Figure[9](https://arxiv.org/html/2606.11192#A5.F9)use onlyρ\\rho\-invariant reference points\.

The dependence ofm​\(x\)m\(x\)onρ\\rhois qualitatively different from its dependence onβ\\betaandκ\\kappa\. The family of curves is not monotone inρ\\rho: for some belief values the index increases withρ\\rho, whereas for others it decreases, and the overall shape of the non\-myopic region changes substantially asρ\\rhovaries\. Figure[10](https://arxiv.org/html/2606.11192#A5.F10), which plotsm​\(x\)m\(x\)as a function ofρ\\rhofor a fixed representative beliefxx, makes this non\-monotone dependence more explicit\.

Asρ\\rhoapproaches the upper feasibility boundary1−p01=0\.751\-p\_\{01\}=0\.75, the MP index appears to approach a limiting profile\. Also, the movement of the transition pointsx1x^\{1\},x0x^\{0\}, andp11p\_\{11\}shows that changes in persistence affect not only the magnitude of the index but also the location and width of the regions where the MP index differs from myopic\.

![Refer to caption](https://arxiv.org/html/2606.11192v1/x9.png)Figure 9:MP indexm​\(x\)m\(x\)versusxxfor several values ofρ\\rho\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x10.png)Figure 10:MP indexm​\(x\)m\(x\)versusρ\\rhofor a fixed beliefxx\.
#### E\.1\.4Dependence of the MP index onp01p\_\{01\}

We next examine the dependence of the MP index on the transition parameterp01p\_\{01\}, using the same base parameter instance as in Appendix[B](https://arxiv.org/html/2606.11192#A2), except thatp01p\_\{01\}is now varied whileρ=0\.6,κ=0\.8,β=0\.95,r=1\\rho=0\.6,\\kappa=0\.8,\\beta=0\.95,r=1are held fixed\. Figure[11](https://arxiv.org/html/2606.11192#A5.F11)plots the MP indexm​\(x\)m\(x\)forp01∈\{0\.001,0\.10,0\.15,0\.20,0\.30,0\.35,0\.399\}\.p\_\{01\}\\in\\\{0\.001,\\,0\.10,\\,0\.15,\\,0\.20,\\,0\.30,\\,0\.35,\\,0\.399\\\}\.

Sinceρ=0\.6\\rho=0\.6is fixed, feasibility requires0<p01<1−ρ=0\.4,0<p\_\{01\}<1\-\\rho=0\.4,so the largest plotted valuep01=0\.399p\_\{01\}=0\.399lies just below the upper admissible boundary\.

As in theρ\\rho\-family case, the key internal reference pointsx1,x0=p01/\(1−ρ\),p11=p01\+ρx^\{1\},x^\{0\}=\{p\_\{01\}\}/\(\{1\-\\rho\}\),p\_\{11\}=p\_\{01\}\+\\rhoall depend on the varying parameter\. Hence they no longer occur at common locations on the horizontal axis, and the defaultxx\-axis ticks in Figure[11](https://arxiv.org/html/2606.11192#A5.F11)display onlyp01p\_\{01\}\-invariant reference points\.

The dependence ofm​\(x\)m\(x\)onp01p\_\{01\}is again non\-monotone: for some belief values the index increases withp01p\_\{01\}, whereas for others it decreases, and the shape and location of the non\-myopic region vary substantially across the family\. Figure[12](https://arxiv.org/html/2606.11192#A5.F12), which plotsm​\(x\)m\(x\)as a function ofp01p\_\{01\}for a fixed representative beliefxx, makes this non\-monotone dependence more explicit\.

Asp01p\_\{01\}approaches the upper feasibility boundary1−ρ=0\.41\-\\rho=0\.4, the MP index appears to approach a limiting profile\. At the same time, the movement ofx1x^\{1\},x0x^\{0\}, andp11p\_\{11\}shows that varyingp01p\_\{01\}changes not only the magnitude of the index but also the location and extent of the regions in which the MP index differs from the myopic index\.

![Refer to caption](https://arxiv.org/html/2606.11192v1/x11.png)Figure 11:MP indexm​\(x\)m\(x\)versusxxfor several values ofp01p\_\{01\}\.![Refer to caption](https://arxiv.org/html/2606.11192v1/x12.png)Figure 12:MP indexm​\(x\)m\(x\)versusp01p\_\{01\}for a fixed beliefxx\.

### E\.2Implementation details and reproducibility

All computations were implemented in Julia\. The numerical evaluation of the discounted metrics and the MP index relied on the no\-ACK skeleton and renewal decompositions developed in Sections[4](https://arxiv.org/html/2606.11192#S4)–[4\.6](https://arxiv.org/html/2606.11192#S4.SS6)\. Infinite pre\-ACK series were evaluated by truncation using the explicit tail bounds derived in \([46](https://arxiv.org/html/2606.11192#S4.E46)\), so that the truncation depth could be chosen to meet a prescribed error tolerance\. Unless otherwise stated, the numerical routines used a toleranceε=10−10\\varepsilon=10^\{\-10\}\.

For dual\-bound computation in the policy\-benchmarking experiments, we used the grouped convex solver described in Section[5\.2](https://arxiv.org/html/2606.11192#S5.SS2), which exploits repeated project types and computes the minimizing Lagrange multiplier by bisection on a subgradient\. For policy simulation, the MP index policy used precomputed type\-specific lookup tables with linear interpolation, while Monte Carlo estimates were based on fixed random seeds to ensure reproducibility\.

The large\-scale numerical tests of \(PCLI1\) and \(PCLI2\), as well as the policy\-benchmarking experiments, were run in multithreaded mode\. The code was checked internally by comparing special cases against the tractable closed forms of Sections[4\.4](https://arxiv.org/html/2606.11192#S4.SS4)–[4\.6](https://arxiv.org/html/2606.11192#S4.SS6), by verifying identities such asF~=r​Θ~\\widetilde\{F\}=r\\widetilde\{\\Theta\}, and by confirming agreement between the original and accelerated dual\-bound solvers on representative test instances\.

The code used to generate the numerical results and figures is available from the author upon request\.

Similar Articles

Catching a Moving Subspace: Low-Rank Bandits Beyond Stationarity

arXiv cs.LG

This paper studies piecewise-stationary low-rank linear contextual bandits, proposes the SPSC algorithm that achieves dynamic regret scaling with the intrinsic rank instead of the ambient dimension, and characterizes the identification boundary for subspace recovery under scalar feedback.

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

arXiv cs.LG

This paper studies cooperative multi-player bandits in continuous Lipschitz action spaces when the Lipschitz constant is unknown, proposing a meta-algorithm (mECAB) that estimates the constant and coordinates discretization across players under different information structures, with regret guarantees.

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback

arXiv cs.AI

This paper introduces a hybrid Track-and-Stop algorithm for best arm identification in generalized linear bandits that unifies absolute and relative feedback. The authors propose a likelihood-ratio-based confidence sequence to adaptively allocate queries, demonstrating improved sample efficiency over baseline methods.

Stochastic Linear Bandits with Partially Observed Actions

arXiv cs.LG

This paper studies stochastic linear bandits where the agent only observes a random subset of action coordinates, proving that sublinear regret is possible when actions have low intrinsic dimension, and proposes the TOFU-POV algorithm with theoretical guarantees.