主-次弱耦合马尔可夫决策过程中的公平策略优化

arXiv cs.LG 论文

摘要

蒙特利尔高等商学院(HEC Montréal)与MILA的一篇论文提出了针对主-次弱耦合马尔可夫决策过程的公平策略优化方法,以单调凹公平函数取代功利主义目标,并引入一种基于计数比例的深度强化学习方法,其中包含基于优先级的采样器,该方法在机器置换任务以及纽约市出租车定价与调度任务上得到了验证。

arXiv:2609.36174v1 Announce Type: new Abstract: We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textit{permutation-invariant} policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:46

# Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes
Source: [https://arxiv.org/html/2609.36174](https://arxiv.org/html/2609.36174)
Xiaohui Tu††thanks:Corresponding author\. Email addresses:[xiaohui\.tu@hec\.ca](mailto:[email protected]),[yossiri\.adulyasak@hec\.ca](mailto:[email protected]), and[erick\.delage@hec\.ca](mailto:[email protected])Affiliation:GERAD & Department of Decision Sciences, HEC Montréal, Montréal, H3T2A7, CanadaAffiliation:MILA \- Quebec AI Institute, Montréal, H2S3H1, CanadaYossiri AdulyasakAffiliation:GERAD & Department of Logistics and Operations Management, HEC Montréal, Montréal, H3T2A7, CanadaErick DelageAffiliation:GERAD & Department of Decision Sciences, HEC Montréal, Montréal, H3T2A7, CanadaAffiliation:MILA \- Quebec AI Institute, Montréal, H2S3H1, Canada

###### Abstract

We consider fair resource allocation in sequential decision\-making environments modeled as major\-minor weakly coupled Markov decision processes \(M2WCMDP\)\. In this framework, resource constraints couple the action spaces of a major sub\-Markov decision process \(sub\-MDP\) and a population of minor sub\-MDPs that would otherwise operate independently\. Instead of using the traditional utilitarian \(total\-sum\) objective, we optimize a general class of monotone, concave, permutation\-invariant, normalized fairness functions\. With homogeneous minor sub\-MDPs, we prove that the problem under symmetry reduces to optimizing the platform\-plus\-mean\-participant utilitarian objective over the class ofpermutation\-invariantpolicies, which allows us to exploit efficient algorithms that optimize the utilitarian\-based objective to solve this fairness\-aware problem\. For more general settings, we introduce a count\-proportion\-based deep reinforcement learning approach with a priority\-based sampler that generates feasible count actions\. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure\. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City\-calibrated dataset\. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness\-aware performance while remaining scalable\.

## 1Introduction

The dynamic allocation of shared resources in large\-scale stochastic systems is a fundamental problem in various domains, including inventory pricing\([Gallego and Van Ryzin, 1994](https://arxiv.org/html/2609.36174#bib.bib32);[Adelman, 2004](https://arxiv.org/html/2609.36174#bib.bib33)\), marketing\([Bertsimas and Mersereau, 2007](https://arxiv.org/html/2609.36174#bib.bib34);[Caro and Gallien, 2007](https://arxiv.org/html/2609.36174#bib.bib4)\), and preemptive asset maintenance\([Nadarajah and Cire, 2025](https://arxiv.org/html/2609.36174#bib.bib68)\)\. These systems are often modeled using weakly coupled Markov decision processes \(WCMDPs\)\([Hawkins, 2003](https://arxiv.org/html/2609.36174#bib.bib73);[Adelman and Mersereau, 2008](https://arxiv.org/html/2609.36174#bib.bib74)\), where each component evolves according to a local sub\-Markov decision process \(sub\-MDP\), and coupling arises through shared resource constraints on the joint action space\.

However, a critical challenge arising in such systems is that system\-level endogenous information and coordination decisions cannot be represented only by local components\. In ride\-hailing platforms, for example, joint decisions on pricing, dispatching, and relocation not only determine current matching outcomes, but also influence future revenue through subsequent vehicle availability and the spatial distribution of demands\([Chen et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib39)\)\. The platform observes system\-level information such as outstanding customer queues and controls coordination decisions such as prices, while individual drivers follow local state transitions determined by their own locations and assigned routes\. Similar structures arise in asset maintenance and public infrastructure systems\([Nadarajah and Cire, 2025](https://arxiv.org/html/2609.36174#bib.bib68)\), where a central operator observes system\-level resource availability and allocates limited budgets across assets\. These applications motivate a major\-minor modeling structure, in which a major sub\-MDP represents system\-level dynamics, while a large number of minor sub\-MDPs represent participants\.

At the same time, maximizing total system utility alone is often insufficient when multiple stakeholders are affected\. A purely total\-sum objective may achieve high overall performance by assigning more profitable opportunities or scarce resources to certain subsets of participants\. However, in public systems, operational decisions are often expected to provide equitable service quality, balanced workload allocation, or fair access to opportunities across participants\. Examples include equitable driver earnings in mobility platforms\([Lesmana et al\., 2019](https://arxiv.org/html/2609.36174#bib.bib11);[Sun et al\., 2022](https://arxiv.org/html/2609.36174#bib.bib38)\), fair access to wireless networks\([Huaizhou et al\., 2013](https://arxiv.org/html/2609.36174#bib.bib10)\), and balanced service levels in healthcare\([Bertsimas et al\., 2013](https://arxiv.org/html/2609.36174#bib.bib18)\)and public infrastructure systems\([Michailidis et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib9)\)\. These considerations have led to growing interest in fairness\-aware sequential decision\-making to promote equitable outcomes among participants\([Hassanzadeh et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib52);[Bhattacharya et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib49);[Aminian et al\., 2026](https://arxiv.org/html/2609.36174#bib.bib2)\)\.

These two considerations motivate the need for new WCMDP formulations that model both endogenous system dynamics and fairness\-aware coordination\. Unfortunately, the important computational difficulties arise from the non\-separable problem structure and the curse of dimensionality\. While classical WCMDP formulations allow tractable analysis through decomposition techniques\([Adelman and Mersereau, 2008](https://arxiv.org/html/2609.36174#bib.bib74)\)or fluid approximation\([Brown and Zhang, 2025](https://arxiv.org/html/2609.36174#bib.bib12)\), they typically assume that sub\-MDPs have additive objectives and separable dynamics\. These assumptions break down when fairness is optimized as part of the objective\. Additionally, existing formulations fail to capture system\-level decisions that evolve separately from the individual behaviour, while affecting the transition and reward structure of the participants\.

We thus introduce a fairness\-aware major\-minor WCMDP \(M2WCMDP\) framework, which consists of one major sub\-MDP and a collection of minor sub\-MDPs\. Conditional on the major state and action, and on each participant’s own state\-action pair, each minor sub\-MDP evolves independently, while their actions remain coupled through shared resource constraints\. Within this framework, the objective balances the expected total discounted reward of the major sub\-MDP with a fairness measure evaluated over the vector of participant expected total discounted rewards, consistent with social welfare theory\([Weymark, 1981](https://arxiv.org/html/2609.36174#bib.bib47);[Chen and Hooker, 2023](https://arxiv.org/html/2609.36174#bib.bib26)\)\. This fairness measure accommodates many commonly used monotone, concave, permutation\-invariant, and normalized welfare measures, such as generalized Gini functions\([Weymark, 1981](https://arxiv.org/html/2609.36174#bib.bib47)\)and max\-min welfare\([Rawls, 1971](https://arxiv.org/html/2609.36174#bib.bib30)\)\.

Our first contribution is a structural reduction result for symmetric M2WCMDPs\. We prove that, under symmetry conditions, optimizing a fairness\-aware objective is equivalent to maximizing a trade\-off between the platform’s expected total discounted rewards and the participants’ average expected total discounted rewards over the class of permutation\-invariant policies without loss of optimality for the fairness\-aware objective\. Additionally, motivated by symmetry reduction of[Gast et al\. \(2022\)](https://arxiv.org/html/2609.36174#bib.bib43), the problem complexity is further reduced by obtaining an equivalent problem on a count\-based representation\.

The second contribution concerns efficient and scalable algorithms\. We develop a deep reinforcement learning \(DRL\) approach that uses count\-proportional representations and a stochastic policy network to generate feasible count actions\. We show mathematically the feasibility\-preserving and expressiveness properties of the proposed priority\-based sampling procedure\. We also provide an approach to translate aggregate decisions into permutation\-invariant participant\-level execution that prevents systematic discrimination among participants\. In particular, a count policy specifies how many participants take each action, while fairness at the individual level requires a permutation\-invariant randomized disaggregation rule that does not systematically favour specific participant labels\.

Finally, we evaluate the proposed framework in two numerical settings\. The first is a machine replacement problem \(MRP\), which provides a simple M2WCMDP environment where optimal solutions can be computed for small instances\. The second is a taxi dispatching and pricing problem \(TDPP\) calibrated with trip records from[New York City Taxi and Limousine Commission \(2026\)](https://arxiv.org/html/2609.36174#bib.bib6)\. This application evaluates the full major\-minor coordination structure\. The results show that the proposed methods achieve significantly improved fairness\-efficiency trade\-offs compared to benchmark approaches while remaining computationally scalable in large\-scale settings\.

Paper structure\.The remainder of this paper is organized as follows\. Section[2](https://arxiv.org/html/2609.36174#S2)introduces the M2WCMDP framework and formalizes the fairness\-aware optimization problem\. Section[3](https://arxiv.org/html/2609.36174#S3)establishes the utilitarian reduction under symmetry and presents the count aggregation reformulation\. In Section[4](https://arxiv.org/html/2609.36174#S4), we develop scalable reinforcement learning algorithms based on count\-proportion representations and stochastic policy networks\. Next, in Section[5](https://arxiv.org/html/2609.36174#S5), we validate the proposed framework and methodologies through extensive numerical experiments for the machine replacement and taxi dispatching applications\. Finally, we conclude the paper in Section[6](https://arxiv.org/html/2609.36174#S6)\. The electronic companions review related literature, scope the boundary of the welfare class, provide proofs and exact count\-model formulations, as well as report implementation details, benchmark formulations, and additional experiments\.

## 2Fairness\-Aware Major\-Minor Weakly Coupled Markov Decision Processes

We start by introducing the labelled infinite\-horizon M2WCMDPs in Section[2\.1](https://arxiv.org/html/2609.36174#S2.SS1)\. In Section[2\.2](https://arxiv.org/html/2609.36174#S2.SS2), we characterize the class of fairness measures considered in this work through a set of properties motivated by social welfare theory, and then formulate the fairness\-aware optimization problem\. Finally, in Section[2\.3](https://arxiv.org/html/2609.36174#S2.SS3), we instantiate the framework with two applications that will be used in the numerical studies\.

Notation\.Let\[N\]:=\{1,…,N\}\[N\]:=\\\{1\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}for any integerNN\. An indicator function𝕀\{x∈A\}\{\\mathbb\{I\}\}\\\{x\\in A\\\}equals 1 ifx∈Ax\\in Aand 0 otherwise\. For a finite setXX, letΔ⁡\(X\)\\Delta\(X\)denote the probability simplex overℝ\|X\|\{\\mathbb\{R\}\}^\{\|X\|\}\. For a vector\-valued mappingξ:ℝN→ℝ\|X\|\\xi:\{\\mathbb\{R\}\}^\{N\}\\rightarrow\{\\mathbb\{R\}\}^\{\|X\|\}with components indexed by a finite setXX, we write\[ξ​\(𝒚\)\]​\(x\)\[\\xi\(\{\\bm\{y\}\}\)\]\(x\)for itsxx\-indexed component, for any𝒚∈ℝN\{\\bm\{y\}\}\\in\{\\mathbb\{R\}\}^\{N\}andx∈Xx\\in X\. We use𝟏\{\\bm\{1\}\}and𝟎\{\\bm\{0\}\}to denote all\-ones and all\-zero arrays, respectively, with their dimensions specified by subscripts\. We define𝒢N\\mathcal\{G\}^\{N\}as the set of allN×NN\\times Npermutation matrices, whereQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}is a linear operator that acts on any𝒗∈ℝN\{\\bm\{v\}\}\\in\{\\mathbb\{R\}\}^\{N\}by reordering its components\.

### 2\.1Model Formulation

We consider a centrally coordinated system evolving over an infinite horizont∈𝒯:=\{0,1,…\}t\\in\{\\mathcal\{T\}\}:=\\\{0\\mathchar 24891\\allowbreak 1\\mathchar 24891\\allowbreak\\dots\\\}\. This system is formulated as a discrete\-time Markov Decision Process \(MDP\) that decomposes into a major\-minor format, which consists of one major sub\-MDPℳ0\\mathcal\{M\}^\{0\}\(referred to as theplatform\) andNNminor sub\-MDPs\{ℳn\}n∈\[N\]\\\{\\mathcal\{M\}^\{n\}\\\}\_\{n\\in\[N\]\}\(referred to asparticipants\)\. A central decision maker jointly optimizes system\-level efficiency and participant\-level welfare\.

Formally, the major sub\-MDPℳ0\\mathcal\{M\}^\{0\}is characterized by the tuple\(𝒮0,𝒜0,p0,r0,μ0,γ\)\(\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak p^\{0\}\\mathchar 24891\\allowbreak r^\{0\}\\mathchar 24891\\allowbreak\\mu^\{0\}\\mathchar 24891\\allowbreak\\gamma\), where𝒮0\\mathcal\{S\}^\{0\}and𝒜0\\mathcal\{A\}^\{0\}are the finite major state and action spaces, respectively\. The major transition probabilityp0​\(⋅\)p^\{0\}\(\\cdot\)and real\-valued reward functionr0​\(⋅\)r^\{0\}\(\\cdot\)depend on the joint state and action of the entire system, and will be defined in Equation[2](https://arxiv.org/html/2609.36174#S2.E2)and the paragraph following it\. Thenn\-th minor sub\-MDPℳn\\mathcal\{M\}^\{n\}is defined by the tuple\(𝒮n,𝒜n,pn,rn,μn,γ\)\(\\mathcal\{S\}^\{n\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{n\}\\mathchar 24891\\allowbreak p^\{n\}\\mathchar 24891\\allowbreak r^\{n\}\\mathchar 24891\\allowbreak\\mu^\{n\}\\mathchar 24891\\allowbreak\\gamma\), where𝒮n\\mathcal\{S\}^\{n\}is a finite set of states with cardinalitySS, and𝒜n\\mathcal\{A\}^\{n\}is a finite set of actions with cardinalityAA\. The stationary transition probability function is denoted bypn​\(s~n∣s0,sn,a0,an\)=ℙ⁡\(st\+1n=s~n∣st0=s0,stn=sn,at0=a0,atn=an\)p^\{n\}\(\\tilde\{s\}^\{n\}\\mid s^\{0\}\\mathchar 24891\\allowbreak s^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a^\{n\}\)=\\mathbb\{P\}\(s^\{n\}\_\{t\+1\}=\\tilde\{s\}^\{n\}\\mid s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak s^\{n\}\_\{t\}=s^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak a^\{n\}\_\{t\}=a^\{n\}\), representing the probability of reaching minor states~n∈𝒮n\\tilde\{s\}^\{n\}\\in\\mathcal\{S\}^\{n\}, after performing minor actionan∈𝒜na^\{n\}\\in\\mathcal\{A\}^\{n\}and major actiona0∈𝒜0a^\{0\}\\in\\mathcal\{A\}^\{0\}in minor statesn∈𝒮ns^\{n\}\\in\\mathcal\{S\}^\{n\}and major states0∈𝒮0s^\{0\}\\in\\mathcal\{S\}^\{0\}at timett\. The reward functionrn​\(s0,sn,a0,an\)r^\{n\}\(s^\{0\}\\mathchar 24891\\allowbreak s^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a^\{n\}\)denotes the immediate real\-valued reward obtained by executing actionana^\{n\}in statesns^\{n\}under the coordination of the major system\(s0,a0\)\(s^\{0\}\\mathchar 24891\\allowbreak a^\{0\}\)\. In the general formulation, the transition probabilities and the reward functions may vary with the sub\-MDPnn, and we assume only that they are stationary over time for simplicity\. The initial state distributions are represented byμ0∈Δ⁡\(𝒮0\)\\mu^\{0\}\\in\\Delta\(\\mathcal\{S\}^\{0\}\)andμn∈Δ⁡\(𝒮n\)\\mu^\{n\}\\in\\Delta\(\\mathcal\{S\}^\{n\}\), while each sub\-MDP’s initial state is assumed to be drawn independently\. The discount factor, common to all sub\-MDPs, is denoted byγ∈\[0,1\)\\gamma\\in\[0\\mathchar 24891\\allowbreak 1\)\.

The state of the system at any time is represented by the pair\(s0,𝒔\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\. The components0∈𝒮0s^\{0\}\\in\\mathcal\{S\}^\{0\}corresponds to the state of the major sub\-MDP\. The joint state space for theNNminor sub\-MDPs is given by𝒔=\(s1,…,sN\)∈𝒮\(N\)\\bm\{s\}=\(s^\{1\}\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak s^\{N\}\)\\in\\mathcal\{S\}^\{\(N\)\}, with𝒮\(N\)\\mathcal\{S\}^\{\(N\)\}as the Cartesian product of the individual state spaces\. There is a set ofKKresource constraints that couple the action selection of the minor sub\-MDPs\. Letdk​\(an∣sn\)d\_\{k\}\(a^\{n\}\\mid s^\{n\}\)be the state\-dependent non\-negative consumption of thekk\-th resource by thenn\-th minor sub\-MDP\. Then, for a given system state\(s0,𝒔\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\), the feasible joint minor action space is a state\-dependent subset of the Cartesian product of minor action spaces, defined as

𝒜\(N\)\(s0,𝒔\):=\{\(a1,…,aN\)∈×n=1N𝒜n\|∑n=1Ndk\(an∣sn\)≤bk\(s0\),∀k∈\[K\]\},\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\):=\\left\\\{\(a^\{1\}\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak a^\{N\}\)\\in\\mathop\{\\vbox\{\\hbox\{\\LARGE$\\times$\}\}\}\_\{n=1\}^\{N\}\\mathcal\{A\}^\{n\}\\,\\Big\|\\,\\sum\_\{n=1\}^\{N\}d\_\{k\}\(a^\{n\}\\mid s^\{n\}\)\\leq b\_\{k\}\(s^\{0\}\)\\mathchar 24891\\allowbreak\\forall k\\in\[K\]\\right\\\}\\mathchar 24891\\allowbreak\(1\)wherebk:𝒮0→ℝ\+b\_\{k\}:\\mathcal\{S\}^\{0\}\\to\{\\mathbb\{R\}\}\_\{\+\}is the budget function mapping the major state to the available resource capacity of typekkto capture the exogenous supply of resources \(e\.g\., passenger demands\) which varies over time\. The major action does not directly restrict the feasible minor action space, but serves as a coordination signal that influences the system through the transition dynamics and the reward function\.

The minor sub\-MDP transitions are conditionally independent\. From an operational perspective, this assumption implies that conditional on the platform’s state \(e\.g\., outstanding customer queues\), the macro\-level coordination signal \(e\.g\., the price multiplier\), and each participant’s own state\-action pair, participants’ state transitions are mutually independent, as a driver’s subsequent travel time and next location depend only on their own realized trajectory\. Specifically, minor sub\-MDPs transition from state𝒔\\bm\{s\}to state𝒔~\\tilde\{\\bm\{s\}\}for a given major actiona0a^\{0\}and feasible minor actions𝒂=\(a1,…,aN\)\\bm\{a\}=\(a^\{1\}\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak a^\{N\}\)at timettwith probability

p\(N\)​\(𝒔~∣s0,𝒔,a0,𝒂\):=∏n=1Npn​\(s~n∣s0,sn,a0,an\)=∏n=1Nℙ⁡\(st\+1n=s~n∣st0=s0,stn=sn,at0=a0,atn=an\),p^\{\(N\)\}\(\\tilde\{\\bm\{s\}\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\prod^\{N\}\_\{n=1\}p^\{n\}\(\\tilde\{s\}^\{n\}\\mid s^\{0\}\\mathchar 24891\\allowbreak s^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a^\{n\}\)=\\prod^\{N\}\_\{n=1\}\\mathbb\{P\}\(s\_\{t\+1\}^\{n\}=\\tilde\{s\}^\{n\}\\mid s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak s\_\{t\}^\{n\}=s^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak a\_\{t\}^\{n\}=a^\{n\}\)\\mathchar 24891\\allowbreakwhile the major sub\-MDP has the transition kernel

p0​\(s~0∣s0,𝒔,a0,𝒂\):=ℙ⁡\(st\+10=s~0∣st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\)\.p^\{0\}\(\\tilde\{s\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\mathbb\{P\}\(s\_\{t\+1\}^\{0\}=\\tilde\{s\}^\{0\}\\mid s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\)\.\(2\)
After choosing an action\(a0,𝒂\)∈𝒜0×𝒜\(N\)​\(s0,𝒔\)\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\in\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)when in state\(s0,𝒔\)∈𝒮0×𝒮\(N\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\\in\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}, the decision maker receives rewards for the minor sub\-MDPs defined as𝒓⁡\(s0,𝒔,a0,𝒂\)=\(r1​\(s0,s1,a0,a1\),…,rN​\(s0,sN,a0,aN\)\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=\(r^\{1\}\(s^\{0\}\\mathchar 24891\\allowbreak s^\{1\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a^\{1\}\)\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak r^\{N\}\(s^\{0\}\\mathchar 24891\\allowbreak s^\{N\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a^\{N\}\)\)with each component representing the reward associated with the respective minor sub\-MDPℳn\\mathcal\{M\}^\{n\}\. We employ a vector form for the rewards to later offer the flexibility for formulating fairness objectives on individual expected total discounted rewards\. The major sub\-MDP receives a real\-valued reward, denoted byr0​\(s0,𝒔,a0,𝒂\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\), which represents the platform\-level performance metric\.

Collectively, the infinite\-horizon M2WCMDP model is completely specified by the collection\(ℳ0,ℳ\(N\)\)\(\\mathcal\{M\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{M\}^\{\(N\)\}\), whereℳ0:=\(𝒮0,𝒜0,p0,r0,μ0,γ\)\\mathcal\{M\}^\{0\}:=\(\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak p^\{0\}\\mathchar 24891\\allowbreak r^\{0\}\\mathchar 24891\\allowbreak\\mu^\{0\}\\mathchar 24891\\allowbreak\\gamma\)andℳ\(N\):=\(𝒮\(N\),𝒜\(N\),p\(N\),𝒓,μ\(N\),γ\)\\mathcal\{M\}^\{\(N\)\}:=\(\\mathcal\{S\}^\{\(N\)\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{\(N\)\}\\mathchar 24891\\allowbreak p^\{\(N\)\}\\mathchar 24891\\allowbreak\\bm\{r\}\\mathchar 24891\\allowbreak\\mu^\{\(N\)\}\\mathchar 24891\\allowbreak\\gamma\)\. Under centralized coordination, the system observes the full joint state and jointly selects the major action and a resource\-feasible minor action for each participant\. LetΠ\\Pidenote the set of all feasible stationary randomized Markov policies\. A joint policy𝝅∈Π\\bm\{\\pi\}\\in\\Piis formally defined by a mapping from the joint state space𝒮0×𝒮\(N\)\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}to a probability distribution over the state\-dependent joint action space𝒜0×𝒜\(N\)​\(s0,𝒔\)\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\. Thus,𝝅\(a0,𝒂∣s0,𝒔\)\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)represents the probability of taking the joint action\(a0,𝒂\)\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)given the joint state\(s0,𝒔\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\. We define this feasible policy space asΠ:=\{𝝅\|𝝅\(⋅,⋅∣s0,𝒔\)∈Δ\(𝒜0×𝒜\(N\)\(s0,𝒔\)\),∀\(s0,𝒔\)\}\\Pi:=\\mathopen\{\\\{\}\\bm\{\\pi\}\\,\\big\|\\,\\bm\{\\pi\}\(\\cdot\\mathchar 24891\\allowbreak\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\\in\\Delta\(\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\)\\mathchar 24891\\allowbreak\\,\\forall\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\\mathclose\{\\\}\}\.

The initial system state\(s00,𝒔0\)\(s^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}\)is sampled from the joint initial distribution𝝁∈Δ⁡\(𝒮0×𝒮\(N\)\)\\bm\{\\mu\}\\in\\Delta\(\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\)\. Using the discounted\-reward criterion, the performance of any joint policy𝝅\\bm\{\\pi\}is evaluated by its expected total discounted rewards\. For the major sub\-MDPℳ0\\mathcal\{M\}^\{0\}under the initial distribution𝝁\\bm\{\\mu\}, this is given by:

V00​\(𝝅\):=𝔼𝝅​\[∑t=0∞γt​r0​\(st0,𝒔t,at0,𝒂t\)\|\(s00,𝒔0\)∼𝝁\]\.V^\{0\}\_\{0\}\(\\bm\{\\pi\}\):=\\mathbb\{E\}\_\{\\bm\{\\pi\}\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r^\{0\}\(s^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}\)\\,\\Big\|\(s^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}\)\\sim\\bm\{\\mu\}\\right\]\.\(3\)
Similarly, the expected total discounted rewardsV0n​\(𝝅\)V^\{n\}\_\{0\}\(\\bm\{\\pi\}\)specific to thenn\-th sub\-MDPℳn\{\\mathcal\{M\}\}^\{n\}, under the initial distribution𝝁\\bm\{\\mu\}and policy𝝅\\bm\{\\pi\}, is defined as

V0n​\(𝝅\):=𝔼𝝅​\[∑t=0∞γt​rn​\(st0,stn,at0,atn\)\|\(s00,𝒔0\)∼𝝁\],∀n∈\[N\],\{V\}^\{n\}\_\{0\}\(\\bm\{\\pi\}\):=\\mathbb\{E\}\_\{\\bm\{\\pi\}\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r^\{n\}\(s^\{0\}\_\{t\}\\mathchar 24891\\allowbreak s^\{n\}\_\{t\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}\\mathchar 24891\\allowbreak a^\{n\}\_\{t\}\)\\,\\Big\|\\,\(s\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}\)\\sim\\bm\{\\mu\}\\right\]\\mathchar 24891\\allowbreak\\forall n\\in\[N\]\\mathchar 24891\\allowbreak\(4\)where\(at0,𝒂t\)∼𝝅\(⋅,⋅∣st0,𝒔t\)\(a^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}\)\\sim\\bm\{\\pi\}\(\\cdot\\mathchar 24891\\allowbreak\\cdot\\mid s^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}\)\. The vectorial expected total discounted rewards𝑽0​\(𝝅\)\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)for all minor sub\-MDPs under policy𝝅\\bm\{\\pi\}is defined as

𝑽0​\(𝝅\):=\[V01​\(𝝅\),…,V0N​\(𝝅\)\]⊤\.\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\):=\[\{V\}\_\{0\}^\{1\}\(\{\\bm\{\\pi\}\}\)\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak\{V\}\_\{0\}^\{N\}\(\{\\bm\{\\pi\}\}\)\]^\{\\top\}\.\(5\)
Traditionally, the performance of such a centrally coordinated system is evaluated by maximizing a purely utilitarian objective\. Expressed through our formulated value functions \([3](https://arxiv.org/html/2609.36174#S2.E3)\) and \([5](https://arxiv.org/html/2609.36174#S2.E5)\), this typically involves maximizing the expected infinite\-horizon discounted total reward of the platform and all participants,V00​\(𝝅\)\+𝟏⊤​𝑽0​\(𝝅\)V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\+\{\\bm\{1\}\}^\{\\top\}\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\([Zhu et al\., 2021](https://arxiv.org/html/2609.36174#bib.bib31)\)\. When no separate platform reward is modelled, this is reduced to𝟏⊤​𝑽0​\(𝝅\)\{\\bm\{1\}\}^\{\\top\}\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\([Chen et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib39)\)\. However, under a strictly efficiency\-driven objective, the utilities of marginalized participants might be systematically sacrificed to achieve a higher system performance\. Fairness concerns therefore arise naturally in various decision\-making contexts and application domains, including healthcare management\([Bertsimas et al\., 2013](https://arxiv.org/html/2609.36174#bib.bib18)\), transportation\([Gutjahr and Fischer, 2018](https://arxiv.org/html/2609.36174#bib.bib16);[Chen et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib39)\), facility location\([Filippi et al\., 2021](https://arxiv.org/html/2609.36174#bib.bib17)\), and resource allocation\([Hassanzadeh et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib52);[Bhattacharya et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib49)\)\. Fairness should be explicitly modeled to reflect equity in the distribution of outcomes across the participants in such systems\([Aziz et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib1)\)\. This concern maps directly to our vector\-valued welfare framework, where we seek to evaluate the joint performance vector𝑽0​\(𝝅\)\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)rather than only the total sum\. This motivates the fairness\-aware optimization problem\.

### 2\.2Fair Optimization Problem

In modern welfare economics, ensuring equity among multiple participants fundamentally reduces to formulating a socially meaningful scalarization of vector\-valued outcomes\. Grounded in the social choice and utility theory literature\([Weymark, 1981](https://arxiv.org/html/2609.36174#bib.bib47);[Moulin, 1991](https://arxiv.org/html/2609.36174#bib.bib46);[Tsang and Shehadeh, 2025](https://arxiv.org/html/2609.36174#bib.bib48)\), we resort to a social welfare function that aggregates the vectorial expected discounted rewards of all minor sub\-MDPs into a scalar representation of system\-level welfare criteria\. This approach provides a unifying mathematical framework for representing welfare objectives suitable for optimization contexts\([Chen and Hooker, 2023](https://arxiv.org/html/2609.36174#bib.bib26);[Fan et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib25)\)\.

Rather than prescriptively defining what constitutes a fair measure, we identify the mathematical properties required to enable tractable fair policy optimization\. We formally defineρ\\rho\-fairness, which serves as the theoretical foundation for our subsequent utilitarian reduction theorem in Section[3](https://arxiv.org/html/2609.36174#S3)\.

###### Definition 1\(ρ\\rho\-fairness\)\.

Letρ:𝒱→ℝ\\rho:\{\\mathcal\{V\}\}\\rightarrow\{\\mathbb\{R\}\}, with𝒱⊆ℝN\{\\mathcal\{V\}\}\\subseteq\{\\mathbb\{R\}\}^\{N\}as the domain ofρ\\rho, be a fairness measure that satisfies:

- •Monotonicity \(MN\): For any𝒗,𝒘∈𝒱\{\\bm\{v\}\}\\mathchar 24891\\allowbreak\{\\bm\{w\}\}\\in\{\\mathcal\{V\}\}such that𝒗≥𝒘\{\\bm\{v\}\}\\geq\{\\bm\{w\}\}componentwise,ρ⁡\[𝒗\]≥ρ⁡\[𝒘\]\\rho\[\{\\bm\{v\}\}\]\\geq\\rho\[\{\\bm\{w\}\}\],
- •Concavity \(CV\): The set𝒱\{\\mathcal\{V\}\}is convex and∀𝒗,𝒘∈𝒱\\forall\{\\bm\{v\}\}\\mathchar 24891\\allowbreak\{\\bm\{w\}\}\\in\{\\mathcal\{V\}\}, andτ∈\[0,1\]\\tau\\in\[0\\mathchar 24891\\allowbreak 1\],ρ⁡\[τ​𝒗\+\(1−τ\)​𝒘\]≥τ​ρ​\[𝒗\]\+\(1−τ\)​ρ​\[𝒘\]\\rho\[\\tau\{\\bm\{v\}\}\+\(1\-\\tau\)\{\\bm\{w\}\}\]\\geq\\tau\\rho\[\{\\bm\{v\}\}\]\+\(1\-\\tau\)\\rho\[\{\\bm\{w\}\}\],
- •Permutation invariance \(PI\):∀𝒗∈𝒱\\forall\{\\bm\{v\}\}\\in\{\\mathcal\{V\}\}and allQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}, bothQ​𝒗∈𝒱Q\{\\bm\{v\}\}\\in\{\\mathcal\{V\}\}andρ⁡\[𝒗\]=ρ⁡\[Q​𝒗\]\\rho\[\{\\bm\{v\}\}\]=\\rho\[Q\{\\bm\{v\}\}\],
- •Constant vector invariance \(CVI\):∀c∈ℝ\\forall c\\in\{\\mathbb\{R\}\}such thatc​𝟏∈𝒱c\{\\bm\{1\}\}\\in\{\\mathcal\{V\}\},ρ⁡\[c​𝟏\]=c\\rho\[c\{\\bm\{1\}\}\]=c\.

MN ensures that improving participant outcomes componentwise cannot reduce social welfare\. CV reflects a preference for equality through decreasing marginal gains and guarantees that redistributing reward among participants cannot decrease welfare\. CV is not only a desirable property in optimization contexts, but also a property satisfied by many commonly used fairness measures as if a fairness measure is non\-concave, its value could paradoxically increase by partitioning the groups of interests into subgroups\([Kolm, 1976](https://arxiv.org/html/2609.36174#bib.bib27);[Williamson and Menon, 2019](https://arxiv.org/html/2609.36174#bib.bib28);[Tsang and Shehadeh, 2025](https://arxiv.org/html/2609.36174#bib.bib48)\)\. Next, PI ensures that welfare does not depend on participant identities\([Moulin, 1991](https://arxiv.org/html/2609.36174#bib.bib46);[Chen and Hooker, 2023](https://arxiv.org/html/2609.36174#bib.bib26)\)\. CVI ensures the measure is normalized to standard individual reward units, which provides a normalized economic scale when trading off against the platform’s utilitarian reward\.

We provide two examples on fairness measures satisfying these conditions that will be used in numerical experiments\. The comparison of additional welfare measures and Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)compliance is provided in[EC\.2](https://arxiv.org/html/2609.36174#A2)to scope the boundary of the proposed welfare class\.

###### Example 1\(Generalized Gini Function\)\.

A typical instance is the Generalized Gini Function \(GGF\)\([Weymark, 1981](https://arxiv.org/html/2609.36174#bib.bib47);[Siddique et al\., 2020](https://arxiv.org/html/2609.36174#bib.bib22)\), defined by assigning ordered weights to the components of a utility vector thatGGF𝒘⁡\[𝒗\]:=min⁡∑n=1Nσ∈𝕊N⁡wn​vσ⁡\(n\)\\GGF\_\{\{\\bm\{w\}\}\}\[\{\\bm\{v\}\}\]:=\\min\_\{\\sigma\\in\{\\mathbb\{S\}\}^\{N\}\}\\sum\_\{n=1\}^\{N\}w\_\{n\}v\_\{\\sigma\(n\)\}, where𝕊N\{\\mathbb\{S\}\}^\{N\}is the set of all permutations ofNNelements, and𝒘∈Δ⁡\(\[N\]\)\{\\bm\{w\}\}\\in\\Delta\(\[N\]\)satisfiesw1≥w2≥⋯≥wNw\_\{1\}\\geq w\_\{2\}\\geq\\dots\\geq w\_\{N\}\. The minimizerσ∗\\sigma^\{\*\}effectively sorts𝒗\{\\bm\{v\}\}in ascending order, prioritizing the lowest earners\. Another prominent special case is the Rawlsian max\-min fairness\([Rawls, 1971](https://arxiv.org/html/2609.36174#bib.bib30)\),ρ⁡\[𝒗\]=minn∈\[N\]⁡\{vn\}\\rho\[\{\\bm\{v\}\}\]=\\min\_\{n\\in\[N\]\}\\\{v\_\{n\}\\\}, which corresponds tow1=1w\_\{1\}=1andwi=0,∀i\>1w\_\{i\}=0\\mathchar 24891\\allowbreak\\forall i\>1\. Both natively satisfy Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)\.□\\square

###### Example 2\(Certainty\-Equivalent Representations ofα\\alpha\-Fairness\)\.

The classicalα\\alpha\-fair welfare family\([Mo and Walrand, 2000](https://arxiv.org/html/2609.36174#bib.bib37);[Ju et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib29)\)is parameterized byα\>0\\alpha\>0and takes the additive formWα​\[𝒗\]=∑n=1Nvn1−α1−αW\_\{\\alpha\}\[\{\\bm\{v\}\}\]=\\sum\_\{n=1\}^\{N\}\\frac\{v\_\{n\}^\{1\-\\alpha\}\}\{1\-\\alpha\}, with∑n=1Nlog⁡\(vn\)\\sum\_\{n=1\}^\{N\}\\log\(v\_\{n\}\)forα=1\\alpha=1as a special instance of Nash Social Welfare\([Fan et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib25);[Mandal and Gan, 2022](https://arxiv.org/html/2609.36174#bib.bib45)\)\. While it satisfies MN, CV and PI, it violates CVI because it is not normalized\. For instance, evaluating a uniform outcome𝑽=v¯​𝟏\\bm\{V\}=\\bar\{v\}\{\\bm\{1\}\}yieldsN​v¯1−α1−αN\\frac\{\\bar\{v\}^\{1\-\\alpha\}\}\{1\-\\alpha\}, explicitly violating CVI that requiresρ⁡\(v¯​𝟏\)=v¯\\rho\(\\bar\{v\}\{\\bm\{1\}\}\)=\\bar\{v\}\. To restore normalization, we apply the certainty\-equivalent alternatives via the inverse utility functionu−1u^\{\-1\}by setting the expected utility modelρ⁡\[𝒗\]=u−1​\(1N​∑n=1Nu⁡\(vn\)\)\\rho\[\{\\bm\{v\}\}\]=u^\{\-1\}\\mathopen\{\(\}\{\\frac\{1\}\{N\}\}\\sum\_\{n=1\}^\{N\}u\(v\_\{n\}\)\\mathclose\{\)\}, whereu⁡\(⋅\)u\(\\cdot\)is a monotone and concave function\. This leads to the alternative ordinally equivalent welfare classρα​\[𝒗\]=\(1N​∑n=1Nvn1−α\)11−α\\rho\_\{\\alpha\}\[\{\\bm\{v\}\}\]=\\mathopen\{\(\}\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}v\_\{n\}^\{1\-\\alpha\}\\mathclose\{\)\}^\{\\frac\{1\}\{1\-\\alpha\}\}forα\>0,α≠1\\alpha\>0\\mathchar 24891\\allowbreak\\alpha\\neq 1, and the geometric meanρ1​\[𝒗\]=\(∏n=1Nvn\)1/N\\rho\_\{1\}\[\{\\bm\{v\}\}\]=\\mathopen\{\(\}\\prod\_\{n=1\}^\{N\}v\_\{n\}\\mathclose\{\)\}^\{1/N\}for Nash Social Welfare\. As limits, this family recovers the utilitarian mean \(α=0\\alpha=0\), the normalized Nash Social Welfare\([Fan et al\., 2023](https://arxiv.org/html/2609.36174#bib.bib25)\)\(α→1\\alpha\\to 1\), and the Rawlsian max\-min\([Rawls, 1971](https://arxiv.org/html/2609.36174#bib.bib30)\)\(α→∞\\alpha\\to\\infty\), all of which satisfy Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)\.□\\square

Given an initial joint state distribution𝝁∈Δ⁡\(𝒮0×𝒮\(N\)\)\\bm\{\\mu\}\\in\\Delta\(\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\), we are seeking a policy that achieves the best tradeoff between the efficiency of the platform and the fairness of participants\([Tsang and Shehadeh, 2025](https://arxiv.org/html/2609.36174#bib.bib48)\)\. Formally, the objective of our fairness\-aware optimization problem \(ρ\\rho\-M2WCMDPs\) under policy𝝅\\bm\{\\pi\}is given by

max𝝅∈Π⁡Gρ​\(𝝅\):=V00​\(𝝅\)\+λ​ρ​\[𝑽0​\(𝝅\)\],\\max\_\{\\bm\{\\pi\}\\in\\Pi\}G\_\{\\rho\}\(\\bm\{\\pi\}\):=V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\+\\lambda\\rho\[\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\]\\mathchar 24891\\allowbreak\(6\)whereV00​\(𝝅\)V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)denotes the expected total discounted rewards of the major sub\-MDP,𝑽0​\(𝝅\)\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)is the vectorial counterpart for the minor sub\-MDPs, both evaluated under the initial distribution𝝁\\bm\{\\mu\}, andλ≥0\\lambda\\geq 0is the weighting factor\. We adopt this weighted\-sum scalarization because it provides a transparent and widely used mechanism for balancing platform efficiency and participant welfare through a single preference parameterλ\\lambda\([Bertsimas et al\., 2012](https://arxiv.org/html/2609.36174#bib.bib20);[Hooker and Williams, 2012](https://arxiv.org/html/2609.36174#bib.bib19);[Tsang and Shehadeh, 2025](https://arxiv.org/html/2609.36174#bib.bib48)\)\. We further note that whenρ\\rhohas a restricted domain, i\.e\.,𝒱≠ℝ\\mathcal\{V\}\\neq\{\\mathbb\{R\}\}, one needs to impose the following assumption to make theρ\\rho\-M2WCMDP well defined\.

###### Assumption 1\.

The expected discounted total reward vector generated by any feasible policy lies entirely within the domain ofρ\\rho\. That is,𝐕0​\(𝛑\)∈𝒱,∀𝛑∈Π\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\\in\{\\mathcal\{V\}\}\\mathchar 24891\\allowbreak\\forall\\bm\{\\pi\}\\in\\Pi\.

This assumption is required to ensure thatGρ​\(𝝅\)G\_\{\\rho\}\(\\bm\{\\pi\}\)is well defined for allπ∈Π\\pi\\in\\Pi\. Alternatively, one should consider using a differentρ\\rhomeasure or formulating a version ofρ\\rho\-M2WCMDP that constrains𝑽0​\(𝝅\)∈𝒱\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\\in\{\\mathcal\{V\}\}, which unfortunately falls outside the scope of our study\.

### 2\.3Two Illustrative Examples

To contextualize the M2WCMDP framework and the tension between system efficiency and equity, we introduce two representative applications that will be extensively studied in the subsequent numerical experiments\.

###### Example 3\(Machine Replacement Problem\)\.

This problem is adapted from[Delage and Mannor \(2010\)](https://arxiv.org/html/2609.36174#bib.bib13);[Akbarzadeh and Mahajan \(2019\)](https://arxiv.org/html/2609.36174#bib.bib14), and represents a natural degenerate case where no active major sub\-MDP is present \(i\.e\.,𝒮0\\mathcal\{S\}^\{0\}and𝒜0\\mathcal\{A\}^\{0\}are singletons\)\. Consider a maintenance system consisting only of a large collection of deteriorating machines \(minor sub\-MDPsℳ\(N\)\\mathcal\{M\}^\{\(N\)\}\), reducing the system to a classical WCMDP\. For each machinenn, the minor statesn∈𝒮ns^\{n\}\\in\\mathcal\{S\}^\{n\}captures its current deterioration level, and the binary action space𝒜n:=\{0,1\}\\mathcal\{A\}^\{n\}:=\\\{0\\mathchar 24891\\allowbreak 1\\\}corresponds to either continuing normal operation \(an=0a^\{n\}=0\) or performing an active replacement \(an=1a^\{n\}=1\) to reset the machine state to the newest condition\. The sub\-MDPs are weakly coupled through the maintenance budgetbkb\_\{k\}\(e\.g\., available repair crews or parts\) that implicitly restricts the decision space for maintenance resources of typek∈\[K\]k\\in\[K\]at each time step\. Formally, the feasible joint minor action space is coupled as𝒜\(N\)\(𝒔\):=\{𝒂∈×n=1N𝒜n\|∑n=1Ndk\(an∣sn\)≤bk,∀k∈\[K\]\}\\mathcal\{A\}^\{\(N\)\}\(\\bm\{s\}\):=\\mathopen\{\\\{\}\\bm\{a\}\\in\\mathop\{\\vbox\{\\hbox\{\\LARGE$\\times$\}\}\}\_\{n=1\}^\{N\}\\mathcal\{A\}^\{n\}\\,\\Big\|\\,\\sum\_\{n=1\}^\{N\}d\_\{k\}\(a^\{n\}\\mid s^\{n\}\)\\leq b\_\{k\}\\mathchar 24891\\allowbreak\\,\\forall k\\in\[K\]\\mathclose\{\\\}\}\.□\\square

In the context of Example[3](https://arxiv.org/html/2609.36174#Thmexample3), if equipment is regionally distributed in cases like electricity or telecommunication networks\([Nadarajah and Cire, 2025](https://arxiv.org/html/2609.36174#bib.bib68)\), a fair policy guarantees equitable operations and replacements, thereby preventing frequent failures in specific areas that lead to unsatisfactory and unfair results for certain customers\. A similar issue occurs in taxi dispatching systems \(as detailed in Example[4](https://arxiv.org/html/2609.36174#Thmexample4)\), purely efficiency\-driven ride assignments may lead to the situation where some drivers consistently get profitable trips, while others systematically receive fewer opportunities\([Dai et al\., 2017](https://arxiv.org/html/2609.36174#bib.bib44)\)\.

###### Example 4\(Taxi Dispatching and Joint Pricing Control Problem\)\.

A mobility\-on\-demand platform coordinatesNNhomogeneous taxis over zonesℳ:=\{1,…,M\}\{\\mathcal\{M\}\}:=\\\{1\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak M\\\}\. The network connectivity is represented by a graph with an adjacency matrix𝑨adj∈\{0,1\}M×M\\bm\{A\}\_\{\\text\{adj\}\}\\in\\\{0\\mathchar 24891\\allowbreak 1\\\}^\{M\\times M\}, whereAadj​\(i,j\)=1A\_\{\\text\{adj\}\}\(i\\mathchar 24891\\allowbreak j\)=1indicates the presence of a direct movement from zoneiitojj\. The travel times between any two nodesi,j∈ℳi\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}are given by a deterministic duration matrixD∈ℕM×MD\\in\{\\mathbb\{N\}\}^\{M\\times M\}, which is proportional to the shortest geodesic distance between two nodes\. The maximum travel duration across the network is defined asDmax:=maxi,j∈ℳ⁡D⁡\(i,j\)D\_\{\\max\}:=\\max\_\{i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\}D\(i\\mathchar 24891\\allowbreak j\), and let𝒟:=\{0,…,Dmax\}\{\\mathcal\{D\}\}:=\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak D\_\{\\max\}\\\}\.

The fleet current state is given by𝒔=\{sn\}n∈\[N\]\\bm\{s\}=\\\{s^\{n\}\\\}\_\{n\\in\[N\]\}, where the state of each taxinnis a tuplesn:=\(i,d\)∈ℳ×𝒟s^\{n\}:=\(i\\mathchar 24891\\allowbreak d\)\\in\\mathcal\{M\}\\times\{\\mathcal\{D\}\}\. This indicates that taxinnis en route to destinationiiwithddremaining time steps\. A new decision is made when the driver is available at the destination zone \(d=0d=0\)\. The platform monitors the origin\-destination \(OD\) queue of unassigned ride requests𝒒t∈ℕM×M\{\\bm\{q\}\}\_\{t\}\\in\{\\mathbb\{N\}\}^\{M\\times M\}and offers a price\-multiplier matrix𝒑t∈ℝ\+M×M\{\\bm\{p\}\}\_\{t\}\\in\{\\mathbb\{R\}\}\_\{\+\}^\{M\\times M\}at timett11endnote:1Here, finiteness of state and action spaces can be obtained by assuming bounded OD queues and discretizing the support of the price multipliers\. In our implementation, we let price multipliers to be continuous with the assumption that our theory from sections[2](https://arxiv.org/html/2609.36174#S2)and[3](https://arxiv.org/html/2609.36174#S3)still holds in continuous spaces\.\. To guarantee price consistency for attracted customers\([Chen et al\., 2024](https://arxiv.org/html/2609.36174#bib.bib39)\), the platform uses a lagged price𝒑t−1\{\\bm\{p\}\}\_\{t\-1\}for order matching and the price\-induced demands are non\-stationary\. Accordingly, we incorporate the pricing multiplier into the augmented major statest0:=\(𝒒t,𝒑t−1,t\)s^\{0\}\_\{t\}:=\(\{\\bm\{q\}\}\_\{t\}\\mathchar 24891\\allowbreak\{\\bm\{p\}\}\_\{t\-1\}\\mathchar 24891\\allowbreak t\)at timett, also written ass0=\(𝒒,𝒑^,t\)s^\{0\}=\(\{\\bm\{q\}\}\\mathchar 24891\\allowbreak\\hat\{\\bm\{p\}\}\\mathchar 24891\\allowbreak t\)when the time index is suppressed\.

Based on the joint augmented states\(s0,𝒔\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\), the platform chooses the major actiona0:=𝒑a^\{0\}:=\{\\bm\{p\}\}and a routing actionan:=\(j,o\)∈ℳ×\{0,1\}a^\{n\}:=\(j\\mathchar 24891\\allowbreak o\)\\in\\mathcal\{M\}\\times\\\{0\\mathchar 24891\\allowbreak 1\\\}for each idle taxi, wherejjis the destination zone andooindicates taxi occupancy \(o=1o=1for a passenger match and00for relocation\)\. The resource budgets are ofK=M2K=M^\{2\}types, each referring to the number of unmatched ordersbi​j​\(s0\):=q⁡\(i,j\)b\_\{ij\}\(s^\{0\}\):=q\(i\\mathchar 24891\\allowbreak j\)for each OD pair\(i,j\)\(i\\mathchar 24891\\allowbreak j\)\. For taxinn, we define the indicator function of accepting a ride fromiitojjasdi​j​\(an∣sn\):=𝕀⁡\{sn=\(i,0\),an=\(j,1\)\}d\_\{ij\}\(a^\{n\}\\mid s^\{n\}\):=\{\\mathbb\{I\}\}\\\{s^\{n\}=\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak 1\)\\\}\. The feasible joint minor action set is𝒜\(N\)\(s0,𝒔\):=\{𝒂∈×n=1N𝒜n\(sn\)\|∑n∈\[N\]𝕀\{sn=\(i,0\),an=\(j,1\)\}≤q\(i,j\),∀i,j∈ℳ\}\.\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\):=\\mathopen\{\\\{\}\\bm\{a\}\\in\\mathop\{\\vbox\{\\hbox\{\\LARGE$\\times$\}\}\}\_\{n=1\}^\{N\}\\mathcal\{A\}^\{n\}\(s^\{n\}\)\\,\\Big\|\\,\\sum\_\{n\\in\[N\]\}\{\\mathbb\{I\}\}\\\{s^\{n\}=\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak 1\)\\\}\\leq q\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\,\\forall i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\\mathclose\{\\\}\}\.Thus, taxis have local routing actions, but their passenger\-matching decisions cannot be selected independently because they compete for the same finite set of ride requests\.

The state transitions involve the evolution of both the fleet and the demand arrivals\. Each idle taxi \(d=0d=0\) takes a routing action, while in\-transit taxis decrement their timers deterministically, i\.e\.,

st\+1n:=\{\(j,D⁡\(i,j\)\)if​d=0,an=\(j,⋅\)\(i,d−1\)if​d≥1,∀n\.s^\{n\}\_\{t\+1\}:=\\begin\{cases\}\(j\\mathchar 24891\\allowbreak D\(i\\mathchar 24891\\allowbreak j\)\)&\\text\{if \}d=0\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak\\cdot\)\\\\ \(i\\mathchar 24891\\allowbreak d\-1\)&\\text\{if \}d\\geq 1\\\\ \\end\{cases\}\\mathchar 24891\\allowbreak\\quad\\forall n\.\(7\)
The customer request queue evolves stochastically according toqt\+1​\(i,j\)=nt\+1​\(i,j\)q\_\{t\+1\}\(i\\mathchar 24891\\allowbreak j\)=n\_\{t\+1\}\(i\\mathchar 24891\\allowbreak j\)that initializes new arrivalsnt\+1​\(i,j\)n\_\{t\+1\}\(i\\mathchar 24891\\allowbreak j\)drawn from independent Poisson distributions with per\-OD ratesθt\+1​\(i,j\)=θt\+1base​\(i,j\)​ψ​\(pt​\(i,j\)\)\\theta\_\{t\+1\}\(i\\mathchar 24891\\allowbreak j\)=\\theta^\{\\text\{base\}\}\_\{t\+1\}\(i\\mathchar 24891\\allowbreak j\)\\psi\(p\_\{t\}\(i\\mathchar 24891\\allowbreak j\)\)for alli,j∈ℳi\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}, whereψ⁡\(⋅\)\\psi\(\\cdot\)is a monotonic price\-response function and the base rate𝜽tbase\\bm\{\\theta\}^\{\\text\{base\}\}\_\{t\}is estimated from historical data\.

The system objective balances platform service fees, driver income, and operational travel costs\. Letcf≥0c\_\{f\}\\geq 0the fixed fare of a served trip,cr≥0c\_\{r\}\\geq 0denote the variable per\-unit ride price,co≥0c\_\{o\}\\geq 0denote the variable per\-unit operational cost for driving, andβ∈\(0,1\)\\beta\\in\(0\\mathchar 24891\\allowbreak 1\)be the platform service rate\. The immediate reward to the drivernnin state\(i,d\)\(i\\mathchar 24891\\allowbreak d\)when taking actionana^\{n\}is given by

r⁡\(s0,sn,an\):=\{\(1−β\)​\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]−co​D​\(i,j\)ifd=0,an=\(j,1\),−co​D​\(i,j\)ifd=0,an=\(j,0\),0if​d≥1\.r\(s^\{0\}\\mathchar 24891\\allowbreak s^\{n\}\\mathchar 24891\\allowbreak a^\{n\}\):=\\begin\{cases\}\(1\-\\beta\)\\left\[c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\right\]\-c\_\{o\}D\(i\\mathchar 24891\\allowbreak j\)&\\text\{if \}d=0\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak 1\)\\mathchar 24891\\\\ \-c\_\{o\}D\(i\\mathchar 24891\\allowbreak j\)&\\text\{if \}d=0\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\\\ 0&\\text\{if \}d\\geq 1\.\\\\ \\end\{cases\}\(8\)
The platform’s net reward is the total service\-fee revenue collected from all matched rides, i\.e\.,r0​\(s0,𝒔,𝒂\):=β​∑i,j∈ℳ∑n∈\[N\]𝕀⁡\{sn=\(i,0\),an=\(j,1\)\}​\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\,\\beta\\sum\_\{i\\mathchar 24891\\allowbreak j\\in\\mathcal\{M\}\}\\sum\_\{n\\in\[N\]\}\{\\mathbb\{I\}\}\\\{s^\{n\}=\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak a^\{n\}=\(j\\mathchar 24891\\allowbreak 1\)\\\}\\mathopen\{\[\}c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\mathclose\{\]\}\. The objective is to maximize the trade\-off between the efficiency of the platform and the fairness of the participants following \([6](https://arxiv.org/html/2609.36174#S2.E6)\)\.□\\square

## 3Utilitarian Reduction under Symmetry

This section develops the structural reduction for symmetric M2WCMDPs\. In Section[3\.1](https://arxiv.org/html/2609.36174#S3.SS1), we start by formally defining the symmetric M2WCMDPs \(Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\) and prove that an optimal policy of the symmetricρ\\rho\-M2WCMDP problem can be obtained by solving the utilitarian M2WCMDP using “permutation\-invariant” policies\. Section[3\.2](https://arxiv.org/html/2609.36174#S3.SS2)proposes an exact count aggregation method to further simplify the model\. Finally, Section[3\.3](https://arxiv.org/html/2609.36174#S3.SS3)instantiates the count formulation for the machine replacement problem of Example[3](https://arxiv.org/html/2609.36174#Thmexample3)and taxi application of Example[4](https://arxiv.org/html/2609.36174#Thmexample4)\. The resulting count aggregation MDP provides the optimization model addressed in Section[4](https://arxiv.org/html/2609.36174#S4)\.

### 3\.1Symmetricρ\\rho\-M2WCMDPs Problem Reduction

We first establish the formal conditions for a M2WCMDP to be considered symmetric\.

###### Definition 2\(Symmetric M2WCMDP\)\.

A M2WCMDP is symmetric if

1. 1\.\(Identical Minor Sub\-MDPs\)Each minor sub\-MDP is identical, i\.e\.,𝒮n=𝒮\\mathcal\{S\}^\{n\}=\\mathcal\{S\},𝒜n=𝒜\\mathcal\{A\}^\{n\}=\\mathcal\{A\},pn=pp^\{n\}=p,rn=rr^\{n\}=r,μn=μ\\mu^\{n\}=\\mu, for alln∈\[N\]n\\in\[N\], and for some\(𝒮,𝒜,p,r,μ,γ\)\(\\mathcal\{S\}\\mathchar 24891\\allowbreak\\mathcal\{A\}\\mathchar 24891\\allowbreak p\\mathchar 24891\\allowbreak r\\mathchar 24891\\allowbreak\\mu\\mathchar 24891\\allowbreak\\gamma\)tuple\.
2. 2\.\(Permutation\-Invariant Initial Distribution\)For any permutation operatorQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}, the probability of selecting the permuted initial minor stateQ​𝒔¯0Q\\bar\{\\bm\{s\}\}\_\{0\}at major states¯00\\bar\{s\}^\{0\}\_\{0\}is equal to that of selecting𝒔¯0\\bar\{\\bm\{s\}\}\_\{0\}, i\.e\.,𝝁⁡\(s¯00,𝒔¯0\)=𝝁⁡\(s¯00,Q​𝒔¯0\),∀\(s¯00,𝒔¯0\)∈𝒮0×𝒮\(N\),∀Q∈𝒢N\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)=\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\mathchar 24891\\allowbreak\\forall\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\in\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\\mathchar 24891\\allowbreak\\forall Q\\in\\mathcal\{G\}^\{N\}\.
3. 3\.\(Permutation\-Invariant Participants Influence on Major sub\-MDP\)The probability of transitioning froms0s^\{0\}tos~0\\tilde\{s\}^\{0\}under actiona0a^\{0\}when the participants apply\(𝒔,𝒂\)\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)satisfiesp0​\(𝒔~0∣s0,𝒔,a0,𝒂\)=p0​\(𝒔~0∣s0,Q​𝒔,a0,Q​𝒂\)p^\{0\}\(\\tilde\{\\bm\{s\}\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=p^\{0\}\(\\tilde\{\\bm\{s\}\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)for allQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}\. The reward function of the major sub\-MDP satisfiesr0​\(s0,𝒔,a0,𝒂\)=r0​\(s0,Q​𝒔,a0,Q​𝒂\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)for allQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}\.

The conditions for a symmetric M2WCMDP define a class of problems that is invariant under any permutation of sub\-MDP indexing\. This invariance naturally leads to the concept ofpermutation\-invariantpolicies \(see Definition 1 in[Cai et al\. \(2021\)](https://arxiv.org/html/2609.36174#bib.bib40)\)\.

###### Definition 3\(Permutation Invariant Policy\)\.

A Markov stationary policy𝝅\\bm\{\\pi\}is said to be permutation\-invariant if the probability of selecting action\(a0,𝒂\)\(a^\{0\}\\mathchar 24891\\allowbreak\{\\bm\{a\}\}\)in state\(s0,𝒔\)\(s^\{0\}\\mathchar 24891\\allowbreak\{\\bm\{s\}\}\)is equal to that of selecting the permuted actionQ​𝒂Q\{\\bm\{a\}\}in the permuted stateQ​𝒔Q\{\\bm\{s\}\}, for allQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}\. Formally, this can be expressed as𝝅\(a0,𝒂∣s0,𝒔\)=𝝅\(a0,Q𝒂∣s0,Q𝒔\)\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)=\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\), for allQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\},𝒔∈𝒮\(N\)\{\\bm\{s\}\}\\in\\mathcal\{S\}^\{\(N\)\}and𝒂∈𝒜\(N\)\{\\bm\{a\}\}\\in\\mathcal\{A\}^\{\(N\)\}\.

This naturally brings us to the question of whether this class of policies is optimal for symmetric M2WCMDPs\. The key observation is that symmetrizing a policy preserves the platform value and the mean participant value while equalizing the participant\-value vector\. The properties PI and CV ofρ\\rhoimply that this equalization cannot decrease welfare, and CVI identifies the welfare of the resulting constant vector with its common component\. This argument connects the fairness\-aware problem to the utilitarian mean in Definition[4](https://arxiv.org/html/2609.36174#Thmdefinition4)\.

###### Definition 4\(Utilitarian Approach\)\.

The utilitarian approach is the uniqueρ\\rho\-fair \(Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)\) linear mapping onℝN\{\\mathbb\{R\}\}^\{N\}, given by

ρU​\[𝒗\]:=1N​∑n=1Nvn=:v¯\.\\rho\_\{U\}\[\{\\bm\{v\}\}\]:=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}v\_\{n\}=:\\bar\{v\}\.

Additionally, the PI property implies that averaging the occupancy measure over all label permutations equalizes participant values\. Using the stationary\-policy construction associated with a feasible discounted occupancy measure \(Theorem 6\.9\.1 from[Puterman \(2005\)](https://arxiv.org/html/2609.36174#bib.bib76)\), we construct a permutation\-invariant policy from any policy, resulting in uniform state\-value representation \(Lemma[1](https://arxiv.org/html/2609.36174#Thmlemma1)\)\.

###### Lemma 1\(Uniform State\-Value Representation\)\.

If a M2WCMDP is symmetric \(Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\), then for any policy𝛑\\bm\{\\pi\}, there exists a corresponding permutation\-invariant policy𝛑¯\\bar\{\\bm\{\\pi\}\}such that the vector of expected total discounted rewards for all sub\-MDPs under𝛑¯\\bar\{\\bm\{\\pi\}\}is equal to the average of the expected total discounted rewards for each sub\-MDP, i\.e\.,𝐕0​\(𝛑¯\)=V¯0​\(𝛑\)​𝟏\\bm\{V\}\_\{0\}\(\\bar\{\\bm\{\\pi\}\}\)=\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\)\{\\bm\{1\}\}, whereV¯0​\(𝛑\):=1N​∑n=1NV0n​\(𝛑\)\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\):=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\\bm\{\\pi\}\), andV00​\(𝛑¯\)=V00​\(𝛑\)V^\{0\}\_\{0\}\(\\bar\{\\bm\{\\pi\}\}\)=V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\.

The proof is detailed in[EC\.3\.3](https://arxiv.org/html/2609.36174#A3.SS3)\. Furthermore, one can use the above lemma to show that the optimal policy for theρ\\rho\-M2WCMDP problem \([6](https://arxiv.org/html/2609.36174#S2.E6)\) under symmetry can be recovered from solving the problem with the utilitarian approach\. Our main result is presented in the following theorem\. See[EC\.3\.4](https://arxiv.org/html/2609.36174#A3.SS4)for a detailed proof\.

###### Theorem 1\(Utilitarian Reduction\)\.

Under Assumption[1](https://arxiv.org/html/2609.36174#Thmassumption1), for a symmetric M2WCMDP, letΠU,PI∗\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}be the set of optimal policies for the utilitarian approach that is permutation\-invariant, thenΠU,PI∗\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}is necessarily non\-empty and all𝛑U,PI∗∈ΠU,PI∗\\bm\{\\pi\}^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}\\in\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}satisfyGρ​\(𝛑U,PI∗\)=max𝛑∈Π⁡Gρ​\(𝛑\)\.G\_\{\\rho\}\(\\bm\{\\pi\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\tiny\{PI\}\}\}^\{\*\}\)=\\max\\limits\_\{\\bm\{\\pi\}\\in\\Pi\}G\_\{\\rho\}\(\\bm\{\\pi\}\)\.

This theorem simplifies solving theρ\\rho\-M2WCMDP problem by reducing it to an equivalent utilitarian problem, showing that any permutation\-invariant policy that is utilitarian optimal is also optimal for the originalρ\\rho\-M2WCMDP problem under the utilitarian reduction that takes the form

\(Utilitarian approach\)max𝝅∈Π⁡GU​\(𝝅\):=V00​\(𝝅\)\+λ​V¯0​\(𝝅\),\\mbox\{\(Utilitarian approach\)\}\\quad\\max\_\{\\bm\{\\pi\}\\in\\Pi\}G\_\{U\}\(\\bm\{\\pi\}\):=V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\+\\lambda\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\)\\mathchar 24891\\allowbreak\(9\)whereV¯0​\(𝝅\):=\(1/N\)​∑n=1NV0n​\(𝝅\)\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\):=\(1/N\)\\sum\_\{n=1\}^\{N\}V^\{n\}\_\{0\}\(\\bm\{\\pi\}\)\. Therefore, we can restrict the search for optimal policies to the class of permutation\-invariant policies\.

### 3\.2The Count Aggregation MDP

Under Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2), the labelled M2WCMDP is invariant under permutations of the minor indices\. This invariance property allows us to further simplify the problem using a more compact count\-based representation motivated by the symmetry simplification representation in[Gast et al\. \(2022\)](https://arxiv.org/html/2609.36174#bib.bib43)for the reduced utilitarian objective\.

Letϕ=\(f,g\)\\phi=\(f\\mathchar 24891\\allowbreak g\)denote the aggregation mapping\. The mappingf:𝒮\(N\)→\{0,…,N\}\|𝒮\|f:\\mathcal\{S\}^\{\(N\)\}\\rightarrow\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\{\\mathcal\{S\}\}\|\}maps from the labelled minor state𝒔\\bm\{s\}to a count state𝒙=f⁡\(𝒔\)\\bm\{x\}=f\(\\bm\{s\}\), where\[f\(𝒔\)\]\(s\):=∑n=1N𝕀\{sn=s\},∀s∈𝒮\[f\(\\bm\{s\}\)\]\(s\):=\\sum\_\{n=1\}^\{N\}\{\\mathbb\{I\}\}\\\{s^\{n\}=s\\\}\\mathchar 24891\\allowbreak\\forall s\\in\{\\mathcal\{S\}\}\. The corresponding count state space is

𝒮f\(N\):=\{𝒙∈\{0,…,N\}\|𝒮\|\|∑s∈𝒮x⁡\(s\)=N\}\.\\mathcal\{S\}^\{\(N\)\}\_\{f\}:=\\left\\\{\\bm\{x\}\\in\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\{\\mathcal\{S\}\}\|\}\\,\\big\|\\,\\sum\_\{s\\in\{\\mathcal\{S\}\}\}x\(s\)=N\\right\\\}\.
Thus,x⁡\(s\)x\(s\)denotes the number of minor sub\-MDPs in thess\-th state\. Similarly, for each feasible state\-action pair\(s,a\)∈𝒮×𝒜\(s\\mathchar 24891\\allowbreak a\)\\in\\mathcal\{S\}\\times\\mathcal\{A\}, the mappingg:𝒮\(N\)×𝒜\(N\)→\{0,…,N\}\|𝒮\|×\|𝒜\|g:\\mathcal\{S\}^\{\(N\)\}\\times\\mathcal\{A\}^\{\(N\)\}\\rightarrow\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}is defined by\[g⁡\(𝒔,𝒂\)\]​\(s,a\):=∑n=1N𝕀⁡\{sn=s,an=a\}\[g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\]\(s\\mathchar 24891\\allowbreak a\):=\\sum\_\{n=1\}^\{N\}\{\\mathbb\{I\}\}\\\{s\_\{n\}=s\\mathchar 24891\\allowbreak a\_\{n\}=a\\\}\. For a major states0∈𝒮0s^\{0\}\\in\\mathcal\{S\}^\{0\}and a count state𝒙∈𝒮f\(N\)\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}, the feasible count action set is

𝒜g\(N\)​\(s0,𝒙\):=\{𝒖∈\{0,…,N\}\|𝒮\|×\|𝒜\|\|∑a∈𝒜u⁡\(s,a\)=x⁡\(s\),∀s∈𝒮∑s∈𝒮∑a∈𝒜dk​\(a∣s\)​u​\(s,a\)≤bk​\(s0\),∀k∈\[K\]\}\.\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\):=\\left\\\{\\bm\{u\}\\in\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}\\,\\big\|\\,\\begin\{aligned\} &\\sum\_\{a\\in\\mathcal\{A\}\}u\(s\\mathchar 24891\\allowbreak a\)=x\(s\)\\mathchar 24891\\allowbreak\\forall s\\in\\mathcal\{S\}\\\\ &\\sum\_\{s\\in\\mathcal\{S\}\}\\sum\_\{a\\in\\mathcal\{A\}\}d\_\{k\}\(a\\mid s\)u\(s\\mathchar 24891\\allowbreak a\)\\leq b\_\{k\}\(s^\{0\}\)\\mathchar 24891\\allowbreak\\forall k\\in\[K\]\\end\{aligned\}\\right\\\}\.\(10\)
Here,u⁡\(s,a\)u\{\(s\\mathchar 24891\\allowbreak a\)\}indicates the number of minor MDPs atss\-th state that performaa\-th action\. In particular, for any labelled state𝒔\\bm\{s\}and any feasible labelled action𝒂∈𝒜\(N\)​\(s0,𝒔\)\\bm\{a\}\\in\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\), iff⁡\(𝒔\)=𝒙f\(\\bm\{s\}\)=\\bm\{x\}, theng⁡\(𝒔,𝒂\)∈𝒜g\(N\)​\(s0,𝒙\)g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\. We can then formulate the count aggregation MDP \(Definition[5](https://arxiv.org/html/2609.36174#Thmdefinition5)\)\.

###### Definition 5\(Count Aggregation MDP\)\.

The count aggregation MDPℳϕ:=\(ℳϕ0,ℳϕ\(N\)\)\\mathcal\{M\}\_\{\\phi\}:=\(\{\\mathcal\{M\}\}^\{0\}\_\{\\phi\}\\mathchar 24891\\allowbreak\{\\mathcal\{M\}\}^\{\(N\)\}\_\{\\phi\}\)derived from M2WCMDP\(ℳ0,ℳ\(N\)\)\(\{\\mathcal\{M\}\}^\{0\}\\mathchar 24891\\allowbreak\{\\mathcal\{M\}\}^\{\(N\)\}\)is composed of:

1. 1\.major aggregation sub\-MDPℳϕ0\{\\mathcal\{M\}\}^\{0\}\_\{\\phi\}with elements\(𝒮0,𝒜0,pϕ0,rϕ0,μ0,γ\)\(\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak p\_\{\\phi\}^\{0\}\\mathchar 24891\\allowbreak r\_\{\\phi\}^\{0\}\\mathchar 24891\\allowbreak\\mu^\{0\}\\mathchar 24891\\allowbreak\\gamma\)derived from major sub\-MDPs\(𝒮0,𝒜0,p0,r0,μ0,γ\)\(\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak p^\{0\}\\mathchar 24891\\allowbreak r^\{0\}\\mathchar 24891\\allowbreak\\mu^\{0\}\\mathchar 24891\\allowbreak\\gamma\)\.
2. 2\.minor aggregation sub\-MDPsℳϕ\(N\)\{\\mathcal\{M\}\}^\{\(N\)\}\_\{\\phi\}with elements\(𝒮f\(N\),𝒜g\(N\),pϕ\(N\),r¯ϕ,μf\(N\),γ\)\(\\mathcal\{S\}^\{\(N\)\}\_\{f\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{\(N\)\}\_\{g\}\\mathchar 24891\\allowbreak p^\{\(N\)\}\_\{\\phi\}\\mathchar 24891\\allowbreak\\bar\{r\}\_\{\\phi\}\\mathchar 24891\\allowbreak\\mu^\{\(N\)\}\_\{f\}\\mathchar 24891\\allowbreak\\gamma\)derived from minor sub\-MDPs\(𝒮\(N\),𝒜\(N\),p\(N\),𝒓,μ\(N\),γ\)\(\\mathcal\{S\}^\{\(N\)\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{\(N\)\}\\mathchar 24891\\allowbreak p^\{\(N\)\}\\mathchar 24891\\allowbreak\\bm\{r\}\\mathchar 24891\\allowbreak\\mu^\{\(N\)\}\\mathchar 24891\\allowbreak\\gamma\)\.

Both representations lead to the same optimization problem as established in[Gast et al\. \(2022\)](https://arxiv.org/html/2609.36174#bib.bib43)when the objective is utilitarian\. Using the count representation, the mean expected total discounted rewardV¯0​\(𝝅\)\\bar\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)for minor sub\-MDPsℳ\(N\)\\mathcal\{M\}^\{\(N\)\}with a permutation\-invariant distribution𝝁\\bm\{\\mu\}under utilitarian reduction \(Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1)\) is then equivalent to the expected total discounted mean rewardV¯0​\(𝝅ϕ\)\\bar\{V\}\_\{0\}\(\{\\bm\{\\pi\}\_\{\\phi\}\}\)for the count aggregation MDPℳϕ\\mathcal\{M\}\_\{\\phi\}given the policy𝝅ϕ:\(𝒮0,𝒮f\(N\)\)→Δ⁡\(𝒜0,𝒜g\(N\)​\(s0,𝒙\)\)\{\\bm\{\\pi\}\}\_\{\\phi\}:\(\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{S\}^\{\(N\)\}\_\{f\}\)\\rightarrow\\Delta\(\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\)under aggregation mapping with initial distribution𝝁f∈Δ⁡\(𝒮0×𝒮f\(N\)\)\\bm\{\\mu\}\_\{f\}\\in\\Delta\(\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\_\{f\}\), i\.e\.,V¯0​\(𝝅\)=1N​∑n=1NV0n​\(𝝅\)=V¯0​\(𝝅ϕ\)\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\)=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}V^\{n\}\_\{0\}\(\\bm\{\\pi\}\)=\\bar\{V\}\_\{0\}\{\(\{\\bm\{\\pi\}\}\_\{\\phi\}\}\)\. Similarly, for the major sub\-MDP,V00​\(𝝅\)=V00​\(𝝅ϕ\)V\_\{0\}^\{0\}\(\\bm\{\\pi\}\)=V\_\{0\}^\{0\}\{\(\{\\bm\{\\pi\}\}\_\{\\phi\}\}\)\. The objective in Equation \([9](https://arxiv.org/html/2609.36174#S3.E9)\) is therefore reformulated as

max𝝅ϕ⁡GUϕ​\(𝝅ϕ\):=V00​\(𝝅ϕ\)\+λ​V¯0​\(𝝅ϕ\),\\max\_\{\\bm\{\\pi\}\_\{\\phi\}\}G^\{\\phi\}\_\{U\}\(\\bm\{\\pi\}\_\{\\phi\}\):=V^\{0\}\_\{0\}\(\\bm\{\\pi\}\_\{\\phi\}\)\+\\lambda\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\_\{\\phi\}\)\\mathchar 24891\\allowbreak\(11\)where

V00​\(𝝅ϕ\):=𝔼𝝅ϕ​\[∑t=0∞γt​rϕ0​\(st0,𝒙t,at0,𝒖t\)\|\(s00,𝒙0\)∼𝝁f\],V^\{0\}\_\{0\}\(\\bm\{\\pi\}\_\{\\phi\}\):=\\mathbb\{E\}\_\{\\bm\{\\pi\}\_\{\\phi\}\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r^\{0\}\_\{\\phi\}\(s^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{x\}\_\{t\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{u\}\_\{t\}\)\\,\\Big\|\\,\(s^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\_\{0\}\)\\sim\\bm\{\\mu\}\_\{f\}\\right\]\\mathchar 24891\\allowbreakand

V¯0​\(𝝅ϕ\):=𝔼𝝅ϕ​\[∑t=0∞γt​r¯ϕ​\(st0,𝒙t,at0,𝒖t\)\|\(s00,𝒙0\)∼𝝁f\]\.\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\_\{\\phi\}\):=\\mathbb\{E\}\_\{\{\\bm\{\\pi\}\}\_\{\\phi\}\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\bar\{r\}\_\{\\phi\}\(s^\{0\}\_\{t\}\\mathchar 24891\\allowbreak\\bm\{x\}\_\{t\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\_\{t\}\)\\,\\Big\|\\,\(s^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\_\{0\}\)\\sim\\bm\{\\mu\}\_\{f\}\\right\]\.
Here, the utilitarian reward aggregation mapping for minor sub\-MDPs isr¯ϕ​\(s0,𝒙,a0,𝒖\):=1N​∑s∈𝒮,a∈𝒜u⁡\(s,a\)​r​\(s0,s,a0,a\)\\bar\{r\}\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\):=\\frac\{1\}\{N\}\\sum\_\{s\\in\\mathcal\{S\}\\mathchar 24891\\allowbreak a\\in\\mathcal\{A\}\}u\(s\\mathchar 24891\\allowbreak a\)r\(s^\{0\}\\mathchar 24891\\allowbreak s\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak a\)\. Similarly, the reward aggregation mapping for the major reward isrϕ0​\(s0,𝒙,a0,𝒖\)=r0​\(s0,𝒔¯,a0,𝒂¯\)r^\{0\}\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)=r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)for any arbitrary labelled state\-action pair\(𝒔¯,𝒂¯\)\(\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)satisfyingf⁡\(𝒔¯\)=𝒙f\(\\bar\{\\bm\{s\}\}\)=\\bm\{x\}andg⁡\(𝒔¯,𝒂¯\)=𝒖g\(\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)=\\bm\{u\}\. Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)ensures that this value is independent of the representative labelled pair\. The count transition kernel and initial distribution ofℳϕ\{\\mathcal\{M\}\}\_\{\\phi\}are derived in[EC\.5](https://arxiv.org/html/2609.36174#A5)\.

By Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1), an optimal policy for the count aggregation MDP, after permutation\-invariant lifting \(see Equation \([EC\.11](https://arxiv.org/html/2609.36174#A6.E11)\) in[EC\.6](https://arxiv.org/html/2609.36174#A6)\), is optimal for the original symmetricρ\\rho\-M2WCMDP\. An exact linear programming \(LP\) method is provided to solve the utilitarian\-reduced count aggregation M2WCMDP in[EC\.6](https://arxiv.org/html/2609.36174#A6)\. However, due to the curse of dimensionality, solving this LP is computationally prohibitive despite the reduction\. We thus design a count\-proportion\-based learning approach in Section[4](https://arxiv.org/html/2609.36174#S4)\.

### 3\.3Examples of Count Reformulation

We now detail the count reformulation to the MRP and TDPP applications in Section[2\.3](https://arxiv.org/html/2609.36174#S2.SS3)\.

###### Example 5\(Count Reformulation of the MRP Example[3](https://arxiv.org/html/2609.36174#Thmexample3)\)\.

For the binary\-action, single\-resource MRP, letx⁡\(s\)x\(s\)denote the number of machines in deterioration statess, and letu⁡\(s\)u\(s\)denote the number replaced\. The count state and feasible count action are𝒮f\(N\):=\{𝒙∈ℕS∣∑s∈𝒮x⁡\(s\)=N\}\\mathcal\{S\}^\{\(N\)\}\_\{f\}:=\\\{\\bm\{x\}\\in\{\\mathbb\{N\}\}^\{S\}\\mid\\sum\_\{s\\in\{\\mathcal\{S\}\}\}x\(s\)=N\\\}, and𝒜g\(N\)\(𝒙\):=\{𝒖∈ℕS∣0≤u\(s\)≤x\(s\),∑s∈𝒮u\(s\)≤b\}\.\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(\\bm\{x\}\):=\\\{\\bm\{u\}\\in\{\\mathbb\{N\}\}^\{S\}\\mid 0\\leq u\(s\)\\leq x\(s\)\\mathchar 24891\\allowbreak\\ \\sum\_\{s\\in\{\\mathcal\{S\}\}\}u\(s\)\\leq b\\\}\.□\\square

###### Example 6\(Count Reformulation of the Taxi Example[4](https://arxiv.org/html/2609.36174#Thmexample4)\)\.

For the taxi application, the count action space is𝒮f\(N\):=\{𝒙∈ℕM×\|𝒟\|\|∑i∈ℳ∑d∈𝒟x⁡\(i,d\)=N\}\\mathcal\{S\}^\{\(N\)\}\_\{f\}:=\\mathopen\{\\\{\}\\bm\{x\}\\in\{\\mathbb\{N\}\}^\{M\\times\|\{\\mathcal\{D\}\}\|\}\\,\\big\|\\,\\sum\_\{i\\in\\mathcal\{M\}\}\\sum\_\{d\\in\{\\mathcal\{D\}\}\}x\(i\\mathchar 24891\\allowbreak d\)=N\\mathclose\{\\\}\}, wherex⁡\(i,d\)x\(i\\mathchar 24891\\allowbreak d\)is the number of taxis that are en route to zoneiiand will be available afterddtime steps\.

Let𝒖:=\(𝒖1,𝒖2\)\\bm\{u\}:=\(\\bm\{u\}^\{1\}\\mathchar 24891\\allowbreak\\bm\{u\}^\{2\}\), whereut1​\(i,j\)u^\{1\}\_\{t\}\(i\\mathchar 24891\\allowbreak j\)andut2​\(i,j\)u^\{2\}\_\{t\}\(i\\mathchar 24891\\allowbreak j\)denote, respectively, the number of idle taxis relocated and matched from origin zonei∈ℳi\\in\{\\mathcal\{M\}\}to target zonej∈ℳj\\in\{\\mathcal\{M\}\}\. For a major states0=\(𝒒,𝒑^,t\)s^\{0\}=\(\{\\bm\{q\}\}\\mathchar 24891\\allowbreak\\hat\{\\bm\{p\}\}\\mathchar 24891\\allowbreak t\), the feasible count action set is𝒜g\(N\)\(s0,𝒙\):=\{\(𝒖1,𝒖2\)∈ℕM×M×ℕM×M\|\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\):=\\mathopen\{\\\{\}\(\\bm\{u\}^\{1\}\\mathchar 24891\\allowbreak\\bm\{u\}^\{2\}\)\\in\{\\mathbb\{N\}\}^\{M\\times M\}\\times\{\\mathbb\{N\}\}^\{M\\times M\}\\,\\big\|\\,\\mathclose\{\}∑j∈ℳ\(u1\(i,j\)\+u2\(i,j\)\)=x\(i,0\),∀i∈ℳ;u2\(i,j\)≤s0\(i,j\),∀i,j∈ℳ\}\\mathopen\{\}\\sum\_\{j\\in\\mathcal\{M\}\}\\bigl\(u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\bigr\)=x\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak\\forall i\\in\\mathcal\{M\}\\mathchar 24635\\allowbreak\\,\\,u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\leq s^\{0\}\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\forall i\\mathchar 24891\\allowbreak j\\in\\mathcal\{M\}\\mathclose\{\\\}\}\. Ford≥1d\\geq 1, the only feasible individual action is the wait\-in\-place actionawaita^\{\\text\{wait\}\}\. The corresponding count therefore must satisfyu⁡\(\(i,d\),await\)=x⁡\(i,d\)u\(\(i\\mathchar 24891\\allowbreak d\)\\mathchar 24891\\allowbreak a^\{\\text\{wait\}\}\)=x\(i\\mathchar 24891\\allowbreak d\)automatically\.

With the convention𝒙t​\(i,Dmax\+1\):=0\\bm\{x\}\_\{t\}\(i\\mathchar 24891\\allowbreak D\_\{\\max\}\+1\):=0andD⁡\(j,i\)=0D\(j\\mathchar 24891\\allowbreak i\)=0if and only ifj=ij=i, the in\-transit counts evolve according to𝒙t\+1\(i,d\):=𝒙t\(i,d\+1\)\+∑j∈ℳ\(ut1\(j,i\)\+ut2\(j,i\)\)𝕀\{D\(j,i\)=d\},\\bm\{x\}\_\{t\+1\}\(i\\mathchar 24891\\allowbreak d\):=\\bm\{x\}\_\{t\}\(i\\mathchar 24891\\allowbreak\{d\+1\}\)\+\{\\sum\_\{j\\in\\mathcal\{M\}\}\\mathopen\{\(\}u^\{1\}\_\{t\}\(j\\mathchar 24891\\allowbreak i\)\+u^\{2\}\_\{t\}\(j\\mathchar 24891\\allowbreak i\)\\mathclose\{\)\}\\mathbb\{I\}\\\{D\(j\\mathchar 24891\\allowbreak i\)=d\\\}\}\\mathchar 24891\\allowbreakfor alld∈𝒟d\\in\{\\mathcal\{D\}\}and alli∈ℳi\\in\{\\mathcal\{M\}\}\.

Under the utilitarian reduction \(Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1)\) and count aggregationϕ\\phi, the platform’s reward is a fraction of all ride rewardsrϕ0​\(s0,𝒙,𝒖\):=β​∑i∈ℳ,j∈ℳ\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]​u2​\(i,j\)r^\{0\}\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak\\bm\{u\}\):=\\beta\\sum\_\{i\\in\\mathcal\{M\}\\mathchar 24891\\allowbreak j\\in\\mathcal\{M\}\}\\mathopen\{\[\}c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\mathclose\{\]\}u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\. The mean participant reward is determined by the total successful ride assignments to locally available taxis and then normalized byNNto get the average per driver reward, i\.e\.,r¯ϕ​\(s0,𝒙,𝒖\):=1N​\[\(1−β\)​∑i,j∈ℳ\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]​u2​\(i,j\)−co​∑i,j∈ℳD⁡\(i,j\)​\(u1​\(i,j\)\+u2​\(i,j\)\)\]\\bar\{r\}\_\{\\phi\}\\mathopen\{\(\}s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak\\bm\{u\}\\mathclose\{\)\}:=\\frac\{1\}\{N\}\\mathopen\{\[\}\(1\-\\beta\)\\sum\_\{i\\mathchar 24891\\allowbreak j\\in\\mathcal\{M\}\}\\mathopen\{\[\}c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\mathclose\{\]\}u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\-c\_\{o\}\\sum\_\{i\\mathchar 24891\\allowbreak j\\in\\mathcal\{M\}\}D\(i\\mathchar 24891\\allowbreak j\)\\mathopen\{\(\}u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\mathclose\{\)\}\\mathclose\{\]\}\. The objective is formulated by \([11](https://arxiv.org/html/2609.36174#S3.E11)\)\.□\\square

## 4Count\-Proportion\-Based Deep Reinforcement Learning

For large\-scale optimization and when the model\(ℳ0,ℳ\(N\)\)\(\{\\mathcal\{M\}\}^\{0\}\\mathchar 24891\\allowbreak\\mathcal\{M\}^\{\(N\)\}\)is unknown, we develop a count\-proportion\-based deep reinforcement learning \(CP\-DRL\) method for approximately solving the utilitarian\-reduced count M2WCMDP in \([11](https://arxiv.org/html/2609.36174#S3.E11)\)\. In Section[4\.1](https://arxiv.org/html/2609.36174#S4.SS1), we introduce the architecture of our stochastic policy network, which is explicitly designed to search within a scalable parameterized subclass of the permutation\-invariant policies\. The fundamental challenge in this setting is bridging the gap between continuous deep neural network outputs and the combinatorial constraints of the feasible discrete count action space\. Therefore, in Section[4\.2](https://arxiv.org/html/2609.36174#S4.SS2), we introduce a parametrized priority\-based sampling procedure\. This ensures operational feasibility while maintaining sufficient stochasticity for exploration\.

### 4\.1Stochastic Policy Neural Network

One key property of the count aggregation MDP is that the dimensions of the state space𝒮f\(N\)\\mathcal\{S\}^\{\(N\)\}\_\{f\}and the action space𝒜g\(N\)\\mathcal\{A\}^\{\(N\)\}\_\{g\}are independent of the numberNNof minor sub\-MDPs, although the cardinalities of the feasible count\-action sets still depend onNN\. For𝒙∈𝒮f\(N\)\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}, we define the count minor state proportion as𝒙¯:=𝒙/N∈\[0,1\]\|𝒮\|\\bar\{\\bm\{x\}\}:=\\bm\{x\}/N\\in\[0\\mathchar 24891\\allowbreak 1\]^\{\|\{\\mathcal\{S\}\}\|\}to further simplify the analysis and eliminate the explicit dependence onNN\. The role of the stochastic policy network is to convert the state\(s0,𝒙¯\)\(s\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{x\}\}\)to a distribution over\(a0,𝒖\)\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)from which a sample is drawn\. We represent the stochastic policy through a chain\-rule decomposition, i\.e\.,𝝅\(a0,𝒖∣s0,𝒙\)=𝝅0\(a0∣s0,𝒙\)𝝅\(N\)\(𝒖∣s0,𝒙,a0\)\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)=\\bm\{\\pi\}^\{0\}\(a^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\bm\{\\pi\}^\{\(N\)\}\(\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\)\. The conditional relationship betweena0a^\{0\}and𝒖\\bm\{u\}is implicitly represented within the joint neural\-network architecture\.

To deal with the exponential size of the support of𝒖∈𝒜g\(N\)\\bm\{u\}\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}and with the challenges associated to the state dependent constraints \([10](https://arxiv.org/html/2609.36174#S3.E10)\), we represent the conditional minor policy𝝅\(N\)​\(𝒖∣s0,𝒙,a0\)\\bm\{\\pi\}^\{\(N\)\}\(\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\)using an intermediate parametrization\(𝒑~,𝑾~,𝑼\)\(\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\), where\(𝒑~,𝑾~,𝑼\)\(\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\)are the state\-dependent network outputs jointly generated with the major\-action outputa0a^\{0\}and used for a priority\-based sampling procedure𝝅𝒑~,𝑾~,𝑼​\(𝒖∣s0,𝒙\)\\bm\{\\pi\}\_\{\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\}\(\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)to be described next\. The neural networks for producinga0a\_\{0\}and\(𝒑~,𝑾~,𝑼\)\(\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\)are combined to reduce the number of learned parameters and share a common representation of the state\(s0,𝒙\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\. Since count\-state proportions alone can erase scale information needed to evaluate resource capacities, the raw count tuple\(s0,𝒙\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)is passed explicitly to the final integer sampler\. Figure[1](https://arxiv.org/html/2609.36174#S4.F1)presents a diagram representing our proposed implementation\.

Figure 1:CP\-based Stochastic Policy Neural Network![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/policy.png)
### 4\.2Parameterized Priority\-based Sampling Procedure

The priority\-based sampling procedure𝝅𝒑~,𝑾~,𝑼​\(𝒖∣s0,𝒙\)\\bm\{\\pi\}\_\{\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\}\(\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)first employs a parameter𝒑~∈\[0,1\]K\\tilde\{\{\\bm\{p\}\}\}\\in\[0\\mathchar 24891\\allowbreak 1\]^\{K\}controlling the proportion of available resources used by the state\-action count plan\. It then exploits a predefined partition of the action space\{𝒜c\}c∈𝒞\\\{\\mathcal\{A\}\_\{c\}\\\}\_\{c\\in\{\\mathcal\{C\}\}\}that identifies\|𝒞\|\|\{\\mathcal\{C\}\}\|categories of actions, in order to control the targeted count of planned actions within each categorycc, for each statess\. This is done using a category\-allocation weight matrix𝑾~∈Δ​\(𝒞\)\|𝒮\|\\widetilde\{\\bm\{W\}\}\\in\\Delta\(\{\\mathcal\{C\}\}\)^\{\|\\mathcal\{S\}\|\}that determines integer category budgetsB⁡\(s,c\)≈𝑾~​\(s,c\)​x​\(s\)B\(s\\mathchar 24891\\allowbreak c\)\\approx\\widetilde\{\\bm\{W\}\}\(s\\mathchar 24891\\allowbreak c\)x\(s\)for everys∈𝒮s\\in\\mathcal\{S\}, satisfying∑c∈𝒞B⁡\(s,c\)=x⁡\(s\)\\sum\_\{c\\in\{\\mathcal\{C\}\}\}B\(s\\mathchar 24891\\allowbreak c\)=x\(s\)\. Finally, a priority score matrix𝑼∈\[0,1\]\|𝒮\|×\|𝒜\|\{\\bm\{U\}\}\\in\[0\\mathchar 24891\\allowbreak 1\]^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}promotes that each state\-action pair be sampled proportionally toU⁡\(s,a\)U\(s\\mathchar 24891\\allowbreak a\)within each category22endnote:2If all eligible priority scores are zero, a uniform value of1/\(\|𝒮\|×\|𝒜\|\)1/\(\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\)is used\.\. Algorithm[1](https://arxiv.org/html/2609.36174#alg1)presents the details of the state\-action count sampling procedure, ensuring a feasible realization of the count action matrix𝒖\\bm\{u\}\.

Algorithm 1Priority\-Score based State\-Action Count Sampling Procedure \(𝒖∼𝝅𝒑~,𝑾~,𝑼\(⋅∣s0,𝒙\)\\bm\{u\}\\sim\\bm\{\\pi\}\_\{\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\widetilde\{\\bm\{W\}\}\\mathchar 24891\\allowbreak\\bm\{U\}\}\(\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\)Input:Major state

s0s^\{0\}, count state

𝒙\\bm\{x\}
Parameters:Resource\-to\-use proportion

𝒑~\\tilde\{\\bm\{p\}\}, category allocation weights

𝑾~\\widetilde\{\\bm\{W\}\}, priority matrix

𝑼\{\\bm\{U\}\}
Assumption:There exists

a^∈𝒜\\hat\{a\}\\in\\mathcal\{A\}such that

dk​\(a^∣s\)=0d\_\{k\}\(\\hat\{a\}\\mid s\)=0for all

s,ks\\mathchar 24891\\allowbreak k
Initialize:

𝒃~←𝒃⁡\(s0\)⋅𝒑~\\tilde\{\{\\bm\{b\}\}\}\\leftarrow\{\\bm\{b\}\}\(s^\{0\}\)\\cdot\\tilde\{\\bm\{p\}\},

𝒖←𝟎\|𝒮\|×\|𝒜\|\\bm\{u\}\\leftarrow\\bm\{0\}\_\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}
\#Identify categories’ count budgets closest to fractional targets

B\(s,ci\)←\[arg​min𝔟∈ℕ\|𝒞\|:∑c∈𝒞𝔟⁡\(c\)=x⁡\(s\)∑c∈𝒞\|𝔟\(c\)−x\(s\)W~\(s,c\)\|\]\(ci\)B\(s\\mathchar 24891\\allowbreak c\_\{i\}\)\\leftarrow\\mathopen\{\[\}\\operatorname\*\{arg\\,min\}\_\{\\mathfrak\{b\}\\in\\mathbb\{N\}^\{\|\\mathcal\{C\}\|\}:\\sum\_\{c\\in\\mathcal\{C\}\}\\mathfrak\{b\}\(c\)=x\(s\)\}\\sum\_\{c\\in\\mathcal\{C\}\}\|\\mathfrak\{b\}\(c\)\-x\(s\)\\widetilde\{W\}\(s\\mathchar 24891\\allowbreak c\)\|\\mathclose\{\]\}\(c\_\{i\}\),

∀ci∈𝒞\\forall c\_\{i\}\\in\\mathcal\{C\},

∀s∈𝒮\\forall s\\in\{\\mathcal\{S\}\}
for

i∈\{1,…,\|𝒞\|\}i\\in\\\{1\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak\|\{\\mathcal\{C\}\}\|\\\}do

ℱ←\{\(s,a\)∈𝒮×𝒜ci∣B⁡\(s,ci\)=0\}\{\\mathcal\{F\}\}\\leftarrow\\\{\(s\\mathchar 24891\\allowbreak a\)\\in\{\\mathcal\{S\}\}\\times\{\\mathcal\{A\}\}\_\{c\_\{i\}\}\\mid\{B\}\(s\\mathchar 24891\\allowbreak c\_\{i\}\)=0\\\}\#Create set of exhausted state\-action pairs in

𝒮×𝒜ci\{\\mathcal\{S\}\}\\times\{\\mathcal\{A\}\}\_\{c\_\{i\}\}
while

\|ℱ\|<\|𝒮\|×\|𝒜ci\|\|\{\\mathcal\{F\}\}\|<\|\\mathcal\{S\}\|\\times\|\{\\mathcal\{A\}\}\_\{c\_\{i\}\}\|do

Sample

\(s,a\)∝𝑼\(s,a\)𝕀\{\(s,a\)∈𝒮×𝒜ci∖ℱ\}\(s\\mathchar 24891\\allowbreak a\)\\propto\{\\bm\{U\}\}\(s\\mathchar 24891\\allowbreak a\)\{\\mathbb\{I\}\}\\\{\(s\\mathchar 24891\\allowbreak a\)\\in\{\\mathcal\{S\}\}\\times\{\\mathcal\{A\}\}\_\{c\_\{i\}\}\\setminus\{\\mathcal\{F\}\}\\\}\#Sample unexhausted pair based on priority

if

dk​\(a∣s\)≤b~kd\_\{k\}\(a\\mid s\)\\leq\\tilde\{b\}\_\{k\}for all

kkthen

B⁡\(s,ci\)←B⁡\(s,ci\)−1B\(s\\mathchar 24891\\allowbreak c\_\{i\}\)\\leftarrow B\(s\\mathchar 24891\\allowbreak c\_\{i\}\)\-1\#Update action count budget

b~k←b~k−dk​\(a∣s\)\\tilde\{b\}\_\{k\}\\leftarrow\\tilde\{b\}\_\{k\}\-d\_\{k\}\(a\\mid s\)for all

kk\#Update left\-over resources

u⁡\(s,a\)←u⁡\(s,a\)\+1u\(s\\mathchar 24891\\allowbreak a\)\\leftarrow u\(\{s\\mathchar 24891\\allowbreak a\}\)\+1\#Add

\(s,a\)\(s\\mathchar 24891\\allowbreak a\)to state\-action count plan

ℱ←ℱ∪\{\(s,a\)∈𝒮×𝒜ci∣B⁡\(s,ci\)=0\}\{\\mathcal\{F\}\}\\leftarrow\{\\mathcal\{F\}\}\\cup\\\{\(s\\mathchar 24891\\allowbreak a\)\\in\{\\mathcal\{S\}\}\\times\{\\mathcal\{A\}\}\_\{c\_\{i\}\}\\mid B\(s\\mathchar 24891\\allowbreak c\_\{i\}\)=0\\\}\#Update

ℱ\{\\mathcal\{F\}\}
else

ℱ←ℱ∪\{\(s,a\)\}\{\\mathcal\{F\}\}\\leftarrow\{\\mathcal\{F\}\}\\cup\\\{\(s\\mathchar 24891\\allowbreak a\)\\\}\#Add

\(s,a\)\(s\\mathchar 24891\\allowbreak a\)to exhausted pairs

endif

endwhile

endfor

u⁡\(s,a^\)←u⁡\(s,a^\)\+x⁡\(s\)−∑au⁡\(s,a\)u\(s\\mathchar 24891\\allowbreak\\hat\{a\}\)\\leftarrow u\(s\\mathchar 24891\\allowbreak\\hat\{a\}\)\+x\(s\)\-\\sum\_\{a\}u\(s\\mathchar 24891\\allowbreak a\)for all

ss\#Complete plan with free action counts

Return:Count action matrix

𝒖\\bm\{u\}

We next establish two properties of the sampler\.[EC\.4](https://arxiv.org/html/2609.36174#A4)provides the proof\.

###### Lemma 2\(Feasibility and deterministic coverage\)\.

Given that there exists a fallback actiona^∈𝒜\\hat\{a\}\\in\\mathcal\{A\}such thatdk​\(a^∣s\)=0d\_\{k\}\(\\hat\{a\}\\mid s\)=0for alls∈𝒮s\\in\\mathcal\{S\}and allk∈\[K\]k\\in\[K\], the sampling procedure described in Algorithm[1](https://arxiv.org/html/2609.36174#alg1), has the following properties:

1. 1\.ℙ⁡\(𝒖∈𝒜g\(N\)​\(s0,𝒙\)\)=1\\mathbb\{P\}\(\\bm\{u\}\\in\{\\mathcal\{A\}^\{\(N\)\}\_\{g\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\)=1for alls0∈𝒮0s^\{0\}\\in\\mathcal\{S\}^\{0\}and𝒙∈𝒮f\(N\)\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\.
2. 2\.If\{𝒜c\}c∈𝒞=𝒜\\\{\{\\mathcal\{A\}\}\_\{c\}\\\}\_\{c\\in\{\\mathcal\{C\}\}\}=\{\\mathcal\{A\}\}, then for any deterministic policy𝝅ϕD:𝒮0×𝒮f\(N\)→\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}:\\mathcal\{S\}^\{0\}\\times\{\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\rightarrow\{0,…,N\}\|𝒮\|×\|𝒜\|\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\{\\mathcal\{S\}\}\|\\times\|\{\\mathcal\{A\}\}\|\}satisfying𝝅ϕD​\(s0,𝒙\)∈𝒜g\(N\)​\(s0,𝒙\)\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\), there exists a mappingh:𝒮0×𝒮f\(N\)→\[0,1\]K×Δ​\(𝒞\)\|𝒮\|×\[0,1\]\|𝒮\|×\|𝒜\|h:\\mathcal\{S\}^\{0\}\\times\{\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\rightarrow\[0\\mathchar 24891\\allowbreak 1\]^\{K\}\\times\\Delta\(\{\\mathcal\{C\}\}\)^\{\|\\mathcal\{S\}\|\}\\times\[0\\mathchar 24891\\allowbreak 1\]^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}such that a sample𝒖∼𝝅h⁡\(s0,𝒙\)\(⋅∣s0,𝒙\)=𝝅ϕD\(s0,𝒙\)\\bm\{u\}\\sim\\bm\{\\pi\}\_\{h\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\}\(\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)=\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)with probability one\.

The assumption of a fallback actiona^\\hat\{a\}ensures that the system always has a valid default operation\. Consequently, the sampling procedure strictly guarantees operational feasibility \(Property 1\) by reverting toa^\\hat\{a\}, so both the count\-conservation constraints and the resource constraints are satisfied for every realization of the sampler\.

Property 2 establishes that the sampling parameterization does not exclude any deterministic feasible count policy when each action is represented by a separate category\. In this case, the category weights can encode the desired action proportions directly, and the sampler can reproduce any feasible integer count action with probability one\. This indicates that, under singleton categories, an optimal deterministic count policy is contained in the sampler\-induced policy class\.

There is, however, a trade\-off in the number of action categories\. Using more categories increases representational flexibility, and Property 2 provides exact deterministic representability in the limiting singleton partition\. On the other hand, a larger number of categories increases the number of policy\-network outputs, the number of sequential sampling stages, and the computational burden of action generation and learning\. Coarser partitions exploit application\-specific structure, but they generally restrict the determinism of the policy class, because actions within the same category are controlled only through their relative priority scores\.

## 5Numerical Results

This section evaluates the proposed CP\-DRL framework \(Section[4](https://arxiv.org/html/2609.36174#S4)\) in two count\-reformulated problem settings of Section[3\.3](https://arxiv.org/html/2609.36174#S3.SS3)\. In all experiments, CP\-DRL is trained to maximize the utilitarian\-reduced count objectiveGUϕ​\(𝝅ϕ\)G^\{\\phi\}\_\{U\}\(\\bm\{\\pi\}\_\{\\phi\}\)in \([11](https://arxiv.org/html/2609.36174#S3.E11)\) and evaluated on the original fairness\-aware objectiveGρ​\(𝝅\)G\_\{\\rho\}\(\\bm\{\\pi\}\)in \([6](https://arxiv.org/html/2609.36174#S2.E6)\) using a permutation\-invariant disaggregation ruleϕ−1\\phi^\{\-1\}\. We first consider the machine replacement problem \(see Example[5](https://arxiv.org/html/2609.36174#Thmexample5)\), where exactρ\\rho\-optimal solutions can be calculated for small instances\. This setting provides a controlled environment for showing the empirical consistency and operational implications of the utilitarian reduction in Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1)\. The taxi dispatching problem \(see Example[6](https://arxiv.org/html/2609.36174#Thmexample6)\) provides a complementary setting to evaluate the full major\-minor coordination with endogenous pricing and routing decisions\.

The MRP does not include a major sub\-MDP and is therefore a classical WCMDP\. This makes it a useful diagnostic experiment to isolate the approximation quality of CP\-DRL and the computational benefit of the count\-proportion representation without confounding effects from major\-state dynamics\. Among the DRL algorithms considered in preliminary experiments, theProximal Policy Optimization\(PPO\) algorithm\([Schulman et al\., 2017](https://arxiv.org/html/2609.36174#bib.bib36)\)consistently delivers the most stable and high\-quality performance, and is thus used as the main algorithm for our CP\-DRL approach\.

### 5\.1The Machine Replacement Problem

The MRP provides a scalable framework for evaluating the CP\-DRL approach as problem size and complexity increase\. We focus on a single resource \(K=1K=1\) and binary action \(A=2A=2\) for each machine, which places the problem within the restless multi\-armed bandits \(RMABs\) subclass of WCMDPs\([Whittle, 1988](https://arxiv.org/html/2609.36174#bib.bib75);[Zhang, 2022](https://arxiv.org/html/2609.36174#bib.bib5)\)\. For selected indexable instances, this structure enables direct comparison with the Whittle index policy, a strong structure\-exploiting benchmark for RMABs\.

#### 5\.1\.1Implementation Details

For each state\-action pair, resource consumptiond⁡\(a∣s\)d\(a\\mid s\\,\)is 1 for replacements and 0 for operations, with up tobbreplacements per time step\. We convert the normalized cost functionc⁡\(s,a\)∈\[0,1\]c\(s\\mathchar 24891\\allowbreak a\)\\in\[0\\mathchar 24891\\allowbreak 1\]into rewards by applying the positive shift affine transformationr⁡\(s,a\):=1−c⁡\(s,a\)r\(s\\mathchar 24891\\allowbreak a\):=1\-c\(s\\mathchar 24891\\allowbreak a\)\. Machines degrade if not replaced and remain in the most deteriorated stateSSuntil replaced\. We choose operational and replacement costs across two presets to capture different cost structures, characterized by different replacement cost constant coefficients \(RCCCs\):i\) Exponential\-RCCCandii\) Quadratic\-RCCC\. Refer to[EC\.7\.1](https://arxiv.org/html/2609.36174#A7.SS1)for cost structures and transition probabilities\.

##### Objectives and Evaluation Protocol

We use the GGF in Example[1](https://arxiv.org/html/2609.36174#Thmexample1)as the primary fairness criterion for testing the symmetry\-based reduction\. GGF weights decay exponentially with a factor of 2, defined aswn=1/2nw\_\{n\}=1/2^\{n\}, and normalized to sum to 1\. The goal is to find a fair policy that maximizes the GGF score over the expected total discounted mean rewards under the count aggregation MDP\. We use a uniform distributionμ\(N\)\\mu^\{\(N\)\}over𝒮\(N\)\\mathcal\{S\}^\{\(N\)\}and set the discount factorγ=0\.95\\gamma=0\.95\. We use Monte Carlo simulations to evaluate policies overNM​CN\_\{MC\}trajectories truncated at time lengthTT\. We chooseNM​CN\_\{MC\}= 1,000 andTT= 300 across all experiments\.

##### CP\-DRL and Benchmarks

We evaluate two CP\-DRL training regimes\. The single\-task CP\-DRL is trained separately for each machine countNN\. CP\-DRL\(MT\) is a multi\-task \(MT\) extension trained jointly overN∈\{2,3,4,5\}N\\in\\\{2\\mathchar 24891\\allowbreak 3\\mathchar 24891\\allowbreak 4\\mathchar 24891\\allowbreak 5\\\}withS=3S=3, randomly switching the configuration at the end of each episode over 2,000 training episodes\. Both variants use the same count\-proportion representation and priority\-based sampler described in Section[4](https://arxiv.org/html/2609.36174#S4)\. Hyperparameters for the CP\-DRL algorithm are in[EC\.7\.2](https://arxiv.org/html/2609.36174#A7.SS2)\. We compare CP\-DRL against seven benchmarks, including optimal solutions \(OPT\) from the utilitarian\-reduced dual LP model \([EC\.10](https://arxiv.org/html/2609.36174#A6.E10)\) for small instances solved with Gurobi 10\.0\.3, the Whittle index policy \(WIP\) for RMABs, and a random \(RDM\) agent that selects actions randomly at each time step and averages the results over 10 independent runs\. Additionally, we implemented a simple DRL baseline, Vanilla\-DRL \(V\-DRL\), with a utilitarian objective\. The stochastic policy network employs a fully connected neural network that maps the vector𝒔\\bm\{s\}to aNN\-dimensional probability vector\. We also implemented two heuristics to complement the random agent approach\. The oldest\-first \(OFT\) approach selects the machine in the worst state, while the myopic \(MYP\) selects the machine that maximizes immediate reward\. We finally implemented an equal\-resources \(EQR\) approach based on[Li and Varakantham \(2022\)](https://arxiv.org/html/2609.36174#bib.bib15), which imposes that each machine be replaced once everyNNsteps to ensure an equal distribution of resources\.

#### 5\.1\.2Results

We designed a series of experiments to test the GGF\-optimality and fairness of our CP\-DRL algorithm\. Additional experiments on scalability and efficiency are provided in[EC\.7\.3](https://arxiv.org/html/2609.36174#A7.SS3)\.

Table 1:GGF Scores \(Exponential\-RCCC\)Note\.CP\-DRL and CP\-DRL\(MT\) entries are mean±\\pmstandard deviation over five random seeds\.

Table 2:GGF Scores \(Quadratic\-RCCC\)Note\.CP\-DRL and CP\-DRL\(MT\) entries are mean±\\pmstandard deviation over five random seeds\.

Tables[1](https://arxiv.org/html/2609.36174#S5.T1.fig1)and[2](https://arxiv.org/html/2609.36174#S5.T2.fig1)report the GGF scores for the Exponential\-RCCC and Quadratic\-RCCC instances, respectively\. WIP and RDM are evaluated with 1,000 Monte Carlo runs, and their standard deviations are negligible and omitted\. Bold font indicating the best GGF scores at each row excluding optimal values\. As shown in Table[1](https://arxiv.org/html/2609.36174#S5.T1.fig1), CP\-DRL\(MT\) consistently achieves scores very close to the OPT values as the number of machines increases from 2 to 4\. For the 5\-machine case, CP\-DRL\(MT\) shows slightly better performance than the single\-task CP\-DRL\. In Table[2](https://arxiv.org/html/2609.36174#S5.T2.fig1), the single\- and multi\-task CP\-DRL agents show slight variations in performance across different machine numbers\. ForN=5N=5, CP\-DRL achieves the best GGF score, slightly outperforming WIP\. In both cases, the CP\-DRL approach outperforms Vanilla\-DRL, the three heuristic methods, and the random benchmark\.

### 5\.2Taxi Dispatching and Dynamic Pricing

We next evaluate the proposed CP\-DRL methods in the taxi dispatching application\. The New York City \(NYC\) calibrated dataset is used to evaluate real\-world applicability and scalability through a calibrated simulator\. In the first setup \(Section[5\.2\.2](https://arxiv.org/html/2609.36174#S5.SS2.SSS2)\), we focus on performance evaluation by comparing fairness objective values and operational metrics on a 15\-node case\. In the second setup \(Section[5\.2\.3](https://arxiv.org/html/2609.36174#S5.SS2.SSS3)\), we conduct mechanism analysis by inspecting the adjusted price distribution and a pricing trajectory\. In the final setup \(Section[5\.2\.4](https://arxiv.org/html/2609.36174#S5.SS2.SSS4)\), we validate the fairness execution on a smaller 4\-zone Midtown network by comparing disaggregation rules and different fairness measures\. Additional details can be found in[EC\.8](https://arxiv.org/html/2609.36174#A8)\.

#### 5\.2\.1Experimental Setup

The simulation environment is built on the NYC network and Yellow Taxi trip records provided by[New York City Taxi and Limousine Commission \(2026\)](https://arxiv.org/html/2609.36174#bib.bib6)\. The environment operates in discrete time steps of lengthΔ​t=3\\Delta t=3minutes, which balances responsiveness and computational tractability, similar to prior mobility system studies\([Sun et al\., 2022](https://arxiv.org/html/2609.36174#bib.bib38)\)\. The spatial domain is discretized into a graph where nodes represent taxi zones and edges represent connectivity based on road network distances\. All operational and economic parameters, including per\-unit ride reward, per\-unit travel cost, and price\-response coefficients, are calibrated to the same setting to ensure cross\-experiment comparability \(see[EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1)for full details of calibration on base arrival rates and other parameters\)\.

##### Pricing Mechanism

We consider four levels of pricing granularity\. OD pricing assigns a distinct multiplier to each OD pair per time step, which offers maximal theoretical flexibility but may complicate learning\. Origin\-only \(O\) uses one multiplier per\-origin, reducing the decision space while still accounting for spatial demand heterogeneity\. Uniform \(U\) pricing implements a single multiplier per time step, focusing mainly on temporal fluctuations\. Static \(S\) pricing maintains a constant price throughout the episode\. The optimal multiplier for this static baseline is determined by an exhaustive search and serves as the primary benchmark to quantify the performance gains of dynamic intervention strategies\. To reduce the price\-output dimension and regularize spatial price variation, low\-rank parameterizations are used, in which two factor matrices𝑳O,𝑳D∈\[0,1\]M×rp\{\\bm\{L\}\}^\{\\text\{O\}\}\\mathchar 24891\\allowbreak\{\\bm\{L\}\}^\{\\text\{D\}\}\\in\[0\\mathchar 24891\\allowbreak 1\]^\{M\\times r\_\{p\}\}are mapped to price multipliers as specified in[EC\.8\.3\.2](https://arxiv.org/html/2609.36174#A8.SS3.SSS2), andrpr\_\{p\}is the chosen pricing rank\.

##### Objectives and Metrics

All policies are trained withλ=1\\lambda=1\. We then evaluate the resulting policies from both welfare and operational perspectives\. Participant welfare is assessed using the GGF \(with the same weight configuration as in Section[5\.1](https://arxiv.org/html/2609.36174#S5.SS1)\), theα\\alpha\-fairness, and the utilitarian mean\. We also report metrics to characterize the mechanisms underlying policy performance\. Occupancy measures productive fleet utilization, admitted orders measure passenger service, and the average and time\-varying OD multipliers are used to interpret the learned pricing policy\.

##### CP\-DRL and Benchmarks

We use PPO and a standard Graph Neural Network \(GNN\) as the feature extractor to capture the underlying spatial dependencies \(see structural details in[EC\.8\.2](https://arxiv.org/html/2609.36174#A8.SS2)\), compared with the following benchmarks: 1\)count\-proportional assignment linear program\(CP\-ALP\) uses a myopic per\-step optimizer to replace the trained priority\-based sampling procedure to generate count actions\. 2\)Random\(RDM\) samples actions uniformly from the feasible parameter space, which isolates the contribution of structured decision\-making\. 3\) Alowest income first\(LIF\) heuristic gives assignment priority to drivers with lower accumulated income\. LIF uses the fixed price multiplier 5\.10 reported in Table[4](https://arxiv.org/html/2609.36174#S5.T4.fig1)and does not optimize price\. LIF is a local execution rule and is included to test the performance of income\-priority matching alone compared with the welfare achieved by coordinated pricing and dispatch\. 4\) We additionally report themultilayer perceptron\(MLP\) benchmark with fully connected layers to show how the spatial feature GNN improves performance\. For every fleet\-size and pricing\-granularity configuration, CP\-DRL and MLP are trained separately over 10 independent seeds\. CP\-ALP, RDM, and LIF are evaluated over 10 independent simulation seeds\. All methods are evaluated with a demand synchronization to CP\-DRL over the corresponding 10 seeds with 50 episodes per seed\. Complete benchmark definitions and implementation details are provided in[EC\.8\.4](https://arxiv.org/html/2609.36174#A8.SS4)\.

#### 5\.2\.2Performance on the 15\-Zone Configuration

We first evaluate CP\-DRL on a 15\-zone NYC\-calibrated network under OD pricing\. The fleet size varies over\{60,120,180,240,300\}\\\{60\\mathchar 24891\\allowbreak 120\\mathchar 24891\\allowbreak 180\\mathchar 24891\\allowbreak 240\\mathchar 24891\\allowbreak 300\\\}, which changes the degree of supply scarcity while keeping the calibrated base demand fixed\. The key managerial question is whether coordinated dynamic pricing and dispatch can simultaneously improve driver\-side welfare, passenger matching, and online decision speed as fleet size changes\.

##### Fairness\-Aware Performance

We first present the main performance on a 15\-zone network \(see details in[EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.SSS0.Px2)\) under OD pricing in Table[3](https://arxiv.org/html/2609.36174#S5.T3.fig1), withρ\\rhoset to GGF andλ=1\\lambda=1\. The number of taxis governs the scarcity of supply\. CP\-DRL obtains the highest mean value at every fleet size\. Relative to CP\-ALP, its mean objective is higher by 18\.9% at 60 taxis, 18\.9% at 120 taxis, 25\.7% at 180 taxis, 23\.5% at 240 taxis, and 19\.7% at 300 taxis\. Since CP\-ALP also learns pricing, this gap should not be attributed merely to the existence of a learned price policy\. Instead, it indicates that myopic per\-step dispatch optimization is insufficient when current assignment decisions shape future taxi availability, future matching opportunities, and the distribution of driver rewards\.

Table 3:Fairness\-Aware Objective Values \(OD Pricing\)LIF remains well below CP\-DRL despite explicitly favoring lower\-income drivers, showing that fairness\-motivated assignment at the execution layer cannot compensate for uncoordinated pricing and future fleet positioning\. RDM deteriorates rapidly as the fleet grows, confirming that feasibility alone does not generate effective coordination\. MLP is competitive for the smaller instances but becomes increasingly variable as fleet size grows\.

##### Operational Performance

Table[4](https://arxiv.org/html/2609.36174#S5.T4.fig1)links the welfare differences to service and fleet utilization\. CP\-DRL sustains substantially higher occupancy and admits many more orders than CP\-ALP, RDM, and LIF\. The gap of admitted orders relative to CP\-ALP grows with fleet size, where CP\-ALP admits 14\.35% fewer orders than CP\-DRL at 60 taxis, and the gap reaches 56\.51% at 300 taxis\. At the same time, CP\-DRL lowers its average multiplier as supply expands, from 3\.04 to 1\.38\. The combined pattern suggests that CP\-DRL does not obtain its welfare gain by suppressing demand through high prices, but uses additional fleet capacity more productively while relaxing prices as supply scarcity falls\.

Table 4:Operational Performance Under OD PricingNote\.Gap vs\. DRL is 100\(Method\-DRL\)/DRL%\. A negative value gap indicates a lower value than CP\-DRL, and vice versa\. For RDM, average price is the sample mean of randomized multipliers\. LIF uses the fixed multiplier 5\.10\.

##### Computational Performance

Table[5](https://arxiv.org/html/2609.36174#S5.T5.fig1)measures online count\-action generation during evaluation\. For CP\-DRL and MLP, the recorded time covers the the count action sampling procedure, whereas CP\-ALP solves its per\-period LP and RDM/LIF construct actions using their respective random or heuristic rules\. Under OD pricing, CP\-DRL’s mean action\-generation time increases linearly with fleet size, from approximately2\.92\.9milliseconds at 60 taxis to6\.76\.7milliseconds at 300 taxis, while CP\-ALP requires approximately 26 to 28 milliseconds\.

Table 5:Computational Time on Count\-Action GenerationNote\.Computation time is measured in milliseconds during policy evaluation\.

#### 5\.2\.3Pricing Mechanism Under Non\-Stationary Demand

The calibrated base arrival rates vary over time, and thus not directly comparable across decision steps\. We next report adjusted price multipliersψ−1​\(θtbase​\(i,j\)​ψ​\(pt​\(i,j\)\)\)\\psi^\{\-1\}\\,\\mathopen\{\(\}\\theta^\{\\mathrm\{base\}\}\_\{t\}\(i\\mathchar 24891\\allowbreak j\)\\psi\(p\_\{t\}\(i\\mathchar 24891\\allowbreak j\)\)\\mathclose\{\)\}according to the base demands to show how the major pricing decision changes with supply scarcity\.

##### Price Adjustment to Supply Scarcity

Figure[2](https://arxiv.org/html/2609.36174#S5.F2)reports the uniform \(U\) pricing policy\. The adjusted learned price distributions provide a demand\-equivalent comparison\. As fleet size increases from 60 to 300, the CP\-DRL’s mean price multipliers are lowered to admit additional demands and allow the platform to use the expanded fleet\. The optimized static multiplier changes across fleet sizes in the same direction as indicated by the dashed line\.

Figure 2:\(Color online\) Distribution of Adjusted Learned Uniform Price Multipliers![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/price_dist_cR.png)

Note\.The letter\-value \(boxen\) plots show the empirical distribution of adjusted price multipliers\. The center line denotes the median\. The nested boxes represent successively wider quantile intervals, from the interquartile range \(25th and 75th percentiles\) to the distribution tails\. The dashed reference line denotes the optimized static multiplier for the corresponding fleet size\.

Figure 3:\(Color online\) Adjusted Uniform Pricing and Dispatching on a Representative OD Pair![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/adjusted_odt_pair_rank1_161_164_63852206_20.png)

Note\.Lines report price multipliers, bars report matched and abandoned orders, and markers report empty relocations over time\.

##### Joint Evolution of Pricing and Matching

Further, in Figure[3](https://arxiv.org/html/2609.36174#S5.F3), we show a representative OD trajectory and how the aggregate patterns arise\. CP\-DRL maintains comparatively stable prices \(blue lines\) with more matched orders on the selected OD pair \(orange bars\) and less abandoned orders \(grey bars\)\. CP\-ALP changes prices more sharply and is accompanied by more abandoned demand, whereas RDM combines highly variable prices with few matches and more taxis relocated to other places while there are demands \(purple dots\)\. This indicates that CP\-DRL’s advantage comes from effectively coordinating pricing with dispatch over time\.

#### 5\.2\.4Fairness Comparison

The previous experiments evaluate count policies at the aggregate level\. However, a count policy specifies how many taxis take each action, not which labelled taxis receive those actions\. This distinction is central for fairness\. If a count action is always assigned to the same subset of taxi labels, individual rewards may become systematically unequal\.

In this section, we focus on small 4\-Node Midtown area \(see details in[EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.SSS0.Px2)\) with varying fleet sizeN∈\{10,20,30,…,60\}N\\in\\\{10\\mathchar 24891\\allowbreak 20\\mathchar 24891\\allowbreak 30\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak 60\\\}to compare the permutation\-invariant randomized disaggregation used by CP\-DRL with a biased \(B\) disaggregation model CP\-DRL\-B that repeatedly selects taxis according to the lowest fixed indices\.

As shown in Figure[4](https://arxiv.org/html/2609.36174#S5.F4), the two execution rules can produce similar mean\-objective values while generating a clear gap in the GGF value\. A permutation\-invariant disaggregation is therefore an operational requirement for transferring the fairness properties of the count policy to individual\-level execution\.

Figure 4:\(Color online\) Effect of Participant\-Level Disaggregation![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/fairness_a_R_bias.png)

Note\.Objectives are evaluated by the utilitarian mean and GGF under permutation\-invariant and biased execution rules across the OD/O/U pricing granularity\. The same learned count policy is executed using either the permutation\-invariant randomized disaggregation rule or the biased low\-index first rule, isolating participant\-level execution from aggregate count decisions\.

Table[6](https://arxiv.org/html/2609.36174#S5.T6.fig1)reports mean, GGF, andα\\alpha\-fairness evaluations withα=1\\alpha=1\. The gap between mean and GGF remains small under permutation\-invariant disaggregation, especially relative to the biased execution experiment\. The table also shows that dynamic OD pricing generally outperforms origin\-only and uniform pricing under CP\-DRL, whereas static pricing can remain competitive in smaller cases because it avoids learning a high\-dimensional policy space while demand variation is limited\. This indicates that OD pricing has a better performance when the platform has sufficient data and computational capacity to learn spatial\-temporal controls\.

Table 6:Evaluation Under Alternative Fairness MeasuresNote\.Entries are reported as mean±\\pmstandard deviation over 10 experiments\. Theα\\alpha\-fairness column usesα=1\\alpha=1and is computed as the geometric meanρ1​\[𝒗\]=\(∏n=1Nvn\)1/N\\rho\_\{1\}\[\{\\bm\{v\}\}\]=\(\\prod\_\{n=1\}^\{N\}v\_\{n\}\)^\{1/N\}\.

## 6Conclusion

We studied fairness\-aware sequential resource allocation in major\-minor weakly coupled Markov decision processes\. The proposed M2WCMDP framework models a centrally coordinated system with one major sub\-MDP representing endogenous information and platform\-level decisions, and a collection of minor sub\-MDPs representing homogeneous participants with joint actions coupled through shared resource constraints\. The objective balances the platform’s expected total discounted reward with a welfare functionρ\\rhodefined over the vector of participant expected total discounted rewards\.

The main theoretical result establishes a utilitarian reduction for the symmetricρ\\rho\-M2WCMDP optimization problem\. We showed that, under symmetry conditions on the minor sub\-MDPs, this fairness\-aware problem can be solved by optimizing the utilitarian\-based objective over permutation\-invariant policies\. We further use this structure to obtain a count\-aggregation reformulation, which removes dependence on participant labels and provides a compact representation for large populations\. To address large\-scale and model\-free settings, we developed a count\-proportion deep reinforcement learning approach\. The proposed CP\-DRL method learns joint count actions using count\-proportion state representations and stochastic policy networks\. Numerical results on the machine replacement problem validate empirically the symmetry\-based reduction in a controlled setting, while experiments on taxi dispatching and pricing show that coordinated pricing, relocation, and matching can improve fairness\-aware welfare relative to various benchmarks\.

Several extensions remain open\. First, action generation remains a computational bottleneck as the current priority sampler is sequential and its representational flexibility depends on the chosen action\-category partition\. A natural next step is to design more efficient feasibility\-preserving samplers and to characterize the approximation loss induced by coarse categories\. Second, the exact utilitarian reduction relies on symmetry of homogeneous minor sub\-MDPs\. An important extension is to heterogeneous populations, for example through finitely many participant types\. This would replace full symmetry by within\-type symmetry, lead to a type\-by\-state count representation, and raise a new fairness question of how welfare should be balanced both within and across heterogeneous groups\.

###### Notes

1. [1](https://arxiv.org/html/2609.36174#endnote1)1 1 1 endnote 1 Here, finiteness of state and action spaces can be obtained by assuming bounded OD queues and discretizing the support of the price multipliers\. In our implementation, we let price multipliers to be continuous with the assumption that our theory from sections and still holds in continuous spaces\.
2. [2](https://arxiv.org/html/2609.36174#endnote2)2 2 2 endnote 2 If all eligible priority scores are zero, a uniform value of / 1 \( × \| S \| \| A \| \) is used\.

## Acknowledgements

Xiaohui Tu was partially funded by GERAD and IVADO\. Yossiri Adulyasak was partially supported by the Canadian Natural Sciences and Engineering Research Council \[Grant RGPIN\-2021\-03264\] and by the Canada Research Chair program \[CRC\-2022\-00087\]\. Erick Delage was partially supported by the Canadian Natural Sciences and Engineering Research Council \[Grant RGPIN\-2022\-05261\] and by the Canada Research Chair program \[950\-230057\]\.

## References

- D\. Adelman and A\. J\. MersereauRelaxations of weakly coupled stochastic dynamic programs\.Operations Research56\(3\),pp\. 712–727\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.36174#S1.p1.1),[§1](https://arxiv.org/html/2609.36174#S1.p4.1)\.
- Adelman \(2004\)D\. AdelmanA price\-directed approach to stochastic inventory/routing\.Operations Research52\(4\),pp\. 499–514\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p1.1)\.
- Akbarzadeh and Mahajan \(2019\)N\. Akbarzadeh and A\. MahajanRestless bandits with controlled restarts: indexability and computation of whittle index\.In2019 IEEE 58th conference on decision and control \(CDC\),pp\. 7294–7300\.Cited by:[§EC\.7\.1](https://arxiv.org/html/2609.36174#A7.SS1.p1.1),[Example 3](https://arxiv.org/html/2609.36174#Thmexample3.p1.1)\.
- Aminianet al\.\(2026\)M\. R\. Aminian, V\. Manshadi, and R\. NiazadehMarkovian search with ex ante constraints: theory and applications to socially aware algorithmic hiring\.Management Science\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1)\.
- Andersson and Djehiche \(2011\)D\. Andersson and B\. DjehicheA maximum principle for sdes of mean\-field type\.Applied Mathematics & Optimization63,pp\. 341–356\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px2.p1.1)\.
- Azizet al\.\(2024\)H\. Aziz, R\. Freeman, N\. Shah, and R\. VaishBest of both worlds: ex ante and ex post fairness in resource allocation\.Operations Research72\(4\),pp\. 1674–1688\.Cited by:[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Bäuerle \(2023\)N\. BäuerleMean field markov decision processes\.Applied Mathematics & Optimization88\(1\),pp\. 12\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px2.p2.1)\.
- Bensoussanet al\.\(2013\)A\. Bensoussan, J\. Frehse, P\. Yam,et al\.Mean field games and mean field type control theory\.Vol\.101,Springer\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px2.p1.1)\.
- Bertsimaset al\.\(2012\)D\. Bertsimas, V\. F\. Farias, and N\. TrichakisOn the efficiency\-fairness trade\-off\.Management Science58\(12\),pp\. 2234–2250\.Cited by:[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p5.2)\.
- Bertsimaset al\.\(2013\)D\. Bertsimas, V\. F\. Farias, and N\. TrichakisFairness, efficiency, and flexibility in organ allocation for kidney transplantation\.Operations research61\(1\),pp\. 73–87\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Bertsimas and Mersereau \(2007\)D\. Bertsimas and A\. J\. MersereauA learning approach for interactive marketing to a customer segment\.Operations Research55\(6\),pp\. 1120–1135\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p1.1)\.
- Bhattacharyaet al\.\(2024\)R\. Bhattacharya, T\. Nguyen, W\. W\. Sun, and M\. TawarmalaniActive learning for fair and stable online allocations\.InProceedings of the 25th ACM Conference on Economics and Computation,pp\. 196–197\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p3.1),[§1](https://arxiv.org/html/2609.36174#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Bistritzet al\.\(2020\)I\. Bistritz, T\. Baharav, A\. Leshem, and N\. BambosMy fair bandit: distributed learning of max\-min fairness with multi\-player bandits\.InInternational Conference on Machine Learning,pp\. 930–940\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p1.1)\.
- Brown and Zhang \(2023\)D\. B\. Brown and J\. ZhangOn the strength of relaxations of weakly coupled stochastic dynamic programs\.Operations Research71\(6\),pp\. 2374–2389\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p2.1)\.
- Brown and Zhang \(2025\)D\. B\. Brown and J\. ZhangFluid policies, reoptimization, and performance guarantees in dynamic resource allocation\.Operations Research73\(2\),pp\. 1029–1045\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p4.1)\.
- Caiet al\.\(2021\)D\. Cai, S\. H\. Lim, and L\. WynterEfficient reinforcement learning in resource allocation problems through permutation invariant multi\-task learning\.In2021 60th IEEE Conference on Decision and Control \(CDC\),pp\. 2270–2275\.Cited by:[§3\.1](https://arxiv.org/html/2609.36174#S3.SS1.p2.1)\.
- Caro and Gallien \(2007\)F\. Caro and J\. GallienDynamic assortment with demand learning for seasonal consumer goods\.Management science53\(2\),pp\. 276–292\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p1.1)\.
- Chenet al\.\(2024\)Q\. Chen, Y\. Lei, and S\. JasinReal\-time spatial–intertemporal pricing and relocation in a ride\-hailing network: near\-optimal policies and the value of dynamic pricing\.Operations Research72\(5\),pp\. 2097–2118\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1),[Example 4](https://arxiv.org/html/2609.36174#Thmexample4.p2.1)\.
- Chen and Hooker \(2023\)V\. X\. Chen and J\. N\. HookerA guide to formulating fairness in an optimization model\.Annals of Operations Research326\(1\),pp\. 581–619\.Cited by:[Appendix EC\.2](https://arxiv.org/html/2609.36174#A2.p2.1),[§1](https://arxiv.org/html/2609.36174#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p3.1),[Example EC\.1](https://arxiv.org/html/2609.36174#Thmexample1a.p1.1)\.
- Chenet al\.\(2025\)Y\. Chen, J\. Dong, Z\. Wang, and C\. ZhangA primal\-dual approach to constrained markov decision processes with applications to queue scheduling and inventory management\.Management Science\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p2.1)\.
- Chenet al\.\(2020\)Y\. Chen, A\. Cuellar, H\. Luo, J\. Modi, H\. Nemlekar, and S\. NikolaidisFair contextual multi\-armed bandits: theory and experiments\.InConference on Uncertainty in Artificial Intelligence,pp\. 181–190\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p3.1)\.
- Cohenet al\.\(2016\)P\. Cohen, R\. Hahn, J\. Hall, S\. Levitt, and R\. MetcalfeUsing big data to estimate consumer surplus: the case of uber\.Technical reportNational Bureau of Economic Research\.Cited by:[§EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.p2.1)\.
- Cousins \(2021\)C\. CousinsAn axiomatic theory of provably\-fair welfare\-centric machine learning\.Advances in Neural Information Processing Systems34,pp\. 16610–16621\.Cited by:[Example EC\.1](https://arxiv.org/html/2609.36174#Thmexample1a.p1.1)\.
- Cuiet al\.\(2024a\)K\. Cui, G\. Dayanıklı, M\. Laurière, M\. Geist, O\. Pietquin, and H\. KoepplLearning discrete\-time major\-minor mean field games\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 9616–9625\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p3.1)\.
- Cuiet al\.\(2024b\)K\. Cui, C\. Fabian, A\. Tahir, and H\. KoepplMajor\-minor mean field multi\-agent reinforcement learning\.International Conference on Machine Learning\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p3.1)\.
- Daiet al\.\(2017\)G\. Dai, J\. Huang, S\. M\. Wambura, and H\. SunA balanced assignment mechanism for online taxi recommendation\.In2017 18th IEEE international conference on mobile data management \(MDM\),pp\. 102–111\.Cited by:[§2\.3](https://arxiv.org/html/2609.36174#S2.SS3.p2.1)\.
- Delage and Mannor \(2010\)E\. Delage and S\. MannorPercentile Optimization for Markov Decision Processes with Parameter Uncertainty\.Operations Research58\(1\),pp\. 203–213\.External Links:[Document](https://dx.doi.org/10.1287/opre.1080.0685)Cited by:[Example 3](https://arxiv.org/html/2609.36174#Thmexample3.p1.1)\.
- El Shar and Jiang \(2023\)I\. El Shar and D\. JiangWeakly coupled deep q\-networks\.Advances in Neural Information Processing Systems36,pp\. 43931–43950\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p3.1)\.
- Fanet al\.\(2023\)Z\. Fan, N\. Peng, M\. Tian, and B\. FainWelfare and fairness in multi\-objective reinforcement learning\.InProceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems,pp\. 1991–1999\.Cited by:[Table EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1.5.1.8.1.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p1.1),[Example 2](https://arxiv.org/html/2609.36174#Thmexample2.p1.1)\.
- Filippiet al\.\(2021\)C\. Filippi, G\. Guastaroba, and M\. G\. SperanzaOn single\-source capacitated facility location with cost and fairness objectives\.European Journal of Operational Research289\(3\),pp\. 959–974\.Cited by:[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Gallego and Van Ryzin \(1994\)G\. Gallego and G\. Van RyzinOptimal dynamic pricing of inventories with stochastic demand over finite horizons\.Management science40\(8\),pp\. 999–1020\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p1.1)\.
- Gastet al\.\(2022\)N\. Gast, B\. Gaujal, and C\. YanReoptimization nearly solves weakly coupled markov decision processes\.arXiv preprint arXiv:2211\.01961\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p2.1),[§1](https://arxiv.org/html/2609.36174#S1.p6.1),[§3\.2](https://arxiv.org/html/2609.36174#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.36174#S3.SS2.p5.1)\.
- Ghalmeet al\.\(2021\)G\. Ghalme, V\. Nair, V\. Patil, and Y\. ZhouLong\-term resource allocation fairness in average markov decision process \(amdp\) environment\.arXiv preprint arXiv:2102\.07120\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p2.1)\.
- Gutjahr and Fischer \(2018\)W\. J\. Gutjahr and S\. FischerEquity and deprivation costs in humanitarian logistics\.European Journal of Operational Research270\(1\),pp\. 185–197\.Cited by:[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Hassanzadehet al\.\(2023\)P\. Hassanzadeh, E\. Kreacic, S\. Zeng, Y\. Xiao, and S\. GaneshSequential fair resource allocation under a markov decision process framework\.InProceedings of the Fourth ACM International Conference on AI in Finance,pp\. 673–680\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p3.1),[§1](https://arxiv.org/html/2609.36174#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.
- Hawkins \(2003\)J\. T\. HawkinsA langrangian decomposition approach to weakly coupled dynamic optimization problems and its applications\.Ph\.D\. Thesis,Massachusetts Institute of Technology\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.36174#S1.p1.1)\.
- Hooker and Williams \(2012\)J\. N\. Hooker and H\. P\. WilliamsCombining equity and utilitarianism in a mathematical programming model\.Management Science58\(9\),pp\. 1682–1693\.Cited by:[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p5.2)\.
- Huaizhouet al\.\(2013\)S\. Huaizhou, R\. V\. Prasad, E\. Onur, and I\. NiemegeersFairness in wireless networks: issues, measures and challenges\.IEEE Communications Surveys & Tutorials16\(1\),pp\. 5–24\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1)\.
- Huanget al\.\(2024\)J\. Huang, B\. Yardim, and N\. HeOn the statistical efficiency of mean\-field reinforcement learning with general function approximation\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 289–297\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p1.1)\.
- Huanget al\.\(2018\)Y\. Huang, D\. Sun, and J\. TangTaxi driver speeding: who, when, where and how? a comparative study between shanghai and new york city\.Traffic injury prevention19\(3\),pp\. 311–316\.Cited by:[§EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.p1.1)\.
- Juet al\.\(2023\)P\. Ju, A\. Ghosh, and N\. B\. ShroffAchieving fairness in multi\-agent markov decision processes using reinforcement learning\.arXiv preprint arXiv:2306\.00324\.Cited by:[Example 2](https://arxiv.org/html/2609.36174#Thmexample2.p1.1)\.
- Jusupet al\.\(2024\)M\. Jusup, B\. Pásztor, T\. Janik, K\. Zhang, F\. Corman, A\. Krause, and I\. BogunovicSafe model\-based multi\-agent mean\-field reinforcement learning\.International conference on autonomous agents and multiagent systems\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p2.1)\.
- Kazemiet al\.\(2019\)S\. M\. Kazemi, R\. Goel, S\. Eghbali, J\. Ramanan, J\. Sahota, S\. Thakur, S\. Wu, C\. Smyth, P\. Poupart, and M\. BrubakerTime2vec: learning a vector representation of time\.arXiv preprint arXiv:1907\.05321\.Cited by:[§EC\.8\.2](https://arxiv.org/html/2609.36174#A8.SS2.p5.2)\.
- Kellyet al\.\(1998\)F\. P\. Kelly, A\. K\. Maulloo, and D\. K\. H\. TanRate control for communication networks: shadow prices, proportional fairness and stability\.Journal of the Operational Research society49\(3\),pp\. 237–252\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p1.1)\.
- Kolm \(1976\)S\. KolmUnequal inequalities\. i\.Journal of economic Theory12\(3\),pp\. 416–442\.Cited by:[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p3.1)\.
- Lehky \(2010\)S\. R\. LehkyDecoding poisson spike trains by gaussian filtering\.Neural computation22\(5\),pp\. 1245–1271\.Cited by:[§EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.SSS0.Px1.p1.1)\.
- Lesmanaet al\.\(2019\)N\. S\. Lesmana, X\. Zhang, and X\. BeiBalancing efficiency and fairness in on\-demand ridesourcing\.Advances in neural information processing systems32\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1)\.
- Li and Varakantham \(2022\)D\. Li and P\. VarakanthamEfficient resource allocation with fairness constraints in restless multi\-armed bandits\.InUncertainty in Artificial Intelligence,pp\. 1158–1167\.Cited by:[§5\.1\.1](https://arxiv.org/html/2609.36174#S5.SS1.SSS1.Px2.p1.1)\.
- Liet al\.\(2019\)M\. Li, Z\. Qin, Y\. Jiao, Y\. Yang, J\. Wang, C\. Wang, G\. Wu, and J\. YeEfficient ridesharing order dispatching with mean field multi\-agent reinforcement learning\.InThe world wide web conference,pp\. 983–994\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p2.1)\.
- Mandal and Gan \(2022\)D\. Mandal and J\. GanSocially fair reinforcement learning\.arXiv preprint arXiv:2208\.12584\.Cited by:[Example 2](https://arxiv.org/html/2609.36174#Thmexample2.p1.1)\.
- Meherrem and Hafayed \(2024\)S\. Meherrem and M\. HafayedA stochastic maximum principle for general mean\-field system with constraints\.Numerical Algebra, Control and Optimization,pp\. 0–0\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px2.p2.1)\.
- Meuleauet al\.\(1998\)N\. Meuleau, M\. Hauskrecht, K\. Kim, L\. Peshkin, L\. P\. Kaelbling, T\. L\. Dean, and C\. BoutilierSolving very large weakly coupled markov decision processes\.AAAI/IAAI8,pp\. 2\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p1.1)\.
- Michailidiset al\.\(2023\)D\. Michailidis, S\. Ghebreab, and F\. P\. SantosBalancing fairness and efficiency in transport network design through reinforcement learning\.InProceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems,pp\. 2532–2534\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1)\.
- Mo and Walrand \(2000\)J\. Mo and J\. WalrandFair end\-to\-end window\-based congestion control\.IEEE/ACM Transactions on networking8\(5\),pp\. 556–567\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p1.1),[Table EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1.5.1.7.1.1),[Example 2](https://arxiv.org/html/2609.36174#Thmexample2.p1.1)\.
- Moulin \(1991\)H\. MoulinAxioms of cooperative decision making\.1,Cambridge university press\.Cited by:[Appendix EC\.2](https://arxiv.org/html/2609.36174#A2.p2.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p3.1)\.
- Moyanoet al\.\(2021\)A\. Moyano, M\. Stępniak, B\. Moya\-Gómez, and J\. C\. García\-PalomaresTraffic congestion and economic context: changes of spatiotemporal patterns of traffic travel times during crisis and post\-crisis periods\.Transportation48\(6\),pp\. 3301–3324\.Cited by:[§EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.SSS0.Px1.p1.1)\.
- Nadarajah and Cire \(2025\)S\. Nadarajah and A\. A\. CireSelf\-adapting network relaxations for weakly coupled markov decision processes\.Management Science71\(2\),pp\. 1779–1802\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p3.1),[§1](https://arxiv.org/html/2609.36174#S1.p1.1),[§1](https://arxiv.org/html/2609.36174#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.36174#S2.SS3.p2.1)\.
- New York City Taxi and Limousine Commission \(2026\)New York City Taxi and Limousine CommissionTLC trip record data\.New York City Taxi and Limousine Commission\.External Links:[Link](https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page)Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p8.1),[§5\.2\.1](https://arxiv.org/html/2609.36174#S5.SS2.SSS1.p1.1)\.
- Parrott \(2024\)J\. A\. ParrottRevised expense model for the nyc taxi and limousine commission’s high\-volume for\-hire vehicle minimum pay standard\.Technical reportCenter for New York City Affairs, The New School\.Note:Report prepared for the New York City Taxi and Limousine CommissionExternal Links:[Link](https://www.nyc.gov/assets/tlc/downloads/pdf/driver_expense_report.pdf)Cited by:[§EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.p2.1)\.
- Puterman \(2005\)M\. L\. PutermanMarkov decision processes: discrete stochastic dynamic programming\.John Wiley & Sons\.Cited by:[Appendix EC\.3](https://arxiv.org/html/2609.36174#A3.p1.1),[Appendix EC\.6](https://arxiv.org/html/2609.36174#A6.p1.1),[§3\.1](https://arxiv.org/html/2609.36174#S3.SS1.p4.1),[Lemma EC\.2](https://arxiv.org/html/2609.36174#Thmlemma2a.p1.1.1)\.
- Rawls \(1971\)J\. RawlsAn egalitarian theory of justice\.Philosophical ethics: An introduction to moral philosophy,pp\. 365–370\.Cited by:[Table EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1.5.1.5.1.1),[§1](https://arxiv.org/html/2609.36174#S1.p5.1),[Example 1](https://arxiv.org/html/2609.36174#Thmexample1.p1.1),[Example 2](https://arxiv.org/html/2609.36174#Thmexample2.p1.1)\.
- Robledoet al\.\(2024\)F\. Robledo, U\. Ayesta, and K\. AvrachenkovDeep reinforcement learning for weakly coupled mdp’s with continuous actions\.InInternational Conference on Analytical and Stochastic Modeling Techniques and Applications,pp\. 67–80\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px1.p3.1)\.
- Schulmanet al\.\(2017\)J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. KlimovProximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§5](https://arxiv.org/html/2609.36174#S5.p2.1)\.
- Sen \(1976\)A\. SenReal national income\.The Review of Economic Studies43\(1\),pp\. 19–39\.Cited by:[Table EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1.5.1.9.1.1),[Example EC\.1](https://arxiv.org/html/2609.36174#Thmexample1a.p1.1)\.
- Siddiqueet al\.\(2020\)U\. Siddique, P\. Weng, and M\. ZimmerLearning fair policies in multi\-objective \(deep\) reinforcement learning with average and discounted rewards\.InInternational Conference on Machine Learning,pp\. 8905–8915\.Cited by:[Example 1](https://arxiv.org/html/2609.36174#Thmexample1.p1.1)\.
- Sinhaet al\.\(2023\)A\. Sinha, A\. Joshi, R\. Bhattacharjee, C\. Musco, and M\. HajiesmailiNo\-regret algorithms for fair resource allocation\.Advances in Neural Information Processing Systems36,pp\. 48083–48109\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p3.1)\.
- Speicheret al\.\(2018\)T\. Speicher, H\. Heidari, N\. Grgic\-Hlaca, K\. P\. Gummadi, A\. Singla, A\. Weller, and M\. B\. ZafarA unified approach to quantifying algorithmic unfairness: measuring individual &group unfairness via inequality indices\.InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining,pp\. 2239–2248\.Cited by:[Example EC\.1](https://arxiv.org/html/2609.36174#Thmexample1a.p1.1)\.
- Sunet al\.\(2022\)J\. Sun, H\. Jin, Z\. Yang, L\. Su, and X\. WangOptimizing long\-term efficiency and fairness in ride\-hailing via joint order dispatching and driver repositioning\.InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining,pp\. 3950–3960\.Cited by:[§1](https://arxiv.org/html/2609.36174#S1.p3.1),[§5\.2\.1](https://arxiv.org/html/2609.36174#S5.SS2.SSS1.p1.1)\.
- Tsang and Shehadeh \(2025\)M\. Y\. Tsang and K\. S\. ShehadehA unified framework for analyzing and optimizing a class of convex fairness measures\.Operations Research\.Cited by:[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p3.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p5.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p5.2)\.
- Wenet al\.\(2021\)M\. Wen, O\. Bastani, and U\. TopcuAlgorithms for fairness in sequential decision making\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 1144–1152\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px4.p2.1)\.
- Weymark \(1981\)J\. A\. WeymarkGeneralized gini inequality indices\.Mathematical social sciences1\(4\),pp\. 409–430\.Cited by:[Table EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1.5.1.4.1.1),[§1](https://arxiv.org/html/2609.36174#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p1.1),[Example 1](https://arxiv.org/html/2609.36174#Thmexample1.p1.1),[Example EC\.1](https://arxiv.org/html/2609.36174#Thmexample1a.p1.1)\.
- Whittle \(1988\)P\. WhittleRestless bandits: activity allocation in a changing world\.Journal of applied probability25\(A\),pp\. 287–298\.Cited by:[§5\.1](https://arxiv.org/html/2609.36174#S5.SS1.p1.1)\.
- Williamson and Menon \(2019\)R\. Williamson and A\. MenonFairness risk measures\.InInternational conference on machine learning,pp\. 6786–6797\.Cited by:[§2\.2](https://arxiv.org/html/2609.36174#S2.SS2.p3.1)\.
- Yanget al\.\(2018\)Y\. Yang, R\. Luo, M\. Li, M\. Zhou, W\. Zhang, and J\. WangMean field multi\-agent reinforcement learning\.InInternational conference on machine learning,pp\. 5571–5580\.Cited by:[Appendix EC\.1](https://arxiv.org/html/2609.36174#A1.SS0.SSS0.Px3.p1.1)\.
- Zhang \(2022\)X\. ZhangNear\-optimality for multi\-action multi\-resource restless bandits with many arms\.Cornell University\.Cited by:[§5\.1](https://arxiv.org/html/2609.36174#S5.SS1.p1.1)\.
- Zhuet al\.\(2021\)Z\. Zhu, J\. Ke, and H\. WangA mean\-field markov decision process model for spatial\-temporal subsidies in ride\-sourcing markets\.Transportation Research Part B: Methodological150,pp\. 540–565\.Cited by:[§2\.1](https://arxiv.org/html/2609.36174#S2.SS1.p10.1)\.

## Electronic Companions of Paper “Fair Policy Optimization in Major\-Minor Weakly Coupled Markov Decision Processes”

## Appendix EC\.1Literature Review

Motivated to address the large\-scale coordination challenges and endogenous system\-level information inherent in taxi dispatching systems, our work extends the standard weakly coupled Markov decision processes \(WCMDP\) framework to a major\-minor formulation\. This formulation is closely connected to the literature on traditional weakly coupled models and modern mean\-field approaches, while explicitly integrating fairness notions\.

##### Weakly Coupled Markov Decision Processes

The foundational work by[Meuleau et al\. \(1998\)](https://arxiv.org/html/2609.36174#bib.bib72)recognized the potential for decomposition in resource\-constrained systems, but the proposed heuristics were task\-specific and without optimality guarantees\. Subsequently,[Hawkins \(2003\)](https://arxiv.org/html/2609.36174#bib.bib73)introduced a Lagrangian relaxation approach for constrained stochastic dynamic programming problems to allow tractable approximations\.[Adelman and Mersereau \(2008\)](https://arxiv.org/html/2609.36174#bib.bib74)extended these ideas through performance bounds and policy design using linear programming \(LP\) relaxations, thus providing theoretical justification for decomposition\-based methods\.

Recent advances continue to build on relaxation and decomposition approaches\. For example,[Gast et al\. \(2022\)](https://arxiv.org/html/2609.36174#bib.bib43)proposed LP\-based policies with sublinear regret and bounded duality gaps as the size of sub\-MDPs grows, while their practicability is based on repeatedly solving relaxed LPs in real time\.[Brown and Zhang \(2023\)](https://arxiv.org/html/2609.36174#bib.bib70)investigated the tightness of various relaxations for weakly coupled stochastic dynamic programs, quantifying the suboptimality introduced by standard linear relaxations and providing conditions under which tighter formulations significantly improve solution quality\.[Chen et al\. \(2025\)](https://arxiv.org/html/2609.36174#bib.bib67)proposed a unified primal\-dual framework that integrates policy gradient methods with dual variable updates to address resource\-constrained planning problems more robustly, offering convergence guarantees under mild assumptions\.

In parallel, several methods have attempted to address the limitations of classical decomposition techniques by introducing learning\-based or more adaptive architectures\. For example,[El Shar and Jiang \(2023\)](https://arxiv.org/html/2609.36174#bib.bib71)introduced a weakly coupled variant of deep Q\-learning that partitions the value function estimation across sub\-tasks and uses dual bounds to coordinate global feasibility\. However, the choice and coverage of multiplier sets are unclear and potentially computationally expensive\. Similarly,[Robledo et al\. \(2024\)](https://arxiv.org/html/2609.36174#bib.bib69)developed a deep reinforcement learning \(RL\) approach tailored to WCMDPs with continuous action spaces, which showed how neural function approximators can be integrated with soft constraint penalties to achieve efficient learning in high\-dimensional problems\.[Nadarajah and Cire \(2025\)](https://arxiv.org/html/2609.36174#bib.bib68)proposed self\-adapting network relaxations that dynamically adjust relaxation tightness to the structure of linking constraints\. These studies provide a promising direction for learning\-based and structure\-aware optimization in weakly coupled systems\.

##### Mean\-Field Control \(MFC\)

Mean\-field approximations provide another powerful tool for modeling and solving large\-population stochastic control problems\. Initiated by[Andersson and Djehiche \(2011\)](https://arxiv.org/html/2609.36174#bib.bib62)and formalized in[Bensoussan et al\. \(2013\)](https://arxiv.org/html/2609.36174#bib.bib61), MFC framed control problems by enabling optimization directly on population\-level distributions instead of individual interactions\.

Recent work has extended MFC to more general and realistic settings\. For example,[Bäuerle \(2023\)](https://arxiv.org/html/2609.36174#bib.bib66)showed that, under weak coupling and ergodicity assumptions, the optimal control converges to the solution of a static optimization over stationary distributions as the population size grows to infinity\.[Meherrem and Hafayed \(2024\)](https://arxiv.org/html/2609.36174#bib.bib58)established stochastic maximum principles for general mean\-field systems with constraints, where coefficients depend nonlinearly on both the state process and its probability law\. However, classical MFC typically assumes full knowledge of system dynamics and often relies on structural conditions such as regularity and convexity, limiting practical applicability\. In contrast, our symmetry reduction is formulated for a finite population, and the count representation removes participant labels without taking an infinite\-population limit\.

##### Mean\-Field Reinforcement Learning \(MFRL\)

MFRL directly extends MFC by integrating data\-driven learning into large\-population control, allowing each agent to interact with mean\-field dynamics\.[Yang et al\. \(2018\)](https://arxiv.org/html/2609.36174#bib.bib59)pioneered MFRL by developing mean\-field Q\-learning and actor\-critic, and proved its convergence to a Nash equilibrium\. Their work showed empirical success in resource allocation and coordination benchmarks, establishing MFRL as a scalable alternative to traditional multi\-agent RL frameworks\.[Huang et al\. \(2024\)](https://arxiv.org/html/2609.36174#bib.bib51)refined the theory of MFRL, analyzing sample efficiency and learning complexity under general function approximation, and identifying limitations of previous convergence results\.

This foundational result was applied in various transportation domains\. For example,[Li et al\. \(2019\)](https://arxiv.org/html/2609.36174#bib.bib60)applied this framework to large\-scale ride\-sharing dispatch, showing that mean\-field interactions in a multi\-agent RL setting could effectively scale\.[Jusup et al\. \(2024\)](https://arxiv.org/html/2609.36174#bib.bib63)further extended this approach by incorporating log\-barrier constraints into mean\-field RL to ensure region\-level safety and coverage, and validated the proposed methods in vehicle repositioning applications\.

Recent MFRL extensions introduce a major\-minor structure, where a dominant agent influences a population of minor agents through mean\-field interaction\.[Cui et al\. \(2024a\)](https://arxiv.org/html/2609.36174#bib.bib65);[Cui et al\. \(2024b\)](https://arxiv.org/html/2609.36174#bib.bib64)established finite\-agent approximation results, proved that stationary policies suffice in the major\-minor mean\-field control setting, and introduced a policy gradient algorithm that effectively solves such structured multi\-agent systems\. This line of research is highly relevant to our work but differs in incorporating fairness objectives into a fully centralized major\-minor setting, where minor actions are constrained by shared resources\.

##### Fairness in Resource Allocation

Fairness concerns arise when allocating limited resources among multiple stakeholders\. Classical notions of fairness in resource allocation have been grounded in network economics and game\-theoretic models, including max\-min fairness\([Bistritz et al\., 2020](https://arxiv.org/html/2609.36174#bib.bib57)\), proportional fairness\([Kelly et al\., 1998](https://arxiv.org/html/2609.36174#bib.bib56)\), andα\\alpha\-fairness\([Mo and Walrand, 2000](https://arxiv.org/html/2609.36174#bib.bib37)\)\.

More recently, fairness has been extended to the sequential decision\-making setting\.[Wen et al\. \(2021\)](https://arxiv.org/html/2609.36174#bib.bib55)and[Ghalme et al\. \(2021\)](https://arxiv.org/html/2609.36174#bib.bib54)explored fairness\-aware formulations of MDPs, introducing objectives that account for equity across states or populations over time\. These efforts have broadened the scope of fairness from static allocation to dynamic and learning\-based environments\.

In parallel, the online learning and bandits community has developed fairness\-aware algorithms to make resource allocation decisions with performance guarantees\. For example,[Chen et al\. \(2020\)](https://arxiv.org/html/2609.36174#bib.bib50)introduced a fairness\-aware approach to contextual bandits that enforces minimum selection rates across users while achieving provable regret\.[Sinha et al\. \(2023\)](https://arxiv.org/html/2609.36174#bib.bib53)developed an online proportional fairness algorithm with regret bounds, while[Hassanzadeh et al\. \(2023\)](https://arxiv.org/html/2609.36174#bib.bib52)introduced policies that optimize the Nash Social Welfare in sequential resource allocation problems with bounded optimality gaps\.[Bhattacharya et al\. \(2024\)](https://arxiv.org/html/2609.36174#bib.bib49)introduced actively sampled fair allocation algorithms that solicit feedback from a subset of agents per period, while achieving logarithmic regret in both fairness and matching stability\. These works show a shift toward algorithmic frameworks that are not only efficient but also equitable\.

## Appendix EC\.2Boundary of the welfare class

To clarify the scope of the welfare classρ\\rhoin Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1), we categorize commonly used welfare functions into two classes: \(i\)directly compliantmeasures that natively satisfy all four properties in Definition \([1](https://arxiv.org/html/2609.36174#Thmdefinition1)\), such as the generalized Gini function \(GGF\) in Example[1](https://arxiv.org/html/2609.36174#Thmexample1), \(ii\)transformablemeasures that satisfy the definition after an appropriate normalization or reformulation, such as the certainty\-equivalent representation ofα\\alpha\-fairness in Example[2](https://arxiv.org/html/2609.36174#Thmexample2), and the widely used mean\-minus penalty introduced next\. Table[EC\.1](https://arxiv.org/html/2609.36174#A2.T1.fig1)summarizes the two categories\.

###### Example EC\.1\(Mean\-Minus Penalty\)\.

Another highly structural and widely applicable class of welfare measure can be constructed by subtracting a convex inequality penaltyφ⁡\(𝒗\)\\varphi\(\{\\bm\{v\}\}\)from the utilitarian meanv¯=1N​∑i=1Nvi\\bar\{v\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}v\_\{i\}, formulated asρ⁡\[𝒗\]=v¯−c⋅φ⁡\(𝒗\)\\rho\[\{\\bm\{v\}\}\]=\\bar\{v\}\-c\\cdot\\varphi\(\{\\bm\{v\}\}\)\. Ifφ⁡\(𝒗\)\\varphi\(\{\\bm\{v\}\}\)is convex, symmetric, and vanishes under perfect equality such thatρ⁡\[c​𝟏\]=c\\rho\[c\{\\bm\{1\}\}\]=c, the resultingρ⁡\[𝒗\]\\rho\[\{\\bm\{v\}\}\]directly satisfies Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)\. This unified form broadly includes Gini dispersion\([Weymark, 1981](https://arxiv.org/html/2609.36174#bib.bib47)\), Gini coefficient\([Sen, 1976](https://arxiv.org/html/2609.36174#bib.bib21)\), mean absolute deviation\([Chen and Hooker, 2023](https://arxiv.org/html/2609.36174#bib.bib26)\), and other inequality indices\([Speicher et al\., 2018](https://arxiv.org/html/2609.36174#bib.bib24);[Cousins, 2021](https://arxiv.org/html/2609.36174#bib.bib23)\)\.□\\square

Table EC\.1:Fairness Measures and Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)ComplianceNote\.For compactness, MN stands for monotonicity, CV for concavity, PI for permutation invariance, and CVI for constant vector invariance, as in Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)\. CE denotes certainty\-equivalent as in Example[2](https://arxiv.org/html/2609.36174#Thmexample2)\.

While our framework is expansive, Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)intentionally excludes certain fairness concepts to preserve mathematical tractability\. For example,threshold fairnessobjectives that seek to maximize the number of participants exceeding a specific performance baseline \(e\.g\.,\|\{n:vn≥c\}\|\|\\\{n:v\_\{n\}\\geq c\\\}\|\) are inherently step\-functions that fundamentally violate CV\. The resulting integer programming formulations are combinatorial and incompatible with gradient\-based policy optimization\([Chen and Hooker, 2023](https://arxiv.org/html/2609.36174#bib.bib26)\)\. Additionally,priority\-weighted welfarestructures where the platform assigns heterogeneous importance weights to specific individuals \(e\.g\., higher weights for important customers such that∑n=1Nwn​vn\\sum\_\{n=1\}^\{N\}w\_\{n\}v\_\{n\}wherewi≠wjw\_\{i\}\\neq w\_\{j\}\) violate PI\([Moulin, 1991](https://arxiv.org/html/2609.36174#bib.bib46)\)\.

## Appendix EC\.3Proofs of Section[3](https://arxiv.org/html/2609.36174#S3)

We start this section with some preliminary results regarding 1\) the effect of replacing a policy with one that has permuted indices on the value function of a symmetric major\-minor weakly coupled Markov decision process \(M2WCMDP\) in[EC\.3\.1](https://arxiv.org/html/2609.36174#A3.SS1); and 2\) a well\-known result from[Puterman \(2005\)](https://arxiv.org/html/2609.36174#bib.bib76)on the equivalency between stationary policies and occupancy measures in[EC\.3\.2](https://arxiv.org/html/2609.36174#A3.SS2)\. This is followed by the proof of Lemma[1](https://arxiv.org/html/2609.36174#Thmlemma1)on the uniform state\-value representation in[EC\.3\.3](https://arxiv.org/html/2609.36174#A3.SS3), which helps establish our main result, Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1)in[EC\.3\.4](https://arxiv.org/html/2609.36174#A3.SS4)\.

### EC\.3\.1Value Function under Permuted Policy for Symmetric M2WCMDP

###### Lemma EC\.1\.

If a M2WCMDP is symmetric, then for any policy𝛑\\bm\{\\pi\}and permutation operatorQQ, we have𝐕0​\(𝛑\)=Q​𝐕0​\(𝛑Q\)\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=Q\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\)andV00​\(𝛑\)=V00​\(𝛑Q\)V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\), where the permuted policy𝛑Q\(a0,𝐚∣s0,𝐬\):=𝛑\(a0,Q𝐚∣s0,Q𝐬\)\\bm\{\\pi\}^\{Q\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\):=\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\)for all\(s0,𝐬,a0,𝐚\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)tuples\.

This lemma implies an important equivalency in symmetric M2WCMDPs with identical minor sub\-MDPs\. If we permute the states and actions of a policy, the permuted version of the resulting value function is equivalent to the original value function\.

Proof\. Fix a permutation operatorQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}\. We can first show that for allt≥0t\\geq 0,

ℙ𝝅Q\(s0t=s0,𝒔t=𝒔,a0t=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)\\displaystyle\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\)\(EC\.1\)=ℙ𝝅\(s0t=s0,𝒔t=Q𝒔,a0t=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)\.\\displaystyle=\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\.
This can be done inductively\. Starting att=0t=0, we have that, for all\(s¯00,𝒔¯0,s0,𝒔,a0,𝒂\)\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\\mathchar 24891\\allowbreak s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):

ℙ𝝅Q\(s00=s0,𝒔0=𝒔,a00=a0,𝒂0=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)\\displaystyle\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\(s\_\{0\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{0\}=\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\)=𝝅Q\(a0,𝒂∣s0,𝒔\)𝕀\{s00=s¯00,𝒔=𝒔¯0\}\\displaystyle=\{\\bm\{\\pi\}^\{Q\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\}\{\\mathbb\{I\}\}\\\{s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}=\\bar\{\\bm\{s\}\}\_\{0\}\\\}=𝝅\(a0,Q𝒂∣s0,Q𝒔\)𝕀\{s00=s¯00,𝒔=𝒔¯0\}\\displaystyle=\{\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\)\{\\mathbb\{I\}\}\\\{s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}=\\bar\{\\bm\{s\}\}\_\{0\}\\\}\}=𝝅\(a0,Q𝒂∣s0,Q𝒔\)𝕀\{s00=s¯00,Q𝒔=Q𝒔¯0\}\\displaystyle=\{\\bm\{\\pi\}\(a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\)\}\{\\mathbb\{I\}\}\\\{s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\\\}=ℙ𝝅\(s00=s0,𝒔0=Q𝒔,a00=a0,𝒂0=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)\.\\displaystyle=\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s\_\{0\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{0\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\.
Next, assuming that Equation \([EC\.1](https://arxiv.org/html/2609.36174#A3.E1)\) holds for all\(s¯00,𝒔¯0,s0,𝒔,a0,𝒂\)\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\\mathchar 24891\\allowbreak s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\. By applying this to an arbitrary tuple\(s¯00,𝒔¯0,s~0,𝒔~,a~0,𝒂~\)\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\\mathchar 24891\\allowbreak\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{s\}\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{a\}\}\), we can show that it is also the case fort\+1t\+1, namely:

ℙ𝝅Q\(st\+10=s~0,𝒔t\+1=𝒔~,at\+10=a~0,𝒂t\+1=𝒂~\|s00=s¯00,𝒔0=𝒔¯0\)\\displaystyle\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\_\{t\+1\}=\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\+1\}=\\tilde\{\\bm\{s\}\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\+1\}=\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\+1\}=\\tilde\{\\bm\{a\}\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\)=𝝅Q\(a~0,𝒂~∣s~0,𝒔~\)∑s0,𝒔,a0,𝒂ℙ\(st\+10=s~0,𝒔t\+1=𝒔~\|st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\)\\displaystyle=\\bm\{\\pi\}^\{Q\}\(\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{a\}\}\\mid\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{s\}\}\)\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\mathbb\{P\}\(s^\{0\}\_\{t\+1\}=\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\+1\}=\\tilde\{\\bm\{s\}\}\|s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\)⋅ℙ𝝅Q\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)\\displaystyle\\qquad\\cdot\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\)=𝝅\(a~0,Q𝒂~∣s~0,Q𝒔~\)∑s0,𝒔,a0,𝒂p\(N\)\(𝒔~\|s0,𝒔,a0,𝒂\)p0\(s~0\|s0,𝒔,a0,𝒂\)\\displaystyle=\\bm\{\\pi\}\(\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak Q\\tilde\{\\bm\{a\}\}\\mid\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak Q\\tilde\{\\bm\{s\}\}\)\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}p^\{\(N\)\}\(\\tilde\{\\bm\{s\}\}\|s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)p^\{0\}\(\\tilde\{s\}^\{0\}\|s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)⋅ℙ𝝅\(st0=s0,𝒔t=Q𝒔,at0=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)\\displaystyle\\qquad\\cdot\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)=𝝅\(a~0,Q𝒂~∣s~0,Q𝒔~\)∑s0,𝒔,a0,𝒂p\(N\)\(Q𝒔~\|s0,Q𝒔,a0,Q𝒂\)p0\(s~0\|s0,Q𝒔,a0,Q𝒂\)\\displaystyle=\\bm\{\\pi\}\(\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak Q\\tilde\{\\bm\{a\}\}\\mid\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak Q\\tilde\{\\bm\{s\}\}\)\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}p^\{\(N\)\}\(Q\\tilde\{\\bm\{s\}\}\|s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)p^\{0\}\(\\tilde\{s\}^\{0\}\|s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)⋅ℙπ\(st0=s0,𝒔t=Q𝒔,at0=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)\\displaystyle\\qquad\\cdot\\mathbb\{P\}^\{\\pi\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)=ℙ𝝅\(st\+10=s~0,𝒔t\+1=Q𝒔~,at\+10=a~0,𝒂t\+1=Q𝒂~\|s00=s¯00,𝒔0=Q𝒔¯0\),\\displaystyle=\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\+1\}=\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\+1\}=Q\\tilde\{\\bm\{s\}\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\+1\}=\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\+1\}=Q\\tilde\{\\bm\{a\}\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\mathchar 24891\\allowbreakwhere we used the fact that the minor sub\-MDPs are identical sop\(N\)​\(𝒔~∣s0,𝒔,a0,𝒂\)=p\(N\)​\(Q​𝒔~∣s0,Q​𝒔,a0,Q​𝒂\)p^\{\(N\)\}\(\\tilde\{\\bm\{s\}\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=p^\{\(N\)\}\(Q\\tilde\{\\bm\{s\}\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)according to minor kernel equivariance \(Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\.1\) andp0​\(s~0∣s0,𝒔,a0,𝒂\)=p0​\(s~0∣s0,Q​𝒔,a0,Q​𝒂\)p^\{0\}\(\\tilde\{s\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=p^\{0\}\(\\tilde\{s\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)according to major kernel equivariance \(Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\.2\)\.

We now have that,

𝑽0\(𝝅Q\)=∑s0,𝒔,a0,𝒂∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅Q\(s0t=s0,𝒔t=𝒔,a0t=a0,𝒂t=𝒂∣s00=s¯00,𝒔0=𝒔¯0\)𝒓\(s0,𝒔,a0,𝒂\)\\displaystyle\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\)=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\\mid s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=∑s0,𝒔,a0,𝒂∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅\(s0t=s0,𝒔t=Q𝒔,a0t=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)𝒓\(s0,𝒔,a0,𝒂\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=∑s0,𝒔,a0,𝒂∑s¯00,𝒔¯0𝝁\(s¯00,Q𝒔¯0\)∑t=0∞γtℙ𝝅\(s0t=s0,𝒔t=Q𝒔,a0t=a0,𝒂t=Qa\|s00=s¯00,𝒔0=Q𝒔¯0\)𝒓\(s0,𝒔,a0,𝒂\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Qa\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=∑s0,𝒔,a0,𝒂∑s¯00,𝒔¯0𝝁\(s¯00,Q𝒔¯0\)∑t=0∞γtℙ𝝅\(s0t=s0,𝒔t=Q𝒔,a0t=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)Q−1𝒓\(s0,Q𝒔,a0,Q𝒂\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)Q^\{\-1\}\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)=Q−1\(∑s0,𝒔,a0,𝒂∑s¯00,𝒔¯0𝝁\(s¯00,Q𝒔¯0\)∑t=0∞γtℙ𝝅\(s0t=s0,𝒔t=Q𝒔,a0t=a0,𝒂t=Q𝒂\|s00=s¯00,𝒔0=Q𝒔¯0\)𝒓\(s0,Q𝒔,a0,Q𝒂\)\)\\displaystyle=Q^\{\-1\}\\left\(\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=Q\\bm\{a\}\|s\_\{0\}^\{0\}=\\bar\{s\}\_\{0\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=Q\\bar\{\\bm\{s\}\}\_\{0\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)\\right\)=Q−1\(∑s0,𝒔′,a0,𝒂′∑s¯00,𝒔¯0′𝝁\(s¯00,𝒔¯0′\)∑t=0∞γtℙ𝝅\(s0t=s0,𝒔t=𝒔′,a0t=a0,𝒂t=𝒂′\|s00=s¯00,𝒔0=𝒔¯0′\)𝒓\(s0,𝒔′,a0,𝒂′\)\)\\displaystyle=Q^\{\-1\}\\left\(\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}^\{\\prime\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}^\{\\prime\}\}\\sum\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}^\{\\prime\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}^\{\\prime\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\(s^\{0\}\_\{t\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}^\{\\prime\}\\mathchar 24891\\allowbreak a^\{0\}\_\{t\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}^\{\\prime\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}^\{\\prime\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}^\{\\prime\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}^\{\\prime\}\)\\right\)=Q−1​𝑽0​\(𝝅\)\\displaystyle=Q^\{\-1\}\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)where we first use the relation betweenℙ𝝅Q\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}and𝝅\\bm\{\\pi\}, then exploit the permutation invariance of𝝁\\bm\{\\mu\}\. We then exploit the permutation invariance𝒓⁡\(s0,Q​𝒔,a0,Q​𝒂\)=Q−1​𝒓​\(s0,𝒔,a0,𝒂\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)=Q^\{\-1\}\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\), maintaining consistent multiplication byQ−1Q^\{\-1\}, and reindex the summations using𝒔′:=Q​𝒔\\bm\{s\}^\{\\prime\}:=Q\\bm\{s\},𝒂′:=Q​𝒂\\bm\{a\}^\{\\prime\}:=Q\\bm\{a\}, and𝒔¯0′:=Q​𝒔¯0\\bar\{\\bm\{s\}\}\_\{0\}^\{\\prime\}:=Q\\bar\{\\bm\{s\}\}\_\{0\}\.

Furthermore,V00​\(𝝅\)=V00​\(𝝅Q\)V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\)follows by the same reindexing argument using the scalar invariancer0​\(s0,𝒔,a0,𝒂\)=r0​\(s0,Q​𝒔,a0,Q​𝒂\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\)as in Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\.□\\square

### EC\.3\.2Mapping between stationary policies and occupancy measures

###### Lemma EC\.2\.

\(Theorem 6\.9\.1 of[Puterman \(2005\)](https://arxiv.org/html/2609.36174#bib.bib76)\) LetΠ\\Pidenote the set of stationary stochastic Markov policies and𝒳\\mathcal\{X\}the set of occupancy measures\.

1. 1\.For any policy𝝅∈Π\\bm\{\\pi\}\\in\\Pi, an occupancy measureq𝝅:𝒮0×𝒮\(N\)×𝒜0×𝒜\(N\)→\[0,1/\(1−γ\)\]∈𝒳q\_\{\\bm\{\\pi\}\}:\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\\times\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\\rightarrow\{\[0\\mathchar 24891\\allowbreak 1/\(1\-\\gamma\)\]\}\\in\\mathcal\{X\}is obtained as q𝝅\(s0,𝒔,a0,𝒂\):=∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\),q\_\{\\bm\{\\pi\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\sum\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\\left\(s\_\{t\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\\right\)\\mathchar 24891\\allowbreak\(EC\.2\)for alls0∈𝒮0,𝒔∈𝒮\(N\)s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\in\\mathcal\{S\}^\{\(N\)\}anda0∈𝒜0,𝒂∈𝒜\(N\)a^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\in\\mathcal\{A\}^\{\(N\)\}\.
2. 2\.For any occupancy measureq⁡\(s0,𝒔,a0,𝒂\):𝒮0×𝒮\(N\)×𝒜0×𝒜\(N\)→\[0,1/\(1−γ\)\]∈𝒳q\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\\times\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\\rightarrow\{\[0\\mathchar 24891\\allowbreak 1/\(1\-\\gamma\)\]\}\\in\\mathcal\{X\}, a policy𝝅q\\bm\{\\pi\}\_\{q\}can be constructed as 𝝅q\(a0,𝒂∣s0,𝒔\):=q⁡\(s0,𝒔,a0,𝒂\)∑a0,𝒂q⁡\(s0,𝒔,a0,𝒂\),\\bm\{\\pi\}\_\{q\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\):=\\frac\{q\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\}\{\\sum\\limits\_\{a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\\left\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\right\)\}\\mathchar 24891\\allowbreak\(EC\.3\)for alls0∈𝒮0,𝒔∈𝒮\(N\)s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\in\\mathcal\{S\}^\{\(N\)\}anda0∈𝒜0,𝒂∈𝒜\(N\)a^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\in\\mathcal\{A\}^\{\(N\)\}such that∑a0,𝒂q⁡\(s0,𝒔,a0,𝒂\)\>0\\sum\_\{a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\\mathopen\{\(\}s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mathclose\{\)\}\>0and letting the policy be arbitrary on unreachable states\.

One necessarily has that for allq∈𝒳q\\in\\mathcal\{X\},q=qπqq=q\_\{\\pi\_\{q\}\}\.

Now, we show that the value functions can be represented using occupancy measures\.

###### Lemma EC\.3\.

For any policy𝛑∈Π\\bm\{\\pi\}\\in\\Pi, and the occupancy measureq𝛑q\_\{\\bm\{\\pi\}\}defined by \([EC\.2](https://arxiv.org/html/2609.36174#A3.E2)\), the expected total discounted rewards under the policy𝛑\\bm\{\\pi\}can be expressed as:

𝑽0​\(𝝅\)=∑s0,𝒔,a0,𝒂q𝝅​\(s0,𝒔,a0,𝒂\)​𝒓​\(s0,𝒔,a0,𝒂\)\.\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bm\{\\pi\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.\(EC\.4\)and

V00​\(𝝅\)=∑s0,𝒔,a0,𝒂q𝝅​\(s0,𝒔,a0,𝒂\)​r0​\(s0,𝒔,a0,𝒂\),V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bm\{\\pi\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\mathchar 24891\\allowbreak\(EC\.5\)

Proof\. Expanding the expected total discounted rewards𝑽0​\(𝝅\)\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\(as defined by Equation[5](https://arxiv.org/html/2609.36174#S2.E5)and[4](https://arxiv.org/html/2609.36174#S2.E4)\), we have:

𝑽0\(𝝅\)=∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑s0,𝒔,a0,𝒂∑t=0∞γtℙ𝝅\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)𝒓\(s0,𝒔,a0,𝒂\)\.\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=\\sum\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\\left\(s\_\{t\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\\right\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.Rearranging the terms:

𝑽0\(𝝅\)=∑s0,𝒔,a0,𝒂\(∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)\)𝒓\(s0,𝒔,a0,𝒂\)\.\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=\\sum\\limits\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\mathopen\{\(\}\\sum\\limits\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\\limits\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}\}\\mathopen\{\(\}s\_\{t\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\\mathclose\{\)\}\\mathclose\{\)\}\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.

By replacing the term in parentheses by the occupancy measureq𝝅​\(s0,𝒔,a0,𝒂\)q\_\{\\bm\{\\pi\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)in \([EC\.2](https://arxiv.org/html/2609.36174#A3.E2)\), we got:

𝑽0​\(𝝅\)=∑s0,𝒔,a0,𝒂q𝝅​\(s0,𝒔,a0,𝒂\)​𝒓​\(s0,𝒔,a0,𝒂\)\.\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}\}\)=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bm\{\\pi\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.Similarly as above, we can prove equation \([EC\.5](https://arxiv.org/html/2609.36174#A3.E5)\)\. This completes the proof\.□\\square

### EC\.3\.3Proof of Lemma[1](https://arxiv.org/html/2609.36174#Thmlemma1)

###### Lemma 1\(Uniform State\-Value Representation\)\.

If a M2WCMDP is symmetric \(Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\), then for any policy𝛑\\bm\{\\pi\}, there exists a corresponding permutation\-invariant policy𝛑¯\\bar\{\\bm\{\\pi\}\}such that the vector of expected total discounted rewards for all sub\-MDPs under𝛑¯\\bar\{\\bm\{\\pi\}\}is equal to the average of the expected total discounted rewards for each sub\-MDP, i\.e\.,𝐕0​\(𝛑¯\)=V¯0​\(𝛑\)​𝟏\\bm\{V\}\_\{0\}\(\\bar\{\\bm\{\\pi\}\}\)=\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\)\{\\bm\{1\}\}, whereV¯0​\(𝛑\):=1N​∑n=1NV0n​\(𝛑\)\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\):=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\\bm\{\\pi\}\), andV00​\(𝛑¯\)=V00​\(𝛑\)V^\{0\}\_\{0\}\(\\bar\{\\bm\{\\pi\}\}\)=\{V\}^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\.

Proof by construction\. First, for any fixedQQ, we characterize its occupancy measureq𝝅Qq\_\{\\bm\{\\pi\}^\{Q\}\}as

q𝝅Q\(s0,𝒔,a0,𝒂\):=∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅Q\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂∣s00=s¯00,𝒔0=𝒔¯0\)\.q\_\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\sum\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bm\{\\pi\}^\{Q\}\}\\left\(s\_\{t\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\\mid s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\\right\)\.\(EC\.6\)
Next, we construct a new measureq¯\\bar\{q\}obtained by averaging all permuted occupancy measuresq𝝅Qq\_\{\\bm\{\\pi\}^\{Q\}\}forQ∈𝒢NQ\\in\\mathcal\{G\}^\{N\}on all\(s0,𝒔,a0,𝒂\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)pairs as

q¯​\(s0,𝒔,a0,𝒂\):=1N\!​∑Qq𝝅Q​\(s0,𝒔,a0,𝒂\)\.\\bar\{q\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=\\frac\{1\}\{N\!\}\\sum\_\{Q\}q\_\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.\(EC\.7\)
One can confirm thatq¯\\bar\{q\}is an occupancy measure, i\.e\.,q¯∈𝒳\\bar\{q\}\\in\\mathcal\{X\}, since∀Q∈𝒢N\\forall Q\\in\\mathcal\{G\}^\{N\}, eachq𝝅Q∈𝒳q\_\{\\bm\{\\pi\}^\{Q\}\}\\in\\mathcal\{X\}and𝒳\\mathcal\{X\}is a convex polytope \(Puterman 6\.9\.2\)\. Crucially, one can show that the averaged policy𝝅¯\\bar\{\\bm\{\\pi\}\}inherits pathwise feasibility\. Because every𝝅Q\(⋅∣s0,𝒔\)\\bm\{\\pi\}^\{Q\}\(\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)only assigns probability mass within𝒜\(N\)​\(s0,Q​𝒔\)=Q​𝒜\(N\)​\(s0,𝒔\)\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\)=Q\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\), and the feasible set is permutation\-invariant \(from Definition[2](https://arxiv.org/html/2609.36174#Thmdefinition2)\), the support of𝝅¯\(⋅∣s0,𝒔\)\\bar\{\\bm\{\\pi\}\}\(\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)remains strictly inside the feasible joint action set𝒜\(N\)​\(s0,𝒔\)\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\.

From Lemma[EC\.2](https://arxiv.org/html/2609.36174#Thmlemma2a), a stationary policy𝝅¯\\bar\{\\bm\{\\pi\}\}can be constructed such that its occupancy measure matchesq¯​\(s0,𝒔,a0,𝒂\)\\bar\{q\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\. Namely, for all\(s0,𝒔,a0,𝒂\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)pairs,

q𝝅¯​\(s0,𝒔,a0,𝒂\)\\displaystyle q\_\{\\bar\{\\bm\{\\pi\}\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\):=∑s¯00,𝒔¯0𝝁\(s¯00,𝒔¯0\)∑t=0∞γtℙ𝝅¯\(st0=s0,𝒔t=𝒔,at0=a0,𝒂t=𝒂\|s00=s¯00,𝒔0=𝒔¯0\)\\displaystyle:=\\sum\_\{\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\}\\bm\{\\mu\}\(\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\_\{0\}\)\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\mathbb\{P\}^\{\\bar\{\\bm\{\\pi\}\}\}\\mathopen\{\(\}s\_\{t\}^\{0\}=s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{t\}=\\bm\{s\}\\mathchar 24891\\allowbreak a\_\{t\}^\{0\}=a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\_\{t\}=\\bm\{a\}\|s^\{0\}\_\{0\}=\\bar\{s\}^\{0\}\_\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\_\{0\}=\\bar\{\\bm\{s\}\}\_\{0\}\\mathclose\{\)\}=q¯​\(s0,𝒔,a0,𝒂\)\.\\displaystyle=\\bar\{q\}\(\{s\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\.
We can then derive the following steps:

𝑽0​\(𝝅¯\)\\displaystyle\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}\}\)=∑s0,𝒔,a0,𝒂q𝝅¯\(s0,𝒔,a0,𝒂\)𝒓\(s0,𝒔,a0,𝒂\)\(By Lemma[EC\.3](https://arxiv.org/html/2609.36174#Thmlemma3)\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bar\{\\bm\{\\pi\}\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\quad\\mbox\{\(By Lemma \\ref\{prop:xr=muv\}\)\}=∑s0,𝒔,a0,𝒂q¯\(s0,𝒔,a0,𝒂\)𝒓\(s0,𝒔,a0,𝒂\)\(By Lemma[EC\.2](https://arxiv.org/html/2609.36174#Thmlemma2a)\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\bar\{q\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\quad\\mbox\{\(By Lemma \\ref\{prop:x\-pi\-x\}\)\}=1N\!∑Q∈𝒢N∑s0,𝒔,a0,𝒂q𝝅Q\(s0,𝒔,a0,𝒂\)𝒓\(s0,𝒔,a0,𝒂\)\(By construction in \([EC\.7](https://arxiv.org/html/2609.36174#A3.E7)\)\)\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\bm\{r\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\quad\\mbox\{\(By construction in \\eqref\{eq:bar\-pi\-x\}\) \}=1N\!∑Q∈𝒢N𝑽0\(𝝅Q\)\(By Lemma[EC\.3](https://arxiv.org/html/2609.36174#Thmlemma3)\)\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\)\\quad\\mbox\{\(By Lemma \\ref\{prop:xr=muv\}\)\}=1N\!∑Q∈𝒢NQ−1𝑽0\(𝝅\)\(By Lemma[EC\.1](https://arxiv.org/html/2609.36174#Thmlemma1a)\)\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}Q^\{\-1\}\\bm\{V\}\_\{0\}\(\\bm\{\\pi\}\)\\quad\\mbox\{\(By Lemma \\ref\{thm:piQVQ\}\)\}=1N\!∑Q∈𝒢NQ−1\[V01​\(𝝅\)⋯V0N​\(𝝅\)\]\(Vector form\)\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}Q^\{\-1\}\\begin\{bmatrix\}V\_\{0\}^\{1\}\(\\bm\{\\pi\}\)\\\\ \\cdots\\\\ V\_\{0\}^\{N\}\(\\bm\{\\pi\}\)\\\\ \\end\{bmatrix\}\\quad\\text\{\(Vector form\)\}=1N\!\[\(N−1\)\!​∑n=1NV0n​\(𝝅\)⋯\(N−1\)\!​∑n=1NV0n​\(𝝅\)\]\(Property of permutation group\)\\displaystyle=\\frac\{1\}\{N\!\}\\begin\{bmatrix\}\(N\-1\)\!\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\\bm\{\\pi\}\)\\\\ \\cdots\\\\ \(N\-1\)\!\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\\bm\{\\pi\}\)\\end\{bmatrix\}\\quad\\text\{\(Property of permutation group\)\}=1N\!​\(N−1\)\!​∑n=1NV0n​\(𝝅\)​𝟏=1N​∑n=1NV0n​\(𝝅\)​𝟏=V¯0​\(𝝅\)​𝟏\.\\displaystyle=\\frac\{1\}\{N\!\}\(N\-1\)\!\\sum\_\{n=1\}^\{N\}V^\{n\}\_\{0\}\(\\bm\{\\pi\}\)\{\\bm\{1\}\}=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\\bm\{\\pi\}\)\{\\bm\{1\}\}=\\bar\{V\}\_\{0\}\(\\bm\{\\pi\}\)\{\\bm\{1\}\}\.
Similarly, we have

V00​\(𝝅¯\)\\displaystyle V^\{0\}\_\{0\}\(\\bar\{\\bm\{\\pi\}\}\)=∑s0,𝒔,a0,𝒂q𝝅¯​\(s0,𝒔,a0,𝒂\)​r0​\(s0,𝒔,a0,𝒂\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bar\{\\bm\{\\pi\}\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=∑s0,𝒔,a0,𝒂q¯​\(s0,𝒔,a0,𝒂\)​r0​\(s0,𝒔,a0,𝒂\)\\displaystyle=\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}\\bar\{q\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=1N\!​∑Q∈𝒢N∑s0,𝒔,a0,𝒂q𝝅Q​\(s0,𝒔,a0,𝒂\)​r0​\(s0,𝒔,a0,𝒂\)\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}\\sum\_\{s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\}q\_\{\\bm\{\\pi\}^\{Q\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)r^\{0\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=1N\!​∑Q∈𝒢NV00​\(𝝅Q\)=1N\!​∑Q∈𝒢NV00​\(𝝅\)=V00​\(𝝅\),\\displaystyle=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}V^\{0\}\_\{0\}\(\\bm\{\\pi\}^\{Q\}\)=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)=V^\{0\}\_\{0\}\(\\bm\{\\pi\}\)\\mathchar 24891\\allowbreakrelying on the scalar major identityV00​\(𝝅Q\)=V00​\(𝝅\)V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{Q\}\}\)=V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}\}\)established in Lemma[EC\.1](https://arxiv.org/html/2609.36174#Thmlemma1a)\.□\\square

### EC\.3\.4Proof of Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1)

###### Theorem 1\(Utilitarian Reduction\)\.

Under Assumption[1](https://arxiv.org/html/2609.36174#Thmassumption1), for a symmetric M2WCMDP, letΠU,PI∗\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}be the set of optimal policies for the utilitarian approach that is permutation\-invariant, thenΠU,PI∗\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}is necessarily non\-empty and all𝛑U,PI∗∈ΠU,PI∗\\bm\{\\pi\}^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}\\in\\Pi^\{\*\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\scriptsize PI\}\}satisfyGρ​\(𝛑U,PI∗\)=max𝛑∈Π⁡Gρ​\(𝛑\)\.G\_\{\\rho\}\(\\bm\{\\pi\}\_\{U\\mathchar 24891\\allowbreak\\mbox\{\\tiny\{PI\}\}\}^\{\*\}\)=\\max\\limits\_\{\\bm\{\\pi\}\\in\\Pi\}G\_\{\\rho\}\(\\bm\{\\pi\}\)\.

Proof\. We define the utilitarian fairness measureρ\\rhoasρU​\[𝒗\]=1N​∑n=1Nvn\\rho\_\{U\}\[\{\\bm\{v\}\}\]=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}v\_\{n\}\. Let the optimal policy to the optimization problem \([6](https://arxiv.org/html/2609.36174#S2.E6)\) with utilitarian objective be𝝅U∗∈arg⁡max𝝅​GU​\(𝝅\)\\bm\{\\pi\}^\{\*\}\_\{U\}\\in\\arg\\max\\limits\_\{\\bm\{\\pi\}\}G\_\{U\}\(\\bm\{\\pi\}\)\.

With Lemma[1](https://arxiv.org/html/2609.36174#Thmlemma1), we can construct a permutation\-invariant policyπ¯U∗\\bar\{\\pi\}^\{\*\}\_\{U\}satisfying

V¯0​\(𝝅U∗\)​𝟏=𝑽0​\(𝝅¯U∗\)∈𝒱,andV00​\(𝝅U∗\)=V00​\(𝝅¯U∗\),\\bar\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\{\\bm\{1\}\}=\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\in\{\\mathcal\{V\}\}\\mathchar 24891\\allowbreak\\quad\\text\{and\}\\quad V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)=V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\mathchar 24891\\allowbreak\(EC\.8\)withV¯0​\(𝝅U∗\):=1N​∑n=1NV0n​\(𝝅U∗\)\\bar\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\):=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}V\_\{0\}^\{n\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\. Then we have that

V00​\(𝝅U∗\)\+λ​ρU​\[𝑽0​\(𝝅U∗\)\]\\displaystyle V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\\right\]=V00​\(𝝅¯U∗\)\+λ​ρU​\[V0¯​\(𝝅U∗\)​𝟏\]\\displaystyle=V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bar\{V\_\{0\}\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\{\\bm\{1\}\}\\right\]=V00​\(𝝅¯U∗\)\+λ​ρ​\[V0¯​\(𝝅U∗\)​𝟏\]\\displaystyle=V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\\left\[\\bar\{V\_\{0\}\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\{\\bm\{1\}\}\\right\]=V00​\(𝝅¯U∗\)\+λ​ρ​\[𝑽0​\(𝝅¯U∗\)\],\\displaystyle=V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\right\]\\mathchar 24891\\allowbreakwhere we exploitρ⁡\[v¯​𝟏\]=v¯\\rho\[\\bar\{v\}\{\\bm\{1\}\}\]=\\bar\{v\}for allv¯∈ℝ\\bar\{v\}\\in\{\\mathbb\{R\}\}such thatv¯​𝟏∈𝒱\\bar\{v\}\{\\bm\{1\}\}\\in\{\\mathcal\{V\}\}\.

Further, let𝝅∗\\bm\{\\pi\}^\{\*\}be any optimal policy to theρ\\rho\-M2WCMDP problem\. One can establish that:

V00​\(𝝅∗\)\+λ​ρ​\[𝑽0​\(𝝅∗\)\]\\displaystyle V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\{\\bm\{\\pi\}\}^\{\*\}\}\)\\right\]≥V00​\(𝝅¯U∗\)\+λ​ρ​\[𝑽0​\(𝝅¯U∗\)\]\\displaystyle\\geq V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\right\]\(EC\.9\)=V00​\(𝝅U∗\)\+λ​ρU​\[𝑽0​\(𝝅U∗\)\]\\displaystyle=V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\\right\]≥V00​\(𝝅∗\)\+λ​ρU​\[𝑽0​\(𝝅∗\)\],\\displaystyle\\geq V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bm\{V\}\_\{0\}\(\{\{\\bm\{\\pi\}\}^\{\*\}\}\)\\right\]\\mathchar 24891\\allowbreakwhere the first inequality holds because𝝅∗\\bm\{\\pi\}^\{\*\}is optimal forGρG\_\{\\rho\}and𝝅¯U∗∈Π\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\\in\\Pi\(it is a stationary Markov policy by Lemma 1, and feasible because it inherits feasibility from𝝅U∗\\bm\{\\pi\}\_\{U\}^\{\*\}\)\.

By Jensen’s inequality and the fact thatρ⁡\[⋅\]\\rho\[\\cdot\]is concave, we have that

ρ⁡\[𝒗\]=1N\!​∑Q∈𝒢Nρ⁡\[Q​𝒗\]≤ρ⁡\[1N\!​∑Q∈𝒢NQ​𝒗\]=ρ⁡\[1N​∑n=1Nvn​𝟏\]=1N​∑n=1Nvn=ρU​\[𝒗\],\\rho\[\{\\bm\{v\}\}\]=\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}\\rho\[Q\{\\bm\{v\}\}\]\\leq\\rho\[\\frac\{1\}\{N\!\}\\sum\_\{Q\\in\\mathcal\{G\}^\{N\}\}Q\{\\bm\{v\}\}\]=\\rho\[\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}v\_\{n\}\{\\bm\{1\}\}\]=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}v\_\{n\}=\\rho\_\{U\}\[\{\\bm\{v\}\}\]\\mathchar 24891\\allowbreakwhere we use permutation invariance ofρ\\rho, followed with its concavity and its constant vector invariance\. Then, it holds thatV00​\(𝝅∗\)\+λ​ρU​\[𝑽0​\(𝝅∗\)\]≥V00​\(𝝅∗\)\+λ​ρ​\[𝑽0​\(𝝅∗\)\]V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\+\\lambda\\rho\_\{U\}\\mathopen\{\[\}\\bm\{V\}\_\{0\}\(\{\{\\bm\{\\pi\}\}^\{\*\}\}\)\\mathclose\{\]\}\\geq V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\+\\lambda\\rho\\mathopen\{\[\}\\bm\{V\}\_\{0\}\(\{\{\\bm\{\\pi\}\}^\{\*\}\}\)\\mathclose\{\]\}\. The inequalities in \([EC\.9](https://arxiv.org/html/2609.36174#A3.E9)\) should therefore all reach equality:

V00​\(𝝅∗\)\+λ​ρ​\[𝑽0​\(𝝅∗\)\]=V00​\(𝝅¯U∗\)\+λ​ρ​\[𝑽0​\(𝝅¯U∗\)\]=V00​\(𝝅U∗\)\+λ​ρU​\[𝑽0​\(𝝅U∗\)\]\.V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\\right\]=V^\{0\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\right\]=V^\{0\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\\right\]\.
Therefore,

Gρ​\(𝝅¯U∗\)=V0​\(𝝅¯U∗\)\+λ​ρ​\[𝑽0​\(𝝅¯U∗\)\]=V0​\(𝝅U∗\)\+λ​ρU​\[𝑽0​\(𝝅U∗\)\]=V0​\(𝝅∗\)\+λ​ρ​\[𝑽0​\(𝝅∗\)\]=maxπ⁡Gρ​\(𝝅\)\.G\_\{\\rho\}\(\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\)=V\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bar\{\\bm\{\\pi\}\}^\{\*\}\_\{U\}\}\)\\right\]=V\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\+\\lambda\\rho\_\{U\}\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\_\{U\}\}\)\\right\]=\{V\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\}\+\\lambda\\rho\\left\[\\bm\{V\}\_\{0\}\(\{\\bm\{\\pi\}^\{\*\}\}\)\\right\]=\\max\_\{\\pi\}G\_\{\\rho\}\(\\bm\{\\pi\}\)\.□\\square

## Appendix EC\.4Proof of Section[4](https://arxiv.org/html/2609.36174#S4)

This section proves the feasibility and deterministic\-coverage properties of the priority\-based sampler stated in Lemma[2](https://arxiv.org/html/2609.36174#Thmlemma2)\.

###### Lemma 2\(Feasibility and deterministic coverage\)\.

Given that there exists a fallback actiona^∈𝒜\\hat\{a\}\\in\\mathcal\{A\}such thatdk​\(a^∣s\)=0d\_\{k\}\(\\hat\{a\}\\mid s\)=0for alls∈𝒮s\\in\\mathcal\{S\}and allk∈\[K\]k\\in\[K\], the sampling procedure described in Algorithm[1](https://arxiv.org/html/2609.36174#alg1), has the following properties:

1. 1\.ℙ⁡\(𝒖∈𝒜g\(N\)​\(s0,𝒙\)\)=1\\mathbb\{P\}\(\\bm\{u\}\\in\{\\mathcal\{A\}^\{\(N\)\}\_\{g\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\)=1for alls0∈𝒮0s^\{0\}\\in\\mathcal\{S\}^\{0\}and𝒙∈𝒮f\(N\)\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\.
2. 2\.If\{𝒜c\}c∈𝒞=𝒜\\\{\{\\mathcal\{A\}\}\_\{c\}\\\}\_\{c\\in\{\\mathcal\{C\}\}\}=\{\\mathcal\{A\}\}, then for any deterministic policy𝝅ϕD:𝒮0×𝒮f\(N\)→\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}:\\mathcal\{S\}^\{0\}\\times\{\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\rightarrow\{0,…,N\}\|𝒮\|×\|𝒜\|\\\{0\\mathchar 24891\\allowbreak\\dots\\mathchar 24891\\allowbreak N\\\}^\{\|\{\\mathcal\{S\}\}\|\\times\|\{\\mathcal\{A\}\}\|\}satisfying𝝅ϕD​\(s0,𝒙\)∈𝒜g\(N\)​\(s0,𝒙\)\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\), there exists a mappingh:𝒮0×𝒮f\(N\)→\[0,1\]K×Δ​\(𝒞\)\|𝒮\|×\[0,1\]\|𝒮\|×\|𝒜\|h:\\mathcal\{S\}^\{0\}\\times\{\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\rightarrow\[0\\mathchar 24891\\allowbreak 1\]^\{K\}\\times\\Delta\(\{\\mathcal\{C\}\}\)^\{\|\\mathcal\{S\}\|\}\\times\[0\\mathchar 24891\\allowbreak 1\]^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}such that a sample𝒖∼𝝅h⁡\(s0,𝒙\)\(⋅∣s0,𝒙\)\)=𝝅ϕD\(s0,𝒙\)\\bm\{u\}\\sim\\bm\{\\pi\}\_\{h\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\}\(\\cdot\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\)=\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)with probability one\.

Proof\. The first property can be verified by following the updates of𝒖\\bm\{u\}made by the algorithm and confirming it never violates the resource consumption condition\. In fact, throughout the procedure whenever\(𝒃~,𝒖\)\(\\tilde\{\{\\bm\{b\}\}\}\\mathchar 24891\\allowbreak\\bm\{u\}\)are updated to\(𝒃~′,𝒖′\)\(\\tilde\{\{\\bm\{b\}\}\}^\{\\prime\}\\mathchar 24891\\allowbreak\\bm\{u\}^\{\\prime\}\), one can confirm that the updated values satisfyb~k′\+∑s,adk​\(a∣s\)​u′​\(s,a\)=b~k\+∑s,adk​\(a∣s\)​u​\(s,a\)\\tilde\{b\}\_\{k\}^\{\\prime\}\+\\sum\_\{s\\mathchar 24891\\allowbreak a\}d\_\{k\}\(a\\mid s\)u^\{\\prime\}\(s\\mathchar 24891\\allowbreak a\)=\\tilde\{b\}\_\{k\}\+\\sum\_\{s\\mathchar 24891\\allowbreak a\}d\_\{k\}\(a\\mid s\)u\(s\\mathchar 24891\\allowbreak a\)and thatb~k′≥0\\tilde\{b\}\_\{k\}^\{\\prime\}\\geq 0\. Since\(𝒃~,𝒖\)\(\\tilde\{\{\\bm\{b\}\}\}\\mathchar 24891\\allowbreak\\bm\{u\}\)is initialized at\(𝒃⁡\(s0\)⋅𝒑~,𝟎\|𝒮\|×\|𝒜\|\)\(\{\\bm\{b\}\}\(s^\{0\}\)\\cdot\\tilde\{\{\\bm\{p\}\}\}\\mathchar 24891\\allowbreak\\bm\{0\}\_\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}\), this implies that when the procedure terminates:

b~k\+∑s,adk​\(a∣s\)​u​\(s,a\)=bk​\(s0\)​p~k≤bk​\(s0\)\.\\tilde\{b\}\_\{k\}\+\\sum\_\{s\\mathchar 24891\\allowbreak a\}d\_\{k\}\(a\\mid s\)u\(s\\mathchar 24891\\allowbreak a\)=b\_\{k\}\(s^\{0\}\)\\tilde\{p\}\_\{k\}\\leq b\_\{k\}\(s^\{0\}\)\.One can further confirm that, at termination,∑au⁡\(s,a\)=x⁡\(s\)\\sum\_\{a\}u\(s\\mathchar 24891\\allowbreak a\)=x\(s\)for allssdue to the last step of the algorithm\.

The second property is achieved by settingh⁡\(s0,𝒙\):=\(𝟏K,hW​\(s0,𝒙\),𝟏\|𝒮\|×\|𝒜\|\)h\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\):=\{\(\{\\bm\{1\}\}\_\{K\}\\mathchar 24891\\allowbreak h\_\{W\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\mathchar 24891\\allowbreak\{\\bm\{1\}\}\_\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{A\}\|\}\)\}, with\[hW​\(s0,𝒙\)\]​\(s,c\):=\[𝝅ϕD​\(s0,𝒙\)\]​\(s,c\)/∑a\[𝝅ϕD​\(s0,𝒙\)\]​\(s,a\)\[h\_\{W\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak c\):=\[\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak c\)/\\sum\_\{a\}\[\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak a\)for∑a\[𝝅ϕD​\(s0,𝒙\)\]​\(s,a\)\>0\\sum\_\{a\}\[\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak a\)\>0\. If the denominator equals 0, the corresponding row ofhW​\(s0,𝒙\)h\_\{W\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)can be chosen arbitrarily\. Fixing some\(s0,𝒙\)\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\), referring toccasaagiven the assumed partition, and using shorthand notationshW​\(s,c\)h\_\{W\}\(s\\mathchar 24891\\allowbreak c\)anduD​\(s,a\)u^\{D\}\(s\\mathchar 24891\\allowbreak a\)for\[hW​\(s0,𝒙\)\]​\(s,c\)\[h\_\{W\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak c\)and\[𝝅ϕD​\(s0,𝒙\)\]​\(s,a\)\[\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\]\(s\\mathchar 24891\\allowbreak a\)respectively, one can follow the steps of the algorithm\. For eacha∈𝒜a\\in\\mathcal\{A\}, the procedure first initializesB⁡\(s,a\):=⌊x⁡\(s\)⋅hW​\(s,a\)⌋=uD​\(s,a\)B\(s\\mathchar 24891\\allowbreak a\):=\\lfloor x\(s\)\\cdot h\_\{W\}\(s\\mathchar 24891\\allowbreak a\)\\rfloor=u^\{D\}\(s\\mathchar 24891\\allowbreak a\)exactly since𝒖D∈𝒜g\(N\)​\(s0,𝒙\)\\bm\{u\}^\{D\}\\in\{\\mathcal\{A\}^\{\(N\)\}\_\{g\}\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)so that∑auD​\(s,a\)=x⁡\(s\)\\sum\_\{a\}u^\{D\}\(s\\mathchar 24891\\allowbreak a\)=x\(s\)\. Consequently,R⁡\(s\)=x⁡\(s\)−∑aB⁡\(s,a\)=0R\(s\)=x\(s\)\-\\sum\_\{a\}B\(s\\mathchar 24891\\allowbreak a\)=0, so the largest\-remainder assignment is vacuous\. Next given that this maximum planned budget on counts is feasible, namely∑s,adk​\(a∣s\)​B​\(s,a\)=∑s,adk​\(a∣s\)​uD​\(s,a\)≤bk​\(s0\)\\sum\_\{s\\mathchar 24891\\allowbreak a\}d\_\{k\}\(a\\mid s\)B\(s\\mathchar 24891\\allowbreak a\)=\\sum\_\{s\\mathchar 24891\\allowbreak a\}d\_\{k\}\(a\\mid s\)u^\{D\}\(s\\mathchar 24891\\allowbreak a\)\\leq b\_\{k\}\(s^\{0\}\)for allkk, we can conclude that samples of type\(s,a\)\(s\\mathchar 24891\\allowbreak a\)will be produced until the budgetB⁡\(s,a\)B\(s\\mathchar 24891\\allowbreak a\)runs out\. Thus, the state\-count planu⁡\(s,a\)u\(s\\mathchar 24891\\allowbreak a\)will increment up to and stop atuD​\(s,a\)u^\{D\}\(s\\mathchar 24891\\allowbreak a\)with probability one\. Consequently,𝒖=𝒖D=𝝅ϕD​\(s0,𝒙\)\{\\bm\{u\}\}=\\bm\{u\}^\{D\}=\{\\bm\{\\pi\}\}\_\{\\phi\}^\{D\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)with probability one\.□\\square

## Appendix EC\.5Count Aggregation Reformulation

Complementing Section[3\.2](https://arxiv.org/html/2609.36174#S3.SS2), the transition kernel and the initial distribution of the count aggregation M2WCMDPℳϕ\{\\mathcal\{M\}\}\_\{\\phi\}under the mappingϕ=\(f,g\)\\phi=\(f\\mathchar 24891\\allowbreak g\)is obtained as follows\.

##### Transition Probability

The minor transition probabilitypϕ\(N\)​\(𝒙~∣𝒙,s0,a0,𝒖\)p^\{\(N\)\}\_\{\\phi\}\(\\tilde\{\\bm\{x\}\}\\mid\\bm\{x\}\\mathchar 24891\\allowbreak s^\{0\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)is the probability that the system moves from count state𝒙\\bm\{x\}to𝒙~\\tilde\{\\bm\{x\}\}given the major state\-action pair\(s0,a0\)\(s^\{0\}\\mathchar 24891\\allowbreak a^\{0\}\)and the minor action counts𝒖\\bm\{u\}\. We define the pre\-imagef−1​\(𝒙~\)f^\{\-1\}\(\\tilde\{\\bm\{x\}\}\)as the set containing all labelled elements𝒔~∈𝒮\(N\)\\tilde\{\\bm\{s\}\}\\in\\mathcal\{S\}^\{\(N\)\}that map to𝒙~\\tilde\{\\bm\{x\}\}\.

Given the equivalence of transitions within the pre\-image set, for an arbitrary state\-action pair\(𝒔¯,𝒂¯\)∈ϕ−1​\(𝒙,𝒖\)\(\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)\\in\\phi^\{\-1\}\(\\bm\{x\}\\mathchar 24891\\allowbreak\\bm\{u\}\), the probability of transitioning from𝒙\\bm\{x\}to𝒙~\\tilde\{\\bm\{x\}\}under action𝒖\\bm\{u\}is the sum of the probabilities of all individual transitions in the original space that correspond to this count state:

pϕ\(N\)​\(𝒙~∣𝒙,s0,a0,𝒖\):=∑𝒔~∈f−1​\(𝒙~\)p\(N\)​\(𝒔~∣s0,𝒔¯,a0,𝒂¯\)=∑𝒔~∈f−1​\(𝒙~\)∏n=1Np⁡\(s~n∣s0,s¯n,a0,a¯n\)\.p^\{\(N\)\}\_\{\\phi\}\(\\tilde\{\\bm\{x\}\}\\mid\\bm\{x\}\\mathchar 24891\\allowbreak s^\{0\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\):=\\sum\_\{\\tilde\{\\bm\{s\}\}\\in f^\{\-1\}\(\\tilde\{\\bm\{x\}\}\)\}p^\{\(N\)\}\(\\tilde\{\\bm\{s\}\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)=\\sum\_\{\\tilde\{\\bm\{s\}\}\\in f^\{\-1\}\(\\tilde\{\\bm\{x\}\}\)\}\\prod\_\{n=1\}^\{N\}p\(\\tilde\{s\}^\{n\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bar\{s\}^\{n\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bar\{a\}^\{n\}\)\.The major transition probability is similarly derived aspϕ0​\(s~0∣s0,𝒙,a0,𝒖\):=p0​\(s~0∣s0,𝒔¯,a0,𝒂¯\)p^\{0\}\_\{\\phi\}\(\\tilde\{s\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\):=p^\{0\}\(\\tilde\{s\}^\{0\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{a\}\}\)\.

##### Initial Distribution

By using a state count representation for the symmetric M2WCMDP, we know that∑s∈𝒮x⁡\(s\)=N\\sum\_\{s\\in\\mathcal\{S\}\}x\(s\)=N\. The cardinality of the pre\-image setf−1​\(𝒙\)f^\{\-1\}\(\\bm\{x\}\)can be obtained through the multinomial expansion of\(s1\+s2\+⋯\+s\|𝒮\|\)N\(s\_\{1\}\+s\_\{2\}\+\\dots\+s\_\{\|\\mathcal\{S\}\|\}\)^\{N\}\. Intuitively, distributingNNidentical sub\-MDPs into\|𝒮\|\|\\mathcal\{S\}\|distinct states such that the counts match𝒙\\bm\{x\}is given by the multinomial coefficient\|f−1​\(𝒙\)\|=N\!∏s∈𝒮x⁡\(s\)\!\|f^\{\-1\}\(\\bm\{x\}\)\|=\\frac\{N\!\}\{\\prod\_\{s\\in\\mathcal\{S\}\}x\(s\)\!\}\. Given that the joint initial distribution𝝁⁡\(s0,𝒔\)\\bm\{\\mu\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)is permutation\-invariant, the pushforward probability of starting from major states0s^\{0\}and minor count state𝒙\\bm\{x\}is

𝝁f​\(s0,𝒙\):=∑𝒔∈f−1​\(𝒙\)𝝁⁡\(s0,𝒔\)=\|f−1​\(𝒙\)\|⋅𝝁⁡\(s0,𝒔¯\)=N\!∏s∈𝒮x⁡\(s\)\!⋅𝝁⁡\(s0,𝒔¯\),\\bm\{\\mu\}\_\{f\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\):=\\sum\_\{\\bm\{s\}\\in f^\{\-1\}\(\\bm\{x\}\)\}\\bm\{\\mu\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)=\|f^\{\-1\}\(\\bm\{x\}\)\|\\cdot\\bm\{\\mu\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\)=\\frac\{N\!\}\{\\prod\_\{s\\in\\mathcal\{S\}\}x\(s\)\!\}\\cdot\\bm\{\\mu\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bar\{\\bm\{s\}\}\)\\mathchar 24891\\allowbreakfor any arbitrary𝒔¯\\bar\{\\bm\{s\}\}such thatf⁡\(𝒔¯\)=𝒙f\(\\bar\{\\bm\{s\}\}\)=\\bm\{x\}\.

## Appendix EC\.6Exact Approaches for Solving Utilitarian\-Reduced Count M2WCMDP

Based on the utilitarian reduction established in Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1), the symmetricρ\\rho\-M2WCMDP problem can be solved through its utilitarian\-reduced count aggregation M2WCMDPℳϕ\{\\mathcal\{M\}\}\_\{\\phi\}\(see Definition[5](https://arxiv.org/html/2609.36174#Thmdefinition5)\)\. Using the exact form constructed in[EC\.5](https://arxiv.org/html/2609.36174#A5)and following the standard dual LP methods for discounted MDPs derived in Section 6\.9\.1 by[Puterman \(2005\)](https://arxiv.org/html/2609.36174#bib.bib76), we define the discounted state\-action occupancy measuresqϕ​\(s0,𝒙,a0,𝒖\)≥0q\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\\geq 0over the aggregated joint state space𝒮0×𝒮f\(N\)\\mathcal\{S\}^\{0\}\\times\\mathcal\{S\}^\{\(N\)\}\_\{f\}and the state\-dependent feasible joint action space𝒜0×𝒜g\(N\)​\(s0,𝒙\)\\mathcal\{A\}^\{0\}\\times\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\.

Under the utilitarian reduction, the objective is to maximize the expected total discounted sum of the platform rewardrϕ0r^\{0\}\_\{\\phi\}and the fairness\-weighted mean participant rewardλ​r¯ϕ\\lambda\\bar\{r\}\_\{\\phi\}\. The exact LP is formulated as follows:

maxqϕ∑s0∈𝒮0,𝒙∈𝒮f\(N\)∑a0∈𝒜0,𝒖∈𝒜g\(N\)​\(s0,𝒙\)\[rϕ0​\(s0,𝒙,a0,𝒖\)\+λ​r¯ϕ​\(s0,𝒙,a0,𝒖\)\]​qϕ​\(s0,𝒙,a0,𝒖\)s\.t\.∑a0∈𝒜0,𝒖∈𝒜g\(N\)​\(s0,𝒙\)qϕ​\(s0,𝒙,a0,𝒖\)−γ∑s~0∈𝒮0,𝒙~∈𝒮f\(N\)∑a~0∈𝒜0,𝒖~∈𝒜g\(N\)​\(s~0,𝒙~\)qϕ\(s~0,𝒙~,a~0,𝒖~\)p0ϕ\(s0∣s~0,𝒙~,a~0,𝒖~\)p\(N\)ϕ\(𝒙∣s~0,𝒙~,a~0,𝒖~\)=𝝁f​\(s0,𝒙\),∀s0∈𝒮0,𝒙∈𝒮f\(N\)qϕ​\(s0,𝒙,a0,𝒖\)≥0∀s0∈𝒮0,𝒙∈𝒮f\(N\),∀a0∈𝒜0,𝒖∈𝒜g\(N\)​\(s0,𝒙\)\\begin\{array\}\[\]\{rl\}\\max\\limits\_\{q\_\{\\phi\}\}&\\sum\\limits\_\{s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\sum\\limits\_\{a^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\}\\left\[r^\{0\}\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\+\\lambda\\bar\{r\}\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\\right\]q\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\\\\ \\textit\{s\.t\.\}&\\sum\\limits\_\{a^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\}q\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\\\\ &\\quad\-\\gamma\\hskip\-5\.0pt\\sum\\limits\_\{\\tilde\{s\}^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{x\}\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\sum\\limits\_\{\\tilde\{a\}^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{x\}\}\)\}\\hskip\-3\.0ptq\_\{\\phi\}\(\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{x\}\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\)p^\{0\}\_\{\\phi\}\(s^\{0\}\\mid\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{x\}\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\)p^\{\(N\)\}\_\{\\phi\}\(\\bm\{x\}\\mid\\tilde\{s\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{x\}\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\)\\\\ &\\quad=\\bm\{\\mu\}\_\{f\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\mathchar 24891\\allowbreak\\qquad\\forall s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\\\\ &q\_\{\\phi\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\\geq 0\\qquad\\forall s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\\mathchar 24891\\allowbreak\\forall a^\{0\}\\in\\mathcal\{A\}^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\in\\mathcal\{A\}^\{\(N\)\}\_\{g\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)\\end\{array\}\(EC\.10\)
Because𝝁f\\bm\{\\mu\}\_\{f\}is the pushforward of the original probability distribution𝝁\\bm\{\\mu\}, it satisfies∑s0∈𝒮0,𝒙∈𝒮f\(N\)𝝁f​\(s0,𝒙\)=1\\sum\_\{s^\{0\}\\in\\mathcal\{S\}^\{0\}\\mathchar 24891\\allowbreak\\,\\bm\{x\}\\in\\mathcal\{S\}^\{\(N\)\}\_\{f\}\}\\bm\{\\mu\}\_\{f\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\)=1\. For an optimal solutionqϕ∗​\(s0,𝒙,a0,𝒖\)q\_\{\\phi\}^\{\*\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\), an optimal count policy is recovered as

𝝅ϕ∗\(a0,𝒖∣s0,𝒙\):=qϕ∗​\(s0,𝒙,a0,𝒖\)∑a~0,𝒖~qϕ∗​\(s0,𝒙,a~0,𝒖~\),∀s0,𝒙,a0,𝒖,\\bm\{\\pi\}\_\{\\phi\}^\{\*\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\):=\\frac\{q\_\{\\phi\}^\{\*\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\)\}\{\\sum\_\{\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\}q\_\{\\phi\}^\{\*\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\)\}\\mathchar 24891\\allowbreak\\quad\\forall s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak a^\{0\}\\mathchar 24891\\allowbreak\\bm\{u\}\\mathchar 24891\\allowbreakfor the utilitarian\-reduced problem such that∑a~0,𝒖~qϕ∗​\(s0,𝒙,a~0,𝒖~\)\>0\\sum\_\{\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\}q\_\{\\phi\}^\{\*\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{x\}\\mathchar 24891\\allowbreak\\tilde\{a\}^\{0\}\\mathchar 24891\\allowbreak\\tilde\{\\bm\{u\}\}\)\>0, with the policy chosen arbitrarily on unreachable states\. LetΓ⁡\(s0,𝒔,𝒖\):=\{𝒂∈𝒜\(N\)​\(s0,𝒔\)∣g⁡\(𝒔,𝒂\)=𝒖\}\\Gamma\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{u\}\):=\\mathopen\{\\\{\}\\bm\{a\}\\in\\mathcal\{A\}^\{\(N\)\}\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)\\mid g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)=\\bm\{u\}\\mathclose\{\\\}\}denote the set of feasible labelled actions inducing the count action𝒖\\bm\{u\}\. The resulting count policy can be lifted to a permutation\-invariant policy in the original labelled M2WCMDP by assigning count actions among exchangeable participants through a permutation\-invariant disaggregation rule that distributes the probability mass equally

𝝅PI\(a0,𝒂∣s0,𝒔\)=𝝅ϕ∗\(a0,g\(𝒔,𝒂\)∣s0,f\(𝒔\)\)\|Γ⁡\(s0,𝒔,g⁡\(𝒔,𝒂\)\)\|,𝒂∈Γ\(s0,𝒔,g\(𝒔,𝒂\)\)\.\\bm\{\\pi\}\_\{\\mbox\{\\scriptsize\{PI\}\}\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)=\\frac\{\\bm\{\\pi\}\_\{\\phi\}^\{\*\}\(a^\{0\}\\mathchar 24891\\allowbreak g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\\mid s^\{0\}\\mathchar 24891\\allowbreak f\(\\bm\{s\}\)\)\}\{\|\\Gamma\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\)\|\}\\mathchar 24891\\allowbreak\\qquad\\bm\{a\}\\in\\Gamma\(s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\\mathchar 24891\\allowbreak g\(\\bm\{s\}\\mathchar 24891\\allowbreak\\bm\{a\}\)\)\.\(EC\.11\)
By construction,𝝅PI\\bm\{\\pi\}\_\{\\mbox\{\\scriptsize\{PI\}\}\}is permutation\-invariant, i\.e\.,𝝅PI\(a0,Q𝒂∣s0,Q𝒔\)=𝝅PI\(a0,𝒂∣s0,𝒔\)\\bm\{\\pi\}\_\{\\mbox\{\\scriptsize\{PI\}\}\}\(a^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak Q\\bm\{s\}\)=\\bm\{\\pi\}\_\{\\mbox\{\\scriptsize\{PI\}\}\}\(a^\{0\}\\mathchar 24891\\allowbreak\\bm\{a\}\\mid s^\{0\}\\mathchar 24891\\allowbreak\\bm\{s\}\)for any permutation matrixQQ\. This explicit exchangeable disaggregation prevents systematic discrimination during execution, addressing the inherent biases of arbitrary deterministic tie\-breaking rules\. By Theorem[1](https://arxiv.org/html/2609.36174#Thmtheorem1), the lifted policy𝝅PI\\bm\{\\pi\}\_\{\\mbox\{\\scriptsize\{PI\}\}\}is therefore optimal for the originalρ\\rho\-M2WCMDP for anyρ\\rhosatisfying Definition[1](https://arxiv.org/html/2609.36174#Thmdefinition1)under the transformationϕ\\phi\.

## Appendix EC\.7Additional Details and Experiments of Machine Replacement Problem

### EC\.7\.1Problem Instance Generation

This section details the construction of the components used to generate the test instances based on[Akbarzadeh and Mahajan \(2019\)](https://arxiv.org/html/2609.36174#bib.bib14), including the cost function, the transition matrix, and the reset probability\. This experiment uses a synthetic data generator implemented on our own that considers a system withSSstates and binary actions \(A=2A=2\), where the two possible actions are to operate or to replace\. After generating the per\-machine state\-action cost matrix of sizeN×S×AN\\times S\\times A, we normalize the costs to the range \[0, 1\] by dividing each entry by the maximum cost over all state\-action pairs\. This ensures that the discounted return always falls within the range\[0,11−γ\]\[0\\mathchar 24891\\allowbreak\\frac\{1\}\{1\-\\gamma\}\]\.

##### Cost Function

The cost functionc⁡\(s\)c\(s\)fors∈\[S\]s\\in\[S\]can be defined in four ways: 1\)Linear:c⁡\(s\)=s−1c\(s\)=s\-1, where the cost increases linearly with the state index; 2\)Quadratic:c⁡\(s\)=\(s−1\)2c\(s\)=\(s\-1\)^\{2\}with a more severe penalty for higher states compared to the linear case; 3\)Exponential:c⁡\(s\)=es−1c\(s\)=e^\{s\-1\}, which leads to exponentially increasing costs; 4\)Replacement Cost Constant Coefficient \(RCCC\):c⁡\(s\)=1\.5​\(S−1\)2c\(s\)=1\.5\(S\-1\)^\{2\}, which is based on a constant ratio of 1\.5 to the maximum quadratic cost\. These components define the state\-action costc⁡\(s,a\)c\(s\\mathchar 24891\\allowbreak a\)in Section[5\.1](https://arxiv.org/html/2609.36174#S5.SS1)\.

##### Transition Function

The transition matrix for the deterioration action is constructed as follows\. Once the machine reaches theSS\-th state, it remains in that worst state indefinitely until being successfully reset by a replacement action\. For thess\-th states∈\[S−1\]s\\in\[S\-1\], the probability of remaining in the same state at the next step is given by a model parameterpm∈\[0,1\]p\_\{m\}\\in\[0\\mathchar 24891\\allowbreak 1\], and the probability of transitioning to the\(s\+1\)\(s\+1\)\-th state is1−pm1\-p\_\{m\}\.

##### Reset Probability

When a replacement occurs, there is a probabilitypsp\_\{s\}that the machine successfully resets to the first state, and a corresponding probability1−ps1\-p\_\{s\}of failing to be repaired and following the deterioration rule\. In our experiments, we only consider a pure reset to the first state with probability 1\.

### EC\.7\.2Hyperparameters

In our experimental setup, we chose the Proximal Policy Optimization \(PPO\) algorithm to implement the count\-proportion\-based deep reinforcement learning \(CP\-DRL\) architecture\. The hidden layers are fully connected and the Tanh activation function is used\. There are two layers, with each layer consisting of 64 units\. The learning rate for the actor is set to5×10−45\\times 10^\{\-4\}and the critic is set to3×10−43\\times 10^\{\-4\}\. The Vanilla\-DRL baseline uses the same network architecture as PPO, with two hidden layers of 64 neurons each, but with a softmax output layer for its stochastic policy\.

### EC\.7\.3Additional Experiments

This subsection complements the small\-instance optimality comparisons in Section[5\.1](https://arxiv.org/html/2609.36174#S5.SS1)\.

##### Scalability

We assess CP\-DRL scalability by increasing the number of machines while keeping the resource proportion at 0\.1 for theExponential\-RCCCinstances\. We refer to this scaled \(SC\) extension as CP\-DRL\(SC\)\. We vary the number of machines from 10 to 100 to evaluate CP\-DRL performance as the problem size grows\. We also use CP\-DRL\(SC\), trained on 10 machines with 1 unit of resource, and scale it to tasks with 20 to 100 machines\. Figure[EC\.1](https://arxiv.org/html/2609.36174#A7.F1)a shows CP\-DRL and CP\-DRL\(SC\) consistently achieve higher GGF values than Whittle index policy \(WIP\) as machine numbers increase\. CP\-DRL\(SC\) delivers results comparable to separately trained CP\-DRL, reducing training time while maintaining similar performance\. Both WIP and CP\-DRL show linear growth in time consumption per episode as machine numbers scale up\.

Figure EC\.1:\(Color online\) Scalability and Time Efficiency of CP\-DRL![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/GGF-SC.png)\(a\)GGF values for the number of machinesN∈\[10,100\]N\\in\[10\\mathchar 24891\\allowbreak 100\]
![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/training_time_SC.png)\(b\)Time per episode in seconds with a resource ratiob/N=0\.1b/N=0\.1
![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/training_time.png)\(c\)Time per episode in seconds with a resource ratiob/N=0\.5b/N=0\.5
Note\.Subfigure \(a\) shows the scalability of CP\-DRL with a fixed resource ratio of 0\.1\. Subfigure \(a\) presents GGF values across different machine counts, with intervals representing the standard deviation over 5 runs\. Subfigure \(b\) and \(c\) depict time per episode in seconds for a fixed resource ratio of 0\.1 and 0\.5, respectively\. In all time plots, the green line represents WIP during MC evaluation, the blue line shows CP\-DRL during training, and the orange line represents CP\-DRL during MC evaluation\.

##### Efficiency

Using the count dual LP model \([EC\.10](https://arxiv.org/html/2609.36174#A6.E10)\) reduces the model size, but constraints still grow as\(N\+S−1S−1\)\\binom\{N\+S\-1\}\{S\-1\}and variables increase by\(N\+S−1S−1\)⋅A\\binom\{N\+S\-1\}\{S\-1\}\\cdot A\. These growth patterns create computational challenges as the problem size increases\. In addition to the time per episode for a fixed ratio of 0\.1 in Figure[EC\.1](https://arxiv.org/html/2609.36174#A7.F1)b, we analyze performance with a 0\.5 ratio \(Figure[EC\.1](https://arxiv.org/html/2609.36174#A7.F1)c\) and varying machine proportions, keeping the number of machines fixed at 10\. We evaluate CP\-DRL over machine proportions from 0\.1 to 0\.9\. The results show that the time per episode increases linearly with the number of machines, while training and evaluation times remain relatively stable\. This indicates that the sampling procedure for legal actions is the primary bottleneck\. Meanwhile, the resource ratio has minimal impact on computing times\.

## Appendix EC\.8Additional Details and Results of Taxi Dispatching and Pricing Problem

We report additional details in the numerical studies\. All experiments were conducted on the Rorqual cluster provided by the Digital Research Alliance of Canada, utilizing Python 3\.10\.13 and Gurobi 13\.0\.0\. To manage the extensive parameter testing, we utilize Slurm job arrays for parallel execution\. Each task was assigned 1 CPU core and 1\.0 GB of RAM\.

### EC\.8\.1Data Source, Preprocessing, and Calibration

We calibrate the fundamental vehicle speed based on[Huang et al\. \(2018\)](https://arxiv.org/html/2609.36174#bib.bib42)\. The average occupancy speed is approximately set at 20 km per hour\. Since that explicit empty\-cruising data is not provided, we assume that the empty speed is identical to the occupancy speed\. Consequently, in a single time stepΔ​t\\Delta t, a vehicle travels a distance of 1 km\. The time window is fixed for 10 hours per day, corresponding toT=T=200 time steps\.

The reward function approximates the New York City \(NYC\) taxi fare structure with a base fare and a distance\-based term\. We aggregate the fixed costs \(base fare $3\.00, metropolitan transportation authority surcharge $0\.50, and improvement surcharge $1\.00\) into a constant interceptcf:=$4\.5c\_\{f\}:=\\$4\.5\. Based on[Parrott \(2024\)](https://arxiv.org/html/2609.36174#bib.bib3), the variable reward proportional to the distance traveled is approximated ascr:=$2\.17c\_\{r\}:=\\$2\.17per km\. The operational cost includes fuel and maintenance costs, and calibrated atco:=$0\.54c\_\{o\}:=\\$0\.54per km\. Further, to model the impact of dynamic pricing on customer decisions, we adopt a log response functionψl​o​g​\(p\):=e−η⁡\(p−1\)\\psi\_\{log\}\(p\):=e^\{\-\\eta\(p\-1\)\}with an elasticity parameterη\\eta:= 0\.51 following empirical studies on ride\-sharing platforms\([Cohen et al\., 2016](https://arxiv.org/html/2609.36174#bib.bib41)\)\. The price range is set asp∈𝒫:=\(0,10\]p\\in\{\\mathcal\{P\}\}:=\(0\\mathchar 24891\\allowbreak 10\]\. We writepmin:=inf𝒫p\_\{\\min\}:=\\inf\{\\mathcal\{P\}\}andpmax:=sup𝒫p\_\{\\max\}:=\\sup\{\\mathcal\{P\}\}for the price bounds used in[EC\.8\.3](https://arxiv.org/html/2609.36174#A8.SS3)\.

Figure EC\.2:Smoothed Aggregate Trip Arrivals for the Four\-Zone Midtown Network![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/tue_arrival_rate_a.png)

Note\.The solid line represents the Gaussian\-smoothed average demand across Tuesdays from July to November 2025, while the dashed lines mark the 25th and 75th percentiles, and the dotted lines mark the 10th and 90th percentiles\. The horizontal axis is hour of day and each observation corresponds to a three\-minute resolution\.

##### Arrival Rate

For empirical validation, we utilize historical trip data from July to November 2025, and focus on Tuesday from 8:00 a\.m\. to 8:00 p\.m\. to capture the representative transition between peak and off\-peak demand patterns\. We model order arrivals as an inhomogeneous Poisson process with time\-varying rates𝜽tbase\\bm\{\\theta\}^\{\\text\{base\}\}\_\{t\}\. To estimate a stable empirical base\-rate profile rather than fitting high\-frequency sampling noise, we first compute mean arrivals across the selected Tuesdays and then apply a Gaussian smoothing filter\([Lehky, 2010](https://arxiv.org/html/2609.36174#bib.bib8)\)withσ:=2\.0\\sigma:=2\.0, approximately spanning a typical city traffic changing window of±15\\pm 15minutes to the raw count data\([Moyano et al\., 2021](https://arxiv.org/html/2609.36174#bib.bib7)\)\. This filter is an implementation choice for estimating the base demand𝜽^tbase\\hat\{\\bm\{\\theta\}\}^\{\\text\{base\}\}\_\{t\}, while reducing variance for preprocessing\. Figure[EC\.2](https://arxiv.org/html/2609.36174#A8.F2)illustrates the summary statistics for the total demand originating from the four\-zone Midtown network on Tuesday \(see zone details in[EC\.8\.1](https://arxiv.org/html/2609.36174#A8.SS1.SSS0.Px2)\)\. The plot reveals a steady increase in demand intensity throughout the day, peaking during the evening commute\.

##### Network Construction and Simulation Environment

We consider two subgraphs of increasing size, each induced from the same NYC zone graph as shown in Figure[EC\.3](https://arxiv.org/html/2609.36174#A8.F3)\. Sub\-figure[EC\.3](https://arxiv.org/html/2609.36174#A8.F3)a comprises 4 nodes covering the densest commercial regions in Midtown Manhattan for mechanism analysis of the joint pricing\-and\-routing decisions\. Sub\-figure[EC\.3](https://arxiv.org/html/2609.36174#A8.F3)b expands to 15 nodes spanning Midtown, Upper Manhattan, and Central Park, etc\., and is mainly used for scalability and real\-world performance evaluation\. Detailed region names are provided in Table[EC\.2](https://arxiv.org/html/2609.36174#A8.T2.fig1)\.

Figure EC\.3:Selected NYC Sub\-Networks![Refer to caption](https://arxiv.org/html/2609.36174v1/figs/selected.png)

Note\.Sub\-figure \(a\) is used for mechanism analysis, and \(b\) for scalability and NYC\-calibrated performance evaluation\. Node labels are taxi\-zone IDs\. The corresponding zone names are listed in Table[EC\.2](https://arxiv.org/html/2609.36174#A8.T2.fig1)\.

Table EC\.2:Node Selection and Corresponding Zone Names

### EC\.8\.2Graph Neural Network Architecture

We encode the spatial relationships among zones using a graph neural network \(GNN\) that operates on the augmented major states0=\(𝒒,𝒑^,t\)s^\{0\}=\(\{\\bm\{q\}\}\\mathchar 24891\\allowbreak\\hat\{\\bm\{p\}\}\\mathchar 24891\\allowbreak t\)and the count state𝒙\\bm\{x\}\. Here,𝒒∈ℝM×M\{\\bm\{q\}\}\\in\{\\mathbb\{R\}\}^\{M\\times M\}is the current origin\-destination \(OD\) request queue,𝒑^∈ℝM×M\\hat\{\\bm\{p\}\}\\in\{\\mathbb\{R\}\}^\{M\\times M\}is the lagged price\-multiplier matrix, and𝒙∈ℝM×\(Dmax\+1\)\\bm\{x\}\\in\{\\mathbb\{R\}\}^\{M\\times\(D\_\{\\max\}\+1\)\}is the zone\-level taxi\-count feature\.

For each edge\-valued input𝒆∈\{𝒒,𝒑^\}\\bm\{e\}\\in\\\{\{\\bm\{q\}\}\\mathchar 24891\\allowbreak\\hat\{\\bm\{p\}\}\\\}, we summarize the local edge structure at every zoneiiby its mean outgoing and incoming edge values over the adjacency matrix𝑨adj\\bm\{A\}\_\{\\text\{adj\}\}:

eout​\(i\):=1max⁡\(1,∑kAadj​\(i,k\)\)​∑j=1Me⁡\(i,j\)​Aadj​\(i,j\),ein​\(i\):=1max⁡\(1,∑kAadj​\(k,i\)\)​∑j=1Me⁡\(j,i\)​Aadj​\(j,i\)\.e\_\{\\operatorname\{out\}\}\(i\):=\\frac\{1\}\{\\max\(1\\mathchar 24891\\allowbreak\\sum\\limits\_\{k\}A\_\{\\text\{adj\}\}\(i\\mathchar 24891\\allowbreak k\)\)\}\\sum\_\{j=1\}^\{M\}e\(i\\mathchar 24891\\allowbreak j\)A\_\{\\text\{adj\}\}\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\quad e\_\{\\operatorname\{in\}\}\(i\):=\\frac\{1\}\{\\max\(1\\mathchar 24891\\allowbreak\\sum\\limits\_\{k\}A\_\{\\text\{adj\}\}\(k\\mathchar 24891\\allowbreak i\)\)\}\\sum\_\{j=1\}^\{M\}e\(j\\mathchar 24891\\allowbreak i\)A\_\{\\text\{adj\}\}\(j\\mathchar 24891\\allowbreak i\)\.

The pair\[ein​\(i\),eout​\(i\)\]∈ℝ2\[e\_\{\\operatorname\{in\}\}\(i\)\\mathchar 24891\\allowbreak e\_\{\\operatorname\{out\}\}\(i\)\]\\in\{\\mathbb\{R\}\}^\{2\}is concatenated to form a compact summary of the underlying edge feature\. Each channel is then independently mapped into a commondhd\_\{h\}\-dimensional node space via a learned linear projection𝑾\{\\bm\{W\}\}followed by a ReLU activation layer:

𝒉iq:=ReLU\(𝑾q\[𝒒in\(i\)∥𝒒out\(i\)\]\),𝒉ip:=ReLU\(𝑾p\[𝒑in\(i\)∥𝒑out\(i\)\]\),𝒉ix:=ReLU\(𝑾x𝒙\(i\)\),\{\\bm\{h\}\}^\{q\}\_\{i\}:=\\operatorname\{ReLU\}\\left\(\{\\bm\{W\}\}\_\{q\}\[\{\\bm\{q\}\}\_\{\\operatorname\{in\}\}\(i\)\\\|\{\\bm\{q\}\}\_\{\\operatorname\{out\}\}\(i\)\]\\right\)\\mathchar 24891\\allowbreak\\,\{\\bm\{h\}\}^\{p\}\_\{i\}:=\\operatorname\{ReLU\}\\left\(\{\\bm\{W\}\}\_\{p\}\[\{\\bm\{p\}\}\_\{\\operatorname\{in\}\}\(i\)\\\|\{\\bm\{p\}\}\_\{\\operatorname\{out\}\}\(i\)\]\\right\)\\mathchar 24891\\allowbreak\\,\{\\bm\{h\}\}^\{x\}\_\{i\}:=\\operatorname\{ReLU\}\\left\(\{\\bm\{W\}\}\_\{x\}\\bm\{x\}\(i\)\\right\)\\mathchar 24891\\allowbreakwhere𝑾q,𝑾p∈ℝdh×2\{\\bm\{W\}\}\_\{q\}\\mathchar 24891\\allowbreak\{\\bm\{W\}\}\_\{p\}\\in\{\\mathbb\{R\}\}^\{d\_\{h\}\\times 2\},𝑾x∈ℝdh×\(Dmax\+1\)\{\\bm\{W\}\}\_\{x\}\\in\{\\mathbb\{R\}\}^\{d\_\{h\}\\times\(D\_\{\\max\}\+1\)\}are learnable weight matrices\.

The three per\-node embeddings are concatenated and passed through a shared linear layer to produce a single fused node representation before any message passing occurs:

𝒉i\(0\):=ReLU⁡\(𝑾fuse​\[𝒉iq​‖𝒉ip‖​𝒉ix\]\),𝑾fuse∈ℝdh×3​dh\.\{\\bm\{h\}\}\_\{i\}^\{\(0\)\}:=\\operatorname\{ReLU\}\\left\(\{\\bm\{W\}\}\_\{\\operatorname\{fuse\}\}\[\{\\bm\{h\}\}^\{q\}\_\{i\}\\\|\{\\bm\{h\}\}^\{p\}\_\{i\}\\\|\{\\bm\{h\}\}^\{x\}\_\{i\}\]\\right\)\\mathchar 24891\\allowbreak\\quad\{\\bm\{W\}\}\_\{\\operatorname\{fuse\}\}\\in\{\\mathbb\{R\}\}^\{d\_\{h\}\\times 3d\_\{h\}\}\.
Fusing the latent features at the node level prior to propagation ensures that the graph convolution layers capture interactions among supply, demand, and pricing signals jointly, rather than aggregating spatially disjoint representations\. Then, the network appliesLLgraph convolution layers over the adjacency matrix𝑨adj\\bm\{A\}\_\{\\text\{adj\}\}\. At each layerℓ=0,…,L−1\\ell=0\\mathchar 24891\\allowbreak\\ldots\\mathchar 24891\\allowbreak L\-1, the embedding of zoneiiis updated by separately transforming the mean neighborhood embedding and the zone’s self\-embedding:

𝒉i\(ℓ\+1\):=ReLU⁡\(𝑾\(ℓ\)​\(1max⁡\(1,∑kAadj​\(i,k\)\)​∑j=1MAadj​\(i,j\)​𝒉j\(ℓ\)\)\+𝑩\(ℓ\)​𝒉i\(ℓ\)\),\{\\bm\{h\}\}\_\{i\}^\{\(\\ell\+1\)\}:=\\operatorname\{ReLU\}\\left\(\{\\bm\{W\}\}^\{\(\\ell\)\}\{\\left\(\\frac\{1\}\{\\max\(1\\mathchar 24891\\allowbreak\\sum\\limits\_\{k\}A\_\{\\text\{adj\}\}\(i\\mathchar 24891\\allowbreak k\)\)\}\\sum\_\{j=1\}^\{M\}A\_\{\\text\{adj\}\}\(i\\mathchar 24891\\allowbreak j\)\{\\bm\{h\}\}\_\{j\}^\{\(\\ell\)\}\\right\)\}\+\\bm\{B\}^\{\(\\ell\)\}\{\\bm\{h\}\}\_\{i\}^\{\(\\ell\)\}\\right\)\\mathchar 24891\\allowbreak\(EC\.12\)where𝑾\(ℓ\),𝑩\(ℓ\)∈ℝdh×dh\{\\bm\{W\}\}^\{\(\\ell\)\}\\mathchar 24891\\allowbreak\\bm\{B\}^\{\(\\ell\)\}\\in\{\\mathbb\{R\}\}^\{d\_\{h\}\\times d\_\{h\}\}are layer\-specific learnable parameters\. The mean aggregation in Equation[EC\.12](https://arxiv.org/html/2609.36174#A8.E12)normalizes by the degree of zoneii, making the update invariant to network size\. The temporal contextttis encoded through a Time2Vec representation𝒆t∈ℝdt\\bm\{e\}\_\{t\}\\in\{\\mathbb\{R\}\}^\{d\_\{t\}\}\([Kazemi et al\., 2019](https://arxiv.org/html/2609.36174#bib.bib35)\)with both sinusoidal and cosinusoidal embeddings, and appended to the output only after spatial propagation is complete\. The final feature vector passed to the downstream policy and value heads is𝒛:=\[𝒉\(L\)∥𝒆t\]∈ℝM​dh\+dt\\bm\{z\}:=\[\{\\bm\{h\}\}^\{\(L\)\}\\\|\\bm\{e\}\_\{t\}\]\\in\{\\mathbb\{R\}\}^\{Md\_\{h\}\+d\_\{t\}\}, where𝒉\(L\):=\[𝒉1\(L\),…,𝒉M\(L\)\]⊤∈ℝM×dh\{\\bm\{h\}\}^\{\(L\)\}:=\[\{\\bm\{h\}\}\_\{1\}^\{\(L\)\}\\mathchar 24891\\allowbreak\\ldots\\mathchar 24891\\allowbreak\{\\bm\{h\}\}\_\{M\}^\{\(L\)\}\]^\{\\top\}\\in\{\\mathbb\{R\}\}^\{M\\times d\_\{h\}\}collects the final node embeddings across all zones\. Putting temporal context at the output rather than at intermediate layers prevents the time signal from interfering with the spatial message\-passing dynamics, while still allowing downstream heads to condition their decisions on the current period\.

### EC\.8\.3Hyperparameters and Reproducibility Protocol

In this section, we report the training configuration required to reproduce the taxi experiments in Section[5\.2](https://arxiv.org/html/2609.36174#S5.SS2)\. Table[EC\.3](https://arxiv.org/html/2609.36174#A8.T3.fig1)summarizes the hyperparameters used for the GNN feature extractor and PPO\.

Table EC\.3:PPO and Simulation Settings for the Taxi Experiments#### EC\.8\.3\.1Normalization

The observation contains the taxi counts𝒙\\bm\{x\}, order queue𝒒\{\\bm\{q\}\}, and normalized timet/Tt/T\. Taxi counts are divided by the fleet sizeNN, and the order queue is divided bymax⁡\{∑i,j∈ℳq⁡\(i,j\),1\}\\max\\\{\\sum\_\{i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\}q\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak 1\\\}\. The per ride reward is divided bymax⁡\{co​Dmax,\|cf\+cr​Dmax​pmax−co​Dmax\|,1\}\\max\\\{c\_\{o\}D\_\{\\max\}\\mathchar 24891\\allowbreak\\,\|c\_\{f\}\+c\_\{r\}D\_\{\\max\}p\_\{\\max\}\-c\_\{o\}D\_\{\\max\}\|\\mathchar 24891\\allowbreak\\,1\\\}to approximately fall between the range\[−1,1\]\[\-1\\mathchar 24891\\allowbreak 1\]without changing its sign\.

#### EC\.8\.3\.2Structured Low\-Rank Pricing

For OD pricing, the policy head outputs origin and destination latent factors𝑳O,𝑳D∈\[0,1\]M×rp\{\\bm\{L\}\}^\{\\text\{O\}\}\\mathchar 24891\\allowbreak\{\\bm\{L\}\}^\{\\text\{D\}\}\\in\[0\\mathchar 24891\\allowbreak 1\]^\{M\\times r\_\{p\}\}, whererpr\_\{p\}is the chosen rank\. Define their row averages byℓO​\(i\):=rp−1​∑q=1rpLO​\(i,q\)\\ell^\{\\text\{O\}\}\(i\):=r\_\{p\}^\{\-1\}\\sum\_\{q=1\}^\{r\_\{p\}\}L^\{\\text\{O\}\}\(i\\mathchar 24891\\allowbreak q\)andℓD​\(j\):=rp−1​∑q=1rpLD​\(j,q\)\\ell^\{\\text\{D\}\}\(j\):=r\_\{p\}^\{\-1\}\\sum\_\{q=1\}^\{r\_\{p\}\}L^\{\\text\{D\}\}\(j\\mathchar 24891\\allowbreak q\)\. The OD price multiplier ispOD​\(i,j\):=pmin\+\(1/2\)​\(pmax−pmin\)​\(ℓO​\(i\)\+ℓD​\(j\)\)p^\{\\text\{OD\}\}\(i\\mathchar 24891\\allowbreak j\):=p\_\{\\min\}\+\{\(1/2\)\}\(p\_\{\\max\}\-p\_\{\\min\}\)\\mathopen\{\(\}\\ell^\{\\text\{O\}\}\(i\)\+\\ell^\{\\text\{D\}\}\(j\)\\mathclose\{\)\}, witha0:=𝒑ODa^\{0\}:=\{\\bm\{p\}\}^\{\\text\{OD\}\}, i\.e\., the normalized OD score matrix is additively separable into origin and destination effects\. Origin\-only pricing usespO​\(i\)=pmin\+\(pmax−pmin\)​ℓO​\(i\)p^\{\\text\{O\}\}\(i\)=p\_\{\\min\}\+\(p\_\{\\max\}\-p\_\{\\min\}\)\\ell^\{\\text\{O\}\}\(i\), witha0:=𝒑O​𝟏MTa^\{0\}:=\{\\bm\{p\}\}^\{\\text\{O\}\}\{\\bm\{1\}\}\_\{M\}^\{T\}\. Uniform pricing takes the scalarsℓO,ℓD\\ell^\{\\text\{O\}\}\\mathchar 24891\\allowbreak\\ell^\{\\text\{D\}\}and averages the scorespU=pmin\+\(pmax−pmin\)​\(ℓO\+ℓD\)/2p^\{\\text\{U\}\}=p\_\{\\min\}\+\(p\_\{\\max\}\-p\_\{\\min\}\)\(\\ell^\{\\text\{O\}\}\+\\ell^\{\\text\{D\}\}\)/2to obtain a single multiplier, witha0:=pU​𝟏M×Ma^\{0\}:=p^\{\\text\{U\}\}\{\\bm\{1\}\}\_\{M\\times M\}\. We setrp=4r\_\{p\}=4for the four\-zone experiments andrp=10r\_\{p\}=10for the 15\-zone experiments\.

### EC\.8\.4Benchmark Details

The benchmarks use the same simulator dynamics, economic parameters, episode horizon, and count\-action feasibility checks as the CP\-DRL\.

##### Count\-Proportional Assignment Linear Program \(CP\-ALP\)

We solve a myopic taxi\-assignment model per\-step\. Notations match the taxi count model in Example[6](https://arxiv.org/html/2609.36174#Thmexample6), with the additional variableν⁡\(i,j\)∈ℤ\+\{\\nu\}\(i\\mathchar 24891\\allowbreak j\)\\in\\mathbb\{Z\}\_\{\+\}representing the number of taxis that are relocated toii, which are still in\-transit but assumed to immediately serve requests tojj\.

We first consider the case \(i\), in which there are more taxis than orders,

max𝒖1,𝒖2,𝝂≥𝟎∑i∈ℳ∑j∈ℳ\(\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]​u2​\(i,j\)−co​D​\(i,j\)​\[u1​\(i,j\)\+u2​\(i,j\)\]\)s\.t\.∑j∈ℳu1\(i,j\)\+∑j∈ℳu2\(i,j\)=x\(i,0\),∀i∈ℳ,∑k∈ℳu1\(k,i\)≥∑j∈ℳν\(i,j\),∀i∈ℳ,ν⁡\(i,j\)\+u2​\(i,j\)≥q⁡\(i,j\),∀i,j∈ℳ,u2​\(i,j\)≤q⁡\(i,j\),∀i,j∈ℳ\.\\begin\{array\}\[\]\{rl\}\\max\\limits\_\{\\bm\{u\}^\{1\}\\mathchar 24891\\allowbreak\\bm\{u\}^\{2\}\\mathchar 24891\\allowbreak\\bm\{\\nu\}\\geq\\bm\{0\}\}&\\sum\\limits\_\{i\\in\{\\mathcal\{M\}\}\}\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}\\left\(\\left\[c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\right\]u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\-c\_\{o\}D\(i\\mathchar 24891\\allowbreak j\)\\left\[u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\right\]\\right\)\\\\ \\text\{s\.t\.\}&\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}u^\{2\}\(i\\mathchar 24891\\allowbreak j\)=x\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak\\hskip 30\.0pt\\forall i\\in\{\\mathcal\{M\}\}\\mathchar 24891\\\\ &\\sum\\limits\_\{k\\in\{\\mathcal\{M\}\}\}u^\{1\}\(k\\mathchar 24891\\allowbreak i\)\\geq\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}\{\\nu\}\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\hskip 76\.0pt\\forall i\\in\{\\mathcal\{M\}\}\\mathchar 24891\\\\ &\{\\nu\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\geq q\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\hskip 78\.0pt\\forall i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\\mathchar 24891\\\\ &u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\leq q\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\hskip 120\.0pt\\forall i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\.\\end\{array\}\(EC\.13\)
The first constraint is the flow conservation of idle taxis at each origin that each available taxi is either relocated or directly assigned\. The second represents that the total number of orders immediately fulfilled atiicannot exceed the number of taxis relocated intoii\. The third guarantees that all customer requests are satisfied, either by direct assignment𝒖2\\bm\{u\}^\{2\}or by post\-relocated taxis𝝂\\bm\{\\nu\}\. The fourth constrains the order capacity\.

We next consider case \(ii\), in which the number of taxis is smaller than the number of orders:

max𝒖1,𝒖2,𝝂≥𝟎∑i∈ℳ∑j∈ℳ\(\[cf\+cr​p^​\(i,j\)​D​\(i,j\)\]​u2​\(i,j\)−co​D​\(i,j\)​\[u1​\(i,j\)\+u2​\(i,j\)\]\)s\.t\.∑j∈ℳu1\(i,j\)\+∑j∈ℳu2\(i,j\)=x\(i,0\),∀i∈ℳ,∑k∈ℳu1\(k,i\)≤∑j∈ℳν\(i,j\),∀i∈ℳ,ν⁡\(i,j\)\+u2​\(i,j\)≤q⁡\(i,j\),∀i,j∈ℳ\.\\begin\{array\}\[\]\{rl\}\\max\\limits\_\{\\bm\{u\}^\{1\}\\mathchar 24891\\allowbreak\\bm\{u\}^\{2\}\\mathchar 24891\\allowbreak\\bm\{\\nu\}\\geq\\bm\{0\}\}&\\sum\\limits\_\{i\\in\{\\mathcal\{M\}\}\}\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}\\left\(\\left\[c\_\{f\}\+c\_\{r\}\\hat\{p\}\(i\\mathchar 24891\\allowbreak j\)D\(i\\mathchar 24891\\allowbreak j\)\\right\]u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\-c\_\{o\}D\(i\\mathchar 24891\\allowbreak j\)\\left\[u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\right\]\\right\)\\\\ \\text\{s\.t\.\}&\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}u^\{1\}\(i\\mathchar 24891\\allowbreak j\)\+\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}u^\{2\}\(i\\mathchar 24891\\allowbreak j\)=x\(i\\mathchar 24891\\allowbreak 0\)\\mathchar 24891\\allowbreak\\hskip 30\.0pt\\forall i\\in\{\\mathcal\{M\}\}\\mathchar 24891\\\\ &\\sum\\limits\_\{k\\in\{\\mathcal\{M\}\}\}u^\{1\}\(k\\mathchar 24891\\allowbreak i\)\\leq\\sum\\limits\_\{j\\in\{\\mathcal\{M\}\}\}\{\\nu\}\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\hskip 76\.0pt\\forall i\\in\{\\mathcal\{M\}\}\\mathchar 24891\\\\ &\{\\nu\}\(i\\mathchar 24891\\allowbreak j\)\+u^\{2\}\(i\\mathchar 24891\\allowbreak j\)\\leq q\(i\\mathchar 24891\\allowbreak j\)\\mathchar 24891\\allowbreak\\hskip 78\.0pt\\forall i\\mathchar 24891\\allowbreak j\\in\{\\mathcal\{M\}\}\.\\end\{array\}\(EC\.14\)

##### Random \(RDM\)

At each step, the entries of the priority\-score matrix are sampled independently from a uniform distribution\[0,1\]\[0\\mathchar 24891\\allowbreak 1\]\.

##### Lowest Income First \(LIF\)

All idle taxis are globally sorted by accumulated discounted reward, with taxi index used only for deterministic tie breaking\. Each taxi is then considered based on the sorted order\. If there are outstanding requests at the current origin, the chosen taxi is matched to the destination maximizing current trip net reward\. If no request is available at the current origin, the taxi remains in place\. LIF neither relocates taxis proactively nor optimizes price\.

##### Multilayer Perceptron \(MLP\)

The configuration is kept the same as CP\-DRL, but replaces graph message passing by 2\-layer fully connected feature processing with a hidden dimension of 128\.

### EC\.8\.5Additional Results on Fairness Preference Factor

Table[EC\.4](https://arxiv.org/html/2609.36174#A8.T4.fig1)reports a robustness check with values of platform components and percentages to the objectives by GGF, complementing the fairness\-aware taxi evaluations in Table[6](https://arxiv.org/html/2609.36174#S5.T6.fig1)\. For OD pricing, the CP\-DRL advantage over CP\-ALP ranges from approximately 37\.6% at 10 taxis to 54\.5% at 50 taxis, showing that the dynamic advantage is not driven only by the fairness term\. Across CP\-DRL variants, OD pricing is generally strongest\. This check indicates that CP\-DRL’s gains arise from intertemporal coordination of pricing and dispatch rather than from the particular welfare weight alone\.

Table EC\.4:Platform Value and the Percentage to the Fairness\-Aware ObjectiveNote\.Entries are mean±\\pmsample standard deviation over available experiments\. Values are recomputed asV00V^\{0\}\_\{0\}and percentages asV00/\(V00\+1⋅ρ⁡\(𝑽0\)\)V^\{0\}\_\{0\}/\(V^\{0\}\_\{0\}\+1\\cdot\\rho\(\\bm\{V\}\_\{0\}\)\)to Table[6](https://arxiv.org/html/2609.36174#S5.T6.fig1)\.

相似文章

面向多目标强化学习的确定性帕累托最优策略综合

arXiv cs.LG

本文引入了一种基于切比雪夫标量化的新颖偏好条件贝尔曼算子,用于计算多目标马尔可夫决策过程中的确定性帕累托最优策略,并证明了该算子的收敛性及其在捕获完整帕累托前沿方面的有效性。

效用约束策略优化

arXiv cs.LG

本文介绍了一种简单而强大的方法,用于效用约束马尔可夫决策过程(UCMDPs),该方法无需预先固定约束界限即可实现风险敏感约束,在Safety Gymnasium基准测试中优于基线方法。

优化训练策略的幻象:单调推理策略作为LLM强化学习的真正目标

Hugging Face Daily Papers

我们介绍了MIPI(单调推理策略改进)及其实例化MIPU,这是一个用于LLM的两步RL框架,通过将优化与推理策略改进明确对齐来解决训练-推理不匹配问题。在FP8量化展开下,MIPU在Qwen3-1.7B和Qwen3-4B模型上实现了改进的推理性能和训练稳定性。

公平强化学习

Reddit r/AI_Agents

公平强化学习引入了民主对齐,以整合来自不同代理的多个竞争性价值集,克服了传统RLHF的局限性,并通过黑盒策略包装器实现了数量级更快的优化。