Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models
Summary
This paper proposes an entropy-regularized reinforcement learning approach to solve linear-quadratic Stackelberg differential games in regime-switching diffusion models, integrating neural networks to approximate value functions and escape suboptimal equilibria.
View Cached Full Text
Cached at: 06/30/26, 05:29 AM
# Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models
Source: [https://arxiv.org/html/2606.28671](https://arxiv.org/html/2606.28671)
Congde HuThis study was funded by the National Key R & D Program of China \(2022YFA1007900\), Natural Science Foundation of Anhui Province \(2408085MA019\), National Natural Science Foundation of China \(12271171\), Shanghai Philosophy Social Science Planning Office Project \(2022ZJB005\), and Fundamental Research Funds for the Central Universities \(2024QKT008\)\. \(Corresponding author: Lin Xu\.\)Congde Hu is with the School of Mathematics and Statistics, Anhui Normal University, Wuhu, Anhui 241002, China \(e\-mail:hcd\.math@gmail\.com\)\.Danping LiDanping Li is with the School of Statistics, Key Laboratory of Advanced Theory and Application in Statistics and Data Science\-MOE, Faculty of Economics and Management, East China Normal University, Shanghai, China \(e\-mail:dpli@fem\.ecnu\.edu\.cn\)\.Lin XuLin Xu is with the School of Mathematics and Statistics, Anhui Normal University, Wuhu, Anhui 241002, China \(e\-mail:xulinahnu@gmail\.com\)\.Wenying XuWenying Xu is with the School of Mathematics, Southeast University, Nanjing, Jiangsu 211189, China \(e\-mail: wyxu@seu\.edu\.cn \)\.
###### Abstract
Stackelberg differential games \(SDGs\) provide a powerful framework for hierarchical decision\-making in stochastic and continuous\-time environments, yet their solution remains computationally challenging due to the complexity of traditional dynamic programming and Hamilton\-Jacobi\-Bellman\-Isaacs \(HJBI\) methods, especially in high\-dimensional systems\. This paper proposes an entropy\-regularized reinforcement learning \(ERRL\) approach for linear\-quadratic SDGs \(LQ\-SDGs\) within a continuous\-time diffusion framework governed by Markovian regime switching\. The key innovation lies in deriving exploratory weakly\-coupled HJBI equations with entropy regularization, which promotes stochastic policies that actively avoid suboptimal equilibria—a limitation of classical SDG methods\. Neural networks are integrated to approximate regime\-dependent value functions and solve high\-dimensional partial differential equations \(PDEs\) efficiently, while a novel sampling technique enhances computational tractability\. Numerical results demonstrate the effectiveness of the framework compared to conventional approaches, particularly in escaping suboptimal traps through exploratory policies\. The study highlights the critical role of entropy regularization and neural network approximations in achieving robust solutions for hierarchical decision\-making problems under abrupt environmental shifts\.
††publicationid:pubid: 0000–0000/00$00\.00 © 2025 IEEE## IIntroduction
### I\-ABackground and motivation
Stackelberg differential games \(SDGs\) represent a cornerstone in the field of game theory, offering a robust framework for modeling hierarchical decision\-making processes\. The concept of Stackelberg games was first introduced by\[[23](https://arxiv.org/html/2606.28671#bib.bib1)\], which sought to describe economic scenarios where one player \(the leader\) has a dominant role over another \(the follower\)\. This hierarchical structure has since been extended to continuous\-time and stochastic settings, giving rise to SDGs, which have found applications in a wide range of disciplines, including economics\[[1](https://arxiv.org/html/2606.28671#bib.bib2)\], engineering\[[32](https://arxiv.org/html/2606.28671#bib.bib3)\], robotics\[[28](https://arxiv.org/html/2606.28671#bib.bib7)\], supply chain optimization\[[18](https://arxiv.org/html/2606.28671#bib.bib8)\], autonomous vehicles\[[8](https://arxiv.org/html/2606.28671#bib.bib4)\], financial markets\[[19](https://arxiv.org/html/2606.28671#bib.bib6)\]\.
Despite their versatility, solving SDGs remains a challenging task due to the complexity of deriving optimal strategies in continuous\-time and stochastic environments\. Traditional methods often rely on dynamic programming principles \(DPP\) and the Hamilton\-Jacobi\-Bellman\-Isaacs \(HJBI\) equations\[[29](https://arxiv.org/html/2606.28671#bib.bib9)\]\. However, these approaches can become computationally intractable in high\-dimensional or nonlinear systems, limiting their applicability to real\-world problems\.
Reinforcement learning \(RL\) has emerged as a powerful paradigm for solving complex decision\-making problems, particularly in scenarios where traditional methods fall short\. Unlike classical optimization techniques, RL leverages data\-driven approaches to learn optimal policies through interaction with the environment\. This makes RL particularly well\-suited for problems involving uncertainty, incomplete information, and dynamic environments\[[21](https://arxiv.org/html/2606.28671#bib.bib11),[6](https://arxiv.org/html/2606.28671#bib.bib19)\]\. entropy\-regularized reinforcement learning \(ERRL\)—a variant of RL that incorporates an entropy term into the objective function, has gained significant attention in recent years\. The entropy term encourages exploration by promoting stochasticity in the policy, thereby*preventing premature convergence to suboptimal solutions*\. This approach has been shown to improve the robustness and stability of learning algorithms, making it particularly effective in high\-dimensional and continuous action spaces\[[14](https://arxiv.org/html/2606.28671#bib.bib18)\]\. The integration of entropy regularization into RL has led to the development of several state\-of\-the\-art algorithms, such as soft actor\-critic \(SAC\)\[[5](https://arxiv.org/html/2606.28671#bib.bib12)\]and maximum entropy deep RL\[[33](https://arxiv.org/html/2606.28671#bib.bib31)\]\. These algorithms have demonstrated superior performance in a wide range of applications, including robotics, game theory, and control systems\. For instance,\[[15](https://arxiv.org/html/2606.28671#bib.bib30)\]demonstrated the effectiveness of ERRL in multi\-agent systems, where the stochasticity of policies facilitated better coordination among agents\.
In the context of SDGs, ERRL offers several advantages\. First, the stochastic nature of entropy\-regularized policies aligns well with the hierarchical structure of SDGs, enabling the leader and follower to adapt their strategies dynamically\. Second, the exploration\-promoting properties of entropy regularization enhance the robustness of learning algorithms, particularly in stochastic environments\. Finally, the probabilistic interpretation of entropy\-regularized policies facilitates derivation of smooth and interpretable strategies, which are crucial for real\-world applications\[[33](https://arxiv.org/html/2606.28671#bib.bib31)\]\.
Despite these advancements, the integration of entropy regularization into RL for SDGs remains underexplored, particularly in the context of continuous\-time diffusion models\. Diffusion models provide a natural framework for modeling stochastic dynamics in continuous time and have recently gained attention in RL due to their ability to handle uncertainty and noise effectively\[[24](https://arxiv.org/html/2606.28671#bib.bib37),[22](https://arxiv.org/html/2606.28671#bib.bib38)\]\. Furthermore, many real\-world systems are not only subject to continuous random fluctuations but also experience abrupt structural changes due to macroeconomic shifts, environmental variations, or component failures\. Markovian regime switching models provide a powerful mathematical tool to capture such sudden transitions\. Integrating regime switching into the ERRL framework for SDGs is therefore practically significant but mathematically demanding, as it introduces weakly\-coupled systems of equations\.
### I\-BMain contributions of this paper
This paper addresses the aforementioned challenges by proposing a novel framework for solving linear\-quadratic SDGs \(LQ\-SDGs\) by using ERRL within a continuous\-time diffusion framework featuring Markovian regime switching\. LQ\-SDGs have evolved since their inception, bridging deterministic and stochastic frameworks with applications in economics, engineering, and finance\. Early works by\[[30](https://arxiv.org/html/2606.28671#bib.bib49)\]established deterministic LQ\-SDGs, while stochastic extensions incorporated Brownian motion, yielding Riccati\-based solutions under adapted strategies\. Recent advances address jump\-diffusion systems, where\[[13](https://arxiv.org/html/2606.28671#bib.bib27)\]investigated LQ\-SDGs for mean\-field switching diffusions,\[[20](https://arxiv.org/html/2606.28671#bib.bib28)\]studied zero\-sum LQ\-SDGs,\[[16](https://arxiv.org/html/2606.28671#bib.bib26)\]derived equilibrium strategies via coupled Riccati equations with Lévy processes, generalizing results to non\-Gaussian noise,\[[12](https://arxiv.org/html/2606.28671#bib.bib29)\]provided closed\-loop solvability of mean\-field LQ\-SDGs\. Contemporary research focuses on mean\-field LQ\-SDGs\[[31](https://arxiv.org/html/2606.28671#bib.bib17),[2](https://arxiv.org/html/2606.28671#bib.bib33)\]and time\-inconsistent problems\[[25](https://arxiv.org/html/2606.28671#bib.bib34)\], leveraging backward stochastic differential equations \(BSDEs\) and leader\-follower asymmetry\.
The work closely related to the problem discussed in this paper is\[[10](https://arxiv.org/html/2606.28671#bib.bib35)\], which established two coupled forms of the HJBI equations and proved that the solutions of these HJBI equations not only stabilize the system but also constitute the Stackelberg equilibrium strategies\. Due to the difficulties in solving the HJBI equations, they also developed an iterative algorithm and conducted simulation examples, which can be found in detail in\[[10](https://arxiv.org/html/2606.28671#bib.bib35)\]\. However, on the one hand, the controlled system therein does not take into account the impact of continuous random disturbances and sudden environmental shifts \(i\.e\., regime switching\)\. On the other hand, finding the optimal strategies often leads to the so\-called “curse of optimality\.” As pointed out in\[[35](https://arxiv.org/html/2606.28671#bib.bib36)\], we strive to seek optimality, but often find ourselves trapped in bad “optimal” solutions that are either local optimizers, or too rigid to leave any room for errors, or based on wrong models\. A way to break this “curse of optimality” is to engage exploration through randomization\. In this paper, we will present a dynamic system equation that takes into account additional Brownian motion and Markovian regime switching, and has an exploration mechanism, as well as performance functions that encourage exploration\.
The main contributions of this work are summarized as follows:
- •An exploratory system of weakly\-coupled HJBI equations for SDGs with Markovian regime switching is derived by applying the DPP\. This derivation establishes a rigorous theoretical foundation for the proposed approach while incorporating an entropy regularization term, ensuring that the resulting policies remain stochastic and exploration\-promoting across different system regimes\.
- •Coupled distributional optimal policies for both the leader and follower are derived, ensuring robustness and adaptability in stochastic environments subject to structural jumps\. A partially model\-free policy improvement algorithm \(PIA\) is designed to approximate the regime\-dependent value functions within an ERRL framework, enhancing exploration and convergence\.
- •The algorithm integrates a novel sampling technique that exploits the structure of the diffusion model, leading to improved computational efficiency\. For fast approximating to the value functions across multiple discrete regimes, a neural network approximation architecture for high\-dimensional partial differential equations \(PDEs\) is incorporated into the procedure design\.
The remainder of this paper is organized as follows\. Section[II](https://arxiv.org/html/2606.28671#S2)introduces the learning framework, including regime\-switching dynamics, and formulates the problem\. Section[III](https://arxiv.org/html/2606.28671#S3)derives the coupled HJBI equations for the ERRL LQ\-SDGs and characterizes the equilibrium strategies\. Section[IV](https://arxiv.org/html/2606.28671#S4)develops a PIA to approximate the value function and compute these strategies\. Finally, Section[V](https://arxiv.org/html/2606.28671#S5)provides numerical examples to demonstrate algorithm convergence, the effect of temperature parameter, and specifically how the framework avoids the “local optimum trap”\.
## Notations
ℝn\\mathbb\{R\}^\{n\}and𝕊n\\mathbb\{S\}^\{n\}denotenn\-dimensional vectors and symmetric matrices, respectively\. For a matrixAA, its transpose, inverse, and trace areATA^\{T\},A−1A^\{\-1\}, andtr\(A\)tr\(A\)\.A\>0A\>0\(A≥0A\\geq 0\) meansAAis positive \(semi\-\)definite, with square rootA1/2=UD1/2UTA^\{1/2\}=UD^\{1/2\}U^\{T\}\.\|⋅\|\|\\cdot\|and‖B‖\|\|B\|\|are the absolute value and Euclidean norm\.δ\\deltais the variational operator, and𝔼x,i\[⋅\]\\mathbb\{E\}\_\{x,i\}\[\\cdot\]is the conditional expectation given statexxand regimeii\. For a functionf\(x\)f\(x\),∇f\(x\)\\nabla f\(x\)andΔf\(x\)\\Delta f\(x\)are its gradient and Hessian\. Finally,a∧ba\\land bis the minimum ofaaandbb, and𝒩\(μ,Σ\)\\mathcal\{N\}\(\\mu,\\Sigma\)represents the Gaussian distribution\.
## IIFormulation and problem
This section aims to construct the two\-player LQ\-SDGs problem with learning\. To clarify our objectives, we first introduce a classical two\-player Stackelberg game problem\. The content in subsection[II\-A](https://arxiv.org/html/2606.28671#S2.SS1)mainly comes from\[[10](https://arxiv.org/html/2606.28671#bib.bib35)\]\. For the convenience of narration, the problem studied by\[[10](https://arxiv.org/html/2606.28671#bib.bib35)\]is referred to as the classical two\-player Stackelberg game problem\.
### II\-AClassical two\-player Stackelberg game problem
Consider a linear dynamic system with two players
dxs=\(Axs\+B1us\+B2vs\)ds,s\>0dx\_\{s\}=\\left\(Ax\_\{s\}\+B\_\{1\}u\_\{s\}\+B\_\{2\}v\_\{s\}\\right\)ds,\\,s\>0\(2\.1\)wherex=\{xs,s\>0\}∈ℝnx=\\\{x\_\{s\},s\>0\\\}\\in\\mathbb\{R\}^\{n\}is the measurable system state,u=\{us,s\>0\}u=\\\{u\_\{s\},s\>0\\\}\(resp\.,v=\{vs,s\>0\}\)\(resp\.,v=\\\{v\_\{s\},s\>0\\\}\)∈ℝp\\in\\mathbb\{R\}^\{p\}is the policy of playerII\(resp\., playerIIII\),B1B\_\{1\}andB2B\_\{2\}are matrices with compatible dimensions\. PlayerIIis the leader, holding a dominant position that enables it to anticipate playerIIII’s response and make its decisionuufirst\. Next, playerIIII\(the follower\) observesuuand responds with actionvv\. Given this decision\-making structure, the performance function for each player is defined as
Jkcl\(x,u,v\)=∫0∞rk\(xs,us,vs\)𝑑s,k=1,2,J^\{cl\}\_\{k\}\(x,u,v\)=\\int\_\{0\}^\{\\infty\}r\_\{k\}\(x\_\{s\},u\_\{s\},v\_\{s\}\)\\,ds,\\,k=1,2,\(2\.2\)with
r1\(x,u,v\)=xTQ1x\+\(u\+θ1v\)TR1\(u\+θ1v\),r\_\{1\}\(x,u,v\)=x^\{T\}Q\_\{1\}x\+\(u\+\\theta\_\{1\}v\)^\{T\}R\_\{1\}\(u\+\\theta\_\{1\}v\),r2\(x,u,v\)=xTQ2x\+\(v\+θ2u\)TR2\(v\+θ2u\),r\_\{2\}\(x,u,v\)=x^\{T\}Q\_\{2\}x\+\(v\+\\theta\_\{2\}u\)^\{T\}R\_\{2\}\(v\+\\theta\_\{2\}u\),andQk≥0,Rk\>0Q\_\{k\}\\geq 0,R\_\{k\}\>0, andθk∈\(0,1\)\\theta\_\{k\}\\in\(0,1\)fork=1,2k=1,2\.
Denote the admissible sets of the two players by𝒰cl\\mathcal\{U\}\_\{cl\}and𝒱cl\\mathcal\{V\}\_\{cl\}respectively with
1. \(i\)∫0t‖us‖2𝑑s<∞\\int\_\{0\}^\{t\}\|\|u\_\{s\}\|\|^\{2\}ds<\\infty,∫0t‖vs‖2𝑑s<∞\\int\_\{0\}^\{t\}\|\|v\_\{s\}\|\|^\{2\}ds<\\infty,∀t≥0\\forall t\\geq 0;
2. \(ii\)with\{xs,s≥0\}\\\{x\_\{s\},s\\geq 0\\\}solving \([2\.1](https://arxiv.org/html/2606.28671#S2.E1)\),∫0∞\|rk\(xs,us,vs\)\|𝑑s<∞\\int\_\{0\}^\{\\infty\}\|r\_\{k\}\(x\_\{s\},u\_\{s\},v\_\{s\}\)\|ds<\\infty,k=1,2k=1,2\.
###### Definition 1
The classical value functions of playerIIIIand playerIIare given by
V2cl\(x;u\)=minv∈𝒱clJ2cl\(x,u,v\)V^\{cl\}\_\{2\}\(x;u\)=\\min\_\{v\\in\\mathcal\{V\}\_\{cl\}\}J^\{cl\}\_\{2\}\(x,u,v\)\(2\.3\)and
V1cl\(x\)=minu∈𝒰clJ1cl\(x,u,v∗\(u\)\),V^\{cl\}\_\{1\}\(x\)=\\min\_\{u\\in\\mathcal\{U\}\_\{cl\}\}J^\{cl\}\_\{1\}\(x,u,v^\{\*\}\(u\)\),\(2\.4\)wherev∗\(u\)=argminv∈𝒱clJ2cl\(x,u,v\)v^\{\*\}\(u\)=\\arg\\min\_\{v\\in\\mathcal\{V\}\_\{cl\}\}J^\{cl\}\_\{2\}\(x,u,v\)\.
For the dynamic system \([2\.1](https://arxiv.org/html/2606.28671#S2.E1)\) and performance function \([2\.2](https://arxiv.org/html/2606.28671#S2.E2)\), the objective of classical SDG aims to design appropriate policiesu∗u^\{\*\}andv∗v^\{\*\}satisfying the following*Stacklelberg equilibrium*conditions,
\{J1cl\(x,u∗,v∗\(u∗\)\)≤J1cl\(x,u,v∗\(u\)\),∀u∈𝒰cl;J2cl\(x,u,v∗\(u\)\)≤J2cl\(x,u,v\),∀v∈𝒱cl,\\left\\\{\\begin\{array\}\[\]\{ll\}J^\{cl\}\_\{1\}\(x,u^\{\*\},v^\{\*\}\(u^\{\*\}\)\)\\leq J^\{cl\}\_\{1\}\(x,u,v^\{\*\}\(u\)\),\\,\\forall u\\in\\mathcal\{U\}\_\{cl\};\\\\ J^\{cl\}\_\{2\}\(x,u,v^\{\*\}\(u\)\)\\leq J^\{cl\}\_\{2\}\(x,u,v\),\\,\\forall v\\in\\mathcal\{V\}\_\{cl\},\\end\{array\}\\right\.\(2\.5\)wherev∗\(u\)=argminv∈𝒱clJ2cl\(x,u,v\)v^\{\*\}\(u\)=\\arg\\min\_\{v\\in\\mathcal\{V\}\_\{cl\}\}J^\{cl\}\_\{2\}\(x,u,v\)\.
### II\-BExploratory two\-player LQ\-SDGs framework
System \([2\.1](https://arxiv.org/html/2606.28671#S2.E1)\) does not incorporate stochastic effects and environmental variations, which motivates us to introduce Brownian motion and regime switching in our analysis\.
Letα=\{αs,s≥0\}\\alpha=\\\{\\alpha\_\{s\},s\\geq 0\\\}be a continuous\-time Markov chain defined on a complete probability space\(Ω,ℱ,\{ℱs\}s≥0,ℙ\)\(\\Omega,\\mathcal\{F\},\\\{\\mathcal\{F\}\_\{s\}\\\}\_\{s\\geq 0\},\\mathbb\{P\}\)with a finite state spaceℳ=\{1,2,…,m\}\\mathcal\{M\}=\\\{1,2,\\dots,m\\\}\. The generator matrix of this Markov chain is defined asΓ=\(qij\)m×m\\Gamma=\(q\_\{ij\}\)\_\{m\\times m\}, whereqij≥0q\_\{ij\}\\geq 0\(i≠ji\\neq j\) represents the transition rate from regimeiito regimejj, satisfyingqii=−∑j=1,j≠imqijq\_\{ii\}=\-\\sum\_\{j=1,j\\neq i\}^\{m\}q\_\{ij\}\. The dynamical system with regime switching is given by
dxs=b\(xs,αs,us,vs\)ds\+σ\(xs,αs,us,vs\)dWs,s\>0,dx\_\{s\}=b\(x\_\{s\},\\alpha\_\{s\},u\_\{s\},v\_\{s\}\)ds\+\\sigma\(x\_\{s\},\\alpha\_\{s\},u\_\{s\},v\_\{s\}\)dW\_\{s\},s\>0,\(2\.6\)where the drift and diffusion terms are determined by the current regimei∈ℳi\\in\\mathcal\{M\}:
b\(x,i,u,v\):=A\(i\)x\+B1\(i\)u\+B2\(i\)v,b\(x,i,u,v\):=A\(i\)x\+B\_\{1\}\(i\)u\+B\_\{2\}\(i\)v,σ\(x,i,u,v\):=C\(i\)x\+D1\(i\)u\+D2\(i\)v\.\\sigma\(x,i,u,v\):=C\(i\)x\+D\_\{1\}\(i\)u\+D\_\{2\}\(i\)v\.Here,W=\{Ws,s≥0\}W=\\\{W\_\{s\},s\\geq 0\\\}is a standardnn\-dimensional Brownian motion, assumed to be independent of the Markov chainα\\alpha\. All system parameters \(e\.g\.,A\(i\),B1\(i\)A\(i\),B\_\{1\}\(i\), etc\.\) are constant matrices dependent on the current regimei∈ℳi\\in\\mathcal\{M\}\. Correspondingly, the classical performance function for each player also depends on the current regimei∈ℳi\\in\\mathcal\{M\}:
r1\(x,i,u,v\)=xTQ1\(i\)x\+\(u\+θ1\(i\)v\)TR1\(i\)\(u\+θ1\(i\)v\),r\_\{1\}\(x,i,u,v\)=x^\{T\}Q\_\{1\}\(i\)x\+\(u\+\\theta\_\{1\}\(i\)v\)^\{T\}R\_\{1\}\(i\)\(u\+\\theta\_\{1\}\(i\)v\),r2\(x,i,u,v\)=xTQ2\(i\)x\+\(v\+θ2\(i\)u\)TR2\(i\)\(v\+θ2\(i\)u\),r\_\{2\}\(x,i,u,v\)=x^\{T\}Q\_\{2\}\(i\)x\+\(v\+\\theta\_\{2\}\(i\)u\)^\{T\}R\_\{2\}\(i\)\(v\+\\theta\_\{2\}\(i\)u\),whereQk\(i\)≥0,Rk\(i\)\>0,θk\(i\)∈\(0,1\)Q\_\{k\}\(i\)\\geq 0,R\_\{k\}\(i\)\>0,\\theta\_\{k\}\(i\)\\in\(0,1\)fork=1,2k=1,2\.
A core principle in RL is environmental exploration through randomized actions\. This involves replacing deterministic policies with probability distributions \(measures\) to facilitate exploration\. Entropy quantifies the dispersion of these distributions: low entropy indicates concentration, while high entropy reflects greater randomness\. This paper focuses on feedback distributional policies, where the random policy’s distribution depends solely on the current statexxand regimeii\.
Let𝝅:\(x,i\)∈ℝn×ℳ→𝝅\(⋅∣x,i\)∈𝒫\(ℝp\)\\boldsymbol\{\\pi\}:\(x,i\)\\in\\mathbb\{R\}^\{n\}\\times\\mathcal\{M\}\\rightarrow\\boldsymbol\{\\pi\}\(\\cdot\\mid x,i\)\\in\\mathcal\{P\}\(\\mathbb\{R\}^\{p\}\)\(resp\.,𝜸:\(x,i\)∈ℝn×ℳ→𝜸\(⋅∣x,i\)∈𝒫\(ℝp\)\\boldsymbol\{\\gamma\}:\(x,i\)\\in\\mathbb\{R\}^\{n\}\\times\\mathcal\{M\}\\rightarrow\\boldsymbol\{\\gamma\}\(\\cdot\\mid x,i\)\\in\\mathcal\{P\}\(\\mathbb\{R\}^\{p\}\)\) be a given stochastic feedback control, where𝒫\(ℝp\)\\mathcal\{P\}\(\\mathbb\{R\}^\{p\}\)is the set of probability density functions defined onℝp\\mathbb\{R\}^\{p\}\. Let the control actionsu∈ℝpu\\in\\mathbb\{R\}^\{p\}andv∈ℝpv\\in\\mathbb\{R\}^\{p\}be sampled from the conditional distributions𝝅\(⋅∣x,i\)\\boldsymbol\{\\pi\}\(\\cdot\\mid x,i\)and𝜸\(⋅∣x,i\)\\boldsymbol\{\\gamma\}\(\\cdot\\mid x,i\), respectively\. For notational brevity, when the current statexxand regimeiiare clear from the context, we will often drop the conditioning and write𝝅\(u\)\\boldsymbol\{\\pi\}\(u\)and𝜸\(v\)\\boldsymbol\{\\gamma\}\(v\)to denote𝝅\(u∣x,i\)\\boldsymbol\{\\pi\}\(u\\mid x,i\)and𝜸\(v∣x,i\)\\boldsymbol\{\\gamma\}\(v\\mid x,i\)\. Inspired by\[[26](https://arxiv.org/html/2606.28671#bib.bib50)\], the drift, diffusion, and performance functions are then extended to an exploratory version \([2\.7](https://arxiv.org/html/2606.28671#S2.E7)\)–\([2\.9](https://arxiv.org/html/2606.28671#S2.E9)\) given by
b~\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpb\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v,\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\):=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}b\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv,\(2\.7\)σ~\(x,i,𝝅,𝜸\)\\displaystyle\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\(2\.8\):=\{∫ℝp\[∫ℝpσ\(x,i,u,v\)σT\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\}1/2,\\displaystyle=\\left\\\{\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\sigma\(x,i,u,v\)\\sigma^\{T\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\\right\\\}^\{1/2\},and
r~1\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpr1\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v,\\displaystyle\\tilde\{r\}\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}r\_\{1\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv,\(2\.9\)r~2\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpr2\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\.\\displaystyle\\tilde\{r\}\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}r\_\{2\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\.For the reader’s convenience, we postpone the derivation of \([2\.7](https://arxiv.org/html/2606.28671#S2.E7)\)–\([2\.9](https://arxiv.org/html/2606.28671#S2.E9)\) to Appendix[A](https://arxiv.org/html/2606.28671#A1)\.
Accordingly, the exploratory dynamic system for \([2\.6](https://arxiv.org/html/2606.28671#S2.E6)\) is given by
dXs𝝅,𝜸=b~\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)ds\+σ~\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)dWs,dX\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}=\\tilde\{b\}\(X^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)ds\+\\tilde\{\\sigma\}\(X^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)dW\_\{s\},\(2\.10\)where we adopt the shorthands𝝅s\(⋅\):=𝝅\(⋅∣Xs𝝅,𝜸,αs\)\\boldsymbol\{\\pi\}\_\{s\}\(\\cdot\):=\\boldsymbol\{\\pi\}\(\\cdot\\mid X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\}\)and𝜸s\(⋅\):=𝜸\(⋅∣Xs𝝅,𝜸,αs\)\\boldsymbol\{\\gamma\}\_\{s\}\(\\cdot\):=\\boldsymbol\{\\gamma\}\(\\cdot\\mid X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\}\)to denote the policies evaluated along the state trajectories\.
Similar to\[[27](https://arxiv.org/html/2606.28671#bib.bib39)\], to encourage exploration, we add an entropy regularization term to the performance function\. Then, with \([2\.10](https://arxiv.org/html/2606.28671#S2.E10)\), the exploratory performance functions of playerIIand playerIIIIare respectively defined by
J1\(x,i,𝝅,𝜸\)=\\displaystyle J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)=𝔼x,i\[∫0∞e−ρ1s\[r~1\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{1\}s\}\\Big\[\\tilde\{r\}\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)\\right\.\(2\.11\)\+λ1∫ℝp𝝅s\(u\)ln𝝅s\(u\)du\]ds\],\\displaystyle\\left\.\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\ln\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du\\,\\Big\]ds\\right\],J2\(x,i,𝝅,𝜸\)=\\displaystyle J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)=𝔼x,i\[∫0∞e−ρ2s\[r~2\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{2\}s\}\\Big\[\\tilde\{r\}\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)\\right\.\(2\.12\)\+λ2∫ℝp𝜸s\(v\)ln𝜸s\(v\)dv\]ds\],\\displaystyle\\left\.\+\\lambda\_\{2\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\gamma\}\_\{s\}\(v\)\\ln\\boldsymbol\{\\gamma\}\_\{s\}\(v\)\\,dv\\Big\]ds\\right\],where𝔼x,i\[⋅\]\\mathbb\{E\}\_\{x,i\}\[\\cdot\]denotes the expectation conditioned on the initial stateX0𝝅,𝜸=xX\_\{0\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}=xand initial regimeα0=i\\alpha\_\{0\}=i\. Here,ρk\>0\\rho\_\{k\}\>0\(k=1,2\)\(k=1,2\)are the*discount factors*, and the exogenous parametersλk\>0\\lambda\_\{k\}\>0\(k=1,2\)\(k=1,2\)are the*temperature parameters*, which reflect the level of exploration\.
###### Assumption 1
Assume that
1. \(i\)for eachs≥0s\\geq 0,𝝅s∈𝒫\(ℝp\)\\boldsymbol\{\\pi\}\_\{s\}\\in\\mathcal\{P\}\(\\mathbb\{R\}^\{p\}\)and𝜸s∈𝒫\(ℝp\)\\boldsymbol\{\\gamma\}\_\{s\}\\in\\mathcal\{P\}\(\\mathbb\{R\}^\{p\}\)a\.s\.;
2. \(ii\)\{∫A𝝅s\(u\)𝑑u,s≥0\}\\\{\\int\_\{A\}\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du,s\\geq 0\\\}and\{∫A𝜸s\(v\)𝑑v,s≥0\}\\\{\\int\_\{A\}\\boldsymbol\{\\gamma\}\_\{s\}\(v\)\\,dv,s\\geq 0\\\}areℱs\\mathcal\{F\}\_\{s\}\-progressively measurable, for eachA∈ℬ\(ℝp\)A\\in\\mathcal\{B\}\(\\mathbb\{R\}^\{p\}\);
3. \(iii\)𝔼x,i\[∫0t\(∫ℝpuTu𝝅s\(u\)𝑑u\+∫ℝpvTv𝜸s\(v\)𝑑v\)𝑑s\]<∞\\mathbb\{E\}\_\{x,i\}\[\\int\_\{0\}^\{t\}\\left\(\\int\_\{\\mathbb\{R\}^\{p\}\}u^\{T\}u\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du\+\\int\_\{\\mathbb\{R\}^\{p\}\}v^\{T\}v\\boldsymbol\{\\gamma\}\_\{s\}\(v\)\\,dv\\right\)ds\]<\\infty, for eacht≥0t\\geq 0;
4. \(iv\)lim infs→∞𝔼x,i\[e−ρksJk\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)\]=0\\liminf\_\{s\\rightarrow\\infty\}\\mathbb\{E\}\_\{x,i\}\\left\[e^\{\-\\rho\_\{k\}s\}J\_\{k\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)\\right\]=0, fork=1,2k=1,2;
5. \(v\)fork=1,2k=1,2,𝔼x,i\[∫0∞e−ρksr~k\(Xs𝝅,𝜸,αs,𝝅s,𝜸s\)𝑑s\]<∞\\mathbb\{E\}\_\{x,i\}\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{k\}s\}\\tilde\{r\}\_\{k\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)ds\]<\\infty\.
Denote by𝒰\\mathcal\{U\}\(resp\.𝒱\\mathcal\{V\}\) the set of control policies𝝅\\boldsymbol\{\\pi\}\(resp\.𝜸\\boldsymbol\{\\gamma\}\) satisfying Assumption[1](https://arxiv.org/html/2606.28671#Thmassumption1)\. A policy in𝒰\\mathcal\{U\}or𝒱\\mathcal\{V\}is called an admissible control with respect to playerIIor playerIIII, respectively\. The following Proposition[1](https://arxiv.org/html/2606.28671#Thmproposition1)guarantees that the controlled system \([2\.10](https://arxiv.org/html/2606.28671#S2.E10)\) is well defined and the proof is postponed to Appendix[B](https://arxiv.org/html/2606.28671#A2)\.
###### Proposition 1
For any𝛑∈𝒰\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}and𝛄∈𝒱\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}, the exploratory dynamic system \([2\.10](https://arxiv.org/html/2606.28671#S2.E10)\) has a unique strong solution\.
###### Definition 2
The value function of playerIIIIis
V2\(x,i;𝝅\)=min𝜸∈𝒱J2\(x,i,𝝅,𝜸\)\.V\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)=\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\.\(2\.13\)The value function of playerIIis
V1\(x,i\)=min𝝅∈𝒰J1\(x,i,𝝅,𝜸∗\(𝝅\)\),V\_\{1\}\(x,i\)=\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\),\(2\.14\)where𝛄∗\(𝛑\)=argmin𝛄∈𝒱J2\(x,i,𝛑,𝛄\)\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)=\\arg\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\.
###### Problem 1
Design appropriate policies𝛑∗\\boldsymbol\{\\pi\}^\{\*\}and𝛄∗\\boldsymbol\{\\gamma\}^\{\*\}to satisfy the following Stackelberg equilibrium conditions, i\.e\.,
\{J1\(x,i,𝝅∗,𝜸∗\(𝝅∗\)\)≤J1\(x,i,𝝅,𝜸∗\(𝝅\)\),∀𝝅∈𝒰;J2\(x,i,𝝅,𝜸∗\(𝝅\)\)≤J2\(x,i,𝝅,𝜸\),∀𝜸∈𝒱\.\\left\\\{\\begin\{array\}\[\]\{ll\}J\_\{1\}\(x,i,\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}^\{\*\}\)\)\\leq J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\),\\,\\forall\\boldsymbol\{\\pi\}\\in\\mathcal\{U\};\\\\ J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\leq J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),\\,\\forall\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\.\\end\{array\}\\right\.\(2\.15\)
For short, we rewrite𝜸∗\(𝝅∗\)\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}^\{\*\}\)as𝜸∗\\boldsymbol\{\\gamma\}^\{\*\}, which will not lead to any confusion in the following analysis\. Then,\(𝝅∗,𝜸∗\)\(\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)satisfying \([2\.15](https://arxiv.org/html/2606.28671#S2.E15)\) is said to be a Stackelberg equilibrium policy pair\. For short, we also rewriteV2\(x,i;𝝅∗\):=V2\(x,i\)V\_\{2\}\(x,i;\\boldsymbol\{\\pi\}^\{\*\}\):=V\_\{2\}\(x,i\)\.
## IIIEquilibrium policies for ERRL LQ\-SDGs
In this section, we employ DPP to elucidate the Stackelberg equilibrium policy pair described in*Problem[1](https://arxiv.org/html/2606.28671#Thmproblem1)*\. First, we derive the HJBI equations for playerIIand playerIIII\(see Lemma[1](https://arxiv.org/html/2606.28671#Thmlemma1)\), and then present the equilibrium policies of both players as functionals of the solution to the HJBI equations \(see Theorem[1](https://arxiv.org/html/2606.28671#Thmtheorem1), Theorem[2](https://arxiv.org/html/2606.28671#Thmtheorem2)\), and finally prove that the identified Stackelberg equilibrium policies solve*Problem[1](https://arxiv.org/html/2606.28671#Thmproblem1)*\(see Theorem[3](https://arxiv.org/html/2606.28671#Thmtheorem3)\)\. For reading ease, the proof for Lemma[1](https://arxiv.org/html/2606.28671#Thmlemma1)is postponed to Appendix[C](https://arxiv.org/html/2606.28671#A3)\.
###### Lemma 1
Suppose that∀i∈ℳ\\forall i\\in\\mathcal\{M\},V1\(⋅,i\),V2\(⋅,i;𝛑\)∈C2\(ℝn\)V\_\{1\}\(\\cdot,i\),V\_\{2\}\(\\cdot,i;\\boldsymbol\{\\pi\}\)\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\), then the formal HJBI equations associated with playerIIIIand playerIIare given by \([3\.1](https://arxiv.org/html/2606.28671#S3.E1)\) and \([3\.2](https://arxiv.org/html/2606.28671#S3.E2)\)\.
ρ2ϕ2\(x,i;𝝅\)=\\displaystyle\\rho\_\{2\}\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)=min𝜸∈𝒱ℋ2\(x,i,∇ϕ2\(x,i;𝝅\),Δϕ2\(x,i;𝝅\),𝝅,𝜸\)\\displaystyle\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\(3\.1\)=\\displaystyle=min𝜸∈𝒱\[r~2\(x,i,𝝅,𝜸\)\+λ2∫ℝp𝜸\(v\)ln𝜸\(v\)dv\\displaystyle\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}\\Big\[\\tilde\{r\}\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+\\lambda\_\{2\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\gamma\}\(v\)\\ln\\boldsymbol\{\\gamma\}\(v\)\\,dv\+∇ϕ2\\displaystyle\+\\nabla\\phi\_\{2\}\(x,i;𝝅\)Tb~\(x,i,𝝅,𝜸\)\+∑j=1mqijϕ2\(x,j;𝝅\)\\displaystyle\(x,i;\\boldsymbol\{\\pi\}\)^\{T\}\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+\\sum\_\{j=1\}^\{m\}q\_\{ij\}\\phi\_\{2\}\(x,j;\\boldsymbol\{\\pi\}\)\+12tr\\displaystyle\+\\frac\{1\}\{2\}tr\(σ~\(x,i,𝝅,𝜸\)σ~T\(x,i,𝝅,𝜸\)Δϕ2\(x,i;𝝅\)\)\],\\displaystyle\\left\(\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\tilde\{\\sigma\}^\{T\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)\\right\)\\Big\],ρ1ϕ1\(x,i\)=\\displaystyle\\rho\_\{1\}\\phi\_\{1\}\(x,i\)=min𝝅∈𝒰ℋ1\(x,i,∇ϕ1\(x,i\),Δϕ1\(x,i\),𝝅,𝜸∗\(𝝅\)\)\\displaystyle\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\mathcal\{H\}\_\{1\}\(x,i,\\nabla\\phi\_\{1\}\(x,i\),\\Delta\\phi\_\{1\}\(x,i\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\(3\.2\)=\\displaystyle=min𝝅∈𝒰\[r~1\(x,i,𝝅,𝜸∗\(𝝅\)\)\+λ1∫ℝp𝝅\(u\)ln𝝅\(u\)du\\displaystyle\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\Big\[\\tilde\{r\}\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\(u\)\\ln\\boldsymbol\{\\pi\}\(u\)\\,du\+∇ϕ1\\displaystyle\+\\nabla\\phi\_\{1\}\(x,i\)Tb~\(x,i,𝝅,𝜸∗\(𝝅\)\)\+∑j=1mqijϕ1\(x,j\)\\displaystyle\(x,i\)^\{T\}\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+\\sum\_\{j=1\}^\{m\}q\_\{ij\}\\phi\_\{1\}\(x,j\)\+12tr\\displaystyle\+\\frac\{1\}\{2\}tr\(σ~\(x,i,𝝅,𝜸∗\(𝝅\)\)σ~T\(x,i,𝝅,𝜸∗\(𝝅\)\)Δϕ1\(x,i\)\)\],\\displaystyle\\left\(\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\Delta\\phi\_\{1\}\(x,i\)\\right\)\\Big\],where
𝜸∗\(𝝅\)=argmin𝜸∈𝒱ℋ2\(x,i,∇ϕ2\(x,i;𝝅\),Δϕ2\(x,i;𝝅\),𝝅,𝜸\)\.\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)=\\arg\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\.
###### Theorem 1
The equilibrium policy of playerIIIIis a Gaussian distribution given by
𝜸∗\(𝝅\)=𝒩\(β,λ2\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)\)−1\),\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)=\\mathcal\{N\}\\left\(\\beta,\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\),\(3\.3\)where
β\\displaystyle\\beta:=−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)\)−1\\displaystyle=\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(3\.4\)\(2θ2\(i\)R2\(i\)\+D2\(i\)TΔϕ2\(x,i;𝝅\)D1\(i\)\)∫ℝpu𝝅\(u\)𝑑u\\displaystyle\\quad\\left\(2\\theta\_\{2\}\(i\)R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{1\}\(i\)\\right\)\\int\_\{\\mathbb\{R\}^\{p\}\}u\\boldsymbol\{\\pi\}\(u\)\\,du−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(D2\(i\)TΔϕ2\(x,i;𝝅\)C\(i\)x\+B2\(i\)T∇ϕ2\(x,i;𝝅\)\)\.\\displaystyle\\quad\\left\(D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)C\(i\)x\+B\_\{2\}\(i\)^\{T\}\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)\\right\)\.
###### Proof 1
For any leader’s policy𝛑\\boldsymbol\{\\pi\}, the follower aims to track optimal response under regimei∈ℳi\\in\\mathcal\{M\}via
𝜸∗\(𝝅\)=argmin𝜸∈𝒱ℋ2\(x,i,∇ϕ2\(x,i;𝝅\),Δϕ2\(x,i;𝝅\),𝝅,𝜸\)\.\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)=\\arg\\min\_\{\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\.Subject to the constraints that𝛄\\boldsymbol\{\\gamma\}is a probability distribution, we have∫ℝp𝛄\(v\)𝑑v=1\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\gamma\}\(v\)\\,dv=1\. Introducing the Lagrange multiplierμ\\muto this problem yields
ℒ2\(x,i,∇ϕ2\(x,i;𝝅\),Δϕ2\(x,i;𝝅\),𝝅,𝜸,μ\)\\displaystyle\\mathcal\{L\}\_\{2\}\(x,i,\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\},\\mu\)\(3\.5\)=\\displaystyle=r~2\(x,i,𝝅,𝜸\)\+λ2∫ℝp𝜸\(v\)ln𝜸\(v\)𝑑v\\displaystyle\\tilde\{r\}\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+\\lambda\_\{2\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\gamma\}\(v\)\\ln\\boldsymbol\{\\gamma\}\(v\)\\,dv\+∇ϕ2\(x,i;𝝅\)Tb~\(x,i,𝝅,𝜸\)\+∑j=1mqijϕ2\(x,j;𝝅\)\\displaystyle\+\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)^\{T\}\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+\\sum\_\{j=1\}^\{m\}q\_\{ij\}\\phi\_\{2\}\(x,j;\\boldsymbol\{\\pi\}\)\+12tr\(σ~\(x,i,𝝅,𝜸\)σ~T\(x,i,𝝅,𝜸\)Δϕ2\(x,i;𝝅\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\tilde\{\\sigma\}^\{T\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)\\right\)−μ\(∫ℝp𝜸\(v\)𝑑v−1\)\.\\displaystyle\-\\mu\\left\(\\int\_\{\\mathbb\{R\}^\{p\}\}\\,\\boldsymbol\{\\gamma\}\(v\)\\,dv\-1\\right\)\.Substitute the specific forms ofb~\(x,i,𝛑,𝛄\)\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),σ~\(x,i,𝛑,𝛄\)\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),r~2\(x,i,𝛑,𝛄\)\\tilde\{r\}\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\), then the stationary condition is
\{δℒ2=∫ℝp\[2θ2\(i\)vTR2\(i\)∫ℝpu𝝅\(u\)du\+vTR2\(i\)v\+∇ϕ2\(x,i;𝝅\)TB2\(i\)v\+vTD2\(i\)TΔϕ2\(x,i;𝝅\)C\(i\)x\+∫ℝpuT𝝅\(u\)𝑑uD1\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)v\+12vTD2\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)v\+λ2\(ln𝜸\+1\)−μ\]δ𝜸\(v\)dv=0,0=∫ℝp𝜸\(v\)𝑑v−1\.\\left\\\{\\begin\{array\}\[\]\{ll\}\\delta\\mathcal\{L\}\_\{2\}=\\int\_\{\\mathbb\{R\}^\{p\}\}\\Big\[2\\theta\_\{2\}\(i\)v^\{T\}R\_\{2\}\(i\)\\int\_\{\\mathbb\{R\}^\{p\}\}u\\boldsymbol\{\\pi\}\(u\)\\,du\\\\ \\qquad\\quad\+v^\{T\}R\_\{2\}\(i\)v\+\\nabla\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)^\{T\}B\_\{2\}\(i\)v\\\\ \\qquad\\quad\+v^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)C\(i\)x\\\\ \\qquad\\quad\+\\int\_\{\\mathbb\{R\}^\{p\}\}u^\{T\}\\boldsymbol\{\\pi\}\(u\)\\,du\\,D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)v\\\\ \\qquad\\quad\+\\frac\{1\}\{2\}v^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)v\\\\ \\qquad\\quad\+\\lambda\_\{2\}\(\\ln\\boldsymbol\{\\gamma\}\+1\)\-\\mu\\Big\]\\,\\delta\\boldsymbol\{\\gamma\}\(v\)dv=0,\\\\ 0=\\int\_\{\\mathbb\{R\}^\{p\}\}\\,\\boldsymbol\{\\gamma\}\(v\)\\,dv\-1\.\\end\{array\}\\right\.Solve for𝛄\\boldsymbol\{\\gamma\}, the equilibrium policy of playerIIIIis
𝜸∗\(𝝅\)=𝒩\(β,λ2\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i;𝝅\)D2\(i\)\)−1\),\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)=\\mathcal\{N\}\\left\(\\beta,\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\),whereβ\\betais given by \([3\.4](https://arxiv.org/html/2606.28671#S3.E4)\)\.
We now turn to track playerII’s equilibrium policy\. For simplicity, letϕ2\(x,i\):=ϕ2\(x,i;𝝅∗\)\\phi\_\{2\}\(x,i\):=\\phi\_\{2\}\(x,i;\\boldsymbol\{\\pi\}^\{\*\}\),
Λ1:=\\displaystyle\\Lambda\_\{1\}=−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(2θ2\(i\)R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D1\(i\)\),\\displaystyle\\quad\\left\(2\\theta\_\{2\}\(i\)R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{1\}\(i\)\\right\),Λ2:=\\displaystyle\\Lambda\_\{2\}=−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(D2\(i\)TΔϕ2\(x,i\)C\(i\)x\+B2\(i\)T∇ϕ2\(x,i\)\),\\displaystyle\\quad\\left\(D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)C\(i\)x\+B\_\{2\}\(i\)^\{T\}\\nabla\\phi\_\{2\}\(x,i\)\\right\),Λ3:=\\displaystyle\\Lambda\_\{3\}=2R1\(i\)\+D1\(i\)TΔϕ1\(x,i\)D1\(i\)\\displaystyle 2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{1\}\(i\)\+D1\(i\)TΔϕ1\(x,i\)D2\(i\)Λ1\+2θ1\(i\)R1\(i\)Λ1\\displaystyle\+D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{2\}\(i\)\\Lambda\_\{1\}\+2\\theta\_\{1\}\(i\)R\_\{1\}\(i\)\\Lambda\_\{1\}\+Λ1TD2\(i\)TΔϕ1\(x,i\)D2\(i\)Λ1\\displaystyle\+\\Lambda\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{2\}\(i\)\\Lambda\_\{1\}\+2θ1\(i\)2Λ1TR1\(i\)Λ1\+2θ1\(i\)Λ1TR1\(i\)\\displaystyle\+2\\theta\_\{1\}\(i\)^\{2\}\\Lambda\_\{1\}^\{T\}R\_\{1\}\(i\)\\Lambda\_\{1\}\+2\\theta\_\{1\}\(i\)\\Lambda\_\{1\}^\{T\}R\_\{1\}\(i\)\+Λ1TD2\(i\)TΔϕ1\(x,i\)D1\(i\)\.\\displaystyle\+\\Lambda\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{1\}\(i\)\.and
α∗:=\\displaystyle\\alpha\_\{\*\}=−Λ3−1\[Λ1TB2\(i\)T∇ϕ1\(x,i\)\+D1\(i\)TΔϕ1\(x,i\)C\(i\)x\\displaystyle\-\\Lambda\_\{3\}^\{\-1\}\\Big\[\\Lambda\_\{1\}^\{T\}B\_\{2\}\(i\)^\{T\}\\nabla\\phi\_\{1\}\(x,i\)\+D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)C\(i\)x\(3\.6\)\+B1\(i\)T∇ϕ1\(x,i\)\+Λ1TD2\(i\)TΔϕ1\(x,i\)C\(i\)x\\displaystyle\+B\_\{1\}\(i\)^\{T\}\\nabla\\phi\_\{1\}\(x,i\)\+\\Lambda\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)C\(i\)x\+\(D1\(i\)TΔϕ1\(x,i\)D2\(i\)\+Λ1TD2\(i\)TΔϕ1\(x,i\)D2\(i\)\\displaystyle\+\\left\(D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{2\}\(i\)\+\\Lambda\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{2\}\(i\)\\right\.\+2θ1\(i\)R1\(i\)\+2θ1\(i\)2Λ1TR1\(i\)\)Λ2\]\.\\displaystyle\\left\.\+2\\theta\_\{1\}\(i\)R\_\{1\}\(i\)\+2\\theta\_\{1\}\(i\)^\{2\}\\Lambda\_\{1\}^\{T\}R\_\{1\}\(i\)\\right\)\\Lambda\_\{2\}\\Big\]\.
###### Assumption 2
The matrixΛ3\\Lambda\_\{3\}is an invertible matrix\.
###### Theorem 2
Under Assumption[2](https://arxiv.org/html/2606.28671#Thmassumption2), the equilibrium policy of playerIIis
𝝅∗=𝒩\(α∗,λ1\(2R1\(i\)\+D1\(i\)TΔϕ1\(x,i\)D1\(i\)\)−1\)\.\\boldsymbol\{\\pi\}^\{\*\}=\\mathcal\{N\}\\left\(\\alpha\_\{\*\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\)\.\(3\.7\)
###### Proof 2
The proof follows a similar approach to that of Theorem[1](https://arxiv.org/html/2606.28671#Thmtheorem1)and is therefore omitted for brevity\.
Substitute \([3\.7](https://arxiv.org/html/2606.28671#S3.E7)\) into \([3\.3](https://arxiv.org/html/2606.28671#S3.E3)\), and the equilibrium policy pair for playerIIand playerIIIIis
\{𝝅∗=𝒩\(α∗,λ1\(2R1\(i\)\+D1\(i\)TΔϕ1\(x,i\)D1\(i\)\)−1\),𝜸∗=𝒩\(β∗,λ2\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D2\(i\)\)−1\),\\left\\\{\\begin\{array\}\[\]\{ll\}\\boldsymbol\{\\pi\}^\{\*\}=\\mathcal\{N\}\\left\(\\alpha\_\{\*\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta\\phi\_\{1\}\(x,i\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\),\\\\ \\boldsymbol\{\\gamma\}^\{\*\}=\\mathcal\{N\}\\left\(\\beta\_\{\*\},\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\),\\end\{array\}\\right\.\(3\.8\)whereα∗\\alpha\_\{\*\}is given by \([3\.6](https://arxiv.org/html/2606.28671#S3.E6)\) and
β∗=\\displaystyle\\beta\_\{\*\}=−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(2θ2\(i\)R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D1\(i\)\)α∗\\displaystyle\\quad\\left\(2\\theta\_\{2\}\(i\)R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{1\}\(i\)\\right\)\\alpha\_\{\*\}−\(2R2\(i\)\+D2\(i\)TΔϕ2\(x,i\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(D2\(i\)TΔϕ2\(x,i\)C\(i\)x\+B2\(i\)T∇ϕ2\(x,i\)\)\.\\displaystyle\\quad\\left\(D\_\{2\}\(i\)^\{T\}\\Delta\\phi\_\{2\}\(x,i\)C\(i\)x\+B\_\{2\}\(i\)^\{T\}\\nabla\\phi\_\{2\}\(x,i\)\\right\)\.Then, substituting \([3\.8](https://arxiv.org/html/2606.28671#S3.E8)\) into \([3\.1](https://arxiv.org/html/2606.28671#S3.E1)\), \([3\.2](https://arxiv.org/html/2606.28671#S3.E2)\) and removing the min operator, we obtain the following coupled HJBI equations:
ρ2ϕ2\(x,i\)=\\displaystyle\\rho\_\{2\}\\phi\_\{2\}\(x,i\)=ℋ2\(x,i,∇ϕ2\(x,i\),Δϕ2\(x,i\),𝝅∗,𝜸∗\),\\displaystyle\\mathcal\{H\}\_\{2\}\(x,i,\\nabla\\phi\_\{2\}\(x,i\),\\Delta\\phi\_\{2\}\(x,i\),\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\),\(3\.9\)ρ1ϕ1\(x,i\)=\\displaystyle\\rho\_\{1\}\\phi\_\{1\}\(x,i\)=ℋ1\(x,i,∇ϕ1\(x,i\),Δϕ1\(x,i\),𝝅∗,𝜸∗\)\.\\displaystyle\\mathcal\{H\}\_\{1\}\(x,i,\\nabla\\phi\_\{1\}\(x,i\),\\Delta\\phi\_\{1\}\(x,i\),\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)\.\(3\.10\)
###### Theorem 3
Under Assumption[2](https://arxiv.org/html/2606.28671#Thmassumption2), ifϕ1\(x,i\)\\phi\_\{1\}\(x,i\)andϕ2\(x,i\)\\phi\_\{2\}\(x,i\)are solutions to the coupled HJBI equations \([3\.9](https://arxiv.org/html/2606.28671#S3.E9)\) and \([3\.10](https://arxiv.org/html/2606.28671#S3.E10)\), then
1. 1\.J1\(x,i,𝝅∗,𝜸∗\)≤J1\(x,i,𝝅,𝜸∗\(𝝅\)\),∀𝝅∈𝒰J\_\{1\}\(x,i,\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)\\leq J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\),\\,\\forall\\boldsymbol\{\\pi\}\\in\\mathcal\{U\},
2. 2\.J1\(x,i,𝝅∗,𝜸∗\)=ϕ1\(x,i\)J\_\{1\}\(x,i,\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)=\\phi\_\{1\}\(x,i\);
3. 3\.J2\(x,i,𝝅,𝜸∗\(𝝅\)\)≤J2\(x,i,𝝅,𝜸\),∀𝜸∈𝒱J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\leq J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),\\,\\forall\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\},
4. 4\.J2\(x,i,𝝅∗,𝜸∗\)=ϕ2\(x,i\)J\_\{2\}\(x,i,\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)=\\phi\_\{2\}\(x,i\)\.
The proof can be found in Appendix[D](https://arxiv.org/html/2606.28671#A4)\.
## IVPIA for ERRL LQ\-SDG
Although \([3\.8](https://arxiv.org/html/2606.28671#S3.E8)\) provides optimal feedback strategies for both players, the HJBI equations are nonlinear second\-order PDEs with integral terms, making analytical solutions difficult to derive\. In this section, we propose a PIA\[[27](https://arxiv.org/html/2606.28671#bib.bib39)\]to approximate the value of the ERRL LQ\-SDGs and the optimal strategies under regime switching\.
The PIA shares similarities with the actor\-critic framework, as it involves both a policy \(actor\) and a value function \(critic\)\. Specifically:
- •*Policy \(Actor\)*: The strategies𝝅\\boldsymbol\{\\pi\}and𝜸\\boldsymbol\{\\gamma\}generate actions\.
- •*Value Function \(Critic\)*:J2\(x,i,𝝅,𝜸\)J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),J1\(x,i,𝝅,𝜸\)J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)estimate state\-regime values and are updated based on policy performance\.
However, PIA extends beyond standard actor\-critic methods by incorporating*robust optimization*via themin\\minoperator over uncertainty sets𝒫\\mathcal\{P\}\. This key distinction makes the PIA a specialized approach tailored to stochastic control problems\.
As the PIA often requires discretizing the state space, which brings about the problem of the “curse of dimensionality”, referring to the relevant research of\[[9](https://arxiv.org/html/2606.28671#bib.bib44),[34](https://arxiv.org/html/2606.28671#bib.bib46)\], we propose a critic neural networks \(critic NNs\) algorithm in subsection[IV\-C](https://arxiv.org/html/2606.28671#S4.SS3)\. This algorithm mainly parameterizes the value function and uses the least\-squares \(LS\) method to determine the parameter weights during the policy evaluation stage\.
### IV\-AConvergence of PIA
###### Theorem 4
\(Policy improvement theorem for playerIIII\)\. Let𝛑\\boldsymbol\{\\pi\}be given and𝛄\(𝛑\)\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)be the corresponding given admissible feedback policy\. Suppose that the corresponding performance functionJ2\(⋅,i,𝛑,𝛄\(𝛑\)\)∈C2\(ℝn\)J\_\{2\}\(\\cdot,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\)and satisfiesΔJ2\(x,i,𝛑,𝛄\(𝛑\)\)\>0\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)\>0, for anyx∈ℝnx\\in\\mathbb\{R\}^\{n\}andi∈ℳi\\in\\mathcal\{M\}, the feedback policy𝛄~\(𝛑\)\\tilde\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)defined by
𝜸~\(𝝅\)=\\displaystyle\\tilde\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)=\(4\.1\)𝒩\(β~,λ2\(2R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\(𝝅\)\)D2\(i\)\)−1\),\\displaystyle\\mathcal\{N\}\\left\(\\tilde\{\\beta\},\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\),where
β~\\displaystyle\\tilde\{\\beta\}=−\(2R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\(𝝅\)\)D2\(i\)\)−1\\displaystyle=\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(2θ2\(i\)R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\(𝝅\)\)D1\(i\)\)∫ℝpu𝝅\(u\)𝑑u\\displaystyle\\left\(2\\theta\_\{2\}\(i\)R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)D\_\{1\}\(i\)\\right\)\\int\_\{\\mathbb\{R\}^\{p\}\}u\\boldsymbol\{\\pi\}\(u\)\\,du−\(2R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\(𝝅\)\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(D2\(i\)TΔJ2\(x,i,𝝅,𝜸\(𝝅\)\)C\(i\)x\+B2\(i\)T∇J2\(x,i,𝝅,𝜸\(𝝅\)\)\)\.\\displaystyle\\quad\\left\(D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)C\(i\)x\+B\_\{2\}\(i\)^\{T\}\\nabla J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)\\right\)\.Then, the performance function under specific policies𝛄~\(𝛑\)\\tilde\{\{\\boldsymbol\{\\gamma\}\}\}\(\\boldsymbol\{\\pi\}\)and𝛄\(𝛑\)\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)satisfies
J2\(x,i,𝝅,𝜸~\(𝝅\)\)≤J2\(x,i,𝝅,𝜸\(𝝅\)\),∀x∈ℝn,i∈ℳ\.J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\tilde\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)\)\\leq J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\),\\,\\forall x\\in\\mathbb\{R\}^\{n\},i\\in\\mathcal\{M\}\.\(4\.2\)
The proof can be found in Appendix[E](https://arxiv.org/html/2606.28671#A5)\.
###### Theorem 5
\(Policy improvement theorem for playerII\)\. Let𝛄\\boldsymbol\{\\gamma\}be given and𝛑\\boldsymbol\{\\pi\}be the corresponding given admissible feedback policy\. Suppose that the corresponding performance functionJ1\(⋅,i,𝛑,𝛄\)∈C2\(ℝn\)J\_\{1\}\(\\cdot,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\)and satisfiesΔJ1\(x,i,𝛑,𝛄\)\>0\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\>0, for anyx∈ℝnx\\in\\mathbb\{R\}^\{n\}andi∈ℳi\\in\\mathcal\{M\}, the feedback policy𝛑~\\tilde\{\\boldsymbol\{\\pi\}\}defined by
𝝅~=\\displaystyle\\tilde\{\\boldsymbol\{\\pi\}\}=\(4\.3\)𝒩\(α~,λ1\(2R1\(i\)\+D1\(i\)TΔJ1\(x,i,𝝅,𝜸\)D1\(i\)\)−1\),\\displaystyle\\mathcal\{N\}\\left\(\\tilde\{\\alpha\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\),where
α~:=\\displaystyle\\tilde\{\\alpha\}=−Λ˙3−1\[B1\(i\)T∇J1\(x,i,𝝅,𝜸\)\\displaystyle\-\\dot\{\\Lambda\}\_\{3\}^\{\-1\}\\Big\[B\_\{1\}\(i\)^\{T\}\\nabla J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+Λ˙1TB2\(i\)T∇J1\(x,i,𝝅,𝜸\)\\displaystyle\\quad\+\\dot\{\\Lambda\}\_\{1\}^\{T\}B\_\{2\}\(i\)^\{T\}\\nabla J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\+D1\(i\)TΔJ1\(x,i,𝝅,𝜸\)C\(i\)x\\displaystyle\\quad\+D\_\{1\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)C\(i\)x\+Λ˙1TD2\(i\)TΔJ1\(x,i,𝝅,𝜸\)C\(i\)x\\displaystyle\\quad\+\\dot\{\\Lambda\}\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)C\(i\)x\+\(D1\(i\)TΔJ1\(x,i,𝝅,𝜸\)D2\(i\)\\displaystyle\\quad\+\\left\(D\_\{1\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\\right\.\+Λ˙1TD2\(i\)TΔJ1\(x,i,𝝅,𝜸\)D2\(i\)\\displaystyle\\quad\+\\dot\{\\Lambda\}\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\+2θ1\(i\)R1\(i\)\+2θ1\(i\)2Λ˙1TR1\(i\)\)Λ˙2\],\\displaystyle\\quad\\left\.\+2\\theta\_\{1\}\(i\)R\_\{1\}\(i\)\+2\\theta\_\{1\}\(i\)^\{2\}\\dot\{\\Lambda\}\_\{1\}^\{T\}R\_\{1\}\(i\)\\right\)\\dot\{\\Lambda\}\_\{2\}\\Big\],Λ˙1:=\\displaystyle\\dot\{\\Lambda\}\_\{1\}=−\(2R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(2θ2\(i\)R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\)D1\(i\)\),\\displaystyle\\left\(2\\theta\_\{2\}\(i\)R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{1\}\(i\)\\right\),Λ˙2:=\\displaystyle\\dot\{\\Lambda\}\_\{2\}=−\(2R2\(i\)\+D2\(i\)TΔJ2\(x,i,𝝅,𝜸\)D2\(i\)\)−1\\displaystyle\-\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\\right\)^\{\-1\}\(D2\(i\)TΔJ2\(x,i,𝝅,𝜸\)C\(i\)x\+B2\(i\)T∇J2\(x,i,𝝅,𝜸\)\),\\displaystyle\\left\(D\_\{2\}\(i\)^\{T\}\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)C\(i\)x\+B\_\{2\}\(i\)^\{T\}\\nabla J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\right\),Λ˙3:=\\displaystyle\\dot\{\\Lambda\}\_\{3\}=2R1\(i\)\+D1\(i\)TΔJ1\(x,i,𝝅,𝜸\)D1\(i\)\\displaystyle 2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{1\}\(i\)\+D1\(i\)TΔJ1\(x,i,𝝅,𝜸\)D2\(i\)Λ˙1\+2θ1\(i\)Λ˙1TR1\(i\)\\displaystyle\+D\_\{1\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\\dot\{\\Lambda\}\_\{1\}\+2\\theta\_\{1\}\(i\)\\dot\{\\Lambda\}\_\{1\}^\{T\}R\_\{1\}\(i\)\+Λ˙1TD2\(i\)TΔJ1\(x,i,𝝅,𝜸\)D2\(i\)Λ˙1\+2θ1\(i\)R1\(i\)Λ˙1\\displaystyle\+\\dot\{\\Lambda\}\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{2\}\(i\)\\dot\{\\Lambda\}\_\{1\}\+2\\theta\_\{1\}\(i\)R\_\{1\}\(i\)\\dot\{\\Lambda\}\_\{1\}\+2θ1\(i\)2Λ˙1TR1\(i\)Λ˙1\+Λ˙1TD2\(i\)TΔJ1\(x,i,𝝅,𝜸\)D1\(i\)\.\\displaystyle\+2\\theta\_\{1\}\(i\)^\{2\}\\dot\{\\Lambda\}\_\{1\}^\{T\}R\_\{1\}\(i\)\\dot\{\\Lambda\}\_\{1\}\+\\dot\{\\Lambda\}\_\{1\}^\{T\}D\_\{2\}\(i\)^\{T\}\\Delta J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)D\_\{1\}\(i\)\.Then the performance function under specific policies𝛑~\\tilde\{\{\\boldsymbol\{\\pi\}\}\}and𝛑\{\\boldsymbol\{\\pi\}\}satisfies
J1\(x,i,𝝅~,𝜸\)≤J1\(x,i,𝝅,𝜸\),∀x∈ℝn,i∈ℳ\.J\_\{1\}\(x,i,\\tilde\{\\boldsymbol\{\\pi\}\},\\boldsymbol\{\\gamma\}\)\\leq J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\),\\,\\forall x\\in\\mathbb\{R\}^\{n\},i\\in\\mathcal\{M\}\.\(4\.4\)
###### Proof 3
The proof follows a similar procedure to that of Theorem[4](https://arxiv.org/html/2606.28671#Thmtheorem4)and is therefore omitted for brevity\.
### IV\-BAlgorithm design
We have developed a PIA in the game setting with regime switching\. The main algorithm consists of two parallel steps: policy improvement and policy evaluation\. For policy improvement, Theorem[4](https://arxiv.org/html/2606.28671#Thmtheorem4)and Theorem[5](https://arxiv.org/html/2606.28671#Thmtheorem5)ensure convergence\. For policy evaluation, we compute the value functions of playerIIand playerIIIIrespectively under any admissible policy pair\(𝝅,𝜸\)\(\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)for all regimesi∈ℳi\\in\\mathcal\{M\}\. The algorithm mainly consists of four steps:
*Step 1 \(Initial value\):*For alli∈ℳi\\in\\mathcal\{M\}, initializeV1\(0\)\(x,i\)V\_\{1\}^\{\(0\)\}\(x,i\),V2\(0\)\(x,i\)V\_\{2\}^\{\(0\)\}\(x,i\),α\(0\)\\alpha^\{\(0\)\},β\(0\)\\beta^\{\(0\)\}\. Then, the initial policy pair\(𝝅\(0\),𝜸\(0\)\)\(\\boldsymbol\{\\pi\}^\{\(0\)\},\\boldsymbol\{\\gamma\}^\{\(0\)\}\)is
𝒩\(α\(0\),λ1\(2R1\(i\)\+D1\(i\)TΔV1\(0\)\(x,i\)D1\(i\)\)−1\),\\displaystyle\\mathcal\{N\}\\left\(\\alpha^\{\(0\)\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta V^\{\(0\)\}\_\{1\}\(x,i\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\),𝒩\(β\(0\),λ2\(2R2\(i\)\+D2\(i\)TΔV2\(0\)\(x,i\)D2\(i\)\)−1\)\.\\displaystyle\\mathcal\{N\}\\left\(\\beta^\{\(0\)\},\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta V^\{\(0\)\}\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\)\.Lets=0s=0andϵ\\epsilonbe a small constant\.
*Step 2 \(Iteration for the leader\):*For eachi∈ℳi\\in\\mathcal\{M\},
- •Value function update step: V1\(s\+1\)\(x,i\)\\displaystyle V\_\{1\}^\{\(s\+1\)\}\(x,i\)=1ρ1ℋ1\(x,i,∇V1\(s\)\(x,i\),ΔV1\(s\)\(x,i\),𝝅\(s\),𝜸\(s\)\)\.\\displaystyle=\\frac\{1\}\{\\rho\_\{1\}\}\\mathcal\{H\}\_\{1\}\(x,i,\\nabla V^\{\(s\)\}\_\{1\}\(x,i\),\\Delta V^\{\(s\)\}\_\{1\}\(x,i\),\\boldsymbol\{\\pi\}^\{\(s\)\},\\boldsymbol\{\\gamma\}^\{\(s\)\}\)\.
- •Policy update step: 𝝅\(s\+1\)=\\displaystyle\\boldsymbol\{\\pi\}^\{\(s\+1\)\}=𝒩\(α\(s\+1\),λ1\(2R1\(i\)\+D1\(i\)TΔV1\(s\+1\)\(x,i\)D1\(i\)\)−1\),\\displaystyle\\mathcal\{N\}\\left\(\\alpha^\{\(s\+1\)\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta V^\{\(s\+1\)\}\_\{1\}\(x,i\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\),whereα\(s\+1\)\\alpha^\{\(s\+1\)\}is updated exactly as derived in Section[III](https://arxiv.org/html/2606.28671#S3), using the known matricesB1\(i\),B2\(i\),D1\(i\),D2\(i\)B\_\{1\}\(i\),B\_\{2\}\(i\),D\_\{1\}\(i\),D\_\{2\}\(i\)and the value function∇V1\(s\+1\)\(x,i\)\\nabla V\_\{1\}^\{\(s\+1\)\}\(x,i\),ΔV1\(s\+1\)\(x,i\)\\Delta V\_\{1\}^\{\(s\+1\)\}\(x,i\),∇V2\(s\)\(x,i\),ΔV2\(s\)\(x,i\)\\nabla V\_\{2\}^\{\(s\)\}\(x,i\),\\Delta V\_\{2\}^\{\(s\)\}\(x,i\)\.
*Step 3 \(Iteration for the follower\):*For eachi∈ℳi\\in\\mathcal\{M\},
- •Value function update step: V2\(s\+1\)\(x,i\)\\displaystyle V\_\{2\}^\{\(s\+1\)\}\(x,i\)=1ρ2ℋ2\(x,i,∇V2\(s\)\(x,i\),ΔV2\(s\)\(x,i\),𝝅\(s\+1\),𝜸\(s\)\)\.\\displaystyle=\\frac\{1\}\{\\rho\_\{2\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla V^\{\(s\)\}\_\{2\}\(x,i\),\\Delta V^\{\(s\)\}\_\{2\}\(x,i\),\\boldsymbol\{\\pi\}^\{\(s\+1\)\},\\boldsymbol\{\\gamma\}^\{\(s\)\}\)\.
- •Policy update step: 𝜸\(s\+1\)=\\displaystyle\\boldsymbol\{\\gamma\}^\{\(s\+1\)\}=𝒩\(β\(s\+1\),λ2\(2R2\(i\)\+D2\(i\)TΔV2\(s\+1\)\(x,i\)D2\(i\)\)−1\),\\displaystyle\\mathcal\{N\}\\left\(\\beta^\{\(s\+1\)\},\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta V^\{\(s\+1\)\}\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\),whereβ\(s\+1\)\\beta^\{\(s\+1\)\}is updated according to the gradients ofV1\(s\+1\)\(x,i\)V\_\{1\}^\{\(s\+1\)\}\(x,i\),V2\(s\+1\)\(x,i\)V\_\{2\}^\{\(s\+1\)\}\(x,i\)as previously established\.
*Step 4 \(Convergence criterion\):*Ifmaxi∈ℳ\|Vk\(s\+1\)\(x,i\)−Vk\(s\)\(x,i\)\|≤ϵ\\max\_\{i\\in\\mathcal\{M\}\}\|V\_\{k\}^\{\(s\+1\)\}\(x,i\)\-V\_\{k\}^\{\(s\)\}\(x,i\)\|\\leq\\epsilonfork=1,2k=1,2, stop and output
\{𝝅∗=𝝅\(s\+1\),𝜸∗=𝜸\(s\+1\),V1\(x,i\)=V1\(s\+1\)\(x,i\),V2\(x,i\)=V2\(s\+1\)\(x,i\)\.\\left\\\{\\begin\{array\}\[\]\{ll\}\\boldsymbol\{\\pi\}^\{\*\}=\\boldsymbol\{\\pi\}^\{\(s\+1\)\},\\,\\boldsymbol\{\\gamma\}^\{\*\}=\\boldsymbol\{\\gamma\}^\{\(s\+1\)\},\\\\ V\_\{1\}\(x,i\)=V\_\{1\}^\{\(s\+1\)\}\(x,i\),\\\\ V\_\{2\}\(x,i\)=V\_\{2\}^\{\(s\+1\)\}\(x,i\)\.\\end\{array\}\\right\.Otherwise, lets=s\+1s=s\+1and go back to execute*Steps 2\-3*in sequence\.
We summarize the results into*Algorithm[1](https://arxiv.org/html/2606.28671#alg1)*\.
Algorithm 1Policy improvement algorithmStep 1 \(Initial value\):For all
i∈ℳi\\in\\mathcal\{M\}, initialize
V1\(0\)\(x,i\)V\_\{1\}^\{\(0\)\}\(x,i\),
V2\(0\)\(x,i\)V\_\{2\}^\{\(0\)\}\(x,i\),
α\(0\)\\alpha^\{\(0\)\},
β\(0\)\\beta^\{\(0\)\}\. Then get the initial policy pair
\(𝝅\(0\),𝜸\(0\)\)\(\\boldsymbol\{\\pi\}^\{\(0\)\},\\boldsymbol\{\\gamma\}^\{\(0\)\}\)\. Let
s=0s=0and
ϵ\\epsilonbe a small constant\.
Step 2 \(Iteration for the leader\):For
i∈ℳi\\in\\mathcal\{M\}:
- •Update value functionV1\(s\+1\)\(x,i\)V\_\{1\}^\{\(s\+1\)\}\(x,i\)\.
- •UpdateΛ1\(s\)\(i\)\\Lambda\_\{1\}^\{\(s\)\}\(i\),Λ2\(s\)\(i\)\\Lambda\_\{2\}^\{\(s\)\}\(i\),Λ3\(s\)\(i\)\\Lambda\_\{3\}^\{\(s\)\}\(i\),α\(s\+1\)\\alpha^\{\(s\+1\)\}to obtain policy𝝅\(s\+1\)\\boldsymbol\{\\pi\}^\{\(s\+1\)\}\.
Step 3 \(Iteration for the follower\):For
i∈ℳi\\in\\mathcal\{M\}:
- •Update value functionV2\(s\+1\)\(x,i\)V\_\{2\}^\{\(s\+1\)\}\(x,i\)\.
- •Updateβ\(s\+1\)\\beta^\{\(s\+1\)\}to obtain policy𝜸\(s\+1\)\\boldsymbol\{\\gamma\}^\{\(s\+1\)\}\.
Step 4 :If
maxi∈ℳ\|Vk\(s\+1\)\(x,i\)−Vk\(s\)\(x,i\)\|≤ϵ\\max\_\{i\\in\\mathcal\{M\}\}\|V\_\{k\}^\{\(s\+1\)\}\(x,i\)\-V\_\{k\}^\{\(s\)\}\(x,i\)\|\\leq\\epsilonfor
k=1,2k=1,2, stop and output the optimal policies and value functions\. Otherwise, let
s=s\+1s=s\+1and go back to executeSteps 2\-3in sequence\.
### IV\-CImplement the algorithm using critic NNs
The implementation process of*Algorithm[1](https://arxiv.org/html/2606.28671#alg1)*faces significant challenges in practical applications\. Traditional numerical solution methods typically require discretizing the continuous state space, a process highly susceptible to the “curse of dimensionality”—specifically, as the dimensionality of the system state increases, the required computational effort grows exponentially\. To better address the challenges of high\-dimensionality, a promising direction is to regard artificial neural networks as a more flexible and efficient function approximation tool\[[34](https://arxiv.org/html/2606.28671#bib.bib46),[9](https://arxiv.org/html/2606.28671#bib.bib44)\]\. In this subsection, we propose a critic NNs algorithm to implement the policy evaluation stage using the least\-squares \(LS\) method\.
Since the regimei∈ℳ=\{1,…,m\}i\\in\\mathcal\{M\}=\\\{1,\\dots,m\\\}takes discrete values, we can approximate the value function for each regime using separate sets of weights\. According to the Weierstrass high\-order approximation theorem, the smooth value functionVk\(x,i\)V\_\{k\}\(x,i\)fork=1,2k=1,2can be approximated using the critic NNs as follows
V¯k\(x,i\)=\(Wk\(i\)\)Tφk\(x\),\\overline\{V\}\_\{k\}\(x,i\)=\(W\_\{k\}\(i\)\)^\{T\}\\varphi\_\{k\}\(x\),\(4\.5\)whereφk\(x\)∈ℝhk\\varphi\_\{k\}\(x\)\\in\\mathbb\{R\}^\{h\_\{k\}\}is the critic NNs activation function vector withhkh\_\{k\}being the number of neurons in the critic NNs hidden layer, andWk\(i\)∈ℝhkW\_\{k\}\(i\)\\in\\mathbb\{R\}^\{h\_\{k\}\}is a constant weight vector corresponding to regimeii\. The corresponding gradient and Hessian matrix with respect to statexxcan be obtained as
∇V¯k\(x,i\)=\(Wk\(i\)\)T∂φk\(x\)∂x,\\nabla\\overline\{V\}\_\{k\}\(x,i\)=\(W\_\{k\}\(i\)\)^\{T\}\\frac\{\\partial\\varphi\_\{k\}\(x\)\}\{\\partial x\},ΔV¯k\(x,i\)=\(Wk\(i\)\)T∂2φk\(x\)∂x2,\\Delta\\overline\{V\}\_\{k\}\(x,i\)=\(W\_\{k\}\(i\)\)^\{T\}\\frac\{\\partial^\{2\}\\varphi\_\{k\}\(x\)\}\{\\partial x^\{2\}\},where∂φk\(x\)/∂x\{\\partial\\varphi\_\{k\}\(x\)\}/\{\\partial x\}is a two\-dimensional tensor ofhk×nh\_\{k\}\\times n, and∂2φk\(x\)/∂x2\{\\partial^\{2\}\\varphi\_\{k\}\(x\)\}/\{\\partial x^\{2\}\}is a three\-dimensional tensor ofhk×n×nh\_\{k\}\\times n\\times n, fork=1,2k=1,2\.
Substitute the approximate formV¯k\(x,i\)\\overline\{V\}\_\{k\}\(x,i\)ofVk\(x,i\)V\_\{k\}\(x,i\)in equation \([4\.5](https://arxiv.org/html/2606.28671#S4.E5)\) into*Algorithm[1](https://arxiv.org/html/2606.28671#alg1)*, so thatΛ1\(s\)\\Lambda\_\{1\}^\{\(s\)\},Λ2\(s\)\\Lambda\_\{2\}^\{\(s\)\},Λ3\(s\)\\Lambda\_\{3\}^\{\(s\)\},α\(s\+1\)\\alpha^\{\(s\+1\)\},𝝅\(s\+1\)\\boldsymbol\{\\pi\}^\{\(s\+1\)\}andβ\(s\+1\)\\beta^\{\(s\+1\)\},𝜸\(s\+1\)\\boldsymbol\{\\gamma\}^\{\(s\+1\)\}in the*Policy update step*can be updated\. Subsequently, the*Value function update step*becomes a key point of concern\. For each regimei∈ℳi\\in\\mathcal\{M\}, the corresponding residual errors for playerIIand playerIIIIare
e~1\(s\)\(i\)\\displaystyle\\tilde\{e\}\_\{1\}^\{\(s\)\}\(i\)=\(W1\(s\+1\)\(i\)\)Tφ1\(x\)\\displaystyle=\(W\_\{1\}^\{\(s\+1\)\}\(i\)\)^\{T\}\\varphi\_\{1\}\(x\)−1ρ1ℋ1\(x,i,∇V¯1\(s\)\(x,i\),ΔV¯1\(s\)\(x,i\),𝝅\(s\),𝜸\(s\)\),\\displaystyle\-\\frac\{1\}\{\\rho\_\{1\}\}\\mathcal\{H\}\_\{1\}\(x,i,\\nabla\\overline\{V\}^\{\(s\)\}\_\{1\}\(x,i\),\\Delta\\overline\{V\}^\{\(s\)\}\_\{1\}\(x,i\),\\boldsymbol\{\\pi\}^\{\(s\)\},\\boldsymbol\{\\gamma\}^\{\(s\)\}\),e~2\(s\)\(i\)\\displaystyle\\tilde\{e\}\_\{2\}^\{\(s\)\}\(i\)=\(W2\(s\+1\)\(i\)\)Tφ2\(x\)\\displaystyle=\(W\_\{2\}^\{\(s\+1\)\}\(i\)\)^\{T\}\\varphi\_\{2\}\(x\)−1ρ2ℋ2\(x,i,∇V¯2\(s\)\(x,i\),ΔV¯2\(s\)\(x,i\),𝝅\(s\+1\),𝜸\(s\)\)\.\\displaystyle\-\\frac\{1\}\{\\rho\_\{2\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla\\overline\{V\}^\{\(s\)\}\_\{2\}\(x,i\),\\Delta\\overline\{V\}^\{\(s\)\}\_\{2\}\(x,i\),\\boldsymbol\{\\pi\}^\{\(s\+1\)\},\\boldsymbol\{\\gamma\}^\{\(s\)\}\)\.
Now our objective is to find the appropriate weight vectorWk\(s\+1\)\(i\)W\_\{k\}^\{\(s\+1\)\}\(i\)such that the residual errore~k\(s\)\(i\)\\tilde\{e\}\_\{k\}^\{\(s\)\}\(i\)approaches zero\. In fact, we can regard this problem as a LS problem, whereWk\(s\+1\)\(i\)W\_\{k\}^\{\(s\+1\)\}\(i\)is the unknown parameter,φk\(x\)\\varphi\_\{k\}\(x\)is the regression vector, and1ρkℋk\\frac\{1\}\{\\rho\_\{k\}\}\\mathcal\{H\}\_\{k\}is a known parameter calculated using the observed system data\. After collecting the system data atMMpoints \(M\>hkM\>h\_\{k\}\), the update law of the critic NNs for each regimeiiis
Wk\(s\+1\)\(i\)=\(ξk\(s\)\(ξk\(s\)\)T\)−1ξk\(s\)χk\(s\)\(i\),W\_\{k\}^\{\(s\+1\)\}\(i\)=\\left\(\\xi\_\{k\}^\{\(s\)\}\(\\xi\_\{k\}^\{\(s\)\}\)^\{T\}\\right\)^\{\-1\}\\xi\_\{k\}^\{\(s\)\}\\chi\_\{k\}^\{\(s\)\}\(i\),\(4\.6\)whereξk\(s\)=\[φk\(xt1\)⋯φk\(xtM\)\]∈ℝhk×M\\xi\_\{k\}^\{\(s\)\}=\\begin\{bmatrix\}\\varphi\_\{k\}\(x\_\{t\_\{1\}\}\)&\\cdots&\\varphi\_\{k\}\(x\_\{t\_\{M\}\}\)\\end\{bmatrix\}\\in\\mathbb\{R\}^\{h\_\{k\}\\times M\}andχk\(s\)\(i\)=1ρk\[ℋk\(t1,i\)⋯ℋk\(tM,i\)\]T∈ℝM×1\\chi\_\{k\}^\{\(s\)\}\(i\)=\\frac\{1\}\{\\rho\_\{k\}\}\\begin\{bmatrix\}\\mathcal\{H\}\_\{k\}\(t\_\{1\},i\)&\\cdots&\\mathcal\{H\}\_\{k\}\(t\_\{M\},i\)\\end\{bmatrix\}^\{T\}\\in\\mathbb\{R\}^\{M\\times 1\}, fork=1,2k=1,2\.
We summarize the results into*Algorithm[2](https://arxiv.org/html/2606.28671#alg2)*\.
Algorithm 2Critic neural networks algorithmStep 1 \(Initial value\):For all
i∈ℳi\\in\\mathcal\{M\}, initialize
W1\(0\)\(i\)W\_\{1\}^\{\(0\)\}\(i\),
W2\(0\)\(i\)W\_\{2\}^\{\(0\)\}\(i\),
α\(0\)\\alpha^\{\(0\)\},
β\(0\)\\beta^\{\(0\)\}\. Collect system data at
MMpoints\. Select appropriate activation functions
φ1\\varphi\_\{1\}and
φ2\\varphi\_\{2\}\. Then,
ΔV¯k\(x,i\)=\(Wk\(i\)\)T∂2φk\(x\)∂x2\\Delta\\overline\{V\}\_\{k\}\(x,i\)=\(W\_\{k\}\(i\)\)^\{T\}\\frac\{\\partial^\{2\}\\varphi\_\{k\}\(x\)\}\{\\partial x^\{2\}\}, for
k=1,2k=1,2\. The initial policy pair is
𝝅\(0\)=𝒩\(α\(0\),λ1\(2R1\(i\)\+D1\(i\)TΔV¯1\(x,i\)D1\(i\)\)−1\),\\displaystyle\\boldsymbol\{\\pi\}^\{\(0\)\}=\\mathcal\{N\}\\left\(\\alpha^\{\(0\)\},\\lambda\_\{1\}\\left\(2R\_\{1\}\(i\)\+D\_\{1\}\(i\)^\{T\}\\Delta\\overline\{V\}\_\{1\}\(x,i\)D\_\{1\}\(i\)\\right\)^\{\-1\}\\right\),𝜸\(0\)=𝒩\(β\(0\),λ2\(2R2\(i\)\+D2\(i\)TΔV¯2\(x,i\)D2\(i\)\)−1\)\.\\displaystyle\\boldsymbol\{\\gamma\}^\{\(0\)\}=\\mathcal\{N\}\\left\(\\beta^\{\(0\)\},\\lambda\_\{2\}\\left\(2R\_\{2\}\(i\)\+D\_\{2\}\(i\)^\{T\}\\Delta\\overline\{V\}\_\{2\}\(x,i\)D\_\{2\}\(i\)\\right\)^\{\-1\}\\right\)\.Let
s=0s=0and
ϵ\\epsilonbe a small constant\.
Step 2 \(Iteration for the leader\):For
i∈ℳi\\in\\mathcal\{M\}:
- •Update the weightW1\(s\+1\)\(i\)W\_\{1\}^\{\(s\+1\)\}\(i\)according to \([4\.6](https://arxiv.org/html/2606.28671#S4.E6)\)\.
- •UpdateΛ1\(s\)\\Lambda\_\{1\}^\{\(s\)\},Λ2\(s\)\\Lambda\_\{2\}^\{\(s\)\},Λ3\(s\)\\Lambda\_\{3\}^\{\(s\)\},α\(s\+1\)\\alpha^\{\(s\+1\)\}to obtain policy𝝅\(s\+1\)\\boldsymbol\{\\pi\}^\{\(s\+1\)\}\.
Step 3 \(Iteration for the follower\):For
i∈ℳi\\in\\mathcal\{M\}:
- •Update the weightW2\(s\+1\)\(i\)W\_\{2\}^\{\(s\+1\)\}\(i\)according to \([4\.6](https://arxiv.org/html/2606.28671#S4.E6)\)\.
- •Updateβ\(s\+1\)\\beta^\{\(s\+1\)\}to obtain policy𝜸\(s\+1\)\\boldsymbol\{\\gamma\}^\{\(s\+1\)\}\.
Step 4 \(Convergence criterion\):If
maxi∈ℳ‖Wk\(s\+1\)\(i\)−Wk\(s\)\(i\)‖≤ϵ\\max\_\{i\\in\\mathcal\{M\}\}\|\|W\_\{k\}^\{\(s\+1\)\}\(i\)\-W\_\{k\}^\{\(s\)\}\(i\)\|\|\\leq\\epsilonfor
k=1,2k=1,2, stop and output
\{𝝅∗=𝝅\(s\+1\),𝜸∗=𝜸\(s\+1\),V1\(x,i\)=\(W1\(s\+1\)\(i\)\)Tφ1\(x\),V2\(x,i\)=\(W2\(s\+1\)\(i\)\)Tφ2\(x\)\.\\left\\\{\\begin\{array\}\[\]\{ll\}\\boldsymbol\{\\pi\}^\{\*\}=\\boldsymbol\{\\pi\}^\{\(s\+1\)\},\\,\\boldsymbol\{\\gamma\}^\{\*\}=\\boldsymbol\{\\gamma\}^\{\(s\+1\)\},\\\\ V\_\{1\}\(x,i\)=\(W\_\{1\}^\{\(s\+1\)\}\(i\)\)^\{T\}\\varphi\_\{1\}\(x\),\\\\ V\_\{2\}\(x,i\)=\(W\_\{2\}^\{\(s\+1\)\}\(i\)\)^\{T\}\\varphi\_\{2\}\(x\)\.\\end\{array\}\\right\.Otherwise, let
s=s\+1s=s\+1and go back to executeSteps 2\-3in sequence\.
## VSimulation results
In this section, we provide a simulation example to illustrate the effectiveness of the proposed critic NNs algorithm, and use this algorithm to analyze some characteristics of the two\-player Stackelberg game under regime switching\.
Suppose the Markov chain has two regimes, i\.e\.,ℳ=\{1,2\}\\mathcal\{M\}=\\\{1,2\\\}\. The parameters in the regime\-switching SDE are selected as follows:
B1\(1\)=\[0\.30\.8\],B2\(1\)=\[0\.80\.2\],D1\(1\)=\[0\.30\.8\],B\_\{1\}\(1\)=\\begin\{bmatrix\}0\.3\\\\ 0\.8\\end\{bmatrix\},\\,B\_\{2\}\(1\)=\\begin\{bmatrix\}0\.8\\\\ 0\.2\\end\{bmatrix\},D\_\{1\}\(1\)=\\begin\{bmatrix\}0\.3\\\\ 0\.8\\end\{bmatrix\},\\,D2\(1\)=\[−0\.80\.2\],A\(1\)=\[010−1\],C\(1\)=\[100−1\];D\_\{2\}\(1\)=\\begin\{bmatrix\}\-0\.8\\\\ 0\.2\\end\{bmatrix\},\\,A\(1\)=\\begin\{bmatrix\}0&1\\\\ 0&\-1\\end\{bmatrix\},\\,C\(1\)=\\begin\{bmatrix\}1&0\\\\ 0&\-1\\end\{bmatrix\};
B1\(2\)=\[0\.40\.7\],B2\(2\)=\[0\.60\.3\],D1\(2\)=\[0\.40\.6\],B\_\{1\}\(2\)=\\begin\{bmatrix\}0\.4\\\\ 0\.7\\end\{bmatrix\},\\,B\_\{2\}\(2\)=\\begin\{bmatrix\}0\.6\\\\ 0\.3\\end\{bmatrix\},\\,D\_\{1\}\(2\)=\\begin\{bmatrix\}0\.4\\\\ 0\.6\\end\{bmatrix\},D2\(2\)=\[−0\.60\.4\],A\(2\)=\[00\.80−1\.2\],C\(2\)=\[0\.800−0\.8\]\.D\_\{2\}\(2\)=\\begin\{bmatrix\}\-0\.6\\\\ 0\.4\\end\{bmatrix\},\\,A\(2\)=\\begin\{bmatrix\}0&0\.8\\\\ 0&\-1\.2\\end\{bmatrix\},\\,C\(2\)=\\begin\{bmatrix\}0\.8&0\\\\ 0&\-0\.8\\end\{bmatrix\}\.
The parameters in the performance functions for each regime are selected as follows:
Q1\(1\)=\[0\.1000\.1\],Q2\(1\)=\[0\.5000\.5\],Q\_\{1\}\(1\)=\\begin\{bmatrix\}0\.1&0\\\\ 0&0\.1\\end\{bmatrix\},\\,Q\_\{2\}\(1\)=\\begin\{bmatrix\}0\.5&0\\\\ 0&0\.5\\end\{bmatrix\},R1\(1\)=5,R2\(1\)=8,θ1\(1\)=0\.4,θ2\(1\)=0\.1;R\_\{1\}\(1\)=5,\\,R\_\{2\}\(1\)=8,\\,\\theta\_\{1\}\(1\)=0\.4,\\,\\theta\_\{2\}\(1\)=0\.1;
Q1\(2\)=\[0\.2000\.2\],Q2\(2\)=\[0\.4000\.4\],Q\_\{1\}\(2\)=\\begin\{bmatrix\}0\.2&0\\\\ 0&0\.2\\end\{bmatrix\},\\,Q\_\{2\}\(2\)=\\begin\{bmatrix\}0\.4&0\\\\ 0&0\.4\\end\{bmatrix\},R1\(2\)=6,R2\(2\)=7,θ1\(2\)=0\.3,θ2\(2\)=0\.2\.R\_\{1\}\(2\)=6,\\,R\_\{2\}\(2\)=7,\\,\\theta\_\{1\}\(2\)=0\.3,\\,\\theta\_\{2\}\(2\)=0\.2\.
The generator matrix of the Markov chain is
Γ=\[−0\.50\.50\.4−0\.4\]\.\\Gamma=\\begin\{bmatrix\}\-0\.5&0\.5\\\\ 0\.4&\-0\.4\\end\{bmatrix\}\.
Other parameters areρ1=3,ρ2=3,λ1=3,λ2=4\\rho\_\{1\}=3,\\rho\_\{2\}=3,\\lambda\_\{1\}=3,\\lambda\_\{2\}=4\. The initial state isx=\[33\]Tx=\\begin\{bmatrix\}3&3\\end\{bmatrix\}^\{T\}, and the initial regime isα0=1\\alpha\_\{0\}=1\. The algorithm starts fromWk\(0\)\(i\)=\[1111\]TW\_\{k\}^\{\(0\)\}\(i\)=\\begin\{bmatrix\}1&1&1&1\\end\{bmatrix\}^\{T\},α\(0\)=0\\alpha^\{\(0\)\}=0, andβ\(0\)=0\\beta^\{\(0\)\}=0fork∈\{1,2\}k\\in\\\{1,2\\\}andi∈\{1,2\}i\\in\\\{1,2\\\}\. Set the convergence criterion asϵ=10−3\\epsilon=10^\{\-3\}\. The activation function is selected as
φ1\(x\)=φ2\(x\)=\[x12x1x2x221\]T\.\\varphi\_\{1\}\(x\)=\\varphi\_\{2\}\(x\)=\\begin\{bmatrix\}x\_\{1\}^\{2\}&x\_\{1\}x\_\{2\}&x\_\{2\}^\{2\}&1\\end\{bmatrix\}^\{T\}\.\(5\.1\)
### V\-AThe convergence of the algorithm
We first implement the critic NNs*Algorithm[2](https://arxiv.org/html/2606.28671#alg2)*\. The weights of the critic NNs are updated after collectingM=50M=50points\. Each point is sampled every0\.010\.01seconds, and the66th point is selected as the observation point\. Fig\.[1](https://arxiv.org/html/2606.28671#S5.F1)illustrates the convergence of the weight vectors in the critic NNs for both the leader and the follower under different regimes\. As depicted, both networks successfully meet the convergence criteria after8181iterations\.
Figure 1:Convergence of the weight vectors\.The loss function is crucial for evaluating the performance of the algorithm\. We define the loss function using the Kullback\-Leibler \(KL\) divergence:
L\(μ,Σ\)=\\displaystyle L\(\\mu,\\Sigma\)=12\(tr\(Σ−1Σ∗\)\+\(μ−μ∗\)TΣ−1\(μ−μ∗\)\\displaystyle\\frac\{1\}\{2\}\\Big\(tr\(\\Sigma^\{\-1\}\\Sigma^\{\*\}\)\+\(\\mu\-\\mu^\{\*\}\)^\{T\}\\Sigma^\{\-1\}\(\\mu\-\\mu^\{\*\}\)−3\+lndet\(Σ∗\)det\(Σ\)\),\\displaystyle\-3\+\\ln\\frac\{\\det\(\\Sigma^\{\*\}\)\}\{\\det\(\\Sigma\)\}\\Big\),whereμ\\muandΣ\\Sigmaare the mean and covariance matrix of the current distribution, andμ∗\\mu^\{\*\}andΣ∗\\Sigma^\{\*\}are the mean and covariance matrices of the target distribution\. As shown in Fig\.[2](https://arxiv.org/html/2606.28671#S5.F2), the value of the loss function drops below10−710^\{\-7\}after a finite number of iterations for both regimes, indicating stable convergence\.
Figure 2:Convergence of the KL loss function\.
### V\-BImpact of temperature parameter on value function
Figure[3](https://arxiv.org/html/2606.28671#S5.F3)illustrates the regime\-dependent value functionsVk\(x,i\)V\_\{k\}\(x,i\)for both the leader \(k=1k=1\) and the follower \(k=2k=2\) under various temperature parameter configurations, specificallyλ1,λ2∈\{\(1,1\),\(4,4\),\(6,6\)\}\\lambda\_\{1\},\\lambda\_\{2\}\\in\\\{\(1,1\),\(4,4\),\(6,6\)\\\}\.
A primary observation from the numerical results is that the amplitude of the value functions exhibits a monotonic decrease as the temperature parameterλ\\lambdaincreases\. This phenomenon is consistent across both regimes\. Within the proposed framework, where negative entropy is incorporated as an auxiliary reward to promote exploration, a largerλ\\lambdaeffectively heightens the agents’ ability to traverse the state space and adapt to environmental transitions\. This enhanced exploration facilitates the discovery of more globally optimal policies, which manifests as a reduction in the magnitude of the value surfaces\. Furthermore, a comparative analysis of the two agents reveals a distinct disparity in sensitivity to the temperature parameter:
- •For the leader, the value functionV1\(x,i\)V\_\{1\}\(x,i\)shows significant vertical displacement asλ\\lambdavaries, indicating that the leader’s strategic value is highly sensitive to the exploration\-exploitation trade\-off\.
- •For the follower, the value functionV2\(x,i\)V\_\{2\}\(x,i\)exhibits much tighter clustering across the differentλ\\lambdasettings\.
This suggests that the follower’s value profile is relatively robust to changes in the temperature parameter, potentially due to its reactive role within the Stackelberg structure\.
Figure 3:Evolution of the leader and follower value surfaces under different temperature parameterλ\\lambda\.Fig\.[4](https://arxiv.org/html/2606.28671#S5.F4)plots the evolution of the leader’s value surfaceV1\(x\)V\_\{1\}\(x\)in Regime11while fixingλ2=4\\lambda\_\{2\}=4\. The results show that the magnitude ofV1\(x\)V\_\{1\}\(x\)decreases monotonically with increasingλ1\\lambda\_\{1\}\. As a key hyperparameter controlling exploration intensity, a largerλ1\\lambda\_\{1\}amplifies the negative entropy reward, enabling the leader to explore the full state space instead of local high\-reward regions\. This leads to the discovery of globally robust policies, evidenced by the downward shift in the value surface—though reducing the overall expected value, the strategy gains superior adaptability to dynamic environments\.
Figure 4:Evolution of the leader’s value surface under different temperature parametersλ\\lambdawith the follower’s temperature parameter fixed in regime11\.Figure[5](https://arxiv.org/html/2606.28671#S5.F5)illustrates a numerical comparison of the leader and follower value functions,V1\(x,i\)V\_\{1\}\(x,i\)andV2\(x,i\)V\_\{2\}\(x,i\), under two distinct regimes \(regime 1 and regime 2\) with the parameterλ=\(4,4\)\\lambda=\(4,4\)\. In terms of surface morphology, both value functions exhibit typical convex quadratic characteristics\. Numerical analysis reveals that the two regime mechanisms exert starkly opposing effects on the players\. For the leader, the value function surface of regime 2 is uniformly higher than that of regime 1, implying that the leader incurs a higher cost under regime 2\. Conversely, the follower’s value function under regime 1 is significantly higher than under regime 2\. Furthermore, its numerical magnitude \(with a peak approaching1010\) is much larger than that of the leader \(with a peak around3\.53\.5\), indicating that the follower is far more sensitive to fluctuations in the state variables\. This ”cross\-domination” phenomenon reveals the asymmetric impact of regime switching on the distribution of interests between the players, suggesting a fundamental drift in the structure of the game equilibrium across different states\.
Figure 5:omparison of leader and follower value functionsV1\(x\)V\_\{1\}\(x\)andV2\(x\)V\_\{2\}\(x\)across Regime 1 and Regime 2 withλ=\(4,4\)\\lambda=\(4,4\)\.
### V\-CImpact of temperature parameter on policy pair of players
As shown in Fig\.[6](https://arxiv.org/html/2606.28671#S5.F6), we examine the equilibrium policy pairs under the following temperature parameter combinations:λ1,λ2∈\{\(1,1\),\(4,4\),\(6,6\)\}\\lambda\_\{1\},\\lambda\_\{2\}\\in\\\{\(1,1\),\(4,4\),\(6,6\)\\\}\. The experimental results reveal the properties of the regime\-dependent equilibrium policies for a given active regimei∈ℳi\\in\\mathcal\{M\}:
- •Mean Stability: The mean of the optimal policy is highly robust to changes in temperature parameters within a specific regime, meaning that exploration intensity has little impact on the policy’s expected behavior, which is primarily dictated by the system’s current structural stateii\.
- •Variance Sensitivity: As the temperature parameter increases, the variance of the policy increases significantly, which reflects the exploration\-exploitation trade\-off in reinforcement learning—a higher temperature parameter promotes exploration by increasing policy entropy, manifested as a wider dispersion of the distribution to better anticipate and react to potential regime shifts\.
Figure 6:Equilibrium policies under varying temperature parameters\.
### V\-DEscaping Local Optima: An Illustrative Case Study
This subsection uses an example to reveal the limitation of classical models, which are prone to falling into local optima\. In contrast, the PIA algorithm adopted under our exploratory framework can effectively escape local optima and achieve the search for global optimal solutions\. For dynamic system \([2\.1](https://arxiv.org/html/2606.28671#S2.E1)\) and performance function
Jkcl\(x,u,v\)=∫0∞e−ρksrk\(xs,us,vs\)𝑑s,k=1,2\.J^\{cl\}\_\{k\}\(x,u,v\)=\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{k\}s\}r\_\{k\}\(x\_\{s\},u\_\{s\},v\_\{s\}\)\\,ds,\\,k=1,2\.\(5\.2\)Referring to the similar steps in Section[III](https://arxiv.org/html/2606.28671#S3), we can obtain the optimal policy pair is
\{u∗=θ1R2−1B2T∇V2cl\(x\)2\(1−θ1θ2\)−R1−1B1T∇V1cl\(x\)2\(1−θ1θ2\)2\+θ2R1−1B1T∇V1cl\(x\)2\(1−θ1θ2\)2,v∗=−θ2u∗−12R2−1B2T∇V2cl\(x\),\\left\\\{\\begin\{array\}\[\]\{ll\}u^\{\*\}=\\frac\{\\theta\_\{1\}R\_\{2\}^\{\-1\}B\_\{2\}^\{T\}\\nabla V^\{cl\}\_\{2\}\(x\)\}\{2\(1\-\\theta\_\{1\}\\theta\_\{2\}\)\}\-\\frac\{R\_\{1\}^\{\-1\}B\_\{1\}^\{T\}\\nabla V^\{cl\}\_\{1\}\(x\)\}\{2\(1\-\\theta\_\{1\}\\theta\_\{2\}\)^\{2\}\}\+\\frac\{\\theta\_\{2\}R\_\{1\}^\{\-1\}B\_\{1\}^\{T\}\\nabla V^\{cl\}\_\{1\}\(x\)\}\{2\(1\-\\theta\_\{1\}\\theta\_\{2\}\)^\{2\}\},\\\\ v^\{\*\}=\-\\theta\_\{2\}u^\{\*\}\-\\frac\{1\}\{2\}R\_\{2\}^\{\-1\}B^\{T\}\_\{2\}\\nabla V^\{cl\}\_\{2\}\(x\),\\end\{array\}\\right\.\(5\.3\)and the corresponding coupled HJB equations are
r2\(x,u∗,v∗\)\+∇V2cl\(x\)T\(Ax\+B1u∗\+B2v∗\)=ρ2V2cl\(x\),\\displaystyle r\_\{2\}\(x,u^\{\*\},v^\{\*\}\)\+\\nabla V^\{cl\}\_\{2\}\(x\)^\{T\}\(Ax\+B\_\{1\}u^\{\*\}\+B\_\{2\}v^\{\*\}\)=\\rho\_\{2\}V^\{cl\}\_\{2\}\(x\),\(5\.4\)r1\(x,u∗,v∗\)\+∇V1cl\(x\)T\(Ax\+B1u∗\+B2v∗\)=ρ1V1cl\(x\)\.\\displaystyle r\_\{1\}\(x,u^\{\*\},v^\{\*\}\)\+\\nabla V^\{cl\}\_\{1\}\(x\)^\{T\}\(Ax\+B\_\{1\}u^\{\*\}\+B\_\{2\}v^\{\*\}\)=\\rho\_\{1\}V^\{cl\}\_\{1\}\(x\)\.In the classical approach, we often suppose that the value function takes the form
Vkcl\(x\)=xTwkx,k=1,2\.V^\{cl\}\_\{k\}\(x\)=x^\{T\}w\_\{k\}x,\\,k=1,2\.\(5\.5\)Substituting \([5\.5](https://arxiv.org/html/2606.28671#S5.E5)\) into \([5\.3](https://arxiv.org/html/2606.28671#S5.E3)\) and \([5\.4](https://arxiv.org/html/2606.28671#S5.E4)\) yields the following coupled Riccati equations
\{ρ1w1=Q1\+u^TR1u^\+2θ1v^TR1u^\+θ12v^TR1v^\+2w1A\+2w1B1u^\+2w1B2v^,ρ2w2=Q2\+v^TR2v^\+2θ2u^TR2v^\+θ22u^TR2u^\+2w2A\+2w2B1u^\+2w2B2v^,\\left\\\{\\begin\{array\}\[\]\{ll\}\\rho\_\{1\}w\_\{1\}=Q\_\{1\}\+\\hat\{u\}^\{T\}R\_\{1\}\\hat\{u\}\+2\\theta\_\{1\}\\hat\{v\}^\{T\}R\_\{1\}\\hat\{u\}\+\\theta\_\{1\}^\{2\}\\hat\{v\}^\{T\}R\_\{1\}\\hat\{v\}\\\\ \\qquad\\quad\+2w\_\{1\}A\+2w\_\{1\}B\_\{1\}\\hat\{u\}\+2w\_\{1\}B\_\{2\}\\hat\{v\},\\\\ \\rho\_\{2\}w\_\{2\}=Q\_\{2\}\+\\hat\{v\}^\{T\}R\_\{2\}\\hat\{v\}\+2\\theta\_\{2\}\\hat\{u\}^\{T\}R\_\{2\}\\hat\{v\}\+\\theta\_\{2\}^\{2\}\\hat\{u\}^\{T\}R\_\{2\}\\hat\{u\}\\\\ \\qquad\\quad\+2w\_\{2\}A\+2w\_\{2\}B\_\{1\}\\hat\{u\}\+2w\_\{2\}B\_\{2\}\\hat\{v\},\\\\ \\end\{array\}\\right\.\(5\.6\)where
u^=\\displaystyle\\hat\{u\}=θ1R2−1B2Tw2\(1−θ1θ2\)−R1−1B1Tw1\(1−θ1θ2\)2\+θ2R1−1B1Tw1\(1−θ1θ2\)2,\\displaystyle\\frac\{\\theta\_\{1\}R\_\{2\}^\{\-1\}B\_\{2\}^\{T\}w\_\{2\}\}\{\(1\-\\theta\_\{1\}\\theta\_\{2\}\)\}\-\\frac\{R\_\{1\}^\{\-1\}B\_\{1\}^\{T\}w\_\{1\}\}\{\(1\-\\theta\_\{1\}\\theta\_\{2\}\)^\{2\}\}\+\\frac\{\\theta\_\{2\}R\_\{1\}^\{\-1\}B\_\{1\}^\{T\}w\_\{1\}\}\{\(1\-\\theta\_\{1\}\\theta\_\{2\}\)^\{2\}\},\(5\.7\)v^=\\displaystyle\\hat\{v\}=−θ2u^−R2−1B2Tw2\.\\displaystyle\-\\theta\_\{2\}\\hat\{u\}\-R\_\{2\}^\{\-1\}B^\{T\}\_\{2\}w\_\{2\}\.Equation \([5\.6](https://arxiv.org/html/2606.28671#S5.E6)\) is a system of nonlinear equations, and its solutions are often non\-unique\. Set
R1=\\displaystyle R\_\{1\}=R2=1\.0,B1=2\.0,B2=1\.0,\\displaystyle R\_\{2\}=0,\\quad B\_\{1\}=0,\\quad B\_\{2\}=0,\(5\.8\)θ1=\\displaystyle\\theta\_\{1\}=0\.6,θ2=0\.4,A=0\.1,\\displaystyle 6,\\quad\\theta\_\{2\}=4,\\quad A=1,Q1=\\displaystyle Q\_\{1\}=0\.2,Q2=0\.1,ρ1=1\.5,ρ2=1\.6\.\\displaystyle 2,\\quad Q\_\{2\}=1,\\quad\\rho\_\{1\}=5,\\quad\\rho\_\{2\}=6\.Equation \([5\.6](https://arxiv.org/html/2606.28671#S5.E6)\) degenerates to
\{1500361w12−1019w1w2\+1310w1−15=0,2400361w1w2−2919w22\+75w2−110=0\.\\left\\\{\\begin\{aligned\} &\\frac\{1500\}\{361\}w\_\{1\}^\{2\}\-\\frac\{10\}\{19\}w\_\{1\}w\_\{2\}\+\\frac\{13\}\{10\}w\_\{1\}\-\\frac\{1\}\{5\}=0,\\\\\[8\.0pt\] &\\frac\{2400\}\{361\}w\_\{1\}w\_\{2\}\-\\frac\{29\}\{19\}w\_\{2\}^\{2\}\+\\frac\{7\}\{5\}w\_\{2\}\-\\frac\{1\}\{10\}=0\.\\end\{aligned\}\\right\.\(5\.9\)In this case, there are two sets of positive solutions
\(w1,w2\)≈\\displaystyle\(w\_\{1\},w\_\{2\}\)\\approx\(0\.1143037,0\.0479209\),\\displaystyle\(1143037,0479209\),\(5\.10\)\(w1,w2\)≈\\displaystyle\(w\_\{1\},w\_\{2\}\)\\approx\(0\.1724645,1\.6282089\)\.\\displaystyle\(1724645,6282089\)\.It is obvious that the first set is the optimal solution, while the second set is not\. This also indicates that classical methods are prone to falling into local optima during the solving process\. At this point, the classical value function is
V1cl\(x\)=0\.1143037x2,V2cl\(x\)=0\.0479209x2\.V^\{cl\}\_\{1\}\(x\)=0\.1143037x^\{2\},\\quad V^\{cl\}\_\{2\}\(x\)=0\.0479209x^\{2\}\.\(5\.11\)
In contrast, we use the PIA method to solve this problem within the RL framework\. Specifically, we consider a single\-regime scenario \(i\.e\., the state space of the Markov chain isℳ=\{1\}\\mathcal\{M\}=\\\{1\\\}\) and setC=D1=D2=0C=D\_\{1\}=D\_\{2\}=0\. The activation function isφ1\(x\)=φ2\(x\)=\[x2x1\]T\\varphi\_\{1\}\(x\)=\\varphi\_\{2\}\(x\)=\\begin\{bmatrix\}x^\{2\}&x&1\\end\{bmatrix\}^\{T\}, whereλ1=0\.01\\lambda\_\{1\}=0\.01andλ2=0\.01\\lambda\_\{2\}=0\.01\. The algorithm starts fromW1\(0\)=\[111\]TW\_\{1\}^\{\(0\)\}=\\begin\{bmatrix\}1&1&1\\end\{bmatrix\}^\{T\},W2\(0\)=\[111\]TW\_\{2\}^\{\(0\)\}=\\begin\{bmatrix\}1&1&1\\end\{bmatrix\}^\{T\},α\(0\)=0\\alpha^\{\(0\)\}=0,β\(0\)=0\\beta^\{\(0\)\}=0, and the initial state isx=0\.1x=0\.1under regimei=1i=1\. One point is sampled every 0\.05 seconds, and after collectingM=1000M=1000points, we obtain the following results\.
Figure 7:Value functionsV1\(x\)V\_\{1\}\(x\),V2\(x\)V\_\{2\}\(x\),V1cl\(x\)V^\{cl\}\_\{1\}\(x\),V2cl\(x\)V^\{cl\}\_\{2\}\(x\)\.TABLE I:Final weight vectors of the critic NNsFig\.[7](https://arxiv.org/html/2606.28671#S5.F7)displays the curve of the value function, while Table[I](https://arxiv.org/html/2606.28671#S5.T1)presents the final weight vectors of the critic NNs\. The results indicate that with a low temperature parameter, the PIA can achieve a good approximation to the optimal value function, which shows that the RL framework plays a crucial role in preventing the algorithm from converging to local optima\.
Mathematically, classical deterministic methods greedily exploit the current best solution and are highly prone to getting trapped in local optima\. In contrast, the introduction of entropy regularization fosters a randomized policy\. This mechanism maintains a certain degree of stochasticity, encouraging the algorithm to continuously explore the action space rather than prematurely committing to a local peak\. By balancing exploration and exploitation, the randomized policy enables the system to evaluate a broader range of state\-action trajectories, successfully escape the attraction of local optima, and ultimately discover the global optimal solution\.
## VIConclusion
This paper investigates the LQ\-SDG problem subject to Markovian regime switching within the ERRL framework\. It begins by introducing the ERRL framework under regime\-switching diffusions and clarifying the equilibrium policy pair problem\. We then characterize the players’ regime\-dependent equilibrium policy pairs through the game’s weakly\-coupled value functions, demonstrating that these policies can effectively solve the formulated LQ\-SDG problem\. Subsequently, a policy improvement algorithm \(PIA\) is proposed to approximate the value functions and equilibrium policies, with a critic neural network architecture introduced to address the high\-dimensional computational challenges inherent in multi\-regime systems\. Finally, the effectiveness of the proposed algorithms is verified through multiple numerical examples, while the model’s convergence, the impact of temperature parameters, and robustness to structural shifts are analyzed\. Specifically, we use a practical example to illustrate how this exploratory framework effectively avoids the “local optimum trap”\.
In addition, the game framework constructed in this paper demonstrates excellent scalability: on one hand, its applicability can be extended from the current linear\-quadratic regime\-switching setting to nonlinear systems, and can also be expanded to multi\-agent interaction scenarios under Markovian environments \(a classical case is provided in\[[11](https://arxiv.org/html/2606.28671#bib.bib43)\]\); on the other hand, this framework provides further space for algorithm optimization, such as exploring fully model\-free RL algorithms that do not require prior knowledge of the control input matrices\.
## Appendix AExploratory dynamic system with regime switching
Consider the dynamic system with regime switching \([2\.6](https://arxiv.org/html/2606.28671#S2.E6)\), and letxsk,lx\_\{s\}^\{k,l\},k=1,⋯,N,l=1,⋯,Mk=1,\\cdots,N,l=1,\\cdots,M, be the copies of the dynamic system respectively under the policies\(usk,vsl\)\(u\_\{s\}^\{k\},v\_\{s\}^\{l\}\), each independently sampled from\(𝝅s,𝜸s\)\(\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\)conditional on the current statexsx\_\{s\}and regimeαs\\alpha\_\{s\}\. Letds\>0ds\>0be small enough and assume that the current regimeαs\\alpha\_\{s\}, as well as the corresponding policies\(usk,vsl\)\(u\_\{s\}^\{k\},v\_\{s\}^\{l\}\), remain fixed fromsstos\+dss\+ds\. Then, the increments of these dynamic system copies are, fork=1,⋯,N,l=1,⋯,Mk=1,\\cdots,N,l=1,\\cdots,M,
dxsk,l≡\\displaystyle dx\_\{s\}^\{k,l\}\\equivxs\+dsk,l−xsk,l\\displaystyle x\_\{s\+ds\}^\{k,l\}\-x\_\{s\}^\{k,l\}≈\\displaystyle\\approxb\(xsk,l,αs,usk,vsl\)ds\\displaystyle b\(x\_\{s\}^\{k,l\},\\alpha\_\{s\},u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)ds\+σ\(xsk,l,αs,usk,vsl\)\(Ws\+dsk,l−Wsk,l\),s≥0\.\\displaystyle\+\\sigma\(x\_\{s\}^\{k,l\},\\alpha\_\{s\},u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)\\left\(W\_\{s\+ds\}^\{k,l\}\-W\_\{s\}^\{k,l\}\\right\),\\,s\\geq 0\.Each such dynamic systemxsk,l,k=1,⋯,N,l=1,⋯,Mx\_\{s\}^\{k,l\},k=1,\\cdots,N,l=1,\\cdots,M, can be viewed as an independent sample from the exploratory dynamic systemXs𝝅,𝜸X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}\.
By conditioning on the current regimeαs=i∈ℳ\\alpha\_\{s\}=i\\in\\mathcal\{M\}, it can be concluded from the law of large numbers that,
limN→∞limM→∞1N1M∑k=1N∑l=1Mdxsk,l\\displaystyle\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}dx\_\{s\}^\{k,l\}≈\\displaystyle\\approxlimN→∞limM→∞1N1M∑k=1N∑l=1Mb\(xsk,l,i,usk,vsl\)ds\\displaystyle\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}b\(x\_\{s\}^\{k,l\},i,u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)ds\+limN→∞limM→∞1N1M∑k=1N∑l=1M\\displaystyle\+\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}σ\(xsk,l,i,usk,vsl\)\(Ws\+dsk,l−Wsk,l\)\\displaystyle\\quad\\sigma\(x\_\{s\}^\{k,l\},i,u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)\\left\(W\_\{s\+ds\}^\{k,l\}\-W\_\{s\}^\{k,l\}\\right\)→a\.s\.\\displaystyle\\xrightarrow\{\\text\{a\.s\.\}\}\{∫ℝp\[∫ℝpb\(xs,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\}ds,\\displaystyle\\left\\\{\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}b\(x\_\{s\},i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\\right\\\}ds,and
limN→∞limM→∞1N1M∑k=1N∑l=1Mdxsk,l\(dxsk,l\)T\\displaystyle\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}dx\_\{s\}^\{k,l\}\\left\(dx\_\{s\}^\{k,l\}\\right\)^\{T\}≈\\displaystyle\\approxlimN→∞limM→∞1N1M∑k=1N∑l=1M\\displaystyle\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}σ\(xsk,l,i,usk,vsl\)σT\(xsk,l,i,usk,vsl\)ds→a\.s\.\\displaystyle\\sigma\(x\_\{s\}^\{k,l\},i,u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)\\sigma^\{T\}\(x\_\{s\}^\{k,l\},i,u\_\{s\}^\{k\},v\_\{s\}^\{l\}\)ds\\xrightarrow\{\\text\{a\.s\.\}\}\{∫ℝp\[∫ℝpσ\(xs,i,u,v\)σT\(xs,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\}ds\.\\displaystyle\\left\\\{\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\sigma\(x\_\{s\},i,u,v\)\\sigma^\{T\}\(x\_\{s\},i,u,v\)\\boldsymbol\{\\pi\}\(u\)du\\right\]\\boldsymbol\{\\gamma\}\(v\)dv\\right\\\}ds\.
Since each statexsk,lx\_\{s\}^\{k,l\}is an independent sample fromXs𝝅,𝜸X\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}, the incrementsdxsk,ldx\_\{s\}^\{k,l\}are independent samples fromdXs𝝅,𝜸dX\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}\. Therefore, using the law of large numbers again, we obtain
limN→∞limM→∞1N1M∑k=1N∑l=1Mdxsk,l=dXs𝝅,𝜸\.\\lim\_\{N\\rightarrow\\infty\}\\lim\_\{M\\rightarrow\\infty\}\\frac\{1\}\{N\}\\frac\{1\}\{M\}\\sum\_\{k=1\}^\{N\}\\sum\_\{l=1\}^\{M\}dx\_\{s\}^\{k,l\}=dX\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\}\.\(A\.1\)Thus, under the regimeαs=i\\alpha\_\{s\}=i, the exploratory drift and diffusion are identified as
b~\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpb\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v,\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\):=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}b\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv,σ~\(x,i,𝝅,𝜸\):=\\displaystyle\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)=\{∫ℝp\[∫ℝpσ\(x,i,u,v\)σT\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\}1/2\.\\displaystyle\\left\\\{\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\sigma\(x,i,u,v\)\\sigma^\{T\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)du\\right\]\\boldsymbol\{\\gamma\}\(v\)dv\\right\\\}^\{1/2\}\.Similarly, utilizing the law of large numbers, the exploratory performance functions under regimeiiare
r~1\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpr1\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v,\\tilde\{r\}\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\):=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}r\_\{1\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv,r~2\(x,i,𝝅,𝜸\):=∫ℝp\[∫ℝpr2\(x,i,u,v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\.\\tilde\{r\}\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\):=\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}r\_\{2\}\(x,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\.
## Appendix BProof of Proposition[1](https://arxiv.org/html/2606.28671#Thmproposition1)
###### Proof 4
For any𝛑∈𝒰\\boldsymbol\{\\pi\}\\in\\mathcal\{U\},𝛄∈𝒱\\boldsymbol\{\\gamma\}\\in\\mathcal\{V\}and fixed regimei∈ℳi\\in\\mathcal\{M\}, we have
- \(11\)Lipschitz condition: ‖b~\(x,i,𝝅,𝜸\)−b~\(y,i,𝝅,𝜸\)‖\\displaystyle\\left\|\\left\|\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\-\\tilde\{b\}\(y,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\right\|\\right\|=\\displaystyle=‖∫ℝp\[∫ℝpb\(x,i,u,v\)−b\(y,i,u,v\)𝝅\(u\)du\]𝜸\(v\)𝑑v‖\\displaystyle\\left\|\\left\|\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}b\(x,i,u,v\)\-b\(y,i,u,v\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\\right\|\\right\|≤\\displaystyle\\leq∫ℝp\[∫ℝp‖b\(x,i,u,v\)−b\(y,i,u,v\)‖𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\\displaystyle\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\|\\left\|b\(x,i,u,v\)\-b\(y,i,u,v\)\\right\|\\right\|\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv≤\\displaystyle\\leq∫ℝp\[∫ℝp‖A\(i\)x−A\(i\)y‖𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\\displaystyle\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\|\\left\|A\(i\)x\-A\(i\)y\\right\|\\right\|\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv≤\\displaystyle\\leq‖A\(i\)‖‖x−y‖\.\\displaystyle\\left\|\\left\|A\(i\)\\right\|\\right\|\\left\|\\left\|x\-y\\right\|\\right\|\.Similarly, ‖σ~\(x,i,𝝅,𝜸\)−σ~\(y,i,𝝅,𝜸\)‖≤‖C\(i\)‖‖x−y‖\.\\left\|\\left\|\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\-\\tilde\{\\sigma\}\(y,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\right\|\\right\|\\leq\\left\|\\left\|C\(i\)\\right\|\\right\|\\left\|\\left\|x\-y\\right\|\\right\|\.
- \(22\)Linear growth condition: ‖b~\(x,i,𝝅,𝜸\)‖\\displaystyle\\left\|\\left\|\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)\\right\|\\right\|=\\displaystyle=‖∫ℝp\[∫ℝp\(A\(i\)x\+B1\(i\)u\+B2\(i\)v\)𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v‖\\displaystyle\\left\|\\left\|\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\(A\(i\)x\+B\_\{1\}\(i\)u\+B\_\{2\}\(i\)v\\right\)\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\\right\|\\right\|≤\\displaystyle\\leq∫ℝp\[∫ℝp‖A\(i\)x\+B1\(i\)u\+B2\(i\)v‖𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\\displaystyle\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\|\\left\|A\(i\)x\+B\_\{1\}\(i\)u\+B\_\{2\}\(i\)v\\right\|\\right\|\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv≤\\displaystyle\\leq∫ℝp\[∫ℝp‖B1\(i\)u\+B2\(i\)v‖𝝅\(u\)𝑑u\]𝜸\(v\)𝑑v\\displaystyle\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\[\\int\_\{\\mathbb\{R\}^\{p\}\}\\left\|\\left\|B\_\{1\}\(i\)u\+B\_\{2\}\(i\)v\\right\|\\right\|\\boldsymbol\{\\pi\}\(u\)\\,du\\right\]\\boldsymbol\{\\gamma\}\(v\)\\,dv\+‖A\(i\)‖‖x‖\.\\displaystyle\+\\left\|\\left\|A\(i\)\\right\|\\right\|\\left\|\\left\|x\\right\|\\right\|\.Similarly, we can proveσ~\(x,i,𝝅,𝜸\)\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\)has linear growth\.
By the classical existence and uniqueness theorem for stochastic differential equations \(c\.f\.\[[17](https://arxiv.org/html/2606.28671#bib.bib48)\], Theorem 5\.2\.1\), \([2\.10](https://arxiv.org/html/2606.28671#S2.E10)\) has a unique strong solution\.
## Appendix CProof of Lemma[1](https://arxiv.org/html/2606.28671#Thmlemma1)
###### Proof 5
To simplify the notation, we denoteXs𝛑,𝛄∗\(𝛑\):=XsX\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\}:=X\_\{s\}\. According to Definition[2](https://arxiv.org/html/2606.28671#Thmdefinition2),
V1\(x,i\)=min𝝅∈𝒰J1\(x,i,𝝅,𝜸∗\(𝝅\)\)\.V\_\{1\}\(x,i\)=\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\.Leth\>0h\>0be arbitrary and applying dynamic programming principle,
V1\(x,i\)\\displaystyle V\_\{1\}\(x,i\)=min𝝅∈𝒰𝔼x,i\[∫0he−ρ1s\[r~1\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle=\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\mathbb\{E\}\_\{x,i\}\\Bigg\[\\int\_\{0\}^\{h\}e^\{\-\\rho\_\{1\}s\}\\Big\[\\tilde\{r\}\_\{1\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\(C\.1\)\+λ1∫ℝp𝝅s\(u\)ln𝝅s\(u\)du\]ds\+e−ρ1hV1\(Xh,αh\)\]\.\\displaystyle\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\ln\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du\\Big\]ds\+e^\{\-\\rho\_\{1\}h\}V\_\{1\}\(X\_\{h\},\\alpha\_\{h\}\)\\Bigg\]\.Applying the generalized Itô’s formula with Markovian switching toe−ρ1hV1\(Xh,αh\)e^\{\-\\rho\_\{1\}h\}V\_\{1\}\(X\_\{h\},\\alpha\_\{h\}\)and integrating from0tohh, we have
e−ρ1h\\displaystyle e^\{\-\\rho\_\{1\}h\}V1\(Xh,αh\)−V1\(x,i\)=∫0he−ρ1s\[−ρ1V1\(Xs,αs\)\\displaystyle V\_\{1\}\(X\_\{h\},\\alpha\_\{h\}\)\-V\_\{1\}\(x,i\)=\\int\_\{0\}^\{h\}e^\{\-\\rho\_\{1\}s\}\\Big\[\-\\rho\_\{1\}V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\(C\.2\)\+∇V1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\+\\nabla V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)σ~T\(⋯\)ΔV1\(Xs,αs\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\cdots\)\\Delta V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\\right\)\+∑j=1mqαsjV1\(Xs,j\)\]ds\+Mh\\displaystyle\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}V\_\{1\}\(X\_\{s\},j\)\\Big\]ds\+M\_\{h\}\+∫0he−ρ1s∇V1\(Xs,αs\)Tσ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)𝑑Ws,\\displaystyle\+\\int\_\{0\}^\{h\}e^\{\-\\rho\_\{1\}s\}\\nabla V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)dW\_\{s\},whereMhM\_\{h\}is a zero\-mean martingale arising from the pure jump process of the Markov chain\. Substituting equation \([C\.2](https://arxiv.org/html/2606.28671#A3.E2)\) back into equation \([C\.1](https://arxiv.org/html/2606.28671#A3.E1)\), taking expectation \(where the integrals with respect todWsdW\_\{s\}and the martingaleMhM\_\{h\}vanish\), yields
min𝝅∈𝒰𝔼x,i\\displaystyle\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\mathbb\{E\}\_\{x,i\}\[∫0he−ρ1s\[r~1\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\\Bigg\[\\int\_\{0\}^\{h\}e^\{\-\\rho\_\{1\}s\}\\Bigg\[\\tilde\{r\}\_\{1\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+λ1∫ℝp𝝅s\(u\)ln𝝅s\(u\)𝑑u−ρ1V1\(Xs,αs\)\\displaystyle\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\ln\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du\-\\rho\_\{1\}V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+∇V1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\+\\nabla V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)σ~T\(⋯\)ΔV1\(Xs,αs\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\cdots\)\\Delta V\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\\right\)\+∑j=1mqαsjV1\(Xs,j\)\]ds\]=0\.\\displaystyle\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}V\_\{1\}\(X\_\{s\},j\)\\Bigg\]ds\\Bigg\]=0\.Leth→0h\\rightarrow 0, divide byhh, and apply the mean value theorem for integrals, and we obtain
ρ1V1\(x,i\)=\\displaystyle\\rho\_\{1\}V\_\{1\}\(x,i\)=min𝝅∈𝒰\[r~1\(x,i,𝝅,𝜸∗\(𝝅\)\)\+λ1∫ℝp𝝅\(u\)ln𝝅\(u\)du\\displaystyle\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\Big\[\\tilde\{r\}\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\(u\)\\ln\\boldsymbol\{\\pi\}\(u\)\\,du\+∇V1\(x,i\)Tb~\(x,i,𝝅,𝜸∗\(𝝅\)\)\\displaystyle\+\\nabla V\_\{1\}\(x,i\)^\{T\}\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(x,i,𝝅,𝜸∗\(𝝅\)\)σ~T\(⋯\)ΔV1\(x,i\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\cdots\)\\Delta V\_\{1\}\(x,i\)\\right\)\+∑j=1mqijV1\(x,j\)\]\.\\displaystyle\+\\sum\_\{j=1\}^\{m\}q\_\{ij\}V\_\{1\}\(x,j\)\\Big\]\.We often rewrite this PDE in the following form
ρ1ϕ1\(x,i\)=min𝝅∈𝒰ℋ1\(x,i,∇ϕ1\(x,i\),Δϕ1\(x,i\),𝝅,𝜸∗\(𝝅\)\),\\rho\_\{1\}\\phi\_\{1\}\(x,i\)=\\min\_\{\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\}\\mathcal\{H\}\_\{1\}\(x,i,\\nabla\\phi\_\{1\}\(x,i\),\\Delta\\phi\_\{1\}\(x,i\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\),whereϕ1\\phi\_\{1\}is the solution to this PDE, and the specific form ofℋ1\\mathcal\{H\}\_\{1\}incorporates the regime switching jumps:
ℋ1\(x,i,∇ϕ1\(x,i\),Δϕ1\(x,i\),𝝅,𝜸∗\(𝝅\)\)\\displaystyle\\mathcal\{H\}\_\{1\}\(x,i,\\nabla\\phi\_\{1\}\(x,i\),\\Delta\\phi\_\{1\}\(x,i\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\):=\\displaystyle=r~1\(x,i,𝝅,𝜸∗\(𝝅\)\)\+λ1∫ℝp𝝅\(u\)ln𝝅\(u\)𝑑u\\displaystyle\\tilde\{r\}\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\(u\)\\ln\\boldsymbol\{\\pi\}\(u\)\\,du\+∇ϕ1\(x,i\)Tb~\(x,i,𝝅,𝜸∗\(𝝅\)\)\+∑j=1mqijϕ1\(x,j\)\\displaystyle\+\\nabla\\phi\_\{1\}\(x,i\)^\{T\}\\tilde\{b\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+\\sum\_\{j=1\}^\{m\}q\_\{ij\}\\phi\_\{1\}\(x,j\)\+12tr\(σ~\(x,i,𝝅,𝜸∗\(𝝅\)\)σ~T\(⋯\)Δϕ1\(x,i\)\)\.\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\cdots\)\\Delta\\phi\_\{1\}\(x,i\)\\right\)\.
We can also prove that ifV2\(⋅,i;𝛑\)∈C2\(ℝn\)V\_\{2\}\(\\cdot,i;\\boldsymbol\{\\pi\}\)\\in C^\{2\}\(\\mathbb\{R\}^\{n\}\)for eachi∈ℳi\\in\\mathcal\{M\}, thenV2\(x,i;𝛑\)V\_\{2\}\(x,i;\\boldsymbol\{\\pi\}\)is a solution to the HJBI equation \([3\.2](https://arxiv.org/html/2606.28671#S3.E2)\)\. The proof is omitted\.
## Appendix DProof of Theorem[3](https://arxiv.org/html/2606.28671#Thmtheorem3)
###### Proof 6
To simplify the notation, we denoteXs𝛑,𝛄∗\(𝛑\):=XsX\_\{s\}^\{\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\}:=X\_\{s\}\. Suppose that𝛑\\boldsymbol\{\\pi\}is an arbitrarily given policy for leader, and𝛄∗\(𝛑\)\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)is the corresponding optimal response policy for follower\.
Applying the generalized Itô’s formula toe−ρ1sϕ1\(Xs,αs\)e^\{\-\\rho\_\{1\}s\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\), we get
de−ρ1sϕ1\(Xs,αs\)=\\displaystyle de^\{\-\\rho\_\{1\}s\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)=e−ρ1s\[−ρ1ϕ1\(Xs,αs\)\+∇ϕ1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle e^\{\-\\rho\_\{1\}s\}\\Big\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(⋅\)σ~T\(⋅\)Δϕ1\(Xs,αs\)\)\+∑j=1mqαsjϕ1\(Xs,j\)\]ds\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(\\cdot\)\\tilde\{\\sigma\}^\{T\}\(\\cdot\)\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\\right\)\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}\\phi\_\{1\}\(X\_\{s\},j\)\\Big\]ds\+e−ρ1s∇ϕ1\(Xs,αs\)Tσ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)dWs\+dMs,\\displaystyle\+e^\{\-\\rho\_\{1\}s\}\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)dW\_\{s\}\+dM\_\{s\},whereMsM\_\{s\}is the associated jump martingale\. LetT\>0T\>0be arbitrary, and define the stopping times
τn:=\\displaystyle\\tau\_\{n\}=\{t≥0:∫0t\|\|e−ρ1s∇ϕ1\(Xs,αs\)T\\displaystyle\\Big\\\{t\\geq 0:\\int\_\{0\}^\{t\}\|\|e^\{\-\\rho\_\{1\}s\}\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}σ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\|\|2ds≥n\},n≥1\.\\displaystyle\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\|\|^\{2\}ds\\geq n\\Big\\\},\\,n\\geq 1\.Integrating from0toT∧τnT\\wedge\\tau\_\{n\}and taking expectation𝔼x,i\[⋅\]\\mathbb\{E\}\_\{x,i\}\[\\cdot\]on both sides, we have
𝔼x,i\[e−ρ1\(T∧τn\)ϕ1\(XT∧τn,αT∧τn\)\]−ϕ1\(x,i\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[e^\{\-\\rho\_\{1\}\(T\\wedge\\tau\_\{n\}\)\}\\phi\_\{1\}\(X\_\{T\\wedge\\tau\_\{n\}\},\\alpha\_\{T\\wedge\\tau\_\{n\}\}\)\\right\]\-\\phi\_\{1\}\(x,i\)=\\displaystyle=𝔼x,i\[∫0T∧τne−ρ1s\[−ρ1ϕ1\(Xs,αs\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Bigg\[\\int\_\{0\}^\{T\\wedge\\tau\_\{n\}\}e^\{\-\\rho\_\{1\}s\}\\Big\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+∇ϕ1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\+\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(⋅\)σ~T\(⋅\)Δϕ1\(Xs,αs\)\)\+∑j=1mqαsjϕ1\(Xs,j\)\]ds\]\.\\displaystyle\+\\frac\{1\}\{2\}tr\(\\tilde\{\\sigma\}\(\\cdot\)\\tilde\{\\sigma\}^\{T\}\(\\cdot\)\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\)\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}\\phi\_\{1\}\(X\_\{s\},j\)\\Big\]ds\\Bigg\]\.Settingn→∞n\\to\\infty, we deduce that
𝔼x,i\[e−ρ1Tϕ1\(XT,αT\)\]−ϕ1\(x,i\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[e^\{\-\\rho\_\{1\}T\}\\phi\_\{1\}\(X\_\{T\},\\alpha\_\{T\}\)\\right\]\-\\phi\_\{1\}\(x,i\)=\\displaystyle=𝔼x,i\[∫0Te−ρ1s\[−ρ1ϕ1\(Xs,αs\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Bigg\[\\int\_\{0\}^\{T\}e^\{\-\\rho\_\{1\}s\}\\Big\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+∇ϕ1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\+\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(⋅\)σ~T\(⋅\)Δϕ1\(Xs,αs\)\)\+∑j=1mqαsjϕ1\(Xs,j\)\]ds\]\.\\displaystyle\+\\frac\{1\}\{2\}tr\(\\tilde\{\\sigma\}\(\\cdot\)\\tilde\{\\sigma\}^\{T\}\(\\cdot\)\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\)\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}\\phi\_\{1\}\(X\_\{s\},j\)\\Big\]ds\\Bigg\]\.Next, using the condition \(iv\) in Assumption[1](https://arxiv.org/html/2606.28671#Thmassumption1)and applying the dominated convergence theorem yields
−ϕ1\(x,i\)=\\displaystyle\-\\phi\_\{1\}\(x,i\)=𝔼x,i\[∫0∞e−ρ1s\[−ρ1ϕ1\(Xs,αs\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Bigg\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{1\}s\}\\Big\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+∇ϕ1\(Xs,αs\)Tb~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\+\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)^\{T\}\\tilde\{b\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)σ~T\(…\)Δϕ1\(Xs,αs\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\(\\tilde\{\\sigma\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\dots\)\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\)\+∑j=1mqαsjϕ1\(Xs,j\)\]ds\]\.\\displaystyle\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}\\phi\_\{1\}\(X\_\{s\},j\)\\Big\]ds\\Bigg\]\.Adding the same item
𝔼x,i\[∫0∞e−ρ1s\[r~1\(Xs,αs,𝝅s,𝜸s∗\(𝝅\)\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{1\}s\}\\left\[\\tilde\{r\}\_\{1\}\(X\_\{s\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}^\{\*\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\right\.\\right\.\+λ1∫ℝp𝝅s\(u\)ln𝝅s\(u\)du\]ds\]\\displaystyle\\qquad\\left\.\\left\.\+\\lambda\_\{1\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\ln\\boldsymbol\{\\pi\}\_\{s\}\(u\)\\,du\\,\\right\]ds\\right\]on both sides, with the definition of the exploratory performance function \([2\.11](https://arxiv.org/html/2606.28671#S2.E11)\) of leader, we have
J1\(x,i,𝝅,𝜸∗\(𝝅\)\)−ϕ1\(x,i\)\\displaystyle J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\-\\phi\_\{1\}\(x,i\)=\\displaystyle=𝔼x,i\[∫0∞e−ρ1s\[−ρ1ϕ1\(Xs,αs\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Big\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{1\}s\}\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\)\+ℋ1\(Xs,αs,∇ϕ1\(Xs,αs\),Δϕ1\(Xs,αs\),𝝅s,𝜸s∗\(𝝅\)\)\]ds\]\.\\displaystyle\+\\mathcal\{H\}\_\{1\}\(X\_\{s\},\\alpha\_\{s\},\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\),\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\),\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\]ds\\Big\]\.On the other hand,
ρ1ϕ1\(Xs𝝅∗,𝜸∗,αs\)\\displaystyle\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\}\)=\\displaystyle=ℋ1\(Xs𝝅∗,𝜸∗,αs,∇ϕ1\(Xs𝝅∗,𝜸∗,αs\),Δϕ1\(Xs𝝅∗,𝜸∗,αs\),𝝅s∗,𝜸s∗\)\\displaystyle\\mathcal\{H\}\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\},\\nabla\\phi\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\}\),\\Delta\\phi\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\}\),\\boldsymbol\{\\pi\}\_\{s\}^\{\*\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\)≤\\displaystyle\\leqℋ1\(Xs,αs,∇ϕ1\(Xs,αs\),Δϕ1\(Xs,αs\),𝝅s,𝜸s∗\(𝝅\)\),∀𝝅∈𝒰\.\\displaystyle\\mathcal\{H\}\_\{1\}\(X\_\{s\},\\alpha\_\{s\},\\nabla\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\),\\Delta\\phi\_\{1\}\(X\_\{s\},\\alpha\_\{s\}\),\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\),\\,\\forall\\boldsymbol\{\\pi\}\\in\\mathcal\{U\}\.Therefore,
J1\(x,i,𝝅,𝜸∗\(𝝅\)\)−ϕ1\(x,i\)\\displaystyle J\_\{1\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\*\}\(\\boldsymbol\{\\pi\}\)\)\-\\phi\_\{1\}\(x,i\)≥\\displaystyle\\geq𝔼x,i\[∫0∞e−ρ1s\[−ρ1ϕ1\(Xs𝝅∗,𝜸∗,αs\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Big\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{1\}s\}\[\-\\rho\_\{1\}\\phi\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\}\)\+ℋ1\(Xs𝝅∗,𝜸∗,αs,⋯\)\]ds\]\\displaystyle\+\\mathcal\{H\}\_\{1\}\(X\_\{s\}^\{\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\},\\alpha\_\{s\},\\cdots\)\]ds\\Big\]=\\displaystyle=0=J1\(x,i,𝝅∗,𝜸∗\)−ϕ1\(x,i\)\.\\displaystyle 0=J\_\{1\}\(x,i,\\boldsymbol\{\\pi\}^\{\*\},\\boldsymbol\{\\gamma\}^\{\*\}\)\-\\phi\_\{1\}\(x,i\)\.The first and second equations of Theorem[3](https://arxiv.org/html/2606.28671#Thmtheorem3)are proved\. The proofs proces for the third and fourth equations of Theorem[3](https://arxiv.org/html/2606.28671#Thmtheorem3)are similar and omitted here\.
## Appendix EProof of Theorem[4](https://arxiv.org/html/2606.28671#Thmtheorem4)
###### Proof 7
LetXs𝛄X\_\{s\}^\{\\boldsymbol\{\\gamma\}\}\(resp\.,Xs𝛄~X\_\{s\}^\{\\tilde\{\\boldsymbol\{\\gamma\}\}\}\),s∈\[0,∞\)s\\in\[0,\\infty\)be the state process under policy pair\(𝛑,𝛄\(𝛑\)\)\(\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)\(resp\.,\(𝛑,𝛄~\(𝛑\)\)\(\\boldsymbol\{\\pi\},\\tilde\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)\)\)\. First applying the generalized Itô’s formula,
de−ρ2sJ2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\\displaystyle de^\{\-\\rho\_\{2\}s\}J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)=\\displaystyle=e−ρ2s\[∇J2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)Tb~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\\displaystyle e^\{\-\\rho\_\{2\}s\}\\Big\[\\nabla J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)^\{T\}\\tilde\{b\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)σ~T\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\right\.ΔJ2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\)−ρ2J2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\\displaystyle\\quad\\left\.\\Delta J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\right\)\-\\rho\_\{2\}J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+∑j=1mqαsjJ2\(Xs𝜸,j,𝝅s,𝜸s\(𝝅\)\)\]ds\+dMs\\displaystyle\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},j,\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\Big\]ds\+dM\_\{s\}\+e−ρ2s∇J2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)Tσ~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)dWs\.\\displaystyle\+e^\{\-\\rho\_\{2\}s\}\\nabla J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)^\{T\}\\tilde\{\\sigma\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)dW\_\{s\}\.Define the stopping times
τn:=\\displaystyle\\tau\_\{n\}=\{t≥0:∫0t\|\|e−ρ2s∇J2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)T\\displaystyle\\Big\\\{t\\geq 0:\\int\_\{0\}^\{t\}\|\|e^\{\-\\rho\_\{2\}s\}\\nabla J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)^\{T\}σ~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\|\|2ds≥n\}\\displaystyle\\tilde\{\\sigma\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\|\|^\{2\}ds\\geq n\\Big\\\}forn≥1n\\geq 1\. Integrating from0toτn\\tau\_\{n\}and taking expectation𝔼x,i\[⋅\]\\mathbb\{E\}\_\{x,i\}\[\\cdot\]on both sides, we have
J2\(x,i,𝝅,𝜸\(𝝅\)\)\\displaystyle J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)=\\displaystyle=𝔼x,i\[−∫0τne−ρ2s\[∇J2\(…\)Tb~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Bigg\[\-\\int\_\{0\}^\{\\tau\_\{n\}\}e^\{\-\\rho\_\{2\}s\}\\Big\[\\nabla J\_\{2\}\(\\dots\)^\{T\}\\tilde\{b\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+12tr\(σ~\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)σ~T\(…\)ΔJ2\(…\)\)\\displaystyle\+\\frac\{1\}\{2\}tr\\left\(\\tilde\{\\sigma\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\tilde\{\\sigma\}^\{T\}\(\\dots\)\\Delta J\_\{2\}\(\\dots\)\\right\)−ρ2J2\(Xs𝜸,αs,𝝅s,𝜸s\(𝝅\)\)\+∑j=1mqαsjJ2\(⋯\)\]ds\\displaystyle\-\\rho\_\{2\}J\_\{2\}\(X\_\{s\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\boldsymbol\{\\gamma\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\+\\sum\_\{j=1\}^\{m\}q\_\{\\alpha\_\{s\}j\}J\_\{2\}\(\\cdots\)\\Big\]ds\+e−ρ2τnJ2\(Xτn𝜸,ατn,𝝅τn,𝜸τn\(𝝅\)\)\]\.\\displaystyle\+e^\{\-\\rho\_\{2\}\\tau\_\{n\}\}J\_\{2\}\(X\_\{\\tau\_\{n\}\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{\\tau\_\{n\}\},\\boldsymbol\{\\pi\}\_\{\\tau\_\{n\}\},\\boldsymbol\{\\gamma\}\_\{\\tau\_\{n\}\}\(\\boldsymbol\{\\pi\}\)\)\\Bigg\]\.On the other hand,
ρ2J2\(x,i,𝝅,𝜸\(𝝅\)\)\\displaystyle\\rho\_\{2\}J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)\(E\.1\)=\\displaystyle=ℋ2\(x,i,∇J2\(x,i,𝝅,𝜸\(𝝅\)\),ΔJ2\(x,i,𝝅,𝜸\(𝝅\)\),𝝅,𝜸\(𝝅\)\)\\displaystyle\\mathcal\{H\}\_\{2\}\(x,i,\\nabla J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\),\\Delta J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)≥\\displaystyle\\geqmin𝜸′ℋ2\(x,i,∇J2\(x,i,𝝅,𝜸′\(𝝅\)\),ΔJ2\(…\),𝝅,𝜸′\(𝝅\)\)\.\\displaystyle\\min\_\{\\boldsymbol\{\\gamma\}^\{\\prime\}\}\\mathcal\{H\}\_\{2\}\(x,i,\\nabla J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\\prime\}\(\\boldsymbol\{\\pi\}\)\),\\Delta J\_\{2\}\(\\dots\),\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}^\{\\prime\}\(\\boldsymbol\{\\pi\}\)\)\.Notice that the minimizer of the Hamiltonian in \([E\.1](https://arxiv.org/html/2606.28671#A5.E1)\) is given by the feedback policy \([4\.1](https://arxiv.org/html/2606.28671#S4.E1)\)\. It then follows that
J2\(x,i,𝝅,𝜸\(𝝅\)\)\\displaystyle J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)≥\\displaystyle\\geq𝔼x,i\[∫0τne−ρ2s\[r~2\(Xs𝜸~,αs,𝝅s,𝜸~s\(𝝅\)\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\Bigg\[\\int\_\{0\}^\{\\tau\_\{n\}\}e^\{\-\\rho\_\{2\}s\}\\left\[\\tilde\{r\}\_\{2\}\(X\_\{s\}^\{\\tilde\{\\boldsymbol\{\\gamma\}\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\right\.\+λ2∫ℝp𝜸~s\(𝝅\)\(v\)ln𝜸~s\(𝝅\)\(v\)dv\]ds\\displaystyle\\left\.\+\\lambda\_\{2\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\(v\)\\ln\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\(v\)\\,dv\\right\]ds\+e−ρ2τnJ2\(Xτn𝜸,ατn,𝝅τn,𝜸τn\(𝝅\)\)\]\.\\displaystyle\+e^\{\-\\rho\_\{2\}\\tau\_\{n\}\}J\_\{2\}\(X\_\{\\tau\_\{n\}\}^\{\\boldsymbol\{\\gamma\}\},\\alpha\_\{\\tau\_\{n\}\},\\boldsymbol\{\\pi\}\_\{\\tau\_\{n\}\},\\boldsymbol\{\\gamma\}\_\{\\tau\_\{n\}\}\(\\boldsymbol\{\\pi\}\)\)\\Bigg\]\.Settingn→∞n\\to\\inftyand applying the condition \(iv\) in Assumption[1](https://arxiv.org/html/2606.28671#Thmassumption1), we obtain
J2\(x,i,𝝅,𝜸\(𝝅\)\)\\displaystyle J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\boldsymbol\{\\gamma\}\(\\boldsymbol\{\\pi\}\)\)≥\\displaystyle\\geq𝔼x,i\[∫0∞e−ρ2s\[r~2\(Xs𝜸~,αs,𝝅s,𝜸~s\(𝝅\)\)\\displaystyle\\mathbb\{E\}\_\{x,i\}\\left\[\\int\_\{0\}^\{\\infty\}e^\{\-\\rho\_\{2\}s\}\\Big\[\\tilde\{r\}\_\{2\}\(X\_\{s\}^\{\\tilde\{\\boldsymbol\{\\gamma\}\}\},\\alpha\_\{s\},\\boldsymbol\{\\pi\}\_\{s\},\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\)\\right\.\+λ2∫ℝp𝜸~s\(𝝅\)\(v\)ln𝜸~s\(𝝅\)\(v\)dv\]ds\]=J2\(x,i,𝝅,𝜸~\(𝝅\)\)\.\\displaystyle\\left\.\+\\lambda\_\{2\}\\int\_\{\\mathbb\{R\}^\{p\}\}\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\(v\)\\ln\\tilde\{\\boldsymbol\{\\gamma\}\}\_\{s\}\(\\boldsymbol\{\\pi\}\)\(v\)\\,dv\\right\]ds\\Big\]=J\_\{2\}\(x,i,\\boldsymbol\{\\pi\},\\tilde\{\\boldsymbol\{\\gamma\}\}\(\\boldsymbol\{\\pi\}\)\)\.
## References
- \[1\]\(1999\)Dynamic noncooperative game theory\.SIAM\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[2\]A\. Bensoussan, M\. Chau, Y\. Lai, and S\. C\. P\. Yam\(2017\)Linear\-quadratic mean field stackelberg games with state and control delays\.SIAM Journal on Control and Optimization55\(4\),pp\. 2748–2781\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[3\]L\. Bo, Y\. Huang, X\. Yu, and T\. Zhang\(2024\)Continuous\-time q\-learning for jump\-diffusion models under tsallis entropy\.arXiv preprint arXiv:2407\.03888\.Cited by:[Remark 1](https://arxiv.org/html/2606.28671#Thmremark1.p1.6.6)\.
- \[4\]Z\. Chen and J\. Gu\(2025\)Exploratory utility maximization problem with tsallis entropy\.arXiv preprint arXiv:2502\.01269\.Cited by:[Remark 1](https://arxiv.org/html/2606.28671#Thmremark1.p1.6.6)\.
- \[5\]T\. Haarnoja, A\. Zhou, P\. Abbeel, and S\. Levine\(2018\)Soft actor\-critic: off\-policy maximum entropy deep reinforcement learning with a stochastic actor\.InInternational conference on machine learning,pp\. 1861–1870\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1)\.
- \[6\]C\. Han, L\. Huo, X\. Tong, H\. Wang, and X\. Liu\(2020\)Spatial anti\-jamming scheme for internet of satellites based on the deep reinforcement learning and stackelberg game\.IEEE Transactions on Vehicular Technology69\(5\),pp\. 5331–5342\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1)\.
- \[7\]X\. Han, R\. Wang, and X\. Y\. Zhou\(2023\)Choquet regularization for continuous\-time reinforcement learning\.SIAM Journal on Control and Optimization61\(5\),pp\. 2777–2801\.Cited by:[Remark 1](https://arxiv.org/html/2606.28671#Thmremark1.p1.6.6)\.
- \[8\]P\. Huang, H\. Ding, Z\. Sun, and H\. Chen\(2024\)A game\-based hierarchical model for mandatory lane change of autonomous vehicles\.IEEE Transactions on Intelligent Transportation Systems25\(9\),pp\. 11256–11268\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[9\]Y\. Jia and X\. Zhou\(2022\)Policy gradient and actor\-critic learning in continuous time and space: theory and algorithms\.Journal of Machine Learning Research23\(275\),pp\. 1–50\.Cited by:[§IV\-C](https://arxiv.org/html/2606.28671#S4.SS3.p1.1),[§IV](https://arxiv.org/html/2606.28671#S4.p3.1)\.
- \[10\]M\. Li, J\. Qin, and L\. Ding\(2019\)Two\-player stackelberg game for linear system via value iteration algorithm\.In2019 IEEE 28th International Symposium on Industrial Electronics \(ISIE\),pp\. 2289–2293\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p2.1),[§II](https://arxiv.org/html/2606.28671#S2.p1.1),[Remark 2](https://arxiv.org/html/2606.28671#Thmremark2.p2.4.4),[Remark 3](https://arxiv.org/html/2606.28671#Thmremark3.p2.5.1)\.
- \[11\]M\. Li, J\. Qin, Q\. Ma, W\. X\. Zheng, and Y\. Kang\(2020\)Hierarchical optimal synchronization for linear systems via reinforcement learning: a stackelberg–nash game perspective\.IEEE Transactions on Neural Networks and Learning Systems32\(4\),pp\. 1600–1611\.Cited by:[§VI](https://arxiv.org/html/2606.28671#S6.p2.1)\.
- \[12\]Z\. Li and J\. Shi\(2024\)Closed\-loop solvability of linear quadratic mean\-field type stackelberg stochastic differential games\.Applied Mathematics & Optimization90\(1\),pp\. 22\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[13\]S\. Lv, J\. Xiong, and X\. Zhang\(2023\)Linear quadratic leader–follower stochastic differential games for mean\-field switching diffusions\.Automatica154,pp\. 111072\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[14\]W\. M\. McEneaney\(2007\)A curse\-of\-dimensionality\-free numerical method for solution of certain hjb pdes\.SIAM journal on Control and Optimization46\(4\),pp\. 1239–1276\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1)\.
- \[15\]C\. Miao, Y\. Cui, H\. Li, and X\. Wu\(2025\)Effective multi\-agent deep reinforcement learning control with relative entropy regularization\.IEEE Transactions on Automation Science and Engineering22\(\),pp\. 3704–3718\.External Links:[Document](https://dx.doi.org/10.1109/TASE.2024.3398712)Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1)\.
- \[16\]J\. Moon\(2023\)Linear–quadratic stochastic leader–follower differential games for markov jump\-diffusion models\.Automatica147,pp\. 110713\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[17\]B\. Oksendal\(2013\)Stochastic differential equations: an introduction with applications\.Springer\.Cited by:[Proof 4](https://arxiv.org/html/2606.28671#Thmproof4.p1.4.1)\.
- \[18\]M\. Pakseresht, B\. Shirazi, I\. Mahdavi, and N\. Mahdavi\-Amiri\(2020\)Toward sustainable optimization with stackelberg game between green product family and downstream supply chain\.Sustainable Production and Consumption23,pp\. 198–211\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[19\]M\. Sockin and W\. Xiong\(2023\)Decentralization through tokenization\.The Journal of Finance78\(1\),pp\. 247–299\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[20\]J\. Sun, H\. Wang, and J\. Wen\(2023\)Zero\-sum stackelberg stochastic linear\-quadratic differential games\.SIAM Journal on Control and Optimization61\(1\),pp\. 252–284\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[21\]R\. S\. Sutton and A\. G\. Barto\(2018\)Reinforcement learning: an introduction\.MIT Press\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1)\.
- \[22\]W\. Tang, Y\. P\. Zhang, and X\. Y\. Zhou\(2022\)Exploratory hjb equations and their convergence\.SIAM Journal on Control and Optimization60\(6\),pp\. 3191–3216\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p5.1)\.
- \[23\]H\. von Stackelberg\(1934\)Market structure and equilibrium\.Springer\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[24\]H\. Wang, T\. Zariphopoulou, and X\. Y\. Zhou\(2019\)Exploration versus exploitation in reinforcement learning: a stochastic control approach\.Social Science Electronic Publishing10\(4\),pp\. 1–35\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p5.1)\.
- \[25\]H\. Wang and R\. Xu\(2023\)Time\-inconsistent lq games for large\-population systems and applications\.Journal of Optimization Theory and Applications197\(3\),pp\. 1249–1268\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[26\]H\. Wang, T\. Zariphopoulou, and X\. Y\. Zhou\(2020\)Reinforcement learning in continuous time and space: a stochastic control approach\.Journal of Machine Learning Research21\(198\),pp\. 1–34\.Cited by:[§II\-B](https://arxiv.org/html/2606.28671#S2.SS2.p4.14)\.
- \[27\]H\. Wang and X\. Y\. Zhou\(2020\)Continuous\-time mean–variance portfolio selection: a reinforcement learning framework\.Mathematical Finance30\(4\),pp\. 1273–1308\.Cited by:[§II\-B](https://arxiv.org/html/2606.28671#S2.SS2.p6.2),[§IV](https://arxiv.org/html/2606.28671#S4.p1.1)\.
- \[28\]W\. Xing, X\. Zhao, Y\. Li, and L\. Liu\(2025\)Denial\-of\-service attacks on cyber\-physical systems against linear quadratic control: a stackelberg\-game analysis\.IEEE Transactions on Automatic Control70\(1\),pp\. 595–602\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[29\]J\. Yong and X\. Y\. Zhou\(1999\)Stochastic controls: hamiltonian systems and hjb equations\.Springer\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p2.1)\.
- \[30\]J\. Yong\(2002\)A leader\-follower stochastic linear quadratic differential game\.SIAM Journal on Control and Optimization41\(4\),pp\. 1015–1041\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[31\]J\. Yong\(2013\)Linear\-quadratic optimal control problems for mean\-field stochastic differential equations\.SIAM Journal on Control and Optimization51\(4\),pp\. 2809–2838\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p1.1)\.
- \[32\]S\. Zhang, X\. Tong, K\. Chi, W\. Gao, X\. Chen, and Z\. Shi\(2025\)Stackelberg game\-based multi\-agent algorithm for resource allocation and task offloading in mec\-enabled c\-its\.IEEE Transactions on Intelligent Transportation Systems26\(10\),pp\. 17940–17951\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2025.3553487)Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p1.1)\.
- \[33\]R\. Zhao, X\. Sun, and V\. Tresp\(2019\)Maximum entropy\-regularized multi\-goal reinforcement learning\.InInternational Conference on Machine Learning,pp\. 7553–7562\.Cited by:[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p3.1),[§I\-A](https://arxiv.org/html/2606.28671#S1.SS1.p4.1)\.
- \[34\]M\. Zhou, J\. Han, and J\. Lu\(2021\)Actor\-critic method for high dimensional static hamilton–jacobi–bellman partial differential equations based on neural networks\.SIAM Journal on Scientific Computing43\(6\),pp\. A4043–A4066\.Cited by:[§IV\-C](https://arxiv.org/html/2606.28671#S4.SS3.p1.1),[§IV](https://arxiv.org/html/2606.28671#S4.p3.1)\.
- \[35\]X\. Y\. Zhou\(2021\)Curse of optimality, and how we break it\.Available at SSRN 3845462\.Cited by:[§I\-B](https://arxiv.org/html/2606.28671#S1.SS2.p2.1)\.Similar Articles
Entropy Regularized Reinforcement Learning for Zero-Sum Stochastic Differential Games in a Regime-Switching Jump-Diffusion Process
This paper introduces an entropy-regularized reinforcement learning framework for zero-sum stochastic differential games in regime-switching jump-diffusion processes, deriving HJBI equations and an actor-critic algorithm with applications to investment games.
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
This paper proposes Adaptive Entropy Regularization (AER), a framework that dynamically balances exploration and exploitation in LLM reinforcement learning by addressing policy entropy collapse through difficulty-aware coefficient allocation and initial-anchored target entropy. Experiments on mathematical reasoning benchmarks demonstrate consistent improvements in both accuracy and exploration capability.
Rethinking the Divergence Regularization in LLM RL
This paper introduces DRPO, which replaces the hard mask in DPPO with a smooth advantage-weighted quadratic regularizer to improve stability and efficiency in LLM reinforcement learning by providing continuous gradient corrections beyond trust-region boundaries.
Self-Play Reinforcement Learning under Imperfect Information in Big 2
This paper presents a self-play reinforcement learning framework for the four-player imperfect-information card game Big 2, comparing policy-gradient and value-based methods and finding that PPO with entropy regularization outperforms others.
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
EMAgnet introduces parameter-space exponential moving average regularization for policy gradient self-play in large two-player zero-sum games, achieving lower exploitability compared to uniform regularization targets.