演化稳定性不保证学习可达性:从多智能体强化学习视角看合作涌现
摘要
本文探讨为什么多智能体系统中的演化稳定合作结果可能无法通过去中心化强化学习算法达到,展示了稳定性和学习可达性之间的不同特性。
arXiv:2609.27664v1 Announce Type: new
Abstract: Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learning agents can reach the same outcome through local reward feedback.
We study this distinction in a transparent three-agent governance-motivated game involving a government, a platform firm, and users. We derive replicator dynamics for the fixed stage-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial-condition grid, and compare it with learning-basin estimates for three decentralized value-based learners. The learning analysis uses independent Q-learning with $\varepsilon$-greedy action selection, scaled Boltzmann exploration, and SA--EA BQL under the same payoff environment and outcome criterion.
The evolutionary basin has volume $V_E=1.00$ on the sampled grid. The empirical learning basin is $0.88$ for $\varepsilon$-IQL and $0.00$ for both scaled Boltzmann and SA--EA BQL. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration.
These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game--learning system. The shared-bike setting is a motivating application; the broader contribution is a framework for comparing population-level stability with the finite-sample accessibility of cooperation under specified multi-agent learning dynamics.
查看缓存全文
缓存时间: 2026/09/24 09:29
# Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence
Source: [https://arxiv.org/html/2609.27664](https://arxiv.org/html/2609.27664)
Yijie WangAffiliation:Liupanshui Normal UniversityAffiliation:Liupanshui, Guizhou, ChinaEmail:[wangyj@lpssy\.edu\.cn](mailto:)
###### Abstract
Cooperation emergence is a central problem in multi\-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others\. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite\-sample learning agents can reach the same outcome through local reward feedback\.
We study this distinction in a transparent three\-agent governance\-motivated game involving a government, a platform firm, and users\. We derive replicator dynamics for the fixed stage\-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial\-condition grid, and compare it with learning\-basin estimates for three decentralized value\-based learners\. The learning analysis uses independent Q\-learning withε\\varepsilon\-greedy action selection, scaled Boltzmann exploration, and SA–EA BQL under the same payoff environment and outcome criterion\.
The evolutionary basin has volumeVE=1\.00V\_\{E\}=1\.00on the sampled grid\. The empirical learning basin is0\.880\.88forε\\varepsilon\-IQL and0\.000\.00for both scaled Boltzmann and SA–EA BQL\. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration\.
These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game–learning system\. The shared\-bike setting is a motivating application; the broader contribution is a framework for comparing population\-level stability with the finite\-sample accessibility of cooperation under specified multi\-agent learning dynamics\.
Keywords:Multi\-agent reinforcement learning; evolutionary game theory; cooperation emergence; learning dynamics; replicator dynamics; multi\-agent systems\.
## 1Introduction
### 1\.1Motivation
Cooperation emergence is a central problem in multi\-agent systems\. Autonomous decision makers increasingly operate in settings in which their returns depend on the simultaneous choices of other agents, including distributed services, robot teams, platform ecosystems, and public–private governance arrangements\. In such settings, specifying a desirable collective outcome is only the first step\. A system must also support the decentralized process through which agents discover and sustain compatible behavior\. Multi\-agent systems research has long treated this challenge as a combination of strategic interaction, information limitations, and adaptation\[[24](https://arxiv.org/html/2609.27664#bib.bib7),[26](https://arxiv.org/html/2609.27664#bib.bib19)\]\. Multi\-agent reinforcement learning \(MARL\) makes the adaptation problem explicit: every learner changes its behavior from reward feedback while the effective environment changes because other learners are also adapting\[[2](https://arxiv.org/html/2609.27664#bib.bib6),[8](https://arxiv.org/html/2609.27664#bib.bib20)\]\.
The difficulty is especially clear in coordination problems\. An agent may have an incentive to select a cooperative action only when its counterparts make compatible choices, yet early observations may be dominated by uncoordinated joint actions\. Decentralized learners must therefore explore, evaluate rewards that depend on others, and form action preferences from histories that can differ substantially across runs\. The shared\-bike governance setting studied here provides a concrete motivating example\. Government regulation, firm compliance, and user responsiveness are mutually dependent, but the substantive question is not limited to this application\. It is whether a cooperative outcome that is strategically attractive can be reached by agents that start with limited information and learn only from local feedback\.
This distinction matters for the design and assessment of artificial multi\-agent systems\. A static incentive analysis may suggest that an outcome is viable, whereas a deployed learning process may still fail to establish the joint behavior required to realize it\. Conversely, a favorable learning run does not establish that a cooperative outcome is stable under changes in the population state\. A useful account of cooperation emergence therefore needs to distinguish the strategic landscape from the dynamics used to navigate that landscape\. This paper makes that distinction observable by evaluating population\-level stability and finite\-sample learning accessibility in the same fixed payoff environment\.
Decentralization makes this separation practically important\. A centralized controller can condition a joint decision on global information and directly enforce a coordinated policy\. Independent agents instead receive local rewards, possess their own action values, and may observe only the realized consequences of a joint action\. In a common\-payoff environment, this does not remove the coordination problem: an individually sampled action can be reasonable under a learner’s current value estimate while being incompatible with the action needed by the group\. The difficulty is amplified when a cooperative reward requires several agents to select their complementary actions within the same period\. Before those compatible actions have occurred repeatedly, learners may receive evidence that is sparse, variable, or dominated by the behavior induced by earlier exploration\. The problem is therefore one of convention formation, not only action optimization\. A system designer needs to know both whether the incentives favor cooperation and whether the specified adaptation process can form a convention that realizes those incentives\.
This framing also avoids two unhelpful simplifications\. First, it does not equate a cooperative equilibrium with a prediction that decentralized agents will coordinate in finite time\. Second, it does not treat a realized learning outcome as a definitive description of the strategic environment\. The two perspectives answer different questions and use different state variables\. Evolutionary analysis tracks the distribution of strategies in a population; independent reinforcement learning tracks each agent’s evolving estimates of action value\. Comparing them in a common payoff environment creates a controlled way to identify where their conclusions coincide and where they diverge\. Such a comparison is particularly useful for AI systems that are intended to operate without persistent central coordination, because implementation depends on the actual learning path rather than on equilibrium existence alone\.
### 1\.2Research gap
Evolutionary game theory offers a principled language for analyzing strategic stability\. Evolutionarily stable strategies and related stability concepts describe whether a strategy can resist invasion under specified strategic conditions\[[18](https://arxiv.org/html/2609.27664#bib.bib1),[33](https://arxiv.org/html/2609.27664#bib.bib10)\]\. Replicator dynamics then supplies a population\-level adjustment process in which strategy shares change according to payoff differences\[[29](https://arxiv.org/html/2609.27664#bib.bib9),[9](https://arxiv.org/html/2609.27664#bib.bib2)\]\. Population\-game formulations further clarify how equilibrium and dynamic behavior depend on the payoff structure and the chosen revision protocol\[[23](https://arxiv.org/html/2609.27664#bib.bib3)\]\. These tools are valuable for identifying cooperative regions of a game, but they do not by themselves specify how a finite collection of independently learning agents will acquire behavior from realized rewards\.
The theory of learning in games makes the adjustment process central rather than incidental\. Equilibrium selection can depend on experience, perturbations, beliefs, and the particular learning rule through which agents revise behavior\[[7](https://arxiv.org/html/2609.27664#bib.bib8),[34](https://arxiv.org/html/2609.27664#bib.bib12)\]\. In reinforcement\-learning settings, an agent evaluates actions from finite histories while other agents alter the reward contingencies it encounters\. This creates a distinction between a state that is attractive under a smooth population dynamic and a state that is reachable by a specified decentralized learning process\. The distinction is not merely terminological\. Population shares, individual action values, information flows, and sources of stochasticity differ across the two descriptions\.
The comparison should consequently be made at matched levels of specificity\. An evolutionary conclusion is conditional on the payoff model, the population state, and the revision dynamic used to define stability\. A learning conclusion is conditional on the learner, action\-selection rule, initialization, random seed, finite horizon, and criterion used to classify an outcome\. Neither conditional statement subsumes the other\. The relevant question is not whether one framework is more fundamental, but whether the accessible outcomes of a stated learning process correspond to the cooperative region indicated by a stated evolutionary process\. This distinction has received less attention than equilibrium characterization or algorithm benchmarking, even though it is essential when an AI system must form cooperation through decentralized adaptation\.
MARL provides a direct framework for examining this gap\. Independent learners update their own value estimates, commonly treating the evolving behavior of other agents as part of the environment\. Such decentralization is useful for studying cooperation, but it also creates non\-stationarity and coordination challenges\[[2](https://arxiv.org/html/2609.27664#bib.bib6),[17](https://arxiv.org/html/2609.27664#bib.bib17)\]\. Work on learning in cooperative multi\-agent systems shows that exploration, game structure, and observability can all affect which conventions are learned\[[3](https://arxiv.org/html/2609.27664#bib.bib15)\]\. Yet evolutionary stability is often used as an intuitive proxy for whether cooperation should emerge, even though the proxy does not specify the learning rule, finite horizon, initialization, or outcome criterion\.
The gap addressed here is therefore precise: does evolutionary stability imply learning accessibility when cooperation must emerge through finite\-sample MARL? We use*learning accessibility*to mean the empirical ability of a specified learning process to reach the pre\-defined cooperative outcome from a defined initial\-condition grid under a fixed experimental configuration\. This definition is deliberately narrower than a universal claim about learnability\. It allows strategic stability and learning accessibility to be compared without treating one as a substitute for the other\.
### 1\.3Research framework and contributions
We study the question in a transparent three\-agent game motivated by post\-subsidy shared\-bike scheduling\. Government, platform firm, and user agents each choose one of two actions, and their original stage payoffs connect regulation, compliance, and responsiveness\. First, we derive replicator dynamics from the expected payoff differences and evaluate the cooperative basin over a frozen symmetric grid\. Second, we keep the stage game unchanged and evaluate three decentralized value\-based learning dynamics:ε\\varepsilon\-greedy independent Q\-learning \(IQL\), scaled Boltzmann exploration, and SA–EA BQL\. Third, we compare their empirical learning basins with the evolutionary basin and use action coverage, policy entropy, and Q\-value separation as diagnostics of the observed divergence\. Figure[1](https://arxiv.org/html/2609.27664#S1.F1)summarizes this comparison\.
The study makes three contributions\. First, it develops an analytical framework that connects evolutionary stability with MARL accessibility while holding strategic incentives fixed\. The framework treats the shared\-bike setting as a motivating application rather than as evidence that its numerical outcomes generalize to every governance environment\. Second, it quantifies the difference between a sampled evolutionary basin and empirical learning basins under a common initial\-condition grid and outcome definition\. This separates a property of the replicator system from a property of the specified learning dynamics\. Third, it analyzes why exploration diversity alone does not guarantee cooperative emergence\. The diagnostics show that continued access to alternative actions can coexist with failure to sustain the compatible joint action required for cooperation in the frozen configuration\. Together, these contributions provide a bounded way to study how population\-level stability, decentralized learning, and coordination interact\.
This work is not intended as an algorithm benchmark study; instead, it investigates the relationship between equilibrium stability and accessibility under different learning dynamics\.
Figure 1:Conceptual framework of evolutionary stability and learning accessibility\. The same multi\-agent payoff environment is examined through two distinct dynamics\. The evolutionary branch updates population strategy frequencies through fitness comparison and replicator adjustment\. The MARL branch maps joint actions to realized payoffs, individual Q\-value updates, policy updates, and the next interaction\.
## 2Related Work
### 2\.1Evolutionary game theory
Evolutionary game theory links strategic interaction to dynamic selection\. The concept of an evolutionarily stable strategy formalizes resistance to strategically relevant deviations\[[18](https://arxiv.org/html/2609.27664#bib.bib1),[33](https://arxiv.org/html/2609.27664#bib.bib10)\]\. Taylor and Jonker established the connection between evolutionary stability and game dynamics through the replicator equation\[[29](https://arxiv.org/html/2609.27664#bib.bib9)\], while Hofbauer and Sigmund developed a broad mathematical treatment of evolutionary games and population dynamics\[[9](https://arxiv.org/html/2609.27664#bib.bib2)\]\. Subsequent work has extended evolutionary analysis beyond simple normal\-form settings\[[4](https://arxiv.org/html/2609.27664#bib.bib11)\], and population\-game theory has made explicit the relationship between payoff functions, revision protocols, and aggregate behavior\[[23](https://arxiv.org/html/2609.27664#bib.bib3)\]\.
This literature supplies the stability lens used in the present paper\. Our contribution is not a new equilibrium concept or a modification of replicator dynamics\. Instead, we use a fixed three\-agent payoff model to ask whether the cooperative region identified under population adjustment is also accessible under a distinct, finite\-sample learning dynamic\.
### 2\.2Learning in games
Learning\-in\-games research studies how adaptive behavior selects, approaches, or fails to approach strategically meaningful outcomes\. Fudenberg and Levine provide a foundational account of learning rules and equilibrium reasoning\[[7](https://arxiv.org/html/2609.27664#bib.bib8)\], and Young shows how stochastic adaptation can shape the emergence of conventions\[[34](https://arxiv.org/html/2609.27664#bib.bib12)\]\. In reinforcement\-learning formulations, the relevant state representation and interaction structure matter: Markov games provide one general framework for multi\-agent learning\[[14](https://arxiv.org/html/2609.27664#bib.bib14)\], while Nash Q\-learning analyzes value\-based learning in general\-sum stochastic games\[[10](https://arxiv.org/html/2609.27664#bib.bib16)\]\. Studies of individual Q\-learning in normal\-form games and of independent learning in multi\-agent settings further emphasize that convergence properties depend on the learning process rather than on payoffs alone\[[13](https://arxiv.org/html/2609.27664#bib.bib13),[17](https://arxiv.org/html/2609.27664#bib.bib17)\]\.
The present study is related to this tradition because it treats the path to a cooperative outcome as an object of analysis\. It differs in its comparison target: the empirical learning basin of a specified MARL procedure is evaluated alongside a basin generated by replicator dynamics under the same strategic incentives\.
### 2\.3Multi\-agent reinforcement learning
MARL addresses decentralized decision making when multiple adaptive agents act in a shared environment\. Early accounts distinguish independent from cooperative learning and identify the interaction between local updates and coordination\[[28](https://arxiv.org/html/2609.27664#bib.bib18),[3](https://arxiv.org/html/2609.27664#bib.bib15)\]\. Surveys organize the field around competing, cooperative, and mixed settings, as well as the non\-stationarity induced by simultaneously learning agents\[[2](https://arxiv.org/html/2609.27664#bib.bib6),[8](https://arxiv.org/html/2609.27664#bib.bib20),[30](https://arxiv.org/html/2609.27664#bib.bib25)\]\. Modern methods often use centralized information during training or structured value decomposition to improve coordination, including multi\-agent actor–critic methods\[[16](https://arxiv.org/html/2609.27664#bib.bib21)\], counterfactual policy gradients\[[6](https://arxiv.org/html/2609.27664#bib.bib22)\], and monotonic value\-factorization approaches\[[21](https://arxiv.org/html/2609.27664#bib.bib23)\]\. Broader theoretical overviews likewise distinguish the information structure, learning objective, and solution concept when characterizing MARL algorithms\[[35](https://arxiv.org/html/2609.27664#bib.bib24),[20](https://arxiv.org/html/2609.27664#bib.bib26)\]\.
Recent MARL research has also studied cooperative exploration, emergent roles, learned communication, and evaluation settings that make coordination challenges visible\[[15](https://arxiv.org/html/2609.27664#bib.bib29),[25](https://arxiv.org/html/2609.27664#bib.bib30),[22](https://arxiv.org/html/2609.27664#bib.bib31),[31](https://arxiv.org/html/2609.27664#bib.bib32),[5](https://arxiv.org/html/2609.27664#bib.bib33),[11](https://arxiv.org/html/2609.27664#bib.bib34),[12](https://arxiv.org/html/2609.27664#bib.bib35)\]\. These directions underscore that coordination can depend on exploration, representation, communication, and population structure\. They are complementary to the present controlled comparison, which does not introduce any of these mechanisms but instead asks how a fixed payoff environment behaves under specified decentralized learning dynamics\.
These studies establish that coordination is sensitive to learning architecture and information design\. Our paper does not introduce a new MARL algorithm or benchmark centralized methods against independent learners\. It instead uses transparent decentralized value\-based rules as diagnostic dynamics, asking whether the cooperative conclusion supplied by evolutionary analysis transfers to finite\-sample learning accessibility\. This focus also connects to the broader study of cooperation, in which repeated interaction and population processes can support cooperative conventions under conditions that must be stated explicitly\[[1](https://arxiv.org/html/2609.27664#bib.bib27),[19](https://arxiv.org/html/2609.27664#bib.bib28)\]\.
## 3Evolutionary Game Model
### 3\.1Three\-party governance game
We model post\-subsidy shared\-bike scheduling as a repeated interaction among a government, a platform firm, and users\. Each player has a binary action\. The government chooses active regulation \(g=1g=1\) or weak regulation \(g=0g=0\); the firm chooses compliant scheduling cooperation \(f=1f=1\) or non\-compliance \(f=0f=0\); and users choose responsive participation \(u=1u=1\) or non\-response \(u=0u=0\)\. Letxx,yy, andzzdenote the population shares choosingg=1g=1,f=1f=1, andu=1u=1, respectively\. This representation uses the evolutionary\-game perspective on strategic stability while retaining a tractable multi\-population payoff structure\[[18](https://arxiv.org/html/2609.27664#bib.bib1),[9](https://arxiv.org/html/2609.27664#bib.bib2)\]\.
The verified stage\-game payoff matrix is given in Table[1](https://arxiv.org/html/2609.27664#S3.T1)\. Regulation incurs costCgC\_\{g\}and yields the government a subsidy\-efficiency returnθS\\theta S\. A compliant firm receives the residual subsidy benefitδB\\delta Bbut bears compliance costcc; a non\-compliant firm obtains the private gainEEand is exposed to a sanctionFFdetected with probabilityλs\\lambda\_\{s\}under active regulation andλw\\lambda\_\{w\}under weak regulation\. User responsiveness produces baseline net utilitya−cua\-c\_\{u\}and may receive firm and government benefitsqqandrr\. The symbolsSS,LL,β\\beta,πe\\pi\_\{e\}, andα\\alpharetain their verified meanings in the implementation: public scheduling benefit, loss from non\-compliance, government responsiveness benefit, the firm’s operating payoff, and the firm–user coordination return\.
### 3\.2Definition: Learning Accessibility
LetMMdenote a specified learning mechanism and letCi=1C\_\{i\}=1if independent run or initial stateiireaches the pre\-defined cooperative outcome under the fixed evaluation criterion, andCi=0C\_\{i\}=0otherwise\. ForNNevaluated runs or initial states, we define the empirical learning\-accessibility measure as
AL\(M\)=1N∑i=1NI\(Ci=1\)\.A\_\{L\}\(M\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathrm\{I\}\(C\_\{i\}=1\)\.\(1\)This quantity records the observed fraction of successful learning trajectories for the stated mechanism, initialization design, finite horizon, and stochastic\-update procedure\. It is an empirical accessibility measure, not a universal metric of learnability\.
The evolutionary basin instead records whether a stable state attracts the population dynamics from the specified initial states\. The two quantities therefore refer to different adaptation processes\. Evolutionary dynamics updates strategy frequencies through population\-level payoff comparison and replicator adjustment\. Learning dynamics updates individual value estimates and policies from finite samples of joint interaction\. Consequently, an evolutionary basin need not equal the learning basin measured byAL\(M\)A\_\{L\}\(M\), even when both analyses use the same payoff environment\.
Table 1:Verified stage\-game payoffs\(ΠG,ΠF,ΠU\)\(\\Pi\_\{G\},\\Pi\_\{F\},\\Pi\_\{U\}\)\.
### 3\.3Replicator dynamics
For each population, strategy shares evolve in proportion to the payoff advantage of the active strategy\. This replicator formulation follows the standard payoff\-difference interpretation of evolutionary adjustment\[[29](https://arxiv.org/html/2609.27664#bib.bib9),[23](https://arxiv.org/html/2609.27664#bib.bib3)\]\. Taking expectations over the other two populations yields
x˙=x\(1−x\)\[−Cg\+θS\+\(λs−λw\)F\(1−y\)\]\.\\dot\{x\}=x\(1\-x\)\\left\[\-C\_\{g\}\+\\theta S\+\(\\lambda\_\{s\}\-\\lambda\_\{w\}\)F\(1\-y\)\\right\]\.\(2\)y˙=y\(1−y\)\[−c\+δB\+αz−E\+\{λw\+\(λs−λw\)x\}F\]\.\\dot\{y\}=y\(1\-y\)\\left\[\-c\+\\delta B\+\\alpha z\-E\+\\left\\\{\\lambda\_\{w\}\+\(\\lambda\_\{s\}\-\\lambda\_\{w\}\)x\\right\\\}F\\right\]\.\(3\)z˙=z\(1−z\)\(a−cu\+qy\+rx\)\.\\dot\{z\}=z\(1\-z\)\\left\(a\-c\_\{u\}\+qy\+rx\\right\)\.\(4\)These equations preserve the binary\-action boundaries and separate the incentives associated with regulation, firm compliance, and user responsiveness\. They are evaluated with the verified parameterization and numerical integration routine used throughout the project\.
### 3\.4Stability and basin evaluation
Local stability characterizes whether a perturbation near a candidate equilibrium decays under equations[2](https://arxiv.org/html/2609.27664#S3.E2)–[4](https://arxiv.org/html/2609.27664#S3.E4); it does not establish that a learning process will reach that equilibrium\. We therefore complement the local analysis with a numerical basin calculation\. Starting from the frozen symmetric grid
\(x0,y0,z0\)=\(p,p,p\),p∈\{0\.1,0\.2,…,0\.9\},\(x\_\{0\},y\_\{0\},z\_\{0\}\)=\(p,p,p\),\\qquad p\\in\\\{0\.1,0\.2,\\ldots,0\.9\\\},we integrate the replicator system under the incentive\-on regime \(α=80\\alpha=80\) and classify a trajectory as cooperative when it converges to the cooperative corner according to the project criterion\. All nine initial conditions converge to that corner, giving the evolutionary basin estimate
VE=99=1\.00\.V\_\{E\}=\\frac\{9\}\{9\}=1\.00\.This is a numerical statement for the sampled symmetric grid, rather than a claim about every point in the continuous state space\.
## 4Multi\-Agent Reinforcement Learning Framework
### 4\.1Multi\-agent environment
The learning environment implements the same verified three\-party stage game as Table[1](https://arxiv.org/html/2609.27664#S3.T1)\. At each round, the government, firm, and user independently select one of their two actions, producing a joint action\(g,f,u\)\(g,f,u\)and the corresponding vector of original stage rewards\(ΠG,ΠF,ΠU\)\(\\Pi\_\{G\},\\Pi\_\{F\},\\Pi\_\{U\}\)\. This decentralized formulation follows the multi\-agent view in which other adaptive agents contribute to the learning environment\[[14](https://arxiv.org/html/2609.27664#bib.bib14),[35](https://arxiv.org/html/2609.27664#bib.bib24)\]\. No reward shaping, auxiliary coordination bonus, or altered payoff is introduced for the learning experiments\.
The static basin experiments use a fixed favourable subsidy state \(δ=1\\delta=1\) and repeatedly sample the same stage\-game incentives\. The environment interface also records subsidy and behavioural summaries for dynamic diagnostic runs, but the Phase G basin comparison is deliberately a fixed\-state accessibility test\. The initial\-preference grid is imposed only at the first decision; subsequent behaviour is determined by each learner’s own updates and exploration rule\.
### 4\.2Independent Q\-learning
Each playeri∈\{G,F,U\}i\\in\\\{G,F,U\\\}maintains action valuesQi\(ai\)Q\_\{i\}\(a\_\{i\}\)and treats the actions of the other two players as part of the environment\. This value\-based update is a repeated\-stage specialization of the Q\-learning tradition\[[32](https://arxiv.org/html/2609.27664#bib.bib5),[27](https://arxiv.org/html/2609.27664#bib.bib4)\]\. After observing its original stage rewardri,tr\_\{i,t\}, it updates the selected action according to
Qi,t\+1\(ai,t\)=Qi,t\(ai,t\)\+ηt\[ri,t−Qi,t\(ai,t\)\],Q\_\{i,t\+1\}\(a\_\{i,t\}\)=Q\_\{i,t\}\(a\_\{i,t\}\)\+\\eta\_\{t\}\\left\[r\_\{i,t\}\-Q\_\{i,t\}\(a\_\{i,t\}\)\\right\],\(5\)while leaving the unselected action unchanged\. This is a repeated\-stage formulation: no continuation\-value term is used in equation[5](https://arxiv.org/html/2609.27664#S4.E5)\. Learning rates, run lengths, random seeds, and other implementation settings are those recorded in the frozen experiment configuration rather than additional manuscript assumptions\.
Theε\\varepsilon\-greedy independent Q\-learning \(IQL\) benchmark chooses a currently highest\-valued action except for its configured exploration probability\. Independent learners are a deliberately transparent MARL baseline, but their local reward updates can encounter coordination and non\-stationarity challenges\[[2](https://arxiv.org/html/2609.27664#bib.bib6),[17](https://arxiv.org/html/2609.27664#bib.bib17)\]\. Its role is descriptive: it provides a reference learning dynamic for comparing basin accessibility, not an optimized policy\.
### 4\.3Scale\-aware Boltzmann exploration
The scale\-aware Boltzmann learner samples an action using a softmax transformation of action values after normalizing for the payoff scale,
Pri\(a∣s\)=exp\(Qi\(s,a\)/\[τidi\]\)∑a′∈\{0,1\}exp\(Qi\(s,a′\)/\[τidi\]\),\\Pr\_\{i\}\(a\\mid s\)=\\frac\{\\exp\\\!\\left\(Q\_\{i\}\(s,a\)/\[\\tau\_\{i\}\\,d\_\{i\}\]\\right\)\}\{\\sum\_\{a^\{\\prime\}\\in\\\{0,1\\\}\}\\exp\\\!\\left\(Q\_\{i\}\(s,a^\{\\prime\}\)/\[\\tau\_\{i\}\\,d\_\{i\}\]\\right\)\},\(6\)whereτi\\tau\_\{i\}is the configured temperature anddid\_\{i\}is the payoff\-scale normalization used by the implementation\. The normalization makes the softmax response interpretable when the three players’ payoff ranges differ\. It is evaluated as a diagnostic comparison with IQL, not presented as a superior algorithm\.
### 4\.4Adaptive exploration variant \(SA–EA BQL\)
The SA–EA BQL adaptive exploration variant combines scale\-aware action selection with an exploration adjustment based on the learner’s action\-value separation\. In generic form, its temperature is allowed to depend on a Q\-value gapΔQi\\Delta Q\_\{i\},
τi=τ0,i\+kiΔQi,\\tau\_\{i\}=\\tau\_\{0,i\}\+k\_\{i\}\\Delta Q\_\{i\},\(7\)using the fixed implementation configuration\. This diagnostic comparison mechanism is included to examine whether adaptive exploration changes learning accessibility\. It is not a claim that adaptive exploration guarantees cooperation, nor are its diagnostic outcomes interpreted as causal evidence beyond the frozen comparison\.
## 5Experimental Protocol
The learning\-basin comparison uses a fixed protocol so that differences across learning mechanisms are not confounded with changes in the environment or evaluation design\. Table[2](https://arxiv.org/html/2609.27664#S5.T2)records the configuration used throughout the reported comparison\.
Table 2:Fixed experimental protocol for the learning\-basin comparison\.All mechanisms interact with the same fixed stage\-game payoffs and are evaluated using the same pre\-specified cooperative\-outcome criterion\. The protocol is reported for reproducibility; it does not represent an additional parameter search or an algorithm\-tuning exercise\.
## 6Results
### 6\.1Evolutionary stability on the sampled grid
Under the incentive\-on specification \(α=80\\alpha=80\), numerical integration of the replicator system converges to the cooperative corner from each of the nine frozen symmetric starting points\. The resulting sampled evolutionary basin is thereforeVE=1\.00V\_\{E\}=1\.00\(9/9\)\. This result establishes evolutionary accessibility for the specified grid and parameterization; it is not a global basin proof\.
Figure 2:Comparison between evolutionary stability and learning accessibility\. Although the replicator dynamics reaches the cooperative basin for all evaluated initial conditions, finite\-sample reinforcement learning exhibits learning\-mechanism\-dependent accessibility\.Figure 3:Replicator stability on the frozen symmetric initial\-condition grid\. \(a\) Three verified trajectories project toward the cooperative corner\(1,1,1\)\(1,1,1\)\. \(b\) Each evaluated value ofppconverges cooperatively; this discrete classification does not infer a continuous basin\.
### 6\.2Learning accessibility
The corresponding reinforcement\-learning experiment contains 1,350 runs: nine initial\-preference settings, 50 random seeds per setting, and three learning rules\. A run is classified using the frozen cooperative\-outcome criterion\. Epsilon\-greedy IQL reaches cooperation in 396 of 450 runs, yielding a learning\-basin estimate of0\.8800\.880with bootstrap 95% confidence interval\[0\.849,0\.909\]\[0\.849,0\.909\]\. The estimated tipping preference isp=0\.1p=0\.1under the recorded rule\.
By contrast, neither scale\-aware Boltzmann nor SA–EA BQL reaches the cooperative outcome in any of its 450 runs: both have an estimated learning basin of0\.0000\.000\. Thus, within this fixed design, evolutionary convergence does not imply that every decentralized learning rule accesses the same cooperative outcome\. Figure[2](https://arxiv.org/html/2609.27664#S6.F2)distinguishes the evolutionary result from the learning\-mechanism\-dependent estimates; Figures[4](https://arxiv.org/html/2609.27664#S6.F4)and[5](https://arxiv.org/html/2609.27664#S6.F5)give the disaggregated learning curves and the theory–learning gap\.
Figure 4:Learning\-basin curves for the three frozen reinforcement\-learning comparisons\. Points are observed success probabilities at the nine evaluated symmetric initial preferences; no interpolation is used\.Figure 5:Theory–learning gap in cooperative basin volume\. The interval forε\\varepsilon\-IQL is the frozen bootstrap 95% confidence interval; the two zero estimates are exact outcomes in their respective 450\-run grids\.
### 6\.3Exploration mechanism diagnosis
Phase H records action trajectories, action coverage, policy entropy, and Q\-value separation for 50 seeds per learning rule over 20,000 rounds\. The diagnostic traces show that the two Boltzmann\-based learners retain substantial action diversity rather than collapsing immediately into a single deterministic policy\. Their Q\-value gaps are also nonzero\. These observations rule out a simple account in which the zero\-basin results arise solely from premature deterministic\-policy collapse\.
A more cautious interpretation is that persistent stochastic action selection can impede convention formation in this coordination environment: learners may continue to visit both actions without jointly sustaining the mutually cooperative action profile\. This is a mechanism diagnosis tied to the observed trajectories, not evidence that exploration is intrinsically harmful or that one algorithm is universally preferable\.
Figure 6:Exploration diversity alone does not determine cooperative emergence\. Points summarize the three players within each seed; error bars show empirical 2\.5th–97\.5th percentiles across 50 seeds\. Action coverage is measured over the recorded run, whereas entropy and Q\-value separation use the terminal logged episode\.
## 7Discussion
### 7\.1Evolutionary stability versus learning accessibility
The central finding is that cooperative stability under the replicator system and cooperative accessibility under finite\-sample reinforcement learning are related but distinct properties\. On the frozen symmetric grid, the evolutionary analysis reaches the cooperative basin for all evaluated initial conditions\. The learning comparison, however, yields method\-dependent basin estimates under the same underlying strategic environment\. This contrast does not imply that either framework is defective\. Instead, it reflects that they instantiate different dynamical processes\. Replicator dynamics updates population shares according to payoff advantages that are evaluated against the current population state\. The learning agents update action values from realized, stochastic stage rewards while their counterparts are simultaneously adapting\. The state variables, information available at each update, and effective averaging of payoff feedback are consequently different\. A cooperative fixed point that attracts nearby population states can therefore remain difficult for decentralized learners to reach from finite histories\.
This distinction clarifies what is and is not established by the present results\. The evolutionary basin estimate of1\.001\.00is a numerical result for nine sampled symmetric initial conditions under the incentive\-on parameterization; it is not a claim that every continuous initial state converges cooperatively\. Likewise, the learning\-basin estimates describe the observed outcomes of the specified finite grids, seeds, learning rates, and exploration rules\. Within those boundaries, strategic incentives sufficient for evolutionary convergence are not sufficient to ensure that every learning mechanism forms the corresponding cooperative convention\. Figure[2](https://arxiv.org/html/2609.27664#S6.F2)makes this separation visible, while Figures[3](https://arxiv.org/html/2609.27664#S6.F3)–[5](https://arxiv.org/html/2609.27664#S6.F5)distinguish the population calculation, learning curves, and IQL uncertainty\. Cooperation should be assessed not only by its stability in a game\-theoretic model, but also by whether a stated adaptation process can access it\.
One interpretation is that learning accessibility depends on a sequence of compatible events rather than solely on the eventual payoff ordering of actions\. A learner must sample an action, experience an informative reward, update its value estimate in a favorable direction, and encounter counterparts whose actions make that experience repeatable\. In a coordination environment, each of these steps is coupled to the other agents’ changing policies\. A temporary cooperative joint action may not be observed often enough, or consistently enough, to become a stable convention in all learning processes\. Conversely, an evolutionary equation effectively summarizes payoff advantages at the population level and does not represent the finite sequence of joint actions through which individual values are acquired\. The results therefore suggest that the path to cooperation is itself part of the scientific object\. Future theoretical work could seek conditions under which properties of a payoff game and properties of a particular learning update are aligned, rather than assuming that equilibrium attraction transfers directly from one dynamic to the other\.
This distinction is consistent with the broader view that stability and learning are properties of specified dynamics, rather than of a payoff table in isolation\[[9](https://arxiv.org/html/2609.27664#bib.bib2),[7](https://arxiv.org/html/2609.27664#bib.bib8)\]\.
### 7\.2Exploration and cooperative convention formation
The mechanism diagnostics refine a simple explanation of the zero\-basin outcomes\. It would be tempting to attribute non\-cooperation by the Boltzmann\-based learners to insufficient exploration or to premature collapse into one deterministic action\. The recorded Phase H traces do not support that simple account\. As summarized in Figure[6](https://arxiv.org/html/2609.27664#S6.F6), the Boltzmann\-based learners retain broader action diversity than theε\\varepsilon\-greedy benchmark on the reported minority\-action, terminal\-entropy, and Q\-value\-gap diagnostics\. Their nonzero Q\-value separation also indicates that the learners are not merely indifferent between their actions\. Yet these properties coexist with an inability to sustain the cooperative outcome under the frozen basin criterion\.
This pattern suggests a more specific, and more cautious, interpretation\. Exploration diversity can preserve access to alternative actions, but access is not the same as coordinated commitment\. In the present game, a cooperative convention requires compatible actions by all three agents\. Continued stochastic action selection may repeatedly interrupt the joint profile that would otherwise reinforce cooperative values, particularly when each learner treats the decisions of the others as part of a changing environment\. The diagnostic evidence is consistent with persistent exploration impeding convention formation in this particular design, but it does not identify a unique causal mechanism\. Other features of the learning process, including relative payoff scales, update trajectories, early joint\-action histories, and the form of value normalization, may also contribute\. The comparison should therefore not be interpreted as evidence that Boltzmann exploration is generally unfavorable, or that lower action diversity is generally desirable\.
The distinction between diversity and coordination has practical consequences for how exploration is evaluated in cooperative MARL\. A high entropy or minority\-action fraction can be useful evidence that agents have not become behaviorally trapped, but these metrics do not measure whether agents are learning compatible policies\. Similarly, a nonzero value gap documents action differentiation without indicating whether the differentiated preferences support the same joint convention across agents\. When cooperation requires synchronized behavior, a diagnostic suite should therefore combine individual\-level indicators with an outcome\-level measure of sustained joint coordination\. The learning\-basin criterion used here provides one such outcome\-level measure for a fixed environment\. More detailed analyses could track the temporal formation and disruption of candidate conventions, distinguish transient coordination from persistent coordination, and examine which joint\-action histories precede successful runs\. Such analyses would help separate exploration that is informative for coordination from exploration that remains behaviorally diverse without building a shared convention\.
This coordination challenge is characteristic of multi\-agent learning, in which each learner faces an environment altered by the adaptation of other learners\[[2](https://arxiv.org/html/2609.27664#bib.bib6)\]\.
### 7\.3Implications for multi\-agent systems
The results suggest several implications for the design and assessment of cooperative multi\-agent systems\. First, equilibrium analysis alone is an incomplete basis for evaluating whether a proposed incentive structure will yield cooperation in a learning population\. It remains valuable for identifying strategically stable states and for understanding how payoff components shape those states\. However, a stable solution does not by itself reveal the learning trajectory by which bounded agents will approach it\. Evaluations of cooperative systems may therefore benefit from reporting both a strategic analysis and a learning\-accessibility analysis, with the latter specifying the initialization, update rule, exploration process, and finite evaluation horizon\. This paired perspective is relevant wherever system designers rely on decentralized adaptation rather than direct coordination\.
Second, mechanism design and learning design should be considered jointly\. A policy instrument, reward structure, or sanction may create a cooperative equilibrium while leaving the learning path fragile or inaccessible for some adaptive agents\. Conversely, a learning rule can alter which strategically available outcomes are observed in finite time without changing the underlying game\. The present framework does not prescribe a universal intervention, but it provides a way to diagnose this mismatch\. A mechanism can be assessed by the basin it creates under evolutionary adjustment and by the extent to which specified learners access that basin\. When the two assessments diverge, the discrepancy points to a design question: whether to modify incentives, information, communication, or the adaptation process itself\. The answer will depend on the application and should be tested rather than inferred from the present three\-agent model\.
Third, learning accessibility is useful as a reporting dimension rather than as a post hoc explanation for an unsuccessful run\. Finite time, finite data, and exploration are operational constraints in many multi\-agent settings\. Basin\-based comparisons connect equilibrium predictions to observed learning outcomes through a common outcome definition while preserving their distinct dynamics\. Claims about cooperative behavior should therefore state the adjustment process under which they are expected to hold\.
Accordingly, strategic and learning dynamics should be reported as complementary objects of analysis when assessing decentralized multi\-agent systems\[[23](https://arxiv.org/html/2609.27664#bib.bib3),[24](https://arxiv.org/html/2609.27664#bib.bib7)\]\.
### 7\.4Relation to learning in games
The comparison can also be positioned within learning in games\. Replicator dynamics describes population\-level adaptation: strategy shares change according to expected payoff differences at the current population state\. MARL, by contrast, describes individual value\-based adaptation: each agent updates action values from realized rewards while the behavior of other agents changes the data\-generating process\. These dynamics can agree in some settings, but their agreement is not guaranteed by the existence of a stable equilibrium\.
This difference helps explain why a cooperative equilibrium may be stable without automatically becoming a reachable learning outcome\. A population equation aggregates payoff advantages, whereas a finite learner must experience compatible joint actions often enough to form and retain a favorable value ordering\. The learning\-in\-games literature similarly treats the adjustment rule and the path of adaptation as part of equilibrium discovery rather than as details that can be omitted\[[7](https://arxiv.org/html/2609.27664#bib.bib8),[34](https://arxiv.org/html/2609.27664#bib.bib12),[13](https://arxiv.org/html/2609.27664#bib.bib13)\]\. In the present fixed configuration, the basin comparison makes this distinction measurable without claiming that one learning rule is generally superior to another\.
### 7\.5Limitations and future directions
Several limitations delimit the scope of the present inference\. The current study focuses on a controlled analytical setting: one transparent three\-agent governance game with binary actions and fixed payoff incentives\. Its value is that it isolates a tractable interaction among regulation, compliance, and responsiveness, but its structure does not represent every form of multi\-agent cooperation\. Larger populations, heterogeneous agent types, network interactions, and richer action spaces may generate different relationships between evolutionary stability and learning accessibility\. The central within\-design comparison remains informative for the specified game, but its transportability to other strategic structures requires additional tests\.
Second, the numerical results depend on finite parameter configurations, a sampled symmetric initial\-condition grid, and a fixed evaluation horizon\. The evolutionary basin is therefore a grid\-based numerical result, and the learning estimates are finite\-sample summaries rather than exhaustive characterizations\. Broader parameter sweeps, asymmetric initial conditions, and sensitivity analyses could test whether the gap persists across other payoff relationships and learning preferences\.
Third, the study does not include human behavioral experiments\. The agents are computational learners whose reward updates and exploration rules are specified by the model\. The results consequently address dynamical accessibility in a MARL setting, not how human participants would respond to regulatory, organizational, or user incentives\. Human experiments, field observations, or hybrid designs could test whether analogous coordination barriers arise when learning is shaped by social beliefs, communication, or institutional knowledge not represented here\.
Finally, the comparison does not establish that any reinforcement\-learning algorithm dominates across settings\. The three learners are diagnostic instruments applied to a common environment, not a comprehensive benchmark suite\. The two zero\-basin outcomes and the positive IQL basin should not be generalized beyond the frozen configuration\. Future extensions may consider continuous state spaces, deep function approximation, larger agent populations, and richer communication mechanisms\. A useful next step is a theoretical characterization of learning basins under these extensions, including when information sharing, commitment devices, or alternative update structures change accessibility\.
## 8Conclusion
This study examines a simple but consequential question for cooperation emergence in multi\-agent systems: does a cooperative state that is stable under evolutionary dynamics necessarily emerge when decentralized agents learn from finite, stochastic experience? We addressed this question by combining a three\-agent evolutionary game with controlled multi\-agent reinforcement\-learning comparisons\. The governance setting supplies a concrete interaction among regulation, firm compliance, and user responsiveness, but the analytical objective is broader\. It is to separate a strategic property of a game from a dynamical property of a learning process\. The resulting framework places evolutionary game theory and MARL in a common environment while preserving the differences between population\-level strategy adjustment and individual value\-based adaptation\.
The evolutionary analysis identifies a cooperative state that is accessible from all nine evaluated symmetric initial conditions in the incentive\-on specification, producing a sampled basin estimate ofVE=1\.00V\_\{E\}=1\.00\. The reinforcement\-learning results do not mirror this outcome uniformly\. Under the frozen learning\-basin design,ε\\varepsilon\-greedy IQL reaches the cooperative criterion in 396 of 450 runs, whereas the scale\-aware Boltzmann and SA–EA BQL comparisons do not reach that criterion in their respective grids\. These estimates should be interpreted within the stated game, initial\-preference grid, seeds, and finite evaluation horizon\. They do not establish a global property of any algorithm\. Nevertheless, their contrast with the evolutionary calculation provides direct evidence that a stable cooperative state can be learning\-mechanism\-dependent in finite\-sample decentralized adaptation\.
The mechanism diagnosis further qualifies how this gap should be interpreted\. The Boltzmann\-based learners retain action diversity and nonzero Q\-value separation in the recorded diagnostic runs\. Their lack of cooperative convergence therefore cannot be reduced to a simple story of insufficient exploration or immediate deterministic\-policy collapse\. Instead, the evidence suggests that exploration diversity alone does not ensure the formation of a stable cooperative convention\. In a coordination problem, learners must not only preserve alternative actions; they must also repeatedly realize compatible joint actions long enough for their individual updates to support the same collective pattern\. Persistent stochasticity may interrupt this process in the present configuration, although the current evidence does not isolate a unique causal route\.
The main conclusion is consequently bounded but general in its conceptual implication: stable cooperation in evolutionary dynamics does not automatically emerge from finite\-sample reinforcement learning\. Equilibrium analysis remains useful for identifying strategically viable cooperative states, and MARL analysis remains useful for observing how specified agents adapt to realized feedback\. Neither perspective substitutes for the other\. Jointly reporting an evolutionary basin and a learning basin makes the difference visible and provides a practical way to assess whether a cooperative mechanism is not only stable in principle but accessible to the learning process expected to implement it\.
Several directions follow from this boundary\. Future work could test larger and more heterogeneous populations, allow richer communication or commitment mechanisms, and characterize learning basins analytically rather than only through finite grids\. It could also compare computational learning with human behavior in settings where beliefs, norms, and institutional information affect coordination\. These extensions are needed before broad claims about governance systems or MARL algorithms can be made\. Within the present design, however, the results clarify a central point: explaining cooperation requires attention both to the incentives that stabilize an outcome and to the learning dynamics that determine whether agents can reach it\.
## Appendix AMathematical details
### A\.1Derivation of the replicator dynamics
For a binary\-action population with active\-strategy shareww, letπ1\\pi\_\{1\}andπ0\\pi\_\{0\}denote the expected payoffs of the active and inactive actions\. The average payoff isπ¯=wπ1\+\(1−w\)π0\\bar\{\\pi\}=w\\pi\_\{1\}\+\(1\-w\)\\pi\_\{0\}\. The standard binary replicator identity is therefore
w˙=w\(π1−π¯\)=w\(1−w\)\(π1−π0\)\.\\dot\{w\}=w\(\\pi\_\{1\}\-\\bar\{\\pi\}\)=w\(1\-w\)\(\\pi\_\{1\}\-\\pi\_\{0\}\)\.Applying this identity to each population and taking expectations over the other two strategy shares yields the three equations in Section 2\.
For the government, active regulation differs from weak regulation only through its direct regulatory return and the sanction applied when the firm is non\-compliant\. Hence,
πG\(g=1\)−πG\(g=0\)=−Cg\+θS\+\(λs−λw\)F\(1−y\)\.\\pi\_\{G\}\(g=1\)\-\\pi\_\{G\}\(g=0\)=\-C\_\{g\}\+\\theta S\+\(\\lambda\_\{s\}\-\\lambda\_\{w\}\)F\(1\-y\)\.Substitution into the binary identity withw=xw=xgives equation[2](https://arxiv.org/html/2609.27664#S3.E2)\.
For the firm, compliance replaces the private gainEEand avoided sanction with the compliance cost and residual subsidy benefit; the firm–user coordination return is received when users are responsive\. Averaging the sanction term over the government share gives
πF\(f=1\)−πF\(f=0\)=−c\+δB\+αz−E\+\{λw\+\(λs−λw\)x\}F\.\\pi\_\{F\}\(f=1\)\-\\pi\_\{F\}\(f=0\)=\-c\+\\delta B\+\\alpha z\-E\+\\left\\\{\\lambda\_\{w\}\+\(\\lambda\_\{s\}\-\\lambda\_\{w\}\)x\\right\\\}F\.Usingw=yw=ygives equation[3](https://arxiv.org/html/2609.27664#S3.E3)\.
For users, non\-response has zero payoff in the specified stage game\. Responsive participation has the baseline net returna−cua\-c\_\{u\}, the firm\-related returnqqwith probabilityyy, and the government\-related returnrrwith probabilityxx\. Thus,
πU\(u=1\)−πU\(u=0\)=a−cu\+qy\+rx\.\\pi\_\{U\}\(u=1\)\-\\pi\_\{U\}\(u=0\)=a\-c\_\{u\}\+qy\+rx\.Usingw=zw=zgives equation[4](https://arxiv.org/html/2609.27664#S3.E4)\. These derivations introduce no additional behavioral assumptions; they only rewrite the payoff differences already encoded in Table[1](https://arxiv.org/html/2609.27664#S3.T1)\.
## Appendix BReinforcement\-learning algorithm details
### B\.1Algorithm 1: Independent Q\-learning training procedure
For each run, initialize each agent’s two action values from the fixed initial\-preference condition and retain the run’s assigned random seed\. Then repeat the following procedure for the configured number of episodes\.
1. 1\.For each agenti∈\{G,F,U\}i\\in\\\{G,F,U\\\}, select a binary actionai,ta\_\{i,t\}using that learner’s configured action\-selection rule\.
2. 2\.Execute the joint action\(g,f,u\)\(g,f,u\)in the unchanged stage game and observe the original individual rewardri,tr\_\{i,t\}for every agent\.
3. 3\.Update only the selected action value according to equation[5](https://arxiv.org/html/2609.27664#S4.E5); leave the unselected action value unchanged\.
4. 4\.Continue until the fixed episode horizon is reached\. Classify the completed run using the pre\-specified learning\-basin criterion\.
This procedure is independent Q\-learning: each agent maintains its own values and does not observe, communicate, or centrally optimize the values of the other agents\. Because the environment is repeated\-stage, the update has no continuation\-value term\.
### B\.2Alternative action\-selection rules
The scaled Boltzmann learner uses the normalized softmax probability in equation[6](https://arxiv.org/html/2609.27664#S4.E6)\. Its temperature and payoff\-scale normalization are the fixed values used in the frozen implementation\. The SA–EA BQL adaptive exploration variant uses the same scale\-aware choice rule while adjusting its temperature from the action\-value separation according to equation[7](https://arxiv.org/html/2609.27664#S4.E7)\. These alternatives are diagnostic comparison mechanisms included to examine whether adaptive exploration changes learning accessibility\. Their description does not imply that either rule is generally superior or more appropriate outside the present fixed configuration\.
## Appendix CExperimental configuration and reproducibility
Table[3](https://arxiv.org/html/2609.27664#A3.T3)records the fixed configuration used for the learning\-basin comparison\. It is provided to make the reported finite\-sample estimates reproducible and should not be read as an additional parameter search\.
Table 3:Fixed configuration for the learning\-basin comparison\.Each algorithm was evaluated over the same grid and seed structure in the unchanged binary\-action payoff environment\. The learning\-basin criterion is the outcome definition used in the main text: it records whether a run reaches the specified cooperative outcome under the fixed evaluation design\.
## References
- \[1\]\(1984\)The evolution of cooperation\.Basic Books,New York\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p3.1)\.
- \[2\]L\. Buşoniu, R\. Babuška, and B\. De Schutter\(2008\)A comprehensive survey of multiagent reinforcement learning\.IEEE Transactions on Systems, Man, and Cybernetics, Part C: Applications and Reviews38\(2\),pp\. 156–172\.Cited by:[§1\.1](https://arxiv.org/html/2609.27664#S1.SS1.p1.1),[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p4.1),[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1),[§4\.2](https://arxiv.org/html/2609.27664#S4.SS2.p2.1),[§7\.2](https://arxiv.org/html/2609.27664#S7.SS2.p4.1)\.
- \[3\]C\. Claus and C\. Boutilier\(1998\)The dynamics of reinforcement learning in cooperative multiagent systems\.InProceedings of the Fifteenth National Conference on Artificial Intelligence,pp\. 746–752\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p4.1),[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[4\]R\. Cressman\(2003\)Evolutionary dynamics and extensive form games\.MIT Press,Cambridge, MA\.Cited by:[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1)\.
- \[5\]J\. Foerster, I\. A\. Assael, N\. de Freitas, and S\. Whiteson\(2016\)Learning to communicate with deep multi\-agent reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.29\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[6\]J\. Foerster, G\. Farquhar, T\. Afouras, N\. Nardelli, and S\. Whiteson\(2018\)Counterfactual multi\-agent policy gradients\.InProceedings of the Thirty\-Second AAAI Conference on Artificial Intelligence,Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[7\]D\. Fudenberg and D\. K\. Levine\(1998\)The theory of learning in games\.MIT Press,Cambridge, MA\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1),[§7\.1](https://arxiv.org/html/2609.27664#S7.SS1.p4.1),[§7\.4](https://arxiv.org/html/2609.27664#S7.SS4.p2.1)\.
- \[8\]P\. Hernandez\-Leal, B\. Kartal, and M\. E\. Taylor\(2019\)A survey and critique of multiagent deep reinforcement learning\.Autonomous Agents and Multi\-Agent Systems33,pp\. 750–797\.Cited by:[§1\.1](https://arxiv.org/html/2609.27664#S1.SS1.p1.1),[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[9\]J\. Hofbauer and K\. Sigmund\(1998\)Evolutionary games and population dynamics\.Cambridge University Press,Cambridge\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.27664#S3.SS1.p1.1),[§7\.1](https://arxiv.org/html/2609.27664#S7.SS1.p4.1)\.
- \[10\]J\. Hu and M\. P\. Wellman\(2003\)Nash Q\-learning for general\-sum stochastic games\.Journal of Machine Learning Research4,pp\. 1039–1069\.Cited by:[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1)\.
- \[11\]J\. Jiang and Z\. Lu\(2018\)Learning attentional communication for multi\-agent cooperation\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[12\]J\. Z\. Leibo, V\. Zambaldi, M\. Lanctot, J\. Marecki, and T\. Graepel\(2017\)Multi\-agent reinforcement learning in sequential social dilemmas\.InProceedings of the Sixteenth Conference on Autonomous Agents and Multiagent Systems,pp\. 464–473\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[13\]D\. S\. Leslie and E\. J\. Collins\(2005\)Individual Q\-learning in normal form games\.SIAM Journal on Control and Optimization44\(2\),pp\. 495–514\.Cited by:[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1),[§7\.4](https://arxiv.org/html/2609.27664#S7.SS4.p2.1)\.
- \[14\]M\. L\. Littman\(1994\)Markov games as a framework for multi\-agent reinforcement learning\.InProceedings of the Eleventh International Conference on Machine Learning,pp\. 157–163\.Cited by:[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.27664#S4.SS1.p1.1)\.
- \[15\]I\. Liu, U\. Jain, R\. A\. Yeh, and A\. G\. Schwing\(2021\)Cooperative exploration for multi\-agent deep reinforcement learning\.InProceedings of the 38th International Conference on Machine Learning,pp\. 6826–6836\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[16\]R\. Lowe, Y\. Wu, A\. Tamar, J\. Harb, P\. Abbeel, and I\. Mordatch\(2017\)Multi\-agent actor\-critic for mixed cooperative\-competitive environments\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[17\]L\. Matignon, G\. J\. Laurent, and N\. Le Fort\-Piat\(2012\)Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems\.The Knowledge Engineering Review27\(1\),pp\. 1–31\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p4.1),[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.27664#S4.SS2.p2.1)\.
- \[18\]J\. Maynard Smith\(1982\)Evolution and the theory of games\.Cambridge University Press,Cambridge\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.27664#S3.SS1.p1.1)\.
- \[19\]M\. A\. Nowak\(2006\)Five rules for the evolution of cooperation\.Science314\(5805\),pp\. 1560–1563\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p3.1)\.
- \[20\]F\. A\. Oliehoek and C\. Amato\(2016\)A concise introduction to decentralized POMDPs\.Springer,Cham\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[21\]T\. Rashid, M\. Samvelyan, C\. Schroeder de Witt, G\. Farquhar, J\. Foerster, and S\. Whiteson\(2018\)QMIX: monotonic value function factorisation for deep multi\-agent reinforcement learning\.InProceedings of the Thirty\-Fifth International Conference on Machine Learning,pp\. 4295–4304\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[22\]M\. Samvelyan, T\. Rashid, C\. Schroeder de Witt, G\. Farquhar, N\. Nardelli, T\. G\. J\. Rudner, C\. Hung, P\. H\. S\. Torr, J\. Foerster, and S\. Whiteson\(2019\)The StarCraft multi\-agent challenge\.InProceedings of the Eighteenth International Conference on Autonomous Agents and Multiagent Systems,pp\. 2186–2188\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[23\]W\. H\. Sandholm\(2010\)Population games and evolutionary dynamics\.MIT Press,Cambridge, MA\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.27664#S3.SS3.p1.1),[§7\.3](https://arxiv.org/html/2609.27664#S7.SS3.p4.1)\.
- \[24\]Y\. Shoham and K\. Leyton\-Brown\(2008\)Multiagent systems: algorithmic, game\-theoretic, and logical foundations\.Cambridge University Press,Cambridge\.Cited by:[§1\.1](https://arxiv.org/html/2609.27664#S1.SS1.p1.1),[§7\.3](https://arxiv.org/html/2609.27664#S7.SS3.p4.1)\.
- \[25\]K\. Son, D\. Kim, W\. J\. Kang, D\. Hostallero, and Y\. Yi\(2019\)QTRAN: learning to factorize with transformation for cooperative multi\-agent reinforcement learning\.InProceedings of the 36th International Conference on Machine Learning,pp\. 5887–5896\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[26\]P\. Stone and M\. Veloso\(2000\)Multiagent systems: a survey from a machine learning perspective\.Autonomous Robots8\(3\),pp\. 345–383\.Cited by:[§1\.1](https://arxiv.org/html/2609.27664#S1.SS1.p1.1)\.
- \[27\]R\. S\. Sutton and A\. G\. Barto\(2018\)Reinforcement learning: an introduction\.2nd edition,MIT Press,Cambridge, MA\.Cited by:[§4\.2](https://arxiv.org/html/2609.27664#S4.SS2.p1.1)\.
- \[28\]M\. Tan\(1993\)Multi\-agent reinforcement learning: independent versus cooperative agents\.InProceedings of the Tenth International Conference on Machine Learning,pp\. 330–337\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[29\]P\. D\. Taylor and L\. B\. Jonker\(1978\)Evolutionarily stable strategies and game dynamics\.Mathematical Biosciences40\(1–2\),pp\. 145–156\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.27664#S3.SS3.p1.1)\.
- \[30\]K\. Tuyls and G\. Weiss\(2012\)Multiagent learning: basics, challenges, and prospects\.AI Magazine33\(3\),pp\. 41–52\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1)\.
- \[31\]T\. Wang, H\. Dong, V\. Lesser, and C\. Zhang\(2020\)ROMA: multi\-agent reinforcement learning with emergent roles\.InProceedings of the 37th International Conference on Machine Learning,pp\. 9876–9886\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p2.1)\.
- \[32\]C\. J\. C\. H\. Watkins and P\. Dayan\(1992\)Q\-learning\.Machine Learning8\(3–4\),pp\. 279–292\.Cited by:[§4\.2](https://arxiv.org/html/2609.27664#S4.SS2.p1.1)\.
- \[33\]J\. W\. Weibull\(1995\)Evolutionary game theory\.MIT Press,Cambridge, MA\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.27664#S2.SS1.p1.1)\.
- \[34\]H\. P\. Young\(1993\)The evolution of conventions\.Econometrica61\(1\),pp\. 57–84\.Cited by:[§1\.2](https://arxiv.org/html/2609.27664#S1.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.27664#S2.SS2.p1.1),[§7\.4](https://arxiv.org/html/2609.27664#S7.SS4.p2.1)\.
- \[35\]K\. Zhang, Z\. Yang, and T\. Başar\(2021\)Multi\-agent reinforcement learning: a selective overview of theories and algorithms\.InHandbook of Reinforcement Learning and Control,pp\. 321–384\.Cited by:[§2\.3](https://arxiv.org/html/2609.27664#S2.SS3.p1.1),[§4\.1](https://arxiv.org/html/2609.27664#S4.SS1.p1.1)\.相似文章
网络拓扑和对手信息在塑造多智能体强化学习系统合作中的作用
本文探讨了网络拓扑和对手信息对多智能体强化学习系统中合作涌现的影响,特别是在重复囚徒困境(IPD)中,发现图结构和信息可用性对合作策略有显著影响。
多智能体交互中出现的工具使用
OpenAI 展示了在躲猫猫环境中训练的智能体能够通过多智能体竞争发现六种不同的突现策略和工具使用行为,而无需明确的对象交互激励。这项工作表明多智能体协同适应可以通过自监督学习产生复杂的智能行为。
超越静态评估:面向对抗博弈的LLM驱动策略演化中的共演化机制
本文提出了三种面向LLM驱动的对抗多智能体博弈代码演化的共演化机制(评估器共演化、分层深度评估和弱点压力),在MCTF 2026海上夺旗任务中取得了最先进的结果。
可扩展的约束多智能体强化学习:通过状态增强与一致性实现可分离动力学
本文提出了一种分布式方法,用于约束多智能体强化学习,该方法采用状态增强策略学习和对偶变量上的邻居间一致性,以在满足全局资源约束的同时实现智能体数量线性扩展。在智能电网需求响应上的实验表明,一致性协调对可行性至关重要:与集中式训练方法不同,它能够扩展到数千个智能体。
学习合作、竞争和沟通
OpenAI 展示了多智能体强化学习环境的研究,其中智能体学习合作、竞争和沟通。该论文介绍了 MADDPG(Multi-Agent DDPG),这是一种集中式评论家方法,能够让智能体比传统的分散式方法更有效地学习协作策略和沟通协议。