基于元多智能体强化学习的交互策略快速适应方法及其在自动驾驶中的应用

arXiv cs.AI 论文

摘要

本文提出了一种元多智能体强化学习(meta-MARL)框架,将多智能体强化学习(MARL)问题建模为马尔可夫博弈,并定义了一种新的求解概念——元纳什均衡(meta-NE),并证明其与基于梯度博弈算法的驻点等价。实验表明,在自动驾驶任务中,该方法相比预训练的 MARL 基线方法能够实现更快的适应。

arXiv:2610.00705v1 Announce Type: new Abstract: This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents' strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.
查看原文
查看缓存全文

缓存时间: 2026/10/02 09:47

# Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
Source: [https://arxiv.org/html/2610.00705](https://arxiv.org/html/2610.00705)
Huiwen YanAffiliation:H\. Yan and M\. Liu are with the Department of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061, USAhuiweny@vt\.edu, mushuang@vt\.eduKyriakos G\. VamvoudakisAffiliation:K\. G\. Vamvoudakis is with the Daniel Guggenheim School of Aerospace Engineering, Georgia Tech, Atlanta, GA 30332, USAkyriakos@gatech\.eduMushuang Liu††thanks:This work was supported in part by the DARPA Young Faculty Award with the grant number D24AP00321, by NSF under grant No\.˜SLES\-2415479, by NASA ULI under grant No\.˜80NSSC25M7104, and by ARO under grant No\.˜W911NF\-24\-1\-0174\.Affiliation:H\. Yan and M\. Liu are with the Department of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061, USAhuiweny@vt\.edu, mushuang@vt\.edu

###### Abstract

This paper develops a meta\-multi\-agent reinforcement learning \(meta\-MARL\) framework to enable fast adaptation of interactive policies in a multi\-agent system \(MAS\)\. Meta\-reinforcement learning \(meta\-RL\) enables agents to rapidly adapt to new tasks/environments using a bi\-level optimization mechanism\. However, existing meta\-RL generally focuses on single\-agent systems\. Extending these frameworks and algorithms to multi\-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents’ strategic interactions\. To address these challenges, we model multi\-agent reinforcement learning \(MARL\) problems as Markov games \(MGs\) and develop a meta\-MARL framework for rapid interactive policy adaptation across a distribution of MGs\. A new concept, called meta\-NE, is defined to describe the desired solution concept in a meta\-MARL problem\. Sufficient conditions for the equivalence between a meta\-NE and a stationary point of the gradient\-play\-based meta\-MARL algorithm are established\. Our evaluation on autonomous\-driving tasks demonstrates that the proposed meta\-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework\.

## IINTRODUCTION

Reinforcement learning \(RL\) has been widely applied to complex decision\-making problems in diverse fields such as robotics\[[1](https://arxiv.org/html/2610.00705#bib.bib19)\], autonomous driving\[[2](https://arxiv.org/html/2610.00705#bib.bib18),[3](https://arxiv.org/html/2610.00705#bib.bib15),[4](https://arxiv.org/html/2610.00705#bib.bib14)\], communication\[[5](https://arxiv.org/html/2610.00705#bib.bib23)\], and smart manufacturing\[[6](https://arxiv.org/html/2610.00705#bib.bib4)\]\. RL problems are commonly formulated as Markov decision processes \(MDPs\), in which an agent interacts with the environment to iteratively optimize its policy\[[7](https://arxiv.org/html/2610.00705#bib.bib16)\]\. However, when the environment changes, the previously learned policies may not be applicable, and retraining from scratch often requires substantial additional data and computation\. It is therefore desirable to develop RL agents that can rapidly adapt to new tasks or environments\.

To enable rapid adaptation to new tasks, meta\-RL has received increasing attention\[[8](https://arxiv.org/html/2610.00705#bib.bib24),[9](https://arxiv.org/html/2610.00705#bib.bib17),[10](https://arxiv.org/html/2610.00705#bib.bib3)\]\. Meta\-RL learns across a distribution of tasks so that an agent can adapt to a new task more quickly or with less additional data, whereas standard RL optimizes a policy for a single task\. Meta\-RL is inherently a bilevel optimization problem: the inner level obtains task\-specific policies or information, and the outer level maximizes the post\-adaptation return\. Thus, meta\-RL optimizes the learning procedure itself to enable rapid adaptation\. Existing meta\-RL methods include gradient\-based approaches such as MAML\[[8](https://arxiv.org/html/2610.00705#bib.bib24)\]and context\-based approaches such as PEARL\[[9](https://arxiv.org/html/2610.00705#bib.bib17)\]\. MAML learns a policy\-parameter initialization that can adapt to a new task through a small number of gradient updates, and it optimizes this initialization based on the performance of the adapted policies\. In contrast, PEARL represents each task using a latent variable inferred from collected data and conditions the policy on this variable\. Adaptation therefore updates the posterior distribution of the latent variable rather than the policy parameters\. Most existing meta\-RL frameworks, however, focus on single\-agent settings\. In a multi\-agent system \(MAS\), a task is defined not only by the environment but also by strategic interactions among agents, creating additional challenges for rapid policy adaptation\.

To characterize strategic interactions in RL problems, multi\-agent reinforcement learning \(MARL\) has recently been investigated\[[11](https://arxiv.org/html/2610.00705#bib.bib10),[12](https://arxiv.org/html/2610.00705#bib.bib21)\]\. In MARL, an agent’s optimal policy depends not only on the environment but also on the policies of the other agents, which leads naturally to a Markov game \(MG\) formulation\[[11](https://arxiv.org/html/2610.00705#bib.bib10)\]\. A standard solution concept in an MG is a Nash equilibrium \(NE\), at which no agent can improve its return by unilaterally changing its policy\. However, finding an NE can be challenging: a deterministic NE may not always exist in a general MG, and learning algorithms, such as gradient play, may not always converge\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\]\. These difficulties motivate the study of Markov potential games \(MPGs\)\[[12](https://arxiv.org/html/2610.00705#bib.bib21),[13](https://arxiv.org/html/2610.00705#bib.bib2)\]\. MPGs are a special class of MGs in which agents’ unilateral incentives are represented by a scalar potential function\. Potential games possess favorable equilibrium and learning properties, including the existence of deterministic NEs and the convergence of best\-response and better\-response dynamics\[[14](https://arxiv.org/html/2610.00705#bib.bib12),[15](https://arxiv.org/html/2610.00705#bib.bib9),[16](https://arxiv.org/html/2610.00705#bib.bib7)\]\. Under the conditions established in\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\], MPGs likewise admit the existence of deterministic NEs and convergence of gradient play to approximate NEs\.

To enable rapid adaptation of interactive policies in complex MASs, we develop a meta\-MARL framework\. As in meta\-RL, the inner level obtains task\-specific policies, while the outer level optimizes post\-adaptation performance\. In an MAS, however, a task must capture both the environment and the agents’ strategic interactions\. Therefore, we formulate each inner\-loop optimization problem as an MG problem, and formulate the outer loop such that the NE to every given MG can be quickly derived\. Along this direction, a few recent studies have considered possible integrations of meta\-RL and MARL\[[17](https://arxiv.org/html/2610.00705#bib.bib8),[18](https://arxiv.org/html/2610.00705#bib.bib13)\]\. However, despite recent advances, formal equilibrium concepts for meta\-MARL, their relationship to the converged points of a learning algorithm, and their effectiveness in autonomous systems such as autonomous driving, still remain open\. The contributions of this paper are fourfold:

1. 1\.We define the meta\-Nash equilibrium \(meta\-NE\) as a solution concept for meta\-MARL\. At a meta\-NE, no agent can improve its expected post\-adaptation return by unilaterally changing its own initialization policy\.
2. 2\.We establish sufficient conditions under which meta\-NEs are equivalent to first\-order stationary policies of the induced meta\-game\.
3. 3\.We develop a MAML\-style meta\-MARL method that exploits potential functions in MPGs to form a meta\-training objective\.
4. 4\.We evaluate the proposed method in autonomous highway forced\-merging scenarios and perform comprehensive comparisons with pretrained MARL baselines\.

The remainder of this paper is organized as follows\. Section[II](https://arxiv.org/html/2610.00705#S2)formulates MGs, MPGs, and MARL\. Section[III](https://arxiv.org/html/2610.00705#S3)defines meta\-NE, establishes sufficient conditions for its equivalence to stationary policies, and presents a MAML\-style meta\-MARL algorithm\. Section[IV](https://arxiv.org/html/2610.00705#S4)evaluates the framework in autonomous driving\. Section[V](https://arxiv.org/html/2610.00705#S5)concludes the paper\.

## IIProblem Formulation and Preliminaries

We formulate Markov games and multi\-agent reinforcement learning in Sections[II\-A](https://arxiv.org/html/2610.00705#S2.SS1)and[II\-B](https://arxiv.org/html/2610.00705#S2.SS2), respectively\.

### II\-AMarkov Games

We define an MG by the tupleℳ=\(𝒩,𝒮,𝒜,P,r,γ,ρ\)\\mathcal\{M\}=\(\\mathcal\{N\},\\mathcal\{S\},\\mathcal\{A\},P,r,\\gamma,\\rho\)\. The agent set is𝒩=\{1,…,N\}\\mathcal\{N\}=\\\{1,\\ldots,N\\\}; the global state space is𝒮=𝒮1×⋯×𝒮N\\mathcal\{S\}=\\mathcal\{S\}\_\{1\}\\times\\cdots\\times\\mathcal\{S\}\_\{N\}; and the joint action space is𝒜=𝒜1×⋯×𝒜N\\mathcal\{A\}=\\mathcal\{A\}\_\{1\}\\times\\cdots\\times\\mathcal\{A\}\_\{N\}, where𝒮i\\mathcal\{S\}\_\{i\}and𝒜i\\mathcal\{A\}\_\{i\}denote the state and action spaces of agentii\. The transition kernelP:𝒮×𝒜→Δ⁡\(𝒮\)P:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\Delta\(\\mathcal\{S\}\)gives the probabilityP⁡\(s′∣s,a\)P\(s^\{\\prime\}\\mid s,a\)of transitioning froms=\(s1,…,sN\)s=\(s\_\{1\},\\ldots,s\_\{N\}\)tos′s^\{\\prime\}under the joint actiona=\(a1,…,aN\)a=\(a\_\{1\},\\ldots,a\_\{N\}\)\. The tupler=\(r1,…,rN\)r=\(r\_\{1\},\\ldots,r\_\{N\}\)collects the reward functionsri:𝒮×𝒜→ℝr\_\{i\}:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathbb\{R\},γ∈\[0,1\)\\gamma\\in\[0,1\)is the discount factor, andρ∈Δ⁡\(𝒮\)\\rho\\in\\Delta\(\\mathcal\{S\}\)is the initial\-state distribution\.

Actions are selected by the joint policyπ:𝒮→Δ⁡\(𝒜\)\\pi:\\mathcal\{S\}\\rightarrow\\Delta\(\\mathcal\{A\}\)\. We assume thatπ\\pifactorizes asπ=π1×⋯×πN\\pi=\\pi\_\{1\}\\times\\cdots\\times\\pi\_\{N\}, whereπi:𝒮→Δ⁡\(𝒜i\)\\pi\_\{i\}:\\mathcal\{S\}\\rightarrow\\Delta\(\\mathcal\{A\}\_\{i\}\)is agentii’s policy\. Thus, each agent selects its own action from the observed global state\. At timett, givenst=\(s1,t,…,sN,t\)s\_\{t\}=\(s\_\{1,t\},\\ldots,s\_\{N,t\}\), the probability of the joint actionat=\(a1,t,…,aN,t\)a\_\{t\}=\(a\_\{1,t\},\\ldots,a\_\{N,t\}\)factorizes as

π⁡\(at\|st\)=∏i=1Nπi​\(ai,t\|st\)\.\\pi\(a\_\{t\}\|s\_\{t\}\)=\\prod\_\{i=1\}^\{N\}\\pi\_\{i\}\(a\_\{i,t\}\|s\_\{t\}\)\.\(1\)We parameterizeπ\\piandπi\\pi\_\{i\}byθ=\(θ1,…,θN\)∈𝒳\\theta=\(\\theta\_\{1\},\\ldots,\\theta\_\{N\}\)\\in\\mathcal\{X\}andθi∈𝒳i\\theta\_\{i\}\\in\\mathcal\{X\}\_\{i\}, respectively, and denote the resulting policies byπθ\\pi\_\{\\theta\}andπi,θi\\pi\_\{i,\\theta\_\{i\}\}, where𝒳=𝒳1×⋯×𝒳N\\mathcal\{X\}=\\mathcal\{X\}\_\{1\}\\times\\cdots\\times\\mathcal\{X\}\_\{N\}\. When the context is clear, we useθ\\thetaandπθ\\pi\_\{\\theta\}, and likewiseθi\\theta\_\{i\}andπi,θi\\pi\_\{i,\\theta\_\{i\}\}, interchangeably\.

For agentii, the value functionViθ:𝒮→ℝV\_\{i\}^\{\\theta\}:\\mathcal\{S\}\\rightarrow\\mathbb\{R\}is the expected discounted sum of future rewards from initial statess:

Viθ\(s\)≔𝔼\[∑t=0∞γtri\(st,at\)\|πθ,s0=s\]\.V\_\{i\}^\{\\theta\}\(s\)\\coloneqq\\mathbb\{E\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{i\}\(s\_\{t\},a\_\{t\}\)\\middle\|\\pi\_\{\\theta\},s\_\{0\}=s\\right\]\.\(2\)The expected return of agentii,Ji:𝒳→ℝJ\_\{i\}:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}, averages its value function overs0∼ρs\_\{0\}\\sim\\rho:

Ji​\(θ\)≔𝔼s0∼ρ​Viθ​\(s0\)\.J\_\{i\}\(\\theta\)\\coloneqq\\mathbb\{E\}\_\{s\_\{0\}\\sim\\rho\}V\_\{i\}^\{\\theta\}\(s\_\{0\}\)\.\(3\)
In an MG, a Nash equilibrium is generally considered a desirable solution concept\. Its definition is given below\.

###### Definition 1

\(Nash equilibrium\[[19](https://arxiv.org/html/2610.00705#bib.bib22)\]\) A policyθ⋆=\(θ1⋆,…,θN⋆\)\\theta^\{\\star\}=\(\\theta\_\{1\}^\{\\star\},\\ldots,\\theta\_\{N\}^\{\\star\}\)is called a Nash equilibrium if

Ji​\(θi⋆,θ−i⋆\)≥Ji​\(θ~i,θ−i⋆\),∀θ~i∈𝒳i,i∈𝒩\.J\_\{i\}\(\\theta\_\{i\}^\{\\star\},\\theta\_\{\-i\}^\{\\star\}\)\\geq J\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}^\{\\star\}\),\\quad\\forall\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},\\quad i\\in\\mathcal\{N\}\.\(4\)

### II\-BMulti\-Agent Reinforcement Learning

Given the conditions established in\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\], we can consider the following gradient play to solve an MG, under which each agent updates its policy according to

θi\(k\+1\)=proj𝒳i⁡\(θi\(k\)\+η​∇θiJi​\(θ\(k\)\)\),\\theta\_\{i\}^\{\(k\+1\)\}=\\operatorname\{proj\}\_\{\\mathcal\{X\}\_\{i\}\}\\\!\\left\(\\theta\_\{i\}^\{\(k\)\}\+\\eta\\nabla\_\{\\theta\_\{i\}\}J\_\{i\}\(\\theta^\{\(k\)\}\)\\right\),\(5\)wherekkindexes the learning iteration andη\>0\\eta\>0denotes the learning step size\.

Gradient play \([5](https://arxiv.org/html/2610.00705#S2.E5)\) may not always converge in a general\-sum MG\. We therefore consider a special subclass, MPGs\.

###### Definition 2

\(Markov potential game\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\]\)\. An MGℳ\\mathcal\{M\}is an MPG if there exists a potential functionϕ:𝒮×𝒜→ℝ\\phi:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathbb\{R\}such that, for everyi∈𝒩i\\in\\mathcal\{N\},\(θ~i,θ−i\),\(θi,θ−i\)∈𝒳\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\),\(\\theta\_\{i\},\\theta\_\{\-i\}\)\\in\\mathcal\{X\}, ands∈𝒮s\\in\\mathcal\{S\}, we have

𝔼\[∑t=0∞γtri\(st,at\)\|π\(θ~i,θ−i\),s0=s\]−𝔼\[∑t=0∞γtri\(st,at\)\|π\(θi,θ−i\),s0=s\]=𝔼\[∑t=0∞γtϕ\(st,at\)\|π\(θ~i,θ−i\),s0=s\]−𝔼\[∑t=0∞γtϕ\(st,at\)\|π\(θi,θ−i\),s0=s\]\.\\begin\{split\}&\\mathbb\{E\}\\Biggl\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{i\}\(s\_\{t\},a\_\{t\}\)\\,\\Biggm\|\\,\\pi\_\{\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\},\\,s\_\{0\}=s\\Biggr\]\\\\ &\\quad\-\\mathbb\{E\}\\Biggl\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}r\_\{i\}\(s\_\{t\},a\_\{t\}\)\\,\\Biggm\|\\,\\pi\_\{\(\\theta\_\{i\},\\theta\_\{\-i\}\)\},\\,s\_\{0\}=s\\Biggr\]\\\\ &=\\mathbb\{E\}\\Biggl\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\phi\(s\_\{t\},a\_\{t\}\)\\,\\Biggm\|\\,\\pi\_\{\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\},\\,s\_\{0\}=s\\Biggr\]\\\\ &\\quad\-\\mathbb\{E\}\\Biggl\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\phi\(s\_\{t\},a\_\{t\}\)\\,\\Biggm\|\\,\\pi\_\{\(\\theta\_\{i\},\\theta\_\{\-i\}\)\},\\,s\_\{0\}=s\\Biggr\]\.\\end\{split\}\(6\)

The total potential function of an MPG is

Φ\(θ\)≔𝔼s0∼ρ\[∑t=0∞γtϕ\(st,at\)\|πθ,s0\]\.\\Phi\(\\theta\)\\coloneqq\\mathbb\{E\}\_\{s\_\{0\}\\sim\\rho\}\\left\[\\sum\_\{t=0\}^\{\\infty\}\\gamma^\{t\}\\phi\(s\_\{t\},a\_\{t\}\)\\Biggm\|\\pi\_\{\\theta\},s\_\{0\}\\right\]\.\(7\)
###### Lemma 1

\(Convergence for gradient play\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\]\)\. Under the conditions established in\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\], gradient play \([5](https://arxiv.org/html/2610.00705#S2.E5)\) reaches anε\\varepsilon\-NE of an MPG for anyε\>0\\varepsilon\>0\.

Here, anε\\varepsilon\-NE is a policy profile at which no agent can improve its expected return by more thanε\\varepsilonthrough a unilateral deviation\.

Given the expected return in \([3](https://arxiv.org/html/2610.00705#S2.E3)\) and the total potential in \([7](https://arxiv.org/html/2610.00705#S2.E7)\), we can rewrite \([6](https://arxiv.org/html/2610.00705#S2.E6)\) as

Ji​\(θ~i,θ−i\)−Ji​\(θi,θ−i\)=Φ⁡\(θ~i,θ−i\)−Φ⁡\(θi,θ−i\),J\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\-J\_\{i\}\(\\theta\_\{i\},\\theta\_\{\-i\}\)=\\Phi\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\-\\Phi\(\\theta\_\{i\},\\theta\_\{\-i\}\),\(8\)Assuming differentiability, this identity gives

∇θiJi​\(θ\)=∇θiΦ​\(θ\)\.\\nabla\_\{\\theta\_\{i\}\}J\_\{i\}\(\\theta\)=\\nabla\_\{\\theta\_\{i\}\}\\Phi\(\\theta\)\.\(9\)Hence, \([5](https://arxiv.org/html/2610.00705#S2.E5)\) is equivalent to the following:

θi\(k\+1\)=proj𝒳i⁡\(θi\(k\)\+η​∇θiΦ​\(θ\(k\)\)\),\\theta\_\{i\}^\{\(k\+1\)\}=\\operatorname\{proj\}\_\{\\mathcal\{X\}\_\{i\}\}\\\!\\left\(\\theta\_\{i\}^\{\(k\)\}\+\\eta\\nabla\_\{\\theta\_\{i\}\}\\Phi\(\\theta^\{\(k\)\}\)\\right\),\(10\)or, jointly,

θ\(k\+1\)=proj𝒳⁡\(θ\(k\)\+η​∇θΦ​\(θ\(k\)\)\)\.\\theta^\{\(k\+1\)\}=\\operatorname\{proj\}\_\{\\mathcal\{X\}\}\\\!\\left\(\\theta^\{\(k\)\}\+\\eta\\nabla\_\{\\theta\}\\Phi\(\\theta^\{\(k\)\}\)\\right\)\.\(11\)

## IIIMeta\-Multi\-Agent Reinforcement Learning

We first consider meta\-MARL under general MGs in Section[III\-A](https://arxiv.org/html/2610.00705#S3.SS1), define the meta\-NE, and study its relationship with first\-order stationary policies\. We then focus on MPGs in Section[III\-B](https://arxiv.org/html/2610.00705#S3.SS2)and develop an MPG\-based MAML\-style meta\-MARL method\.

### III\-AMeta\-MARL with General MGs

The goal of meta\-MARL is to enable rapid policy adaptation to any given task in MASs, where a task is represented by an MG characterizing both the environment and the strategic interactions of agents\. Specifically, we consider a distribution of MGs, denoted byp⁡\(ℳ\)p\(\\mathcal\{M\}\), from which a taskℳm\\mathcal\{M\}\_\{m\}is sampled\. We usemmto denote taskℳm\\mathcal\{M\}\_\{m\}when no ambiguity arises\.

We first introduce the meta\-NE solution concept\. Define agentii’s expected post\-adaptation return overp⁡\(ℳ\)p\(\\mathcal\{M\}\),Ψi:𝒳→ℝ\\Psi\_\{i\}\\colon\\mathcal\{X\}\\rightarrow\\mathbb\{R\}, by

Ψi​\(θ\)≔𝔼ℳm∼p⁡\(ℳ\)​Ji,m​\(𝕌m​\(θ\)\),\\Psi\_\{i\}\(\\theta\)\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}J\_\{i,m\}\(\\mathbb\{U\}\_\{m\}\(\\theta\)\),\(12\)where𝕌m:𝒳→𝒳\\mathbb\{U\}\_\{m\}\\colon\\mathcal\{X\}\\rightarrow\\mathcal\{X\}is the joint adaptation rule for taskmm\. We consider decentralized adaptation𝕌m=𝕌1,m×⋯×𝕌N,m\\mathbb\{U\}\_\{m\}=\\mathbb\{U\}\_\{1,m\}\\times\\cdots\\times\\mathbb\{U\}\_\{N,m\}\. Further, we define𝕌i,m:𝒳→𝒳i\\mathbb\{U\}\_\{i,m\}\\colon\\mathcal\{X\}\\rightarrow\\mathcal\{X\}\_\{i\}as the adaptation rule, with agentii’s adaptation depending on the joint policyπ\\pi, which we refer to as a coupled adaptation\. This formulation induces the meta\-game𝒢=\(𝒩,\{𝒳i\}i=1N,\{Ψi\}i=1N\)\\mathcal\{G\}=\(\\mathcal\{N\},\\\{\\mathcal\{X\}\_\{i\}\\\}\_\{i=1\}^\{N\},\\\{\\Psi\_\{i\}\\\}\_\{i=1\}^\{N\}\)over policy initializations\. A meta\-NE is a Nash equilibrium of this induced game, at which no agent can improve its expected post\-adaptation return by changing its initialization while the other agents’ initialization policies remain fixed\.

###### Definition 3

\(Meta\-Nash equilibrium\)\. A meta\-policy initializationθ⋆=\(θ1⋆,…,θN⋆\)\\theta^\{\\star\}=\(\\theta\_\{1\}^\{\\star\},\\ldots,\\theta\_\{N\}^\{\\star\}\)is a meta\-Nash equilibrium with respect top⁡\(ℳ\)p\(\\mathcal\{M\}\)and𝕌m\\mathbb\{U\}\_\{m\}if

Ψi​\(θi⋆,θ−i⋆\)≥Ψi​\(θ~i,θ−i⋆\),∀θ~i∈𝒳i,i∈𝒩\.\\Psi\_\{i\}\(\\theta\_\{i\}^\{\\star\},\\theta\_\{\-i\}^\{\\star\}\)\\geq\\Psi\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}^\{\\star\}\),\\quad\\forall\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},\\quad i\\in\\mathcal\{N\}\.\(13\)

Next, we give sufficient conditions under which meta\-NEs are equivalent to first\-order stationary policies of the meta\-game𝒢\\mathcal\{G\}\.

###### Definition 4

\(First\-order stationary policy in meta\-MARL\)\. A policyθ⋆=\(θ1⋆,…,θN⋆\)\\theta^\{\\star\}=\(\\theta\_\{1\}^\{\\star\},\\ldots,\\theta\_\{N\}^\{\\star\}\)is first\-order stationary for the induced meta\-game if

\(θ~i−θi⋆\)⊤​∇θiΨi​\(θ⋆\)≤0,∀θ~i∈𝒳i,i∈𝒩\.\(\\tilde\{\\theta\}\_\{i\}\-\\theta\_\{i\}^\{\\star\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\leq 0,\\quad\\forall\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},\\quad i\\in\\mathcal\{N\}\.\(15\)

###### Assumption 1

\(Regularity\)\. For eachi∈𝒩i\\in\\mathcal\{N\},𝒳i\\mathcal\{X\}\_\{i\}is nonempty, compact, and convex\. Moreover,Ψi\\Psi\_\{i\}is continuously differentiable on an open set containing𝒳\\mathcal\{X\}\.

###### Assumption 2

\(Meta\-gradient domination\)\. For eachi∈𝒩i\\in\\mathcal\{N\}, there exists a finite constantCi\>0C\_\{i\}\>0such that, for everyθ∈𝒳\\theta\\in\\mathcal\{X\}andθ~i∈𝒳i\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},

Ψi​\(θ~i,θ−i\)−Ψi​\(θ\)≤Ci​maxθ¯i∈𝒳i​\(θ¯i−θi\)⊤​∇θiΨi​\(θ\)\.\\Psi\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\-\\Psi\_\{i\}\(\\theta\)\\leq C\_\{i\}\\max\_\{\\bar\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\}\}\(\\bar\{\\theta\}\_\{i\}\-\\theta\_\{i\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta\)\.\(16\)

Assumption[2](https://arxiv.org/html/2610.00705#Thmassump2)is analogous to the gradient\-domination property established in\[[12](https://arxiv.org/html/2610.00705#bib.bib21)\]\.

###### Theorem 1

Under Assumptions[1](https://arxiv.org/html/2610.00705#Thmassump1)and[2](https://arxiv.org/html/2610.00705#Thmassump2), a policy is first\-order stationary for the induced meta\-game if and only if it is a meta\-NE\.

###### Proof:

Suppose first thatθ⋆\\theta^\{\\star\}is a meta\-NE\. Fix anyi∈𝒩i\\in\\mathcal\{N\}andθ~i∈𝒳i\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\}\. Since𝒳i\\mathcal\{X\}\_\{i\}is convex,\(1−τ\)​θi⋆\+τ​θ~i∈𝒳i,∀τ∈\(0,1\]\(1\-\\tau\)\\theta\_\{i\}^\{\\star\}\+\\tau\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},\\forall\\tau\\in\(0,1\]\. By the meta\-NE condition,

Ψi​\(\(1−τ\)​θi⋆\+τ​θ~i,θ−i⋆\)−Ψi​\(θ⋆\)≤0\.\\Psi\_\{i\}\\bigl\(\(1\-\\tau\)\\theta\_\{i\}^\{\\star\}\+\\tau\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}^\{\\star\}\\bigr\)\-\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\leq 0\.\(17\)Dividing byτ\>0\\tau\>0and lettingτ↓0\\tau\\downarrow 0, differentiability ofΨi\\Psi\_\{i\}gives

\(θ~i−θi⋆\)⊤​∇θiΨi​\(θ⋆\)≤0,\(\\tilde\{\\theta\}\_\{i\}\-\\theta\_\{i\}^\{\\star\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\leq 0,\(18\)satisfying first\-order stationarity\.

Conversely, suppose thatθ⋆\\theta^\{\\star\}is first\-order stationary\. Then, for everyi∈𝒩i\\in\\mathcal\{N\},

\(θ¯i−θi⋆\)⊤​∇θiΨi​\(θ⋆\)≤0,∀θ¯i∈𝒳i\.\(\\bar\{\\theta\}\_\{i\}\-\\theta\_\{i\}^\{\\star\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\leq 0,\\qquad\\forall\\bar\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\}\.\(19\)Sinceθ¯i=θi⋆\\bar\{\\theta\}\_\{i\}=\\theta\_\{i\}^\{\\star\}is feasible, it follows thatmaxθ¯i∈𝒳i⁡\(θ¯i−θi⋆\)⊤​∇θiΨi​\(θ⋆\)=0\\max\_\{\\bar\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\}\}\(\\bar\{\\theta\}\_\{i\}\-\\theta\_\{i\}^\{\\star\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta^\{\\star\}\)=0\. Hence, by Assumption[2](https://arxiv.org/html/2610.00705#Thmassump2), for anyθ~i∈𝒳i\\tilde\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\},

Ψi​\(θ~i,θ−i⋆\)−Ψi​\(θ⋆\)≤Ci​maxθ¯i∈𝒳i​\(θ¯i−θi⋆\)⊤​∇θiΨi​\(θ⋆\)=0\.\\Psi\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}^\{\\star\}\)\-\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\leq C\_\{i\}\\max\_\{\\bar\{\\theta\}\_\{i\}\\in\\mathcal\{X\}\_\{i\}\}\(\\bar\{\\theta\}\_\{i\}\-\\theta\_\{i\}^\{\\star\}\)^\{\\top\}\\nabla\_\{\\theta\_\{i\}\}\\Psi\_\{i\}\(\\theta^\{\\star\}\)=0\.\(20\)Therefore,Ψi​\(θ⋆\)≥Ψi​\(θ~i,θ−i⋆\)\\Psi\_\{i\}\(\\theta^\{\\star\}\)\\geq\\Psi\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}^\{\\star\}\), which is a meta\-NE\. ∎

### III\-BMeta\-MARL with MPGs

We now focus on MPGs\. We first establish the potential structure of the induced meta\-game under independent adaptation and then formulate an MPG\-based, MAML\-style meta\-MARL framework\.

For a distribution of MPGs, define the expected post\-adaptation total potentialΓ:𝒳→ℝ\\Gamma:\\mathcal\{X\}\\rightarrow\\mathbb\{R\}by

Γ⁡\(θ\)≔𝔼ℳm∼p⁡\(ℳ\)​Φm​\(𝕌m​\(θ\)\)\.\\Gamma\(\\theta\)\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\Phi\_\{m\}\(\\mathbb\{U\}\_\{m\}\(\\theta\)\)\.\(21\)
###### Proposition 1

For every taskmm, if each agent’s adaptation rule is independent of the other agents’ policies, i\.e\.,

𝕌i,m​\(θi,θ−i\)=𝕌i,m​\(θi,θ−i′\),\\displaystyle\\mathbb\{U\}\_\{i,m\}\(\\theta\_\{i\},\\theta\_\{\-i\}\)=\\mathbb\{U\}\_\{i,m\}\(\\theta\_\{i\},\\theta\_\{\-i\}^\{\\prime\}\),\(22\)∀i∈𝒩,∀θi∈𝒳i,∀θ−i,θ−i′∈𝒳−i,\\displaystyle\\forall i\\in\\mathcal\{N\},\\;\\forall\\theta\_\{i\}\\in\\mathcal\{X\}\_\{i\},\\;\\forall\\theta\_\{\-i\},\\theta\_\{\-i\}^\{\\prime\}\\in\\mathcal\{X\}\_\{\-i\},and if the sampled tasks are MPGs with probability one underp⁡\(ℳ\)p\(\\mathcal\{M\}\), then the game𝒢\\mathcal\{G\}is an exact potential game with potential functionΓ\\Gamma\.

###### Proof:

Under the condition in \([22](https://arxiv.org/html/2610.00705#S3.E22)\), the following holds for every sampled MPG \(cf\. \([8](https://arxiv.org/html/2610.00705#S2.E8)\)\):

Ji,m​\(𝕌m​\(θ~i,θ−i\)\)−Ji,m​\(𝕌m​\(θi,θ−i\)\)\\displaystyle J\_\{i,m\}\(\\mathbb\{U\}\_\{m\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\)\-J\_\{i,m\}\(\\mathbb\{U\}\_\{m\}\(\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\)\(23\)=Φm​\(𝕌m​\(θ~i,θ−i\)\)−Φm​\(𝕌m​\(θi,θ−i\)\)\.\\displaystyle=\\Phi\_\{m\}\(\\mathbb\{U\}\_\{m\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\)\-\\Phi\_\{m\}\(\\mathbb\{U\}\_\{m\}\(\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\)\.Taking expectations over the task distribution yields

Ψi​\(θ~i,θ−i\)−Ψi​\(θi,θ−i\)=Γ⁡\(θ~i,θ−i\)−Γ⁡\(θi,θ−i\),\\Psi\_\{i\}\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\-\\Psi\_\{i\}\(\\theta\_\{i\},\\theta\_\{\-i\}\)=\\Gamma\(\\tilde\{\\theta\}\_\{i\},\\theta\_\{\-i\}\)\-\\Gamma\(\\theta\_\{i\},\\theta\_\{\-i\}\),\(24\)which completes the proof\. ∎

Together with Assumption[1](https://arxiv.org/html/2610.00705#Thmassump1), Proposition[1](https://arxiv.org/html/2610.00705#Thmprop1)implies thatΓ\\Gammaattains a global maximum on𝒳\\mathcal\{X\}, and that every global maximizer is a meta\-NE of𝒢\\mathcal\{G\}\.

In the remainder of this subsection, we study the gradient\-based \(i\.e\., MAML\-style\[[8](https://arxiv.org/html/2610.00705#bib.bib24)\]\) adaptation rule and parameterize the policy with a neural network \(NN\)\. Letθ\\thetadenote the parameters of an NN, and therefore, the projection operator can be omitted\. We consider one gradient step:

θi,m′=𝕌i,m​\(θ\)=θi\+η1​∇θiJi,m​\(θ\),\\theta\_\{i,m\}^\{\\prime\}=\\mathbb\{U\}\_\{i,m\}\(\\theta\)=\\theta\_\{i\}\+\\eta\_\{1\}\\nabla\_\{\\theta\_\{i\}\}J\_\{i,m\}\(\\theta\),\(25\)whereη1\\eta\_\{1\}is the inner step size\.

Because taskmmis an MPG, \([25](https://arxiv.org/html/2610.00705#S3.E25)\) is equivalent to

θi,m′=θi\+η1​∇θiΦm​\(θ\),\\theta\_\{i,m\}^\{\\prime\}=\\theta\_\{i\}\+\\eta\_\{1\}\\nabla\_\{\\theta\_\{i\}\}\\Phi\_\{m\}\(\\theta\),\(26\)and stacking the updates across agents yields

θm′=θ\+η1​∇θΦm​\(θ\),η1\>0,\\theta\_\{m\}^\{\\prime\}=\\theta\+\\eta\_\{1\}\\nabla\_\{\\theta\}\\Phi\_\{m\}\(\\theta\),\\quad\\eta\_\{1\}\>0,\(27\)whereΦm\\Phi\_\{m\}is the total potential in taskmm\. Equation \([27](https://arxiv.org/html/2610.00705#S3.E27)\) has the same form as the MAML inner update because an MPG admits a scalar total potential\.

The meta\-objective therefore becomes

maxθ⁡Γ⁡\(θ\)\\displaystyle\\max\_\{\\theta\}\\Gamma\(\\theta\)=maxθ⁡𝔼ℳm∼p⁡\(ℳ\)​Φm​\(θm′\)\\displaystyle=\\max\_\{\\theta\}\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\Phi\_\{m\}\(\\theta\_\{m\}^\{\\prime\}\)\(28\)=maxθ⁡𝔼ℳm∼p⁡\(ℳ\)​Φm​\(θ\+η1​∇θΦm​\(θ\)\)\.\\displaystyle=\\max\_\{\\theta\}\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\Phi\_\{m\}\\bigl\(\\theta\+\\eta\_\{1\}\\nabla\_\{\\theta\}\\Phi\_\{m\}\(\\theta\)\\bigr\)\.The meta\-optimization is performed overθ\\theta, whereas the objective is evaluated with the adapted parametersθm′\\theta\_\{m\}^\{\\prime\}\. Thus, the MAML\-style procedure seeks an initialization from which one or a small number of gradient steps produce high total potential on a new task\.

If eachΦm\\Phi\_\{m\}is twice differentiable and differentiation can be interchanged with expectation, the exact meta\-gradient is

∇θΓ​\(θ\)\\displaystyle\\nabla\_\{\\theta\}\\Gamma\(\\theta\)=𝔼ℳm∼p⁡\(ℳ\)\[\(𝐈\+η1∇θ2Φm\(θ\)\)⊤\\displaystyle=\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\Bigl\[\\bigl\(\\mathbf\{I\}\+\\eta\_\{1\}\\nabla\_\{\\theta\}^\{2\}\\Phi\_\{m\}\(\\theta\)\\bigr\)^\{\\top\}\(29\)×∇θm′Φm\(θm′\)\],\\displaystyle\}\{\\displaystyle\\times\\nabla\_\{\\theta\_\{m\}^\{\\prime\}\}\\Phi\_\{m\}\(\\theta\_\{m\}^\{\\prime\}\)\\Bigr\],where𝐈\\mathbf\{I\}is an identity matrix\. Exact\-gradient ascent would updateθ\\thetaaccording to

θ\(k\+1\)=θ\(k\)\+η2​∇θΓ​\(θ\(k\)\),η2\>0,\\theta^\{\(k\+1\)\}=\\theta^\{\(k\)\}\+\\eta\_\{2\}\\nabla\_\{\\theta\}\\Gamma\(\\theta^\{\(k\)\}\),\\quad\\eta\_\{2\}\>0,\(30\)whereη2\\eta\_\{2\}is the outer step size\. Following first\-order MAML\[[8](https://arxiv.org/html/2610.00705#bib.bib24)\], we omit the Hessian term and use the approximation

∇~θ​Γ​\(θ\)≔𝔼ℳm∼p⁡\(ℳ\)​\[∇θm′Φm​\(θm′\)\]\.\\widetilde\{\\nabla\}\_\{\\theta\}\\Gamma\(\\theta\)\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\left\[\\nabla\_\{\\theta\_\{m\}^\{\\prime\}\}\\Phi\_\{m\}\(\\theta\_\{m\}^\{\\prime\}\)\\right\]\.\(31\)Algorithm[1](https://arxiv.org/html/2610.00705#alg1)summarizes the above implementation\.

Algorithm 1MAML\-style MPG\-based Meta\-Multi\-Agent Reinforcement LearningRequire:p⁡\(ℳ\)p\(\\mathcal\{M\}\): distribution over tasks,η1,η2\\eta\_\{1\},\\eta\_\{2\}: inner and outer step sizes, andEE: number of meta\-training epochs Procedures:

1:Initialize policy parameter

θ\(0\)\\theta^\{\(0\)\}\.

2:for

k=0,1,⋯,E−1k=0,1,\\cdots,E\-1do

3:Sample a batch of tasksℳm∼p⁡\(ℳ\)\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\.

4:foreach sampled task

ℳm\\mathcal\{M\}\_\{m\}do

5:Samples0∼ρs\_\{0\}\\sim\\rhoand collect the induced trajectories usingπθ\(k\)\\pi\_\{\\theta^\{\(k\)\}\}inℳm\\mathcal\{M\}\_\{m\}\.

6:Compute∇θΦm​\(θ\(k\)\)\\nabla\_\{\\theta\}\\Phi\_\{m\}\(\\theta^\{\(k\)\}\)\.

7:Computeθm′=θ\(k\)\+η1​∇θΦm​\(θ\(k\)\)\\theta\_\{m\}^\{\\prime\}=\\theta^\{\(k\)\}\+\\eta\_\{1\}\\nabla\_\{\\theta\}\\Phi\_\{m\}\(\\theta^\{\(k\)\}\)\.

8:Samples0′∼ρs\_\{0\}^\{\\prime\}\\sim\\rhoand collect the induced trajectories usingπθm′\\pi\_\{\\theta\_\{m\}^\{\\prime\}\}inℳm\\mathcal\{M\}\_\{m\}\.

9:Compute∇θm′Φm​\(θm′\)\\nabla\_\{\\theta\_\{m\}^\{\\prime\}\}\\Phi\_\{m\}\(\\theta\_\{m\}^\{\\prime\}\)\.

10:endfor

11:Compute∇~θ​Γ​\(θ\(k\)\)\\widetilde\{\\nabla\}\_\{\\theta\}\\Gamma\(\\theta^\{\(k\)\}\)according to \([31](https://arxiv.org/html/2610.00705#S3.E31)\) and update:θ\(k\+1\)=θ\(k\)\+η2​∇~θ​Γ​\(θ\(k\)\)\\theta^\{\(k\+1\)\}=\\theta^\{\(k\)\}\+\\eta\_\{2\}\\widetilde\{\\nabla\}\_\{\\theta\}\\Gamma\(\\theta^\{\(k\)\}\)\.

12:endfor

## IVApplications to Autonomous Driving

In this section, we evaluate the MPG\-based meta\-MARL method in autonomous highway forced\-merging scenarios\.

### IV\-ASimulation Setup

\(i\) Environment and agent configuration

Consider the multi\-agent highway forced\-merging scenario shown in Fig\.[1](https://arxiv.org/html/2610.00705#S4.F1)\. The on\-ramp vehicle is the ego vehicle, and the other vehicles travel in the target lane\. We define a vehicle as leading \(following\) if its longitudinal position is ahead of \(behind\) that of the ego vehicle along the direction of travel\. Then, we select the four nearest leading and four nearest following vehicles relative to the ego vehicle as the game players\. Together with the ego vehicle, they form a nine\-agent game \(i\.e\.,N=9N=9\)\.

![Refer to caption](https://arxiv.org/html/2610.00705v1/HighwayMerging_ConflictPoint.png)Fig\. 1:Single\-lane highway forced merge scenario\.We use the kinematic bicycle model\[[20](https://arxiv.org/html/2610.00705#bib.bib11)\]to describe the dynamics of each vehicle\. At timett, vehicleiihas control input\[ui,t,δi,f,t\]⊤\[u\_\{i,t\},\\,\\delta\_\{i,f,t\}\]^\{\\top\}, whereui,tu\_\{i,t\}is the longitudinal acceleration andδi,f,t\\delta\_\{i,f,t\}is the front\-wheel steering angle\. The statesi,t=\[xi,t,yi,t,vi,t,ϕi,t\]⊤s\_\{i,t\}=\[x\_\{i,t\},\\,y\_\{i,t\},\\,v\_\{i,t\},\\,\\phi\_\{i,t\}\]^\{\\top\}contains the longitudinal and lateral center\-of\-mass \(CoM\) positions in anxx–yyinertial frame, the CoM speed, and the heading angle\. The kinematic bicycle model is\[[20](https://arxiv.org/html/2610.00705#bib.bib11)\]:

xi,t\+1=xi,t\+vi,t​cos⁡\(ϕi,t\+βi,t\)​Δ​t,yi,t\+1=yi,t\+vi,t​sin⁡\(ϕi,t\+βi,t\)​Δ​t,vi,t\+1=vi,t\+ui,t​Δ​t,ϕi,t\+1=ϕi,t\+vi,tlr​sin⁡\(βi,t\)​Δ​t,βi,t=arctan⁡\(lrlr\+lf​tan⁡\(δi,f,t\)\),\\begin\{split\}&x\_\{i,t\+1\}=x\_\{i,t\}\+v\_\{i,t\}\\cos\\bigl\(\\phi\_\{i,t\}\+\\beta\_\{i,t\}\\bigr\)\\,\\Delta t,\\\\ &y\_\{i,t\+1\}=y\_\{i,t\}\+v\_\{i,t\}\\sin\\bigl\(\\phi\_\{i,t\}\+\\beta\_\{i,t\}\\bigr\)\\,\\Delta t,\\\\ &v\_\{i,t\+1\}=v\_\{i,t\}\+u\_\{i,t\}\\,\\Delta t,\\\\ &\\phi\_\{i,t\+1\}=\\phi\_\{i,t\}\+\\frac\{v\_\{i,t\}\}\{l\_\{r\}\}\\sin\\bigl\(\\beta\_\{i,t\}\\bigr\)\\,\\Delta t,\\\\ &\\beta\_\{i,t\}=\\arctan\\Biggl\(\\frac\{l\_\{r\}\}\{l\_\{r\}\+l\_\{f\}\}\\tan\\bigl\(\\delta\_\{i,f,t\}\\bigr\)\\Biggr\),\\end\{split\}\(32\)whereβi,t\\beta\_\{i,t\}is the sideslip angle between the CoM velocity vector and the vehicle’s longitudinal axis, andlfl\_\{f\}andlrl\_\{r\}are the distances from the CoM to the front and rear axles, respectively\. Here, we setΔ​t=0\.5​s\\Delta t=0\.5~\\mathrm\{s\}\.

In the simulation, we let each vehicle follow a reference path defined according to the road structure \(i\.e\., vehicles have to be “on\-road”\), and thus steering is generated by the reference\-path tracking\. The action here is thus reduced to a scalar longitudinal acceleration, i\.e\.,ai,t=ui,t∈\[−g,g\]a\_\{i,t\}=u\_\{i,t\}\\in\[\-g,g\], whereg=9\.81​m/s2g=9\.81~\\mathrm\{m\}/\\mathrm\{s\}^\{2\}\. Vehicle speed is constrained to0​m/s≤vi,t≤30​m/s0~\\mathrm\{m/s\}\\leq v\_\{i,t\}\\leq 30~\\mathrm\{m/s\}\.

![Refer to caption](https://arxiv.org/html/2610.00705v1/HighwayMerging_MetaMARL.png)Fig\. 2:Policy variation under different tasks\.![Refer to caption](https://arxiv.org/html/2610.00705v1/Aggr2Cons2.png)\(a\)Vehicles 5–6 aggressive; 7–8 conservative\.
![Refer to caption](https://arxiv.org/html/2610.00705v1/Aggr3Cons1.png)\(b\)Vehicles 5–7 aggressive; 8 conservative\.
![Refer to caption](https://arxiv.org/html/2610.00705v1/Aggr4Cons0.png)\(c\)Vehicles 5–8 aggressive\.

Fig\. 3:Policy adaptation under different tasks\.\(ii\) Reward function design and task definition

We follow the sufficient conditions for MPG construction in\[[21](https://arxiv.org/html/2610.00705#bib.bib20)\]and design the reward functions accordingly\. The ego vehicle is considered in the target lane once it passes the merging conflict point \(see Fig\.[1](https://arxiv.org/html/2610.00705#S4.F1)\), whose longitudinal position,xc=180​mx\_\{c\}=180~\\mathrm\{m\}, is determined by the road geometry\.

Let𝒩i,s⊂𝒩\\mathcal\{N\}\_\{i,\\mathrm\{s\}\}\\subset\\mathcal\{N\}and𝒩i,c⊂𝒩\\mathcal\{N\}\_\{i,\\mathrm\{c\}\}\\subset\\mathcal\{N\}denote the sets of vehicles traveling in the same lane as vehicleiiand in different lanes from vehicleii, respectively\. In taskmm, vehicleiireceives

ri,m​\(st,at\)\\displaystyle r\_\{i,m\}\(s\_\{t\},a\_\{t\}\)=αi,1\(m\)​riself,1\+αi,2​riself,2\\displaystyle=\\alpha\_\{i,1\}^\{\(m\)\}r\_\{i\}^\{\\mathrm\{self\},1\}\+\\alpha\_\{i,2\}r\_\{i\}^\{\\mathrm\{self\},2\}\(33\)\+∑j∈𝒩i,swi​j,1ri​jjoint,1\+∑j∈𝒩i,cwi​j,2ri​jjoint,2,\\displaystyle\+\\sum\_\{j\\in\\mathcal\{N\}\_\{i,\\mathrm\{s\}\}\}w\_\{ij,1\}r\_\{ij\}^\{\\mathrm\{joint\},1\}\+\\sum\_\{j\\in\\mathcal\{N\}\_\{i,\\mathrm\{c\}\}\}w\_\{ij,2\}r\_\{ij\}^\{\\mathrm\{joint\},2\},where we omit the arguments of each reward component on the right\-hand side and consider nonnegative weights\. Each term on the right\-hand side of \([33](https://arxiv.org/html/2610.00705#S4.E33)\) is specified as follows\.

The first term rewards desired\-speed tracking:

riself,1​\(si,t,ai,t\)=−\(vi,t−vi,d\)2,r\_\{i\}^\{\\mathrm\{self\},1\}\(s\_\{i,t\},a\_\{i,t\}\)=\-\(v\_\{i,t\}\-v\_\{i,d\}\)^\{2\},\(34\)wherevi,d=15​m/sv\_\{i,d\}=15~\\mathrm\{m/s\}is the desired speed for each vehicle\.

The second term rewards small accelerations:

riself,2​\(si,t,ai,t\)=−\(ui,t\)2\.r\_\{i\}^\{\\mathrm\{self\},2\}\(s\_\{i,t\},a\_\{i,t\}\)=\-\(u\_\{i,t\}\)^\{2\}\.\(35\)
The third and fourth terms are for collision avoidance\. The third term rewards same\-lane collision avoidance:

ri​jjoint,1​\(si,t,sj,t\)\\displaystyle r\_\{ij\}^\{\\mathrm\{joint\},1\}\(s\_\{i,t\},s\_\{j,t\}\)=−1τi​j,t\+ϵτ,\\displaystyle=\-\\frac\{1\}\{\\tau\_\{ij,t\}\+\\epsilon\_\{\\tau\}\},\(36\)τi​j,t\\displaystyle\\tau\_\{ij,t\}=\|xi,t−xj,t\|max⁡\{\|vi,t−vj,t\|,vc\},\\displaystyle=\\frac\{\|x\_\{i,t\}\-x\_\{j,t\}\|\}\{\\max\\\{\|v\_\{i,t\}\-v\_\{j,t\}\|,v\_\{c\}\\\}\},whereτi​j,t\\tau\_\{ij,t\}is a timescale for same\-lane longitudinal interactions, andvc\>0v\_\{c\}\>0is a relative\-speed threshold andϵτ\>0\\epsilon\_\{\\tau\}\>0is a regularizer to prevent division by zero\.

The fourth term rewards larger arrival\-time gaps between vehicles in different lanes that have not yet passed the conflict point\. The reward is designed as:

ri​jjoint,2​\(si,t,sj,t\)=−𝟏\{di​\(t\)\>0,dj​\(t\)\>0\}Ti​\(t\)​Tj​\(t\)​\(Ti​\(t\)−Tj​\(t\)\)2\+ϵ3,\\begin\{split\}&r\_\{ij\}^\{\\mathrm\{joint\},2\}\(s\_\{i,t\},s\_\{j,t\}\)\\\\ &=\-\\frac\{\\mathbf\{1\}\_\{\\\{d\_\{i\}\(t\)\>0,\\;d\_\{j\}\(t\)\>0\\\}\}\}\{\\sqrt\{T\_\{i\}\(t\)T\_\{j\}\(t\)\}\(T\_\{i\}\(t\)\-T\_\{j\}\(t\)\)^\{2\}\+\\epsilon\_\{3\}\},\\end\{split\}\(37\)whereTi​\(t\)=di​\(t\)/\(vi,t\+ϵv\)T\_\{i\}\(t\)=d\_\{i\}\(t\)/\(v\_\{i,t\}\+\\epsilon\_\{v\}\)andTj​\(t\)=dj​\(t\)/\(vj,t\+ϵv\)T\_\{j\}\(t\)=d\_\{j\}\(t\)/\(v\_\{j,t\}\+\\epsilon\_\{v\}\)estimate the times required for vehiclesiiandjjto reach the conflict point, anddi​\(t\)d\_\{i\}\(t\)is vehicleii’s remaining distance along its reference path\. Here,ϵv\>0\\epsilon\_\{v\}\>0andϵ3\>0\\epsilon\_\{3\}\>0are to prevent the denominator from being zero\. The factorTi​\(t\)​Tj​\(t\)\\sqrt\{T\_\{i\}\(t\)T\_\{j\}\(t\)\}weights the arrival\-time gap by a “risk level”: for a fixedTi​\(t\)−Tj​\(t\)T\_\{i\}\(t\)\-T\_\{j\}\(t\), a smaller value ofTi​\(t\)​Tj​\(t\)\\sqrt\{T\_\{i\}\(t\)T\_\{j\}\(t\)\}indicates greater risk\.

Letℐs\\mathcal\{I\}\_\{\\mathrm\{s\}\}be the undirected set of same\-lane interaction pairs andℐc\\mathcal\{I\}\_\{\\mathrm\{c\}\}the undirected set of different\-lane interaction pairs with merging conflicts\. Hence, the potential for taskmmis

ϕm​\(st,at\)\\displaystyle\\phi\_\{m\}\(s\_\{t\},a\_\{t\}\)=∑i∈𝒩\[αi,1\(m\)​riself,1\+αi,2​riself,2\]\\displaystyle=\\sum\_\{i\\in\\mathcal\{N\}\}\\left\[\\alpha\_\{i,1\}^\{\(m\)\}r\_\{i\}^\{\\mathrm\{self\},1\}\+\\alpha\_\{i,2\}r\_\{i\}^\{\\mathrm\{self\},2\}\\right\]\(38\)\+∑\{i,j\}∈ℐswi​j,1ri​jjoint,1\+∑\{i,j\}∈ℐcwi​j,2ri​jjoint,2,\\displaystyle\+\\sum\_\{\\\{i,j\\\}\\in\\mathcal\{I\}\_\{\\mathrm\{s\}\}\}w\_\{ij,1\}r\_\{ij\}^\{\\mathrm\{joint\},1\}\+\\sum\_\{\\\{i,j\\\}\\in\\mathcal\{I\}\_\{\\mathrm\{c\}\}\}w\_\{ij,2\}r\_\{ij\}^\{\\mathrm\{joint\},2\},where we omit the arguments on the right\-hand side\.

Given the above setup, the componentrrinℳ=\(𝒩,𝒮,𝒜,P,r,γ,ρ\)\\mathcal\{M\}=\(\\mathcal\{N\},\\mathcal\{S\},\\mathcal\{A\},P,r,\\gamma,\\rho\)is task\-dependent\. Hence, we define the task distribution by varying the weighting parameters\{αi,1\(m\)\}i∈𝒩∖\{1\}\\\{\\alpha\_\{i,1\}^\{\(m\)\}\\\}\_\{i\\in\\mathcal\{N\}\\setminus\\\{1\\\}\}of the surrounding vehicles, while fixing the ego vehicle’s weightα1,1\(m\)\\alpha\_\{1,1\}^\{\(m\)\}\. In other words, we use the eight\-dimensional vector\[α2,1\(m\),…,α9,1\(m\)\]⊤\[\\alpha\_\{2,1\}^\{\(m\)\},\\ldots,\\alpha\_\{9,1\}^\{\(m\)\}\]^\{\\top\}, to specify a task\. Here we select the parameterαi,1\(m\)\\alpha\_\{i,1\}^\{\(m\)\}because it characterizes vehicleii’s driving aggressiveness: a largerαi,1\(m\)\\alpha\_\{i,1\}^\{\(m\)\}indicates a stronger desire to keep its speed \(or equivalently, less emphasis on maintaining a safe distance\)\.

\(iii\) Policy network and parameter sharing

Each vehicle is controlled by a deterministic policy represented by a NN\. Because the vehicles have identical dynamics, we adopt parameter sharing\[[22](https://arxiv.org/html/2610.00705#bib.bib6),[23](https://arxiv.org/html/2610.00705#bib.bib5)\]to improve training efficiency: all agents share a common policy network, while a one\-hot vehicle index enables distinct actions\. We note that such an architecture may restrict the parameter space and lead to potential suboptimality; however, the potential suboptimality is not observed in our simulations\. For vehicleii, the network receives a global observation containing its normalized position, velocity, and lane information; the relative positions, velocities, and lane information of all other vehicles; and a one\-hot encoding of its index\. As shown in Fig\.[4](https://arxiv.org/html/2610.00705#S4.F4), the policy network comprises an input layer, two hidden layers, and an output layer:

1. 1\.Input layer:a global observation of all vehicles\.
2. 2\.Hidden layers:two fully connected layers, each followed by a leaky rectified linear unit \(Leaky ReLU\) activation function\.
3. 3\.Output layer:a fully connected layer followed by a hyperbolic tangent \(tanh\) activation function\. The output is then scaled by9\.81​m/s29\.81~\\mathrm\{m/s^\{2\}\}\.

Fig\. 4:Policy network architecture\.![Refer to caption](https://arxiv.org/html/2610.00705v1/convergence.png)Fig\. 5:Convergence of policy during meta\-training\.\(iv\) Training setup and convergence results

During meta\-training, we setγ=0\.99\\gamma=0\.99and estimateΓ\\Gammawith a finite horizon ofH=60H=60decision steps, which corresponds toH​Δ​t=30​sH\\Delta t=30~\\mathrm\{s\}:

ΓH​\(θ\)\\displaystyle\\Gamma^\{H\}\(\\theta\)≔𝔼ℳm∼p⁡\(ℳ\)​ΦmH​\(θm′\)\\displaystyle\\coloneqq\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\Phi\_\{m\}^\{H\}\(\\theta\_\{m\}^\{\\prime\}\)\(39\)=𝔼ℳm∼p⁡\(ℳ\)𝔼s0∼ρ\[∑t=0H−1γtϕm\(st,at\)\|πθm′,s0\]\.\\displaystyle=\\mathbb\{E\}\_\{\\mathcal\{M\}\_\{m\}\\sim p\(\\mathcal\{M\}\)\}\\mathbb\{E\}\_\{s\_\{0\}\\sim\\rho\}\\left\[\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\\phi\_\{m\}\(s\_\{t\},a\_\{t\}\)\\Biggm\|\\pi\_\{\\theta\_\{m\}^\{\\prime\}\},s\_\{0\}\\right\]\.
The initial states are generated using stratified sampling to ensure broad coverage of the state space\. Vehicle longitudinal positions and velocities are sampled within predefined ranges\. For each pair of consecutive vehicles in the same lane, the initial longitudinal CoM distance is at least7​m7~\\mathrm\{m\}, and, when the rear vehicle is closing on the vehicle ahead, the time\-to\-collision \(TTC\), computed as the longitudinal CoM distance divided by the closing speed, is at least4​s4~\\mathrm\{s\}\.

\(a\)\(b\)
Fig\. 6:Illustration of the merging behavior in Task 1\. \(a\) Ego vehicle before merging\. \(b\) Ego vehicle after merging\.\(a\)\(b\)
Fig\. 7:Illustration of the merging behavior in Task 2\. \(a\) Ego vehicle before merging\. \(b\) Ego vehicle after merging\.\(a\)\(b\)
Fig\. 8:Illustration of the merging behavior in Task 3\. \(a\) Ego vehicle before merging\. \(b\) Ego vehicle after merging\.Meta\-training is performed for 1500 epochs\. At each meta\-training epoch, we sample 16 tasks and perform one inner\-loop stochastic gradient ascent \(SGA\) step with learning rateη1=10−3\\eta\_\{1\}=10^\{\-3\}using trajectories induced by the meta\-policy to obtain the adapted policies\. We then evaluate the adapted policies to estimate the first\-order meta\-gradient and update the meta\-policy in the outer\-loop using the adaptive moment estimation \(Adam\) optimizer\[[24](https://arxiv.org/html/2610.00705#bib.bib1)\]with learning rateη2=10−3\\eta\_\{2\}=10^\{\-3\}\. Fig\.[5](https://arxiv.org/html/2610.00705#S4.F5)plots the estimated meta\-objectiveΓH\\Gamma^\{H\}during meta\-training and shows empirical convergence over the reported training run\.

TABLE I:Statistical test results for the evaluated policies
### IV\-BTask\-Specific Evaluation

To evaluate the adaptation performance of the learned meta\-policy for the ego vehicle when interacting with surrounding vehicles exhibiting different driving styles, we consider three specific tasks \(see Fig\.[2](https://arxiv.org/html/2610.00705#S4.F2)\)\. In each task, the vehicle policies are adapted from: \(a\) themeta\-policyinitialization and \(b\) apretrained policy, which is an NE in a task different from the target test task\. After initialization, both policies adapt according to gradient ascent\. We compare the adapted policies with \(c\) anoracle policy, which is the NE in the target test task\. The “average return on test task” in Fig\.[3](https://arxiv.org/html/2610.00705#S4.F3)refers to the average estimated total potentialΦmH\\Phi\_\{m\}^\{H\}evaluated on the target test task\. The ego vehicle behaviors described below are from the adapted meta\-policy\.

Task 1:Vehicles 5 and 6 exhibit aggressive behavior with largeαi,1\\alpha\_\{i,1\}, while vehicles 7 and 8 exhibit conservative behavior with smallαi,1\\alpha\_\{i,1\}\. As the gap between vehicles 6 and 7 opens, the ego decelerates to let vehicles 5 and 6 pass, then merges into the gap\. Two key moments are shown in Fig\.[6](https://arxiv.org/html/2610.00705#S4.F6), and the corresponding adaptation curves in Fig\.[3](https://arxiv.org/html/2610.00705#S4.F3)\.

Task 2:Vehicles 5, 6, and 7 exhibit aggressive behavior, while vehicle 8 exhibits conservative behavior\. The ego vehicle first decelerates to let vehicles 5–7 pass and then merges into the gap between vehicles 7 and 8\. Two key moments are shown in Fig\.[7](https://arxiv.org/html/2610.00705#S4.F7)\. The corresponding adaptation curves are presented in Fig\.[3](https://arxiv.org/html/2610.00705#S4.F3)\.

Task 3:Vehicles 5–8 all exhibit aggressive behavior\. The ego vehicle decelerates to let them pass, then merges behind vehicle 8\. Two key moments are shown in Fig\.[8](https://arxiv.org/html/2610.00705#S4.F8), with the corresponding adaptation curves in Fig\.[3](https://arxiv.org/html/2610.00705#S4.F3)\.

Figs\.[6](https://arxiv.org/html/2610.00705#S4.F6)–[8](https://arxiv.org/html/2610.00705#S4.F8)show that the meta\-policy can enable the ego vehicle to quickly adapt to various surrounding vehicles’ driving styles and perform reasonably well in terms of “where” and “how” to merge\. Fig\.[3](https://arxiv.org/html/2610.00705#S4.F3), in addition, shows that \(i\) the meta\-policy adapts faster than the pretrained policy, and that \(ii\) the meta\-policy converges to the optimal \(i\.e\., the oracle\) within around1010gradient steps\.

### IV\-CStatistical Evaluation

To provide a more comprehensive evaluation and to assess the robustness of the meta\-policy, we conduct a statistical comparison\. Specifically, we evaluate 50 randomly selected tasks, and for each task, we consider 100 randomly selected scenarios \(i\.e\., initial states\)\. Therefore, in total,50005000random scenarios are tested\. The performance of the \(a\) meta\-policy, \(b\) pretrained policy, and \(c\) oracle policy is shown in Table[I](https://arxiv.org/html/2610.00705#S4.T1)\. In Table[I](https://arxiv.org/html/2610.00705#S4.T1), a collision occurs when the ego vehicle contacts a surrounding vehicle during a test scenario\. The remaining four metrics are computed for each scenario and averaged over all50005000scenarios\. We observe the following:

1. 1\.Safety:No collisions were observed for the adapted meta\-policy, whereas 6 collisions were observed for the pretrained policy\. In addition, the adapted meta\-policy leads to a larger minimum inter\-vehicle distance\. Both metrics suggest that the adapted meta\-policy performs better than the pretrained policy in terms of safety\.
2. 2\.Travel efficiency:The adapted meta\-policy leads to a higher ego\-vehicle speed than the pretrained policy, indicating that the better safety performance does not compromise the travel efficiency\.
3. 3\.Energy efficiency:The adapted meta\-policy exhibited lower absolute acceleration than the pretrained policy, indicating less control effort\.
4. 4\.Ride comfort:The adapted meta\-policy exhibited a higher absolute jerk than the pretrained policy, indicating that the better safety, travel efficiency, and energy efficiency may come at the price of ride comfort\.

## VCONCLUSIONS

Meta\-MARL enables rapid adaptation of interactive policy across tasks in MASs, including changes in both the environment and agents’ interaction patterns\. We defined a meta\-NE to represent the desired policy profile in which no agent can improve its expected post\-adaptation return through a unilateral change\. We established sufficient conditions under which meta\-NEs are equivalent to first\-order stationary policies of the induced meta\-game\. For MPG tasks with independent adaptation, we showed that the induced meta\-game has an exact potential; for coupled MAML\-style adaptation, we formulated a MPG\-based meta\-MARL framework using a first\-order MAML approximation\. Numerical studies in autonomous driving applications show that the developed meta\-policy exhibits fast adaptations to various driving styles of the surrounding vehicles\. Statistical studies suggest that the meta\-policy performs better than the pretrained policy in terms of safety, travel efficiency, and energy efficiency\.

## References

- \[1\]\(2023\)Champion\-level drone racing using deep reinforcement learning\.Nature620\(7976\),pp\. 982–987\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06419-4)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[2\]J\. Chen, S\. E\. Li, and M\. Tomizuka\(2022\)Interpretable end\-to\-end urban autonomous driving with latent deep reinforcement learning\.IEEE Transactions on Intelligent Transportation Systems23\(6\),pp\. 5068–5078\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2020.3046646)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[3\]L\. Wang, J\. Liu, H\. Shao, W\. Wang, R\. Chen, Y\. Liu, and S\. L\. Waslander\(2023\)Efficient reinforcement learning for autonomous driving with parameterized skills and priors\.InRobotics: Science and Systems XIX,External Links:[Document](https://dx.doi.org/10.15607/RSS.2023.XIX.102)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[4\]Q\. Li, X\. Jia, S\. Wang, and J\. Yan\(2025\)Think2Drive: efficient reinforcement learning by thinking with latent world model for autonomous driving \(in CARLA\-v2\)\.InComputer Vision – ECCV 2024,Cham,pp\. 142–158\.External Links:ISBN 978\-3\-031\-72995\-9Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[5\]M\. Liu, Y\. Wan, S\. Li, F\. L\. Lewis, and S\. Fu\(2019\)Learning and uncertainty\-exploited directional antenna control for robust long\-distance and broad\-band aerial communication\.IEEE Transactions on Vehicular Technology69\(1\),pp\. 593–606\.Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[6\]C\. Zhang, W\. Song, Z\. Cao, J\. Zhang, P\. S\. Tan, and X\. Chi\(2020\)Learning to dispatch for job shop scheduling via deep reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1621–1632\.Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[7\]M\. L\. Puterman\(1994\)Markov decision processes: discrete stochastic dynamic programming\.John Wiley & Sons,New York, NY, USA\.Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p1.1)\.
- \[8\]C\. Finn, P\. Abbeel, and S\. Levine\(2017\)Model\-agnostic meta\-learning for fast adaptation of deep networks\.External Links:1703\.03400,[Link](https://arxiv.org/abs/1703.03400)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p2.1),[§III\-B](https://arxiv.org/html/2610.00705#S3.SS2.p5.1.1),[§III\-B](https://arxiv.org/html/2610.00705#S3.SS2.p8.3.1)\.
- \[9\]K\. Rakelly, A\. Zhou, C\. Finn, S\. Levine, and D\. Quillen\(2019\)Efficient off\-policy meta\-reinforcement learning via probabilistic context variables\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 5331–5340\.Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p2.1)\.
- \[10\]L\. Zintgraf, K\. Shiarlis, M\. Igl, S\. Schulze, Y\. Gal, K\. Hofmann, and S\. Whiteson\(2020\)VariBAD: a very good method for bayes\-adaptive deep rl via meta\-learning\.InICLR 2020: Proceedings of the Eighth International Conference on Learning Representations,Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p2.1)\.
- \[11\]S\. V\. Albrecht, F\. Christianos, and L\. Schäfer\(2024\)Multi\-agent reinforcement learning: foundations and modern approaches\.MIT Press\.External Links:[Link](https://www.marl-book.com/)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1)\.
- \[12\]R\. Zhang, Z\. Ren, and N\. Li\(2024\)Gradient play in stochastic games: stationary points, convergence, and sample complexity\.IEEE Transactions on Automatic Control69\(10\),pp\. 6499–6514\.External Links:[Document](https://dx.doi.org/10.1109/TAC.2024.3387208)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1),[§II\-B](https://arxiv.org/html/2610.00705#S2.SS2.p1.1),[§III\-A](https://arxiv.org/html/2610.00705#S3.SS1.p4.1),[Definition 2](https://arxiv.org/html/2610.00705#Thmdefn2.p1.1.1),[Lemma 1](https://arxiv.org/html/2610.00705#Thmlemma1.p1.1.1)\.
- \[13\]S\. Leonardos, W\. Overman, I\. Panageas, and G\. Piliouras\(2022\)Global convergence of multi\-agent policy gradient in Markov potential games\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=gfwON7rAm4)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1)\.
- \[14\]D\. Monderer and L\. S\. Shapley\(1996\)Potential games\.Games and Economic Behavior14\(1\),pp\. 124–143\.External Links:ISSN 0899\-8256,[Document](https://dx.doi.org/https%3A//doi.org/10.1006/game.1996.0044)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1)\.
- \[15\]M\. Liu, I\. Kolmanovsky, H\. E\. Tseng, S\. Huang, D\. Filev, and A\. Girard\(2023\)Potential game\-based decision\-making for autonomous driving\.IEEE Transactions on Intelligent Transportation Systems24\(8\),pp\. 8014–8027\.External Links:[Document](https://dx.doi.org/10.1109/TITS.2023.3264665)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1)\.
- \[16\]M\. Liu, H\. E\. Tseng, D\. Filev, A\. Girard, and I\. Kolmanovsky\(2024\)Safe and human\-like autonomous driving: a predictor–corrector potential game approach\.IEEE Transactions on Control Systems Technology32\(3\),pp\. 834–848\.External Links:[Document](https://dx.doi.org/10.1109/TCST.2023.3332438)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p3.1)\.
- \[17\]D\. K\. Kim, M\. Liu, M\. D\. Riemer, C\. Sun, M\. Abdulhai, G\. Habibi, S\. Lopez\-Cot, G\. Tesauro, and J\. How\(2021\)A policy gradient algorithm for learning to learn in multiagent reinforcement learning\.InProceedings of the 38th International Conference on Machine Learning,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 5541–5550\.Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p4.1)\.
- \[18\]W\. Mao, H\. Qiu, C\. Wang, H\. Franke, Z\. Kalbarczyk, R\. Iyer, and T\. Basar\(2023\)Multi\-agent meta\-reinforcement learning: sharper convergence rates with task similarity\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 66556–66570\.External Links:[Document](https://dx.doi.org/10.52202/075280-2906)Cited by:[§I](https://arxiv.org/html/2610.00705#S1.p4.1)\.
- \[19\]D\. Fudenberg\(1991\)Game theory\.MIT press\.Cited by:[Definition 1](https://arxiv.org/html/2610.00705#Thmdefn1.p1.1.1)\.
- \[20\]R\. Rajamani\(2012\)Vehicle dynamics and control\.2nd edition,Springer,New York, NY, USA\.External Links:ISBN 978\-1\-4614\-1432\-2,[Document](https://dx.doi.org/10.1007/978-1-4614-1433-9)Cited by:[§IV\-A](https://arxiv.org/html/2610.00705#S4.SS1.p3.1)\.
- \[21\]H\. Yan and M\. Liu\(2026\)Markov potential game and multi\-agent reinforcement learning for autonomous driving\.IEEE Transactions on Vehicular Technology\(\),pp\. 1–14\.Cited by:[§IV\-A](https://arxiv.org/html/2610.00705#S4.SS1.p6.1)\.
- \[22\]J\. K\. Terry, N\. Grammel, S\. Son, B\. Black, and A\. Agrawal\(2023\)Revisiting parameter sharing in multi\-agent deep reinforcement learning\.Note:arXiv:2005\.13625Cited by:[§IV\-A](https://arxiv.org/html/2610.00705#S4.SS1.p15.1)\.
- \[23\]F\. Christianos, G\. Papoudakis, M\. A\. Rahman, and S\. V\. Albrecht\(2021\)Scaling multi\-agent reinforcement learning with selective parameter sharing\.InProceedings of the 38th International Conference on Machine Learning,Vol\.139,pp\. 1989–1998\.Cited by:[§IV\-A](https://arxiv.org/html/2610.00705#S4.SS1.p15.1)\.
- \[24\]D\. P\. Kingma and J\. Ba\(2015\)Adam: a method for stochastic optimization\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/1412.6980)Cited by:[§IV\-A](https://arxiv.org/html/2610.00705#S4.SS1.p19.1)\.

相似文章

自适应多时间视野强化学习

arXiv cs.LG

本文提出一种多时间视野强化学习方法,能够自适应地选择并组合时间视野,无需手动调整折扣因子即可鲁棒地适应变化的奖励结构,并在MiniGrid环境中进行了实验验证。

用于摔跤的元学习

OpenAI Blog

OpenAI 研究人员开发了元学习智能体,能够在多轮竞争性游戏中持续调整其策略,相比固定策略智能体展现出优异性能,并对环境和身体变化具有鲁棒性。