Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach

arXiv cs.LG Papers

Summary

This paper proposes VGG-MADiffRL, a value-gradient-guided multi-agent diffusion reinforcement learning algorithm, and MDCA, a hierarchical control architecture, for cooperative target tracking in multi-AUV ad-hoc networks under constrained acoustic communication and dynamic underwater disturbances.

arXiv:2608.12436v1 Announce Type: new Abstract: Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking. To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion?based hierarchical control architecture. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP. The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence. Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:29 AM

# Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach
Source: [https://arxiv.org/html/2608.12436](https://arxiv.org/html/2608.12436)
Chuan LinGuangjie HanShengchao ZhuQian ZhuYing LiuZhenyu WangThanks:*Corresponding author: Guangjie Han*Thanks:Jiaao Ma, Chuan Lin, Qian Zhu, Ying Liu and Zhenyu Wang are with the Software College, Northeastern University, Shenyang, China\. \(e\-mails: 2727746375@qq\.com; chuanlin1988@gmail\.com; zhuq@swc\.neu\.edu\.cn; liuy@swc\.neu\.edu\.cn;larrywang1019@outlook\.com\)\.Thanks:Guangjie Han is with the Key Laboratory of Maritime Intelligent Network Information Technology, Ministry of Education, Hohai University, Changzhou, China \(e\-mail: hanguangjie@gmail\.com\)\.Thanks:Shengchao Zhu is with the College of Computer Science and Software Engineering, Hohai University, Nanjing, 210013, China \(e\-mail: zhushengchao77@gmail\.com\)\.

###### Abstract

Multi\-AUV ad\-hoc network\-based target tracking requires networked autonomous underwater vehicles \(AUVs\) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances\. Although multi\-agent reinforcement learning \(MARL\) enables decentralized coordination through centralized training, existing methods suffer from high\-dimensional joint state\-action modeling, noise\-sensitive policy generation, leading to unstable training and degraded tracking\. To address these issues, we propose VGG\-MADiffRL, a value\-gradient\-guided multi\-agent diffusion RL algorithm, and MDCA, a diffusion\-based hierarchical control architecture\. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi\-AUV ad\-hoc networks as an MDP\. The proposed MDCA constitutes a three\-tier closed\-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer\. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback\. Within MDCA, the local online training layer is the policy learning framework; VGG\-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns\. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence\. Experimental results show that VGG\-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings\.

###### Index Terms:

Autonomous underwater vehicle, multi\-AUV ad\-hoc network system, underwater target tracking, multi\-agent reinforcement learning, diffusion model\.

## IIntroduction

The ocean contains abundant biological and mineral resources and is a critical domain for sustaining human society and advancing deep ocean strategic initiatives\[[2](https://arxiv.org/html/2608.12436#bib.bib1)\]\. As underwater communication technologies\[[15](https://arxiv.org/html/2608.12436#bib.bib2)\], particularly acoustic communication, have advanced, the research focus has shifted from single AUV systems to AUV swarm systems\[[16](https://arxiv.org/html/2608.12436#bib.bib3)\]\. In such systems, AUVs form autonomously organized, distributed, and intelligent Multi\-AUV ad\-hoc Networks \(MAANs\)\[[6](https://arxiv.org/html/2608.12436#bib.bib4)\]that enable collaborative execution of complex underwater missions, such as multiple target tracking and encirclement\[[8](https://arxiv.org/html/2608.12436#bib.bib6),[24](https://arxiv.org/html/2608.12436#bib.bib5)\], even in GPS\-denied or otherwise unavailable environments\.

However, underwater environments are dynamic and uncertain\[[7](https://arxiv.org/html/2608.12436#bib.bib7),[14](https://arxiv.org/html/2608.12436#bib.bib8)\]\. Conventional data\-transmission\-focused architectures treat the network as a passive conduit, making operational consistency difficult given acoustic channel challenges: rapidly shifting topologies and severely constrained bandwidth\[[10](https://arxiv.org/html/2608.12436#bib.bib9),[3](https://arxiv.org/html/2608.12436#bib.bib10)\]\. Overcoming this bottleneck requires a fundamental change in AUV network management: from passive data reception to active perception and reasoning about the environment\[[17](https://arxiv.org/html/2608.12436#bib.bib11)\]\. Although Centralized Training with Decentralized Execution \(CTDE\) has been widely used to address nonstationarity in multi\-agent settings\[[19](https://arxiv.org/html/2608.12436#bib.bib16),[9](https://arxiv.org/html/2608.12436#bib.bib12)\], its policy learning and optimization struggle with precise continuous cooperative tracking, especially when node relationships in the autonomous network change frequently\.

In Multi\-AUV ad\-hoc network cooperative tracking tasks, existing methods based on MARL still face three main challenges: \(1\) Insufficient modeling of complex continuous action distributions: traditional deterministic policies struggle to capture the strong interdependencies and couplings of continuous joint actions needed for collaborative decision\-making under dynamic topologies, often leading to limited expressiveness and suboptimal decisions\[[13](https://arxiv.org/html/2608.12436#bib.bib13)\]; \(2\) Training instability caused by relying on a single optimization objective: when policy updates depend solely on a single reward signal or local supervisory signal, they become highly sensitive to intermittent communication failures, which can cause oscillations in policy learning, slow convergence, and performance fluctuations\[[27](https://arxiv.org/html/2608.12436#bib.bib14),[26](https://arxiv.org/html/2608.12436#bib.bib26)\]; \(3\) Lack of value guidance during sampling: in the reverse denoising process of diffusion models, the absence of explicit value constraints allows sampled actions to drift away from regions of high expected return and introduces ineffective noise, reducing decision consistency, a problem that is especially harmful in low\-bandwidth underwater environments\[[4](https://arxiv.org/html/2608.12436#bib.bib15)\]\.

This work addresses training instability, restricted policy representation, and suboptimal action sampling in cooperative tracking with multi\-AUV ad\-hoc networks\. We build on the generative modeling paradigm of diffusion processes\. Using a joint optimization scheme that combines loss functions informed by value estimates with policy gradients, we propose the Value Gradient Guided Multi\-Agent Diffusion Reinforcement Learning \(VGG\-MADiffRL\) algorithm\. By replacing deterministic actors with diffusion policies and steering the reverse denoising trajectory via value gradients, this framework allows multi\-AUV networks to achieve robust policy optimization, faster convergence, and precise cooperative actions under shifting topologies\. Its core components are as follows\. \(1\) A diffusion policy architecture that improves the modeling of complex, interdependent cooperative actions\. \(2\) A dual\-objective framework that combines value signals with policy gradients to stabilize training when the topology changes\. \(3\) A sampling mechanism that uses value gradients to refine denoising trajectories, prioritizing actions with high expected returns\.

1. 1\.Framework for policy learning via diffusion models: VGG\-MADiffRL replaces the conventional deterministic actor with a diffusion model to construct a generative policy architecture closely integrated with value estimation\. This enables stable and robust policy optimization in multi\-AUV ad\-hoc networks, thereby enhancing the policy capability to model complex, continuous, and highly interdependent action distributions inherent in dynamic cooperative scenarios\.
2. 2\.Combined optimization mechanism for actor networks: VGG\-MADiffRL incorporates a dual objective optimization scheme that jointly minimizes a loss function informed by value estimates alongside a policy gradient loss\. The component guided by value signals leverages global value estimates to constrain the diffusion policy update direction, effectively mitigating policy oscillations caused by single objective optimization under volatile network topologies\.
3. 3\.Diffusion sampling mechanism guided by value gradients: VGG\-MADiffRL integrates value signals into the reverse denoising process to iteratively refine the denoising trajectory, steering action sampling toward regions of high expected return from the outset\. This mechanism reduces the influence of suboptimal or irrelevant actions and enhances training stability in underwater networks with limited bandwidth and autonomously organized topologies\.

The remainder of this paper is organized as follows\. Section[II](https://arxiv.org/html/2608.12436#S2)reviews the related work\. Section[III](https://arxiv.org/html/2608.12436#S3)displays the formulation of the problem and the system preliminaries\. Section[IV](https://arxiv.org/html/2608.12436#S4)proposes the MDCA framework\. Section[V](https://arxiv.org/html/2608.12436#S5)details the implementation of the multi\-AUV cooperative tracking algorithm based on VGG\-MADiffRL\. Section[VI](https://arxiv.org/html/2608.12436#S6)presents experimental evaluations and results\. Finally, Section[VII](https://arxiv.org/html/2608.12436#S7)concludes the paper and outlines future research directions\.

## IIRelated Works

In this section, we mainly review the latest research works related to the subject, specifically divided into the following two areas: 1\) AUV target tracking algorithms based on reinforcement learning/multi\-agent reinforcement learning; 2\) reinforcement learning methods based on diffusion models\.

### II\-AAUV Tracking Algorithm Based on RL

In\[[11](https://arxiv.org/html/2608.12436#bib.bib17)\], ACL\-SAC, a target tracking framework for AUVs that combines an attention\-based convolutional LSTM with Soft Actor\-Critic \(SAC\), was proposed\. It fuses multi\-sensor data through attention\-weighted temporal feature extraction and optimizes stochastic policies to balance entropy maximization and reward accumulation\. A composite reward function and prioritized experience replay improve training efficiency, enabling robust tracking under ocean current disturbances, sound speed uncertainty, and sparse observations\.

In\[[20](https://arxiv.org/html/2608.12436#bib.bib18)\], DDPG\-SAC was introduced for high\-precision path following of underactuated AUVs under ocean currents\. The method decouples surge and heading control, uses an enhanced Line\-of\-Sight \(LOS\) law to compensate for drift, and applies a low\-pass filter to suppress control chattering\. By redesigning state/action spaces and the reward function, it handles nonlinear dynamics, time\-varying hydrodynamics, and environmental disturbances, achieving strong generalization and stability in simulations\.

In\[[5](https://arxiv.org/html/2608.12436#bib.bib19)\], AMAML was developed for trajectory tracking under unknown time\-varying dynamics\. It decomposes the problem into fixed\-dynamics subtasks, models tracking as a Markov decision process using LOS guidance, and embeds an attention module to capture hidden dynamic features\. Integrating proximal policy optimization with maximum\-entropy meta\-learning enables fast adaptation across tasks, overcoming poor generalization and model dependency in conventional RL\.

Building on these single\-agent advances, research has shifted toward multi\-agent coordination, where diffusion models are increasingly integrated with multi\-agent reinforcement learning to address complex cooperative decision\-making in dynamic underwater environments\.

### II\-BDiffusion Models\-Based MARL

In\[[30](https://arxiv.org/html/2608.12436#bib.bib20)\], MADiff introduces a diffusion\-based offline multi\-agent reinforcement learning framework that models inter\-agent coordination via latent\-space attention and unifies decentralized execution with centralized training and teammate modeling\. It uses classifier\-free guidance to generate high\-return trajectories and recovers actions through an inverse dynamics model, effectively addressing extrapolation errors and limited expressivity in offline settings\.

In\[[18](https://arxiv.org/html/2608.12436#bib.bib21)\], HGCD extends this idea to heterogeneous teams by integrating a heterogeneous graph attention mechanism into the diffusion process, enabling adaptation to unseen team compositions\. Combined with offline meta\-RL, it achieves policy generalization across teams while supporting decentralized deployment and tackling challenges in data diversity, compositional generalization, and real\-world applicability\.

In\[[25](https://arxiv.org/html/2608.12436#bib.bib22)\], DMADRL applies diffusion models to online decision\-making in semantic vehicular edge computing\. It formulates denoising as an MDP to jointly optimize semantic task offloading and resource allocation\. Using a composite reward \(semantic fidelity, priority, energy\) and Gumbel\-Softmax reparameterization, it handles mixed discrete\-continuous action spaces and dynamic communication constraints, improving system utility and latency\.

Despite these advances, existing diffusion\-based MARL methods are either offline \(MADiff, HGCD\) or target discrete\-continuous hybrid domains like vehicular networks \(DMADRL\), making them unsuitable for online, purely continuous, fixed\-formation multi\-AUV tracking in dynamic underwater environments\. To bridge this gap, this paper proposes VGG\-MADiffRL and the MDCA architecture, which integrate value\-gradient\-guided diffusion sampling, dual\-objective policy optimization, and hierarchical coordination to enhance training stability, action quality, and robustness in complex cooperative underwater tracking tasks\.

## IIIPreliminary Materials

To accurately emulate the obstacle avoidance and target tracking behaviors of multi\-AUV ad\-hoc networks in dynamic underwater environments, we incorporate ocean current disturbances and inter\-node communication constraints into the modeling process, thus capturing realistic underwater network dynamics\. Based on this model, the cooperative operation problem of the multi\-AUV system is formulated as a Markov decision process \(MDP\)\.

### III\-AOcean Circulation Modeling

Given the uncertain and dynamic underwater environment, sonar is employed for precise relative positioning between AUVs and targets\. The AUV emits acoustic waves to scan its surroundings, and a sector coverage receiver array captures echoes from multiple directions, enabling target localization via intensity analysis\. The target detection process is modeled using the active sonar equation in Eq\. \([1](https://arxiv.org/html/2608.12436#S3.E1)\):

EM=SL−2​TL\+TS−\(NL−DI\)−DT,\\mathrm\{EM\}=\\mathrm\{SL\}\-2\\mathrm\{TL\}\+\\mathrm\{TS\}\-\(\\mathrm\{NL\}\-\\mathrm\{DI\}\)\-\\mathrm\{DT\},\(1\)whereEM\\mathrm\{EM\}is the excess margin,SL\\mathrm\{SL\}is the source level,TL\\mathrm\{TL\}is the transmission loss,TS\\mathrm\{TS\}is the target strength,NL\\mathrm\{NL\}is the ambient noise level,DI\\mathrm\{DI\}is the directivity index, andDT\\mathrm\{DT\}is the detection threshold \(in dB\)\.

The Navier–Stokes equation is a fundamental governing equation for fluid motion\. It accurately characterizes the dynamic behavior of fluids and can be used to analyze and calculate the hydrodynamic forces acting on underwater vehicles in marine environments\. The mathematical form of the equation is given in Eq\. \([2](https://arxiv.org/html/2608.12436#S3.E2)\):

ρ⁡\(∂u∂t\+u⋅∇u\)=−∇p\+μ​∇2u\+F,\\rho\\left\(\\frac\{\\partial\\mathrm\{u\}\}\{\\partial t\}\+\\mathrm\{u\}\\cdot\\nabla\\mathrm\{u\}\\right\)=\-\\nabla p\+\\mu\\nabla^\{2\}\\mathrm\{u\}\+\\mathrm\{F\},\(2\)whereρ\\rhodenotes the density of the fluid,u\\mathrm\{u\}is the velocity field,∂u∂t\\frac\{\\partial\\mathrm\{u\}\}\{\\partial t\}represents the temporal variation of the velocity,u⋅∇u\\mathrm\{u\}\\cdot\\nabla\\mathrm\{u\}is the convective term,∇p\\nabla pis the pressure gradient,μ\\mudenotes the dynamic viscosity,∇2u\\nabla^\{2\}\\mathrm\{u\}is the diffusion term, andF\\mathrm\{F\}is the external force term\.

### III\-BKalman Filter\-Based State Estimation

We define the AUV state by its three\-dimensional position and velocity, forming a state\-space model for Kalman filtering\. The state vector is given by Eq\. \([3](https://arxiv.org/html/2608.12436#S3.E3)\):

xk=\[xk,yk,zk,vx,k,vy,k,vz,k\]⊤,x\_\{k\}=\\left\[x\_\{k\},\\;y\_\{k\},\\;z\_\{k\},\\;v\_\{x,k\},\\;v\_\{y,k\},\\;v\_\{z,k\}\\right\]^\{\\top\},\(3\)wherexkx\_\{k\},yky\_\{k\},zkz\_\{k\}are the position components at stepkk, andvx,kv\_\{x,k\},vy,kv\_\{y,k\},vz,kv\_\{z,k\}the corresponding velocities\.

The discrete\-time dynamics with control inputs and process noise are described by the state transition model in Eq\. \([4](https://arxiv.org/html/2608.12436#S3.E4)\):

xk\+1=A​xk\+B​uk\+wk,x\_\{k\+1\}=Ax\_\{k\}\+Bu\_\{k\}\+w\_\{k\},\(4\)whereAAis the state transition matrix,BBthe control input matrix,uku\_\{k\}the control input, andwkw\_\{k\}the process noise\.

Measurements are related to the true state through the observation model with additive noise, as given in Eq\. \([5](https://arxiv.org/html/2608.12436#S3.E5)\):

zk=H​xk\+vk,z\_\{k\}=Hx\_\{k\}\+v\_\{k\},\(5\)wherezkz\_\{k\}is the measurement vector,HHthe observation matrix, andvkv\_\{k\}the observation noise\.

A Kalman filter is used to suppress underwater noise and improve the accuracy of the state estimate\. In the prediction step, the prior estimate and its covariance are computed from the previous posterior according to Eq\. \([6](https://arxiv.org/html/2608.12436#S3.E6)\):

\{xk\|k−1=A​xk−1\|k−1\+B​uk−1,Pk\|k−1=A​Pk−1\|k−1​A⊤\+Q,\\left\\\{\\begin\{aligned\} x\_\{k\|k\-1\}&=Ax\_\{k\-1\|k\-1\}\+Bu\_\{k\-1\},\\\\ P\_\{k\|k\-1\}&=AP\_\{k\-1\|k\-1\}A^\{\\top\}\+Q,\\end\{aligned\}\\right\.\(6\)wherexk\|k−1x\_\{k\|k\-1\}andPk\|k−1P\_\{k\|k\-1\}are the predicted state and covariance, andQQis the process noise covariance\.

In the measurement update, the Kalman gainKkK\_\{k\}is computed from the prior error covariance and the observation noise covariance, balancing the relative confidence in the model prediction and the measurement, as given in Eq\. \([7](https://arxiv.org/html/2608.12436#S3.E7)\):

Kk=Pk\|k−1​H⊤​\(H​Pk\|k−1​H⊤\+R\)−1,K\_\{k\}=P\_\{k\|k\-1\}H^\{\\top\}\\left\(HP\_\{k\|k\-1\}H^\{\\top\}\+R\\right\)^\{\-1\},\(7\)wherePk\|k−1P\_\{k\|k\-1\}is the prior estimation error covariance matrix,HHthe observation matrix, andRRthe observation noise covariance matrix\.

The prior state estimate is then updated by the observation innovation, which combines the model prediction with the real\-time measurement to form the posterior estimate\. This step is described by Eq\. \([8](https://arxiv.org/html/2608.12436#S3.E8)\):

xk\|k=xk\|k−1\+Kk​\(zk−H​xk\|k−1\),x\_\{k\|k\}=x\_\{k\|k\-1\}\+K\_\{k\}\\left\(z\_\{k\}\-Hx\_\{k\|k\-1\}\\right\),\(8\)wherexk\|k−1x\_\{k\|k\-1\}andxk\|kx\_\{k\|k\}are the prior and posterior state estimate vectors,zkz\_\{k\}is the observation vector at stepkk, and the termzk−H​xk\|k−1z\_\{k\}\-Hx\_\{k\|k\-1\}is the observation innovation \(residual\) that is mapped byKkK\_\{k\}into the state correction\.

Finally, the posterior error covariance matrix is updated to reflect the residual uncertainty after incorporating the measurement and to serve as the starting point for the next prediction cycle, as shown in Eq\. \([9](https://arxiv.org/html/2608.12436#S3.E9)\):

Pk\|k=\(I−Kk​H\)​Pk\|k−1,P\_\{k\|k\}=\\left\(I\-K\_\{k\}H\\right\)P\_\{k\|k\-1\},\(9\)wherePk\|kP\_\{k\|k\}andPk\|k−1P\_\{k\|k\-1\}are the posterior and prior estimation error covariance matrices,IIis the identity matrix, andKk​HK\_\{k\}Hrepresents the combined gain and observation mapping\.

### III\-CMarkov Process Modeling

In the underwater cooperative decision\-making process of a multi\-AUV system, the interaction with the environment is formalized as a Markov Decision Process \(MDP\)\. The MDP is defined by the tuple in Eq\. \([10](https://arxiv.org/html/2608.12436#S3.E10)\):

ℳ=\(𝒮,𝒜,𝒫,ℛ,γ\)\\mathcal\{M\}=\(\\mathcal\{S\},\\mathcal\{A\},\\mathcal\{P\},\\mathcal\{R\},\\gamma\)\(10\)where𝒮\\mathcal\{S\}is the state space,𝒜\\mathcal\{A\}the action space,𝒫\\mathcal\{P\}the state transition probability,ℛ\\mathcal\{R\}the reward function, andγ\\gammathe discount factor\.

The state space𝒮=\{s1,s2,…,sn\}\\mathcal\{S\}=\\\{s\_\{1\},s\_\{2\},\\ldots,s\_\{n\}\\\}contains the states of all AUVs\. The state of theii\-th AUV,sis\_\{i\}, combines its ego stateηi\\eta\_\{i\}, an environmental perceptionϕi\\phi\_\{i\}, and an observation vectoroio\_\{i\}\. The ego state includes positionpi∈ℝ3p\_\{i\}\\in\\mathbb\{R\}^\{3\}and velocityvi∈ℝ3v\_\{i\}\\in\\mathbb\{R\}^\{3\}\. The observation vectoroi=\(κi,σi,λi\)o\_\{i\}=\(\\kappa\_\{i\},\\sigma\_\{i\},\\lambda\_\{i\}\)captures the relative position information of targets, neighboring AUVs, and environmental landmarks, respectively\. Its dimension depends on the number of targetsNκN\_\{\\kappa\}, the number of neighborsNσN\_\{\\sigma\}, and the number of landmarksNλN\_\{\\lambda\}\.

The action space𝒜=\{a1,a2,…,an\}\\mathcal\{A\}=\\\{a\_\{1\},a\_\{2\},\\ldots,a\_\{n\}\\\}represents the continuous actions of the multi\-AUV system\. For AUVii, the action vectorai∈ℝ3a\_\{i\}\\in\\mathbb\{R\}^\{3\}provides control inputs along thexx,yy, andzzaxes, i\.e\.,ai=\[ax,i,ay,i,az,i\]⊤a\_\{i\}=\[a\_\{x,i\},a\_\{y,i\},a\_\{z,i\}\]^\{\\top\}\. Actions are generated by a diffusion policy\. To meet execution constraints, the policy outputs undergo range clipping and scaling in the environment execution layer before being converted into executable control commands\.

The state transition probability𝒫\\mathcal\{P\}captures the dynamic evolution of system states and is built on a physical kinematic model\. The discount factorγ∈\[0,1\)\\gamma\\in\[0,1\)balances the weight of future returns\. Together, these components define the long\-term cumulative return used in policy evaluation\.

The reward functionℛ\\mathcal\{R\}accounts for several factors: it explicitly incorporates tracking accuracy, collision avoidance, and environmental constraints to improve cooperative multi\-AUV tracking performance\. The precise formulation is given in Section[V](https://arxiv.org/html/2608.12436#S5)\.

## IVHierarchical Multi\-Agent Collaborative Control Architecture Based on Value Gradient\-Guided Diffusion Strategy MARL

In this section, we present the proposed multi\-agent diffusion\-based collaborative architecture \(MDCA\), a CTDE reinforcement learning framework specifically designed to address the highly dynamic topology and communication\-constrained nature of multi\-AUV ad\-hoc networks\. Building on this architecture, we introduce the value gradient\-guided multi\-agent diffusion reinforcement learning algorithm \(VGG\-MADiffRL\)\.

### IV\-AOverview of MDCA

This subsection presents a hierarchical collaborative control architecture for multi\-agent systems under the centralized training with decentralized execution framework\. By decomposing global coordination objectives into local operational domains, the proposed design strengthens cooperative consistency within each cluster while decoupling interactions between distinct regions\.

![Refer to caption](https://arxiv.org/html/2608.12436v1/Structure.png)Fig\. 1:Multi\-Agent Diffusion\-Based Collaborative ArchitectureAs illustrated in Fig\.[1](https://arxiv.org/html/2608.12436#S4.F1), the collaborative control architecture for multi\-AUV ad\-hoc networks comprises three layers: the global intelligent control layer, the local online training layer, and the physical action execution layer\. This design aligns well with ad\-hoc network characteristics: decentralization, self\-organization, dynamic topology, and distributed collaboration\. The functional principles and implementation details of these layers are described in the following subsections\.

Global Intelligent Control Layer:The global intelligent control layer, centered on the Unmanned Surface Vessel\-based cooperative gateway \(USV\-CG\), serves as the global coordination entry point of the multi\-AUV ad\-hoc network\. It receives global task commands from the satellite\. Through the northbound interface with the underwater controller, this layer performs information parsing, forwarding, and scheduling via a hierarchical block mechanism, enabling efficient interaction with the lower\-level local online training layer\. The global intelligent control layer also receives the ad\-hoc network topology, node states, and task progress reported by the local online training layer in real time, dynamically maintaining the global network topology and jointly performing global task decomposition and policy scheduling based on the mission scenario, link quality, and node resources\. This design ensures the multi\-AUV ad\-hoc network can still stably execute cooperative tracking tasks under topology fluctuations\.

Local Online Training Layer:The local online training layer formulates cooperative decisions for dynamically formed AUV clusters in designated operational domains\. It maintains a reliable data link with the global intelligent control layer via the northbound interface to acquire and interpret global mission directives, then validate commands and encapsulate local protocols\. Leveraging current node distribution, link connectivity, and available resource margins within its assigned domain, this layer routes mission instructions to designated AUV units through the southbound interface, enabling distributed cooperative execution\. Concurrently, it acquires navigation states, channel quality indicators, and execution deviations from the physical actuation layer to stabilize the local network topology and optimize policies online\. The layer transmits aggregated local states and mission outcomes to the global tier, forming a continuous feedback loop supplying precise operational data for strategic planning\.

Physical Action Execution Layer:As the terminal layer, the physical action execution tier comprises the AUV swarm for mission execution and perception\. Via the southbound interface, it receives trajectory planning and cooperative constraint directives from the local online training layer\. Each AUV generates control actions using an Actor diffusion policy: the forward process injects noise to characterize environmental and channel uncertainties, while the reverse process reconstructs feasible control sequences guided by critic value gradients\. These actions satisfy cooperative constraints and enable high\-precision target tracking\. Operational data \(navigation states and topological variations\) is uploaded to an experience replay buffer\. This feedback drives online policy updates in the local training layer, ensuring robust cooperative execution under dynamic conditions\.

In summary, the proposed hierarchical architecture operates under the centralized training with decentralized execution \(CTDE\) framework and is suitable for Multi\-AUV ad\-hoc networks\. Decomposing global coordination into localized domains, the design strengthens cooperative consistency within each cluster while decoupling interactions between distinct regions\. This structure enhances target tracking performance and system robustness under low bandwidth and highly dynamic underwater conditions\.

### IV\-BProposed Multi\-Agent Reinforcement Learning Algorithm

To address training instability and convergence difficulties in multi\-AUV cooperative continuous control tasks caused by non\-stationarity in multi\-agent environments, Q\-value estimation bias, and distributional mismatches in experience replay, this paper proposes a value gradient\-guided multi\-agent diffusion reinforcement learning algorithm \(VGG\-MADiffRL\)\.

![Refer to caption](https://arxiv.org/html/2608.12436v1/algorithm.png)Fig\. 2:Forward and Critic\-Guided Reverse Diffusion Processes of VGG\-MADiffRL#### IV\-B1Diffusion Architecture Guided by Value Gradients

Under the constraints of multi\-AUV ad\-hoc networks, the forward and reverse diffusion processes of the VGG\-MADiffRL algorithm are illustrated in Fig\.[2](https://arxiv.org/html/2608.12436#S4.F2)\. The framework explicitly accounts for key underwater characteristics of ad\-hoc networks—namely, limited bandwidth, highly dynamic topology, and high communication latency\. During the forward diffusion phase, progressive noise injection is applied to construct state\-action sequences that emulate the dual uncertainties arising from both the complex marine environment and volatile ad\-hoc network conditions\. In the reverse diffusion phase, a multilayer perceptron \(MLP\) network performs iterative denoising guided by value gradients, enabling accurate reconstruction of the original action sequence that satisfies cooperative tracking requirements—all under distributed execution with minimal communication overhead\.

Forward Diffusion Process:The forward diffusion process defines how noise is gradually added to the initial action, enabling the marginal distribution and sampling form of the noisy action at any diffusion step to be computed in closed form\. The noise schedule coefficient and cumulative signal retention are defined in Eq\. \([11](https://arxiv.org/html/2608.12436#S4.E11)\):

αt=1−βt,α¯t=∏i=1tαi,\\alpha\_\{t\}=1\-\\beta\_\{t\},\\qquad\\bar\{\\alpha\}\_\{t\}=\\prod\_\{i=1\}^\{t\}\\alpha\_\{i\},\(11\)whereβt\\beta\_\{t\}is the noise variance coefficient in steptt,αt\\alpha\_\{t\}denotes the signal retention ratio per step, andα¯t\\bar\{\\alpha\}\_\{t\}denotes the cumulative signal retention ratio from the initial timestep to steptt\.

Given the initial action sample, the corresponding marginal distribution is described in Eq\. \([12](https://arxiv.org/html/2608.12436#S4.E12)\):

q⁡\(xt∣x0\)=𝒩⁡\(xt,α¯t​x0,\(1−α¯t\)​I\),q\(\\mathrm\{x\}\_\{t\}\\mid\\mathrm\{x\}\_\{0\}\)=\\mathcal\{N\}\\\!\\bigl\(\\mathrm\{x\}\_\{t\};\\,\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathrm\{x\}\_\{0\},\\,\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathrm\{I\}\\bigr\),\(12\)where𝒩\\mathcal\{N\}denotes a Gaussian distribution, the mean termα¯t​x0\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathrm\{x\}\_\{0\}represents the scaled initial action, and the covariance term\(1−α¯t\)​I\(1\-\\bar\{\\alpha\}\_\{t\}\)\\mathrm\{I\}represents the accumulated noise covariance over time\.

The sampling form based on the reparameterization trick is given in Eq\. \([13](https://arxiv.org/html/2608.12436#S4.E13)\):

xt=α¯t​x0\+1−α¯t​ϵ,ϵ∼𝒩⁡\(𝟎,I\),\\mathrm\{x\}\_\{t\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\\,\\mathrm\{x\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\},\\qquad\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathrm\{I\}\),\(13\)wherext\\mathrm\{x\}\_\{t\}is the noisy action sample at steptt, andϵ\\boldsymbol\{\\epsilon\}is Gaussian noise drawn from the standard normal distribution\. This formulation enables direct sampling from the initial action to any diffusion step through reparameterization\.

Reverse Diffusion Process Guided by Value Gradients:The reverse sampling process progressively reconstructs the initial action from noisy actions through denoising, while further improving action quality with value\-function\-guided gradients\. The reconstruction estimate of the clean action is computed in Eq\. \([14](https://arxiv.org/html/2608.12436#S4.E14)\):

x^0=1α¯t​\(xt−1−α¯t​ϵθ​\(xt,t,st\)\),\\hat\{\\mathrm\{x\}\}\_\{0\}=\\frac\{1\}\{\\sqrt\{\\bar\{\\alpha\}\_\{t\}\}\}\\left\(\\mathrm\{x\}\_\{t\}\-\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\}\}\\,\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathrm\{x\}\_\{t\},t,\\mathrm\{s\}\_\{t\}\)\\right\),\(14\)wherex^0\\hat\{\\mathrm\{x\}\}\_\{0\}is the denoised action estimate reconstructed from the reverse process, andϵθ​\(xt,t,st\)\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathrm\{x\}\_\{t\},t,\\mathrm\{s\}\_\{t\}\)is the output of the noise prediction network at timesteptt, which is used to recover the initial action estimate from the noisy actionxt\\mathrm\{x\}\_\{t\}\.

The value\-function\-guided correction is given in Eq\. \([15](https://arxiv.org/html/2608.12436#S4.E15)\):

x^0=x^0\+λ​∇x^0Qϕ​\(s,x^0\),\\hat\{\\mathrm\{x\}\}\_\{0\}=\\hat\{\\mathrm\{x\}\}\_\{0\}\+\\lambda\\,\\nabla\_\{\\hat\{\\mathrm\{x\}\}\_\{0\}\}Q\_\{\\phi\}\(\\mathrm\{s\},\\hat\{\\mathrm\{x\}\}\_\{0\}\),\(15\)wherex^0\\hat\{\\mathrm\{x\}\}\_\{0\}is the action estimate after value gradient guidance,Qϕ​\(s,x^0\)Q\_\{\\phi\}\(\\mathrm\{s\},\\hat\{\\mathrm\{x\}\}\_\{0\}\)is the Critic value function,∇x^0Qϕ\\nabla\_\{\\hat\{\\mathrm\{x\}\}\_\{0\}\}Q\_\{\\phi\}denotes the gradient of the value function with respect to the action, pointing toward the direction that most rapidly increases the expected return, andλ\\lambdais the guidance strength coefficient that controls the influence of the value gradient\.

The value gradient correction in Eq\. \([15](https://arxiv.org/html/2608.12436#S4.E15)\) applies only during sampling\. Gradients through this correction term are truncated, preventing direct parameter updates for the diffusion policy\. Instead, this guidance is captured by the joint Actor loss in Eq\. \([23](https://arxiv.org/html/2608.12436#S4.E23)\), whereℒQ​\-​g​u​i​d​e\\mathcal\{L\}\_\{Q\\text\{\-\}guide\}steers the value\-guided action estimatex^0\\hat\{\\mathrm\{x\}\}\_\{0\}toward high\-value regions, whileℒp​g\\mathcal\{L\}\_\{pg\}enhances expected returns through differentiable sampled actions\. This design circumvents explicit higher\-order gradient computation from the value signal during Actor optimization, improving training stability\.

The reverse diffusion sampling update is completed by Eq\. \([16](https://arxiv.org/html/2608.12436#S4.E16)\):

xt−1=α¯t−1​x^0\+1−α¯t−1−σt2​ϵθ​\(xt,t,st\)\+σt​𝐳,\\mathrm\{x\}\_\{t\-1\}=\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\hat\{\\mathrm\{x\}\}\_\{0\}\+\\sqrt\{1\-\\bar\{\\alpha\}\_\{t\-1\}\-\\sigma\_\{t\}^\{2\}\}\\,\\boldsymbol\{\\epsilon\}\_\{\\theta\}\(\\mathrm\{x\}\_\{t\},t,\\mathrm\{s\}\_\{t\}\)\+\\sigma\_\{t\}\\mathbf\{z\},\(16\)wherext−1\\mathrm\{x\}\_\{t\-1\}is the action sample at stept−1t\-1in the reverse process, the first termα¯t−1​x^0\\sqrt\{\\bar\{\\alpha\}\_\{t\-1\}\}\\,\\hat\{\\mathrm\{x\}\}\_\{0\}is the guided action estimate, the second term is the predicted noise component, and the third termσt​𝐳\\sigma\_\{t\}\\mathbf\{z\}is the injected random noise, where𝐳\\mathbf\{z\}follows the standard normal distribution\. This equation realizes the denoising transition from timestepttto timestept−1t\-1\.

The noise standard deviation of the reverse process is determined by Eq\. \([17](https://arxiv.org/html/2608.12436#S4.E17)\):

σt=η​1−α¯t−11−α¯t​1−α¯tα¯t−1,\\sigma\_\{t\}=\\eta\\sqrt\{\\frac\{1\-\\bar\{\\alpha\}\_\{t\-1\}\}\{1\-\\bar\{\\alpha\}\_\{t\}\}\}\\sqrt\{1\-\\frac\{\\bar\{\\alpha\}\_\{t\}\}\{\\bar\{\\alpha\}\_\{t\-1\}\}\},\(17\)whereσt\\sigma\_\{t\}is the noise standard deviation used in reverse sampling at steptt, andη\\etais the randomness control coefficient \(DDPM parameter\)\. Whenη=0\\eta=0, the sampling is deterministic; whenη=1\\eta=1, the sampling is fully stochastic\. This parameter balances the diversity and stability of the generation process\.

Value Gradient Estimation:The value gradient is defined in Eq\. \([18](https://arxiv.org/html/2608.12436#S4.E18)\) to quantify the sensitivity of the Critic value function to action variations:

∇a0Qϕ​\(s,a0\)=\[∂Qϕ∂a\(1\)⋯∂Qϕ∂a\(d\)\]⊤\|a=a0,\\nabla\_\{\\mathrm\{a\}\_\{0\}\}Q\_\{\\phi\}\(\\mathrm\{s\},\\mathrm\{a\}\_\{0\}\)=\\left\.\\begin\{bmatrix\}\\frac\{\\partial Q\_\{\\phi\}\}\{\\partial a^\{\(1\)\}\}&\\cdots&\\frac\{\\partial Q\_\{\\phi\}\}\{\\partial a^\{\(d\)\}\}\\end\{bmatrix\}^\{\\top\}\\right\|\_\{\\mathrm\{a\}=\\mathrm\{a\}\_\{0\}\},\(18\)where∂Qϕ∂a\(i\)\\frac\{\\partial Q\_\{\\phi\}\}\{\\partial a^\{\(i\)\}\}is the partial derivative of the value function with respect to theii\-th action component, and\(⋅\)⊤\(\\cdot\)^\{\\top\}denotes vector transpose\. This gradient vector points toward the direction of the fastest increase in the value function\.

#### IV\-B2Joint Optimization Loss of VGG\-MADiffRL

In the design of the loss function, we explicitly account for both algorithmic stability and the generative characteristics of the diffusion model, decomposing the overall optimization objective into two components: the dual critic loss and the diffusion\-based actor loss\.

Dual Critic Loss:The critic loss is constructed based on the temporal difference \(TD\) error\. First, the target Q\-value for theii\-th mini\-batch sample is computed as given in Eq\. \([19](https://arxiv.org/html/2608.12436#S4.E19)\), and then the mean squared error is used to measure the deviation between the online critics and the target Q\-value, as shown in Eq\. \([20](https://arxiv.org/html/2608.12436#S4.E20)\):

yi=ri\+γ⁡\(1−di\)​min⁡\{Q1,ϕi′​\(s′,a′\),Q2,ϕi′​\(s′,a′\)\},y\_\{i\}=r\_\{i\}\+\\gamma\(1\-d\_\{i\}\)\\min\\left\\\{Q^\{\\prime\}\_\{1,\\phi\_\{i\}\}\(s^\{\\prime\},a^\{\\prime\}\),\\,Q^\{\\prime\}\_\{2,\\phi\_\{i\}\}\(s^\{\\prime\},a^\{\\prime\}\)\\right\\\},\(19\)ℒcritic\(i\)=‖Q1,ϕi​\(s,a\)−yi‖2\+‖Q2,ϕi​\(s,a\)−yi‖2,\\mathcal\{L\}\_\{\\text\{critic\}\}^\{\(i\)\}=\\left\\\|Q\_\{1,\\phi\_\{i\}\}\(s,a\)\-y\_\{i\}\\right\\\|^\{2\}\+\\left\\\|Q\_\{2,\\phi\_\{i\}\}\(s,a\)\-y\_\{i\}\\right\\\|^\{2\},\(20\)wheress,s′s^\{\\prime\}denote the current and next global states input to the critics;aais the current joint action of all agents;a′a^\{\\prime\}is the next joint action generated by the target actor networks;rir\_\{i\}anddid\_\{i\}are the immediate reward and termination flag of agentii;γ\\gammais the reward discount factor;Q1,ϕiQ\_\{1,\\phi\_\{i\}\},Q2,ϕiQ\_\{2,\\phi\_\{i\}\}are the dual online critic networks for agentii; andQ1,ϕi′Q^\{\\prime\}\_\{1,\\phi\_\{i\}\},Q2,ϕi′Q^\{\\prime\}\_\{2,\\phi\_\{i\}\}are the corresponding dual target critic networks\. Under the CTDE framework, the centralized critics receive the global statessand joint actionaa, while the diffusion\-based actors condition on each agent’s local observationoio\_\{i\}to generate actions\.

Diffusion Actor Loss:The Q\-guided loss term of the diffusion\-based actor guides action generation using the Q\-values output by the critics, with its mathematical form given in Eq\. \([21](https://arxiv.org/html/2608.12436#S4.E21)\):

ℒQ​\-guide=−λq⋅1N∑i=1Nmin\{Q1\(s,a0,i\),Q2\(s,a0,i\)\},\\mathcal\{L\}\_\{Q\\text\{\-guide\}\}=\-\\lambda\_\{q\}\\cdot\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\min\\left\\\{Q\_\{1\}\(s,a\_\{0,i\}\),\\,Q\_\{2\}\(s,a\_\{0,i\}\)\\right\\\},\(21\)whereλq\\lambda\_\{q\}is the Q\-guidance coefficient, anda0,ia\_\{0,i\}is the denoised action estimate used in the Q\-guidance term\.

The policy gradient loss directly maximizes the action value evaluated by the critics, with its mathematical expression given in Eq\. \([22](https://arxiv.org/html/2608.12436#S4.E22)\):

ℒpg=−1N∑i=1Nmin\{Q1\(s,ai\),Q2\(s,ai\)\},\\mathcal\{L\}\_\{\\text\{pg\}\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\min\\left\\\{Q\_\{1\}\(s,a\_\{i\}\),\\,Q\_\{2\}\(s,a\_\{i\}\)\\right\\\},\(22\)whereaia\_\{i\}is the differentiable sampled action used in the policy gradient term\.

The total actor loss combines both components:

ℒactor=ℒQ​\-guide\+ℒpg,\\mathcal\{L\}\_\{\\text\{actor\}\}=\\mathcal\{L\}\_\{Q\\text\{\-guide\}\}\+\\mathcal\{L\}\_\{\\text\{pg\}\},\(23\)whereℒactor\\mathcal\{L\}\_\{\\text\{actor\}\}is the final actor loss,ℒQ​\-guide\\mathcal\{L\}\_\{Q\\text\{\-guide\}\}is the Q\-guided loss defined in Eq\. \([21](https://arxiv.org/html/2608.12436#S4.E21)\), andℒpg\\mathcal\{L\}\_\{\\text\{pg\}\}is the policy gradient loss defined in Eq\. \([22](https://arxiv.org/html/2608.12436#S4.E22)\)\.

## VProposed Multi\-AUV Cooperative Target Tracking Scheme

This paper considers cooperative target tracking in multi\-AUV ad\-hoc networks\. Using the proposed VGG\-MADiffRL algorithm, we present the key tracking strategy for ad\-hoc network constraints and describe the algorithm’s complete execution pipeline\.

### V\-AProposed Reward Function for Multi\-AUV Tracking

We present a composite reward function for multi\-AUV ad\-hoc networks\. In the RL framework, cooperative tracking under ad\-hoc network constraints is a Markov decision process \(MDP\) maximizing cumulative reward, guiding AUVs toward stable coordination in bandwidth\-limited, dynamic, and communication\-impaired underwater environments\. We design three reward components for tracking fidelity, inter\-agent safety, and environmental constraints, ensuring efficiency and long\-term policy stability in complex ad\-hoc network scenarios\.

The total reward for theii\-th AUV at the current timestep is given by Eq\. \([24](https://arxiv.org/html/2608.12436#S5.E24)\):

Ri=α⋅rpos\+β⋅rcol\+δ⋅rland,R\_\{i\}=\\alpha\\cdot r\_\{\\text\{pos\}\}\+\\beta\\cdot r\_\{\\text\{col\}\}\+\\delta\\cdot r\_\{\\text\{land\}\},\(24\)whereRiR\_\{i\}denotes the instantaneous scalar reward received by theii\-th AUV;α\\alpha,β\\beta, andδ\\deltaare positive weighting coefficients that balance the influence of the target position reward \(rposr\_\{\\text\{pos\}\}\), the collision penalty \(rcolr\_\{\\text\{col\}\}\), and the landmark constraint term \(rlandr\_\{\\text\{land\}\}\) on policy learning\.

To emulate realistic underwater tracking environments, the target position reward is defined as given in Eq\. \([25](https://arxiv.org/html/2608.12436#S5.E25)\):

rpos=\{−dt,dt\>dt,min,w​dt−\(w\+1\)​dt,min,dt≤dt,min,r\_\{\\text\{pos\}\}=\\begin\{cases\}\-d\_\{t\},&d\_\{t\}\>d\_\{t,\\min\},\\\\ wd\_\{t\}\-\(w\+1\)d\_\{t,\\min\},&d\_\{t\}\\leq d\_\{t,\\min\},\\end\{cases\}\(25\)wheredtd\_\{t\}denotes the distance between the agent and the target,dt,mind\_\{t,\\min\}is the threshold distance defining the target’s proximity zone, andwwis a reward modulation parameter within this zone\.

To prevent collisions among nodes during dense cooperative maneuvers in the ad\-hoc network, a safety constraint penalty is designed based on relative inter\-agent distances\. The collision penalty is formulated as shown in Eq\. \([26](https://arxiv.org/html/2608.12436#S5.E26)\):

rcol=\{−λ1​\(do,min−do\)2,do<do,min,−min⁡\(λ2,λ3​do\),do≥do,min,r\_\{\\text\{col\}\}=\\begin\{cases\}\-\\lambda\_\{1\}\(d\_\{o,\\min\}\-d\_\{o\}\)^\{2\},&d\_\{o\}<d\_\{o,\\min\},\\\\ \-\\min\(\\lambda\_\{2\},\\lambda\_\{3\}d\_\{o\}\),&d\_\{o\}\\geq d\_\{o,\\min\},\\end\{cases\}\(26\)wheredod\_\{o\}represents the relative distance between two AUVs,do,mind\_\{o,\\min\}is the minimum allowable safe distance,λ1\\lambda\_\{1\}controls the penalty intensity for close\-range proximity,λ2\\lambda\_\{2\}sets the upper bound for long\-range penalties, andλ3\\lambda\_\{3\}is the linear slope coefficient governing the penalty decay\.

To smoothly model landmark region constraints within the reward function, a Sigmoid function is introduced, as expressed in Eq\. \([27](https://arxiv.org/html/2608.12436#S5.E27)\):

σ⁡\(x\)=11\+e−x,\\sigma\(x\)=\\frac\{1\}\{1\+e^\{\-x\}\},\(27\)whereσ⁡\(x\)\\sigma\(x\)serves as a smooth activation function that transforms hard\-threshold penalty relationships into a continuously differentiable form, thereby enhancing training stability\.

Building upon Eq\. \([27](https://arxiv.org/html/2608.12436#S5.E27)\), the obstacle avoidance penalty is defined as:

rland=−λl∑k=13σ\(τk−dl,ksk\),r\_\{\\text\{land\}\}=\-\\lambda\_\{l\}\\sum\_\{k=1\}^\{3\}\\sigma\\left\(\\frac\{\\tau\_\{k\}\-d\_\{l,k\}\}\{s\_\{k\}\}\\right\),\(28\)whereλl\\lambda\_\{l\}is the landmark penalty weight,dl,kd\_\{l,k\}denotes the distance between the agent and thekk\-th landmark,τk\\tau\_\{k\}represents the constraint threshold for the corresponding landmark, andsks\_\{k\}is a smoothing adjustment parameter\.

In summary, the proposed composite reward function comprises three core components: target tracking accuracy, swarm safety constraints, and underwater obstacle avoidance penalties\. This formulation comprehensively captures the primary objectives and operational constraints of multi\-AUV ad\-hoc networks in dynamic underwater scenarios\. By closely mirroring real\-world underwater task dynamics, this reward mechanism significantly improves simulation fidelity and effectively enhances the learning efficiency, convergence stability, and cooperative robustness of reinforcement learning models under bandwidth\-limited and topologically volatile network conditions\.

### V\-BProposed Tracking Algorithm Based on VGG\-MADiffRL

Algorithm 1Proposed Multi\-AUV Cooperative Target Tracking Algorithm Based on VGG\-MADiffRL1:Number of episodes

NepN\_\{\\text\{ep\}\}, episode length

LepL\_\{\\text\{ep\}\}, batch size

BB, update interval

IupdateI\_\{\\text\{update\}\}, minimal buffer size

SminS\_\{\\min\}, soft update coefficient

τ\\tau, guidance interval

IguideI\_\{\\text\{guide\}\}
2:Trained diffusion policies and twin critics for multi\-AUV cooperative tracking

3:Initialize environment

𝑒𝑛𝑣\\mathit\{env\}and VGG\-MADiffRL model

4:Initialize replay buffer

𝒟\\mathcal\{D\}
5:Set global step counter

ttotal←0t\_\{\\text\{total\}\}\\leftarrow 0
6:for

e=1e=1to

NepN\_\{\\text\{ep\}\}do

7:Reset environment and obtain initial state

𝐬\\mathbf\{s\}
8:for

t=1t=1to

LepL\_\{\\text\{ep\}\}do

9:Set critic\-guidance flag

gt←𝕀⁡\(tmodIguide=0\)g\_\{t\}\\leftarrow\\mathbb\{I\}\(t\\bmod I\_\{\\text\{guide\}\}=0\)
10:foreach agent

i=1i=1to

NNdo

11:if

gt=1g\_\{t\}=1then

12:Generate action

𝐚i\\mathbf\{a\}\_\{i\}by guided diffusion sampling \(Eqs\. \([14](https://arxiv.org/html/2608.12436#S4.E14)\)–\([18](https://arxiv.org/html/2608.12436#S4.E18)\)\)

13:else

14:Generate action

𝐚i\\mathbf\{a\}\_\{i\}by diffusion policy without guidance

15:Execute joint action

𝐚=\(𝐚1,…,𝐚N\)\\mathbf\{a\}=\(\\mathbf\{a\}\_\{1\},\\ldots,\\mathbf\{a\}\_\{N\}\)in

𝑒𝑛𝑣\\mathit\{env\}
16:Observe next state

𝐬′\\mathbf\{s\}^\{\\prime\}, reward

𝐫\\mathbf\{r\}, and done flag

𝐝\\mathbf\{d\}
17:Store transition

\(𝐬,𝐚,𝐫,𝐬′,𝐝,0\)\(\\mathbf\{s\},\\mathbf\{a\},\\mathbf\{r\},\\mathbf\{s\}^\{\\prime\},\\mathbf\{d\},0\)in

𝒟\\mathcal\{D\}
18:Update current state

𝐬←𝐬′\\mathbf\{s\}\\leftarrow\\mathbf\{s\}^\{\\prime\}
19:

ttotal←ttotal\+1t\_\{\\text\{total\}\}\\leftarrow t\_\{\\text\{total\}\}\+1
20:if

\|𝒟\|≥Smin\|\\mathcal\{D\}\|\\geq S\_\{\\min\}and

ttotalmodIupdate=0t\_\{\\text\{total\}\}\\bmod I\_\{\\text\{update\}\}=0then

21:Sample minibatch

ℬ\\mathcal\{B\}from

𝒟\\mathcal\{D\}
22:foreach agent

i=1i=1to

NNdo

23:Construct target joint action

𝐚′\\mathbf\{a\}^\{\\prime\}from target diffusion policies

24:Compute target value

yiy\_\{i\}using Eq\. \([19](https://arxiv.org/html/2608.12436#S4.E19)\)

25:Update twin critics by minimizing

ℒcritic\(i\)\\mathcal\{L\}\_\{\\text\{critic\}\}^\{\(i\)\}in Eq\. \([20](https://arxiv.org/html/2608.12436#S4.E20)\)

26:Update diffusion actor by minimizing

ℒactor\(i\)\\mathcal\{L\}\_\{\\text\{actor\}\}^\{\(i\)\}in Eq\. \([23](https://arxiv.org/html/2608.12436#S4.E23)\)

27:where

ℒQ​\-guide\(i\)\\mathcal\{L\}\_\{Q\\text\{\-guide\}\}^\{\(i\)\}and

ℒp​g\(i\)\\mathcal\{L\}\_\{pg\}^\{\(i\)\}are defined in Eqs\. \([21](https://arxiv.org/html/2608.12436#S4.E21)\)–\([22](https://arxiv.org/html/2608.12436#S4.E22)\)

28:Soft update Actor and Critic target networks for all agents by

θtarget←\(1−τ\)​θtarget\+τ​θ\\theta\_\{\\text\{target\}\}\\leftarrow\(1\-\\tau\)\\theta\_\{\\text\{target\}\}\+\\tau\\theta\.

29:if

𝐝=True\\mathbf\{d\}=\\text\{True\}then

30:break

In this section, the proposed VGG\-MADiffRL\-based tracking algorithm is formalized in Algorithm[1](https://arxiv.org/html/2608.12436#alg1)\. The complete workflow for ad\-hoc networks is decomposed into three steps, as illustrated in Fig\.[3](https://arxiv.org/html/2608.12436#S5.F3)\.

Step 1: Simulation Environment Construction\.The underwater simulation framework for multi\-AUV ad\-hoc networks integrates 3D environmental modeling, acoustic channel simulation, and sonar\-based state representation, capturing hydrodynamic and topological uncertainties within an MDP and providing a robust validation platform\.

Step 2: Diffusion Policy Architecture Design\.A diffusion\-based policy learning framework is deployed where the forward process injects progressive noise to emulate environmental and network fluctuations\. During the reverse process, Critic value gradients guide iterative denoising to reconstruct high\-quality cooperative action sequences, ensuring precise distributed decision\-making under bandwidth constraints\.

Step 3: Closed\-Loop Cooperative Execution\.In distributed execution, each AUV acts autonomously on local observations and generates actions via the trained diffusion policy\. Interaction trajectories are stored in the experience replay buffer\. Periodically, buffered samples are used to update the diffusion policy and dual Critic networks through soft target updates\. This closed\-loop paradigm enables the multi\-AUV system to sustain adaptive coordination and robust target tracking in dynamic underwater environments\.

The proposed VGG\-MADiffRL\-based tracking algorithm appears in Algorithm[1](https://arxiv.org/html/2608.12436#alg1), integrating guided diffusion sampling, twin\-critic learning, and soft target updates into a unified multi\-AUV loop\. Algorithm[1](https://arxiv.org/html/2608.12436#alg1)initializes the environment, replay buffer, and global timestep counterttotalt\_\{\\text\{total\}\}\(Lines 1–3\) and iterates over episodes \(Line 4\)\. Each episode resets environment \(Line 5\) and proceeds over timesteps \(Line 6\)\. At each timestep,tmodIguidet\\bmod I\_\{\\text\{guide\}\}determines guidance flaggtg\_\{t\}\(Line 7\), and each agent generates actions by either value\-gradient\-guided diffusion sampling \(Eqs\. \([14](https://arxiv.org/html/2608.12436#S4.E14)\)–\([18](https://arxiv.org/html/2608.12436#S4.E18)\), Lines 9–10\) or unguided diffusion policy sampling \(Lines 11–12\)\. The joint action is executed in the environment, and the next state, reward, and done flag are observed \(Lines 13–14\)\. The transition\(𝐬,𝐚,𝐫,𝐬′,𝐝,0\)\(\\mathbf\{s\},\\mathbf\{a\},\\mathbf\{r\},\\mathbf\{s\}^\{\\prime\},\\mathbf\{d\},0\)is stored in replay buffer𝒟\\mathcal\{D\}, and the current state and global timestep counter are updated \(Lines 15–17\)\. When\|𝒟\|≥Smin\|\\mathcal\{D\}\|\\geq S\_\{\\min\}andttotalmodIupdate=0t\_\{\\text\{total\}\}\\bmod I\_\{\\text\{update\}\}=0\(Line 18\), a minibatch is sampled from𝒟\\mathcal\{D\}\(Line 19\), and each agent is updated \(Line 20\): target joint actions constructed from target diffusion policies \(Line 21\), target values computed via Eq\. \([19](https://arxiv.org/html/2608.12436#S4.E19)\) \(Line 22\), twin critics optimized via Eq\. \([20](https://arxiv.org/html/2608.12436#S4.E20)\) \(Line 23\), and diffusion actors optimized via Eq\. \([23](https://arxiv.org/html/2608.12436#S4.E23)\), withℒQ​\-guide\\mathcal\{L\}\_\{Q\\text\{\-guide\}\}andℒp​g\\mathcal\{L\}\_\{pg\}defined in Eqs\. \([21](https://arxiv.org/html/2608.12436#S4.E21)\)–\([22](https://arxiv.org/html/2608.12436#S4.E22)\) \(Lines 24–25\)\. Each update cycle concludes with soft updates of Actor and Critic target networks viaθtarget←\(1−τ\)​θtarget\+τ​θ\\theta\_\{\\text\{target\}\}\\leftarrow\(1\-\\tau\)\\theta\_\{\\text\{target\}\}\+\\tau\\theta\(Line 26\)\. The episode terminates early if the done flag is true \(Lines 27–28\)\. This procedure enables stable, sample\-efficient cooperative tracking by coupling value\-guided diffusion action generation with twin\-critic\-based policy optimization\.

![Refer to caption](https://arxiv.org/html/2608.12436v1/Process.png)Fig\. 3:Workflow of the Generative Multi\-AUV MARL Algorithm

## VIEvaluations

This section presents a comprehensive experimental evaluation of the proposed algorithm\. The cooperative tracking performance of the multi\-AUV ad\-hoc network is analyzed across multiple metrics, including average reward, convergence stability, tracking accuracy, and robustness\. Comparative experiments with state\-of\-the\-art multi\-agent reinforcement learning algorithms are conducted to validate the effectiveness of VGG\-MADiffRL\.

### VI\-ASimulation Setup

All experiments are conducted on a computational platform equipped with an AMD Ryzen 9 8940HX processor, RTX 5060 GPU, and 16 GB RAM\. All code is implemented in Python 3\.10\.

The simulation environment is built on OceanGym\[[23](https://arxiv.org/html/2608.12436#bib.bib30)\], a benchmark environment for underwater embodied agents\. We adopt its modular agent\-environment interface for standardized multi\-agent underwater interaction and extend it with custom hydrodynamic effects and acoustic communication constraints to model the physical dynamics and network conditions of multi\-AUV ad\-hoc networks\.

In the evaluations, the target moves at a predefined constant velocity, while AUVs are initially distributed in a circular ring\-shaped region approximately 3\.5 to 5 km from the target\.

To comprehensively evaluate algorithm performance under varying network scales, four distinct tracking scenarios are employed in the assessment: 10 AUVs tracking 3 targets, 8 AUVs tracking 3 targets, 6 AUVs tracking 2 targets, and 4 AUVs tracking 2 targets\.

All the parameters in the evaluations are detailed in Table[I](https://arxiv.org/html/2608.12436#S6.T1)\.

### VI\-BResults and Discussion

We compare VGG\-MADiffRL against two groups of MARL methods\. The first group consists of five MARL algorithms for continuous control: MASAC\[[22](https://arxiv.org/html/2608.12436#bib.bib25)\], MAPPO\[[26](https://arxiv.org/html/2608.12436#bib.bib26)\], MAAC\[[28](https://arxiv.org/html/2608.12436#bib.bib27)\], MATD3\[[1](https://arxiv.org/html/2608.12436#bib.bib28)\], and MADDPG\[[12](https://arxiv.org/html/2608.12436#bib.bib29)\]\. The second group consists of two MARL methods designed for underwater AUV scenarios, DSBM\[[21](https://arxiv.org/html/2608.12436#bib.bib23)\]and MA\-A3C\[[29](https://arxiv.org/html/2608.12436#bib.bib24)\]\. Our approach is evaluated mainly from the following aspects: 1\) convergence speed; 2\) tracking accuracy; 3\) mean tracking error; 4\) error standard deviation; 5\) diffusion steps required for strategy generation in our framework; 6\) ablation studies validating each component’s contribution to system performance, and 7\) system availability under dynamic ad\-hoc network conditions\.

TABLE I:Simulation Parameters![Refer to caption](https://arxiv.org/html/2608.12436v1/4track2.png)\(a\)Scenario of 4 AUVs Tracking 2 Targets
![Refer to caption](https://arxiv.org/html/2608.12436v1/6track2.png)\(b\)Scenario of 6 AUVs Tracking 2 Targets
![Refer to caption](https://arxiv.org/html/2608.12436v1/8track3.png)\(c\)Scenario of 8 AUVs Tracking 3 Targets
![Refer to caption](https://arxiv.org/html/2608.12436v1/10track3.png)\(d\)Scenario of 10 AUVs Tracking 3 Targets

Fig\. 4:Convergence Speed EvaluationTABLE II:Tracking Accuracy ComparisonTABLE III:Mean Tracking Error ComparisonTABLE IV:Tracking Error Standard Deviation Comparison#### VI\-B1System Convergence Speed

To evaluate the training convergence of VGG\-MADiffRL, we compare its convergence performance against the two categories of methods described above across four multi\-AUV ad hoc network scenarios\. The convergence curves are shown in Figs\.[4](https://arxiv.org/html/2608.12436#S6.F4)\(a\)–[4](https://arxiv.org/html/2608.12436#S6.F4)\(d\)\.

Among the general\-purpose MARL baselines, MAPPO achieves the strongest convergence, benefiting from its centralized training with Generalized Advantage Estimation \(GAE\) and a stochastic Actor policy that balances global coordination with policy diversity\. MASAC, MAAC, MATD3, and MADDPG converge more slowly, particularly in larger\-scale scenarios \(8 AUVs Tracking 3 Targets and 10 AUVs Tracking 3 Targets\), where increased agent interactions amplify multi\-agent non\-stationarity\. Their deterministic or entropy\-regularized policies cannot adequately model the interdependent action distributions needed for coordinated tracking under dynamic ad hoc topologies\.

Among the underwater\-specific methods, DSBM and MA\-A3C train stably but converge to lower returns than VGG\-MADiffRL\. DSBM uses dynamic\-switching attention for multi\-target tracking; it performs moderately but plateaus early because its discrete switching mechanism restricts the expressiveness of continuous cooperative actions\. MA\-A3C adopts a hierarchical software\-defined architecture with advantage\-attention actor\-critic and advantage resampling; it converges steadily but reaches a lower final return because its deterministic policy gradient limits action expressiveness and its reward\-weighted attention compression discards fine\-grained coordination signals needed under fast\-changing ad hoc topologies\.

VGG\-MADiffRL converges faster than all compared methods across all scenarios\. During early training, the critic value\-guided mechanism drives rapid policy improvement: the diffusion policy generates actions via differentiable sampling, and a joint loss combining policy gradient objectives with Q\-guidance terms from global dual\-Q network outputs steers updates toward high\-value regions\. Batch updates that start after the replay buffer reaches a minimum size suppress small\-sample bias and improve sample reuse\. In later training, the algorithm remains smooth and stable\. The dual\-Q target networks take the minimum of two independent estimates to reduce overestimation, while soft updates avoid abrupt parameter shifts\. Gradient clipping limits update magnitudes, and the inherent action continuity of diffusion\-based generation helps avoid the oscillations seen in several baseline methods\.

#### VI\-B2Tracking Accuracy

In multi\-AUV ad\-hoc network cooperative tracking tasks, tracking accuracy serves as a core metric for evaluating training effectiveness and policy optimization\. To validate the effectiveness of the proposed algorithm, experiments are configured with AUVs initially deployed at a distance of 4\.5 km from the target, and a tracking error threshold of 0\.8 km is adopted for performance assessment\. As summarized in Table[II](https://arxiv.org/html/2608.12436#S6.T2), the comparative results demonstrate that VGG\-MADiffRL achieves the highest tracking accuracy under this scenario, significantly outperforming existing baseline methods\. These results fully verify the robustness and reliability of the proposed algorithm in achieving high\-precision, sustained, and stable tracking within complex, dynamic underwater ad\-hoc networks\.

#### VI\-B3Mean Tracking Error \(MTE\)

In multi\-AUV ad\-hoc network cooperative tracking, the mean tracking error quantifies overall temporal tracking deviation and serves as a core metric for evaluating policy accuracy and stability\. Under a unified experimental setup, MTE is defined as the sample mean of Euclidean distances between each agent and its target, across all timesteps and agents per evaluation episode\. This formulation comprehensively reflects the error level throughout execution, rather than solely on terminal states\. Table[III](https://arxiv.org/html/2608.12436#S6.T3)shows that VGG\-MADiffRL achieves the lowest MTE value among all compared methods, demonstrating a more pronounced error advantage\. These results indicate the proposed method enables higher\-precision and more robust continuous cooperative tracking in complex dynamic underwater ad\-hoc networks\.

#### VI\-B4Error Standard Deviation \(Error Std\)

To characterize the stability and fluctuation of tracking errors in multi\-AUV ad\-hoc network cooperative tracking, this study computes the standard deviation of distance error samples between each agent and its target across all timesteps and agents per evaluation episode, defined as the Error Std metric\. This indicator quantifies error dispersion: a smaller value signifies more temporally consistent tracking performance with reduced fluctuations\. Table[IV](https://arxiv.org/html/2608.12436#S6.T4)shows that VGG\-MADiffRL achieves the lowest Error Std among all compared methods, demonstrating that the proposed algorithm not only maintains a low mean tracking error but also exhibits superior error suppression and more stable dynamic tracking performance in complex underwater ad\-hoc networks\. In Table[IV](https://arxiv.org/html/2608.12436#S6.T4), standard deviations below the reported decimal precision are denoted as±0\.0000\\pm 0\.0000\.

![Refer to caption](https://arxiv.org/html/2608.12436v1/diffusion_numbers.png)Fig\. 5:Convergence Speed Across Different Numbers of Diffusion Time Steps in VGG\-MADiffRL![Refer to caption](https://arxiv.org/html/2608.12436v1/ablation_4-2.png)\(a\)Ablation analysis in 4 AUVs Tracking 2 Targets
![Refer to caption](https://arxiv.org/html/2608.12436v1/ablation_8-3.png)\(b\)Ablation analysis in 8 AUVs Tracking 3 Targets

Fig\. 6:Ablation Evaluation![Refer to caption](https://arxiv.org/html/2608.12436v1/Middle_8.png)![Refer to caption](https://arxiv.org/html/2608.12436v1/End_8.png)\(a\) Mid\-Phase of 8 AUVs Tracking 3 Targets\(b\) Final Phase of 8 AUVs Tracking 3 Targets![Refer to caption](https://arxiv.org/html/2608.12436v1/Middle_10.png)![Refer to caption](https://arxiv.org/html/2608.12436v1/End_10.png)\(c\) Mid\-Phase of 10 AUVs Tracking 3 Targets\(d\) Final Phase of 10 AUVs Tracking 3 TargetsFig\. 7:Availability Evaluation
#### VI\-B5Number of Diffusion Time Steps

The number of diffusion steps, the count of reverse sampling iterations in the diffusion\-based policy generation process, is a critical hyperparameter affecting strategy accuracy and computational overhead\. A larger number of diffusion steps enables more thorough denoising in the reverse process, theoretically yielding higher\-quality action distributions and more refined policy representations, yet concurrently increases inference overhead and latency\. Fig\.[5](https://arxiv.org/html/2608.12436#S6.F5)illustrates a scenario of 4 AUVs in an ad\-hoc network Tracking 2 Targets, where comparative evaluations under different diffusion step settings are conducted under unified experimental conditions to analyze their trade\-off effects on control precision and efficiency\. Experimental results demonstrate that the adopted diffusion step configuration achieves a better balance between tracking effectiveness and computational cost, while maintaining satisfactory cooperative tracking performance in multi\-AUV ad\-hoc networks\.

#### VI\-B6Ablation Evaluation

To systematically evaluate the contribution of each key module in the proposed method for multi\-AUV ad\-hoc networks, an ablation study is conducted\. The following comparative variants are configured: \(1\) removal of the value\-gradient\-guided reverse diffusion mechanism \(excluding value function guidance during action sampling\); and \(2\) replacement of the diffusion policy module with a conventional deterministic policy network to examine the individual impact of diffusion modeling on performance\. All other training configurations remain identical to ensure a fair comparison\. Fig\.[6](https://arxiv.org/html/2608.12436#S6.F6)shows the complete method consistently outperforms both ablated variants in convergence stability and cumulative return\. These results demonstrate that both the value\-gradient guidance mechanism and the diffusion\-based policy module play critical roles in enhancing cooperative tracking performance, thereby validating the effectiveness of the proposed architectural design\.

#### VI\-B7Availability Evaluation

To show the training convergence and cooperative tracking performance of the proposed method, we build a high\-fidelity underwater simulation environment using the 3D modeling and physics engine of Unity\. Fig\.[7](https://arxiv.org/html/2608.12436#S6.F7)visualizes the cooperative tracking process and environment configuration, showing the mid and late stages of 10 AUVs tracking three targets and 8 AUVs tracking three targets\. Yellow moving entities represent AUVs, glowing spheres represent dynamic targets, spirals represent obstacles, and colored trajectory lines mark the historical path of each AUV\. The simulation reproduces complex underwater dynamics \(acoustic communication constraints, ocean current disturbances, and time\-varying network topologies\) and provides a reliable platform for validating the stability and effectiveness of the algorithm\. All AUV nodes, targets, and obstacles are modeled and rendered in real time, enabling direct assessment of the algorithm’s operation in dynamic ad\-hoc networks\.

## VIIConclusion

This paper investigated cooperative target tracking in multi\-AUV ad\-hoc networks and proposed the MDCA hierarchical control architecture together with the VGG\-MADiffRL algorithm\. MDCA decomposes global coordination into three layers \(global intelligent control, local online training, and physical action execution\), enabling synergistic optimization under the CTDE paradigm\. VGG\-MADiffRL introduces three key innovations: a diffusion\-based policy that replaces the deterministic Actor to model complex continuous action distributions; a dual\-objective joint optimization mechanism that combines Q\-guided and policy gradient losses to stabilize training under volatile topologies; and a value\-gradient\-guided reverse sampling mechanism that steers the denoising process toward high\-return action regions, reducing ineffective sampling and improving policy robustness\.

Extensive experiments across four multi\-AUV tracking scenarios demonstrate that VGG\-MADiffRL consistently outperforms seven state\-of\-the\-art MARL algorithms in convergence speed, tracking accuracy, mean tracking error, and error stability\. Ablation studies confirm that both the value\-gradient guidance and the diffusion policy module contribute substantially to overall performance\.

Several directions warrant further work: \(1\) optimizing underwater obstacle avoidance to reduce potential AUV damage; \(2\) balancing energy consumption among AUVs to extend system endurance; and \(3\) designing robust control frameworks that explicitly account for unstable underwater acoustic communication\.

## References

- \[1\]J\. Ackermann, V\. Gabler, T\. Osa, and M\. Sugiyama\(2019\)Reducing overestimation bias in multi\-agent domains using double centralized critics\.External Links:1910\.01465Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[2\]D\. M\. Bailey and C\. R\. Hopkins\(2023\)Sustainable use of ocean resources\.Marine Policy154,pp\. 105672\.External Links:ISSN 0308\-597X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.marpol.2023.105672)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[3\]J\. Huang, X\. Ye, Y\. Wang, and L\. Fu\(2025\)Leveraging propagation delays: a delay\-aware multiagent reinforcement learning mac protocol for underwater acoustic networks\.IEEE Internet of Things Journal12\(20\),pp\. 42076–42089\.External Links:[Document](https://dx.doi.org/10.1109/JIOT.2025.3595133)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[4\]M\. Janner, Y\. Du, J\. Tenenbaum, and S\. Levine\(2022\)Planning with diffusion for flexible behavior synthesis\.InProceedings of the 39th International Conference on Machine Learning,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvari, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 9902–9915\.Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p3.1)\.
- \[5\]P\. Jiang, S\. Song, and G\. Huang\(2022\)Attention\-based meta\-reinforcement learning for tracking control of auv with time\-varying dynamics\.IEEE Transactions on Neural Networks and Learning Systems33\(11\),pp\. 6388–6401\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2021.3079148)Cited by:[§II\-A](https://arxiv.org/html/2608.12436#S2.SS1.p3.1)\.
- \[6\]S\. Jiang\(2019\)On securing underwater acoustic networks: a survey\.IEEE Communications Surveys & Tutorials21\(1\),pp\. 729–752\.External Links:[Document](https://dx.doi.org/10.1109/COMST.2018.2864127)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[7\]Y\. Jiang, K\. Zhang, M\. Zhao, and H\. Qin\(2024\)Adaptive meta\-reinforcement learning for auvs 3d guidance and control under unknown ocean currents\.Ocean Engineering309,pp\. 118498\.External Links:ISSN 0029\-8018,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.oceaneng.2024.118498)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[8\]J\. Li and Q\. Chen\(2026\)A multi\-auv adaptive collaborative target coverage algorithm for unknown environment\.Ad Hoc Networks180,pp\. 104033\.External Links:ISSN 1570\-8705,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.adhoc.2025.104033)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[9\]L\. Li, R\. An, Z\. Guo, and J\. Gao\(2025\)Multi\-auv cooperative search for moving targets based on multi\-agent reinforcement learning\.Journal of Marine Science and Engineering13\(11\)\.External Links:ISSN 2077\-1312,[Document](https://dx.doi.org/10.3390/jmse13112072)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[10\]Z\. Li, J\. Du, C\. Jiang, W\. Mi, and Y\. Ren\(2024\)HA\-marl: heuristic and apf assisted multi\-agent reinforcement learning for wireless data sharing in auv swarms\.InICC 2024 \- IEEE International Conference on Communications,Vol\.,pp\. 5401–5406\.External Links:[Document](https://dx.doi.org/10.1109/ICC51166.2024.10622437)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[11\]D\. Liang, J\. Chu, Y\. Cui, D\. Liang, and Y\. Feng\(2024\)Underwater dynamic tracking control of auv based on complex environment simulation and acl\-sac deep reinforcement learning\.IEEE Transactions on Intelligent Vehicles\(\),pp\. 1–16\.External Links:[Document](https://dx.doi.org/10.1109/TIV.2024.3443393)Cited by:[§II\-A](https://arxiv.org/html/2608.12436#S2.SS1.p1.1)\.
- \[12\]R\. Lowe, Y\. Wu, A\. Tamar, J\. Harb, P\. Abbeel, and I\. Mordatch\(2017\)Multi\-agent actor\-critic for mixed cooperative\-competitive environments\.InProceedings of the 31st International Conference on Neural Information Processing Systems,NIPS’17,Red Hook, NY, USA,pp\. 6382–6393\.External Links:ISBN 9781510860964Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[13\]S\. Luo, Y\. Li, S\. Liu, X\. Zhang, Y\. Shao, and C\. Wu\(2024\)Multi\-agent continuous control with generative flow networks\.Neural Networks174,pp\. 106243\.External Links:ISSN 0893\-6080,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neunet.2024.106243)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p3.1)\.
- \[14\]A\. Luvisutto, A\. Celani, F\. Renda, C\. Stefanini, and G\. De Masi\(2025\)Enhancing collaboration in uncertain environment: multi\-agent reinforcement learning for underwater monitoring\.Expert Systems with Applications277,pp\. 127256\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.eswa.2025.127256)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[15\]S\. A\. H\. Mohsan, Y\. Li, M\. Sadiq, J\. Liang, and M\. A\. Khan\(2023\)Recent advances, future trends, applications and challenges of internet of underwater things \(iout\): a comprehensive review\.Journal of Marine Science and Engineering11\(1\)\.External Links:ISSN 2077\-1312,[Document](https://dx.doi.org/10.3390/jmse11010124)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[16\]L\. Paull, S\. Saeedi, M\. Seto, and H\. Li\(2014\)AUV navigation and localization: a review\.IEEE Journal of Oceanic Engineering39\(1\),pp\. 131–149\.External Links:[Document](https://dx.doi.org/10.1109/JOE.2013.2278891)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[17\]H\. Peng, K\. Jiang, D\. Yuan, Z\. Zeng, and Z\. Wu\(2026\)Bio\-inspired hierarchical multi\-agent reinforcement learning for auv swarm energy weakest\-link mitigation\.Ocean Engineering343,pp\. 123385\.External Links:ISSN 0029\-8018,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.oceaneng.2025.123385)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[18\]L\. Pimentel, S\. Ye, J\. E\. G\. Pagan, and M\. Gombolay\(2025\)Diverse heterogeneous graph conditioned diffusion for multi\-agent teaming\.InProceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems,AAMAS ’25,Richland, SC,pp\. 2714–2716\.External Links:ISBN 9798400714269Cited by:[§II\-B](https://arxiv.org/html/2608.12436#S2.SS2.p2.1)\.
- \[19\]T\. Rashid, M\. Samvelyan, C\. S\. De Witt, G\. Farquhar, J\. Foerster, and S\. Whiteson\(2020\)Monotonic value function factorisation for deep multi\-agent reinforcement learning\.J\. Mach\. Learn\. Res\.21\(1\)\.External Links:ISSN 1532\-4435Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p2.1)\.
- \[20\]C\. Wang, J\. Du, J\. Wang, and Y\. Ren\(2021\)AUV path following control using deep reinforcement learning under the influence of ocean currents\.InProceedings of the 2021 5th International Conference on Digital Signal Processing,ICDSP ’21,New York, NY, USA,pp\. 225–231\.External Links:ISBN 9781450389365,[Document](https://dx.doi.org/10.1145/3458380.3459041)Cited by:[§II\-A](https://arxiv.org/html/2608.12436#S2.SS1.p2.1)\.
- \[21\]S\. Wang, C\. Lin, G\. Han, S\. Zhu, Z\. Li, Z\. Wang, and Y\. Ma\(2025\)Multi\-auv cooperative underwater multi\-target tracking based on dynamic\-switching\-enabled multi\-agent reinforcement learning\.IEEE Transactions on Mobile Computing24\(5\),pp\. 4296–4311\.External Links:[Document](https://dx.doi.org/10.1109/TMC.2024.3521889)Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[22\]X\. Wu, X\. Li, J\. Li, P\. C\. Ching, V\. C\. M\. Leung, and H\. V\. Poor\(2021\)Caching transient content for iot sensing: multi\-agent soft actor\-critic\.IEEE Transactions on Communications69\(9\),pp\. 5886–5901\.External Links:[Document](https://dx.doi.org/10.1109/TCOMM.2021.3086535)Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[23\]Y\. Xue, M\. Mao, X\. Ru, Y\. Zhu, B\. Ren, S\. Qiao, M\. Wang, S\. Deng, X\. An, N\. Zhang, Y\. Chen, and H\. Chen\(2025\)OceanGym: a benchmark environment for underwater embodied agents\.External Links:2509\.26536,[Link](https://arxiv.org/abs/2509.26536)Cited by:[§VI\-A](https://arxiv.org/html/2608.12436#S6.SS1.p2.1)\.
- \[24\]X\. Yan, W\. Luo, J\. Jia, D\. Jiang, and T\. Zhang\(2025\)Static consensus analysis of multi\-auv systems with impulsive protocol and time delays\.Ocean Engineering331,pp\. 121370\.External Links:ISSN 0029\-8018,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.oceaneng.2025.121370)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p1.1)\.
- \[25\]Y\. Yang, W\. Ma, W\. Sun, J\. He, Y\. Fu, C\. Yuen, and Y\. Zhang\(2025\)Diffusion\-based multi\-agent reinforcement learning for semantic vehicular edge computing\.IEEE Transactions on Services Computing18\(6\),pp\. 3668–3681\.External Links:[Document](https://dx.doi.org/10.1109/TSC.2025.3618082)Cited by:[§II\-B](https://arxiv.org/html/2608.12436#S2.SS2.p3.1)\.
- \[26\]C\. Yu, A\. Velu, E\. Vinitsky, J\. Gao, Y\. Wang, A\. Bayen, and Y\. WU\(2022\)The surprising effectiveness of ppo in cooperative multi\-agent games\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 24611–24624\.Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p3.1),[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[27\]K\. Zhang, Z\. Yang, and T\. Başar\(2021\)Multi\-agent reinforcement learning: a selective overview of theories and algorithms\.InHandbook of Reinforcement Learning and Control,K\. G\. Vamvoudakis, Y\. Wan, F\. L\. Lewis, and D\. Cansever \(Eds\.\),pp\. 321–384\.External Links:ISBN 978\-3\-030\-60990\-0,[Document](https://dx.doi.org/10.1007/978-3-030-60990-0%5F12)Cited by:[§I](https://arxiv.org/html/2608.12436#S1.p3.1)\.
- \[28\]J\. Zhao, T\. Zhu, S\. Xiao, Z\. Gao, and H\. Sun\(2022\)Actor\-critic for multi\-agent reinforcement learning with self\-attention\.International Journal of Pattern Recognition and Artificial Intelligence36\(09\),pp\. 2252014\.External Links:[Document](https://dx.doi.org/10.1142/S0218001422520140)Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[29\]S\. Zhu, G\. Han, C\. Lin, and Q\. Tao\(2024\)Underwater target tracking based on hierarchical software\-defined multi\-auv reinforcement learning: a multi\-auv advantage\-attention actor\-critic approach\.IEEE Transactions on Mobile Computing23\(12\),pp\. 13639–13653\.External Links:[Document](https://dx.doi.org/10.1109/TMC.2024.3437376)Cited by:[§VI\-B](https://arxiv.org/html/2608.12436#S6.SS2.p1.1)\.
- \[30\]Z\. Zhu, M\. Liu, L\. Mao, B\. Kang, M\. Xu, Y\. Yu, S\. Ermon, and W\. Zhang\(2024\)MADiff: offline multi\-agent learning with diffusion models\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,Cited by:[§II\-B](https://arxiv.org/html/2608.12436#S2.SS2.p1.1)\.

![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/Jiaao_Ma.png)Jiaao Mais currently pursuing a Bachelor’s degree at the Software College, Northeastern University, Shenyang, China\. His research interests include reinforcement learning, diffusion models, and supervised learning\.![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/ChuanLin.png)Chuan Lin\[S’17, M’20\] is currently an associate professor with the Software College, Northeastern University, Shenyang, China\. He received the B\.S\. degree in Computer Science and Technology from Liaoning University, Shenyang, China in 2011, the M\.S\. degree in Computer Science and Technology from Northeastern University, Shenyang, China in 2013, and the Ph\.D\. degree in computer architecture in 2018\. From Nov\. 2018 to Nov\. 2020, he was a Postdoctoral Researcher with the School of Software, Dalian University of Technology, Dalian, China\. His research interests include UWSNs, industrial IoT, software\-defined networking\.![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/GuangjieHan.png)Guangjie Han\(Fellow, IEEE\) is currently a Professor with the Department of Internet of Things Engineering, Hohai University, Changzhou, China\. He received his Ph\.D\. degree from Northeastern University, Shenyang, China, in 2004\. In February 2008, he finished his work as a Postdoctoral Researcher with the Department of Computer Science, Chonnam National University, Gwangju, Korea\. From October 2010 to October 2011, he was a Visiting Research Scholar with Osaka University, Suita, Japan\. From January 2017 to February 2017, he was a Visiting Professor with City University of Hong Kong, China\. From July 2017 to July 2020, he was a Distinguished Professor with Dalian University of Technology, China\. His current research interests include Internet of Things, Industrial Internet, Machine Learning and Artificial Intelligence, Mobile Computing, Security and Privacy\. Dr\. Han has over 500 peer\-reviewed journal and conference papers, in addition to 160 granted and pending patents\. Currently, his H\-index is 82 and i10\-index is 400 in Google Citation \(Google Scholar\)\. The total citation count of his papers raises above 25000 times\. Dr\. Han is a Fellow of the UK Institution of Engineering and Technology \(FIET\)\. He has served on the Editorial Boards of up to 10 international journals, including the IEEE TII, IEEE TCCN, IEEE TVT, IEEE TNSM, IEEE Systems, etc\. He has guest\-edited several special issues in IEEE Journals and Magazines, including the IEEE JSAC, IEEE Communications, IEEE Wireless Communications, Computer Networks, etc\. Dr\. Han has also served as chair of organizing and technical committees in many international conferences\. He has been awarded 2020 IEEE Systems Journal Annual Best Paper Award and the 2017\-2019 IEEE ACCESS Outstanding Associate Editor Award\. He is a Fellow of IEEE\.![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/Zhu.jpg)Shengchao Zhu\(Student member, IEEE\) received his B\.S\. degree in Internet of Things Engineering from Hohai University, Changzhou, China, in 2023\. He is currently pursuing the Ph\.D\. degree with the Department of Computer Science and Technology at Hohai University, Nanjing, China\. His current research interests include swarm intelligence, swarm ocean, Multi\-Agent Reinforcement Learning\.![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/QianZhu.png)Qian Zhuis an associate professor with the Software College, Northeastern University, Shenyang, China\. She received the B\.S\. degree in Information and Computing Science \(2006\), the M\.S\. degree in Operation Science and Control Theory \(2008\), and the Ph\.D\. degree in Communication and Information System \(2018\), all from Northeastern University, Shenyang, China\. Her research interests include artificial intelligence optimization algorithms, Unmanned Aerial Vehicle \(UAV\) technology and software development for applications\.![[Uncaptioned image]](https://arxiv.org/html/2608.12436v1/YingLiu.jpg)Ying Liureceived the B\.S\., M\.S\., and Ph\.D\. degrees from Northeastern University, Shenyang, China, in 2003, 2006, and 2012, respectively, all in computer science\. She is currently an Associate Professor with the College of Software, Northeastern University\. She has published over 50 articles, and refereed conference papers\. Her current research interests include Service Computing and Edge Computing\.

Similar Articles

Learning to cooperate, compete, and communicate

OpenAI Blog

OpenAI presents research on multi-agent reinforcement learning environments where agents learn to cooperate, compete, and communicate. The paper introduces MADDPG (Multi-Agent DDPG), a centralized critic approach that enables agents to learn collaborative strategies and communication protocols more effectively than traditional decentralized methods.