Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks

arXiv cs.AI Papers

Summary

This paper proposes JUROR, a reinforcement learning-based framework that jointly optimizes UAV flight paths and decentralized opportunistic routing in delay-tolerant networks under centralized training and decentralized execution.

arXiv:2608.04590v1 Announce Type: new Abstract: The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connectivity. However, intermittent contacts, finite buffers, and limited message time-to-live (TTL) often give rise to sparse delivery and congestion, leading to substantial end-to-end performance degradation. To address this challenge, this study explores the joint optimization of decentralized opportunistic routing and controllable unmanned aerial vehicle (UAV) flight, aiming to enlarge future contacts through discrete UAV headings while enabling per-node replication under contact-limited observations. Building upon this architecture, we study cooperative factored routing--UAV control under centralized training and decentralized execution (CTDE) and propose JUROR (Joint UAV flight and Opportunistic Routing, based on the proximal policy optimization (PPO) framework. In our design, we first cast the problem as a factored partially observable Markov decision process with sequential motion--routing coupling and a per-step team reward; subsequently, decentralized actors act on local observations while a training-time critic uses global statistics, and an optional multi-horizon hotspot predictor provides auxiliary supervision. Simulation results over four traffic modes demonstrate effective gains over PRoPHET and MaxProp, while retaining contact-limited decentralized execution.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:42 AM

# Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks
Source: [https://arxiv.org/html/2608.04590](https://arxiv.org/html/2608.04590)
Xiao Wang and Shun\-Ren YangX\. Wang and S\.\-R\. Yang are with the Department of Computer Science and the Institute of Communications Engineering, National Tsing Hua University, Hsinchu 30013, Taiwan \(e\-mail: sryang@cs\.nthu\.edu\.tw\)\.

###### Abstract

The growing deployment of delay\-tolerant networks \(DTNs\) has made store\-carry\-forward \(SCF\) communication indispensable under sparse connectivity\. However, intermittent contacts, finite buffers, and limited message time\-to\-live \(TTL\) often give rise to sparse delivery and congestion, leading to substantial end\-to\-end performance degradation\. To address this challenge, this study explores the joint optimization of decentralized opportunistic routing and controllable unmanned aerial vehicle \(UAV\) flight, aiming to enlarge future contacts through discrete UAV headings while enabling per\-node replication under contact\-limited observations\. Building upon this architecture, we study cooperative factored routing–UAV control under centralized training and decentralized execution \(CTDE\) and propose JUROR \(JointUAV flight andOpportunisticRouting\), based on the proximal policy optimization \(PPO\) framework\. In our design, we first cast the problem as a factored partially observable Markov decision process with sequential motion–routing coupling and a per\-step team reward; subsequently, decentralized actors act on local observations while a training\-time critic uses global statistics, and an optional multi\-horizon hotspot predictor provides auxiliary supervision\. Simulation results over four traffic modes demonstrate effective gains over PRoPHET and MaxProp, while retaining contact\-limited decentralized execution\.

## IIntroduction

In recent years, delay\-tolerant networks \(DTNs\) have attracted extensive attention as a communication paradigm for environments in which persistent end\-to\-end connectivity cannot be guaranteed\. Applications span sparse vehicular networks, disaster response\[[1](https://arxiv.org/html/2608.04590#bib.bib1)\], and remote sensing, wherein SCF relays opportunistically exchange messages under intermittent contacts, finite buffer occupancy, and limited message time\-to\-live \(TTL\)\. Classical utility\-based protocols such as PRoPHET\[[2](https://arxiv.org/html/2608.04590#bib.bib2)\]estimate delivery likelihood from encounter histories and forward toward nodes with higher predicted utility\. Although such hand\-crafted metrics remain effective in selected scenarios, they often require re\-tuning once mobility patterns, traffic intensity, or network topology evolve, which limits robustness across heterogeneous deployment conditions\.

Unmanned aerial vehicles \(UAVs\) enlarge contact opportunities through extended radio range and controllable trajectories\[[3](https://arxiv.org/html/2608.04590#bib.bib3),[4](https://arxiv.org/html/2608.04590#bib.bib4)\]\. Unlike fixed ground mobility, however, each UAV heading decision reshapes the contact graph at subsequent steps, so relay placement and forwarding become tightly coupled control problems\. Contacts remain range\-limited and transient even with UAV assistance, and heterogeneous ground/UAV communication ranges further complicate pairwise node encounters along road\-constrained routes\. Consequently, UAV\-assisted DTN control must reason about*future*topology rather than only about the current encounter set\.

Joint UAV\-assisted DTN control faces three coupled difficulties\. First,time\-varying topology: UAV motion and ground mobility jointly change the contact graph, so routing at one step alters future node contacts\[[4](https://arxiv.org/html/2608.04590#bib.bib4),[3](https://arxiv.org/html/2608.04590#bib.bib3)\]\. Second,decentralized execution versus team learning: cooperative credit assignment for joint routing and UAV control is difficult when execution\-time agents act only on local and contact\-exchangeable inputs rather than a global network state\[[5](https://arxiv.org/html/2608.04590#bib.bib5),[6](https://arxiv.org/html/2608.04590#bib.bib6)\]\. Third,weak and delayed learning signals: delivery events are sparse relative to forwarding attempts, while congestion hotspots evolve with buffer pressure and traffic injection, leaving open whether optional hotspot auxiliary learning should complement delivery\-centric team rewards\[[7](https://arxiv.org/html/2608.04590#bib.bib7),[3](https://arxiv.org/html/2608.04590#bib.bib3),[8](https://arxiv.org/html/2608.04590#bib.bib8)\]\.

These challenges remain unaddressed by three mainstream research branches\. First, classical utility\-based routing protocols such as PRoPHET\[[2](https://arxiv.org/html/2608.04590#bib.bib2)\]solely optimize terrestrial forwarding strategies and assume UAV flight paths are given in advance rather than controllable routing decisions\. Second, hand\-crafted UAV\-assisted routing and disaster\-targeted UAV–DTN route discovery schemes\[[4](https://arxiv.org/html/2608.04590#bib.bib4),[3](https://arxiv.org/html/2608.04590#bib.bib3),[9](https://arxiv.org/html/2608.04590#bib.bib9)\]refine encounter scoring metrics and path construction logic, yet they fully decouple UAV flight adjustments from decentralized per\-node SCF forwarding decisions\. Third, routing\-only RL and tabular Q\-learning variants\[[10](https://arxiv.org/html/2608.04590#bib.bib10),[11](https://arxiv.org/html/2608.04590#bib.bib11),[12](https://arxiv.org/html/2608.04590#bib.bib12)\]learn forwarding policies under intermittent connectivity; they neither co\-optimize discrete UAV headings alongside per\-node routing actions nor offer modular auxiliary hotspot prediction modules\. Furthermore, existing multi\-UAV MADRL methods for trajectory and transmission control such as\[[13](https://arxiv.org/html/2608.04590#bib.bib13)\]focus solely on generic wireless relaying and neglect DTN\-native SCF message replication mechanisms\. Collectively, the above limitations inspire the proposed learning framework that integrates joint routing–UAV control, deployment\-adaptive observation inputs, and a systematically evaluated design space for optional hotspot modules\.

This framework is namedJUROR\(JointUAV flight andOpportunistic Routing\): a discrete\-time SCF system trained withCentralized Training and Decentralized Execution\(CTDE\) andProximal Policy Optimization\(PPO\)\[[14](https://arxiv.org/html/2608.04590#bib.bib14),[15](https://arxiv.org/html/2608.04590#bib.bib15),[16](https://arxiv.org/html/2608.04590#bib.bib16)\]\. It employs factored cooperative policies: every agent chooses constrained local forwarding candidates and discrete UAV headings, all optimized under a unified episodic objective decomposed into per\-step team rewards\. The below are our contributions\.

- •Episodic optimization and joint MDP formulation:We formalize joint UAV–routing control as a cooperative episodic optimization problem \(P1\) over delivery, congestion, and fleet\-placement metrics under SCF constraints, cast it as a factored partially observable MDP with sequential routing–motion coupling, and derive a per\-step reward proxy that bridges \(P1\) to CTDE–PPO learning\.
- •CTDE factorization architecture:Fig\.[2](https://arxiv.org/html/2608.04590#S5.F2)presents JUROR as an MDP interface, a four\-stage CTDE–PPO loop, and explicit actor–critic layers with an optional LSTM branch\. At inference, routing and UAV actors use local on\-board and contact\-exchangeable inputs only\.
- •Optional hotspot auxiliaries:JUROR optionally attaches a UAV LSTM predictor trained by multi\-horizon replay supervision \(LpL\_\{p\}\), and optionally activates hotspot\-guided alignment \(HGA\) reward shaping that scores agreement between executed UAV headings and sector\-stress directions\. The default experimental stack disables both \(λ=0\\lambda\{=\}0, HGA off\); they are evaluated as ablations because their delivery benefit is traffic\-dependent \(Sec\.[VI](https://arxiv.org/html/2608.04590#S6)\)\.

Paper organization\.Sec\.[II](https://arxiv.org/html/2608.04590#S2)reviews related work on DTN routing, UAV\-assisted forwarding, and multi\-agent reinforcement learning\. Secs\.[III](https://arxiv.org/html/2608.04590#S3)–[V](https://arxiv.org/html/2608.04590#S5)present the system model, episodic optimization and MDP formulation, and the JUROR CTDE–PPO framework \(Fig\.[2](https://arxiv.org/html/2608.04590#S5.F2)\)\. Sec\.[VI](https://arxiv.org/html/2608.04590#S6)presents simulation results and discussions along a staged pipeline: simulation settings, ablation study of core innovations, baseline routing comparison, and mechanism analysis\. Sec\.[VII](https://arxiv.org/html/2608.04590#S7)concludes the paper\.

## IIRelated Work

### II\-ADTN Routing: Classical and Learning\-Based

DTN routing protocols span flooding and epidemic families, quota\-based spray methods, and utility\-based forwarders\[[17](https://arxiv.org/html/2608.04590#bib.bib17)\]\. Utility\-based schemes such as PRoPHET\[[2](https://arxiv.org/html/2608.04590#bib.bib2)\]and MaxProp\[[18](https://arxiv.org/html/2608.04590#bib.bib18)\]rank next hops from encounter histories, meeting recency, and buffer pressure; recent variants add cache\-state awareness\[[19](https://arxiv.org/html/2608.04590#bib.bib19)\], social/geographic context\[[6](https://arxiv.org/html/2608.04590#bib.bib6),[20](https://arxiv.org/html/2608.04590#bib.bib20)\], or comparative benchmarks under contact\-plan uncertainty\[[21](https://arxiv.org/html/2608.04590#bib.bib21)\]\. These methods provide interpretable baselines used in our comparison, but they optimize hand\-crafted forwarding scores while treating relay trajectories as exogenous\.

Learning\-based routing extends this line with tabular and deep RL\[[5](https://arxiv.org/html/2608.04590#bib.bib5),[22](https://arxiv.org/html/2608.04590#bib.bib22)\]\. FQLRP\[[10](https://arxiv.org/html/2608.04590#bib.bib10)\], Q\-learning spray\-and\-wait variants\[[11](https://arxiv.org/html/2608.04590#bib.bib11)\], and disaster\-recovery multi\-agent DRL\[[8](https://arxiv.org/html/2608.04590#bib.bib8)\]adapt forwarding from buffer, mobility, or social\-context proxies\. Predictive routers such as PF\-DTN\[[7](https://arxiv.org/html/2608.04590#bib.bib7)\]couple LSTM trajectory forecasting with hybrid relay ranking\. Across both classical and learned routers, the dominant pattern is*routing\-only*control: per\-node forwarding may adapt, but UAV headings are not co\-optimized jointly with replication in a shared SCF simulator\.

### II\-BUAV\-Assisted DTN and Vehicular Routing

UAV relays enlarge contact opportunities through extended range and controllable trajectories, yet routing must still account for heterogeneous mobility, link persistence, and fleet coordination\[[23](https://arxiv.org/html/2608.04590#bib.bib23),[1](https://arxiv.org/html/2608.04590#bib.bib1)\]\. Du et al\.\[[3](https://arxiv.org/html/2608.04590#bib.bib3)\]weight meeting probability and persistent connection time in UAV\-assisted vehicular DTNs; Fan et al\.\[[4](https://arxiv.org/html/2608.04590#bib.bib4)\]rank relays with multi\-attribute utilities and dynamic buffer prioritization for post\-disaster VANETs\. Disaster\-targeted UAV–DTN route discovery\[[9](https://arxiv.org/html/2608.04590#bib.bib9)\]and Q\-learning anycast routing over multi\-base\-station UAV graphs\[[12](https://arxiv.org/html/2608.04590#bib.bib12)\]further improve forwarding under mobility\. These protocols refine encounter scoring, trajectory utilities, or path construction, but they decouple flight planning from decentralized SCF decisions at every graph node and do not jointly learn headings with per\-node replication under road\-constrained mobility and heterogeneous communication ranges\.

### II\-CCooperative MARL, CTDE, and Training

Cooperative multi\-agent reinforcement learning \(MARL\) addresses distributed control with coupled dynamics; independent learners can be unstable when each agent treats partners as non\-stationary environment noise\. Centralized training and decentralized execution \(CTDE\) trains centralized critics on global or summarized state while keeping actors decentralized at execution time\[[14](https://arxiv.org/html/2608.04590#bib.bib14),[15](https://arxiv.org/html/2608.04590#bib.bib15)\]\. Recent GNN\-based multiagent DRL for interplanetary networks\[[24](https://arxiv.org/html/2608.04590#bib.bib24),[25](https://arxiv.org/html/2608.04590#bib.bib25)\]applies PPO with graph attention for distributed routing and scheduling on planned contact graphs\. Multi\-UAV MADRL for wireless relay networks\[[13](https://arxiv.org/html/2608.04590#bib.bib13)\]co\-optimizes trajectory and transmission but models forwarding as generic relay service rather than decentralized message replication with buffer pressure and delivery\-centric team rewards\. These MARL lines scale cooperative control on graphs or aerial relays, yet they rarely joint\-learn discrete UAV headings with per\-node SCF actions in DTN systems\.

JUROR adopts CTDE–PPO\[[16](https://arxiv.org/html/2608.04590#bib.bib16)\]for factored discrete routing logits and UAV headings: routing actors use local observations, and a centralized critic trains on fixed\-length network aggregates rather than per\-message states\. When delivery rewards are sparse, auxiliary losses can stabilize representation learning\[[26](https://arxiv.org/html/2608.04590#bib.bib26)\]; JUROR optionally supervises per\-UAV hotspot predictors with multi\-horizon replay\-aligned targets and ablates ground\-truth versus predicted hotspot\-guided alignment \(HGA\) reward coupling\.

### II\-DPaper Positioning Against Existing Works

Table I:Positioning of JUROR relative to four representative research lines\. Column headers:*Joint UAV & routing*, joint UAV motion and opportunistic routing;*SCF sim\.*, SCF simulator;*CTDE/MARL*, centralized training with decentralized execution;*Hotspot aux\.*, optional hotspot predictor and HGA shaping;*Local deploy*, contact\-limited decentralized actors\. ✓: core native capability; \(✓\): partial, optional or conditionally effective function; –: not supported\.Research lineJoint UAV & routingSCF sim\.CTDE/MARLHotspot aux\.Local deployClassical DTN routing\[[2](https://arxiv.org/html/2608.04590#bib.bib2),[18](https://arxiv.org/html/2608.04590#bib.bib18)\]–\(✓\)––✓Hand\-crafted UAV DTN/VANET\[[3](https://arxiv.org/html/2608.04590#bib.bib3),[4](https://arxiv.org/html/2608.04590#bib.bib4)\]\(✓\)\(✓\)––✓Learning\-based DTN routing\[[10](https://arxiv.org/html/2608.04590#bib.bib10),[11](https://arxiv.org/html/2608.04590#bib.bib11),[8](https://arxiv.org/html/2608.04590#bib.bib8),[7](https://arxiv.org/html/2608.04590#bib.bib7)\]–\(✓\)\(✓\)\(✓\)\(✓\)Generic MARL routing\[[24](https://arxiv.org/html/2608.04590#bib.bib24)\]\(✓\)–✓–\(✓\)JUROR \(this work\)✓✓✓\(✓\)✓Table[I](https://arxiv.org/html/2608.04590#S2.T1)quantitatively contrasts JUROR against four mainstream research categories\. 1\) Classical DTN routing delivers interpretable forwarding rules without learnable UAV motion control; 2\) Handcrafted UAV\-DTN protocols add relay scoring logic but separate flight and forwarding optimization; 3\) Learning\-based DTN solutions improve adaptive forwarding but remain routing\-centric without joint UAV control; 4\) Generic MARL schedulers scale for contact graphs yet lack complete SCF modeling, deployment\-aware local observation design\. JUROR bridges UAV\-assisted DTN and cooperative MARL research\. Built on a unified discrete\-time SCF system under CTDE\-PPO, it jointly optimizes per\-node forwarding and UAV headings and supports contact\-limited decentralized inference for real\-world deployment\.

## IIISystem Model

This section first introduces the network model and the UAV system model\. Second, the integrated routing and UAV motion control decision model in the DTN is presented\. Third, we define stress and geometry fields that characterize congestion distribution and fleet placement around UAV relays\.

Table II:Unified notation and actor/critic observation inventory\. Parts I–II, IV: system/MDP symbols\. Part III: observation features with deployment scope \(*Local*: deployable on\-board / contact\-limited signals, including UAV navigation features;*Global*/*Local\-ctx*: mutually exclusive UAV\-actor contextsguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}\(network\-wide; default base\) andguL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\(contact\-limited\);*Train\-only*: critics\(t\)s^\{\\,\(t\)\}, omitted at deployment\)\. Observation layouts follow Eqs\. \([27](https://arxiv.org/html/2608.04590#S4.E27)\), \([26](https://arxiv.org/html/2608.04590#S4.E26)\), and \([16](https://arxiv.org/html/2608.04590#S3.E16)\); experimentalα\\alphavalues in Sec\.[VI\-A](https://arxiv.org/html/2608.04590#S6.SS1)\.ScopeSymbolDescriptionI\. Network, mobility, contacts, buffers, and messages—tt,TmaxT\_\{\\max\};i,ji,j;u,vu,v;iu=Ngr\+ui\_\{u\}\{=\}N\_\{\\mathrm\{gr\}\}\{\+\}u;NN,NgrN\_\{\\mathrm\{gr\}\},NuavN\_\{\\mathrm\{uav\}\}Discrete step and episode horizon; node / UAV indices; graph index of UAVuu; total / ground / UAV counts \(N=Ngr\+NuavN\{=\}N\_\{\\mathrm\{gr\}\}\{\+\}N\_\{\\mathrm\{uav\}\}\)—vgrv\_\{\\mathrm\{gr\}\},vuavv\_\{\\mathrm\{uav\}\};vminv\_\{\\min\},vmaxv\_\{\\max\};𝐩i\(t\)\\mathbf\{p\}\_\{i\}^\{\\,\(t\)\},𝐯i\(t\)\\mathbf\{v\}\_\{i\}^\{\\,\(t\)\};WW,HH,LwL\_\{\\mathrm\{w\}\};di​j\(t\)d\_\{ij\}^\{\\,\(t\)\},ri​jr\_\{ij\};rgrr\_\{\\mathrm\{gr\}\},ruavr\_\{\\mathrm\{uav\}\}Nominal ground/UAV speeds; ground speed bounds; 2\-D position/velocity; map width/height and world spanLwL\_\{\\mathrm\{w\}\}\(distance normalizer\); pairwise distance; type\-dependent communication range \(rgrr\_\{\\mathrm\{gr\}\}orruavr\_\{\\mathrm\{uav\}\}\)—kk,B=8B\{=\}8;𝐝k\\mathbf\{d\}\_\{k\},𝐝u\(t\)\\mathbf\{d\}\_\{u\}^\{\\,\(t\)\};KK,KσK\_\{\\sigma\},ℓ\\ellHeading\-bin index and count; unit vector of binkk; executed UAV heading; max routing candidates; Top\-KσK\_\{\\sigma\}stress neighbors in𝐞u\\mathbf\{e\}\_\{u\}; candidate slot index—Ci​j\(t\)C\_\{ij\}^\{\\,\(t\)\},𝐂\(t\)\\mathbf\{C\}^\{\\,\(t\)\};νi\(t\)\\nu\_\{i\}^\{\\,\(t\)\},ν¯i\(t\)\\bar\{\\nu\}\_\{i\}^\{\\,\(t\)\};bi\(t\)b\_\{i\}^\{\\,\(t\)\},BbB\_\{b\};βig​\(t\)\\beta\_\{i\}^\{\\mathrm\{g\}\\,\(t\)\},βid​\(t\)\\beta\_\{i\}^\{\\mathrm\{d\}\\,\(t\)\}Contact indicator and matrix; raw / normalized degreeν¯i=νi/\(N−1\)\\bar\{\\nu\}\_\{i\}\{=\}\\nu\_\{i\}/\(N\{\-\}1\); buffer fill ratio and capacity; fractions of messages generated / delivered at nodeii—b~i​j\\tilde\{b\}\_\{ij\},ν~i​j\\tilde\{\\nu\}\_\{ij\},τi​j\\tau\_\{ij\};ωex\\omega\_\{\\mathrm\{ex\}\};ℬi\(t\)\\mathcal\{B\}\_\{i\}^\{\\,\(t\)\},ℳj\\mathcal\{M\}\_\{j\};qq,τq\\tau\_\{q\},ιq\(t\)\\iota\_\{q\}^\{\\,\(t\)\},hq\(t\)h\_\{q\}^\{\\,\(t\)\};T0T\_\{0\},ςq\(t\)\\varsigma\_\{q\}^\{\\,\(t\)\}ii’s cached estimate ofjj’s buffer/degree, steps since last exchange, and decay weight; local buffer set / undelivered set atjj; message id, remaining TTL, age, hop count; TTL reference and normalized size—qp\(t\)q\_\{\\mathrm\{p\}\}^\{\\,\(t\)\},t~\(t\)\\tilde\{t\}^\{\\,\(t\)\}Pending\-message ratio \(undelivered / created\) and normalized episode timet/Tmaxt/T\_\{\\max\}; used inguGg\_\{u\}^\{\\mathrm\{G\}\}ands\(t\)s^\{\\,\(t\)\}II\. Delivery stress, stress fields, geometry, and optional HGA—σj\(t\)\\sigma\_\{j\}^\{\\,\(t\)\},υ¯j\(t\)\\bar\{\\upsilon\}\_\{j\}^\{\\,\(t\)\};ζu​v\(t\)\\zeta\_\{uv\}^\{\\,\(t\)\},ξ\(t\)\\xi^\{\\,\(t\)\},κ\(t\)\\kappa^\{\\,\(t\)\}Ground delivery stressbj​\(1\+υ¯j\)b\_\{j\}\(1\{\+\}\\bar\{\\upsilon\}\_\{j\}\)and mean TTL urgency \(Eqs\. \([11](https://arxiv.org/html/2608.04590#S3.E11)\)–\([12](https://arxiv.org/html/2608.04590#S3.E12)\)\); pairwise / fleet\-mean UAV separation \(Eq\. \([10](https://arxiv.org/html/2608.04590#S3.E10)\)\); heading\-switch fraction—ρiu\(t\)\\rho\_\{i\_\{u\}\}^\{\\,\(t\)\},ρ\(t\)\\rho^\{\\,\(t\)\},ρf\(t\)\\rho\_\{f\}^\{\\,\(t\)\}In\-range stress density at relayuu, fleet mean \(Eq\. \([13](https://arxiv.org/html/2608.04590#S3.E13)\)\), and optional EMA forecast stress density \(Eq\. \([15](https://arxiv.org/html/2608.04590#S3.E15)\)\)—𝐰\\mathbf\{w\},𝐯​\(𝐰\)\\mathbf\{v\}\(\\mathbf\{w\}\);α\(t\)\\alpha^\{\\,\(t\)\};ki​j∗\(t\)\{k^\{\*\}\_\{ij\}\}^\{\(t\)\}Nonnegative sector weights and induced unit reference heading \(Eq\. \([29](https://arxiv.org/html/2608.04590#S5.E29)\)\); fleet HGA alignment score; best heading bin of nodejjrelative toii\(Eq\. \([14](https://arxiv.org/html/2608.04590#S3.E14)\)\)III\-A\. Routing actor —oi\(t\)=\(𝐱i\(t\),\{𝐜i,ℓ\(t\)\}\)o\_\{i\}^\{\\,\(t\)\}=\(\\mathbf\{x\}\_\{i\}^\{\\,\(t\)\},\\\{\\mathbf\{c\}\_\{i,\\ell\}^\{\\,\(t\)\}\\\}\)Local𝐱i\(t\)∈ℝ9\\mathbf\{x\}\_\{i\}^\{\\,\(t\)\}\\in\\mathbb\{R\}^\{9\}\(bi,βig,βid,v¯i,c¯i,𝐩¯i,ν¯i,b¯inb\)\(b\_\{i\},\\beta\_\{i\}^\{\\mathrm\{g\}\},\\beta\_\{i\}^\{\\mathrm\{d\}\},\\bar\{v\}\_\{i\},\\bar\{c\}\_\{i\},\\bar\{\\mathbf\{p\}\}\_\{i\},\\bar\{\\nu\}\_\{i\},\\bar\{b\}\_\{i\}^\{\\mathrm\{nb\}\}\): buffer; gen\./del\. ratios;‖𝐯i‖/vmax\\\|\\mathbf\{v\}\_\{i\}\\\|/v\_\{\\max\}; contact count/\(N−1\)\(N\{\-\}1\); position/\(W,H\)/\(W,H\); degree/\(N−1\)\(N\{\-\}1\); exchange\-weighted neighbor buffer \(Sec\.[IV\-C](https://arxiv.org/html/2608.04590#S4.SS3)\)Local𝐜i,ℓ\(t\)∈ℝ7\\mathbf\{c\}\_\{i,\\ell\}^\{\\,\(t\)\}\\in\\mathbb\{R\}^\{7\}For candidateℓ\\elltoward destinationdd: livebdb\_\{d\}or cachedb~i​d\\tilde\{b\}\_\{id\}; cachedν~i​d\\tilde\{\\nu\}\_\{id\};τq/T0\\tau\_\{q\}/T\_\{0\};ιq/T0\\iota\_\{q\}/T\_\{0\};hq/\(N−1\)h\_\{q\}/\(N\{\-\}1\);ςq\\varsigma\_\{q\};di​d/Lwd\_\{id\}/L\_\{\\mathrm\{w\}\}; infeasible slots maskedIII\-B\. UAV motion actor — navigation𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\}/ supervision𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\\,\(t\)\}\(hotspot aux\.\)Local𝐳u\(t\)=\[𝐒u\(t\);𝐞u\(t\);𝐧^u\(t\)\]∈ℝDy\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\}=\[\\mathbf\{S\}\_\{u\}^\{\\,\(t\)\};\\mathbf\{e\}\_\{u\}^\{\\,\(t\)\};\\hat\{\\mathbf\{n\}\}\_\{u\}^\{\\,\(t\)\}\]\\in\\mathbb\{R\}^\{D\_\{y\}\}Navigation vector \(policy input\), assembled from sector stress𝐒u\\mathbf\{S\}\_\{u\}\(Eq\. \([14](https://arxiv.org/html/2608.04590#S3.E14)\)\), Top\-KσK\_\{\\sigma\}offsets𝐞u\\mathbf\{e\}\_\{u\}, and centroid direction𝐧^u\\hat\{\\mathbf\{n\}\}\_\{u\};Dy=8\+2​Kσ\+2D\_\{y\}\{=\}8\{\+\}2K\_\{\\sigma\}\{\+\}2Local𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\\,\(t\)\};𝐲^u\(t\)\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\\,\(t\)\}\(opt\.\);λ\\lambdaSupervision vector \(same layout as𝐳u\\mathbf\{z\}\_\{u\}; GT target for hotspot auxiliary\); LSTM prediction concatenated into the direction head iffλ\>0\\lambda\{\>\}0III\-C\. CTDE context — UAV\-actorguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}/guL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\(xor\) and critics\(t\)s^\{\\,\(t\)\}\(Train\-only\)GlobalguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}Network\-wide UAV context \(default*base*\): shared MLP over mean/std of\[𝐟i;𝐩i\]\[\\mathbf\{f\}\_\{i\};\\mathbf\{p\}\_\{i\}\], contact\-density & degree moments,qp\(t\)q\_\{\\mathrm\{p\}\}^\{\\,\(t\)\},t~\(t\)\\tilde\{t\}^\{\\,\(t\)\};𝐟i=\(bi,βig,βid,v¯i,c¯i\)\\mathbf\{f\}\_\{i\}=\(b\_\{i\},\\beta\_\{i\}^\{\\mathrm\{g\}\},\\beta\_\{i\}^\{\\mathrm\{d\}\},\\bar\{v\}\_\{i\},\\bar\{c\}\_\{i\}\)\. Mutually exclusive withguLg\_\{u\}^\{\\mathrm\{L\}\}\.Local\-ctxguL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}Contact\-limited UAV context \(*base local\-ctx*/ deploy\-near\-real\): mean/std of ground\[𝐟j;𝐩j\]\[\\mathbf\{f\}\_\{j\};\\mathbf\{p\}\_\{j\}\]over\{j:Ciu​j\(t\)=1\}\\\{j:C\_\{i\_\{u\}j\}^\{\\,\(t\)\}\{=\}1\\\}only \(noqp/t~q\_\{\\mathrm\{p\}\}/\\tilde\{t\}\)\.Train\-onlys\(t\)s^\{\\,\(t\)\}Critic input \(distinct fromguG/guLg\_\{u\}^\{\\mathrm\{G\}\}/g\_\{u\}^\{\\mathrm\{L\}\}\): network\-wide statistics akin toguGg\_\{u\}^\{\\mathrm\{G\}\}, plusmeanu​\(𝐳u\)\\mathrm\{mean\}\_\{u\}\(\\mathbf\{z\}\_\{u\}\)ifNuav\>0N\_\{\\mathrm\{uav\}\}\{\>\}0; omitted at deploymentIV\. Actions, team reward, episode flags, and reward weightsα⋅\\alpha\_\{\\cdot\}—𝒜R\\mathcal\{A\}\_\{R\},𝒜M\\mathcal\{A\}\_\{M\},ℱi\(t\)\\mathcal\{F\}\_\{i\}^\{\\,\(t\)\};ai\(t\)a\_\{i\}^\{\\,\(t\)\},mu\(t\)m\_\{u\}^\{\\,\(t\)\}Routing / UAV heading action sets; feasible \(unmasked\) routing subset atii; chosen next\-hop\-or\-idle and heading\-bin actions—oi\(t\)o\_\{i\}^\{\\,\(t\)\},om,u\(t\)o\_\{m,u\}^\{\\,\(t\)\};gu\(t\)∈\{guG,guL\}g\_\{u\}^\{\\,\(t\)\}\\in\\\{g\_\{u\}^\{\\mathrm\{G\}\},g\_\{u\}^\{\\mathrm\{L\}\}\\\};s\(t\)s^\{\\,\(t\)\}See Part III: routing obs\.; UAV\-motion obs\.\(𝐳u,gu\)\(\\mathbf\{z\}\_\{u\},g\_\{u\}\); active UAV context; critic statistics—Ih\(t\)I\_\{h\}^\{\\,\(t\)\},r\(t\)r^\{\\,\(t\)\};D\(t\)D^\{\\,\(t\)\},E\(t\)E^\{\\,\(t\)\},R\(t\)R^\{\\,\(t\)\},B¯\(t\)\\bar\{B\}^\{\\,\(t\)\};gs,gm,geg\_\{s\},g\_\{m\},g\_\{e\};δ\(t\)\\delta^\{\\,\(t\)\}Routing\-activity indicator \(Eq\. \([17](https://arxiv.org/html/2608.04590#S4.E17)\)\); team reward; delivered / expired / dropped counts and mean buffer; stage gates on separation, heading smoothness, and forecast density \(Sec\.[V\-A](https://arxiv.org/html/2608.04590#S5.SS1)\); episode\-end flag—αd,αe,αd​r,αb,αs,αh\\alpha\_\{d\},\\alpha\_\{e\},\\alpha\_\{dr\},\\alpha\_\{b\},\\alpha\_\{s\},\\alpha\_\{h\};αρ,αc,αx,αk,αf\\alpha\_\{\\rho\},\\alpha\_\{c\},\\alpha\_\{x\},\\alpha\_\{k\},\\alpha\_\{f\};αa,αa​d,αa​r\\alpha\_\{a\},\\alpha\_\{ad\},\\alpha\_\{ar\}Delivery/queue weights; stress/UAV\-geometry weights; optional HGA weights onα\(t\)\\alpha^\{\\,\(t\)\},α\(t\)​D\(t\)\\alpha^\{\\,\(t\)\}D^\{\\,\(t\)\},α\(t\)​Ih\(t\)\\alpha^\{\\,\(t\)\}I\_\{h\}^\{\\,\(t\)\}\(default0\)![Refer to caption](https://arxiv.org/html/2608.04590v1/x1.png)

![Refer to caption](https://arxiv.org/html/2608.04590v1/x2.png)

Figure 1:System model: \(a\) spatial layout of ground and UAV nodes; \(b\) per\-step execution pipeline\.### III\-ASystem Model

In Fig\.[1](https://arxiv.org/html/2608.04590#S3.F1), we demonstrate the SCF DTN architecture that integrates road\-constrained ground mobility with controllable UAV relay flight over the plane, wherein heterogeneous nodes execute opportunistic message replication under range\-limited contacts and finite per\-node buffers\. In this architecture, each node maintains a local message queue subject to a maximum buffer capacityBbB\_\{b\}, and the UAV relay fleet furnishes aerial contact opportunities whose spatial distribution is shaped by discrete heading decisions specified in the DTN routing and UAV motion decision model subsection\. We consider the network to consist ofN=Ngr\+NuavN\{=\}N\_\{\\mathrm\{gr\}\}\{\+\}N\_\{\\mathrm\{uav\}\}nodes indexed byi∈\{0,1,…,N−1\}i\\in\\\{0,1,\\ldots,N\{\-\}1\\\}, where ground vehicles occupy indices\{0,…,Ngr−1\}\\\{0,\\ldots,N\_\{\\mathrm\{gr\}\}\{\-\}1\\\}and theuu\-th UAV relay is mapped to indexiu=Ngr\+ui\_\{u\}\{=\}N\_\{\\mathrm\{gr\}\}\{\+\}u\. Communication feasibility at stepttis encoded by the contact matrix𝐂\(t\)\\mathbf\{C\}^\{\\,\(t\)\}, and message replication is permitted only whenCi​j\(t\)=1C\_\{ij\}^\{\\,\(t\)\}\{=\}1\. Table[II](https://arxiv.org/html/2608.04590#S3.T2)summarizes the unified notation for all subsequent sections\.

We model a discrete\-time slot structure in which each episode horizon comprisesTmaxT\_\{\\max\}equal intervals indexed byt∈\{0,1,…,Tmax\}t\\in\\\{0,1,\\ldots,T\_\{\\max\}\\\}\. Fig\.[1](https://arxiv.org/html/2608.04590#S3.F1)\(a\) depicts the spatial domain at steptt, wherein road\-constrained ground vehicles move at speedvgrv\_\{\\mathrm\{gr\}\}within communication rangergrr\_\{\\mathrm\{gr\}\}, freely moving UAV relays move at speedvuavv\_\{\\mathrm\{uav\}\}within communication rangeruavr\_\{\\mathrm\{uav\}\}, and per\-relay congestion stress fields \(𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\},𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\\,\(t\)\}\) which are derived from ground\-node message backlog, influence the moving directions of UAV relays\. In our scenario, ground mobility is modeled as an exogenous process independent of the control policy, whereas UAV headings and opportunistic forwarding constitute the joint controllable decision space optimized by our system\.

1\) Integrated Network and UAV Relay Model:The network comprisesNNheterogeneous nodes partitioned intoNgrN\_\{\\mathrm\{gr\}\}ground vehicles andNuavN\_\{\\mathrm\{uav\}\}UAV relays such thatNgr\+Nuav=NN\_\{\\mathrm\{gr\}\}\+N\_\{\\mathrm\{uav\}\}\{=\}N\. We assume all nodes equipped with onboard Global Navigation Satellite System \(GNSS\) localization and all ground vehicles are restricted to moving along the predefined road network\. Each ground vehicle travels within bounded speed limitsvg​r∈\[vm​i​n,vmax\]v\_\{gr\}\\in\[v\_\{min\},v\_\{\\mathrm\{max\}\}\], and it may either select new cross\-district driving paths from the set of all valid road segments or repeat previously traversed road routes, and its position is updated along the shortest road segments within each geographic region\. For aerial relay nodes, the mobility model differs substantially\. UAV relays move freely on the two\-dimensional plane at constant speedvuavv\_\{\\mathrm\{uav\}\}, and their positions are updated by discrete heading actions defined in the DTN routing and UAV motion decision model subsection\. Consequently, ground vehicles provide routing decisions only, whereas each UAV simultaneously participates in SCF replication and controllable aerial motion\.

2\) Communication Contact Model:Let𝐩i\(t\)∈ℝ2\\mathbf\{p\}\_\{i\}^\{\\,\(t\)\}\\in\\mathbb\{R\}^\{2\}denote the planar coordinate of nodeiiat time steptt\. The Euclidean distance between nodesiiandjjis defined as

di​j\(t\)=‖𝐩i\(t\)−𝐩j\(t\)‖2d\_\{ij\}^\{\\,\(t\)\}=\\bigl\\\|\\mathbf\{p\}\_\{i\}^\{\\,\(t\)\}\-\\mathbf\{p\}\_\{j\}^\{\\,\(t\)\}\\bigr\\\|\_\{2\}\(1\)We employ a type\-heterogeneous communication disk model, wherein the effective communication rangeri​jr\_\{ij\}depends on the types of the two communication endpoints:

ri​j=\{rgr,both​i​and​j​are ground vehiclesruav,at least one endpoint is a UAVr\_\{ij\}=\\begin\{cases\}r\_\{\\mathrm\{gr\}\},&\\text\{both \}i\\text\{ and \}j\\text\{ are ground vehicles\}\\\\\[2\.0pt\] r\_\{\\mathrm\{uav\}\},&\\text\{at least one endpoint is a UAV\}\\end\{cases\}\(2\)At time steptt, a communication link between nodesiiandjjis established if and only ifdi​j\(t\)≤ri​jd\_\{ij\}^\{\\,\(t\)\}\\leq r\_\{ij\}\. Instantaneous pairwise connectivity is characterized by binary contact indicatorsCi​j\(t\)∈\{0,1\}C\_\{ij\}^\{\\,\(t\)\}\\in\\\{0,1\\\}, which constitute the time\-varying contact matrix𝐂\(t\)=\[Ci​j\(t\)\]i,j=0N−1\\mathbf\{C\}^\{\\,\(t\)\}=\\bigl\[C\_\{ij\}^\{\\,\(t\)\}\\bigr\]\_\{i,j=0\}^\{N\-1\}:

Ci​j\(t\)=\{1,i≠j​and​di​j\(t\)≤ri​j0,otherwiseC\_\{ij\}^\{\\,\(t\)\}=\\begin\{cases\}1,&i\\neq j\\text\{ and \}d\_\{ij\}^\{\\,\(t\)\}\\leq r\_\{ij\}\\\\\[2\.0pt\] 0,&\\text\{otherwise\}\\end\{cases\}\(3\)A diagonal constraintCi​i\(t\)=0C\_\{ii\}^\{\\,\(t\)\}\{=\}0is imposed for alliito eliminate self\-loop contacts\. The raw contact degree of nodeiicounts the number of reachable neighbors at steptt,

νi\(t\)=∑j=0N−1Ci​j\(t\)\\nu\_\{i\}^\{\\,\(t\)\}=\\sum\_\{j=0\}^\{N\-1\}C\_\{ij\}^\{\\,\(t\)\}\(4\)and the normalized contact degree

ν¯i\(t\)=νi\(t\)N−1∈\[0,1\]\\bar\{\\nu\}\_\{i\}^\{\\,\(t\)\}=\\frac\{\\nu\_\{i\}^\{\\,\(t\)\}\}\{N\-1\}\\in\[0,1\]\(5\)measures neighbor density within the unit interval\.

3\) Traffic Injection and Message State Model:Messages are injected by an exogenous traffic process that is independent of per\-step routing and UAV actions\. The concrete injection schedules used in evaluation are specified in Sec\.[VI](https://arxiv.org/html/2608.04590#S6)\. Each message records its source, destination, normalized payload size, hop count, and remaining TTL\.

4\) Buffer Queueing and Forwarding Constraint Model:Due to the limited buffer space at each node, unrestrained replication may induce buffer saturation and delivery failure\. Therefore, optimal forwarding must incorporate real\-time buffer occupancy during packet replication operations\. The buffer utilization of nodeiiat time stepttis defined as

bi\(t\)=\|ℬi\(t\)\|Bb∈\[0,1\]b\_\{i\}^\{\\,\(t\)\}=\\frac\{\\bigl\|\\mathcal\{B\}\_\{i\}^\{\\,\(t\)\}\\bigr\|\}\{B\_\{b\}\}\\in\[0,1\]\(6\)whereℬi\(t\)\\mathcal\{B\}\_\{i\}^\{\\,\(t\)\}denotes the set of messages buffered at nodeii\. All forwarding and buffer update operations described in the DTN routing and UAV motion decision model comply with the following constraints:

1. 1\.Contact Feasibility: A message transfer from nodeiito nodejjat time stepttis permitted only ifCi​j\(t\)=1C\_\{ij\}^\{\\,\(t\)\}\{=\}1\.
2. 2\.Single Reception Constraint: Each node can receive at most one replicated message copy per time step, reflecting a per\-step receiver capacity in the discrete\-time formulation\.
3. 3\.Buffer Admission Rule: A replicated message is accepted only if the receiver possesses available buffer space\. Otherwise, the oldest message will be dropped\.
4. 4\.Replication and Delivery Logic: A valid routing action replicates the selected message to an eligible neighboring node\. Delivery is recorded once the destination is reached, and the hop counter is incremented on each forwarding operation\.
5. 5\.TTL Decay Mechanism: After all forwarding operations complete at time steptt, the TTL of every buffered message is decremented and expired messages are purged\.

### III\-BDTN Routing and UAV Motion Decision Model

In the DTN routing and UAV motion decision model, we adopt the previously constructed mobility and contact model and formulate joint per\-node forwarding together with discrete UAV heading control as the integrated decision space, enabling the policy to adjust future contact opportunities\. Under intermittent connectivity, each node must select at most one replication action from a capped candidate set constructed from its local buffer and instantaneous contacts, denoted by\{j:Ci​j\(t\)=1\}\\\{j:C\_\{ij\}^\{\\,\(t\)\}\{=\}1\\\}, while each UAV relay additionally selects a heading bin that updates its planar position and thereby modifies𝐂\(t\+1\)\\mathbf\{C\}^\{\\,\(t\+1\)\}\.

LetKKdenote the maximum number of eligible forwarding candidates maintained per node\. Each candidate is a*\(message, next\-hop neighbor\)*pair ranked according to three priority criteria:\(a\)one\-hop direct delivery when the neighbor is identical to the message destination;\(b\)shorter Euclidean distance from the neighbor to the target destination for non\-direct transfers;\(c\)smaller remaining TTL as tiebreaker\. Only the top\-KKranked candidates are retained after truncation\.

Each nodei∈\{0,…,N−1\}i\\in\\\{0,\\ldots,N\{\-\}1\\\}selects a routing action from the discrete action space

ai\(t\)∈𝒜R=\{0,1,…,K\}a\_\{i\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{R\}=\\\{0,1,\\ldots,K\\\}\(7\)where index0denotes the idle operation and indices11toKKcorrespond to prioritized forwarding candidates\. Ground vehicle positions evolve exogenously along predefined road trajectories without controllable motion inputs\.

UAV relaysu∈\{0,…,Nuav−1\}u\\in\\\{0,\\ldots,N\_\{\\mathrm\{uav\}\}\{\-\}1\\\}are appended after theNgrN\_\{\\mathrm\{gr\}\}ground nodes\. In addition to routing decisionaiu\(t\)∈𝒜Ra\_\{i\_\{u\}\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{R\}, each relay selects a discrete heading action

mu\(t\)∈𝒜M=\{0,1,…,B−1\},m\_\{u\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{M\}=\\\{0,1,\\ldots,B\-1\\\},\(8\)with fixed bin numberB=8B\{=\}8, where each binkkcorresponds to a unit direction vector𝐝k\\mathbf\{d\}\_\{k\}aligned with one of eight compass orientations\. The joint control tuple\(aiu\(t\),mu\(t\)\)∈𝒜R×𝒜M\(a\_\{i\_\{u\}\}^\{\\,\(t\)\},\\,m\_\{u\}^\{\\,\(t\)\}\)\\in\\mathcal\{A\}\_\{R\}\\times\\mathcal\{A\}\_\{M\}is generated independently at each step, and the UAV position update follows

𝐩iu\(t\+1\)=𝐩iu\(t\)\+vuav​𝐝mu\(t\)\\mathbf\{p\}\_\{i\_\{u\}\}^\{\\,\(t\+1\)\}=\\mathbf\{p\}\_\{i\_\{u\}\}^\{\\,\(t\)\}\+v\_\{\\mathrm\{uav\}\}\\,\\mathbf\{d\}\_\{m\_\{u\}^\{\\,\(t\)\}\}\(9\)
By combining per\-node message replication with UAV heading assignment, we derive the step\-wise execution pipeline visualized in Fig\.[1](https://arxiv.org/html/2608.04590#S3.F1)\(b\), which follows three sequential stages:\(a\)Execute all UAV heading actionsmu\(t\)m\_\{u\}^\{\(t\)\};\(b\)Carry out per\-node routing actionsai\(t\)a\_\{i\}^\{\(t\)\}under the contact matrix𝐂\(t\)\\mathbf\{C\}^\{\(t\)\}and aforementioned buffer constraints;\(c\)Refresh ground and UAV positions according to Eq\. \([9](https://arxiv.org/html/2608.04590#S3.E9)\), decrement message TTL values, generate exogenous packet arrivals where applicable, and recalculate the contact matrix𝐂\(t\+1\)\\mathbf\{C\}^\{\(t\+1\)\}alongside the set of eligible forwarding candidates\.

### III\-CStress Fields and Navigation Vectors

we construct a hierarchical set of congestion descriptors for decentralized UAV heading control, including scalar delivery stress on ground nodes, derived stress fields around each relay \(isotropic stress density and directional sector stress\), fleet geometry signals, and finally navigation and supervision vectors of unified dimensionDyD\_\{y\}for policy input\. Within each step, delivery stress is defined strictly on ground nodesj<Ngrj<N\_\{\\mathrm\{gr\}\}, and only connected pairs satisfyingCiu​j\(t\)=1C\_\{i\_\{u\}j\}^\{\\,\(t\)\}\{=\}1contribute to relay\-local stress fields\.

1\) UAV Fleet Geometry and Heading Stability:Under decentralized heading control, independent UAV motion decisions may induce relay clustering or frequent heading switches\. We define the pairwise normalized separation between distinct relaysuuandvvas

ζu​v\(t\)=min⁡\(1,‖𝐩iu\(t\)−𝐩iv\(t\)‖2ruav\),u≠v,\\zeta\_\{uv\}^\{\\,\(t\)\}=\\min\\left\(1,\\,\\frac\{\\left\\\|\\mathbf\{p\}\_\{i\_\{u\}\}^\{\\,\(t\)\}\-\\mathbf\{p\}\_\{i\_\{v\}\}^\{\\,\(t\)\}\\right\\\|\_\{2\}\}\{r\_\{\\mathrm\{uav\}\}\}\\right\),\\quad u\\neq v,whereruavr\_\{\\mathrm\{uav\}\}denotes the homogeneous UAV communication range in Eq\. \([2](https://arxiv.org/html/2608.04590#S3.E2)\)\. The fleet\-averaged separation is

ξ\(t\)=meanu≠v⁡ζu​v\(t\)\\xi^\{\\,\(t\)\}=\\operatorname\{mean\}\_\{u\\neq v\}\\zeta\_\{uv\}^\{\\,\(t\)\}\(10\)where the mean is taken over all distinct relay pairs withu,v∈\{0,…,Nuav−1\}u,v\\in\\\{0,\\ldots,N\_\{\\mathrm\{uav\}\}\{\-\}1\\\}andu≠vu\\neq v\. We further define the heading\-switch ratioκ\(t\)∈\[0,1\]\\kappa^\{\\,\(t\)\}\\in\[0,1\]as the fraction of UAVs satisfyingmu\(t\)≠mu\(t−1\)m\_\{u\}^\{\\,\(t\)\}\\neq m\_\{u\}^\{\\,\(t\-1\)\}, withκ\(0\)=0\\kappa^\{\\,\(0\)\}\{=\}0, which quantifies temporal consistency of fleet heading selections\.

2\) Ground\-Node Delivery Stress:Each ground vehiclejjis assigned a scalar delivery stress to measure message backlog urgency:

σj\(t\)=bj\(t\)​\(1\+υ¯j\(t\)\)\\sigma\_\{j\}^\{\\,\(t\)\}=b\_\{j\}^\{\\,\(t\)\}\\left\(1\+\\bar\{\\upsilon\}\_\{j\}^\{\\,\(t\)\}\\right\)\(11\)wherebj\(t\)b\_\{j\}^\{\\,\(t\)\}denotes buffer utilization in Eq\. \([6](https://arxiv.org/html/2608.04590#S3.E6)\) andυ¯j\(t\)\\bar\{\\upsilon\}\_\{j\}^\{\\,\(t\)\}denotes the average TTL urgency factor

υ¯j\(t\)=1\|ℳj\|​∑q∈ℳjmax⁡\(0,1−τqT0\)\\bar\{\\upsilon\}\_\{j\}^\{\\,\(t\)\}=\\frac\{1\}\{\|\\mathcal\{M\}\_\{j\}\|\}\\sum\_\{q\\in\\mathcal\{M\}\_\{j\}\}\\max\\left\(0,\\,1\-\\frac\{\\tau\_\{q\}\}\{T\_\{0\}\}\\right\)\(12\)Here,ℳj\\mathcal\{M\}\_\{j\}denotes undelivered messages at ground nodejj,τq\\tau\_\{q\}is the remaining TTL of packetqq, andT0T\_\{0\}is the global maximum TTL reference\. We setσiu\(t\)=0\\sigma\_\{i\_\{u\}\}^\{\\,\(t\)\}\{=\}0for all UAV nodes, which mean Uavs do not participate in the calculation of delivery stress\.

3\) Isotropic and Directional stress Fields:For relayuumapped toiui\_\{u\}, only contact\-valid ground nodes contribute to the stress fields\. The per\-relay and fleet\-mean isotropic stress densities are

ρiu\(t\)\\displaystyle\\rho\_\{i\_\{u\}\}^\{\\,\(t\)\}=min⁡\(1,12​Ngr​∑j=0Ngr−1Ciu​j\(t\)​σj\(t\)\),\\displaystyle=\\min\\left\(1,\\,\\frac\{1\}\{2N\_\{\\mathrm\{gr\}\}\}\\sum\_\{j=0\}^\{N\_\{\\mathrm\{gr\}\}\-1\}C\_\{i\_\{u\}j\}^\{\\,\(t\)\}\\sigma\_\{j\}^\{\\,\(t\)\}\\right\),ρ\(t\)\\displaystyle\\rho^\{\\,\(t\)\}=1Nuav​∑u=0Nuav−1ρiu\(t\)\.\\displaystyle=\\frac\{1\}\{N\_\{\\mathrm\{uav\}\}\}\\sum\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\rho\_\{i\_\{u\}\}^\{\\,\(t\)\}\.\(13\)The fleet\-average termρ\(t\)\\rho^\{\\,\(t\)\}enters the global team reward via coefficientαρ\\alpha\_\{\\rho\}\(Sec\.[V\-A](https://arxiv.org/html/2608.04590#S5.SS1)\)\.

To obtain directional sector stress, delivery\-stress values are aggregated into eight compass heading bins\. Let𝐫^iu​j\(t\)\\hat\{\\mathbf\{r\}\}\_\{i\_\{u\}j\}^\{\\,\(t\)\}denote the unit vector from relayiui\_\{u\}to ground nodejj,

𝐫^iu​j\(t\)=𝐩j\(t\)−𝐩iu\(t\)‖𝐩j\(t\)−𝐩iu\(t\)‖2,\\hat\{\\mathbf\{r\}\}\_\{i\_\{u\}j\}^\{\\,\(t\)\}=\\frac\{\\mathbf\{p\}\_\{j\}^\{\\,\(t\)\}\-\\mathbf\{p\}\_\{i\_\{u\}\}^\{\\,\(t\)\}\}\{\\left\\\|\\mathbf\{p\}\_\{j\}^\{\\,\(t\)\}\-\\mathbf\{p\}\_\{i\_\{u\}\}^\{\\,\(t\)\}\\right\\\|\_\{2\}\},and letkiu​j∗\(t\)=arg⁡maxk′⁡\(𝐫^iu​j\(t⊤\)​𝐝k′\)\{k\_\{i\_\{u\}j\}^\{\*\}\}^\{\(t\)\}=\\arg\\max\_\{k^\{\\prime\}\}\\bigl\(\\hat\{\\mathbf\{r\}\}\_\{i\_\{u\}j\}^\{\\,\(t\\top\)\}\\mathbf\{d\}\_\{k^\{\\prime\}\}\\bigr\)denote the heading bin best aligned with nodejj\. The sector\-stress component for binkkof relayuuis

Su,k\(t\)=min⁡\(1,12​Ngr​∑j=0Ngr−1Ciu​j\(t\)​σj\(t\)⋅𝕀​\[k=kiu​j∗\(t\)\]\),S\_\{u,k\}^\{\\,\(t\)\}=\\min\\left\(1,\\,\\frac\{1\}\{2N\_\{\\mathrm\{gr\}\}\}\\sum\_\{j=0\}^\{N\_\{\\mathrm\{gr\}\}\-1\}C\_\{i\_\{u\}j\}^\{\\,\(t\)\}\\sigma\_\{j\}^\{\\,\(t\)\}\\cdot\\mathbb\{I\}\\left\[k=\{k\_\{i\_\{u\}j\}^\{\*\}\}^\{\(t\)\}\\right\]\\right\),\(14\)where𝕀​\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\. ScalarsSu,k\(t\)S\_\{u,k\}^\{\\,\(t\)\}are stacked into the sector\-stress vector𝐒u\(t\)=\[Su,0\(t\),…,Su,7\(t\)\]\\mathbf\{S\}\_\{u\}^\{\\,\(t\)\}=\\bigl\[S\_\{u,0\}^\{\\,\(t\)\},\\ldots,S\_\{u,7\}^\{\\,\(t\)\}\\bigr\]around relayuu\.

4\) Optional Exponential Moving Average \(EMA\) Stress Forecast:The environment optionally applies EMA and trend extrapolation over\{σj\(t\)\}\\\{\\sigma\_\{j\}^\{\\,\(t\)\}\\\}to produceσ~j\(t\)\\tilde\{\\sigma\}\_\{j\}^\{\\,\(t\)\}\. Substitutingσ~j\(t\)\\tilde\{\\sigma\}\_\{j\}^\{\\,\(t\)\}forσj\(t\)\\sigma\_\{j\}^\{\\,\(t\)\}in Eq\. \([13](https://arxiv.org/html/2608.04590#S3.E13)\) yields optional forecast stress densities

ρf,iu\(t\)\\displaystyle\\rho\_\{f,i\_\{u\}\}^\{\\,\(t\)\}=min⁡\(1,12​Ngr​∑j=0Ngr−1Ciu​j\(t\)​σ~j\(t\)\),\\displaystyle=\\min\\\!\\left\(1,\\ \\frac\{1\}\{2N\_\{\\mathrm\{gr\}\}\}\\sum\_\{j=0\}^\{N\_\{\\mathrm\{gr\}\}\-1\}C\_\{i\_\{u\}j\}^\{\\,\(t\)\}\\,\\tilde\{\\sigma\}\_\{j\}^\{\\,\(t\)\}\\right\),ρf\(t\)\\displaystyle\\rho\_\{f\}^\{\\,\(t\)\}=1Nuav​∑u=0Nuav−1ρf,iu\(t\)\.\\displaystyle=\\frac\{1\}\{N\_\{\\mathrm\{uav\}\}\}\\sum\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\rho\_\{f,i\_\{u\}\}^\{\\,\(t\)\}\.\(15\)
5\) UAV Navigation and Supervision Vectors:Based on the sector stress𝐒u\(t\)\\mathbf\{S\}\_\{u\}^\{\(t\)\}defined in Eq\. \([14](https://arxiv.org/html/2608.04590#S3.E14)\), we construct two vectors of identical length for each relay, both mapped to the unified dimensionDyD\_\{y\}\. The navigation vector𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\(t\)\}acts as the input to the UAV heading policy, while the supervision vector𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\(t\)\}supplies the ground\-truth hotspot layout for optional auxiliary learning\. The navigation vector aggregates the following components:\(i\)the eight\-dimensional sector stress𝐒u\(t\)\\mathbf\{S\}\_\{u\}^\{\(t\)\};\(ii\)the stacked offset block𝐞u\(t\)∈ℝ2​Kσ\\mathbf\{e\}\_\{u\}^\{\(t\)\}\\in\\mathbb\{R\}^\{2K\_\{\\sigma\}\}containing normalized\(Δ​x,Δ​y\)\(\\Delta x,\\Delta y\)displacements from the relay to theKσK\_\{\\sigma\}ground nodes with the highest delivery stress;\(iii\)the unit centroid direction𝐧^u\(t\)∈ℝ2\\hat\{\\mathbf\{n\}\}\_\{u\}^\{\(t\)\}\\in\\mathbb\{R\}^\{2\}pointing toward the urgent centroid of in\-range ground nodes\.

Accordingly,

𝐳u\(t\)=\[𝐒u\(t\);𝐞u\(t\);𝐧^u\(t\)\],Dy=8\+2​Kσ\+2,\\mathbf\{z\}\_\{u\}^\{\(t\)\}=\\bigl\[\\mathbf\{S\}\_\{u\}^\{\(t\)\};\\,\\mathbf\{e\}\_\{u\}^\{\(t\)\};\\,\\hat\{\\mathbf\{n\}\}\_\{u\}^\{\(t\)\}\\bigr\],\\qquad D\_\{y\}=8\+2K\_\{\\sigma\}\+2,\(16\)and𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\(t\)\}adopts the same dimensional structure with ground\-truth sector, offset, and centroid components\.

## IVProblem Formulation

In this section, we outline the optimization goals of the joint UAV–routing network\. Then we cast SCF control as a finite\-horizon cooperative sequential decision process with factored policies, and we derive a per\-step team reward that converts the episodic objective into a form amenable to centralized training and decentralized execution learning under partial observability\.

### IV\-AObjective Equation

Section[III](https://arxiv.org/html/2608.04590#S3)formalizes how the DTN evolves under range\-limited contacts, finite buffers, and joint routing–UAV control\. This subsection states multiple objectives the cooperative control should optimize\.

We aim to maximize the successful message delivery of SCF messages while limiting TTL expiry, buffer\-induced loss, and uncontrolled UAV congestion, and furthermore to shape UAV headings and fleet positioning so that the relay fleet remains dispersed and headings remain stable\. At each steptt, we defineD\(t\)D^\{\\,\(t\)\},E\(t\)E^\{\\,\(t\)\}andR\(t\)R^\{\\,\(t\)\}as the respective counts of delivered, expired and dropped messages,B¯\(t\)\\bar\{B\}^\{\\,\(t\)\}as the average node buffer utilization, andIh\(t\)∈\{0,1\}I\_\{h\}^\{\\,\(t\)\}\\in\\\{0,1\\\}as the routing\-activity indicator:

Ih\(t\)=\{1,if at least one replication succeeds in step​t0,otherwise;I\_\{h\}^\{\\,\(t\)\}=\\begin\{cases\}1,&\\text\{if at least one replication succeeds in step \}t\\\\ 0,&\\text\{otherwise\};\\end\{cases\}\(17\)The variablesρ\(t\)\\rho^\{\\,\(t\)\},ξ\(t\)\\xi^\{\\,\(t\)\},κ\(t\)\\kappa^\{\\,\(t\)\}andρf\(t\)\\rho\_\{f\}^\{\\,\(t\)\}denote fleet stress\-density and geometry signals\. Average buffer utilizationB¯\(t\)\\bar\{B\}^\{\\,\(t\)\}and ground\-node delivery stressσj\(t\)\\sigma\_\{j\}^\{\\,\(t\)\}\(Eqs\. \([6](https://arxiv.org/html/2608.04590#S3.E6)\), \([11](https://arxiv.org/html/2608.04590#S3.E11)\)\) quantify congestion hotspots to be mitigated via routing and UAV positioning\. The stress fieldsρ\(t\)\\rho^\{\\,\(t\)\}and sector stress𝐒u\(t\)\\mathbf\{S\}\_\{u\}^\{\\,\(t\)\}, together with the supervision vector𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\\,\(t\)\}, render such pressure observable to UAV agents and to the per\-step surrogate objective derived below\. State\-dependent gating functionsgs​\(D\(t\)\),gm​\(D\(t\)\),ge​\(E\(t\)\)∈\[0,1\]g\_\{s\}\(D^\{\\,\(t\)\}\),g\_\{m\}\(D^\{\\,\(t\)\}\),g\_\{e\}\(E^\{\\,\(t\)\}\)\\in\[0,1\]modulate UAV geometry and forecast\-density shaping without altering the primary delivery terms:gsg\_\{s\}andgmg\_\{m\}reduce separation and heading\-smoothness incentives when step deliveries occur, whereasgeg\_\{e\}strengthens forecast\-density shaping when expirations rise\.

Based on this priority ordering, we formulate the principal delivery objective as

Jdel​\(π\)\\displaystyle J\_\{\\mathrm\{del\}\}\(\\pi\)=𝔼π\[∑t=0Tmax\(αdD\(t\)\+αeE\(t\)\+αd​rR\(t\)\\displaystyle=\\mathbb\{E\}\_\{\\pi\}\\\!\\Biggl\[\\sum\_\{t=0\}^\{T\_\{\\max\}\}\\bigl\(\\alpha\_\{d\}D^\{\\,\(t\)\}\+\\alpha\_\{e\}E^\{\\,\(t\)\}\+\\alpha\_\{dr\}R^\{\\,\(t\)\}\(18\)\+αbB¯\(t\)\+αs\+αhIh\(t\)\)\],\\displaystyle\\qquad\\qquad\+\\alpha\_\{b\}\\bar\{B\}^\{\\,\(t\)\}\+\\alpha\_\{s\}\+\\alpha\_\{h\}I\_\{h\}^\{\\,\(t\)\}\\bigr\)\\Biggr\],whereα⋅\\alpha\_\{\\cdot\}are scalar weights withαe,αd​r,αb,αs,αh≤0\\alpha\_\{e\},\\alpha\_\{dr\},\\alpha\_\{b\},\\alpha\_\{s\},\\alpha\_\{h\}\\leq 0for penalty terms, thereby rewarding successful delivery while penalizing TTL expiry, buffer drops, mean congestion, idle step, and unnecessary replication\. We further formulate the UAV placement objective as

Juav​\(π\)\\displaystyle J\_\{\\mathrm\{uav\}\}\(\\pi\)=𝔼π\[∑t=0Tmax\(αρρ\(t\)\+αcρ\(t\)D\(t\)\+αxgs\(D\(t\)\)ξ\(t\)\\displaystyle=\\mathbb\{E\}\_\{\\pi\}\\\!\\Biggl\[\\sum\_\{t=0\}^\{T\_\{\\max\}\}\\Bigl\(\\alpha\_\{\\rho\}\\,\\rho^\{\\,\(t\)\}\+\\alpha\_\{c\}\\,\\rho^\{\\,\(t\)\}D^\{\\,\(t\)\}\+\\alpha\_\{x\}\\,g\_\{s\}\\\!\\bigl\(D^\{\\,\(t\)\}\\bigr\)\\,\\xi^\{\\,\(t\)\}\(19\)−αk​gm​\(D\(t\)\)​κ\(t\)\+αf​ge​\(E\(t\)\)​ρf\(t\)\\displaystyle\\qquad\\qquad\-\\alpha\_\{k\}\\,g\_\{m\}\\\!\\bigl\(D^\{\\,\(t\)\}\\bigr\)\\,\\kappa^\{\\,\(t\)\}\+\\alpha\_\{f\}\\,g\_\{e\}\\\!\\bigl\(E^\{\\,\(t\)\}\\bigr\)\\,\\rho\_\{f\}^\{\\,\(t\)\}\+αaα\(t\)\+αa​dα\(t\)D\(t\)\+αa​rα\(t\)Ih\(t\)\)\],\\displaystyle\\qquad\\qquad\+\\alpha\_\{a\}\\,\\alpha^\{\\,\(t\)\}\+\\alpha\_\{ad\}\\,\\alpha^\{\\,\(t\)\}D^\{\\,\(t\)\}\+\\alpha\_\{ar\}\\,\\alpha^\{\\,\(t\)\}I\_\{h\}^\{\\,\(t\)\}\\Bigr\)\\Biggr\],Hereαρ\\alpha\_\{\\rho\}rewards in\-range stress density,αc\\alpha\_\{c\}couples that density to step deliveries,αx\\alpha\_\{x\}encourages pairwise separation,αk\\alpha\_\{k\}penalizes heading switches, andαf\\alpha\_\{f\}shapes placement by forecast density under expiry\. The remaining coefficientsαa,αa​d,αa​r\\alpha\_\{a\},\\alpha\_\{ad\},\\alpha\_\{ar\}optionally activate hotspot\-guided alignment \(HGA\):α\(t\)∈\[0,1\]\\alpha^\{\\,\(t\)\}\\in\[0,1\]measures how well executed UAV headings match the directional congestion layout induced by sector stress \(Sec\.[III\-C](https://arxiv.org/html/2608.04590#S3.SS3)\), and the productsα\(t\)​D\(t\)\\alpha^\{\\,\(t\)\}D^\{\\,\(t\)\}andα\(t\)​Ih\(t\)\\alpha^\{\\,\(t\)\}I\_\{h\}^\{\\,\(t\)\}couple that geometric agreement to instantaneous delivery and routing activity\. By defaultαa=αa​d=αa​r=0\\alpha\_\{a\}\{=\}\\alpha\_\{ad\}\{=\}\\alpha\_\{ar\}\{=\}0\(HGA off\); when enabled, these terms provide an auxiliary geometric prior that steers relays toward congested sectors without replacing the primary delivery objective\. The full construction ofα\(t\)\\alpha^\{\\,\(t\)\}is deferred to Sec\.[V\-A](https://arxiv.org/html/2608.04590#S5.SS1)\. Although merely optimizing the delivery objectiveJdel​\(π\)J\_\{\\mathrm\{del\}\}\(\\pi\)suffers from sparse, delayed step rewards caused by the SCF transmission paradigm, optimizing only the UAV placement objectiveJuav​\(π\)J\_\{\\mathrm\{uav\}\}\(\\pi\)will drive relays to gather around congestion hotspots without tangible delivery gains\. To avoid the drawbacks of single\-objective optimization, we integrate the two metrics into a unified cooperative control objective as

\(P1\)maxπ⁡J​\(π\)=Jdel​\(π\)\+Juav​\(π\)\\text\{\(P1\)\}\\quad\\max\_\{\\pi\}\\;J\(\\pi\)=J\_\{\\mathrm\{del\}\}\(\\pi\)\+J\_\{\\mathrm\{uav\}\}\(\\pi\)\(20\)subject to the SCF and action rules of Secs\.[III](https://arxiv.org/html/2608.04590#S3)where Ground mobility is exogenous, UAV motion and all routing decisions are policy\-controlled throughπθ\\pi\_\{\\theta\}\.

### IV\-BJoint Decision Process and Factored Policy

We recast the system model of Sec\.[III](https://arxiv.org/html/2608.04590#S3)as a cooperative sequential decision process under the fixed per\-step pipeline\. According to the control pipeline, each step applies UAV heading assignment, opportunistic replication, and network\-state refresh, returns decentralized observations𝐨\(t\)\\mathbf\{o\}^\{\\,\(t\)\}, and terminates episodes att=Tmaxt\{=\}T\_\{\\max\}\. Letx\(t\)x^\{\\,\(t\)\}denote the full environment state: GNSS\-localized positions\{𝐩i\(t\)\}\\\{\\mathbf\{p\}\_\{i\}^\{\\,\(t\)\}\\\}, buffers\{ℬi\(t\)\}\\\{\\mathcal\{B\}\_\{i\}^\{\\,\(t\)\}\\\}, contact matrix𝐂\(t\)\\mathbf\{C\}^\{\\,\(t\)\}, and the stress fields of Sec\.[III\-C](https://arxiv.org/html/2608.04590#S3.SS3)\. The joint action is

𝐚\(t\)=\(a0\(t\),…,aN−1\(t\),m0\(t\),…,mNuav−1\(t\)\),\\mathbf\{a\}^\{\\,\(t\)\}=\\bigl\(a\_\{0\}^\{\\,\(t\)\},\\ldots,a\_\{N\-1\}^\{\\,\(t\)\},\\,m\_\{0\}^\{\\,\(t\)\},\\ldots,m\_\{N\_\{\\mathrm\{uav\}\}\-1\}^\{\\,\(t\)\}\\bigr\),\(21\)whereai\(t\)∈𝒜Ra\_\{i\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{R\}selects a routing candidate or idle action andmu\(t\)∈𝒜Mm\_\{u\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{M\}selects a UAV heading bin \(Eq\. \([9](https://arxiv.org/html/2608.04590#S3.E9)\)\)\. The per\-step transition follows

x\(t\+1\)=Φ​\(Ψ​\(x\(t\),\{mu\(t\)\}u=0Nuav−1\),\{ai\(t\)\}i=0N−1\),x^\{\\,\(t\+1\)\}=\\Phi\\\!\\Big\(\\Psi\\big\(x^\{\\,\(t\)\},\\,\\\{m\_\{u\}^\{\\,\(t\)\}\\\}\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\big\),\\,\\\{a\_\{i\}^\{\\,\(t\)\}\\\}\_\{i=0\}^\{N\-1\}\\Big\),\(22\)whereΨ\\Psiapplies UAV motion before routing transfers, andΦ\\Phicompletes replication, TTL decay, ground mobility, and contact recomputation\. Two bidirectional structural couplings exist between UAV mobility and routing subsystems, as detailed below\.

\(C1\) Topology\-to\-routing couplingFor any nodeii, the set of reachable forwarding neighbors is defined by contact indicators\{j:Ci​j\(t\)=1\}\\\{j:C\_\{ij\}^\{\(t\)\}=1\\\}\. Contact availability relies on pairwise Euclidean distancesdi​j\(t\)d\_\{ij\}^\{\(t\)\}, which are directly controlled by UAV positions\. When a relayuuadjusts its headingmu\(t\)m\_\{u\}^\{\(t\)\}to move toward congested ground zones, it can establish new communication links that were unavailable in the previous time step\. This creates extra forwarding opportunities that routing strategies cannot access under a static, fixed network topology\.

\(C2\) Routing\-to\-stress\-field feedback couplingMessage replication and successful deliveries modify node buffer occupancy and the spatial distribution of pending packets\. These buffer updates feed into ground\-node delivery stressσj\(t\)\\sigma\_\{j\}^\{\(t\)\}, as well as fleet\-wide stress\-density and geometry signals\(ρ\(t\),ξ\(t\),κ\(t\)\)\(\\rho^\{\(t\)\},\\xi^\{\(t\)\},\\kappa^\{\(t\)\}\)embedded in the joint optimization objective \(P1\)\. Routing operations reshape congestion hotspots encoded by the stress fields, which subsequently guide UAV position adjustments\. Meanwhile, the UAV movement decisions taken in prior time steps determine which node pairs can exchange messages in the current step\.

The state transition operatorsΨ\\Psi\(UAV mobility update\) andΦ\\Phi\(routing & network refresh\) follow a rigid execution order: UAV positions are updated before any message replication within one time step\. The two operators cannot be swapped, which introduces sequential cross\-coupling between subsystems\. As a result, we cannot optimize UAV mobility and routing via separate, decoupled iterative loops\.

Factored Policy and Action Spaces\.Instead of learning a single joint policy for global network flow routing, we factor the control task intoN\+NuavN\+N\_\{\\mathrm\{uav\}\}cooperative decision units:NNrouting units \(one per nodeii\) andNuavN\_\{\\mathrm\{uav\}\}motion units \(one per UAVuu\)\. Each UAV participates twice in Eq\. \([21](https://arxiv.org/html/2608.04590#S4.E21)\): as routing unit at nodeiui\_\{u\}throughaiu\(t\)a\_\{i\_\{u\}\}^\{\\,\(t\)\}, and as motion unituuthroughmu\(t\)m\_\{u\}^\{\\,\(t\)\}, with distinct observationsoiu\(t\)o\_\{i\_\{u\}\}^\{\\,\(t\)\}versusom,u\(t\)o\_\{m,u\}^\{\\,\(t\)\}but shared parametersθ\\theta\. A relay must both reach contacts through motion and use them for SCF replication\. Each node selects at most one local transfer from a capped candidate set \(\|𝒜R\|=K\+1\|\\mathcal\{A\}\_\{R\}\|\{=\}K\{\+\}1\), and each UAV selects one ofB=8B\{=\}8headings\. Letℱi\(t\)⊆𝒜R\\mathcal\{F\}\_\{i\}^\{\\,\(t\)\}\\subseteq\\mathcal\{A\}\_\{R\}denote the masked feasible set at nodeii, the per\-step factored action space is

\|𝒜joint\(t\)\|=∏i=0N−1\|ℱi\(t\)\|×BNuav,\|\\mathcal\{A\}\_\{\\mathrm\{joint\}\}^\{\\,\(t\)\}\|\\;=\\;\\prod\_\{i=0\}^\{N\-1\}\|\\mathcal\{F\}\_\{i\}^\{\\,\(t\)\}\|\\;\\times\\;B^\{N\_\{\\mathrm\{uav\}\}\},\(23\)where\|ℱi\(t\)\|\|\\mathcal\{F\}\_\{i\}^\{\\,\(t\)\}\|depends on𝐂\(t\)\\mathbf\{C\}^\{\\,\(t\)\}and therefore depends on prior motion, while each factor conditions only on local observations\. Given decentralized observations𝐨\(t\)\\mathbf\{o\}^\{\\,\(t\)\}, nodeiichoosesai\(t\)∈𝒜Ra\_\{i\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{R\}fromπθ,i\(⋅∣oi\(t\)\)\\pi\_\{\\theta,i\}\(\\cdot\\mid o\_\{i\}^\{\\,\(t\)\}\), and relayuuchoosesmu\(t\)∈𝒜Mm\_\{u\}^\{\\,\(t\)\}\\in\\mathcal\{A\}\_\{M\}fromπθ,m,u\(⋅∣om,u\(t\)\)\\pi\_\{\\theta,m,u\}\(\\cdot\\mid o\_\{m,u\}^\{\\,\(t\)\}\)\. The joint policy factorizes as

πθ​\(𝐚\(t\)∣𝐨\(t\)\)\\displaystyle\\pi\_\{\\theta\}\(\\mathbf\{a\}^\{\\,\(t\)\}\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)=∏i=0N−1πθ,i​\(ai\(t\)∣oi\(t\)\)\\displaystyle=\\prod\_\{i=0\}^\{N\-1\}\\pi\_\{\\theta,i\}\\\!\\left\(a\_\{i\}^\{\\,\(t\)\}\\mid o\_\{i\}^\{\\,\(t\)\}\\right\)×∏u=0Nuav−1πθ,m,u\(mu\(t\)∣om,u\(t\)\)\.\\displaystyle\\quad\\times\\prod\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\pi\_\{\\theta,m,u\}\\\!\\left\(m\_\{u\}^\{\\,\(t\)\}\\mid o\_\{m,u\}^\{\\,\(t\)\}\\right\)\.\(24\)All routing and mobility policy heads share the same parameter weightsθ\\theta\. Although the factorized policy achieves conditional independence under decentralized local observations, strong bidirectional coupling still exists within the environment state transition law formulated in Eq\. \([22](https://arxiv.org/html/2608.04590#S4.E22)\)\.

### IV\-CObservation Model and Per\-Step Decomposition

Direct policy search over objective \(P1\) is computationally intractable\. We adopt a decentralized observation paradigm under the CTDE framework and perform an exact stepwise decomposition of the \(P1\)\. The decomposed decision agents form a fully cooperative, partially observable multi\-agent team under the CTDE paradigm introduced in Section[II](https://arxiv.org/html/2608.04590#S2)\. Both the distributed actors and the training\-only critic module take input from the standardized set of observation features defined in Part III of Table[II](https://arxiv.org/html/2608.04590#S3.T2)\. The partial observability limitation arises from three practical communication constraints:\(a\)The buffer occupancy and neighbor count of remote nodes can only be retrieved from outdated neighbor state tables exchanged when nodes come into communication range, without access to real\-time ground\-truth values;\(b\)When the destination node of a message cannot be connected at the current time step, the routing logic has to use locally cached historical information instead of up\-to\-date status of the target;\(c\)The heading\-control agent deployed on each UAV cannot directly obtain network\-wide stress fields and must rely on the local navigation vector and CTDE context\.

#### Neighbor exchange summaries\.

Each routing nodeiiobserves its ownbi\(t\)b\_\{i\}^\{\\,\(t\)\}andν¯i\(t\)\\bar\{\\nu\}\_\{i\}^\{\\,\(t\)\}\(Eqs\. \([6](https://arxiv.org/html/2608.04590#S3.E6)\) and \([5](https://arxiv.org/html/2608.04590#S3.E5)\)\) from on\-board statistics, but cannot read other nodes’ current buffer utilization or contact degree unless they meet\. To support contact\-limited routing features, each nodeiimaintains a neighbor exchange table\. Entry\(b~i​j,ν~i​j\)\(\\tilde\{b\}\_\{ij\},\\tilde\{\\nu\}\_\{ij\}\)storesii’s most recent received estimate of neighborjj’s buffer utilization and normalized contact degree, refreshed to\(bj\(t\),ν¯j\(t\)\)\(b\_\{j\}^\{\\,\(t\)\},\\bar\{\\nu\}\_\{j\}^\{\\,\(t\)\}\)wheneverCi​j\(t\)=1C\_\{ij\}^\{\\,\(t\)\}\{=\}1\. Letτi​j≥0\\tau\_\{ij\}\\geq 0denote the steps elapsed since that refresh\. Between meetings,b~i​j\\tilde\{b\}\_\{ij\}andν~i​j\\tilde\{\\nu\}\_\{ij\}remain stale summaries ofjj’s state at the last exchange rather than the live quantitiesbj\(t\)b\_\{j\}^\{\\,\(t\)\}andν¯j\(t\)\\bar\{\\nu\}\_\{j\}^\{\\,\(t\)\}\. The neighbor\-buffer componentb¯inb​\(t\)\\bar\{b\}\_\{i\}^\{\\mathrm\{nb\}\\,\(t\)\}of𝐱i\(t\)\\mathbf\{x\}\_\{i\}^\{\\,\(t\)\}is

b¯inb​\(t\)=∑j:Ci​j\(t\)=1exp⁡\(−ωex​τi​j\)​b~i​j∑j:Ci​j\(t\)=1exp⁡\(−ωex​τi​j\)\\bar\{b\}\_\{i\}^\{\\mathrm\{nb\}\\,\(t\)\}=\\frac\{\\sum\_\{j:C\_\{ij\}^\{\\,\(t\)\}\{=\}1\}\\exp\(\-\\omega\_\{\\mathrm\{ex\}\}\\tau\_\{ij\}\)\\,\\tilde\{b\}\_\{ij\}\}\{\\sum\_\{j:C\_\{ij\}^\{\\,\(t\)\}\{=\}1\}\\exp\(\-\\omega\_\{\\mathrm\{ex\}\}\\tau\_\{ij\}\)\}\(25\)Here,ωex\>0\\omega\_\{\\mathrm\{ex\}\}\>0is the exchange decay factor, andb¯inb​\(t\)=0\\bar\{b\}\_\{i\}^\{\\mathrm\{nb\}\(t\)\}=0when node i has no connected neighbors\. For a candidate at nodeiitoward destinationdd, the destination\-buffer feature in𝐜i,ℓ\(t\)\\mathbf\{c\}\_\{i,\\ell\}^\{\\,\(t\)\}\(Part III\-A\) usesbd\(t\)b\_\{d\}^\{\\,\(t\)\}whenCi​d\(t\)=1C\_\{id\}^\{\\,\(t\)\}\{=\}1and otherwiseb~i​d\\tilde\{b\}\_\{id\}; the destination\-degree feature readsν~i​d\\tilde\{\\nu\}\_\{id\}fromii’s table\.

#### UAV motion observations and CTDE context\.

The UAV agent can only perceive congestion hotspots indirectly via two types of inputs: the navigation vector𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\(t\)\}\(assembled from sector stress and related blocks\) and a CTDE contextgu\(t\)g\_\{u\}^\{\\,\(t\)\}that is instantiated asguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}orguL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\. Each motion unituuattached to graph nodeiui\_\{u\}receives

om,u\(t\)=\(𝐳u\(t\),gu\(t\)\),gu\(t\)∈\{guG​\(t\),guL​\(t\)\},o\_\{m,u\}^\{\\,\(t\)\}=\\bigl\(\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\},\\,g\_\{u\}^\{\\,\(t\)\}\\bigr\),\\quad g\_\{u\}^\{\\,\(t\)\}\\in\\bigl\\\{g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\},\\,g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\\bigr\\\},\(26\)where the navigation vector is the concatenation

𝐳u\(t\)=\[𝐒u\(t\);𝐞u\(t\);𝐧^u\(t\)\]∈ℝDy\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\}=\\bigl\[\\mathbf\{S\}\_\{u\}^\{\\,\(t\)\};\\,\\mathbf\{e\}\_\{u\}^\{\\,\(t\)\};\\,\\hat\{\\mathbf\{n\}\}\_\{u\}^\{\\,\(t\)\}\\bigr\]\\in\\mathbb\{R\}^\{D\_\{y\}\}\(27\)
The dimension of𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\(t\)\}follows Eq\. \([16](https://arxiv.org/html/2608.04590#S3.E16)\)\. For the same physical UAV relay, next\-hop forwarding control adopts the routing observationoiu\(t\)o\_\{i\_\{u\}\}^\{\(t\)\}, which shares an identical structure with that of ground vehicles\. The environment additionally generates a ground\-truth supervision vector𝐲u\(t\)∈ℝDy\\mathbf\{y\}\_\{u\}^\{\(t\)\}\\in\\mathbb\{R\}^\{D\_\{y\}\}with the same three\-block layout as𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\(t\)\}for optional auxiliary training\. Whenλ\>0\\lambda\>0, the LSTM prediction𝐲^u\(t\)\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\(t\)\}can be concatenated into the UAV heading policy head\.

#### Per\-step decomposition\.

Every term in Eqs\. \([18](https://arxiv.org/html/2608.04590#S4.E18)\)–\([19](https://arxiv.org/html/2608.04590#S4.E19)\) depends only on quantities available at steptt, so we instantiate \(P1\) by a common per\-step team rewardr\(t\)r^\{\\,\(t\)\}shared by allN\+NuavN\{\+\}N\_\{\\mathrm\{uav\}\}decision units, with each reward component corresponds to the corresponding terms in Eqs\. \([18](https://arxiv.org/html/2608.04590#S4.E18)\)–\([19](https://arxiv.org/html/2608.04590#S4.E19)\)\. Learning parameterizes the factored policy asπθ\\pi\_\{\\theta\}and maximizes the discounted return

J​\(θ\)=𝔼πθ​\[∑t=0Tmaxγt​r\(t\)\],J\(\\theta\)=\\mathbb\{E\}\_\{\\pi\_\{\\theta\}\}\\\!\\left\[\\sum\_\{t=0\}^\{T\_\{\\max\}\}\\gamma^\{t\}\\,r^\{\\,\(t\)\}\\right\],\(28\)Whenγ=1\\gamma=1, Eq\. \([28](https://arxiv.org/html/2608.04590#S4.E28)\) is equivalent to the undiscounted finite\-horizon objective \(P1\)\. This discounted objective acts as a standard reinforcement\-learning surrogate for \(P1\)\.

Our JUROR that parameterizes the factorized policyπθ\\pi\_\{\\theta\}and optimizes the above team return is elaborated in Sec\.[V](https://arxiv.org/html/2608.04590#S5)\.

## VProposed JUROR Algorithm

This section introduces JUROR, our joint opportunistic routing and UAV motion control mechanism under the CTDE paradigm\. Building on the finite\-horizon cooperative decision process formulated in Sec\.[IV](https://arxiv.org/html/2608.04590#S4), this subsection details the algorithm implementation\.

### V\-APer\-Step Team Reward Instantiation

The observation, action and reward specifications for all cooperative agents adhere to the definitions provided in Sec\.[IV](https://arxiv.org/html/2608.04590#S4)\. Following the formulation of the joint optimization problem \(P1\), we construct a multi\-objective single\-step team reward signal\. This reward is designed to maximize the number of successfully delivered messages, while penalizing TTL expiration events, buffer overflow and unregulated network congestion\. Additionally, it incorporates shaping terms to regularize UAV fleet spatial distribution and supports an optional hotspot\-guided alignment \(HGA\) regularization configuration\.

1\) Optional Hotspot\-Guided Alignment \(HGA\):Sparse SCF deliveries provide weak motion cues for where relays should fly next\. HGA complements the primary delivery reward by scoring whether each UAV’s executed heading agrees with the instantaneous directional congestion layout around that relay, thereby translating the sector\-stress field of Sec\.[III\-C](https://arxiv.org/html/2608.04590#S3.SS3)into an explicit geometric prior for heading control\. The optional hotspot\-guided alignment module calculates a fleet\-level matching scoreα\(t\)∈\[0,1\]\\alpha^\{\(t\)\}\\in\[0,1\]\. This metric is constructed from the eight\-dimensional sector\-stress subvector of either the ground\-truth supervision vector𝐲u\(t\)\\mathbf\{y\}\_\{u\}^\{\(t\)\}or the gradient\-detached LSTM prediction𝐲^u\(t\)\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\(t\)\}; under the default ablation*base\+\+HGA*, ground\-truth sector masses are used and the LSTM branch remains off \(λ=0\\lambda\{=\}0\)\. Let𝐝u\(t\)=𝐝mu\(t\)\\mathbf\{d\}\_\{u\}^\{\(t\)\}=\\mathbf\{d\}\_\{m\_\{u\}^\{\(t\)\}\}denote the unit vector of the UAV’s actual heading, and let𝐰=\(w0,…,w7\)∈ℝ≥08\\mathbf\{w\}=\(w\_\{0\},\\dots,w\_\{7\}\)\\in\\mathbb\{R\}\_\{\\geq 0\}^\{8\}be a non\-negative sector weight vector\. We define the normalized reference heading vector as

𝐯​\(𝐰\)=∑k=07wk​𝐝k‖∑k=07wk​𝐝k‖2\\mathbf\{v\}\(\\mathbf\{w\}\)=\\frac\{\\sum\_\{k=0\}^\{7\}w\_\{k\}\\,\\mathbf\{d\}\_\{k\}\}\{\\left\\\|\\sum\_\{k=0\}^\{7\}w\_\{k\}\\,\\mathbf\{d\}\_\{k\}\\right\\\|\_\{2\}\}\(29\)If∑kwk​𝐝k=𝟎\\sum\_\{k\}w\_\{k\}\\mathbf\{d\}\_\{k\}=\\mathbf\{0\}, we set𝐯​\(𝐰\)=𝟎\\mathbf\{v\}\(\\mathbf\{w\}\)=\\mathbf\{0\}\. Here,𝐝k\\mathbf\{d\}\_\{k\}stands for the unit direction vector of heading bin k, as specified in Sec\.[III](https://arxiv.org/html/2608.04590#S3)\. The weight vector𝐰\\mathbf\{w\}is populated with the sector\-wise subvector𝐲u,s\(t\)\\mathbf\{y\}\_\{u,s\}^\{\(t\)\}or𝐲^u,s\(t\)\\hat\{\\mathbf\{y\}\}\_\{u,s\}^\{\(t\)\}\. Accordingly, we derive two fleet\-averaged alignment scores for ground\-truth \(GT\) and predicted \(Pred\) hotspot targets separately:

αg\(t\)\\displaystyle\\alpha\_\{g\}^\{\\,\(t\)\}=1Nuav​∑u=0Nuav−1max⁡\(0,𝐝u\(t\)⊤​𝐯​\(𝐲u,s\(t\)\)\),\\displaystyle=\\frac\{1\}\{N\_\{\\mathrm\{uav\}\}\}\\sum\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\max\\\!\\bigl\(0,\\,\{\\mathbf\{d\}\_\{u\}^\{\\,\(t\)\}\}^\{\\top\}\\mathbf\{v\}\(\\mathbf\{y\}\_\{u,s\}^\{\\,\(t\)\}\)\\bigr\),\(30\)αp\(t\)\\displaystyle\\alpha\_\{p\}^\{\\,\(t\)\}=1Nuav​∑u=0Nuav−1max⁡\(0,𝐝u\(t\)⊤​𝐯​\(𝐲^u,s\(t\)\)\)\.\\displaystyle=\\frac\{1\}\{N\_\{\\mathrm\{uav\}\}\}\\sum\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\max\\\!\\bigl\(0,\\,\{\\mathbf\{d\}\_\{u\}^\{\\,\(t\)\}\}^\{\\top\}\\mathbf\{v\}\(\\hat\{\\mathbf\{y\}\}\_\{u,s\}^\{\\,\(t\)\}\)\\bigr\)\.\(31\)When HGA is activated, we assignα\(t\)=αg\(t\)\\alpha^\{\\,\(t\)\}=\\alpha\_\{g\}^\{\\,\(t\)\}orα\(t\)=αp\(t\)\\alpha^\{\\,\(t\)\}=\\alpha\_\{p\}^\{\\,\(t\)\}according to whether ground\-truth or predicted hotspot supervision is adopted\. The resultingα\(t\)\\alpha^\{\\,\(t\)\}enters the team reward throughαa​α\(t\)\+αa​d​α\(t\)​D\(t\)\+αa​r​α\(t\)​Ih\(t\)\\alpha\_\{a\}\\alpha^\{\\,\(t\)\}\+\\alpha\_\{ad\}\\alpha^\{\\,\(t\)\}D^\{\\,\(t\)\}\+\\alpha\_\{ar\}\\alpha^\{\\,\(t\)\}I\_\{h\}^\{\\,\(t\)\}in Eq\. \([32](https://arxiv.org/html/2608.04590#S5.E32)\), so alignment is rewarded more strongly when deliveries or successful replications occur in the same step\.

2\) Per\-Step Team RewardFor the multi\-agent cooperative team, the single\-step scalar rewardr\(t\)r^\{\(t\)\}in Eq\. \([28](https://arxiv.org/html/2608.04590#S4.E28)\) is defined as a linear combination of the component terms in Eqs\. \([18](https://arxiv.org/html/2608.04590#S4.E18)\)–\([19](https://arxiv.org/html/2608.04590#S4.E19)\), as follows\.

r\(t\)\\displaystyle r^\{\(t\)\}=αd​D\(t\)\+αe​E\(t\)\+αd​r​R\(t\)\+αb​B¯\(t\)\+αs\+Ih\(t\)​αh\\displaystyle=\\alpha\_\{d\}D^\{\(t\)\}\+\\alpha\_\{e\}E^\{\(t\)\}\+\\alpha\_\{dr\}R^\{\(t\)\}\+\\alpha\_\{b\}\\bar\{B\}^\{\(t\)\}\+\\alpha\_\{s\}\+I\_\{h\}^\{\(t\)\}\\alpha\_\{h\}\+αρ​ρ\(t\)\+αc​ρ\(t\)​D\(t\)\\displaystyle\\quad\+\\alpha\_\{\\rho\}\\rho^\{\(t\)\}\+\\alpha\_\{c\}\\rho^\{\(t\)\}D^\{\(t\)\}\+αx​gs​\(D\(t\)\)​ξ\(t\)−αk​gm​\(D\(t\)\)​κ\(t\)\+αf​ge​\(E\(t\)\)​ρf\(t\)\\displaystyle\\quad\+\\alpha\_\{x\}g\_\{s\}\\bigl\(D^\{\(t\)\}\\bigr\)\\xi^\{\(t\)\}\-\\alpha\_\{k\}g\_\{m\}\\bigl\(D^\{\(t\)\}\\bigr\)\\kappa^\{\(t\)\}\+\\alpha\_\{f\}g\_\{e\}\\bigl\(E^\{\(t\)\}\\bigr\)\\rho\_\{f\}^\{\(t\)\}\+αa​α\(t\)\+αa​d​α\(t\)​D\(t\)\+αa​r​α\(t\)​Ih\(t\),\\displaystyle\\quad\+\\alpha\_\{a\}\\alpha^\{\(t\)\}\+\\alpha\_\{ad\}\\alpha^\{\(t\)\}D^\{\(t\)\}\+\\alpha\_\{ar\}\\alpha^\{\(t\)\}I\_\{h\}^\{\(t\)\},\(32\)where the stage\-dependent gating functions adopt the definitions provided in the preceding subsection\. The step\-wise state metricsD\(t\),E\(t\),R\(t\),B¯\(t\),Ih\(t\),ρ\(t\),ξ\(t\),κ\(t\)D^\{\(t\)\},E^\{\(t\)\},R^\{\(t\)\},\\bar\{B\}^\{\(t\)\},I\_\{h\}^\{\(t\)\},\\rho^\{\(t\)\},\\xi^\{\(t\)\},\\kappa^\{\(t\)\}as well as the optional predicted congestion termρf\(t\)\\rho\_\{f\}^\{\(t\)\}are defined in Sec\.[IV\-A](https://arxiv.org/html/2608.04590#S4.SS1), Sec\.[III\-C](https://arxiv.org/html/2608.04590#S3.SS3), and Sec\.[III\-C](https://arxiv.org/html/2608.04590#S3.SS3), respectively\.

### V\-BCTDE–PPO Learning Algorithm

Fig\.[2](https://arxiv.org/html/2608.04590#S5.F2)summarizes the JUROR structure: a factorized MDP framework, decentralized actors for routing and UAV control, and a centralized critic optimized with PPO\. Under CTDE, each agent acts from decentralized observations, while the critic conditions on global statistics only during training\. This design preserves deployability under intermittent DTN contacts and uses privileged global information to reduce training variance\. The structure includes four modules:\(S1\)constructs structured multi\-agent observations from simulation states,\(S2\)outputs UAV mobility and routing actions with a training\-only centralized critic,\(S3\)steps the simulator forward and stores on\-policy transition samples, and\(S4\)computes composite PPO\-auxiliary loss to update actor\-critic network parameters\. At deployment, only the actors related with \(S1\)–\(S2\) are retained\.

![Refer to caption](https://arxiv.org/html/2608.04590v1/x3.png)Figure 2:JUROR end\-to\-end architecture under CTDE–PPO\. Factored MDP interface with decentralized routing and UAV actions, training\-only critic ons\(t\)s^\{\\,\(t\)\}, and team rewardr\(t\)r^\{\\,\(t\)\}from the SCF simulator \(Sec\.[IV](https://arxiv.org/html/2608.04590#S4)\)\. One training epoch comprises stages S1–S4 \(environment interface, CTDE policy, on\-policy buffer, PPO update with optional multi\-horizon auxiliary lossλ​Lp\\lambda L\_\{p\}\); dashed paths denote delayed LSTM supervision\. Actor–critic layers use shared MLP width 256 and an optional per\-UAV LSTM \(hidden 128\) producing𝐲^u\\hat\{\\mathbf\{y\}\}\_\{u\}for concatenation into the UAV direction head and forLpL\_\{p\};Dy=8\+2​Kσ\+2D\_\{y\}\{=\}8\{\+\}2K\_\{\\sigma\}\{\+\}2\(Eq\. \([16](https://arxiv.org/html/2608.04590#S3.E16)\)\)\. At deployment, only S1–S2 actors run with contact\-limited observations; the critic and S3–S4 are omitted\. Default base usesλ=0\\lambda\{=\}0\(LSTM branch off\)\.JUROR parameterizes a shared policyπθ\\pi\_\{\\theta\}for all routing and motion factors\. The joint log\-probability factorizes as

log⁡πθ​\(𝐚\(t\)∣𝐨\(t\)\)\\displaystyle\\log\\pi\_\{\\theta\}\(\\mathbf\{a\}^\{\\,\(t\)\}\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)=∑i=0N−1log⁡πθ,i​\(ai\(t\)∣oi\(t\)\)\\displaystyle=\\sum\_\{i=0\}^\{N\-1\}\\log\\pi\_\{\\theta,i\}\(a\_\{i\}^\{\\,\(t\)\}\\mid o\_\{i\}^\{\\,\(t\)\}\)\(33\)\+∑u=0Nuav−1log⁡πθ,m,u​\(mu\(t\)∣om,u\(t\)\),\\displaystyle\\quad\+\\sum\_\{u=0\}^\{N\_\{\\mathrm\{uav\}\}\-1\}\\log\\pi\_\{\\theta,m,u\}\(m\_\{u\}^\{\\,\(t\)\}\\mid o\_\{m,u\}^\{\\,\(t\)\}\),Using the standard policy\-gradient estimator with centralized\-critic advantage, the shared\-parameter update is

∇θJ​\(θ\)≈𝔼​\[A^\(t\)​∇θlog⁡πθ​\(𝐚\(t\)∣𝐨\(t\)\)\]\.\\nabla\_\{\\theta\}J\(\\theta\)\\;\\approx\\;\\mathbb\{E\}\\\!\\left\[\\hat\{A\}^\{\\,\(t\)\}\\,\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(\\mathbf\{a\}^\{\\,\(t\)\}\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)\\right\]\.\(34\)whereA^\(t\)\\hat\{A\}^\{\\,\(t\)\}is computed by the centralized critic from global state\. Substituting Eq\. \([33](https://arxiv.org/html/2608.04590#S5.E33)\) into Eq\. \([34](https://arxiv.org/html/2608.04590#S5.E34)\) yields a sum of routing and UAV\-motion log\-gradient terms, all scaled by the same team\-levelA^\(t\)\\hat\{A\}^\{\\,\(t\)\}\. Hence, both routing actor and UAV actor are optimized jointly toward one cooperative objective while preserving decentralized action selection\.

1\) Decentralized Routing Actor:For each routing nodeii\(ground vehicle or UAV graph nodeiui\_\{u\}\), the decentralized observation isoi\(t\)=\(𝐱i\(t\),\{𝐜i,ℓ\(t\)\}ℓ=1K\)o\_\{i\}^\{\\,\(t\)\}=\(\\mathbf\{x\}\_\{i\}^\{\\,\(t\)\},\\\{\\mathbf\{c\}\_\{i,\\ell\}^\{\\,\(t\)\}\\\}\_\{\\ell=1\}^\{K\}\), i\.e\., one local state vector and up toKKcandidate rows\. The routing head encodes each local observationoi\(t\)o\_\{i\}^\{\\,\(t\)\}and its feasible candidates with MLP encoders\. Node and candidate embeddings are fused to produceK\+1K\{\+\}1masked logits per node\.

2\) UAV Direction Head:The motion head encodesom,u\(t\)=\(𝐳u\(t\),gu\(t\)\)o\_\{m,u\}^\{\\,\(t\)\}=\(\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\},g\_\{u\}^\{\\,\(t\)\}\)withgu\(t\)∈\{guG​\(t\),guL​\(t\)\}g\_\{u\}^\{\\,\(t\)\}\\in\\\{g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\},g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\\\}, whenλ\>0\\lambda\{\>\}0, it additionally uses the optional LSTM prediction𝐲^u\(t\)\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\\,\(t\)\}\(dashed path in Fig\.[2](https://arxiv.org/html/2608.04590#S5.F2)\)\. Letℓu\(t\)∈ℝB\\ell\_\{u\}^\{\\,\(t\)\}\\in\\mathbb\{R\}^\{B\}denote direction logits andfmf\_\{m\}denote the motion MLP\. Whenλ\>0\\lambda\{\>\}0,

ℓu\(t\)=fm​\(\[gu\(t\);𝐳u\(t\);𝐲^u\(t\)\]\),\\ell\_\{u\}^\{\\,\(t\)\}=f\_\{m\}\\\!\\bigl\(\[g\_\{u\}^\{\\,\(t\)\};\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\};\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\\,\(t\)\}\]\\bigr\),\(35\)andπθ,m,u\\pi\_\{\\theta,m,u\}samplesmu\(t\)m\_\{u\}^\{\\,\(t\)\}fromsoftmax⁡\(ℓu\(t\)\)\\operatorname\{softmax\}\(\\ell\_\{u\}^\{\\,\(t\)\}\)\. Ifλ=0\\lambda\{=\}0, the𝐲^u\(t\)\\hat\{\\mathbf\{y\}\}\_\{u\}^\{\\,\(t\)\}is omitted\. Routing and motion branches each use a product of categoricals with masked feasible sets\.

3\) Centralized Critic:The centralized critic module in Fig\.[2](https://arxiv.org/html/2608.04590#S5.F2)approximates the state valueV​\(s\(t\)\)V\\\!\\bigl\(s^\{\\,\(t\)\}\\bigr\)from global statisticss\(t\)s^\{\\,\(t\)\}\(Sec\.[IV\-C](https://arxiv.org/html/2608.04590#S4.SS3)\) through summary MLPs and a value head\. Leveraging the full global states\(t\)s^\{\(t\)\}effectively cuts the variance of the generalized advantage estimateA^\(t\)\\hat\{A\}^\{\(t\)\}, since complete network\-wide statistics eliminate the information deficit caused by partial local observations and produce more accurate state\-value predictions, whereas all decentralized actors make decisions solely based on local, contact\-limited observations\. The global states\(t\)s^\{\\,\(t\)\}is assembled from Table[II](https://arxiv.org/html/2608.04590#S3.T2), Part III\-C, and is available only in the training phase\.

4\) Joint Learning Objective:JUROR uses standard PPO\[[16](https://arxiv.org/html/2608.04590#bib.bib16)\]as the policy optimizer and an optional auxiliary multi\-horizon congestion hotspot forecasting task is integrated into the overall loss function via a scaled auxiliary loss termλ​Lp\\lambda L\_\{p\}\. From on\-policy tuples\{𝐨\(t\),𝐚\(t\),r\(t\),s\(t\)\}\\\{\\mathbf\{o\}^\{\\,\(t\)\},\\mathbf\{a\}^\{\\,\(t\)\},r^\{\\,\(t\)\},s^\{\\,\(t\)\}\\\}, we compute GAE advantagesA^\(t\)\\hat\{A\}^\{\\,\(t\)\}and form the ratio

ϱ\(t\)​\(θ\)=πθ​\(𝐚\(t\)∣𝐨\(t\)\)πθ′​\(𝐚\(t\)∣𝐨\(t\)\),\\varrho^\{\\,\(t\)\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(\\mathbf\{a\}^\{\\,\(t\)\}\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)\}\{\\pi\_\{\\theta^\{\\prime\}\}\(\\mathbf\{a\}^\{\\,\(t\)\}\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)\},\(36\)with the clipped surrogate

Lclip​\(θ\)=𝔼​\[min⁡\(ϱ\(t\)​\(θ\)​A^\(t\),clip⁡\(ϱ\(t\)​\(θ\),1−ϵ,1\+ϵ\)​A^\(t\)\)\],L\_\{\\mathrm\{clip\}\}\(\\theta\)=\\mathbb\{E\}\\\!\\Bigl\[\\min\\bigl\(\\varrho^\{\\,\(t\)\}\(\\theta\)\\,\\hat\{A\}^\{\\,\(t\)\},\\,\\operatorname\{clip\}\\bigl\(\\varrho^\{\\,\(t\)\}\(\\theta\),1\-\\epsilon,1\+\\epsilon\\bigr\)\\hat\{A\}^\{\\,\(t\)\}\\bigr\)\\Bigr\],\(37\)For JUROR, the clipped PPO surrogate loss in Eq\. \([37](https://arxiv.org/html/2608.04590#S5.E37)\) is computed with the factorized joint policy formulated in Eq\. \([33](https://arxiv.org/html/2608.04590#S5.E33)\)\. The training objective is

L\\displaystyle L=LP\+λ​Lp,\\displaystyle=L\_\{P\}\+\\lambda L\_\{p\},LP\\displaystyle L\_\{P\}=−Lclip\+cv​LV−ce​LH,\\displaystyle=\-L\_\{\\mathrm\{clip\}\}\+c\_\{v\}L\_\{V\}\-c\_\{e\}L\_\{H\},\(38\)with

LV\\displaystyle L\_\{V\}=𝔼​\[\(Vψ​\(s\(t\)\)−V^ret\(t\)\)2\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\bigl\(V\_\{\\psi\}\(s^\{\\,\(t\)\}\)\-\\hat\{V\}\_\{\\mathrm\{ret\}\}^\{\\,\(t\)\}\\bigr\)^\{2\}\\right\],\(39\)LH\\displaystyle L\_\{H\}=𝔼\[ℋ\(πθ\(⋅∣𝐨\(t\)\)\)\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\mathcal\{H\}\\\!\\left\(\\pi\_\{\\theta\}\(\\cdot\\mid\\mathbf\{o\}^\{\\,\(t\)\}\)\\right\)\\right\],\(40\)whereV^ret\(t\)\\hat\{V\}\_\{\\mathrm\{ret\}\}^\{\\,\(t\)\}is the return target andℋ​\(⋅\)\\mathcal\{H\}\(\\cdot\)denotes entropy\.

The complete CTDE–PPO training workflow of JUROR is summarized in Algorithm[2](https://arxiv.org/html/2608.04590#alg2)\. Algorithm[1](https://arxiv.org/html/2608.04590#alg1)implements the multi\-horizon target alignment procedure, which retrieves congestion hotspot ground truth states delayed byΔ\\Deltatime steps to construct supervision labels for the auxiliary lossLpL\_\{p\}\. Detailed implementations of vectorized environment rollouts and experience replay buffers are provided in Section[VI\-A](https://arxiv.org/html/2608.04590#S6.SS1)\.

Algorithm 1Gather multi\-horizon hotspot targets from rollout buffer1:On\-policy rollout buffer

ℛ\\mathcal\{R\}; minibatch index set

𝒮\\mathcal\{S\}; horizon set

ℋ\\mathcal\{H\}; episode\-end flags

\{δb\}\\\{\\delta\_\{b\}\\\}stored in

ℛ\\mathcal\{R\}
2:Target tensor

𝐘tgt∈ℝ\|𝒮\|×\|ℋ\|×Nuav×Dy\\mathbf\{Y\}^\{\\mathrm\{tgt\}\}\\in\\mathbb\{R\}^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{H\}\|\\times N\_\{\\mathrm\{uav\}\}\\times D\_\{y\}\}; validity mask

𝐌val∈\{0,1\}\|𝒮\|×\|ℋ\|\\mathbf\{M\}^\{\\mathrm\{val\}\}\\in\\\{0,1\\\}^\{\|\\mathcal\{S\}\|\\times\|\\mathcal\{H\}\|\}
3:foreach horizon

Δ∈ℋ\\Delta\\in\\mathcal\{H\}do

4:foreach sample index

ν∈𝒮\\nu\\in\\mathcal\{S\}do

5:

b←νb\\leftarrow\\nu;

Mν,Δval←1M^\{\\mathrm\{val\}\}\_\{\\nu,\\Delta\}\\leftarrow 1
6:for

ℓ=1\\ell=1to

Δ\\Deltado

7:if

δb=1\\delta\_\{b\}=1then

8:

Mν,Δval←0M^\{\\mathrm\{val\}\}\_\{\\nu,\\Delta\}\\leftarrow 0;break

9:endif

10:

b←SuccInEpisode​\(ℛ,b\)b\\leftarrow\\textsc\{SuccInEpisode\}\(\\mathcal\{R\},b\)//next slot in same episode

11:endfor

12:if

Mν,Δval=1M^\{\\mathrm\{val\}\}\_\{\\nu,\\Delta\}=1then

13:

𝐘ν,Δtgt←\\mathbf\{Y\}^\{\\mathrm\{tgt\}\}\_\{\\nu,\\Delta\}\\leftarrowhotspot

𝐲\\mathbf\{y\}stored at buffer index

bb
14:endif

15:endfor

16:endfor

Algorithm 2JUROR CTDE–PPO training \(one epoch\)1:Initial policy parameters

θ\\theta; PPO coefficients

ϵ\\epsilon,

cvc\_\{v\},

cec\_\{e\}; auxiliary weight

λ\\lambda; horizons

ℋ\\mathcal\{H\}with weights

\{wΔ\}\\\{w\_\{\\Delta\}\\\}; minibatch size

BmbB\_\{\\mathrm\{mb\}\};

NpassN\_\{\\mathrm\{pass\}\}update passes \(Sec\.[VI\-A](https://arxiv.org/html/2608.04590#S6.SS1)\)

2:Phase 1 \(S3\):collect rollouts

3:Roll out parallel environments into on\-policy buffer

ℛ\\mathcal\{R\}
4:Store

\(𝐨\(t\),𝐚\(t\),r\(t\),s\(t\),δ\(t\),hu\(t\),cu\(t\)\)\(\\mathbf\{o\}^\{\\,\(t\)\},\\mathbf\{a\}^\{\\,\(t\)\},r^\{\\,\(t\)\},s^\{\\,\(t\)\},\\delta^\{\\,\(t\)\},h\_\{u\}^\{\\,\(t\)\},c\_\{u\}^\{\\,\(t\)\}\)per step;

δ\(t\)=1\\delta^\{\\,\(t\)\}\{=\}1marks episode end

5:Phase 2 \(S4\):estimate advantages

6:Compute

V​\(s\(t\)\)V\(s^\{\\,\(t\)\}\)and GAE advantages

A^\(t\)\\hat\{A\}^\{\\,\(t\)\}from rewards

\{r\(t\)\}\\\{r^\{\\,\(t\)\}\\\}
7:Phase 3 \(S4\):joint update

8:for

p=1p=1to

NpassN\_\{\\mathrm\{pass\}\}do

9:foreach minibatch index set

𝒮\\mathcal\{S\}of size

BmbB\_\{\\mathrm\{mb\}\}from

ℛ\\mathcal\{R\}do

10:

LP←−Lclip​\(θ\)\+cv​LV−ce​LHL\_\{P\}\\leftarrow\-L\_\{\\mathrm\{clip\}\}\(\\theta\)\+c\_\{v\}L\_\{V\}\-c\_\{e\}L\_\{H\}
11:

\(𝐘tgt,𝐌val\)←GatherTargets​\(ℛ,𝒮,ℋ\)\(\\mathbf\{Y\}^\{\\mathrm\{tgt\}\},\\mathbf\{M\}^\{\\mathrm\{val\}\}\)\\leftarrow\\textsc\{GatherTargets\}\(\\mathcal\{R\},\\mathcal\{S\},\\mathcal\{H\}\)//Alg\.[1](https://arxiv.org/html/2608.04590#alg1)

12:

Lp←∑Δ∈ℋwΔ​∑ν∈𝒮Mν,Δval​‖𝐲^\(tν\)−𝐘ν,Δtgt‖22L\_\{p\}\\leftarrow\\sum\_\{\\Delta\\in\\mathcal\{H\}\}w\_\{\\Delta\}\\sum\_\{\\nu\\in\\mathcal\{S\}\}M^\{\\mathrm\{val\}\}\_\{\\nu,\\Delta\}\\bigl\\\|\\hat\{\\mathbf\{y\}\}^\{\\,\(t\_\{\\nu\}\)\}\-\\mathbf\{Y\}^\{\\mathrm\{tgt\}\}\_\{\\nu,\\Delta\}\\bigr\\\|\_\{2\}^\{2\}
13:

θ←θ−η​∇θ\(LP\+λ​Lp\)\\theta\\leftarrow\\theta\-\\eta\\nabla\_\{\\theta\}\(L\_\{P\}\+\\lambda L\_\{p\}\)with gradient clipping

14:endfor

15:endfor

## VISimulation Results and Discussions

This section presents experimental results that substantiate the performance of the proposed JUROR method\. Subsequent subsections first specify the shared simulation settings, then conduct an ablation study that verifies the core contribution in our work, next we compare classical and recent routing baselines with JUROR, and finally we analyze performance trends from a mechanism perspective\.

### VI\-ASimulation Settings

In the experiment, JUROR is implemented in Python whose environment interface conforms to Gymnasium\[[27](https://arxiv.org/html/2608.04590#bib.bib27)\]and is trained with Tianshou\[[28](https://arxiv.org/html/2608.04590#bib.bib28)\]under vectorized parallel environments that execute the models of Secs\.[III](https://arxiv.org/html/2608.04590#S3)–[V](https://arxiv.org/html/2608.04590#S5)\. We evaluate on a Helsinki\-medium WKT testbed with defaultN=70N\{=\}70nodes \(Ngr=65N\_\{\\mathrm\{gr\}\}\{=\}65,Nuav=5N\_\{\\mathrm\{uav\}\}\{=\}5\), communication rangesrgr=300r\_\{\\mathrm\{gr\}\}\{=\}300m andruav=900r\_\{\\mathrm\{uav\}\}\{=\}900m, and episode horizonTmax=5000T\_\{\\max\}\{=\}5000steps\. Four traffic modes M1–M4 \(Table[III](https://arxiv.org/html/2608.04590#S6.T3)\) are organized along two orthogonal axes of source multiplicity and temporal injection pattern, thereby characterizing robustness under relay planning, sustained congestion, multi\-source scheduling, and concurrent burst load\. M1 is adopted as the primary ablation setting, whereas M2–M4 are designed to stress congestion\-aware forwarding, multi\-source competition, and concurrent burst coverage, respectively\.

Table III:Traffic modes employed in the experimental evaluation \(Sec\.[VI](https://arxiv.org/html/2608.04590#S6)\)\. The modes are organized along two orthogonal axes: source multiplicity \(single\-source versus multi\-source\) and temporal injection pattern \(one\-shot burst versus sustained injection\)\.ModeAxesInjection schedule and evaluation focusM1Single\-source; one\-shot burstAtt=0t\{=\}0, one source injects 69 messages toward the remaining nodes, stressing sparse connectivity and UAV relay planning; primary ablation setting\.M2Network\-wide; sustained injectionInjection every 20 steps yields∼\\sim4300 messages/episode, stressing congestion\-aware forwarding, buffer pressure, and delivery–auxiliary interactions under persistent load\.M3Multi\-source; sustained injectionStochastic multi\-source arrivals every 20 steps yield∼\\sim306 messages/episode, stressing multi\-source scheduling and competition for shared contacts\.M4Multi\-source; one\-shot burstAtt=0t\{=\}0, five sources concurrently inject∼\\sim345 messages, stressing replication control and UAV fleet coverage under simultaneous demand\.Unless otherwise noted, each run trains PPO for 80 epochs with seed 42, minibatch size 128, clipϵ=0\.2\\epsilon\{=\}0\.2, value\-loss weightcv=0\.5c\_\{v\}\{=\}0\.5, and entropy weightce=0\.01c\_\{e\}\{=\}0\.01, using eight parallel training and ten evaluation environments, and a shared actor MLP of hidden width 256 for routing and UAV direction heads\. The default base disables the LSTM auxiliary branch \(λ=0\\lambda\{=\}0\); optional base\+\+LSTM usesλ=0\.05\\lambda\{=\}0\.05with horizonsℋ=\{1,10,25\}\\mathcal\{H\}\{=\}\\\{1,10,25\\\}and weightswΔ∈\{0\.34,0\.33,0\.33\}w\_\{\\Delta\}\\in\\\{0\.34,0\.33,0\.33\\\}, retaining forecast in the on\-policy buffer for multi\-horizon supervision\. Held\-out metrics are epoch meansDtest\(e\)D\_\{\\text\{test\}\}^\{\\,\(e\)\},Etest\(e\)E\_\{\\text\{test\}\}^\{\\,\(e\)\}, andRtest\(e\)R\_\{\\text\{test\}\}^\{\\,\(e\)\}\. Unless otherwise noted, the team\-reward weights areαd=4\.0\\alpha\_\{d\}\{=\}4\.0,αe=−1\.5\\alpha\_\{e\}\{=\}\{\-\}1\.5,αd​r=−0\.5\\alpha\_\{dr\}\{=\}\{\-\}0\.5,αb=−0\.1\\alpha\_\{b\}\{=\}\{\-\}0\.1,αs=−0\.03\\alpha\_\{s\}\{=\}\{\-\}0\.03,αh=−0\.005\\alpha\_\{h\}\{=\}\{\-\}0\.005,αρ=0\.05\\alpha\_\{\\rho\}\{=\}0\.05,αc=0\.015\\alpha\_\{c\}\{=\}0\.015,αx=0\.006\\alpha\_\{x\}\{=\}0\.006,αk=0\.003\\alpha\_\{k\}\{=\}0\.003,αf=0\.01\\alpha\_\{f\}\{=\}0\.01, andαa=αa​d=αa​r=0\\alpha\_\{a\}\{=\}\\alpha\_\{ad\}\{=\}\\alpha\_\{ar\}\{=\}0\(HGA off, ablations will retuneαa,αa​d,αa​r\\alpha\_\{a\},\\alpha\_\{ad\},\\alpha\_\{ar\}orαf\\alpha\_\{f\}\)\. Delivery weights satisfy\|αd\|≫\|αρ\|,\|αx\|,\|αk\|\|\\alpha\_\{d\}\|\\gg\|\\alpha\_\{\\rho\}\|,\|\\alpha\_\{x\}\|,\|\\alpha\_\{k\}\|so that \(P1\) remains delivery\-centric\. Instantiating the gates of Sec\.[IV\-A](https://arxiv.org/html/2608.04590#S4.SS1), we setgs​\(D\(t\)\)=1g\_\{s\}\(D^\{\\,\(t\)\}\)\{=\}1ifD\(t\)=0D^\{\\,\(t\)\}\{=\}0and0\.30\.3otherwise,gm​\(D\(t\)\)=1g\_\{m\}\(D^\{\\,\(t\)\}\)\{=\}1ifD\(t\)=0D^\{\\,\(t\)\}\{=\}0and0\.50\.5otherwise, andge​\(E\(t\)\)=min⁡\(1,E\(t\)/2\)g\_\{e\}\(E^\{\\,\(t\)\}\)=\\min\(1,E^\{\\,\(t\)\}/2\)\. Since the four traffic modes generate distinct volumes of injected packets, cross\-mode comparison of absolute delivered messages mixes the impact of traffic load and algorithm performance, only intra\-mode comparisons are valid for performance analysis\. The ablation experiments record peak and terminal absolute delivery volumes of the joint learning policy on held\-out test environments, while baseline evaluations compute and focus on delivery ratios\.

### VI\-BAblation Study

The ablation study evaluates learned JUROR policies through held\-out testing at the end of each training epoch over parallel environments, LetDtest​\(e\)D\_\{\\text\{test\}\}\(e\)denote the mean delivered count on the evaluation environments after epochee, we report the peak valuemaxe⁡Dtest\(e\)\\max\_\{e\}D\_\{\\text\{test\}\}^\{\\,\(e\)\}and the terminal valueDtest\(80\)D\_\{\\text\{test\}\}^\{\\,\(80\)\}\. Fig\.[4](https://arxiv.org/html/2608.04590#S6.F4)shows the held\-out delivery learning curves under M1, which injects one message from source node 0 to each of the other 69 nodes att=0t\{=\}0and serves as the primary setting for interpreting ablation trends\. Table[IV](https://arxiv.org/html/2608.04590#S6.T4)consolidates six core runs across M1–M4 along a CTDE observation axis and a structure–shaping axis\. We verify the contributions stated in Sec\.[I](https://arxiv.org/html/2608.04590#S1): joint UAV–routing necessity, CTDE observation factorization from deploy\-realistic to privileged training, and optional hotspot auxiliary learning and HGA coupling\. TheBaseconfiguration disables the LSTM auxiliary branch \(λ=0\\lambda=0\) and hotspot\-guided alignment \(HGA\), and adopts global UAV contextguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}together with the complete navigation vector𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\(t\)\}\. The centralized critic always ingests global network statisticss\(t\)s^\{\(t\)\}for all training runs\.Two extended variants are defined:Base \+ LSTM: activates multi\-horizon auxiliary loss withλ=0\.05\\lambda=0\.05;Base \+ HGA: enables HGA with hyperparameters\(αa,αa​d,αa​r\)=\(0\.003,0\.012,0\.004\)\(\\alpha\_\{a\},\\alpha\_\{ad\},\\alpha\_\{ar\}\)=\(0\.003,0\.012,0\.004\)\.

Table IV:Cross\-traffic ablation on Helsinki\-medium \(N=70N\{=\}70,Nuav=5N\_\{\\mathrm\{uav\}\}\{=\}5\): peak and terminal held\-outDtestD\_\{\\text\{test\}\}after 80 training epochs \(seed 42\)\. The table retains six core configurations along a CTDE observation axis \(deploy\-realistic→\\rightarrowprivileged\) and a structure–shaping axis \(fleet composition; optional HGA\)\.*Base*\(ref\.\):λ=0\\lambda\{=\}0\(LSTM off\), global UAV contextguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}, full𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\}, HGA off\. Optional*base\+\+LSTM*usesλ=0\.05\\lambda\{=\}0\.05;*base\+\+HGA*uses\(αa,αa​d,αa​r\)=\(0\.003,0\.012,0\.004\)\(\\alpha\_\{a\},\\alpha\_\{ad\},\\alpha\_\{ar\}\)\{=\}\(0\.003,0\.012,0\.004\)\(defaults in Sec\.[VI\-A](https://arxiv.org/html/2608.04590#S6.SS1)\)\. Peak and terminal optima need not coincide\. Deploy\-near\-real: Sec\.[3](https://arxiv.org/html/2608.04590#S6.F3)\.M1M2M3M4ConfigurationPeakTerm\.PeakTerm\.PeakTerm\.PeakTerm\.A\. CTDE observation axis \(deploy\-realistic→\\rightarrowprivileged\)deploy\-near\-real49\.331\.02051\.31989\.0203\.1195\.0228\.0184\.8base local\-ctx54\.852\.02562\.72366\.8211\.8198\.6290\.8282\.2*base**\(ref\.\)*57\.652\.82948\.82648\.5214\.8214\.8302\.8298\.0base\+\+LSTM \(λ=0\.05\\lambda\{=\}0\.05\)51\.147\.82592\.52592\.5210\.4197\.3294\.5281\.8B\. Structure–shaping axis \(fleet / optional HGA\)base−\-UAV \(Nuav=0N\_\{\\mathrm\{uav\}\}\{=\}0\)13\.211\.8815\.5788\.270\.968\.598\.290\.8base\+\+HGA55\.451\.52782\.52438\.8217\.0200\.4315\.9289\.41\) Joint UAV Relaying and Fleet Scale:Fig\.[4](https://arxiv.org/html/2608.04590#S6.F4)shows that removing controllable aerial relays \(base−\-UAV\) yields substantially lower delivery than the CTDE observation\-axis configurations \(M1 peak about 13 versus about 51–58\)\. The main reason is that ground\-only contacts withinrgr=300r\_\{\\mathrm\{gr\}\}\{=\}300m cannot span the road\-constrained Helsinki geometry, whereas UAV links operate atruav=900r\_\{\\mathrm\{uav\}\}\{=\}900m and reshape the contact matrix as formalized in Eq\. \([22](https://arxiv.org/html/2608.04590#S4.E22)\)\. Fig\.[3](https://arxiv.org/html/2608.04590#S6.F3)and Table[V](https://arxiv.org/html/2608.04590#S6.T5)further report a fleet\-size sweepNuav∈\{0,1,2,3,5,8\}N\_\{\\mathrm\{uav\}\}\\in\\\{0,1,2,3,5,8\\\}on M1 under a fixed training recipe \(LSTM on,λ=0\.05\\lambda\{=\}0\.05, HGA on; not the default base\)\. It can be found that the first UAV provides the largest incremental gain and that marginal returns diminish beyondNuav∈\[3,5\]N\_\{\\mathrm\{uav\}\}\{\\in\}\[3,5\], which motivates the default fleet of five relays\. This is because additional relays expand contact opportunities only until coverage overlap and multi\-agent credit assignment begin to offset the geometric benefit under the tested horizon\. Consequently, joint motion–routing control is a structural prerequisite rather than an incremental booster under the evaluated scenario\.

Table V:UAV fleet\-size sensitivity on M1 \(80 epochs, seed 42\)\. Fixed recipe with LSTM auxiliary \(λ=0\.05\\lambda\{=\}0\.05\) and HGA on \(not the default base\); onlyNuavN\_\{\\mathrm\{uav\}\}varies\.NuavN\_\{\\mathrm\{uav\}\}PeakDtestD\_\{\\text\{test\}\}TerminalDtestD\_\{\\text\{test\}\}EtestE\_\{\\text\{test\}\}at PeakDtestD\_\{\\text\{test\}\}Nuav=0N\_\{\\mathrm\{uav\}\}\{=\}012\.811\.756\.2Nuav=1N\_\{\\mathrm\{uav\}\}\{=\}145\.745\.723\.3Nuav=2N\_\{\\mathrm\{uav\}\}\{=\}247\.739\.821\.3Nuav=3N\_\{\\mathrm\{uav\}\}\{=\}349\.844\.719\.2Nuav=5N\_\{\\mathrm\{uav\}\}\{=\}5\(default\)57\.153\.311\.9Nuav=8N\_\{\\mathrm\{uav\}\}\{=\}860\.958\.28\.1![Refer to caption](https://arxiv.org/html/2608.04590v1/figures/fig_uav_num_sweep_delivered.png)Figure 3:M1 UAV fleet\-size sweep on Helsinki\-medium \(N=70N\{=\}70,Nuav∈\{0,1,2,3,5,8\}N\_\{\\mathrm\{uav\}\}\\in\\\{0,1,2,3,5,8\\\}\): mean held\-outDtestD\_\{\\text\{test\}\}vs\. training epoch with LSTM on \(λ=0\.05\\lambda\{=\}0\.05\) and HGA on \(fixed recipe; not the default base\)\. 80 epochs, seed 42\.2\) CTDE Information Factorization:Fig\.[4](https://arxiv.org/html/2608.04590#S6.F4)also compares CTDE observation\-axis configurations that vary UAV\-actor privileged information fromdeploy\-near\-real\(contact\-limitedguL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}, eight\-sector𝐳u\(t\)\\mathbf\{z\}\_\{u\}^\{\\,\(t\)\}, LSTM and HGA disabled\) throughbase local\-ctxto*base*, with optionalbase\+\+LSTM\(λ=0\.05\\lambda\{=\}0\.05\)\. In the different traffic\-mode environments of Table[IV](https://arxiv.org/html/2608.04590#S6.T4), it can be found that deploy\-near\-real trails*base*and base local\-ctx in every mode, with the largest gap under bursty M4, whereas contact\-limited context alone preserves most of the base\-stack benefit\. In addition, enabling the LSTM auxiliary \(base\+\+LSTM\) does not improve over base: on M1 the peak falls from 57\.6 to 51\.1, and under heavy M2 load from 2948\.8 to 2592\.5\. This is because noisy hotspot targets can conflict with sparse delivery rewards underλ=0\.05\\lambda\{=\}0\.05, so that exclusive reliance on the team rewardr\(t\)r^\{\\,\(t\)\}is more stable than joint optimization with multi\-horizon supervision\. It is also observed that CTDE training can exploit global statisticss\(t\)s^\{\\,\(t\)\}and optionalguG​\(t\)g\_\{u\}^\{\\mathrm\{G\}\\,\(t\)\}, yet decentralized execution remains competitive when actor inputs stay near contact\-limited observations \(guL​\(t\)g\_\{u\}^\{\\mathrm\{L\}\\,\(t\)\}\), which is consistent with the Local / Local\-ctx scopes in Table[II](https://arxiv.org/html/2608.04590#S3.T2), Part III\.

![Refer to caption](https://arxiv.org/html/2608.04590v1/figures/fig_node1_stageA_delivered.png)Figure 4:Helsinki\-medium M1 \(N=70N\{=\}70,Nuav=5N\_\{\\mathrm\{uav\}\}\{=\}5; 69 messages att=0t\{=\}0\): mean held\-outDtestD\_\{\\text\{test\}\}vs\. training epoch for base−\-UAV and CTDE observation\-axis rows base local\-ctx,*base*, and base\+\+LSTM\. 80 epochs, seed 42\.3\) Optional Hotspot Auxiliary Learning and HGA:The structure–shaping axis isolates optional modules that are disabled in the default base stack\. Comparing*base*with base\+\+HGA in Table[IV](https://arxiv.org/html/2608.04590#S6.T4)shows that enabling HGA alone can raise peak delivery on M3–M4 \(217\.0 and 315\.9\), whereas on M1–M2*base*remains strongest and terminal delivery still favors*base*on most modes\. Together with the CTDE\-axis finding that base\+\+LSTM underperforms*base*, these results indicate that optional hotspot modules are traffic\-dependent extras rather than default ingredients\. Therefore, when delivery transfer from auxiliary supervision is uncertain, keepingλ=0\\lambda\{=\}0and HGA off \(default base\) is more reliable than enabling optional shaping by default\.

4\) Cross\-Traffic Learning Dynamics:Fig\.[5](https://arxiv.org/html/2608.04590#S6.F5)shows learning curves for the optional base\+\+LSTM stack across traffic modes M1–M4, and the detailed peak and terminal values for the core ablation rows are reported in Table[IV](https://arxiv.org/html/2608.04590#S6.T4)\. In the different traffic environments, it can be found that the same patterns persist under heavier injection: base−\-UAV collapses in every mode; base\+\+LSTM underperforms base on sustained M2 load; and on M3–M4, base\+\+HGA achieves the highest peak delivery while terminal delivery favors base\. This is mainly due to the interaction between traffic\-dependent reward sparsity and optional shaping terms, which can elevate mid\-training peaks without guaranteeing terminal optimality under a fixed hyperparameter schedule\.

![Refer to caption](https://arxiv.org/html/2608.04590v1/figures/fig_full_delivered_four_traffic.png)Figure 5:Optional base\+\+LSTM configuration \(λ=0\.05\\lambda\{=\}0\.05, HGA off\) on Helsinki\-medium \(N=70N\{=\}70,Nuav=5N\_\{\\mathrm\{uav\}\}\{=\}5\): mean held\-outDtestD\_\{\\text\{test\}\}vs\. training epoch on M1–M4\. 80 epochs, seed 42\. Default base usesλ=0\\lambda\{=\}0\(Table[IV](https://arxiv.org/html/2608.04590#S6.T4)\)\.
### VI\-CBaseline Routing Comparison

To underscore the effectiveness of opportunistic forwarding under the same discrete\-time simulator, we compare five routing\-only baseline methods that select one transfer per source node from the candidate sets of Sec\.[IV](https://arxiv.org/html/2608.04590#S4), namely:

1. 1\.PRoPHET:The PRoPHET\-based routing algorithm predicts a node’s future contacts from encounter history and transitivity and formulates forwarding decisions accordingly\[[2](https://arxiv.org/html/2608.04590#bib.bib2)\]\. The fundamental principle of the algorithm is that if two nodes have often encountered each other in the past, then the probability of their encounter in the future is also high\. Our implementation realizes aging, encounter record updates, transitive delivery predictability, and threshold\-based forwarding, but omits full implementation of the wireless transceiver protocol stack native to The ONE simulator\[[29](https://arxiv.org/html/2608.04590#bib.bib29)\]\.
2. 2\.MaxProp:The MaxProp\-based routing algorithm prioritizes transfers that improve delivery prospects by combining destination meeting recency, hop\-count, TTL urgency, and destination delivery likelihood under buffer pressure\[[18](https://arxiv.org/html/2608.04590#bib.bib18)\]\. Our implementation scores candidates from recency, TTL, and destination meeting recency weights, while omitting full The ONE queue management\[[29](https://arxiv.org/html/2608.04590#bib.bib29)\]\.
3. 3\.Fan DPUVR:The Fan DPUVR algorithm ranks relay candidates by a multiplicative utility function over trajectory similarity, surplus energy, link survival, remaining distance, and queuing delay, with adaptive attribute weights\[[4](https://arxiv.org/html/2608.04590#bib.bib4)\]\. We implement the core utility for the five attributes and omit Dijkstra reference trajectories and three\-level priority queues\.
4. 4\.ICC Q\-learning:The ICC Q\-learning algorithm formulates forwarding as tabular Q\-learning over\(src,relay,dest\)\(\\text\{src\},\\text\{relay\},\\text\{dest\}\)with delivery\-coupled rewards andϵ\\epsilon\-greedy exploration\[[10](https://arxiv.org/html/2608.04590#bib.bib10)\]\. Our adapted implementation employs per\-step online updates, uses buffer occupancy as proxy metrics for residual node energy, and enforces episodic Q\-table reset\.
5. 5\.ICC FQLRP:The ICC FQLRP algorithm extends the Q\-learning core with a fuzzy classification layer before action selection\[[10](https://arxiv.org/html/2608.04590#bib.bib10)\]\.

All five baseline protocols are partially reimplemented on our discrete\-time SCF simulator\. We preserve their core scoring and learning mechanisms documented in original papers, while omitting The ONE’s native transceiver stack\[[29](https://arxiv.org/html/2608.04590#bib.bib29)\]and substituting simplified proxy metrics as described earlier\. We evaluate all standalone routing baselines under a fixed random seed of 42\. Each table entry corresponds to one complete simulation trial, which records three metrics: the count of successfully delivered packets D, the count of expired packets E, and delivery ratioηdel=D/Mcreated\\eta\_\{\\mathrm\{del\}\}=D/M\_\{\\text\{created\}\}\. HereMcreatedM\_\{\\text\{created\}\}stands for the total number of packets injected within one full episode\. For traffic modes M1 to M4, the total injected packets per simulation trial are 69, 4321, 306 and 345, respectively\.

Table[VI](https://arxiv.org/html/2608.04590#S6.T6)summarizes the resulting delivery metrics\. Rows marked†list selected JUROR peak held\-out deliveries after 80\-epoch training \(seed 42\)—*base\+\+LSTM*\(λ=0\.05\\lambda\{=\}0\.05\),*base−\-UAV*,Nuav∈\{1,2\}N\_\{\\mathrm\{uav\}\}\{\\in\}\\\{1,2\\\}, and*deploy\-near\-real*—for qualitative reference only and are not directly comparable to the routing\-only protocol; default*base*and*base\+\+HGA*peaks remain in Table[IV](https://arxiv.org/html/2608.04590#S6.T4)\.1\) Delivery Ratio AnalysisNo single baseline outperforms others on all traffic modes\. PRoPHET performs best under regular ground mobility, yet all baselines see severe delivery degradation under the sustained congestion of M2, as encounter\-based scoring cannot resolve buffer saturation\. By jointly optimizing UAV trajectories and forwarding rules, JUROR effectively relieves such congestion limitations and achieves more stable delivery performance\.2\) Congestion\-Aware & Learning\-Based SchemesMaxProp and Fan DPUVR gain marginal benefits under multi\-source congestion via multi\-dimensional scoring\. The two tabular Q\-learning methods struggle with continuous traffic due to periodic Q\-table resets\. Unlike all routing\-only baselines, JUROR coordinates UAV movement to expand contact opportunities, delivering consistently higher delivery ratios across all four traffic scenarios\.

Table VI:Routing\-only baselines and selected JUROR peak references \(†\) on Helsinki\-medium under M1–M4 \(Table[III](https://arxiv.org/html/2608.04590#S6.T3)\)\.ModeMethodDel\.Exp\.ηdel\\eta\_\{\\mathrm\{del\}\}M1PRoPHET34\.934\.10\.505MaxProp13\.455\.60\.194Fan DPUVR12\.057\.00\.174ICC Q\-learning17\.052\.00\.246ICC FQLRP29\.040\.00\.420base \+LSTM†51\.1017\.90\.741base−\-UAV†13\.2055\.80\.191Nuav=1†N\_\{\\mathrm\{uav\}\}\{=\}1^\{\\dagger\}45\.7023\.30\.662Nuav=2†N\_\{\\mathrm\{uav\}\}\{=\}2^\{\\dagger\}47\.7021\.30\.691deploy\-near\-real†49\.3319\.70\.715M2PRoPHET1611\.0371\.00\.373MaxProp911\.0553\.00\.211Fan DPUVR944\.0516\.00\.218ICC Q\-learning706\.0593\.00\.163ICC FQLRP724\.0639\.00\.168base \+LSTM†2592\.50171\.30\.600base−\-UAV†815\.50642\.70\.189Nuav=1†N\_\{\\mathrm\{uav\}\}\{=\}1^\{\\dagger\}1601\.67383\.20\.371Nuav=2†N\_\{\\mathrm\{uav\}\}\{=\}2^\{\\dagger\}2014\.67280\.30\.466deploy\-near\-real†2051\.33315\.30\.475
ModeMethodDel\.Exp\.ηdel\\eta\_\{\\mathrm\{del\}\}M3PRoPHET157\.017\.00\.513MaxProp92\.030\.00\.301Fan DPUVR73\.033\.00\.239ICC Q\-learning65\.043\.00\.212ICC FQLRP58\.047\.00\.190base \+LSTM†210\.4011\.80\.688base−\-UAV†70\.9043\.40\.232Nuav=1†N\_\{\\mathrm\{uav\}\}\{=\}1^\{\\dagger\}133\.0027\.50\.435Nuav=2†N\_\{\\mathrm\{uav\}\}\{=\}2^\{\\dagger\}183\.5012\.80\.600deploy\-near\-real†203\.1013\.30\.664M4PRoPHET207\.0138\.00\.600MaxProp104\.0241\.00\.301Fan DPUVR107\.0238\.00\.310ICC Q\-learning85\.0260\.00\.246ICC FQLRP140\.0205\.00\.406base \+LSTM†294\.5050\.50\.854base−\-UAV†98\.25246\.80\.285Nuav=1†N\_\{\\mathrm\{uav\}\}\{=\}1^\{\\dagger\}106\.62238\.40\.309Nuav=2†N\_\{\\mathrm\{uav\}\}\{=\}2^\{\\dagger\}299\.6245\.40\.868deploy\-near\-real†228\.00117\.00\.661

### VI\-DMechanism Analysis and Discussions

As analyzed above, the ablation and baseline results admit a mechanism\-level reading along two dimensions\.

1\) Baseline Performance Under Traffic Loads:PRoPHET outperforms peers with regular node encounters, yet all routing\-only baselines degrade severely under persistent congestion as full buffers block new packet copies\. MaxProp and Fan DPUVR yield minor gains via queue\-aware scoring under multi\-copy competition, while episodic tabular RL becomes unstable with heavy traffic injection\. Unlike these schemes that only reorder existing contacts, JUROR co\-optimizes forwarding and UAV mobility to proactively create new relay opportunities\. Treating UAV movement as a controllable decision variable, our method improves delivery via joint relay scheduling and long\-term topology reshaping\. The sharp performance decline in the UAV\-free ablation case, and consistent recovery with aerial relays enabled, verifies joint routing\-UAV optimization as the core source of JUROR’s performance gains\.

2\) CTDE and Auxiliary Effects AnalysisCTDE improves training stability while preserving decentralized execution, and the deploy\-near\-real results show that JUROR retains most of its benefit even with contact\-limited actor inputs\. Optional LSTM and HGA are traffic\-dependent refinements rather than mandatory components, so the default configuration remains the most reliable general setting unless mode\-specific retuning is available\.

## VIIConclusion

In this paper, we investigate core bottlenecks restricting joint routing and UAV motion control in delay\-tolerant aerial networks, in attaining reliable message delivery while supporting adaptive congestion mitigation and dynamic topology optimization\. To tackle these limitations, a CTDE multi\-agent learning architecture and a joint flight\-forwarding control framework built upon PPO cooperative optimization are developed\.

First, the multi\-agent decentralized execution paradigm is integrated into the DTN SCF system and deployed over the UAV relay fleet to expand intermittent contact opportunities\. Second, multiple network observation features are fused to capture the time\-varying contact evolution of heterogeneous ground and aerial nodes\. CTDE training rules are then leveraged to model the partially observable sequential decision process of all network agents, constructing the shared team reward function and converting the joint topology shaping task into a cooperative reinforcement learning optimization problem\. Finally, to accommodate the intermittent connectivity property inherent to DTN environments, a joint UAV\-routing decision algorithm is formulated\. Each relay agent builds individual policy output conditioned on local congestion and neighbor observations to produce adaptive movement and forwarding actions, realizing stable end\-to\-end data delivery and supporting global collaborative resource allocation across the entire UAV fleet\.

Simulation results validate that the proposed JUROR framework delivers superior overall performance, and its advantage over baseline routing algorithms expands significantly under heavier traffic loads\.

## References

- \[1\]Badawi, Salma, Ahmad, Norulhusna, Dziyauddin, Rudzidatul Akmam, Mohamed, Norliza, and Sam, Suriani Mohd, “Routing Protocols in FANET for Disaster Area Networks: A Review,”*ASEAN Engineering Journal*, vol\. 15, no\. 3, pp\. 81–100, 2025, doi: 10\.11113/aej\.V15\.22954\.
- \[2\]Lindgren, Anders, Doria, Avri, and Schelén, Olov, “Probabilistic Routing in Intermittently Connected Networks,”*ACM SIGMOBILE Mobile Computing and Communications Review*, vol\. 7, no\. 3, pp\. 19–26, 2003, doi: 10\.1145/961268\.961272\.
- \[3\]Du, Zhaoyang, Wu, Celimuge, Yoshinaga, Tsutomu, Chen, Xianfu, Wang, Xiaoyan, Yau, Kok\-Lim Alvin, and Ji, Yusheng, “A Routing Protocol for UAV\-Assisted Vehicular Delay Tolerant Networks,”*IEEE Open Journal of the Computer Society*, vol\. 2, pp\. 85–98, 2021, doi: 10\.1109/OJCS\.2021\.3054759\.
- \[4\]Fan, Zhijie, Zhang, Mansi, Cao, Yue, Liu, Zilong, Kaiwartya, Omprakash, Javed, Yasir, and Hussain, Faisal Bashir, “A Novel UAV\-assisted VANET Routing Protocol for Post\-Disaster Emergency Communications,”*IEEE Transactions on Network Science and Engineering*, vol\. 13, pp\. 4863–4882, 2026, doi: 10\.1109/TNSE\.2025\.3644432\.
- \[5\]Mammeri, Zoubir, “Reinforcement Learning Based Routing in Networks: Review and Classification of Approaches,”*IEEE Access*, vol\. 7, pp\. 55916–55950, 2019, doi: 10\.1109/ACCESS\.2019\.2913776\.
- \[6\]Fernandes, Inês and Pereira, Paulo Rogério, “Social and Geographical Routing for Vehicular Delay\-Tolerant Networks,”*Proc\. Int\. Young Engineers Forum Electr\. Comput\. Eng\. \(YEF\-ECE\)*, 2025, doi: 10\.1109/YEF\-ECE66503\.2025\.11117524\.
- \[7\]Sammou, El Mastapha, “PF\-DTN: Predictive Routing for Intelligent Delay Tolerant Networks Using RNN\-LSTM Deep Learning With Monte Carlo Dropout Uncertainty Estimation and Hybrid Deterministic, Probabilistic, and Uncertain Routing Strategies,”*IEEE Access*, vol\. 14, pp\. 10841–10859, 2026, doi: 10\.1109/ACCESS\.2026\.3655507\.
- \[8\]Wen, Xin and Tan, Long, “Enhancing DTN Routing Strategies with Deep Reinforcement Learning in Disaster Recovery Networks,”*Proc\. Int\. Conf\. Frontier Technol\. Inf\. Comput\. \(ICFTIC\)*, pp\. 450–456, 2024, doi: 10\.1109/ICFTIC64248\.2024\.10913424\.
- \[9\]Chakrabarti, Chandrima, “A Noble Route Discovery Technique in UAV Based Delay Tolerant Network for Disaster Management,”*Proc\. Int\. Conf\. Res\. Methodol\. Knowl\. Manag\., Artif\. Intell\. Telecommun\. Eng\. \(RMKMATE\)*, 2025, doi: 10\.1109/RMKMATE64874\.2025\.11042440\.
- \[10\]Dhurandher, Sanjay K\., Singh, Jagdeep, Woungang, Isaac, Srivastava, Sanjeev, and Rodrigues, Joel J\. P\. C\., “Reinforcement Learning\-Based Routing Protocol for Opportunistic Networks,”*Proc\. IEEE Int\. Conf\. Commun\. \(ICC\)*, 2020, doi: 10\.1109/ICC40277\.2020\.9149039\.
- \[11\]Yao, Lei, Bai, Xiangyu, and Zhou, Kexin, “An Improved Spray and Wait Algorithm Based on Q\-learning in Delay Tolerant Network,”*Proc\. Int\. Joint Conf\. Neural Netw\. \(IJCNN\)*, pp\. 1–8, 2024, doi: 10\.1109/IJCNN60899\.2024\.10650780\.
- \[12\]Xiang, Yuhong, Gao, Shuai, Wang, Hongchao, Yang, Dong, Zhang, Yuming, and Zhang, Hongke, “Performance Evaluation for Q\-Learning Based Anycast Routing Protocol in Unmanned Aerial Vehicle Networks with Multiple Base Stations,”*Ad Hoc Networks*, vol\. 168, pp\. 103719, 2025, doi: 10\.1016/j\.adhoc\.2024\.103719\.
- \[13\]Fan, Zesong, Gong, Shimin, Long, Yusi, Li, Lanhua, Gu, Bo, and Luong, Nguyen Cong, “Delay\-Tolerant Multi\-Agent DRL for Trajectory Planning and Transmission Control in UAV\-Assisted Wireless Networks,”*Proc\. IEEE Veh\. Technol\. Conf\. \(VTC Spring\)*, pp\. 1–5, 2024, doi: 10\.1109/VTC2024\-Spring59600\.2024\.10555325\.
- \[14\]Lowe, Ryan, Wu, Yi, Tamar, Aviv, Harb, Jean, Abbeel, Pieter, and Mordatch, Igor, “Multi\-Agent Actor\-Critic for Mixed Cooperative\-Competitive Environments,”*Advances in Neural Information Processing Systems*, vol\. 30, 2017\.
- \[15\]Foerster, Jakob N\., Farquhar, Gregory, Afouras, Triantafyllos, Nardelli, Nantas, and Whiteson, Shimon, “Counterfactual Multi\-Agent Policy Gradients,”*Proceedings of the AAAI Conference on Artificial Intelligence*, vol\. 32, no\. 1, 2018\.
- \[16\]Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg, “Proximal Policy Optimization Algorithms,”*arXiv preprint arXiv:1707\.06347*, 2017\.
- \[17\]Spyropoulos, Thrasyvoulos, Psounis, Konstantinos, and Raghavendra, C\. S\., “Routing for Disruption Tolerant Networks: Taxonomy and Design,”*Wireless Communications*, vol\. 16, no\. 6, pp\. 39–46, 2009, doi: 10\.1109/MWC\.2009\.5302297\.
- \[18\]Burgess, John, Gallagher, Brian, Jensen, David, and Levine, Brian Neil, “MaxProp: Routing for Vehicle\-Based Disruption\-Tolerant Networks,”*Proc\. IEEE INFOCOM*, 2006, doi: 10\.1109/INFOCOM\.2006\.228\.
- \[19\]Zhang, Yu, “Epidemic Routing Optimization Algorithm Based on Node Cache Status,”*EBNCS flooding\-family variant with cache\-state classification and ACK cleanup; author manuscript*, 2025\.
- \[20\]Ullah, Saif, Muhammad, Asif, Ali, Zulfiqar, Waqar, Muhammad, and Kim, Ajung, “SR\-SAAD: A Social Rank\-Based Routing Protocol for Enhanced Efficiency in Delay Tolerant Networks,”*IET Communications*, vol\. 19, no\. 1, pp\. e70099, 2025, doi: 10\.1049/cmu2\.70099\.
- \[21\]D’Argenio, Pedro R\., Fraire, Juan, Hartmanns, Arnd, and Raverta, Fernando, “Comparing Statistical, Analytical, and Learning\-Based Routing Approaches for Delay\-Tolerant Networks,”*ACM Transactions on Modeling and Computer Simulation*, vol\. 35, no\. 2, pp\. 1–26, 2025, doi: 10\.1145/3665927\.
- \[22\]Salam, Mohammad Abdus, Saif, A\. F\. M\. Saifuddin, Katroju, Priya Hamsa, and Kassouf\-Short, Robert, “A Constructive Analysis on Machine Learning Integration in High Delay Tolerant Networking \(HDTN\),”*Proc\. Int\. Conf\. Adv\. Commun\. Technol\. Netw\. \(CommNet\)*, pp\. 1–6, 2024, doi: 10\.1109/CommNet63022\.2024\.10793376\.
- \[23\]Zhou, Xuan, Lin, Haitao, and Chen, Jin, “A Review of Research on Routing Protocols for Unmanned Aerial Vehicle Cluster Self\-Organizing Networks,”*Proc\. Int\. Conf\. Intell\. Syst\., Commun\. Comput\. Netw\. \(ISCCN\)*, pp\. 246–255, 2025, doi: 10\.1145/3732945\.3732981\.
- \[24\]Zhou, Xixuan, Xu, Ruoyu, Tian, Xiaojian, Zhang, Yueyue, Liang, Yu, Chen, Xiaoliang, and Zhu, Zuqing, “Distributed Routing and Data Scheduling in IPNs With GNN\-Based Multiagent DRL,”*IEEE Internet of Things Journal*, vol\. 12, no\. 12, pp\. 21565–21576, 2025, doi: 10\.1109/JIOT\.2025\.3547341\.
- \[25\]Tian, Xiaojian, Chen, Xiaoliang, Zhou, Xixuan, and Zhu, Zuqing, “On the Scaling of Reliable Interplanetary Networks with Deep Reinforcement Learning,”*Proc\. Int\. Conf\. Design Rel\. Commun\. Netw\. \(DRCN\)*, 2025, doi: 10\.1109/DRCN65040\.2025\.11046163\.
- \[26\]Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z\., Silver, David, and Kavukcuoglu, Koray, “Reinforcement Learning with Unsupervised Auxiliary Tasks,”*International Conference on Learning Representations*, 2017\.
- \[27\]Brockman, Greg, Cheung, Vicki, Petrov, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech, “OpenAI Gym,”*arXiv preprint arXiv:1606\.01540*, 2016\.
- \[28\]Weng, Jiayi, Chen, Huan, Yan, Dong, You, Kaichao, Duburcq, Antoine, Zhang, Minghao, Su, Hang, and Zhu, Jun, “Tianshou: A Highly Modularized Deep Reinforcement Learning Library,”*Journal of Machine Learning Research*, vol\. 22, no\. 267, pp\. 1–6, 2021\.
- \[29\]Keränen, Ari, Ott, Jörg, and Kärkkäinen, Teemu, “The ONE Simulator for DTN Protocol Evaluation,”*Proceedings of the 2nd International Conference on Simulation Tools and Techniques*, pp\. 1–10, 2009\.

Similar Articles