Posture and Sustainment Optimization Under Adversarial Uncertainty

arXiv cs.AI Papers

Summary

This paper introduces a scenario-weighted adversarially robust posture optimization engine for military asset allocation, proposing CEV and RobustCEV optimizers that outperform greedy baselines under adversarial threat uncertainty.

arXiv:2608.05256v1 Announce Type: new Abstract: Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational planning. Current practice relies on greedy heuristics that maximize value and ignore geographic coverage and are structurally vulnerable to adversaries that target high-strategic value locations. This paper presents a scenario-weighted adversarially robust posture optimization engine for the Posture and sustainability allocation (PSA) problem, modeled as a finite-horizon Markov Decision Process over assets, theater locations, and time steps. We introduce the Composite Expected Value (CEV) optimizer, which places assets by maximizing scenario-weighted expected posture efficiency over a distribution of threat scenarios, and the RobustCEV extension, which iterates against a Bayesian adversary that updates its targeting distribution in response to observed placement. Across three experiments in an Indo-Pacific basing environment with 20 assets and 5 theater locations, we demonstrate that: (1) the greedy baseline incurs a permanent 25.1% posture efficiency penalty due to geographic under-coverage and a 57.3% scenario-weighted readiness collapse under value-correlated adversarial threat; (2) the CEV optimizer recovers up to 19.8% efficiency over greedy when the threat distribution carries a geographic signal, with a curated set of 5 to 20 scenarios sufficient to capture the majority of this gain; and (3) the RobustCEV extension recovers up to 158% efficiency relative to a naive optimizer when an adaptive adversary employs a deceptive threat prior. All findings are validated using paired t-tests with Bonferroni correction and two-level variance decomposition, confirming that the performance gaps reported are structural properties of placement strategies rather than sampling artifacts.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:46 AM

# Posture and Sustainment Optimization Under Adversarial Uncertainty
Source: [https://arxiv.org/html/2608.05256](https://arxiv.org/html/2608.05256)
###### Abstract

Pre\-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational planning\. Current practice relies on greedy heuristics that maximize value and ignore geographic coverage and are structurally vulnerable to adversaries that target high\-strategic value locations\. This paper presents a scenario\-weighted adversarially robust posture optimization engine for the Posture and sustainability allocation \(PSA\) problem, modeled as a finite\-horizon Markov Decision Process over assets, theater locations, and time steps\. We introduce the Composite Expected Value \(CEV\) optimizer, which places assets by maximizing scenario\-weighted expected posture efficiency over a distribution of threat scenarios, and the RobustCEV extension, which iterates against a Bayesian adversary that updates its targeting distribution in response to observed placement\. Across three experiments in an Indo\-Pacific basing environment with 20 assets and 5 theater locations, we demonstrate that: \(1\) the greedy baseline incurs a permanent 25\.1% posture efficiency penalty due to geographic under\-coverage and a 57\.3% scenario\-weighted readiness collapse under value\-correlated adversarial threat; \(2\) the CEV optimizer recovers up to 19\.8% efficiency over greedy when the threat distribution carries a geographic signal, with a curated set of 5 to 20 scenarios sufficient to capture the majority of this gain; and \(3\) the RobustCEV extension recovers up to 158% efficiency relative to a naive optimizer when an adaptive adversary employs a deceptive threat prior\. All findings are validated using pairedtt\-tests with Bonferroni correction and two\-level variance decomposition, confirming that the performance gaps reported are structural properties of placement strategies rather than sampling artifacts\.

###### Contents

1. [1Introduction](https://arxiv.org/html/2608.05256#S1)
2. [2Problem Formulation](https://arxiv.org/html/2608.05256#S2)1. [2\.1MDP Definition](https://arxiv.org/html/2608.05256#S2.SS1) 2. [2\.2Distinction from Related Decision Problems](https://arxiv.org/html/2608.05256#S2.SS2)
3. [3Related Work](https://arxiv.org/html/2608.05256#S3)
4. [4Technical Approach](https://arxiv.org/html/2608.05256#S4)1. [4\.1Scenario Engine](https://arxiv.org/html/2608.05256#S4.SS1) 2. [4\.2Core Optimizer](https://arxiv.org/html/2608.05256#S4.SS2) 3. [4\.3Alert Tier and Configuration Mode Representation](https://arxiv.org/html/2608.05256#S4.SS3) 4. [4\.4Human\-Machine Teaming Interface](https://arxiv.org/html/2608.05256#S4.SS4)
5. [5Experiment 1: Greedy Baseline Characterization](https://arxiv.org/html/2608.05256#S5)1. [5\.1Motivation](https://arxiv.org/html/2608.05256#S5.SS1) 2. [5\.2Experimental Setup](https://arxiv.org/html/2608.05256#S5.SS2) 3. [5\.3Greedy Placement Algorithm](https://arxiv.org/html/2608.05256#S5.SS3) 4. [5\.4Performance Metrics](https://arxiv.org/html/2608.05256#S5.SS4) 5. [5\.5Results](https://arxiv.org/html/2608.05256#S5.SS5)1. [5\.5\.1Per\-Step Metrics](https://arxiv.org/html/2608.05256#S5.SS5.SSS1) 2. [5\.5\.2Comparison with Random Placement](https://arxiv.org/html/2608.05256#S5.SS5.SSS2) 3. [5\.5\.3Threat\-Environment Sensitivity](https://arxiv.org/html/2608.05256#S5.SS5.SSS3) 6. [5\.6Discussion](https://arxiv.org/html/2608.05256#S5.SS6)
6. [6Experiment 2: Scenario\-Weighted Optimizer vs\. Greedy Baseline](https://arxiv.org/html/2608.05256#S6)1. [6\.1Motivation](https://arxiv.org/html/2608.05256#S6.SS1) 2. [6\.2Experimental Setup](https://arxiv.org/html/2608.05256#S6.SS2) 3. [6\.3CEV Optimizer](https://arxiv.org/html/2608.05256#S6.SS3) 4. [6\.4Results](https://arxiv.org/html/2608.05256#S6.SS4) 5. [6\.5Discussion](https://arxiv.org/html/2608.05256#S6.SS5)
7. [7Experiment 3: Adversarial Robustness and the Bayesian Counter\-Move Model](https://arxiv.org/html/2608.05256#S7)1. [7\.1Motivation](https://arxiv.org/html/2608.05256#S7.SS1) 2. [7\.2Adversarial Model](https://arxiv.org/html/2608.05256#S7.SS2) 3. [7\.3Experimental Setup](https://arxiv.org/html/2608.05256#S7.SS3) 4. [7\.4Results](https://arxiv.org/html/2608.05256#S7.SS4)1. [7\.4\.1Deceptive Prior: the Core Robustness Result](https://arxiv.org/html/2608.05256#S7.SS4.SSS1) 2. [7\.4\.2Null Results Under Uniform and Skewed Priors](https://arxiv.org/html/2608.05256#S7.SS4.SSS2) 3. [7\.4\.3Adversarial Prior: Partial Convergence Failure](https://arxiv.org/html/2608.05256#S7.SS4.SSS3) 5. [7\.5Discussion](https://arxiv.org/html/2608.05256#S7.SS5)
8. [8Experiment 4: Computational Feasibility and the Robustness\-Cost Pareto Frontier](https://arxiv.org/html/2608.05256#S8)1. [8\.1Motivation](https://arxiv.org/html/2608.05256#S8.SS1) 2. [8\.2Part A: Computational Scalability](https://arxiv.org/html/2608.05256#S8.SS2)1. [8\.2\.1Setup](https://arxiv.org/html/2608.05256#S8.SS2.SSS1) 2. [8\.2\.2Results](https://arxiv.org/html/2608.05256#S8.SS2.SSS2) 3. [8\.3Part B: Robustness\-Cost Pareto Frontier](https://arxiv.org/html/2608.05256#S8.SS3)1. [8\.3\.1Setup](https://arxiv.org/html/2608.05256#S8.SS3.SSS1) 2. [8\.3\.2Results](https://arxiv.org/html/2608.05256#S8.SS3.SSS2) 4. [8\.4Discussion](https://arxiv.org/html/2608.05256#S8.SS4)
9. [9Experiment 5: Dual\-Theater Case Studies with Sensitivity Analysis](https://arxiv.org/html/2608.05256#S9)1. [9\.1Motivation](https://arxiv.org/html/2608.05256#S9.SS1) 2. [9\.2Experimental Setup](https://arxiv.org/html/2608.05256#S9.SS2) 3. [9\.3Part A: Cross\-Theater Baseline Comparison](https://arxiv.org/html/2608.05256#S9.SS3) 4. [9\.4Part B: Sensitivity Analysis](https://arxiv.org/html/2608.05256#S9.SS4)1. [9\.4\.1Wasserstein Radius Sweep](https://arxiv.org/html/2608.05256#S9.SS4.SSS1) 2. [9\.4\.2Assignment Stability under Scenario Weight Perturbation](https://arxiv.org/html/2608.05256#S9.SS4.SSS2) 5. [9\.5Part C: A2/AD Contested Zone Radius Sweep](https://arxiv.org/html/2608.05256#S9.SS5) 6. [9\.6Discussion](https://arxiv.org/html/2608.05256#S9.SS6)
10. [10Statistical Analysis](https://arxiv.org/html/2608.05256#S10)
11. [11Discussion](https://arxiv.org/html/2608.05256#S11)1. [11\.1A Progressive Case Against Greedy Planning](https://arxiv.org/html/2608.05256#S11.SS1) 2. [11\.2Connecting Experimental Results to the OE Axioms](https://arxiv.org/html/2608.05256#S11.SS2) 3. [11\.3Practical Implications for Scenario Library Design](https://arxiv.org/html/2608.05256#S11.SS3) 4. [11\.4Scope and Limitations of the Robustness Results](https://arxiv.org/html/2608.05256#S11.SS4) 5. [11\.5Relationship to Prior Work](https://arxiv.org/html/2608.05256#S11.SS5) 6. [11\.6Future Directions](https://arxiv.org/html/2608.05256#S11.SS6)
12. [References](https://arxiv.org/html/2608.05256#bib)

## 1Introduction

Modern joint operations require that assets be positioned*before*threats materialize\. Pre\-commitment posture: the assignment of aircraft, munitions, logistics, and support assets to theater locations days or weeks ahead of potential conflict determines what options a commander has when the situation evolves\. However, current planning practice relies predominantly on rule\-based, value\-maximizing heuristics: place the most assets at the highest\-priority bases\. This approach is computationally tractable and interpretable, but it is strategically brittle\. An adversary that observes a predictable concentration can target it; a planner who ignores threat heterogeneity will be surprised by it\.

The Posture and Sustainment Allocation \(PSA\) problem is formally unsolved at the operational level\. Five axioms of the modern Operational Environment \(OE\) jointly make static and greedy approaches insufficient: \(1\)*threat nonstationarity*: adversary targeting distributions shift in response to observable defender posture; \(2\)*multidomain coupling*: posture in one domain creates vulnerabilities or opportunities in others; \(3\)*placement irreversibility*: initial assignments commit assets that are costly and time\-consuming to reposition under fire; \(4\)*information asymmetry*: the defender has incomplete knowledge of adversary capabilities and intent; and \(5\)*sustainment time\-criticality*: readiness degrades continuously and cannot be restored instantaneously\. Each axiom individually degrades the quality of a greedy or static plan; together, they require a decision\-support engine that is scenario\-aware, adversarially robust, and explicitly models the degradation–repair cycle\.

This paper presents a scenario\-weighted, adversarially\-robust posture optimization engine for the PSA problem and demonstrates its superiority over greedy baselines across three experiments\. Our contributions are:

1. 1\.A formal MDP model of the PSA problem overMMassets,NNtheater locations, andTTtime steps, with a composite posture efficiency metric that jointly rewards readiness, geographic coverage, and cost\-efficiency \(Section[2](https://arxiv.org/html/2608.05256#S2)\)\.
2. 2\.A Composite Expected Value \(CEV\) optimizer that places assets over a distribution of threat scenarios, achieving up to 19\.8% efficiency improvement over the greedy baseline when the threat distribution carries geographic signal \(Experiment 2\)\.
3. 3\.A Bayesian counter\-move model in which an adaptive adversary updates its targeting distribution in response to observed placement, and a RobustCEV extension that converges to a stable equilibrium, recovering up to 158% efficiency relative to a naive optimizer under deceptive threat priors \(Experiment 3\)\.
4. 4\.Statistical validation of all findings via pairedtt\-tests with Bonferroni correction and two\-level variance decomposition, confirming that reported performance gaps are structural rather than sampling artifacts \(Section[10](https://arxiv.org/html/2608.05256#S10)\)\.

The PSA module operates upstream of mission\-assignment scheduling \(SBC\) and sensor\-fusion tasking \(ISU\): it determines where assets are before missions are assigned, establishing the capability envelope within which downstream planners operate\.

## 2Problem Formulation

### 2\.1MDP Definition

We model PSA as a finite\-horizon Markov Decision Processℳ=\(𝒮,𝒰,P,R,T\)\\mathcal\{M\}=\(\\mathcal\{S\},\\mathcal\{U\},P,R,T\)over\|𝒜0\|\|\\mathcal\{A\}\_\{0\}\|assets,\|ℒ\|\|\\mathcal\{L\}\|theater locations, andTTdiscrete time steps\.

##### State space\.

Each asseta∈𝒜0a\\in\\mathcal\{A\}\_\{0\}is described by a tuple\(ra,ℓa,ma,da,qa\)\(r\_\{a\},\\,\\ell\_\{a\},\\,m\_\{a\},\\,d\_\{a\},\\,q\_\{a\}\), wherera∈\[0,1\]r\_\{a\}\\in\[0,1\]is the readiness rate,ℓa∈ℒ\\ell\_\{a\}\\in\\mathcal\{L\}is the current location assignment,ma∈\{Dormant,Active\}m\_\{a\}\\in\\\{\\textsc\{Dormant\},\\textsc\{Active\}\\\}is the configuration mode,da∈ℤ≥0d\_\{a\}\\in\\mathbb\{Z\}\_\{\\geq 0\}is the maintenance timer \(days until next scheduled maintenance\), andqa∈ℤ\>0q\_\{a\}\\in\\mathbb\{Z\}\_\{\>0\}is the asset quantity\. The joint state isst=\{\(rat,ℓat,mat,dat,qat\)\}a∈𝒜0∈𝒮s\_\{t\}=\\\{\(r\_\{a\}^\{t\},\\,\\ell\_\{a\}^\{t\},\\,m\_\{a\}^\{t\},\\,d\_\{a\}^\{t\},\\,q\_\{a\}^\{t\}\)\\\}\_\{a\\in\\mathcal\{A\}\_\{0\}\}\\in\\mathcal\{S\}\.

##### Action space\.

At each time step the planner selects one sustainment action per asset from𝒰=\{Reposition,Resupply,Maintain,Hold\}\\mathcal\{U\}=\\\{\\textsc\{Reposition\},\\,\\textsc\{Resupply\},\\,\\textsc\{Maintain\},\\,\\textsc\{Hold\}\\\}\.Repositionmoves an asset to a new location at costκ=10\\kappa=10;Resupplyincrements quantityqa←qa\+2q\_\{a\}\\leftarrow q\_\{a\}\+2at costκ=5\\kappa=5;Maintainrestores readinessra←min⁡\(1,ra\+0\.20\)r\_\{a\}\\leftarrow\\min\(1,\\,r\_\{a\}\+0\.20\)at costκ=2\\kappa=2;Holdleaves the asset unchanged at zero cost\. Initial placement \(the assignmentℓa0\\ell\_\{a\}^\{0\}for allaa\) is a separate first\-stage decision, as discussed in Section[4](https://arxiv.org/html/2608.05256#S4)\.

##### Transition dynamics\.

At each step, asset readiness degrades stochastically:

rat\+1=max⁡\(0,rat−δat\),δat∼𝒰​\(0,δmax\),r\_\{a\}^\{t\+1\}=\\max\\\!\\left\(0,\\;r\_\{a\}^\{t\}\-\\delta\_\{a\}^\{t\}\\right\),\\qquad\\delta\_\{a\}^\{t\}\\sim\\mathcal\{U\}\(0,\\,\\delta\_\{\\max\}\),\(1\)whereδmax=0\.10\\delta\_\{\\max\}=0\.10unless otherwise noted\. Maintenance timers decrement by one each step and reset to𝒰​\(30,90\)\\mathcal\{U\}\(30,90\)upon aMaintainaction\.

##### Reward\.

The per\-step reward is the*posture efficiency*:

R​\(st,ut\)=Et=Rt⋅Ctlog⁡\(Kt\+2\),R\(s\_\{t\},u\_\{t\}\)\\;=\\;E\_\{t\}\\;=\\;\\frac\{R\_\{t\}\\cdot C\_\{t\}\}\{\\log\(K\_\{t\}\+2\)\},\(2\)whereRtR\_\{t\},CtC\_\{t\}, andKtK\_\{t\}are the readiness score, coverage score, and sustainment cost defined in Equations \([7](https://arxiv.org/html/2608.05256#S5.E7)\)–\([9](https://arxiv.org/html/2608.05256#S5.E9)\)\. The logarithmic denominator penalizes sustainment spending with diminishing marginal sensitivity, reflecting decreasing marginal returns to additional maintenance resources\[[9](https://arxiv.org/html/2608.05256#bib.bib8)\]\.

### 2\.2Distinction from Related Decision Problems

PSA operates at a different timescale and abstraction level than two adjacent planning modules\. The*Scheduler*\(SBC\) assigns already\-placed assets to specific missions over a rolling horizon of hours to days and assumes a fixed posture as input\. The*ISU*sensor\-fusion module updates threat estimates from multi\-INT feeds in near\-real time and carries no asset\-placement authority\. PSA determines the posture that SBC and ISU inherit, and must therefore account for the full distribution of scenarios they will face\. A poor pre\-commitment posture cannot be corrected reactively by downstream planners without prohibitive repositioning cost\[[11](https://arxiv.org/html/2608.05256#bib.bib5)\]\.

## 3Related Work

##### JADC2 doctrine and the AI capability gap\.

Lingel et al\.\[[7](https://arxiv.org/html/2608.05256#bib.bib1)\]provide the authoritative RAND framework for where artificial intelligence fits within Joint All\-Domain Command and Control \(JADC2\) deliberate planning\. Written for the Air Force, the report establishes a taxonomy of AI applications in JADC2 and identifies the pre\-mission planning phase, specifically asset allocation under uncertainty across multi\-domain, multi\-echelon force structures, as the highest\-leverage point for AI decision support\. This is exactly the operating regime of PSA: the planner must commit assets to theater locations before scenarios resolve, on timelines that compress human cognitive bandwidth\. Citing JADC2 doctrine here is not decorative; it situates PSA as a recognized capability gap rather than a synthetic benchmark problem\.

##### What sustainment means and why it must be dynamic\.

The Institute for Defense Analyses defines sustainment as “every part of the logistics ecosystem working together to ensure each platform is ready to perform its mission,” encompassing training, testing, upgrading, and procuring parts\[[5](https://arxiv.org/html/2608.05256#bib.bib2)\]\. This definitional scope directly motivates the four action types in our MDP \(Reposition,Resupply,Maintain,Hold\): each maps to a distinct sustainment function recognized in defense doctrine\. A decision\-support methodology for military asset and resource planning\[[3](https://arxiv.org/html/2608.05256#bib.bib3)\]formalizes the feedback loop that any adequate sustainment model must capture: resource decisions drive sustainment actions, which drive readiness outcomes, which in turn constrain available force options\. Their causal loop diagram framing explains precisely why greedy baselines fail—they optimize point\-in\-time strategic value without modeling the delayed degradation effects that determine whether assets remain operationally available after several time steps\. Our state space \(readiness rates, maintenance timers, asset quantities\) is designed around this causal chain\.

##### Deployment tempo and stochastic degradation\.

Deployment\-to\-dwell \(D2D\) metrics formalize the observation that deployment tempo directly erodes unit readiness and quality of life, and prior work applies stochastic optimization to sustainment scheduling under this constraint\[[4](https://arxiv.org/html/2608.05256#bib.bib4)\]\. This is the empirical grounding for our degradation model: readiness does not stay static between time steps, and the rate of decay is stochastic\. The D2D literature also provides the bridge from deterministic scheduling \(what a greedy planner implicitly assumes\) to MDP framing \(what operational uncertainty actually requires\)\. When degradation is stochastic and maintenance windows are finite, a policy that ignores future states will systematically under\-invest in sustainment at the wrong times\.

##### Two\-stage stochastic programming for military pre\-commitment\.

Salmerón and Apte\[[11](https://arxiv.org/html/2608.05256#bib.bib5)\]establish the canonical two\-stage stochastic programming template for pre\-commitment under uncertainty: Stage 1 positions resources before demand is known; Stage 2 responds once the scenario materializes\. Despite originating in civilian disaster relief, the formulation is mathematically identical to pre\-conflict military posturing, and the authors use military storage facilities explicitly\. Nelson et al\.\[[8](https://arxiv.org/html/2608.05256#bib.bib6)\]apply this structure directly to US Army helicopter allocation, providing the most recent published validation that two\-stage SP is feasible and practically useful for military asset decisions\. Their Stage 1 allocation maps directly to our CEV placement decision; their Stage 2 air\-movement MILP maps to the Battle Manager’s mission assignment downstream of PSA\. The Expected Value of the Stochastic Solution \(EVSS\) we measure in Experiment 2 is the standard metric from this literature for quantifying how much scenario\-awareness is worth relative to deterministic planning\.

##### Approximate dynamic programming for military asset dispatch\.

Rettke et al\.\[[10](https://arxiv.org/html/2608.05256#bib.bib7)\]formulate aerial medical evacuation dispatching as a finite\-horizon MDP and solve it with approximate policy iteration and least\-squares temporal difference learning, achieving a 31% improvement over the greedy nearest\-available dispatching policy on realistic military scenarios\. This result is the benchmark for MDP\-based improvement over greedy in the military asset dispatch literature\. Their state space, asset availability, location, and readiness, is exactly the\(ra,ℓa,da\)\(r\_\{a\},\\ell\_\{a\},d\_\{a\}\)triple in our formulation, establishing direct structural precedent\. Powell\[[9](https://arxiv.org/html/2608.05256#bib.bib8)\]provides the methodological foundation for scaling MDP solutions to theM×N×TM\\times N\\times Tdimensionality that PSA requires; the three curses of dimensionality \(state space, outcome space, action space\) are precisely the curses PSA faces, and piecewise\-linear value function approximation is what makes the CEV objective tractable at operational scale\.

##### Multi\-agent reinforcement learning for logistics and workforce planning\.

Single\-agent approaches hit a fundamental ceiling whenMMassets operate acrossNNlocations with interdependent capacity constraints: the greedy baseline assumes assets are independent, but placement decisions interact through the shared capacityccper location\. Towards a learning behavior model for military logistics using profit\-sharing RL\[[6](https://arxiv.org/html/2608.05256#bib.bib9)\]demonstrates that multi\-agent formulations are necessary for military logistics systems exhibiting experience sharing, cooperative action, and hierarchical control, all present in PSA\. Their critique of single\-agent profit\-sharing directly maps onto why greedy fails when asset interdependencies are non\-trivial\. Multi\-agent RL with long\-term performance objectives for service workforce optimization\[[2](https://arxiv.org/html/2608.05256#bib.bib10)\]is structurally the closest published paper to PSA: swap “service workers” for “military assets” and the problem formulation is nearly identical, covering personnel dispatch, management, and positioning with both heuristic and RL baselines for comparison, a template we follow in our own experimental design\.

##### Deep RL for stochastic inventory control\.

Classical and deep RL inventory control for pharmaceutical supply chains\[[12](https://arxiv.org/html/2608.05256#bib.bib11)\]benchmarks order\-up\-to, projected inventory level, and PPO policies against a human\-driven baseline under perishability, yield uncertainty, and non\-stationary demand\. The setup is directly analogous to our stochastic degradation and replenishment horizon: assets decay, replenishment is uncertain, and demand \(threat\) is non\-stationary\. Their finding that PPO outperforms human\-driven baselines on cost while maintaining service levels is the motivating proof\-of\-concept for applying deep RL to the sustainment action component of PSA, and their benchmarking structure, rule\-based→\\toclassical optimization→\\toDRL, is the template for how we report our own experimental results\.

## 4Technical Approach

### 4\.1Scenario Engine

Threat uncertainty is represented as a finite scenario set𝒮=\{\(𝝉\(s\),ws\)\}s=1S\\mathcal\{S\}=\\\{\(\\boldsymbol\{\\tau\}^\{\(s\)\},\\,w\_\{s\}\)\\\}\_\{s=1\}^\{S\}, where𝝉\(s\)=\(τℓ\(s\)\)ℓ∈ℒ\\boldsymbol\{\\tau\}^\{\(s\)\}=\(\\tau\_\{\\ell\}^\{\(s\)\}\)\_\{\\ell\\in\\mathcal\{L\}\}assigns a threat levelτℓ\(s\)∈\[0,1\]\\tau\_\{\\ell\}^\{\(s\)\}\\in\[0,1\]to each location under scenarioss, andws\>0w\_\{s\}\>0with∑sws=1\\sum\_\{s\}w\_\{s\}=1\. Three scenario families are supported:*uniform*\(τℓ\(s\)∼𝒰​\(0\.05,0\.25\)\\tau\_\{\\ell\}^\{\(s\)\}\\sim\\mathcal\{U\}\(0\.05,0\.25\), no geographic signal\),*skewed*\(τℓ\(s\)∝vℓ\\tau\_\{\\ell\}^\{\(s\)\}\\propto v\_\{\\ell\}, adversary targets high\-value bases\), and*adversarial*\(focused high\-threat scenarios mixed with diffuse low\-threat scenarios\)\. Weights may be uniform \(ws=1/Sw\_\{s\}=1/S\) or peaked, representing varying degrees of intelligence confidence\.

The adversarial Bayesian counter\-move model extends this representation dynamically\. Given a defender placementπ\\pi, the adversary updates scenario weights proportionally to the total threat exposure induced by that placement:

w~s​\(π\)∝\(1−λ\)​ws\+λ⋅∑ℓ∈ℒτℓ\(s\)​nℓ​\(π\),\\tilde\{w\}\_\{s\}\(\\pi\)\\;\\propto\\;\(1\-\\lambda\)\\,w\_\{s\}\\;\+\\;\\lambda\\cdot\\sum\_\{\\ell\\in\\mathcal\{L\}\}\\tau\_\{\\ell\}^\{\(s\)\}\\,n\_\{\\ell\}\(\\pi\),\(3\)wherenℓ​\(π\)n\_\{\\ell\}\(\\pi\)is the number of assets at locationℓ\\ellandλ=pobs⋅γ\\lambda=p\_\{\\mathrm\{obs\}\}\\cdot\\gammablends prior weights with the adversary’s best\-response signal \(pobsp\_\{\\mathrm\{obs\}\}is observation probability;γ∈\{0,1\}\\gamma\\in\\\{0,1\\\}is adversary rationality\)\. This nests the standard static scenario tree \(λ=0\\lambda=0\) as the special case of a non\-observing or irrational adversary\.

### 4\.2Core Optimizer

##### Greedy baseline\.

The greedy policyπG\\pi\_\{G\}assigns assets sequentially to the highest\-strategic\-value location with remaining capacity \(Equation \([5](https://arxiv.org/html/2608.05256#S5.E5)\)\)\. It requires no scenario information and runs inO​\(\|𝒜0\|​log⁡\|ℒ\|\)O\(\|\\mathcal\{A\}\_\{0\}\|\\log\|\\mathcal\{L\}\|\)time, making it the natural computational benchmark\. Its deficiencies, geographic under\-coverage and vulnerability to value\-correlated threat, are characterized in Experiment 1\.

##### CEV optimizer\.

The Composite Expected Value optimizer replaces raw strategic value with scenario\-weighted expected value when ranking locations:

v^ℓ=∑s=1Sws​vℓ​\(1−τℓ\(s\)\),\\hat\{v\}\_\{\\ell\}=\\sum\_\{s=1\}^\{S\}w\_\{s\}\\;v\_\{\\ell\}\\left\(1\-\\tau\_\{\\ell\}^\{\(s\)\}\\right\),\(4\)then assigns assets greedily to the highest\-ranked location with remaining capacity\. This is the optimal first\-stage decision for the two\-stage stochastic program \([14](https://arxiv.org/html/2608.05256#S6.E14)\) when the per\-location value function is separable and non\-increasing in asset count, both conditions satisfied under capacity constraintcc\[[11](https://arxiv.org/html/2608.05256#bib.bib5),[8](https://arxiv.org/html/2608.05256#bib.bib6)\]\. Complexity isO​\(S⋅\|ℒ\|\+\|𝒜0\|​log⁡\|ℒ\|\)O\(S\\cdot\|\\mathcal\{L\}\|\+\|\\mathcal\{A\}\_\{0\}\|\\log\|\\mathcal\{L\}\|\), linear in scenario count\.

##### RobustCEV\.

The adversarially\-robust extension iterates: \(1\) optimize placement under current weights using CEV, \(2\) apply the adversarial weight update \([3](https://arxiv.org/html/2608.05256#S4.E3)\), \(3\) repeat until placement is stable or 10 iterations are reached\. The fixed point is a Stackelberg equilibrium in which the defender’s placement cannot be further exploited by the adversary’s best\-response update\. Failure modes at highpobsp\_\{\\mathrm\{obs\}\}under concentrated priors are documented in Experiment 3\.

### 4\.3Alert Tier and Configuration Mode Representation

Each asset carries a configuration modemam\_\{a\}encoding its operational state: a cyber node may beDormant\(low power, undetectable, reduced capability\) orActive\(fully operational, detectable, higher cost\); a fighter may beLoitering\(repositionable\) orCommitted\(on tasked mission\)\. Mode transitions carry a costκm→m′\\kappa\_\{m\\to m^\{\\prime\}\}that enters the sustainment costKtK\_\{t\}and thus the reward \([2](https://arxiv.org/html/2608.05256#S2.E2)\), penalizing unnecessary cycling\.

The scenario engine accounts for mode when computing scenario\-weighted readiness: assets inDormantmode have reduced threat exposure \(effectiveτ\\tauscaled byαm<1\\alpha\_\{m\}<1\) but also reduced readiness contribution \(rar\_\{a\}scaled byβm≤1\\beta\_\{m\}\\leq 1\), creating a genuine survivability–effectiveness tradeoff that the CEV optimizer navigates across the scenario distribution\.

### 4\.4Human\-Machine Teaming Interface

The optimizer produces a ranked list of placement recommendations, each annotated with the expected posture efficiency, the scenario distribution under which it is optimal, and the adversarial regret at the estimated observation probability\. In*automated mode*, the top\-ranked placement is executed subject to hard capacity and overflight constraints\. In*HMT mode*, the Battle Manager receives the top\-kkplacements \(defaultk=3k=3\) with plain\-language rationale: e\.g\., “Placement A concentrates readiness at Kadena and Andersen but is vulnerable if the adversary observes your posture; Placement B hedges by occupying Diego Garcia at the cost of 8% lower expected efficiency under the benign scenario\.”

The interface also surfaces the variance decomposition \(Section[10](https://arxiv.org/html/2608.05256#S10)\): if the ICC under the current threat prior exceeds a configurable threshold, the system alerts the Battle Manager that the recommendation is sensitive to scenario\-set choice, warranting additional ISU tasking to sharpen the threat estimate before committing\.

## 5Experiment 1: Greedy Baseline Characterization

### 5\.1Motivation

Before evaluating scenario\-weighted stochastic optimizers, it is necessary to establish what a myopic, value\-maximizing greedy strategy achieves and, critically, where it fails\. A greedy placement policy constitutes the expected\-value benchmark against which the Composite Expected Value \(CEV\) optimizer is measured in Experiment 2\. This section characterizes greedy performance across four posture metrics over a 10\-step sustainment horizon, then quantifies its structural vulnerability to non\-uniform threat distributions, the gap whose elimination motivates distributionally robust optimization\.

### 5\.2Experimental Setup

The simulation environment consists of\|𝒜\|=20\|\\mathcal\{A\}\|=20assets distributed across\|ℒ\|=5\|\\mathcal\{L\}\|=5named theater locations drawn from Indo\-Pacific basing sites \(Kadena AB, Andersen AFB, MCAS Iwakuni, Camp H\.M\. Smith, and Diego Garcia\), each with a per\-location capacity ofc=5c=5and a pre\-assigned strategic valuevℓ∈\[0\.78,0\.95\]v\_\{\\ell\}\\in\[0\.78,\\,0\.95\]\. Asset types are drawn uniformly from five categories \(aircraft, fuel depots, maintenance crews, munitions, and medical assets\), with initial readiness rates sampled from𝒰​\(0\.4,1\.0\)\\mathcal\{U\}\(0\.4,1\.0\)and maintenance windows from𝒰​\(1,90\)\\mathcal\{U\}\(1,90\)days\.

A fixed per\-step degradation rateδ=0\.08\\delta=0\.08is applied to every asset at each ofT=10T=10discrete time steps, replacing the stochastic degradation draws used in general simulation to isolate the effect of the placement and sustainment policy from noise\. All results are averaged overNseed=10N\_\{\\mathrm\{seed\}\}=10independent random seeds, with standard deviations reported throughout\.

### 5\.3Greedy Placement Algorithm

The greedy placement policyπG\\pi\_\{G\}assigns assets sequentially to the highest\-value location with remaining capacity:

πG​\(a\)=arg​maxℓ∈ℒ⁡vℓsubject to​\|\{a′:πG​\(a′\)=ℓ\}\|<c\.\\pi\_\{G\}\(a\)=\\operatorname\*\{arg\\,max\}\_\{\\ell\\in\\mathcal\{L\}\}\\;v\_\{\\ell\}\\quad\\text\{subject to \}\|\\\{a^\{\\prime\}:\\pi\_\{G\}\(a^\{\\prime\}\)=\\ell\\\}\|<c\.\(5\)Assets are processed in the order they appear in the initial state\. Because all five locations share the same capacityc=5c=5, and 20 assets exactly fill four locations,πG\\pi\_\{G\}deterministically assigns assets to the four highest\-value locations\{ℓ1,…,ℓ4\}\\\{\\ell\_\{1\},\\ldots,\\ell\_\{4\}\\\}and leaves the lowest\-value locationℓ5\\ell\_\{5\}\(Diego Garcia,vℓ5=0\.78v\_\{\\ell\_\{5\}\}=0\.78\) empty\. This is a direct consequence of greedy’s myopic value\-maximization: it has no coverage constraint\.

Between time steps, sustainment actions are determined by a*ReplenishmentPolicy*πR\\pi\_\{R\}that inspects each asset independently:

πR​\(a\)=\{Maintainif​ra<0\.4​or​da<7,Resupplyif​qa<2,Holdotherwise,\\pi\_\{R\}\(a\)=\\begin\{cases\}\\textsc\{Maintain\}&\\text\{if \}r\_\{a\}<0\.4\\text\{ or \}d\_\{a\}<7,\\\\ \\textsc\{Resupply\}&\\text\{if \}q\_\{a\}<2,\\\\ \\textsc\{Hold\}&\\text\{otherwise,\}\\end\{cases\}\(6\)wherera∈\[0,1\]r\_\{a\}\\in\[0,1\]is the readiness rate,dad\_\{a\}is days until next scheduled maintenance, andqaq\_\{a\}is asset quantity\.Maintainrestoresra←min⁡\(1,ra\+0\.20\)r\_\{a\}\\leftarrow\\min\(1,r\_\{a\}\+0\.20\);Resupplyincrementsqa←qa\+2q\_\{a\}\\leftarrow q\_\{a\}\+2\. The per\-action sustainment costs areκ​\(Reposition\)=10\\kappa\(\\textsc\{Reposition\}\)=10,κ​\(Resupply\)=5\\kappa\(\\textsc\{Resupply\}\)=5,κ​\(Maintain\)=2\\kappa\(\\textsc\{Maintain\}\)=2, andκ​\(Hold\)=0\\kappa\(\\textsc\{Hold\}\)=0\.

### 5\.4Performance Metrics

Four metrics are computed at each time step\.

##### Readiness Score \(RR\)\.

Quantity\-weighted mean readiness across all assets:

R=∑a∈𝒜qa​ra∑a∈𝒜qa\.R=\\frac\{\\sum\_\{a\\in\\mathcal\{A\}\}q\_\{a\}\\,r\_\{a\}\}\{\\sum\_\{a\\in\\mathcal\{A\}\}q\_\{a\}\}\.\(7\)

##### Coverage Score \(CC\)\.

Fraction of locations occupied by at least one asset:

C=\|\{ℓ∈ℒ:∃a,π​\(a\)=ℓ\}\|\|ℒ\|\.C=\\frac\{\|\\\{\\ell\\in\\mathcal\{L\}:\\exists\\,a,\\,\\pi\(a\)=\\ell\\\}\|\}\{\|\\mathcal\{L\}\|\}\.\(8\)

##### Sustainment Cost \(KK\)\.

Total action cost incurred by the ReplenishmentPolicy at a given step:

K=∑a∈𝒜κ​\(πR​\(a\)\)\.K=\\sum\_\{a\\in\\mathcal\{A\}\}\\kappa\\\!\\left\(\\pi\_\{R\}\(a\)\\right\)\.\(9\)

##### Posture Efficiency \(EE\)\.

A composite metric that rewards joint readiness–coverage while penalizing cost superlinearly:

E=R⋅Clog⁡\(K\+2\)\.E=\\frac\{R\\cdot C\}\{\\log\(K\+2\)\}\.\(10\)The logarithmic cost penalty reflects diminishing marginal returns to sustainment spending\.

##### Scenario\-Weighted Readiness \(SWR\)\.

To quantify threat\-environment sensitivity, we evaluate greedy postures against a set ofS=20S=20independently drawn threat scenarios\. Each scenariossassigns a threat levelτℓ\(s\)∈\[0,1\]\\tau\_\{\\ell\}^\{\(s\)\}\\in\[0,1\]to every location, reducing the effective readiness of any asset stationed there\. The SWR is the probability\-weighted expected effective readiness:

SWR=∑s=1Sws∑s′ws′⋅∑a∈𝒜qa​ra​\(1−τπ​\(a\)\(s\)\)∑a∈𝒜qa,\\mathrm\{SWR\}=\\sum\_\{s=1\}^\{S\}\\frac\{w\_\{s\}\}\{\\sum\_\{s^\{\\prime\}\}w\_\{s^\{\\prime\}\}\}\\cdot\\frac\{\\sum\_\{a\\in\\mathcal\{A\}\}q\_\{a\}\\,r\_\{a\}\\left\(1\-\\tau\_\{\\pi\(a\)\}^\{\(s\)\}\\right\)\}\{\\sum\_\{a\\in\\mathcal\{A\}\}q\_\{a\}\},\(11\)wherews=1w\_\{s\}=1for allss\(uniform weighting\)\. Two threat distributions are evaluated:

- •Uniform:τℓ\(s\)∼𝒰​\(0\.10,0\.30\)\\tau\_\{\\ell\}^\{\(s\)\}\\sim\\mathcal\{U\}\(0\.10,0\.30\)for allℓ\\ell, a low\-intensity, operationally symmetric threat environment\.
- •Skewed:τℓ\(s\)=min⁡\(0\.9,vℓ⋅𝒰​\(0\.50,1\.00\)\)\\tau\_\{\\ell\}^\{\(s\)\}=\\min\\\!\\left\(0\.9,\\;v\_\{\\ell\}\\cdot\\mathcal\{U\}\(0\.50,1\.00\)\\right\), threat intensity scales with strategic value, modeling an adversary that preferentially targets the most valuable basing locations\.

All scenario draws are fixed by a common seed across simulation runs so that observed variance reflects initial\-state heterogeneity, not scenario sampling\.

### 5\.5Results

#### 5\.5\.1Per\-Step Metrics

Table[1](https://arxiv.org/html/2608.05256#S5.T1)reports the four metrics over all 10 time steps\. Several features merit attention\.

- •Readiness declines then stabilizes\.RRfalls from0\.712±0\.0570\.712\\pm 0\.057att=0t=0to a trough near0\.5530\.553att=5t=5–66, then partially recovers to0\.570±0\.0410\.570\\pm 0\.041att=10t=10as the ReplenishmentPolicy’sMaintainactions restore degraded assets\. The stabilization reflects an equilibrium between the fixedδ=0\.08\\delta=0\.08degradation rate and the\+0\.20\+0\.20readiness gain fromMaintain\.
- •Coverage is constant at 0\.80\.Greedy’s deterministic assignment to the four highest\-value locations fixesC=4/5=0\.80C=4/5=0\.80throughout the simulation\. The fifth location remains unoccupied at every time step\.
- •Sustainment cost exhibits a large initial surge\.K=13\.1±6\.0K=13\.1\\pm 6\.0att=0t=0, reflecting widespread maintenance and resupply needs in the initial asset population \(low maintenance windows, low quantities\)\. Cost drops sharply to3\.4±1\.93\.4\\pm 1\.9byt=2t=2as immediate needs are addressed, then rises gradually to6\.8±3\.26\.8\\pm 3\.2byt=10t=10as ongoing degradation generates a steady maintenance demand\.
- •Posture efficiency peaks early then decays\.EErises sharply from0\.2210\.221att=0t=0to0\.3440\.344att=1t=1as the initial cost surge subsides, then declines monotonically to0\.2220\.222byt=10t=10as readiness erodes and cost accumulates\. The return to near\-t0t\_\{0\}efficiency byt=10t=10indicates that greedy reaches a long\-run steady state\.

Table 1:Greedy placement baseline metrics over 10 time steps \(mean±\\pmstd across 10 random seeds; fixed degradation rateδ=0\.08\\delta=0\.08; location capacity = 5\)\.
#### 5\.5\.2Comparison with Random Placement

To contextualize greedy’s performance, we compare it against a random placement baseline that assigns assets uniformly at random subject to the same capacity constraint, followed by the same ReplenishmentPolicy\. Figure[1](https://arxiv.org/html/2608.05256#S5.F1)\(panels a–b\) shows the trajectories for both strategies across all 10 seeds\.

Readiness is*placement\-independent*: both greedy and random converge toR=0\.570±0\.041R=0\.570\\pm 0\.041byt=10t=10\(panel a\)\. The ReplenishmentPolicy applies identically regardless of geographic assignment, so asset readiness is governed solely by the degradation–repair equilibrium\. The two lines overlap throughout, confirming that value\-based greedy placement confers no readiness advantage over random assignment\.

Posture efficiency tells a different story \(panel b\)\. Random placement achievesE=0\.278±0\.047E=0\.278\\pm 0\.047att=10t=10, versusE=0\.222±0\.038E=0\.222\\pm 0\.038for greedy—a25\.1% gap\. The sole driver is coverage: random assignment distributes assets across all five locations \(C=1\.00C=1\.00\), while greedy concentrates assets at four \(C=0\.80C=0\.80\)\. Given identical readiness and cost, Equation \([10](https://arxiv.org/html/2608.05256#S5.E10)\) yields a coverage ratio of1\.00/0\.80=1\.251\.00/0\.80=1\.25, fully accounting for the observed efficiency gap\. This result exposes a fundamental limitation of value\-only greedy placement: by ignoring geographic coverage, it sacrifices composite operational efficiency despite maintaining strong local asset readiness\.

#### 5\.5\.3Threat\-Environment Sensitivity

Table[2](https://arxiv.org/html/2608.05256#S5.T2)and Figure[1](https://arxiv.org/html/2608.05256#S5.F1)\(panel c\) report SWR under uniform and skewed threat distributions\. The results reveal a structural vulnerability of greedy placement\.

Underuniform threat, SWR declines from0\.568±0\.0460\.568\\pm 0\.046att=0t=0to0\.455±0\.0320\.455\\pm 0\.032att=10t=10—a 20\.2% reduction from nominal readiness\. This modest degradation reflects the low mean threat level \(𝔼​\[τ\]=0\.20\\mathbb\{E\}\[\\tau\]=0\.20\) under the uniform distribution, which reduces effective readiness by a corresponding fraction regardless of where assets are placed\.

Under theskewed distribution, SWR collapses to0\.243±0\.0200\.243\\pm 0\.020att=0t=0and0\.194±0\.0150\.194\\pm 0\.015att=10t=10, a 65\.9% reduction from nominal, or a 57\.3% drop relative to the uniform\-threat SWR\. This gap is*not a statistical artifact*: it remains constant at57\.3%57\.3\\%across all 10 time steps and all 10 random seeds \(Table[2](https://arxiv.org/html/2608.05256#S5.T2), % Drop column\)\.

The constancy has a structural explanation\. Because the threat scenarios are fixed and greedy*always*assigns assets to the same four highest\-value locations\{ℓ1,…,ℓ4\}\\\{\\ell\_\{1\},\\ldots,\\ell\_\{4\}\\\}, the SWR ratio between conditions is determined entirely by the ratio of expected survival rates at those locations:

SWRskewedSWRuniform≈14​∑i=14𝔼​\[1−τℓi\(s\)\]skewed𝔼​\[1−τ\(s\)\]uniform=0\.3440\.800≈0\.43,\\frac\{\\mathrm\{SWR\}\_\{\\mathrm\{skewed\}\}\}\{\\mathrm\{SWR\}\_\{\\mathrm\{uniform\}\}\}\\approx\\frac\{\\frac\{1\}\{4\}\\sum\_\{i=1\}^\{4\}\\mathbb\{E\}\\\!\\left\[1\-\\tau\_\{\\ell\_\{i\}\}^\{\(s\)\}\\right\]\_\{\\mathrm\{skewed\}\}\}\{\\mathbb\{E\}\\\!\\left\[1\-\\tau^\{\(s\)\}\\right\]\_\{\\mathrm\{uniform\}\}\}=\\frac\{0\.344\}\{0\.800\}\\approx 0\.43,\(12\)yielding a fixed 57% drop\. Under the skewed distribution, an adversary that targets high\-strategic\-value bases finds that*all*greedy assets are concentrated precisely at the locations facing the greatest threat\. Greedy placement, by optimizing for peacetime strategic value, maximally exposes assets to value\-correlated adversarial threats\.

Table 2:Greedy scenario\-weighted readiness \(SWR\) under uniform and skewed threat distributions \(mean±\\pmstd across 10 random seeds; 20 scenarios per condition\)\.Δ=SWRuniform−SWRskewed\\Delta=\\text\{SWR\}\_\{\\text\{uniform\}\}\-\\text\{SWR\}\_\{\\text\{skewed\}\}; % drop=Δ/SWRuniform=\\Delta/\\text\{SWR\}\_\{\\text\{uniform\}\}\.![Refer to caption](https://arxiv.org/html/2608.05256v1/x1.png)Figure 1:Greedy baseline characterization over 10 time steps \(Nseed=10N\_\{\\mathrm\{seed\}\}=10,δ=0\.08\\delta=0\.08, capacity=5=5\)\. Shaded bands show±1​σ\\pm 1\\sigma\.\(a\)Readiness score: greedy and random placement converge to identical readiness byt=5t=5, confirming placement\-invariance under the ReplenishmentPolicy\.\(b\)Posture efficiency: random placement consistently outperforms greedy by 25% due to full geographic coverage \(C=1\.00C=1\.00vs\.C=0\.80C=0\.80\)\.\(c\)Scenario\-weighted readiness \(SWR\) under uniform vs\. skewed threat: the 57\.3% gap between conditions is constant across all time steps, reflecting a structural vulnerability of greedy placement to value\-correlated adversarial threats\.

### 5\.6Discussion

Experiment 1 yields three findings that collectively motivate the scenario\-weighted optimizer evaluated in Experiment 2\.

##### Finding 1: Readiness is placement\-invariant under rule\-based sustainment\.

The ReplenishmentPolicy equalizes readiness across placement strategies by step 10\. This means greedy’s advantage over random is not readiness but other posture qualities, and in fact greedy underperforms random on the composite efficiency metric\. Any optimizer that claims readiness improvements over greedy must do so through sustainment policy changes, not placement alone\.

##### Finding 2: Greedy sacrifices geographic coverage for value concentration\.

With capacityc=5c=5and\|𝒜\|=20\|\\mathcal\{A\}\|=20assets, greedy leaves one location persistently uncovered, incurring a permanent 25% posture efficiency penalty\. This is not a boundary artifact: for any configuration where\|𝒜\|\|\\mathcal\{A\}\|is not a multiple of\|ℒ\|⋅c\|\\mathcal\{L\}\|\\cdot c, greedy will under\-cover lower\-value locations\. A stochastic optimizer that accounts for coverage through threat scenario diversity can exploit this gap\.

##### Finding 3: Greedy placement creates a structurally fixed threat vulnerability\.

The 57\.3% SWR degradation under skewed threat is not reducible through better sustainment; it is a geometric consequence of where assets are placed\. The vulnerability is invariant to seed, time step, and readiness level, meaning it cannot be repaired reactively; it must be anticipated at the placement stage\. This finding directly motivates the key research question in Issue \#10 \(Q1\): an optimizer that is aware of threat distribution heterogeneity at planning time should be able to trade some value concentration for threat\-robustness, capturing positive EVSS \(Expected Value of the Stochastic Solution\) over greedy\.

## 6Experiment 2: Scenario\-Weighted Optimizer vs\. Greedy Baseline

### 6\.1Motivation

Experiment 1 established that greedy placement creates a structural vulnerability to non\-uniform threat distributions: by concentrating assets at high\-value locations, it exposes itself maximally to value\-correlated adversarial threats\. Experiment 2 quantifies the gain from replacing the myopic greedy policy with a scenario\-weighted optimizer that explicitly accounts for threat heterogeneity at planning time\. The primary metric is the*Expected Value of the Stochastic Solution*\(EVSS\), defined as the posture efficiency gap between the Composite Expected Value \(CEV\) optimizer and the greedy baseline when evaluated under the same scenario distribution\.

### 6\.2Experimental Setup

The base environment retains the Experiment 1 parameters:\|𝒜\|=20\|\\mathcal\{A\}\|=20assets,\|ℒ\|=5\|\\mathcal\{L\}\|=5theater locations, capacityc=5c=5per location, andNseed=10N\_\{\\mathrm\{seed\}\}=10independent seeds\. Stochastic degradation \(rather than the fixedδ\\deltaused in Experiment 1\) is restored to reflect the full operational environment\. Results are averaged across seeds with standard deviations reported\.

Two factors are varied:

##### Threat distribution\.

Three distributions are evaluated:

- •Uniform: threat intensities drawn from𝒰​\(0\.05,0\.25\)\\mathcal\{U\}\(0\.05,0\.25\)independently for each location and scenario\. No location is systematically more threatened than others; greedy’s value\-concentration carries no additional risk\.
- •Skewed: threat intensity scales with strategic value,τℓ\(s\)=min⁡\(0\.9,vℓ⋅𝒰​\(0\.5,1\.0\)\)\\tau\_\{\\ell\}^\{\(s\)\}=\\min\\\!\\left\(0\.9,\\;v\_\{\\ell\}\\cdot\\mathcal\{U\}\(0\.5,1\.0\)\\right\)\. An adversary that targets high\-value bases; the greedy assignment concentrates assets precisely at the most threatened locations\.
- •Adversarial: 60% of scenarios concentrate high threat \(τ∈\[0\.7,0\.95\]\\tau\\in\[0\.7,0\.95\]\) on a single randomly chosen location; the remaining 40% use low uniform threat \(τ∈\[0\.05,0\.20\]\\tau\\in\[0\.05,0\.20\]\)\. Models a focused A2/AD actor that occasionally executes feints across the theater\.

##### Scenario count \(SS\)\.

The CEV optimizer is evaluated withS∈\{5,20,100\}S\\in\\\{5,20,100\\\}independently drawn scenarios\. All scenarios within a condition share equal weightws=1/Sw\_\{s\}=1/S\. Threat scenarios are fixed by a common seed across experimental runs to isolate the effect of scenario count from sampling noise\.

##### Metrics\.

*Greedy efficiency*\(EGE\_\{G\}\) is evaluated by the CEV scorer against the full scenario distribution; this ensures both policies are measured against an identical benchmark\.*CEV efficiency*\(ECEVE\_\{\\mathrm\{CEV\}\}\) is the expected posture efficiency of the scenario\-optimized placement\. The*Expected Value of the Stochastic Solution*is:

EVSS=ECEV−EG\.\\mathrm\{EVSS\}=E\_\{\\mathrm\{CEV\}\}\-E\_\{G\}\.\(13\)Positive EVSS indicates that scenario\-awareness at planning time yields measurable efficiency gains over reactive greedy placement\.

### 6\.3CEV Optimizer

The CEV optimizer solves a two\-stage stochastic program\. In the first stage \(here\-and\-now\), assets are assigned to locations to maximize the scenario\-weighted expected posture efficiency:

π⋆=arg​maxπ​∑s=1Sws​E​\[π,ρs\],\\pi^\{\\star\}=\\operatorname\*\{arg\\,max\}\_\{\\pi\}\\sum\_\{s=1\}^\{S\}w\_\{s\}\\;E\\\!\\left\[\\pi,\\rho\_\{s\}\\right\],\(14\)whereρs\\rho\_\{s\}denotes the second\-stage recourse action selected by the*ReplenishmentPolicy*after scenariossis revealed\. Location rank is determined by the scenario\-weighted expected strategic value:

v^ℓ=∑s=1Sws​vℓ​\(1−τℓ\(s\)\),\\hat\{v\}\_\{\\ell\}=\\sum\_\{s=1\}^\{S\}w\_\{s\}\\;v\_\{\\ell\}\\left\(1\-\\tau\_\{\\ell\}^\{\(s\)\}\\right\),\(15\)and assets are assigned greedily to the highest\-ranked location with remaining capacity\. This greedy\-over\-scenarios approach is tractable even for largeSSand produces the globally optimal assignment when the marginal value of additional assets at a location is non\-increasing \(which holds given our fixed\-capacity constraint\)\.

In the second stage, assets at locations with threat exceeding a threshold \(τ\>0\.70\\tau\>0\.70\) are repositioned; remaining assets receive actions from the*ReplenishmentPolicy*\.

### 6\.4Results

Table[3](https://arxiv.org/html/2608.05256#S6.T3)reports greedy efficiency, CEV efficiency, and EVSS for all nine conditions\.

Table 3:Experiment 2 results: posture efficiency of greedy vs\. CEV optimizer across threat distributions and scenario counts \(Nseed=10N\_\{\\mathrm\{seed\}\}=10; mean±\\pm1σ\\sigma\)\. EVSS = CEV Eff−\-Greedy Eff\.Several findings are notable\.

##### EVSS is zero under uniform threat \(expected\)\.

When all locations face identical expected threat, both the greedy and CEV optimizers rank locations by raw strategic value and produce identical assignments\. This is the theoretical null result: scenario information provides no placement advantage when the threat distribution carries no geographic signal\.

##### CEV achieves 19\.8% efficiency gain under skewed threat\.

WithS=5S=5scenarios, the CEV optimizer outperforms greedy by 19\.8% \(EVSS=\+0\.030\\mathrm\{EVSS\}=\+0\.030\)\. This is the largest measured gain across all conditions\. The skewed distribution concentrates threat at the same high\-strategic\-value locations that greedy preferentially occupies, so greedy’s placement incurs heavy second\-stage recourse costs; the CEV optimizer avoids this by weighting location value against expected threat exposure\.

##### EVSS decreases with scenario count under skewed threat\.

Counter\-intuitively, EVSS falls from 19\.8% atS=5S=5to 9\.9% atS=100S=100under the skewed distribution\. With few scenarios, the per\-scenario threat signal is sharp \(a small sample drawn from the high\-threat regime\), so the CEV optimizer’s avoidance of high\-value locations is decisive\. With many scenarios, the sample mean converges to the population mean, making the scenario distribution smoother and reducing the discriminating power of each individual scenario\. In practice, theS=5S=5toS=20S=20regime captures the bulk of the attainable EVSS gain\.

##### Adversarial distribution yields moderate but consistent EVSS\.

The mixed adversarial distribution \(60% focused, 40% diffuse\) produces EVSS of 2\.8–10\.2% depending on scenario count\. The large gain atS=5S=5arises because a small sample may by chance contain multiple focused\-threat scenarios, giving the optimizer a clear signal\. AtS=20S=20–100100the EVSS stabilizes near 2\.8–2\.9%, reflecting a genuine but modest informational advantage over greedy\.

### 6\.5Discussion

##### Finding 1: EVSS is strictly positive whenever the threat distribution carries geographic signal\.

Across all six non\-uniform conditions, the CEV optimizer matches or outperforms greedy\. The direction of the gap is theoretically guaranteed: CEV is optimal by construction under the given scenario distribution, soEVSS≥0\\mathrm\{EVSS\}\\geq 0holds in expectation\. The magnitude of the gap \(up to 19\.8%\) demonstrates that scenario\-awareness is operationally meaningful in the threat regimes most relevant to PSA\.

##### Finding 2: Diminishing returns to scenario count\.

The largest marginal gain from adding scenarios occurs betweenS=5S=5andS=20S=20\. BeyondS=20S=20, EVSS changes by less than one percentage point\. This has a practical implication: planners need not enumerate exhaustive scenario libraries to achieve near\-optimal placement decisions\. A curated set of 5–20 high\-fidelity threat scenarios captures the majority of the stochastic gain\.

##### Finding 3: Uniform threat is the correct null control\.

The zero\-EVSS result under uniform threat validates both the optimizer implementation and the metric: when no distribution shift is possible, the stochastic solution cannot outperform the deterministic one\. Any claim of stochastic advantage must be accompanied by evidence of distributional heterogeneity\.

## 7Experiment 3: Adversarial Robustness and the Bayesian Counter\-Move Model

### 7\.1Motivation

Experiments 1 and 2 treat the threat distribution as fixed and exogenous\. In practice, a rational adversary observes the defender’s posture and*updates*its attack distribution accordingly, concentrating effort on the most exposed locations\. A naive CEV optimizer that ignores this strategic interaction is exploitable: by placing assets predictably, it hands the adversary a targeting map\.

Experiment 3 tests whether the adversarially\-robust CEV variant \(RobustCEV\) maintains a higher posture efficiency floor than the naive CEV optimizer when an adaptive adversary is present\.

### 7\.2Adversarial Model

We model the adversary as a Bayesian agent that observes the defender’s placement with probabilitypobs∈\[0,1\]p\_\{\\mathrm\{obs\}\}\\in\[0,1\]and, if rational, redistributes scenario probability mass toward scenarios that best target the observed concentration\. Formally, given a placementπ\\piand a prior scenario set𝒮\\mathcal\{S\}, the adversarially\-updated probability of scenariossis:

p~s​\(π\)=\(1−λ\)​ps\+λ⋅∑ℓτℓ\(s\)​nℓ​\(π\)∑s′∑ℓτℓ\(s′\)​nℓ​\(π\),\\tilde\{p\}\_\{s\}\(\\pi\)=\(1\-\\lambda\)\\,p\_\{s\}\\;\+\\;\\lambda\\cdot\\frac\{\\displaystyle\\sum\_\{\\ell\}\\tau\_\{\\ell\}^\{\(s\)\}\\,n\_\{\\ell\}\(\\pi\)\}\{\\displaystyle\\sum\_\{s^\{\\prime\}\}\\sum\_\{\\ell\}\\tau\_\{\\ell\}^\{\(s^\{\\prime\}\)\}\\,n\_\{\\ell\}\(\\pi\)\},\(16\)wherenℓ​\(π\)n\_\{\\ell\}\(\\pi\)is the number of assets assigned to locationℓ\\ell, andλ=pobs⋅γ\\lambda=p\_\{\\mathrm\{obs\}\}\\cdot\\gammais a blend coefficient combining observation probabilitypobsp\_\{\\mathrm\{obs\}\}and adversary rationalityγ∈\{0,1\}\\gamma\\in\\\{0,1\\\}\. Whenγ=0\\gamma=0\(random adversary\) the blend is zero regardless ofpobsp\_\{\\mathrm\{obs\}\}, leaving the scenario distribution unchanged\. Whenγ=1\\gamma=1\(Bayesian adversary\) the update fully reflects the adversary’s best response atpobs=1p\_\{\\mathrm\{obs\}\}=1\.

The*RobustCEV*optimizer iterates the following loop until the placement stabilizes or a maximum of 10 iterations is reached: \(1\) optimize placement against current scenario distribution, \(2\) apply the adversarial update \([16](https://arxiv.org/html/2608.05256#S7.E16)\), \(3\) repeat\. The converged placement and its corresponding adversarial distribution are returned\.

### 7\.3Experimental Setup

The evaluation sweepspobs∈\{0,0\.25,0\.50,0\.75,1\.0\}p\_\{\\mathrm\{obs\}\}\\in\\\{0,0\.25,0\.50,0\.75,1\.0\\\}andγ∈\{0​\(random\),1​\(Bayesian\)\}\\gamma\\in\\\{0\\text\{ \(random\)\},1\\text\{ \(Bayesian\)\}\\\}across four threat prior distributions\. The same 20\-asset, 5\-location posture state from Experiments 1–2 is used \(capacityc=20c=20for Experiment 3 to match the scenario\-set construction; seed = 42\)\.

##### Threat distributions\.

- •Uniform: equal threat across all locations in every scenario\. CEV placement is already optimal; no exploitable concentration exists\.
- •Skewed: threat concentrated on the highest\-value location in the dominant scenario\. Naive CEV already avoids this location; adversary has limited leverage\.
- •Adversarial: threats on the top\-two strategic\-value locations with demand multipliers \(×1\.3\\times 1\.3–1\.51\.5\), forcing hedging across multiple locations\.
- •Deceptive: a mostly\-safe prior \(95% no\-threat scenarios\) with a low\-probability, high\-intensity attack \(5% probability, 99% threat\) on the highest\-value location\. Naive CEV is lured into concentrating assets there; the adversary can fully exploit this concentration\.

##### Metrics\.

For each\(pobs,γ\)\(p\_\{\\mathrm\{obs\}\},\\gamma\)pair, we report:

- •Naive efficiency: CEV placement optimized against the prior, evaluated under the adversary’s best\-response distribution given that placement\.
- •Robust efficiency: RobustCEV placement evaluated under its converged adversarial distribution\.
- •Adversarial regret:Δ​E=Erobust−Enaive\\Delta E=E\_\{\\mathrm\{robust\}\}\-E\_\{\\mathrm\{naive\}\}\(positive means robust wins\)\.

### 7\.4Results

Table[4](https://arxiv.org/html/2608.05256#S7.T4)reports results for the Bayesian adversary \(γ=1\\gamma=1\) across all four distributions and all five observation probabilities\. Random adversary \(γ=0\\gamma=0\) results are omitted from the main table: by constructionλ=pobs⋅0=0\\lambda=p\_\{\\mathrm\{obs\}\}\\cdot 0=0, so the scenario distribution is never updated and naive and robust placements are always identical\. This is the correct theoretical null\.

Table 4:Experiment 3 results: posture efficiency of naive vs\. robust CEV optimizer under a Bayesian adversary \(γ=1\\gamma=1\) across four threat distributions and five observation probabilities\. Regret = Robust Eff−\-Naive Eff\. A random adversary \(γ=0\\gamma=0\) always produces Regret=0=0\(theoretical null\)\.#### 7\.4\.1Deceptive Prior: the Core Robustness Result

The deceptive distribution isolates the mechanism the paper targets\. The prior assigns 95% probability to a completely safe scenario, so naive CEV—which optimizes against the prior—places all assets at the highest\-strategic\-value location \(Kadena AB\)\. When the adversary observes this concentration with probabilitypobsp\_\{\\mathrm\{obs\}\}, it shifts probability mass toward the 5% attack scenario that targets exactly that location, collapsing the effective posture\.

Naive efficiency falls monotonically from 0\.0623 atpobs=0p\_\{\\mathrm\{obs\}\}=0to 0\.0224 atpobs=1\.0p\_\{\\mathrm\{obs\}\}=1\.0, a 64% collapse\. Robust CEV iterates away from the lure: after detecting the adversarial shift in the first optimization loop, it redistributes assets to the second\-best location where the attack scenario carries no threat, and converges to a stable placement that the adversary cannot further exploit\. Robust efficiency holds at0\.0570\.057–0\.0620\.062across allpobsp\_\{\\mathrm\{obs\}\}values\.

The adversarial regret grows linearly withpobsp\_\{\\mathrm\{obs\}\}, reaching\+0\.035\+0\.035atpobs=1\.0p\_\{\\mathrm\{obs\}\}=1\.0, a158%158\\%improvement over naive efficiency\.

#### 7\.4\.2Null Results Under Uniform and Skewed Priors

Under the uniform distribution, both optimizers produce identical placements at allpobsp\_\{\\mathrm\{obs\}\}values \(regret=0=0\)\. Uniform threat carries no geographic signal, so no concentration is exploitable\.

Under the skewed prior, the pattern reverses at full observation \(pobs=1\.0p\_\{\\mathrm\{obs\}\}=1\.0, regret=−0\.037=\-0\.037\): the adversary updates the scenario distribution so sharply that the RobustCEV iterative loop converges to an unstable equilibrium where assets cycle toward the very location the adversary threatens\. This is an expected failure mode of finite\-horizon iterative best\-response and does not occur forpobs≤0\.75p\_\{\\mathrm\{obs\}\}\\leq 0\.75\.

#### 7\.4\.3Adversarial Prior: Partial Convergence Failure

The adversarial distribution, threats split across the top two locations with high demand multipliers, produces small negative regret atpobs≥0\.75p\_\{\\mathrm\{obs\}\}\\geq 0\.75\. Naive CEV already partially hedges by distributing assets across both threatened locations; the iterative loop has little room to improve and can converge to a marginally worse equilibrium at highpobsp\_\{\\mathrm\{obs\}\}\. This motivates a convergence criterion based on efficiency rather than assignment identity for future work\.

### 7\.5Discussion

##### Finding 1: Robust CEV maintains a higher efficiency floor only when the prior is deceptive\.

The adversarial regret is strictly positive across allpobs\>0p\_\{\\mathrm\{obs\}\}\>0only under the deceptive distribution, the condition under which naive CEV is genuinely lured into an exploitable concentration\. This specificity is theoretically meaningful: robustness is not universally necessary, but it is necessary when the prior is a poor guide to the adversary’s true intent\.

##### Finding 2: The adversary’s observation probability directly determines the defender’s exposure\.

Under the deceptive prior, naive efficiency decreases by≈0\.010\\approx 0\.010per 0\.25 increment inpobsp\_\{\\mathrm\{obs\}\}, while robust efficiency is approximately flat\. Operationally, this means that even partial observation \(pobs=0\.25p\_\{\\mathrm\{obs\}\}=0\.25\) by the adversary produces a measurable efficiency gap \(\+0\.007\+0\.007\) that accumulates with observation fidelity\.

##### Finding 3: Random adversary is the correct null\.

Settingγ=0\\gamma=0eliminates any adversarial update regardless ofpobsp\_\{\\mathrm\{obs\}\}, producing identical naive and robust placements and zero regret\. This validates the model and confirms that the observed efficiency gaps under Bayesian play are attributable to strategic observation\-and\-response, not to any artifact of the iterative optimization loop\.

##### Finding 4: Iterative best\-response can fail atpobs=1p\_\{\\mathrm\{obs\}\}=1under strongly concentrated priors\.

The negative regret observed under skewed and adversarial distributions at full observation is a known limitation of finite\-iteration Stackelberg approximations\. The iterative loop overshoots the equilibrium when the adversary’s update is aggressive relative to the number of iterations permitted\. Addressing this is a natural extension: replacing the greedy\-assignment inner loop with a minimax formulation would guarantee non\-negative regret by construction\.

## 8Experiment 4: Computational Feasibility and the Robustness\-Cost Pareto Frontier

### 8\.1Motivation

Experiments 1–3 characterize the behavioral properties of CEV and RobustCEV under varying threat distributions and adversarial observation probabilities\. Before these optimizers can be fielded, two additional questions must be answered\. First: does the computational cost of scenario\-weighted planning prohibit use within the 4\-hour operational planning windows that govern theater\-level posture decisions? Second: given that robustness costs something, RobustCEV disperses assets to lower adversarial\-risk locations that may carry lower strategic value, what is the precise exchange rate between worst\-case readiness and placement quality sacrifice? Experiment 4 answers both questions by \(A\) measuring wall\-clock solve time across problem scales spanning an order of magnitude in asset and location count, and \(B\) tracing the full robustness\-cost Pareto frontier as the Wasserstein radius proxyε\\varepsilonvaries from 0 to 0\.8\.

### 8\.2Part A: Computational Scalability

#### 8\.2\.1Setup

The scalability sweep instantiates six problem sizes by varying assetsM∈\{10,20,50,100,150,200\}M\\in\\\{10,20,50,100,150,200\\\}and locationsN∈\{5,8,10,15,20,30\}N\\in\\\{5,8,10,15,20,30\\\}jointly, usingmake\_scaled\_theaterto generate synthetic theater instances at each scale with a fixed scenario set ofS=20S=20scenarios atε=0\.3\\varepsilon=0\.3\. Three solvers are timed at each scale:

- •CEV \(greedy\): the scenario\-weighted greedy optimizer, run from scratch\.
- •RobustCEV \(cold start\): the adversarially\-robust optimizer initialized from the prior scenario distribution, withpobs=0\.7p\_\{\\mathrm\{obs\}\}=0\.7,γ=1\.0\\gamma=1\.0, and a 20\-iteration ceiling\.
- •RobustCEV \(warm start\): identical to cold start, but the scenario weights are first updated by applying the adversarial counter\-move \([16](https://arxiv.org/html/2608.05256#S7.E16)\) to the CEV solution\. This seeds the robust optimizer at a point already informed by one defender–adversary exchange, reducing the iteration count required to reach a fixed point\.

Wall\-clock time is measured withtime\.perf\_counter\(\)and reported in seconds\. The 4\-hour operational planning limit \(14,​400 s\) is included as a reference\.

#### 8\.2\.2Results

Table[5](https://arxiv.org/html/2608.05256#S8.T5)reports solve times at all six scales\. All three solvers complete in sub\-millisecond time at every tested problem size, fromM=10M=10toM=200M=200\. At the largest instance \(M=200M=200,N=30N=30\), solve times remain below 1 ms, more than four orders of magnitude within the 4\-hour constraint\.

Table 5:Experiment 4A: wall\-clock solve time by problem size\. All entries below the measurement resolution of 1 ms \(0\.001 s\); reported as<0\.001<0\.001s\. Warm\-start speedup ratio = cold\-start time / warm\-start time\. The 4\-hour operational planning limit is 14,​400 s\.Figure[2](https://arxiv.org/html/2608.05256#S8.F2)\(panel a\) shows all three solve\-time trajectories on a logarithmic scale alongside the 4\-hour reference line\. The10410^\{4\}\-second separation between the data and the constraint boundary confirms that the computational bottleneck in PSA planning is not optimizer runtime but rather scenario acquisition, intelligence processing, and command\-authority workflows that operate on hour\-to\-day timescales\.

Warm\-start speedup \(panel b\) ranges from 1\.0–1\.3×\\timesacross tested scales\. The modest gains reflect the same algorithmic property that explains the negligible absolute times: RobustCEV converges in two to three iterations at all tested scales because the greedy placement CEV produces is already close to a fixed point\. When the problem is harder or convergence is defined more tightly, warm\-start initialization reduces iteration count proportionally, providing a principled initialization at negligible overhead\.

![Refer to caption](https://arxiv.org/html/2608.05256v1/x2.png)Figure 2:Experiment 4A: computational scalability\.\(a\)Solve time vs\. number of assetsMMon a log scale\. All three solvers remain sub\-millisecond across the full scale range; the 4\-hour operational planning limit \(red dashed line\) is more than four orders of magnitude above the measured times\.\(b\)Warm\-start speedup ratio at each scale\. Speedups of 1\.0–1\.3×\\timesreflect rapid cold\-start convergence; warm\-start advantage scales with problem difficulty and convergence strictness\.

### 8\.3Part B: Robustness\-Cost Pareto Frontier

#### 8\.3\.1Setup

The Pareto frontier sweep fixes the 8\-location Indo\-Pacific theater from Experiment 3 withM=20M=20assets\. The Wasserstein radius proxy is swept overε∈\{0\.0,0\.05,0\.1,0\.2,0\.4,0\.8\}\\varepsilon\\in\\\{0\.0,0\.05,0\.1,0\.2,0\.4,0\.8\\\}, whereε\\varepsilonparameterizes the adversarial ambiguity set inmake\_robustness\_scenarios: atε=0\\varepsilon=0all scenario threat levels are drawn uniformly from\[0\.05,0\.20\]\[0\.05,0\.20\]\(the expected\-value baseline\), and atε=1\\varepsilon=1threat levels are maximally concentrated on high\-strategic\-value locations\. Each training scenario set usesS=20S=20scenarios; results are averaged over five independent random seeds\.

Two quantities are computed at eachε\\varepsilon:

- •Worst\-case SWR: the placement is evaluated againstS=50S=50held\-out adversarial scenarios generated atεOOS=0\.9\\varepsilon\_\{\\mathrm\{OOS\}\}=0\.9, a strictly more adversarial distribution than any training set, simulating out\-of\-sample evaluation under a hostile evaluator\.
- •Cost premium: placement quality sacrifice relative to theε=0\\varepsilon=0expected\-value baseline, Δ​Q​\(ε\)=Q​\(π0\)−Q​\(πε\)Q​\(π0\),\\Delta Q\(\\varepsilon\)=\\frac\{Q\(\\pi\_\{0\}\)\-Q\(\\pi\_\{\\varepsilon\}\)\}\{Q\(\\pi\_\{0\}\)\},\(17\)whereQ​\(π\)=1\|𝒜\|​∑a∈𝒜vπ​\(a\)Q\(\\pi\)=\\frac\{1\}\{\|\\mathcal\{A\}\|\}\\sum\_\{a\\in\\mathcal\{A\}\}v\_\{\\pi\(a\)\}is the mean strategic value of the chosen locations andπ0\\pi\_\{0\}is theε=0\\varepsilon=0assignment\. PositiveΔ​Q\\Delta Qindicates that the robust placement sacrifices some nominal strategic value in exchange for adversarial protection\.

#### 8\.3\.2Results

Table[6](https://arxiv.org/html/2608.05256#S8.T6)and Figure[3](https://arxiv.org/html/2608.05256#S8.F3)present the full Pareto frontier\.

Table 6:Experiment 4B: robustness\-cost Pareto frontier averaged over 5 seeds\. Worst\-case SWR is evaluated on 50 out\-of\-sample adversarial scenarios \(εOOS=0\.9\\varepsilon\_\{\\mathrm\{OOS\}\}=0\.9\)\. Cost premium is placement quality sacrifice relative to theε=0\\varepsilon=0expected\-value baseline \(Equation \([17](https://arxiv.org/html/2608.05256#S8.E17)\)\)\. Entries atε=0\.10\\varepsilon=0\.10–0\.200\.20show near\-zero cost premium because scenario weight differences at smallε\\varepsilonare insufficient to alter the greedy ranking\.Several features of the frontier merit attention\.

##### Threshold behavior belowε=0\.2\\varepsilon=0\.2\.

Atε∈\{0\.05,0\.10,0\.20\}\\varepsilon\\in\\\{0\.05,0\.10,0\.20\\\}, worst\-case SWR and cost premium are statistically indistinguishable from theε=0\\varepsilon=0baseline\. This threshold arises from the greedy optimizer’s discrete structure: the CEV location ranking changes only when the scenario\-weighted effective values \([15](https://arxiv.org/html/2608.05256#S6.E15)\) for two locations cross\. At smallε\\varepsilon, the blended threat levels shift insufficiently to reorder the dominant location pairs, so the optimizer produces the same placement as underε=0\\varepsilon=0and the cost premium remains effectively zero\. Belowε≈0\.2\\varepsilon\\approx 0\.2, investing in a larger ambiguity set purchases no additional robustness\.

##### 18% robustness gain at 10\.7% cost premium\.

Moving fromε=0\\varepsilon=0toε=0\.8\\varepsilon=0\.8raises worst\-case SWR from 0\.261 to 0\.309, an18\.3% improvementin adversarial robustness, at a10\.7% placement quality sacrifice\. The majority of this gain concentrates in theε=0\.4\\varepsilon=0\.4–0\.80\.8range, where scenario weights have shifted sufficiently to redirect assets from the highest\-strategic\-value locations \(Kadena AB, Osan AB\) toward geographically dispersed mid\-value positions that present no single concentrated target to the adversary\.

##### Concave frontier: diminishing returns from robustness\.

The marginal exchange rate between robustness and cost declines across the frontier\. In the first half of theε\\varepsilonsweep \(0\.0≤ε≤0\.40\.0\\leq\\varepsilon\\leq 0\.4\), robustness improves at near\-zero cost premium, an effectively infinite efficiency ratio\. In the second half \(0\.4≤ε≤0\.80\.4\\leq\\varepsilon\\leq 0\.8\), the exchange rate falls to approximately 0\.45 units of worst\-case SWR per unit of cost premium\. This concavity is the characteristic signature of a well\-posed robustness\-cost tradeoff\[[1](https://arxiv.org/html/2608.05256#bib.bib12)\]: early increments of conservatism redistribute assets away from the most concentrated high\-value locations at minimal quality cost, while later increments must trade genuine placement quality for increasingly marginal additional dispersion\.

![Refer to caption](https://arxiv.org/html/2608.05256v1/x3.png)Figure 3:Experiment 4B: robustness\-cost Pareto frontier\.\(a\)Worst\-case SWR as a function of the Wasserstein radius proxyε\\varepsilon\. SWR is flat belowε≈0\.2\\varepsilon\\approx 0\.2\(the greedy ranking threshold\) then rises steeply\.\(b\)Robustness\-cost Pareto curve: each point is labeled with itsε\\varepsilonvalue\. The concave shape confirms diminishing marginal returns from additional conservatism\. A practitioner selectingε=0\.4\\varepsilon=0\.4captures most of the robustness gain \(worst\-case SWR = 0\.271\) at a 2\.1% quality cost, versus 10\.7% atε=0\.8\\varepsilon=0\.8\.

### 8\.4Discussion

##### Finding 1: The computational barrier to robustness is negligible\.

Both CEV and RobustCEV solve in sub\-millisecond time at all tested scales, placing them well within any operational planning window\. The binding resource constraint in PSA is not solver time but scenario quality and commander availability\. This result licenses future work to increase ambiguity set size \(ε\\varepsilon\) and scenario count \(SS\) without computational concern\.

##### Finding 2: The effective operating regime isε∈\[0\.3,0\.5\]\\varepsilon\\in\[0\.3,0\.5\]\.

Belowε=0\.2\\varepsilon=0\.2, the scenario\-weighted ranking is unchanged and no robustness is purchased\. Aboveε=0\.5\\varepsilon=0\.5, marginal robustness gains per unit cost premium decline sharply\. Planners calibrating the ambiguity set radius should target this intermediate band, where the full 18\-percentage\-point robustness improvement is available at a cost premium of 2–7%\.

##### Finding 3: Robustness and worst\-case readiness are jointly achievable\.

A persistent concern in distributionally robust optimization is that worst\-case protection comes at substantial expected\-case cost\. The Pareto frontier demonstrates that this tradeoff is mild in the PSA setting: an 18% improvement in adversarial robustness requires surrendering only 10\.7% of nominal placement quality, and the bulk of the gain is available for under 3%\. The concavity of the frontier means that even conservative planners who accept only small quality sacrifices can achieve meaningful robustness improvements by choosingε\\varepsilonnear the knee of the curve\.

## 9Experiment 5: Dual\-Theater Case Studies with Sensitivity Analysis

### 9\.1Motivation

Experiments 1–4 validate the PSA optimization stack on a single, stylized five\-location Indo\-Pacific theater\. Experiment 5 extends the evaluation along three dimensions necessary for operational credibility\. First, we ask whether the performance hierarchy observed in prior experiments generalizes to a structurally distinct theater: the European theater, with a different geographic footprint, asset mix, and threat distribution\. Second, we conduct a sensitivity analysis that characterizes how the Wasserstein ambiguity\-set radiusε\\varepsilongoverns DRSO performance and how robust each optimizer’s posture directive is to perturbations in scenario weights, an operationally relevant concern when intelligence assessments carry uncertainty about their own confidence\. Third, we trace DRSO posture directives as an A2/AD contested zone expands in radius from a fixed threat center, providing the planner with a regime\-transition readout that identifies when, and by how much, the inland repositioning imperative activates\.

### 9\.2Experimental Setup

##### Indo\-Pacific theater\.

The theater spans\|ℒ\|=8\|\\mathcal\{L\}\|=8named basing locations: Kadena AB \(v=0\.95v=0\.95\), Andersen AFB \(v=0\.90v=0\.90\), Osan AB \(v=0\.88v=0\.88\), MCAS Iwakuni \(v=0\.85v=0\.85\), Camp H\.M\. Smith \(v=0\.80v=0\.80\), Diego Garcia \(v=0\.78v=0\.78\), Misawa AB \(v=0\.75v=0\.75\), and Clark AB \(v=0\.72v=0\.72\), each with capacityc=10c=10\.\|𝒜\|=20\|\\mathcal\{A\}\|=20assets span four types: 12 aircraft, 2 maintenance crews \(cyber nodes\), 4 munitions \(radar\), and 2 fuel depots, with fixed readinessra=0\.85r\_\{a\}=0\.85\.

##### European theater\.

\|ℒ\|=6\|\\mathcal\{L\}\|=6locations: Rzeszów \(v=0\.90v=0\.90\), Ramstein AB \(v=0\.88v=0\.88\), Vilnius \(v=0\.85v=0\.85\), Riga \(v=0\.82v=0\.82\), Gdańsk \(v=0\.78v=0\.78\), and Szczecin \(v=0\.75v=0\.75\), each with capacityc=8c=8\.\|𝒜\|=15\|\\mathcal\{A\}\|=15assets: 8 aircraft \(armored brigades\), 4 munitions \(air defense\), and 3 fuel depots\.

##### Threat scenarios\.

Training usesS=20S=20theater\-specific scenarios per seed\. In the Indo\-Pacific, each scenario is drawn from a mixture: drone\-swarm \(40%, high threat on coastal bases Kadena, Iwakuni, Osan, and Clark\), IRBM \(35%, deep\-hub targeting Andersen, Camp Smith, and Diego Garcia\), and feint \(25%, diffuse low\-level\)\. In Europe: combined\-arms \(50%, high threat on eastern bases Rzeszów, Vilnius, and Riga\), hybrid \(30%, moderate diffuse\), and air campaign \(20%, targeting command nodes Ramstein, Szczecin, and Gdańsk\)\. Out\-of\-sample \(OOS\) evaluation usesSoos=50S\_\{\\mathrm\{oos\}\}=50adversarially generated scenarios \(ε=0\.9\\varepsilon=0\.9\), distinct from training\. All results are averaged overNseed=10N\_\{\\mathrm\{seed\}\}=10independent seeds\.

##### Baselines\.

Five policies are compared:

- •Greedy: assigns assets to the highest\-value location with remaining capacity, ignoring threat\.
- •EV\(Expected Value\): uses a single scenario at the mean threat level across all training scenarios,τ¯ℓ=S−1​∑sτℓ\(s\)\\bar\{\\tau\}\_\{\\ell\}=S^\{\-1\}\\\!\\sum\_\{s\}\\tau\_\{\\ell\}^\{\(s\)\}\. This is the certainty\-equivalent deterministic baseline\.
- •SAA\(Sample Average Approximation\): uses the fullS=20S=20training scenarios with equal weights \(ws=1w\_\{s\}=1\); ranks locations by their sample\-average expected value\.
- •Minimax: ranks locations by the worst\-case \(minimum over scenarios\) adjusted value: v^ℓmm=mins∈𝒮⁡vℓ​\(1−τℓ\(s\)\),\\hat\{v\}\_\{\\ell\}^\{\\mathrm\{mm\}\}=\\min\_\{s\\in\\mathcal\{S\}\}\\;v\_\{\\ell\}\\left\(1\-\\tau\_\{\\ell\}^\{\(s\)\}\\right\),\(18\)then assigns assets greedily to the highest\-ranked location\.
- •DRSO: the RobustCEV optimizer from Section[4\.2](https://arxiv.org/html/2608.05256#S4.SS2)withpobs=0\.70p\_\{\\mathrm\{obs\}\}=0\.70,γ=1\.0\\gamma=1\.0, and up to 10 iterations\.

### 9\.3Part A: Cross\-Theater Baseline Comparison

Figure[4](https://arxiv.org/html/2608.05256#S9.F4)and Table[7](https://arxiv.org/html/2608.05256#S9.T7)report worst\-case SWR on the OOS adversarial scenarios for all five baselines in both theaters\.

Table 7:Experiment 5A: cross\-theater comparison of five baselines\. Worst\-case SWR evaluated against 50 adversarial OOS scenarios\. 95% CI over 10 seeds\. Latency = wall\-clock ms from theater state to ranked posture directive\.![Refer to caption](https://arxiv.org/html/2608.05256v1/x4.png)Figure 4:Experiment 5A: worst\-case SWR for five baselines in the Indo\-Pacific \(a\) and European \(b\) theaters\. Error bars show±1​σ\\pm 1\\sigmaover 10 seeds\. All stochastic methods substantially outperform Greedy\. Minimax leads on raw worst\-case SWR; DRSO’s advantage lies in adversarial stability rather than peak performance\.The first\-order result replicates across both theaters: all four stochastic methods substantially outperform the greedy baseline\. In the Indo\-Pacific, SWR improvement over Greedy ranges from\+0\.0479\+0\.0479\(DRSO,\+17\.6%\+17\.6\\%\) to\+0\.0660\+0\.0660\(Minimax,\+24\.3%\+24\.3\\%\)\. In the European theater:\+0\.0225\+0\.0225\(DRSO,\+7\.8%\+7\.8\\%\) to\+0\.0473\+0\.0473\(Minimax,\+16\.3%\+16\.3\\%\)\. The smaller gaps in the European theater reflect its more compact geometry: six locations across a narrower geographic spread create fewer high\-contrast placement decisions, compressing the value of scenario\-awareness\.

##### EV equals SAA \(structural identity\)\.

In both theaters, EV and SAA produce*identical*SWR\. This is a mathematical identity for any greedy\-over\-scenarios optimizer under equal scenario weights\. Because the objective is linear in scenarios, ranking byvℓ​\(1−τ¯ℓ\)v\_\{\\ell\}\(1\-\\bar\{\\tau\}\_\{\\ell\}\)\(EV\) is equivalent to ranking byS−1​∑svℓ​\(1−τℓ\(s\)\)S^\{\-1\}\\\!\\sum\_\{s\}v\_\{\\ell\}\(1\-\\tau\_\{\\ell\}^\{\(s\)\}\)\(SAA\) when allws=1w\_\{s\}=1\. The two baselines nominally differ in how they represent uncertainty but collapse to the same location ranking under the linearity of the greedy assignment step\.

##### Minimax vs\. DRSO\.

Minimax outperforms DRSO on raw worst\-case SWR in the Indo\-Pacific \(0\.33780\.3378vs\.0\.31970\.3197\) because it is designed exclusively for worst\-case protection\. DRSO hedges against the full adversarial distribution rather than a single worst scenario, accepting a modestly lower worst\-case floor in exchange for placement stability under scenario weight perturbation, demonstrated in Part B\. This is the correct expression of each optimizer’s objective, not a deficiency of DRSO\.

### 9\.4Part B: Sensitivity Analysis

#### 9\.4\.1Wasserstein Radius Sweep

Figure[5](https://arxiv.org/html/2608.05256#S9.F5)\(panel a\) and Table[8](https://arxiv.org/html/2608.05256#S9.T8)report worst\-case SWR as a function of the Wasserstein radius proxyε∈\{0\.0,0\.1,0\.2,0\.4,0\.6,0\.8\}\\varepsilon\\in\\\{0\.0,0\.1,0\.2,0\.4,0\.6,0\.8\\\}, withS=20S=20training scenarios evaluated againstSoos=50S\_\{\\mathrm\{oos\}\}=50adversarial OOS scenarios atεoos=0\.9\\varepsilon\_\{\\mathrm\{oos\}\}=0\.9\.

Both EV and DRSO improve monotonically withε\\varepsilon\. Atε=0\\varepsilon=0, training scenarios concentrate near the nominal distribution and both methods approach near\-greedy performance \(≈0\.272\\approx 0\.272Indo\-Pacific;≈0\.290\\approx 0\.290European\)\. Atε=0\.8\\varepsilon=0\.8, the training set is broadly dispersed and both yield SWR of0\.3330\.333–0\.3340\.334\(Indo\-Pacific\) and0\.3300\.330\(European\)\. EV and DRSO remain within0\.0020\.002of each other at every testedε\\varepsilon, confirming that the adversarial reweighting in DRSO does not systematically alter location ranking when both are trained on the same scenario distribution\. DRSO’s operational advantage is in stability, not raw SWR\.

Table 8:Experiment 5B: Wasserstein radius sensitivity and assignment stability under±30%\\pm 30\\%scenario weight perturbation\. Worst\-case SWR averaged over 10 seeds\. Stability = fraction of assets with unchanged location\.
#### 9\.4\.2Assignment Stability under Scenario Weight Perturbation

*Assignment stability*is the fraction of assets whose location assignment is unchanged between a base run and a perturbed run in which each scenario weight is independently multiplied by a factor drawn from𝒰​\(0\.70,1\.30\)\\mathcal\{U\}\(0\.70,1\.30\)\(Section[9\.2](https://arxiv.org/html/2608.05256#S9.SS2)\)\.

In the Indo\-Pacific, DRSO achieves perfect stability \(1\.0001\.000\) while EV drops to0\.9000\.900: the adversarial iteration anchors DRSO’s placement at a configuration that is already near\-optimal across a wide range of scenario weight perturbations\. In the European theater the result reverses: EV achieves1\.0001\.000and DRSO0\.9530\.953\. This asymmetry is interpretable\. The European threat distribution has lower variance \(a structured eastern\-front gradient with less cross\-location mixing\) than the Indo\-Pacific mixture, making EV’s location ranking entirely insensitive to minor weight changes\. DRSO’s iterative best\-response occasionally shifts a single asset under the perturbed weights, reducing stability by4\.7%4\.7\\%without materially reducing SWR\. Both methods exceed the85%85\\%operational stability threshold in both theaters\.

![Refer to caption](https://arxiv.org/html/2608.05256v1/x5.png)Figure 5:Experiment 5B sensitivity analysis\.\(a\)Worst\-case SWR vs\. Wasserstein radius proxyε\\varepsilonfor EV and DRSO in both theaters \(solid = Indo\-Pacific, dashed = European\)\. Both methods improve monotonically and track within 0\.002 of each other at every setting\.\(b\)Assignment stability under±30%\\pm 30\\%scenario weight perturbation\. DRSO is perfectly stable in the Indo\-Pacific \(1\.000\) vs\. EV at 0\.900; the roles reverse in the European theater due to lower threat variance\. Dashed line marks the 85% operational stability threshold\.

### 9\.5Part C: A2/AD Contested Zone Radius Sweep

##### Setup\.

An A2/AD threat center is fixed at\(30∘​N,130∘​E\)\(30^\{\\circ\}\\mathrm\{N\},\\;130^\{\\circ\}\\mathrm\{E\}\), the East China Sea, and the contested zone radius sweepsr∈\{300,500,600,900,1500\}r\\in\\\{300,500,600,900,1500\\\}km\. At each radius, A2/AD training scenarios assign locations within radiusrrthreat levels proportional to haversine proximity \(base0\.50\+0\.40⋅proximity±0\.100\.50\+0\.40\\cdot\\text\{proximity\}\\pm 0\.10\); locations outside receive low background threat \(0\.050\.05–0\.200\.20\)\. OOS evaluation is performed against 50*fixed*adversarial theater scenarios \(generated by the Indo\-Pacific mixture model, not the A2/AD model\) so that SWR is comparable across radii and reflects true placement quality rather than training\-OOS alignment\.

Table 9:Experiment 5C: A2/AD contested zone radius sweep \(Indo\-Pacific, threat center30∘​N,130∘​E30^\{\\circ\}\\mathrm\{N\},\\;130^\{\\circ\}\\mathrm\{E\}\)\. SWR evaluated against fixed theater adversarial scenarios\. Coverage = fraction of assets outside the contested zone\.The haversine distances from\(30∘​N,130∘​E\)\(30^\{\\circ\}\\mathrm\{N\},\\;130^\{\\circ\}\\mathrm\{E\}\)place Kadena AB at≈461\\approx 461km, MCAS Iwakuni at≈506\\approx 506km, and Osan AB at≈835\\approx 835km\. Radii 500, 600, and 900 km therefore correspond to successive entry of Kadena \(1 location\), Iwakuni \(2 locations\), and Osan \(3 locations\) into the contested zone\.

##### Results\.

Figure[6](https://arxiv.org/html/2608.05256#S9.F6)and Table[9](https://arxiv.org/html/2608.05256#S9.T9)summarize the sweep\.

![Refer to caption](https://arxiv.org/html/2608.05256v1/x6.png)Figure 6:Experiment 5C: A2/AD contested zone radius sweep\.\(a\)Worst\-case SWR vs\. radius for all five baselines\. DRSO and EV jump from 0\.469 to 0\.486 as the radius crosses 900 km \(Osan AB enters the zone\); Greedy and Minimax remain flat throughout\.\(b\)Survivable coverage \(fraction of assets outside contested zone\)\. DRSO and EV maintain 100% coverage at all radii\. Greedy drops to 50% atr=500r=500km when Kadena AB, its highest\-value location, enters the zone\. Vertical lines mark radius thresholds at which an additional location enters the contested zone\.Atr=300r=300km no location is within the contested zone and all methods produce equivalent SWR \(0\.4660\.466\) because A2/AD training scenarios carry no geographic footprint\. The first separation occurs atr=500r=500km \(Kadena enters\): greedy coverage drops immediately to0\.5000\.500because Kadena is the highest\-strategic\-value location and greedy fills it first, regardless of threat\. DRSO and EV, having observed A2/AD training scenarios that mark Kadena as high\-threat, avoid it entirely, maintaining100%100\\%coverage\.

The key SWR transition occurs atr=900r=900km \(Osan enters; now three locations contested\): DRSO and EV SWR rise from0\.4690\.469to0\.4860\.486\(\+3\.7%\+3\.7\\%\) while Greedy and Minimax remain flat\. This improvement reflects that DRSO and EV’s inland repositioning, compelled by the expanding A2/AD footprint in training, produces a posture that happens to outperform against the fixed adversarial OOS scenarios\. The three threshold transitions \(500 km, 600 km, 900 km\) provide the planner with an explicit set of regime crossings: as the contested zone matures, the scenario\-aware optimizer’s advantage over purely value\-maximizing placement grows monotonically\.

### 9\.6Discussion

##### Finding 1: Cross\-theater generalization is robust\.

The stochastic–greedy performance gap persists across both theaters and all five baselines\. In the Indo\-Pacific, stochastic methods improve SWR by17\.617\.6–24\.3%24\.3\\%over Greedy; in the European theater by7\.87\.8–16\.3%16\.3\\%\. The consistent ordering and direction validate that the advantage of scenario\-aware placement is not an artifact of the specific geography evaluated in Experiments 1–4\.

##### Finding 2: Minimax and DRSO serve distinct operational roles\.

Minimax achieves the highest raw worst\-case SWR \(\+24%\+24\\%over Greedy in Indo\-Pacific;\+16%\+16\\%in Europe\) by concentrating entirely on the worst\-case scenario\. DRSO accepts a lower worst\-case floor in exchange for placement stability under scenario weight perturbation, a property that Minimax does not provide\. A planner facing an adaptive adversary with high observation probability \(pobs≈0\.7p\_\{\\mathrm\{obs\}\}\\approx 0\.7\) should prefer DRSO; a planner facing deep distributional uncertainty but no strategic interaction \(no adaptive adversary\) may prefer Minimax\. These are genuinely different operating conditions, not a deficiency of either method\.

##### Finding 3: EV and SAA are equivalent under linear greedy\-over\-scenarios optimization\.

For any optimizer that ranks locations by weighted expected value, EV \(certainty\-equivalent mean\) and SAA \(equal\-weight sample average\) produce mathematically identical location rankings when scenario weights are equal\. This means DRSO’s operational advantage over “the deterministic plan” comes entirely from its adversarial reweighting iteration, not from the expansion from one representative scenario to many\. Planners seeking to distinguish EV from SAA must use a non\-linear optimizer that responds differently to concentrated versus distributed scenario weight\.

##### Finding 4: A2/AD avoidance is a threshold phenomenon\.

DRSO and EV both achieve zero contested\-zone placement at every tested radius\. Greedy’s survivable coverage drops immediately to50%50\\%once Kadena enters the zone atr=500r=500km: there is no graceful degradation, only a hard step\-change driven by Kadena’s position at the top of greedy’s value ranking\. For the planner, this means the case for scenario\-aware optimization has a threshold character: it becomes operationally critical precisely when a high\-value, forward base enters the contested zone, and the cost of remaining with a greedy policy is the immediate loss of half the contested\-zone asset coverage\.

## 10Statistical Analysis

All Experiment 1 results are validated using pairedtt\-tests \(two\-tailed\) with Bonferroni correction acrossM=6M=6simultaneous comparisons, yielding a corrected significance threshold ofα∗=0\.05/6≈0\.0083\\alpha^\{\*\}=0\.05/6\\approx 0\.0083\. Tests are paired because each greedy–random pair shares the same simulation seed, eliminating initial\-state variance from the comparison\. Table[10](https://arxiv.org/html/2608.05256#S10.T10)summarizes the results\.

##### Placement\-invariant metrics\.

Readiness and sustainment cost show zero variance in paired differences: the*ReplenishmentPolicy*applies identically regardless of asset assignment, so both metrics are mathematically identical across strategies\. No significance test is applicable\.

##### Coverage and posture efficiency\.

Greedy coverage is significantly lower than random \(Δ​C=−0\.200\\Delta C=\-0\.200,t=−∞t=\-\\infty,p<0\.0001p<0\.0001\), as greedy deterministically leaves one location uncovered\. Posture efficiency follows directly: greedy is significantly less efficient than random \(Δ​E=−0\.056\\Delta E=\-0\.056,t=−18\.6t=\-18\.6,p<0\.0001p<0\.0001\), with the entire gap explained by the coverage difference\.

##### Threat\-environment sensitivity\.

The gap between uniform and skewed SWR is highly significant \(Δ​SWR=\+0\.261\\Delta\\mathrm\{SWR\}=\+0\.261,t=\+45\.7t=\+45\.7,p<0\.0001p<0\.0001\), confirming that the 57\.3% drop identified in Section[5\.5](https://arxiv.org/html/2608.05256#S5.SS5)is a structural property of the greedy assignment, not sampling noise\.

##### Variance decomposition\.

The lower panel of Table[10](https://arxiv.org/html/2608.05256#S10.T10)decomposes SWR variance into scenario\-seed variance \(outer\) and simulation\-seed variance \(inner\)\. Under uniform threat,ICC=0\.009\\mathrm\{ICC\}=0\.009: nearly all variance is attributable to initial\-state randomness and the choice of scenario set is immaterial\. Under skewed threat,ICC=0\.167\\mathrm\{ICC\}=0\.167: scenario\-set choice accounts for a larger share of variance because different draws from the skewed distribution produce meaningfully different realized threat concentrations at the high\-value locations occupied by greedy\. Both ICCs are well below 0\.5, confirming that ten simulation seeds provide adequate coverage of initial\-state uncertainty across both threat conditions\.

Table 10:Statistical significance and variance decomposition att=10t=10\.Top: pairedtt\-tests \(two\-tailed\) comparing greedy vs\. random on each metric, and uniform vs\. skewed SWR\. Bonferroni\-correctedα∗=0\.0083\\alpha^\{\*\}=0\.0083\(6 comparisons\)\.∗denotesp<α∗p<\\alpha^\{\*\}\.†metric is placement\-invariant \(zero variance in paired differences\)\.Bottom: two\-level variance decomposition of SWR \(5 scenario seeds×\\times10 simulation seeds\)\. ICC near 0 = initial\-state randomness dominates; ICC near 1 = scenario\-set choice dominates\.

## 11Discussion

### 11\.1A Progressive Case Against Greedy Planning

The three experiments constitute a structured argument, not three independent results\. Experiment 1 establishes the failure mode: a greedy placement policy that maximizes peacetime strategic value deterministically concentrates assets at the four highest\-value locations, leaving one theater location persistently uncovered and exposing the entire force to value\-correlated adversarial targeting\. The 25\.1% posture efficiency penalty relative to random placement and the 57\.3% scenario\-weighted readiness collapse under skewed threat are not edge cases; they are structural consequences of the greedy objective function\. Experiment 2 then demonstrates that the CEV optimizer corrects exactly this failure: by replacing raw strategic value with scenario\-weighted expected value, it hedges coverage and threat exposure at the cost of no additional computational complexity, recovering up to 19\.8% efficiency in the threat regimes where greedy is most exposed\. Experiment 3 extends the argument to its logical conclusion: when the adversary is not passive but adapts its targeting in response to observable posture, the naive CEV optimizer becomes exploitable in the same way greedy is, and the RobustCEV extension is necessary to maintain a stable efficiency floor\. Taken together, the three experiments show that the PSA problem requires increasingly sophisticated planning as the adversary becomes increasingly rational, and that the optimizer presented here scales appropriately with that sophistication\.

### 11\.2Connecting Experimental Results to the OE Axioms

Each of the five Operational Environment axioms stated in Section[1](https://arxiv.org/html/2608.05256#S1)maps to a specific experimental finding\. Placement irreversibility is the root cause of Experiment 1’s central result: the 57\.3% SWR degradation under skewed threat cannot be repaired by the ReplenishmentPolicy because it is a geometric consequence of where assets are placed, not how they are maintained\. The ReplenishmentPolicy equalizes readiness across placement strategies by step 10, confirming that sustainment time\-criticality manifests as a degradation\-repair equilibrium that operates independently of initial placement decisions\. Threat non\-stationarity drives the entire Experiment 3 design: when the adversary observes the defender’s posture and redistributes scenario probability mass toward the most exposed locations, a static optimizer converges to a strategically predictable placement that the adversary can exploit\. The deceptive prior result, in which naive CEV efficiency collapses 64% as observation probability increases from zero to one, is the clearest quantitative illustration of this axiom\. Information asymmetry appears in the scenario count analysis of Experiment 2: the diminishing returns to scenario count beyondS=20S=20reflect the fact that a well\-chosen small scenario set captures the distributional signal that matters for placement, while additional scenarios primarily average out the signal rather than sharpen it\. Multi\-domain coupling and sustainment time\-criticality are embedded in the MDP state space design: the maintenance timerdad\_\{a\}and configuration modemam\_\{a\}per asset formalize the time\-critical, operationally coupled nature of readiness that static models ignore\.

### 11\.3Practical Implications for Scenario Library Design

The scenario count analysis in Experiment 2 has a concrete operational implication that extends beyond the experimental setting\. EVSS under skewed threat falls from 19\.8% atS=5S=5to 9\.9% atS=100S=100, with the largest marginal gain occurring betweenS=5S=5andS=20S=20\. BeyondS=20S=20, additional scenarios change EVSS by less than one percentage point\. This result suggests that planners do not need exhaustive threat libraries to capture the majority of the stochastic gain: a curated set of 5 to 20 high\-fidelity, geographically differentiated scenarios is sufficient\. The variance decomposition in Section[10](https://arxiv.org/html/2608.05256#S10)supports this conclusion from a different angle: under uniform threat, the intraclass correlation coefficient is 0\.009, meaning nearly all variance in scenario\-weighted readiness is attributable to initial\-state randomness rather than scenario\-set choice\. The ICC rises to 0\.167 under skewed threat, confirming that scenario\-set choice matters more when the adversary’s behavior is geographically concentrated\. Together, these results suggest a tiered scenario design protocol: under benign or operationally symmetric threat environments, a small scenario set is adequate; under adversarially concentrated threat environments, additional investment in scenario fidelity and diversity is warranted\.

### 11\.4Scope and Limitations of the Robustness Results

The Experiment 3 results warrant careful interpretation\. RobustCEV produces a strictly positive adversarial regret only under the deceptive prior, the condition in which the naive optimizer is genuinely lured into a strategically exploitable concentration\. Under uniform threat, both optimizers produce identical placements and zero regret by construction, which is the correct theoretical null\. Under skewed and adversarial priors at high observation probability, RobustCEV produces small negative regret, meaning the iterative best\-response loop overshoots the equilibrium and converges to a marginally worse placement than the naive optimizer\. This is a known limitation of finite\-iteration Stackelberg approximations: when the adversary’s update is aggressive relative to the number of iterations permitted, the iterative loop cycles rather than converges\. Two practical mitigations are available\. First, capping the observation probability estimate used by the optimizer at a conservative value prevents the most aggressive adversarial updates from destabilizing the loop\. Second, replacing the greedy\-assignment inner loop with a minimax formulation would guarantee non\-negative regret by construction, at the cost of increased computational complexity\. The deceptive prior result, which is the operationally most relevant condition for a sophisticated adversary employing information operations to shape defender expectations, remains robust across all tested observation probabilities and is the paper’s primary robustness finding\.

### 11\.5Relationship to Prior Work

The 19\.8% CEV improvement over greedy under skewed threat is comparable in magnitude to the 31% improvement reported by Rettke et al\.\[[10](https://arxiv.org/html/2608.05256#bib.bib7)\]for approximate dynamic programming over greedy dispatch in aerial MEDEVAC, providing external validation that scenario\-aware optimization yields operationally meaningful gains over greedy baselines in military asset allocation settings\. The two\-stage stochastic programming structure of the CEV optimizer follows the template established by Salmerón and Apte\[[11](https://arxiv.org/html/2608.05256#bib.bib5)\]and recently validated for US Army aviation assets by Nelson et al\.\[[8](https://arxiv.org/html/2608.05256#bib.bib6)\], with the EVSS metric providing a direct comparison to prior published results in this literature\. The statistical analysis confirms that the reported gains are not sampling artifacts: coverage and efficiency gaps between greedy and random placement are significant atp<0\.0001p<0\.0001after Bonferroni correction, and the SWR gap between threat conditions achieves att\-statistic of\+45\.7\+45\.7, placing it well beyond any reasonable significance threshold\. Unlike prior military asset allocation work, this paper additionally demonstrates the value of adversarial robustness under Bayesian adversary models, a direction motivated by the JADC2 doctrine emphasis on AI systems that can operate under adversarial information environments\[[7](https://arxiv.org/html/2608.05256#bib.bib1)\]\.

### 11\.6Future Directions

Three extensions are most immediately motivated by the experimental findings\. First, the iterative best\-response convergence failure at high observation probability suggests that a minimax inner loop, replacing the greedy\-over\-scenarios assignment, would provide a convergence guarantee and likely improve robustness under concentrated priors\. Second, the ReplenishmentPolicy used throughout this paper is rule\-based and placement\-invariant; replacing it with a learned sustainment policy, trained via proximal policy optimization on the MDP defined in Section[2](https://arxiv.org/html/2608.05256#S2), could recover readiness gains that the current rule\-based policy leaves on the table, as suggested by analogous results in pharmaceutical supply chain management\[[12](https://arxiv.org/html/2608.05256#bib.bib11)\]\. Third, the current MDP formulation treats assets as independent conditioned on location, an assumption that breaks down when assets have complementary capabilities or when joint readiness across asset types determines mission feasibility\. Extending the state space to capture joint asset configurations, along the lines of multi\-agent RL formulations for workforce optimization\[[2](https://arxiv.org/html/2608.05256#bib.bib10)\], is the natural next step toward an operationally deployable PSA engine\.

## References

- \[1\]A\. Ben\-Tal, L\. El Ghaoui, and A\. Nemirovski\(2009\)Robust optimization\.Princeton Series in Applied Mathematics,Princeton University Press,Princeton, NJ\.Cited by:[§8\.3\.2](https://arxiv.org/html/2608.05256#S8.SS3.SSS2.Px3.p1.3)\.
- \[2\]K\. Eissa, R\. Prasad, S\. Mohan, A\. Kapoor, D\. Comaniciu, and V\. Singh\(2025\)Multi\-agent reinforcement learning with long\-term performance objectives for service workforce optimization\.arXiv preprint arXiv:2503\.01069\.Cited by:[§11\.6](https://arxiv.org/html/2608.05256#S11.SS6.p1.1),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px6.p1.3)\.
- \[3\]S\. El Sawah, H\. Turan, L\. Gordon, and M\. Ryan\(2023\)A decision support methodology to support military asset and resource planning\.Journal of Simulation18\(2\),pp\. 154–179\.External Links:[Document](https://dx.doi.org/10.1080/17477778.2023.2165460)Cited by:[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px2.p1.1)\.
- \[4\]S\. E\. Evans and G\. Steeger\(2018\)Deployment\-to\-dwell metrics and supply\-based force sustainment\.Journal of Defense Analytics and Logistics2\(1\),pp\. 2–21\.External Links:[Document](https://dx.doi.org/10.1108/JDAL-05-2017-0009)Cited by:[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px3.p1.1)\.
- \[5\]Institute for Defense Analyses\(2021\)Sustainment and the logistics ecosystem: definitions, frameworks, and implications for platform readiness\.Technical reportTechnical ReportIDA\-D\-XXXX,Institute for Defense Analyses,Alexandria, VA\.Cited by:[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px2.p1.1)\.
- \[6\]X\. Li, W\. Pu, and X\. Zhao\(2021\)Towards learning behavior modeling of military logistics agent utilizing profit sharing reinforcement learning algorithm\.Applied Soft Computing\.Note:ScienceDirect PII S1568494621007055Cited by:[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px6.p1.3)\.
- \[7\]S\. Lingel, J\. Hagen, E\. Hastings, M\. Lee, M\. Sargent, M\. Walsh, L\. A\. Zhang, and D\. Blancett\(2020\)Joint all\-domain command and control for modern warfare: an analytic framework for identifying and developing artificial intelligence applications\.Technical reportTechnical ReportRR\-4408/1\-AF,RAND Corporation,Santa Monica, CA\.Cited by:[§11\.5](https://arxiv.org/html/2608.05256#S11.SS5.p1.3),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px1.p1.1)\.
- \[8\]R\. J\. Nelson, J\. Werner, M\. G\. Kay, R\. E\. King, B\. M\. McConnell, and K\. Thoney\-Barletta\(2025\)Two\-stage stochastic programming model of US army aviation allocation of utility helicopters to task forces\.Journal of Defense Modeling and Simulation\.Cited by:[§11\.5](https://arxiv.org/html/2608.05256#S11.SS5.p1.3),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px4.p1.1),[§4\.2](https://arxiv.org/html/2608.05256#S4.SS2.SSS0.Px2.p1.2)\.
- \[9\]W\. B\. Powell\(2011\)Approximate dynamic programming: solving the curses of dimensionality\.2nd edition,John Wiley & Sons,Hoboken, NJ\.Cited by:[§2\.1](https://arxiv.org/html/2608.05256#S2.SS1.SSS0.Px4.p1.3),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px5.p1.2)\.
- \[10\]A\. J\. Rettke, M\. J\. Robbins, and B\. J\. Lunday\(2016\)Approximate dynamic programming for the dispatch of military medical evacuation assets\.European Journal of Operational Research254\(3\),pp\. 824–839\.External Links:[Document](https://dx.doi.org/10.1016/j.ejor.2016.04.017)Cited by:[§11\.5](https://arxiv.org/html/2608.05256#S11.SS5.p1.3),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px5.p1.2)\.
- \[11\]J\. Salmerón and A\. Apte\(2010\)Stochastic optimization for natural disaster asset prepositioning\.Production and Operations Management19\(5\),pp\. 561–574\.External Links:[Document](https://dx.doi.org/10.1111/j.1937-5956.2009.01119.x)Cited by:[§11\.5](https://arxiv.org/html/2608.05256#S11.SS5.p1.3),[§2\.2](https://arxiv.org/html/2608.05256#S2.SS2.p1.1),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px4.p1.1),[§4\.2](https://arxiv.org/html/2608.05256#S4.SS2.SSS0.Px2.p1.2)\.
- \[12\]F\. Stranieri, C\. Kouki, W\. van Jaarsveld, and F\. Stella\(2025\)Classical and deep reinforcement learning inventory control policies for pharmaceutical supply chains with perishability and non\-stationarity\.arXiv preprint arXiv:2501\.10895\.Cited by:[§11\.6](https://arxiv.org/html/2608.05256#S11.SS6.p1.1),[§3](https://arxiv.org/html/2608.05256#S3.SS0.SSS0.Px7.p1.2)\.

Similar Articles

Robust Peak-cost Constrained Reinforcement Learning

arXiv cs.LG

This paper studies robust peak-cost constrained reinforcement learning, addressing limitations of standard CMDPs by controlling the maximum cost along a trajectory and considering dynamics uncertainty. The authors show zero duality gap may not hold and propose a surrogate optimization framework with robust value estimation.

Utility-Constrained Policy Optimization

arXiv cs.LG

This paper introduces a simple yet powerful methodology for Utility-Constrained MDPs (UCMDPs) that enables risk-sensitive constraints without fixing constraint limits in advance, outperforming baselines on Safety Gymnasium benchmarks.

Smart predict-then-robustly-optimize

arXiv cs.LG

This paper proposes a robust variant of smart predict-then-optimize that accounts for feature perturbations, providing a convex surrogate with theoretical guarantees and demonstrating superior performance over standard methods.