基于概率鲁棒性的可解释通用对抗性扰动针对基于深度强化学习的入侵检测系统

arXiv cs.LG 论文

摘要

论文提出PX-UAP方法,该方法结合概率鲁棒性与可解释人工智能,生成针对基于深度强化学习的入侵检测系统的通用对抗性扰动,并在实验中展示了攻击有效性的提升。

arXiv:2609.30605v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the universality objective of UAPs, i.e., PR quantifies the prevalence of misclassification in the input space, making it a natural signal for guiding UAP generation. Hence, we propose PR-based UAP, which represents the first integration of an explicit PR-driven objective into generating UAPs against DRL-based IDS. Building on this formulation, we introduce PX-UAP, which leverages explainable artificial intelligence (XAI) to guide perturbation shaping under realistic domain constraints, and provides a rigorous theoretical analysis of its design. Extensive experiments demonstrate that PX-UAP consistently outperforms state-of-the-art UAP methods in attack effectiveness.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:40

# Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
Source: [https://arxiv.org/html/2609.30605](https://arxiv.org/html/2609.30605)
Hongsen ZhangAffiliation:WMGAffiliation:University of WarwickAffiliation:Coventry, CV4 7ALEmail:[Hongsen\.Zhang@warwick\.ac\.uk](mailto:)Lu ZhangAffiliation:Wolfson School of Mechanical, Electrical and Manufacturing EngineeringAffiliation:Loughborough UniversityAffiliation:Loughborough, LE11 3TUEmail:[L\.Zhang12@lboro\.ac\.uk](mailto:)Mingjing XuAffiliation:Faculty of Science and EngineeringAffiliation:Swansea universityAffiliation:Swansea, SA2 8PPEmail:[2596785@swansea\.ac\.uk](mailto:)Yi ZhangAffiliation:WMGAffiliation:University of WarwickAffiliation:Coventry, CV4 7ALEmail:[Yi\.Zhang\.16@warwick\.ac\.uk](mailto:)Gregory EpiphaniouAffiliation:WMGAffiliation:University of WarwickAffiliation:Coventry, CV4 7ALEmail:[Gregory\.Epiphaniou@warwick\.ac\.uk](mailto:)Carsten MapleAffiliation:WMGAffiliation:University of WarwickAffiliation:Coventry, CV4 7ALEmail:[CM@warwick\.ac\.uk](mailto:)

###### Abstract

Deep reinforcement learning \(DRL\) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems \(IDS\) to adversarial threats such as universal adversarial perturbations \(UAPs\), which apply a single input\-agnostic perturbation to degrade detection performance across traffic\. Probabilistic Robustness \(PR\), as a post\-hoc evaluation metric, provides a principled, population\-level measure of adversarial impact that conceptually aligns with the universality objective of UAPs, i\.e\., PR quantifies the prevalence of misclassification in the input space, making it a natural signal for guiding UAP generation\. Hence, we propose PR\-based UAP, which represents the first integration of an explicit PR\-driven objective into generating UAPs against DRL\-based IDS\. Building on this formulation, we introduce PX\-UAP, which leverages explainable artificial intelligence \(XAI\) to guide perturbation shaping under realistic domain constraints, and provides a rigorous theoretical analysis of its design\. Extensive experiments demonstrate that PX\-UAP consistently outperforms state\-of\-the\-art UAP methods in attack effectiveness\.

## 1Introduction

Deep learning \(DL\) has been extensively applied to intrusion detection systems \(IDS\) to model high\-dimensional network\-flow features and improve detection performance\[[1](https://arxiv.org/html/2609.30605#bib.bib14),[2](https://arxiv.org/html/2609.30605#bib.bib35)\]\. More recently, deep reinforcement learning \(DRL\) has emerged as an alternative paradigm, enabling IDS to learn adaptive decision policies through interaction\-driven reward feedback and to better cope with dynamic network environments and evolving attack patterns\[[3](https://arxiv.org/html/2609.30605#bib.bib36),[4](https://arxiv.org/html/2609.30605#bib.bib15)\]\. Despite these advantages, DL\-based IDS are susceptible to adversarial attacks, where carefully crafted input perturbations can induce erroneous predictions\[[5](https://arxiv.org/html/2609.30605#bib.bib16)\]\. For DRL\-based IDS, this vulnerability has been empirically demonstrated in prior work by applying adversarial examples \(AEs\), exposing critical security risks at inference time\[[6](https://arxiv.org/html/2609.30605#bib.bib21)\]\. In a typical deployment \(Fig\.[1](https://arxiv.org/html/2609.30605#S2.F1)\), an IDS continuously monitors network flows and classifies each flow as benign or malicious\. An adversary can generate adversarial perturbations by manipulating observed traffic through practical actions, such as packet padding, packet fragmentation, traffic injection, or transmission\-timing adjustments, thereby perturbing flow\-level features and inducing misclassifications\[[7](https://arxiv.org/html/2609.30605#bib.bib22)\]\.

Universal adversarial perturbations \(UAPs\) constitute a practical and severe threat, by using a single input\-agnostic perturbation to induce misclassification across a large fraction of samples, in contrast to instance\-specific attacks that require per\-sample optimization\[[8](https://arxiv.org/html/2609.30605#bib.bib2)\]\. In intrusion detection, a UAP corresponds to a universal flow\-level manipulation applied at inference time, enabling an adversary to consistently degrade detection performance across observed traffic\. This universality makes UAP attacks even more serious in real\-world IDS deployments, as they can be precomputed and applied efficiently across traffic samples to compromise IDS even under latency constraints and limited attacker resources\[[9](https://arxiv.org/html/2609.30605#bib.bib23),[10](https://arxiv.org/html/2609.30605#bib.bib1)\]\.

Probabilistic Robustness \(PR\) has been proposed as a principled framework for characterizing robustness properties of DL models by quantifying the probability that a model’s prediction remains invariant within a specified perturbation region\[[11](https://arxiv.org/html/2609.30605#bib.bib3)\]\. Unlike conventional adversarial robustness that focuses on worst\-case guarantees, PR provides a distribution\-aware, population\-level perspective on adversarial impact\. Recent studies incorporate PR into adversarial\-learning pipelines in the computer vision area, providing instance\-specific AE analysis\[[12](https://arxiv.org/html/2609.30605#bib.bib12),[13](https://arxiv.org/html/2609.30605#bib.bib4)\]and leveraging PR as a post\-hoc selection or evaluation criterion in adversarial training\[[14](https://arxiv.org/html/2609.30605#bib.bib5),[15](https://arxiv.org/html/2609.30605#bib.bib6)\]\. Importantly, PR can capture adversarial effectiveness as a probability over the data distribution rather than on individual samples\. This conceptually aligns with the objective of UAPs, i\.e\., both quantify the portion of misclassification in the input space: PR focuses on a neighborhood of a given input and UAPs target the whole input space, which makes PR a natural signal for guiding UAP generation\. Hence, for the first time, we propose an explicit PR\-driven objective for generating effective UAP attacks against DRL\-based IDS\.

Explainable artificial intelligence \(XAI\)\[[16](https://arxiv.org/html/2609.30605#bib.bib13)\]provides feature\-level attributions that reveal how DL models make decisions and has been adopted in IDS to analyze feature relevance\[[17](https://arxiv.org/html/2609.30605#bib.bib9),[18](https://arxiv.org/html/2609.30605#bib.bib8)\]\. In adversarial settings, such attributions can guide perturbations toward the most influential features under a limited adversarial budget\. However, XAI\-guided adversarial attacks remain largely unexplored for UAPs, particularly in IDS scenarios\. Hence, in this work, to enhance the attack performance of the proposed PR\-based UAP, we further advance XAI to guide the perturbation refinement in UAP generation\. We also provide a theoretical analysis of PX\-UAP, showing that the refinement admits an entropy\-regularized perturbation budget allocation interpretation and is theoretically aligned with the targeted universal evasion objective\. In summary, the key contributions of this work include: ∙\\bulletWe propose a Probabilistic Robustness–driven Universal Adversarial Perturbation \(PR\-based UAP\) framework for attacking DRL\-based IDS, introducing the concept of PR into UAP generation for the first time\. ∙\\bulletBuilding on the PR\-based UAP framework, we further enhance attack performance by leveraging XAI, denoted asPX\-UAP, which is the first approach to use feature attributions to shape a universal perturbation attack strategy in IDS\. ∙\\bulletWe provide a rigorous theoretical justification for PX\-UAP, by establishing a constrained universal evasion objective and an attribution\-guided, entropy\-regularized perturbation scheme with a first\-order improvement guarantee\. ∙\\bulletWe conduct extensive experiments demonstrating that PX\-UAP consistently outperforms state\-of\-the\-art UAP methods for intrusion detection in terms of attack effectiveness\.

## 2Related Works

### 2\.1Adversarial Attack on Network IDS

In network IDS applications, network traffic is modeled as flow\-based samples, each representing a network connection with temporal behaviors aggregated into statistical features\[[19](https://arxiv.org/html/2609.30605#bib.bib19)\]\. Adversaries can induce IDS misclassification by perturbing such features at inference time, thereby undermining the reliability of IDS predictions\[[20](https://arxiv.org/html/2609.30605#bib.bib17),[21](https://arxiv.org/html/2609.30605#bib.bib18)\]\. Beyond instance\-specific attacks \(e\.g\., FGSM, PGD\)\[[22](https://arxiv.org/html/2609.30605#bib.bib24),[23](https://arxiv.org/html/2609.30605#bib.bib25)\], recent works also demonstrate IDS vulnerability to input\-agnostic UAPs\[[24](https://arxiv.org/html/2609.30605#bib.bib38),[9](https://arxiv.org/html/2609.30605#bib.bib23)\]\. However, these two UAP\-on\-IDS researches do not consider realistic domain constraints, including protocol compliance, valid feature ranges, and inter\-feature mathematical dependencies\[[25](https://arxiv.org/html/2609.30605#bib.bib27),[7](https://arxiv.org/html/2609.30605#bib.bib22)\]\. Zhang et al\.\[[10](https://arxiv.org/html/2609.30605#bib.bib1)\]address this gap by integrating inter\-feature relationships with feature grouping to maintain validity during UAP generation\. Despite the above progress and the inherent universality target of UAPs, no existing work explicitly incorporates a distribution\-level objective into the UAP design for a principled generalizable adversarial impact across the distribution\.

### 2\.2Probabilistic Robustness

PR is a distribution\-aware robustness notion defined over a bounded perturbation region\[[11](https://arxiv.org/html/2609.30605#bib.bib3)\]\. In contrast to worst\-case robustness, which considers themaximumclassification loss of any AE in the region, PR quantifies theprobabilitythat the model’s prediction remains unchanged over the region\[[14](https://arxiv.org/html/2609.30605#bib.bib5)\]\. The concept of PR was first introduced by Webb et al\.\[[11](https://arxiv.org/html/2609.30605#bib.bib3)\]for AE detection in a white\-box threat model and later evolved into a principled framework for robustness evaluation, with extensions to black\-box settings\[[12](https://arxiv.org/html/2609.30605#bib.bib12),[26](https://arxiv.org/html/2609.30605#bib.bib28)\]and broader multimodal tasks\[[13](https://arxiv.org/html/2609.30605#bib.bib4)\]\. However, these works are limited to using PR for robustness assessment\. More recently, Zhang et al\.\[[15](https://arxiv.org/html/2609.30605#bib.bib6)\]for the first time combined PR with adversarial attacks to evaluate the attack effectiveness of candidate AEs for adversarial training\. To the best of our knowledge, despite the conceptual compatibility between PR’s distribution\-aware characterization of adversarial impact and UAPs’ universality objective, existing research has used PR only in a post\-hoc manner to assess robustness or attack performance, leaving PR\-driven UAP design unexplored\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure1.png)

Figure 1:\(a\): Adversarial scenario for a network\-flow\-based IDS\. \(b\): Comparison between worst\-case robustness \(i\) and probabilistic robustness \(ii\), adapted from\[[15](https://arxiv.org/html/2609.30605#bib.bib6)\]\.

## 3Methodology

In this section, we describe the proposed PR\-based UAP, which is the first attempt to integrate PR into UAP generation against DRL\-based IDS under realistic domain constraints\. We first outline the rationale and intuition behind the proposed PR\-based UAP, and then present the detailed attack framework and algorithm\. Building upon the PR\-based UAP, we introduce PX\-UAP, which leverages XAI signals to guide perturbation shaping and further enhance attack performance, representing the first XAI\-guided UAP approach in IDS\. Furthermore, we provide a rigorous theoretical justification to substantiate the effectiveness of the proposed attack method\.

### 3\.1Threat model

In this work, the IDS task is formulated as a binary benign–malicious classification problem\. We consider awhite\-boxattacker who has full access to the deployed IDS model \(architecture, parameters, and gradients\) and the feature preprocessing pipeline for the publicly available CICIDS2018 dataset\[[27](https://arxiv.org/html/2609.30605#bib.bib20)\], seeking a targeted evasion attack that mapsmaliciousflows to thebenigndecision\. The attacker is constrained toL2L\_\{2\}\-boundedfeature\-space perturbations with realistic domain constraints\. Specifically, following Zhang et al\.\[[10](https://arxiv.org/html/2609.30605#bib.bib1)\], we partition features into Modified Features \(MF\), Related Features \(RF\), and Unmodified Features \(UF\): only MF can be perturbed within feasible feature ranges, then the values of RF are recalculated based on the modified MF, and the UF remain unchanged\. This MF\-RF recalculation rule is enforced throughout this work\. More details are provided in Appx\.[C](https://arxiv.org/html/2609.30605#A3)\.

### 3\.2Rationale behind the proposed PR\-based UAP

Given an inputxxwith the ground\-truth labelyy, we consider perturbationsδ∼Pr\(⋅∣x\)\\delta\\sim\\Pr\(\\cdot\\mid x\)constrained within anLpL\_\{p\}\-norm ball of radiusϵ\\epsilon\(i\.e\.,‖δ‖p≤ϵ\\\|\\delta\\\|\_\{p\}\\leq\\epsilon\), yielding the AExadv=x\+δx\_\{\\mathrm\{adv\}\}=x\+\\delta\. PR is defined as follows\[[14](https://arxiv.org/html/2609.30605#bib.bib5)\]:

PR\(x,ϵ\)=𝔼δ∼Pr\(⋅∣x\)‖δ‖≤ϵ\[𝕀\{fθ\(x\+δ\)=y\}\(x\+δ\)\]\\mathrm\{PR\}\(x,\\epsilon\)=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}\\delta\\sim\\Pr\(\\cdot\\mid x\)\\\\ \\\|\\delta\\\|\\leq\\epsilon\\end\{subarray\}\}\\\!\\left\[\\mathbb\{I\}\_\{\\\{f\_\{\\theta\}\(x\+\\delta\)=y\\\}\}\(x\+\\delta\)\\right\]\(1\)where the indicator𝕀\{fθ\(x\+δ\)=y\}\\mathbb\{I\}\_\{\\\{f\_\{\\theta\}\(x\+\\delta\)=y\\\}\}equals11if classifierfθf\_\{\\theta\}predicts ground truth labelyyforxadvx\_\{\\mathrm\{adv\}\}, and00 otherwise\. Intuitively, in contrast to worst\-case robustness \(Fig\.[1](https://arxiv.org/html/2609.30605#S2.F1)\.b\(i\)\),PR⁡\(x,ϵ\)\\mathrm\{PR\}\(x,\\epsilon\)measures the fraction of the perturbation region for whichfθ​\(x\+δ\)=yf\_\{\\theta\}\(x\+\\delta\)=y, with perturbation distributionPr\(⋅∣x\)\\Pr\(\\cdot\\mid x\)\[[12](https://arxiv.org/html/2609.30605#bib.bib12),[15](https://arxiv.org/html/2609.30605#bib.bib6)\], as illustrated in Fig\.[1](https://arxiv.org/html/2609.30605#S2.F1)\.b\(ii\)\. Thus, a smallerPR⁡\(x,ϵ\)\\mathrm\{PR\}\(x,\\epsilon\)implies that more data samples are misclassified, indicating a larger AE\-dominatedϵ\\epsilon\-bounded neighborhood ofxx\. Motivated by this interpretation, prior studies\[[14](https://arxiv.org/html/2609.30605#bib.bib5),[15](https://arxiv.org/html/2609.30605#bib.bib6)\]employ a PR\-based variablekkas a post\-hoc criterion to select the most effective AE among multiple candidates generated by Projected Gradient Descent attack\[[28](https://arxiv.org/html/2609.30605#bib.bib7)\], as formalized in Eq\.[2](https://arxiv.org/html/2609.30605#S3.E2):

k:PR⁡\(x\+δ,k\)=0,∀δ∈PGD⁡\(x\),‖δ‖≤ϵ\.k\\;:\\;\\mathrm\{PR\}\(x\+\\delta,k\)=0,\\quad\\forall\\,\\delta\\in\\mathrm\{PGD\}\(x\),\\ \\\|\\delta\\\|\\leq\\epsilon\.\(2\)
As illustrated in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(a\), two PGD runs yield perturbationsδ1\\delta\_\{1\}andδ2\\delta\_\{2\}under anϵ\\epsilonbudget\. Centered at the corresponding AEsx\+δx\+\\delta, the conditionPR⁡\(x\+δ,k\)=0\\mathrm\{PR\}\(x\+\\delta,k\)=0implies that the neighborhood with radiuskk\(denoted asℬ⁡\(x\+δ,k\)\\mathcal\{B\}\(x\+\\delta,k\)later\) contains no label\-preserving samples, i\.e\., all input points within it are AEs\. Hence, a largerkkindicates a stronger attack, inducing misclassification over a larger neighborhood\. However, in such an approach,kkis a post\-hoc metric and does not participate in the perturbation generation process\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure2.png)

Figure 2:\(a\): Illustration of mapping between PR\-based variablekkand Misclassification\-Guaranteed Region in the UAP setting\. Two examplesδ1\\delta\_\{1\}andδ2\\delta\_\{2\}are provided in \(i\) and \(ii\), respectively\. UAP byδ1\\delta\_\{1\}\(larger k\) induces a larger size Misclassification\-Guaranteed Regionℬ⁡\(x\+δ1,k1\)\\mathcal\{B\}\(x\+\\delta\_\{1\},k\_\{1\}\), while smaller sizeℬ⁡\(x\+δ2,k2\)\\mathcal\{B\}\(x\+\\delta\_\{2\},k\_\{2\}\)byδ2\\delta\_\{2\}\.\(b\): Main Procedures of the PR\-based UAP: \(1\) Attack initialization\. \(2\) AE candidate selection\. \(3\)kk\-driven optimization\.Now we present the rationale for the proposed PR\-based UAP attack by providing an intuitive explanation of the compatibility between the PR\-driven optimization and UAP generation\. Specifically, usingδ1\\delta\_\{1\}andδ2\\delta\_\{2\}in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(a\) as representative examples, a UAP targets misclassification for most data samplesxxin the distribution𝒳\\mathcal\{X\}with a shared perturbation\[[8](https://arxiv.org/html/2609.30605#bib.bib2)\]\. Letδn\\delta\_\{n\}\(n∈\{1,2\}n\\in\\\{1,2\\\}\) be such a shared perturbation\. Applyingδn\\delta\_\{n\}to any input induces adistribution\-level translationx↦x\+δnx\\mapsto x\+\\delta\_\{n\}\(shaded regions in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(a\)\), which shifts the dashed\-circle neighborhood aroundxxto the solid\-circle neighborhood aroundx\+δnx\+\\delta\_\{n\}\(follow the two arrows, respectively\)\. In this context, PR also provides adistribution\-awareview of adversarial impact, naturally aligning with the distribution\-level target and translation mapping of UAPs\. By Eq\.[2](https://arxiv.org/html/2609.30605#S3.E2), the notion ofPR⁡\(x\+δ,k\)=0\\mathrm\{PR\}\(x\+\\delta,k\)=0implies∀z∈ℬ⁡\(x\+δ,k\),fθ​\(z\)≠y\\forall z\\in\\mathcal\{B\}\(x\+\\delta,k\),\\\\ \\,f\_\{\\theta\}\(z\)\\neq y, i\.e\., all input points withinℬ⁡\(x\+δ,k\)\\mathcal\{B\}\(x\+\\delta,k\)are AEs, which yields a converse implication:

\(∀z∈ℬ\(x\+δn,kn\),fθ\(z\)≠y\)∧\(ℬ\(x,kn\)\+δn⊆ℬ\(x\+δn,kn\)\)\\displaystyle\(\\forall z\\in\\mathcal\{B\}\(x\+\\delta\_\{n\},k\_\{n\}\),\\ f\_\{\\theta\}\(z\)\\neq y\)\\wedge\\ \(\\mathcal\{B\}\(x,k\_\{n\}\)\+\\delta\_\{n\}\\subseteq\\mathcal\{B\}\(x\+\\delta\_\{n\},k\_\{n\}\)\)\(3\)⇒∀x′∈ℬ\(x,kn\),fθ\(x′\+δn\)≠y\.\\displaystyle\\Rightarrow\\ \\forall x^\{\\prime\}\\in\\mathcal\{B\}\(x,k\_\{n\}\),\\ f\_\{\\theta\}\(x^\{\\prime\}\+\\delta\_\{n\}\)\\neq y\.Given the UAP mapping fromℬ⁡\(x,kn\)\\mathcal\{B\}\(x,k\_\{n\}\)toℬ⁡\(x\+δn,kn\)\\mathcal\{B\}\(x\+\\delta\_\{n\},k\_\{n\}\), any pointx′∈ℬ⁡\(x,kn\)x^\{\\prime\}\\in\\mathcal\{B\}\(x,k\_\{n\}\)will satisfyfθ​\(x′\+δn\)≠yf\_\{\\theta\}\(x^\{\\prime\}\+\\delta\_\{n\}\)\\neq y\. Therefore,δn\\delta\_\{n\}induces amisclassification\-guaranteedregionℬ⁡\(x,kn\)\\mathcal\{B\}\(x,k\_\{n\}\)aroundxx\(dashed circles in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(a\)\)\. A largerkkyields a broader affected neighborhood aroundxx, as shown by the larger dashed circle forδ1\\delta\_\{1\}compared with the circle forδ2\\delta\_\{2\}, thereby a larger portion of the input space𝒳\\mathcal\{X\}will be misclassified, which directly aligns with the objective of UAP generation\. Thus, it is reasonable and intuitive to incorporate PR as an explicit objective in the optimization process for UAP generation\.

### 3\.3The proposed PR\-based UAP attack

In this section, we detail the PR\-based UAP attack, which reformulates the PR\-derived variablekkas an explicit optimization objective and integrates it into the UAP update\. This turns the post\-hoc discrete selection of maximalkkinto a direct continuous optimization problem, yielding a more principled UAP generation process\. Our method comprises two algorithms: a UAP generation framework \(details shown in Alg\.[2](https://arxiv.org/html/2609.30605#alg2)in Appx\.[B](https://arxiv.org/html/2609.30605#A2)\) and PR\-based gradient computation \(Alg\.[1](https://arxiv.org/html/2609.30605#alg1)\)\. Specifically, Alg\.[1](https://arxiv.org/html/2609.30605#alg1)computes the PR\-based update gradient and forwards it to Alg\.[2](https://arxiv.org/html/2609.30605#alg2)in Appx\.[B](https://arxiv.org/html/2609.30605#A2), which iteratively updates the perturbation to obtain the final PR\-based UAP\.

Alg\.[2](https://arxiv.org/html/2609.30605#alg2)in Appx\.[B](https://arxiv.org/html/2609.30605#A2)displays the UAP generation framework\. In line 1, we sample a seed setS​e​e​d​s​e​t⊂T​rSeedset\\subset Trand initializeu​a​puap=𝟎∈ℝd=\\mathbf\{0\}\\in\\mathbb\{R\}^\{d\}withd=dim\(x\)d=\\dim\(x\)\. During UAP generation,S​e​e​d​s​e​tSeedsetis shuffled at the beginning of each iteration in line 4\. The target classifierCCdenotes the pre\-trained DRL\-based agent\. For eachx∈S​e​e​d​s​e​tx\\in Seedset, the predictionsC⁡\(x\)C\(x\)andC⁡\(x\+u​a​p\)C\(x\+uap\)are computed at line 5\. If the predicted labels are identical \(line 6\), the currentu​a​puapfails to foolCConxx, and we then perform one update step by generating an update perturbationp​t​bptb\. Specifically, we compute the gradientggvia the PR\-based gradient calculation in Alg\.[1](https://arxiv.org/html/2609.30605#alg1)at line 7\. Usingggas the update direction, we apply an element\-wise mask to restrict modifications to MF only, and formp​t​bptbby anL2L\_\{2\}\-normalized FGM\-style step with magnitudeϵ\\epsilon\[[5](https://arxiv.org/html/2609.30605#bib.bib16)\]\. The accumulatedu​a​puapis then projected onto theL2L\_\{2\}ball to maintain the perturbation budget\. The loop continues by updating the iteration numberi​t​e​r​\_​n​u​miter\\\_numuntil reaching the max iterationm​a​x​\_​i​t​e​rmax\\\_iter, after which the finalu​a​puapis returned \(line 12\)\.

Algorithm 1UAP Direction via the Gradient ofk∗k^\{\*\}Input:target modelCC, data instancexxwith labelyy, perturbation budgetϵ\\epsilon, PGD attack numberNN, PGD iteration stepnstepn\_\{\\mathrm\{step\}\}, PGD step sizeα\\alpha, SGD max iterationS​G​D​\_​n​u​mSGD\\\_num Output:the gradient ofk∗k^\{\*\}with respect toxx:∇xk∗\\nabla\_\{x\}k^\{\*\}

1:Initialize

𝒳adv←\[\]\\mathcal\{X\}\_\{\\mathrm\{adv\}\}\\leftarrow\[\\;\],

𝒟←\[\]\\mathcal\{D\}\\leftarrow\[\\;\]
2:

𝒳adv​\[0\]←PGD⁡\(x,y,nstep,α,ϵ,ℓ2\)\\mathcal\{X\}\_\{\\mathrm\{adv\}\}\[0\]\\leftarrow\\mathrm\{PGD\}\(x,y,n\_\{\\mathrm\{step\}\},\\alpha,\\epsilon,\\ell\_\{2\}\)
3:while

len⁡\(𝒳adv\)<N\\mathrm\{len\}\(\\mathcal\{X\}\_\{\\mathrm\{adv\}\}\)<Ndo

4:

z∼𝒩⁡\(0,I\),r∼Uniform⁡\(0,ϵ\)z\\sim\\mathcal\{N\}\(0,I\),\\quad r\\sim\\mathrm\{Uniform\}\(0,\\epsilon\),

xinit←x\+r⋅z∥z∥2x\_\{\\mathrm\{init\}\}\\leftarrow x\+r\\cdot\\frac\{z\}\{\\lVert z\\rVert\_\{2\}\}
5:

ϵ′←Unif⁡\(0\.98∗ϵ,ϵ\)\\epsilon^\{\\prime\}\\leftarrow\\mathrm\{Unif\}\(0\.98\*\\epsilon,\\epsilon\),

α′←Unif⁡\(0\.9∗α,1\.1∗α\)\\alpha^\{\\prime\}\\leftarrow\\mathrm\{Unif\}\(0\.9\*\\alpha,\\,1\.1\*\\alpha\),

nstep′←Unif⁡\(0\.9∗nstep,1\.1∗nstep\)n^\{\\prime\}\_\{\\mathrm\{step\}\}\\leftarrow\\mathrm\{Unif\}\(0\.9\\,\*n\_\{\\mathrm\{step\}\},\\,1\.1\\,\*n\_\{\\mathrm\{step\}\}\)
6:Append

PGD⁡\(xinit,y,ϵ′,α′,nstep′\)\\mathrm\{PGD\}\(x\_\{\\mathrm\{init\}\},y,\\epsilon^\{\\prime\},\\alpha^\{\\prime\},n^\{\\prime\}\_\{\\mathrm\{step\}\}\)to

𝒳adv\\mathcal\{X\}\_\{\\mathrm\{adv\}\}
7:endwhile

8:foreach

xadv∈𝒳advx\_\{\\mathrm\{adv\}\}\\in\\mathcal\{X\}\_\{\\mathrm\{adv\}\}do

9:

x′←xadvx^\{\\prime\}\\leftarrow x\_\{\\mathrm\{adv\}\},

i←0i\\leftarrow 0
10:while

C⁡\(x′\)≠y∧i<S​G​D​\_​n​u​mC\(x^\{\\prime\}\)\\neq y\\;\\land\\;i<SGD\\\_numdo

11:

x′←x′−α​∇x′ℒ​\(Clogit​\(x′\),y\)‖∇x′ℒ​\(Clogit​\(x′\),y\)‖2x^\{\\prime\}\\leftarrow x^\{\\prime\}\-\\alpha\\,\\frac\{\\nabla\_\{x^\{\\prime\}\}\\mathcal\{L\}\\big\(C\_\{\\mathrm\{logit\}\}\(x^\{\\prime\}\),y\\big\)\}\{\\left\\lVert\\nabla\_\{x^\{\\prime\}\}\\mathcal\{L\}\\big\(C\_\{\\mathrm\{logit\}\}\(x^\{\\prime\}\),y\\big\)\\right\\rVert\_\{2\}\},

i←i\+1i\\,\\leftarrow\\,i\+1
12:endwhile

13:Append

∥xadv−x′∥2\\lVert x\_\{\\mathrm\{adv\}\}\-x^\{\\prime\}\\rVert\_\{2\}to

𝒟\\mathcal\{D\}
14:endfor

15:

k∗←max⁡𝒟k^\{\*\}\\leftarrow\\max\\mathcal\{D\}
16:return

∇xk∗\\nabla\_\{x\}k^\{\*\}

Next, we detail Alg\.[1](https://arxiv.org/html/2609.30605#alg1), which computes the PR\-driven update gradient for our PR\-based UAP\. As shown in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(b\), it consists of three main steps\. Concretely, aiming to maximize the radiuskkin Eq\.[2](https://arxiv.org/html/2609.30605#S3.E2), we first employ the candidate\-search method proposed by Zhang et al\.\[[15](https://arxiv.org/html/2609.30605#bib.bib6)\]to identify the AE corresponding to local maximum loss in Steps \(1\)–\(2\):

\(1\) Attack initialization:Multiple random PGD attacks are launched on inputxxto obtain a set of AE candidates𝒳adv\\mathcal\{X\}\_\{\\mathrm\{adv\}\}\. \(2\) AE candidate selection:For each AE candidate, the distance fromxadvx\_\{\\mathrm\{adv\}\}to the decision boundary is estimated and treated as the correspondingkkin Eq\.[2](https://arxiv.org/html/2609.30605#S3.E2)\. The candidate with the largestkk\(denotedk∗k^\{\*\}\) is selected as the starting point for the subsequent step\.

The key difference from Zhang et al\.\[[15](https://arxiv.org/html/2609.30605#bib.bib6)\]lies in Step \(3\)\. The discrete selection ofk∗k^\{\*\}may be constrained by limited search resolution and finite exploration steps\. Therefore, motivated by our rationale, selecting the largest value in𝒟\\mathcal\{D\}ask∗k^\{\*\}chooses the candidate that causes the largestmisclassification\-guaranteedregionℬ⁡\(x,k∗\)\\mathcal\{B\}\(x,k^\{\*\}\), inducing the most promising local region for UAP generation\. Consequently, we formulatek∗k^\{\*\}as an explicit optimization objective and adopt a continuouskk\-driven optimization scheme, enabling more effective and principled radius maximization\.

\(3\)kk\-driven optimization:The gradient ofk∗k^\{\*\}with respect toxx,∇xk∗\\nabla\_\{x\}k^\{\*\}is calculated and served as the update gradientggin Alg\.[2](https://arxiv.org/html/2609.30605#alg2)\(Eq\.[4](https://arxiv.org/html/2609.30605#S3.E4)\)\. By perturbingxxalong the direction ofggto maximize the radiusk∗k^\{\*\}, a more effective perturbation can be obtained, corresponding toδ′,k′\\delta^\{\\prime\},k^\{\\prime\}andxadv′x\_\{\\mathrm\{adv\}\}^\{\\prime\}in Fig\.[2](https://arxiv.org/html/2609.30605#S3.F2)\(a\)\.

∇xk∗​\(x\)=∇xk​\(x,δ∗\),\\displaystyle\\nabla\_\{x\}k^\{\*\}\(x\)=\\nabla\_\{x\}k\\bigl\(x,\\delta^\{\*\}\\bigr\),\(4\)where​δ∗≜arg⁡maxδ⁡k⁡\(x,δ\)​s\.t\.​δ∈PGD⁡\(x\),‖δ‖≤ϵ\.\\displaystyle\\text\{where \}\\delta^\{\*\}\\triangleq\\arg\\max\_\{\\delta\}\\,k\(x,\\delta\)\\ \\text\{s\.t\. \}\\delta\\in\\mathrm\{PGD\}\(x\),\\ \\\|\\delta\\\|\\leq\\epsilon\.Specifically, Alg\.[1](https://arxiv.org/html/2609.30605#alg1)illustrates the computation of∇xk∗\\nabla\_\{x\}k^\{\*\}\. After initialization, multiple PGD attacks from random starting pointsxinitx\_\{\\mathrm\{init\}\}are launched \(line 4\), wherexinitx\_\{\\mathrm\{init\}\}is sampled byz∼𝒩⁡\(0,I\)z\\sim\\mathcal\{N\}\(0,I\),r∼Uniform⁡\(0,ϵ\)r\\sim\\mathrm\{Uniform\}\(0,\\epsilon\), andxinit=x\+r⋅z‖z‖2x\_\{\\mathrm\{init\}\}=x\+r\\cdot\\frac\{z\}\{\\\|z\\\|\_\{2\}\}\. We further randomize the attack hyperparameters\(ϵ′,α′,nstep′\)\(\\epsilon^\{\\prime\},\\alpha^\{\\prime\},n^\{\\prime\}\_\{\\mathrm\{step\}\}\)within predefined ranges to generate an AE set𝒳adv\\mathcal\{X\}\_\{\\mathrm\{adv\}\}\(lines 5\-6\)\. For each AE candidatexadvx\_\{\\mathrm\{adv\}\}, we iteratively update the current pointx′x^\{\\prime\}via a gradient descent\-based method until it becomes non\-adversarial or the maximum number of stepsS​G​D​\_​n​u​mSGD\\\_numis reached \(lines 10–12\)\. This procedure yields a new pointx′x^\{\\prime\}that can be regarded as an approximate “projection” ofxadvx\_\{\\mathrm\{adv\}\}onto the decision boundary\. At line 13, we compute∥xadv−x′∥2\\lVert x\_\{\\mathrm\{adv\}\}\-x^\{\\prime\}\\rVert\_\{2\}as the radiuskkand store thesekkvalues in a distance set𝒟\\mathcal\{D\}, from which we take the maximum ask∗k^\{\*\}\(line 15\)\. Thus, we compute∇xk∗\\nabla\_\{x\}k^\{\*\}\(obtained by automatic differentiation sincek∗k^\{\*\}is computed without hard clipping orarg⁡max\\arg\\max\) and return it as the update gradientggto Alg\.[2](https://arxiv.org/html/2609.30605#alg2)\. Thenxxis perturbed along the direction of∇xk∗\\nabla\_\{x\}k^\{\*\}, withk∗k^\{\*\}serving as the optimization objective to be maximized, ultimately generating the proposed PR\-based UAP\.

### 3\.4PX\-UAP attack

This section introduces the PX\-UAP attack, which leverages XAI to guide perturbation shaping and further enhance attack effectiveness\. Before presenting the technical details, we first provide the rationale behind PX\-UAP\. XAI comprises techniques that improve the interpretability of ML/DL models by providing human\-understandable explanations, typically in the form of feature\-level attributions that identify decision\-critical inputs for a given prediction\[[29](https://arxiv.org/html/2609.30605#bib.bib29),[30](https://arxiv.org/html/2609.30605#bib.bib30)\]\. In adversarial settings, such attributions offer a principled mechanism for prioritizing the perturbations under a limited adversarial budget, which is especially relevant for IDS where realistic domain constraints restrict arbitrary feature modifications\[[18](https://arxiv.org/html/2609.30605#bib.bib8),[17](https://arxiv.org/html/2609.30605#bib.bib9)\]\.

Under these constraints, XAI\-derived signals enable the identification of influential and permissible features, thereby facilitating more effective UAPs\. Prior studies show that attribution\-guided pertur\- bations facilitate per\-instance AE generation in IDS\[[17](https://arxiv.org/html/2609.30605#bib.bib9),[31](https://arxiv.org/html/2609.30605#bib.bib31)\]\. However, these efforts do not consider realistic network domain constraints and have not explored XAI’s role in input\-agnostic UAP generation\. Motivated by this observation, we incorporate XAI to guide the proposed PR\-based UAP, resulting in the PX\-UAP attack, representing the first XAI\-guided universal adversarial attack in IDS\.

Now we introduce PX\-UAP \(Alg\.[3](https://arxiv.org/html/2609.30605#alg3)in Appx\.[B](https://arxiv.org/html/2609.30605#A2)\), which calculates the feature\-level importance scores and selects effective perturbations based on detection performance\. Specifically, the PR\-based UAPP​R​u​a​pPRuapis first recalculated for realistic constraints and projected onto theL2L\_\{2\}budgetϵ\\epsilon\(line 1\)\. We use the predictions of realistic AEs\(T​r\+P​R​u​a​p\)\(Tr\+PRuap\)and clean samplesT​rTrto compute the false negative rate \(FNR\) as a reference performance metric for later comparison \(line 2\)\. A higher FNR indicates that more malicious samples are misclassified as benign, aligning with the adversary’s objective in this work\. At line 3, a batchℬ\\mathcal\{B\}withBBinstances is sampled from the training setT​rTrto estimate feature\-level importance\. In this research, thebenignclass is used as the target class, and Integrated Gradients\[[32](https://arxiv.org/html/2609.30605#bib.bib10)\]asXAImtd\\operatorname\{XAImtd\}to compute feature attribution scores, as shown in Eq\.[5](https://arxiv.org/html/2609.30605#S3.E5):

IG⁡\(x,x′\)\\displaystyle\\mathrm\{IG\}\(x;x^\{\\prime\}\)≜\(x−x′\)⊙∫01∇xClogit​\(x′\+α⁡\(x−x′\)\)​𝑑α\.\\displaystyle\\triangleq\(x\\\!\-\\\!x^\{\\prime\}\)\\odot\\textstyle\\int\\nolimits\_\{0\}^\{1\}\\nabla\_\{x\}\\,C\_\{\\mathrm\{logit\}\}\\\!\(x^\{\\prime\}\\\!\+\\\!\\alpha\(x\\\!\-\\\!x^\{\\prime\}\)\)\\,\\mathrm\{d\}\\alpha\.\(5\)IG is well suited to our white\-box threat model and the case of tabular data, as it provides gradient\-based attributions with low computational overhead\[[33](https://arxiv.org/html/2609.30605#bib.bib11)\]\. In practice, we use the discretemm\-step approximation in Eq\.[6](https://arxiv.org/html/2609.30605#S3.E6)\. The attribution scores for eachx∈ℬx\\in\\mathcal\{B\}are averaged to obtain an aggregated importance vectori​m​pimp\(lines 4–5\)\.

IG^​\(x,x′\)\\displaystyle\\widehat\{\\mathrm\{IG\}\}\(x;x^\{\\prime\}\)≜\(x−x′\)⊙1m∑t=1m∇xClogit\(x′\+tm\(x−x′\)\)\.\\displaystyle\\triangleq\(x\\\!\-\\\!x^\{\\prime\}\)\\odot\\frac\{1\}\{m\}\\textstyle\\sum\\limits\_\{t=1\}^\{m\}\\nabla\_\{x\}\\,C\_\{\\mathrm\{logit\}\}\\\!\(x^\{\\prime\}\\\!\+\\\!\\tfrac\{t\}\{m\}\(x\\\!\-\\\!x^\{\\prime\}\)\)\.\(6\)
In line 6, we normalizei​m​pimpby∥i​m​p∥∞\\lVert imp\\rVert\_\{\\infty\}, scaling its values to a standard range\[−1,1\]\[\-1,1\]\. This mitigates batch\-sampling variability in estimatingi​m​pimp, ensuring a standardized XAI\-based weighting and a fair comparison across attacks\. We then obtainP​X​u​a​pPXuapby XAI\-weightingP​R​u​a​pPRuapas in Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7):

imp←mask⊙i​m​p∥i​m​p∥∞,imp∈\[−1,1\]d,u\\displaystyle imp\\leftarrow\\mathrm\{mask\}\\odot\\textstyle\\frac\{imp\}\{\\lVert imp\\rVert\_\{\\infty\}\},\\;imp\\in\[\-1,1\]^\{d\},u←PRuap⊙Softmax\(imp\),PXuap←ϵ⋅u∥u∥2\.\\displaystyle\\leftarrow\\mathrm\{PRuap\}\\odot\\mathrm\{Softmax\}\(imp\),\\quad\\mathrm\{PXuap\}\\leftarrow\\epsilon\\cdot\\textstyle\\frac\{u\}\{\\lVert u\\rVert\_\{2\}\}\.\(7\)
Specifically, takingi​m​pimpas input, a feature mask is applied to restrict weighting only on the top\-nnmost important features\. We then transformi​m​pimpviaSoftmax⁡\(⋅\)\\operatorname\{Softmax\}\(\\cdot\)to obtain feature\-wise scaling coefficients, which nonlinearly map relative differences in importance scores and concentrate the limited perturbation budget on the most influential features\. For example,i​m​p=\[1,−1,0\]imp=\[1,\-1,0\]yieldsSoftmax⁡\(i​m​p\)≈\[0\.66,0\.09,0\.24\]\\operatorname\{Softmax\}\(imp\)\\approx\[0\.66,0\.09,0\.24\]\. After the projection and recalculation, the FNR ofP​X​u​a​pPXuapis computed at line 7\. Due to the stochasticity introduced by batch sampling, we compare the performance ofP​X​u​a​pPXuapwithP​R​u​a​pPRuapand select the one with a higher FNR as the final PX\-UAP\.

### 3\.5Theoretical analysis

This subsection provides a rigorous analysis of the PX\-UAP design in Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7)by \(i\) formalizing the targeted universal evasion objective under feasibility constraints, \(ii\) establishing IG as a principled feature\-priority signal, \(iii\) showing that Softmax weighting arises as an entropy\-regularized allocation rule, and \(iv\) giving a first\-order improvement guarantee for the resulting weighted perturbations\.

Formal objective: targeted universal evasion under feasibility\.The adversary targetsmalicious→\\tobenignevasion in white\-box setting\. Let𝒳1⊆T​r\\mathcal\{X\}\_\{1\}\\subseteq Trdenote the malicious subset \(xxwith ground\-truthy=1y=1\), and let the target class beyt=0y\_\{t\}=0\(benign\)\. We denote byu~≜recalculate⁡\(u,ϵ,ℓ2\)\\widetilde\{u\}\\triangleq\\operatorname\{recalculate\}\(u,\\epsilon,\\ell\_\{2\}\)the feasibility\-enforced perturbation returned by the realistic constraints in Alg\.[3](https://arxiv.org/html/2609.30605#alg3)\.

The targeted universal evasion objective can be written as

maxu∈ℝd𝔼x∼𝒳1\[𝕀\{C\(x\+u~\)=yt\}\]\\displaystyle\\max\_\{u\\in\\mathbb\{R\}^\{d\}\}\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{X\}\_\{1\}\}\\Big\[\\mathbb\{I\}\_\{\\\{C\(x\+\\widetilde\{u\}\)=y\_\{t\}\\\}\}\\Big\]s\.t\.​‖u~‖2≤ϵ,u~​satisfies realistic constraints\.\\displaystyle\\text\{s\.t\.\}\\;\\;\\\|\\widetilde\{u\}\\\|\_\{2\}\\leq\\epsilon,\\;\\;\\widetilde\{u\}\\ \\text\{satisfies realistic constraints\.\}\(8\)
Eq\.[8](https://arxiv.org/html/2609.30605#S3.E8)is aligned with the metric used in Alg\.[3](https://arxiv.org/html/2609.30605#alg3): maximizing the probability of predicting benign on malicious inputs corresponds to increasing FNR\. Moreover, the PR\-based construction \(Sec\.[3\.2](https://arxiv.org/html/2609.30605#S3.SS2)\) aims to enlarge the*misclassification\-guaranteed*neighborhood via the PR\-derived radiuskk\(Eq\.[2](https://arxiv.org/html/2609.30605#S3.E2)\)\. It increases the measure of inputs mapped into the adversarial region, consistent with the universal objective above\.

Integrated Gradients as a principled feature priority signal\.PX\-UAP uses IG\-based feature attributions asi​m​pimp\(Eq\.[5](https://arxiv.org/html/2609.30605#S3.E5)\)\. The following formalizes why IG is appropriate for deciding which features deserve more perturbation\.

###### Proposition 1\(Completeness of Integrated Gradients\)

LetF:ℝd→ℝF:\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}be differentiable andx,x′∈ℝdx,x^\{\\prime\}\\in\\mathbb\{R\}^\{d\}\. DefineIG⁡\(x,x′\)\\mathrm\{IG\}\(x;x^\{\\prime\}\)as in Eq\.[5](https://arxiv.org/html/2609.30605#S3.E5)\. Then IG satisfies the completeness identity∑i=1dIGi​\(x,x′\)=F⁡\(x\)−F⁡\(x′\)\\sum\_\{i=1\}^\{d\}\\mathrm\{IG\}\_\{i\}\(x;x^\{\\prime\}\)\\;=\\;F\(x\)\-F\(x^\{\\prime\}\)\. WhenF​\(⋅\)=Clogit​\(⋅\)F\(\\cdot\)=C\_\{\\mathrm\{logit\}\}\(\\cdot\)for the target classyty\_\{t\}, positiveIGi​\(x,x′\)\\mathrm\{IG\}\_\{i\}\(x;x^\{\\prime\}\)identifies features whose increase along the path fromx′x^\{\\prime\}toxxraises the target logit, and thus are natural candidates for priority perturbation in targeted evasion\.

In lines 5–6 in Alg\.[3](https://arxiv.org/html/2609.30605#alg3), we aggregate IG overℬ\\mathcal\{B\}to obtaini​m​pimp, so that \(by Prop\.[1](https://arxiv.org/html/2609.30605#Thmproposition1)\)i​m​pimpacts as an empirical estimate of which dimensions contribute to increasingClogit​\(⋅\)C\_\{\\mathrm\{logit\}\}\(\\cdot\)toward the target classyty\_\{t\}across the data distribution\.

Why Softmax weighting: an entropy\-regularized allocation view\.PX\-UAP transformsi​m​pimpinto nonnegative coefficients via Softmax and uses them to reallocate the perturbation ofP​R​u​a​pPRuap\(Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7)\)\. This can be derived as the unique solution of an entropy\-regularized linear allocation\. LetΔmask\\Delta\_\{\\mathrm\{mask\}\}denote masked probability:Δmask≜\{w∈ℝ≥0d:1⊤w=1,\(𝟏mask\)⊙w=𝟎\}\\Delta\_\{\\mathrm\{mask\}\}\\triangleq\\\{w\\in\\mathbb\{R\}^\{d\}\_\{\\geq 0\}:\\;\\mathbf\{1\}^\{\\top\}w=1,\\;\\;\(\\mathbf\{1\}\\mathrm\{mask\}\)\\odot w=\\mathbf\{0\}\\\}\.

Table 1:UAP baseline with loss functions for comparison with PR\-based UAP\.IndexBaseline UAPLoss function to maximize\(1\)PD\_mean\_UAP\[[34](https://arxiv.org/html/2609.30605#bib.bib32)\]log⁡\(∏i=1kmean⁡\(li\+eps\)\)\\log\\\!\\left\(\\prod\_\{i=1\}^\{k\}\\mathrm\{mean\}\(l\_\{i\}\+\\mathrm\{eps\}\)\\right\)where​li=activation⁡\(x\+δ\),eps=1×10−8\\text\{where \}l\_\{i\}=\\mathrm\{activation\}\(x\+\\delta\),\\ \\mathrm\{eps\}=1\\times 10^\{\-8\}\(2\)PD\_L2\_UAP\[[35](https://arxiv.org/html/2609.30605#bib.bib33)\]log⁡\(∏i=1k‖li‖2\+eps\)\\log\\\!\\left\(\\prod\_\{i=1\}^\{k\}\\left\\lVert l\_\{i\}\\right\\rVert\_\{2\}\+\\mathrm\{eps\}\\right\)where​li=activation⁡\(x\+δ\),eps=1×10−8\\text\{where \}l\_\{i\}=\\mathrm\{activation\}\(x\+\\delta\),\\ \\mathrm\{eps\}=1\\times 10^\{\-8\}\(3\)COSSIM\_L3\_UAP\[[36](https://arxiv.org/html/2609.30605#bib.bib34)\]−cossim⁡\(activationli​\(x\+δ\),activationli​\(x\)\)\-\\mathrm\{cossim\}\\\!\\left\(\\mathrm\{activation\}\_\{l\_\{i\}\}\(x\+\\delta\),\\ \\mathrm\{activation\}\_\{l\_\{i\}\}\(x\)\\right\),where​li=l3\\text\{where \}l\_\{i\}=l\_\{3\}\(4\)COSSIM\_L4\_UAP\[[36](https://arxiv.org/html/2609.30605#bib.bib34)\]−cossim⁡\(activationli​\(x\+δ\),activationli​\(x\)\)\-\\mathrm\{cossim\}\\\!\\left\(\\mathrm\{activation\}\_\{l\_\{i\}\}\(x\+\\delta\),\\ \\mathrm\{activation\}\_\{l\_\{i\}\}\(x\)\\right\),where​li=l4\\text\{where \}l\_\{i\}=l\_\{4\}\(5\)PCC\_UAP\[[10](https://arxiv.org/html/2609.30605#bib.bib1)\]PCCsim⁡\(activationli​\(x\+δ\),activationli​\(δ\)\)\\mathrm\{PCCsim\}\\\!\\left\(\\mathrm\{activation\}\_\{l\_\{i\}\}\(x\+\\delta\),\\ \\mathrm\{activation\}\_\{l\_\{i\}\}\(\\delta\)\\right\),where​li=l4\\text\{where \}l\_\{i\}=l\_\{4\}Define Shannon entropyH\(w\)≜−∑i=1dwilogwiH\(w\)\\triangleq\-\\sum\_\{i=1\}^\{d\}w\_\{i\}\\log w\_\{i\}forw∈Δmaskw\\in\\Delta\_\{\\mathrm\{mask\}\}\[[37](https://arxiv.org/html/2609.30605#bib.bib37)\]\. Then:

###### Lemma 1\(Softmax as the entropy\-regularized maximizer\)

Fix a score vectors∈ℝds\\in\\mathbb\{R\}^\{d\}and temperatureτ\>0\\tau\>0\. The optimization problemmaxw∈Δmask⁡⟨s,w⟩\+τ​H​\(w\)\\max\_\{w\\in\\Delta\_\{\\mathrm\{mask\}\}\}\\;\\;\\langle s,w\\rangle\+\\tau H\(w\)has a unique solution given by

wi⋆=exp⁡\(si/τ\)∑j:maskj=1exp\(sj/τ\)\\displaystyle w\_\{i\}^\{\\star\}=\\frac\{\\exp\(s\_\{i\}/\\tau\)\}\{\\sum\_\{j:\\,\\mathrm\{mask\}\_\{j\}=1\}\\exp\(s\_\{j\}/\\tau\)\}for alliwithmaski=1,wi⋆=0ifmaski=0,\\displaystyle\\text\{for all \}i\\text\{ with \}\\mathrm\{mask\}\_\{i\}=1,\\quad w\_\{i\}^\{\\star\}=0\\;\\text\{if \}\\mathrm\{mask\}\_\{i\}=0,\(9\)i\.e\., the Softmax distribution on the masked coordinates\.

Lemma[1](https://arxiv.org/html/2609.30605#Thmlemma1)provides an interpretation of the Softmax step in Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7): it allocates a unit budget across feasible features by trading off exploitation \(large⟨s,w⟩\\langle s,w\\rangle\) and diversity/stability \(large entropy\)\. In PX\-UAP we sets≡i​m​ps\\equiv imp\(after normalization and masking\), and useτ=1\\tau=1, yielding the weighting in Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7)\.Due to space limits, the rest of the theoretical analysis is provided in Appx\.[A](https://arxiv.org/html/2609.30605#A1)\.

## 4Experiments

In this section, we describe the experimental setup and provide the experimental results to validate the effectiveness of our algorithm design\. More results are provided in Appx\.[F](https://arxiv.org/html/2609.30605#A6)\.

### 4\.1Experimental Settings

In this study, we train a DQN agent\[[38](https://arxiv.org/html/2609.30605#bib.bib40)\]on CICIDS2018 dataset\[[27](https://arxiv.org/html/2609.30605#bib.bib20)\], keepingbenigntraffic as thebenignclass and merging all remaining categories into themaliciousclass\. Then we use the agent’s policy network as the IDS classifierCC, which is a fully\-connected MLP with four hidden layers \(64 units each\) and ReLU activations\.CCoutputsQ⁡\(x,a\)Q\(x,a\)fora∈\{0,1\}a\\in\\\{0,1\\\}corresponding to the two classes, and predictsy^=arg⁡maxa⁡Q⁡\(x,a\)\\hat\{y\}=\\arg\\max\_\{a\}Q\(x,a\)\. Each network flow is represented as a feature vectorx∈ℝ76x\\in\\mathbb\{R\}^\{76\}, where each feature is normalized to\[0,1\]\[0,1\], with one\-hot encoding for categorical features, and under\- \-sampling for class imbalance\. We evaluate the attack effectiveness in all experiments by FNR and accuracy \(ACC\), with the adversarial budgetϵ\\epsilonon the x\-axis under theL2L\_\{2\}constraint\. For allocating sufficient perturbation budget for effective attacks while preserving theimperceptibleproperty in the normalized feature space, we setϵ∈\[0\.04,0\.15\]\\epsilon\\in\[0\.04,\\,0\.15\]\. To mitigate potential bias from algorithmic ran\- \-domness, we report results averaged over 150 runs, ensuring a fair comparison across methods\. The hyperparameter settings for all algorithms are provided in Appx\.[D](https://arxiv.org/html/2609.30605#A4)\.

### 4\.2PR\-based UAP result

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure3.png)

Figure 3:\(a\): False Negative Rate \(FNR\) and \(b\): Accuracy of the proposed PR\-based UAP and other UAP baselines\. \(c\): FNR and \(d\): ACC of PX\-UAP with differentTopn\\mathrm\{Topn\}values\. \(e\): FNR and \(f\): ACC for ablation study\.We now report the results of the proposed PR\-based UAP against several state\-of\-the\-art UAP baselines, whose details are summarized in Tab\.[1](https://arxiv.org/html/2609.30605#S3.T1)\. As shown in Fig\.[3](https://arxiv.org/html/2609.30605#S4.F3)\(a\) and \(b\), the PR\-based UAP consistently outperforms all baselines across the evaluatedϵ\\epsilonrange\. Starting fromϵ=0\.04\\epsilon=0\.04, all evaluated UAP methods exhibit comparable effectiveness, withFNR≈0\.2\\mathrm\{FNR\}\\approx 0\.2andACC≈0\.93\\mathrm\{ACC\}\\approx 0\.93\. Asϵ\\epsilonincreases to 0\.09, all attacks become progressively more effective, while the PR\-based UAP remains the top performer\. Notably, PCC\_UAP \(pink\) rises sharply nearϵ=0\.09\\epsilon=0\.09and matches the proposed method atϵ=0\.09\\epsilon=0\.09, achievingFNR=0\.72\\mathrm\{FNR\}=0\.72andACC=0\.73\\mathrm\{ACC\}=0\.73\. After this point, the proposed PR\-based UAP continues to improve and stabilizes forϵ≥0\.13\\epsilon\\geq 0\.13\. Overall, these results demonstrate that our PR\-based UAP framework is effective and successfully surpasses other state\-of\-the\-art UAPs\.

### 4\.3PX\-UAP result

Building upon the PR\-based UAP, we report the results of PX\-UAP in Fig\.[3](https://arxiv.org/html/2609.30605#S4.F3)\(c\) and \(d\) with differentTopn\\mathrm\{Topn\}values, i\.e\., only theTopn\\mathrm\{Topn\}most important features ini​m​pimpare used for XAI\-based weighting in Alg\.[3](https://arxiv.org/html/2609.30605#alg3)\. During the early “cold\-start” stage \(ϵ∈\[0\.04,0\.07\]\\epsilon\\\!\\in\\\!\[0\.04,0\.07\]\), all PX\-UAP variants consistently strengthen the attack, with an≈\\approx10% improvement in effectiveness \(higher FNR and lower ACC\) over the PR\-based baseline\. In the mid\-range \(ϵ∈\[0\.07,0\.10\]\\epsilon\\in\[0\.07,0\.10\]\), the XAI\-induced benefit is comparatively limited\. OnlyTopn=3\\mathrm\{Topn\}\\\!=\\\!3\(red\) achieves a≈\\approx5% improvement aroundϵ≈0\.075\\epsilon\\approx 0\.075\. Later, XAI guidance further improves the PR\-based UAP, even though it already outperforms state\-of\-the\-art baselines inϵ∈\[0\.10,0\.15\]\\epsilon\\in\[0\.10,0\.15\]\. All four PX\-UAP variants provide additional gains in this “stable” range, withTopn=3\\mathrm\{Topn\}\\\!=\\\!3achieving the best overall performance\. Overall, the results show that incorporating XAI enables PX\-UAP to outperform the proposed PR\-based UAP, with particularly pronounced gains at small budgets and in the high\-ϵ\\epsilonrange where performance saturates\. In these cases, XAI helps allocate the limited perturbation budget to influential features, thereby improving attack effectiveness\.

### 4\.4Ablation Studies

As described in Sec\.[3\.3](https://arxiv.org/html/2609.30605#S3.SS3), Alg\.[1](https://arxiv.org/html/2609.30605#alg1)comprises three key components: \(i\) PGD\-based attack initialization, \(ii\)kk\-based selection of AE candidates, and \(iii\)k∗k^\{\*\}\-driven optimization\. To verify that the observed effectiveness stems from the combination of the components rather than any single factor, we conduct ablation studies by comparing the proposed PR\-based UAP with several ablated variants in which specific components are removed or modified\. Details and explanations of the variants are summarized in Tab\.[4](https://arxiv.org/html/2609.30605#A5.T4)in Appx\.[E](https://arxiv.org/html/2609.30605#A5)\. The results of the ablation study are reported in Fig\.[3](https://arxiv.org/html/2609.30605#S4.F3)\(e\) and \(f\)\.

Forϵ<0\.06\\epsilon<0\.06, all candidates exhibit similar performance, withFNR≈0\.28\\mathrm\{FNR\}\\approx 0\.28andACC≈0\.90\\mathrm\{ACC\}\\approx 0\.90\. Then, PR\-based UAP gradually outperforms the others across the remainingϵ\\epsilonrange\. Among the variants,AE selectionimproves effectiveness under FGM initialization, as evidenced by FUAP \(1\) versus KFUAP \(2\), but limited benefit under PGD initialization \(PUAP and KPUAP; \(3\) and \(4\)\)\. All var\- iants indicate that changingAttack initializationandAE selectionhas a limited impact on overall attack effectiveness\. Taking the best\-performing PR\-based UAP and the basic FUAP as the references, the ablation study indicates that combining all three components jointly in Alg\. 3 yields a substantial performance improvement, with thek∗k^\{\*\}\-driven optimizationplaying a critical role in this performance gain, which further validates the effectiveness of our proposed rationale and algorithm design\.

### 4\.5Additional Results

Beyond the results above, we conduct extensive additional experiments for a comprehensive evaluation of the proposed attack method, detailed results and analysis are shown in Appx\.[F](https://arxiv.org/html/2609.30605#A6)\. In summary: ∙\\bulletWe report the F1 scores of the above three experiments, providing a more balanced and comprehensive metric to evaluate the effectiveness of the proposed PR\-based UAP and PX\-UAP\. ∙\\bulletWe apply the proposed attack methods to two other dataset and further use different DRL agents and network architectures as target classifiers on CICIDS2018\. The results demonstrate the generalizability of the proposed attack across datasets, DRL agents, and model architectures\. ∙\\bulletWe further evaluate the proposed attack methods under a black\-box threat model by generating UAPs on a surrogate model and applying them to different target models\. The results demonstrate that the proposed attacks exhibit strong cross\-model transferability under the black\-box setting\. ∙\\bulletWe conduct an ablation study onϵ\\epsilon\-steps to verify that the significant performance gain of our proposed PR\-based UAP stems from the attack design itself rather than favorable step\-size selection\. The results show that PR\-based UAP consistently outperforms the baselines under both the fixed unifiedϵ\\epsilon\-step setting and the setting whereϵ\\epsilon\-steps are tuned for each method\.

## 5Conclusion

In this work, for the first time, we integrate the concept of PR into the UAP generation process and propose a novel UAP attack against DRL–based IDS\. We further enhance attack performance by advancing XAI techniques to guide the perturbation shaping for the proposed PR\-based UAP framework under realistic domain constraints, with providing a rigorous theoretical justification\. Extensive experiments demonstrating that PX\-UAP consistently outperforms state\-of\-the\-art UAP methods for intrusion detection in terms of attack effectiveness\.

## References

- \[1\]\(2019\)Survey of intrusion detection systems: techniques, datasets and challenges\.Cybersecurity2\(1\),pp\. 1–22\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1)\.
- \[2\]J\. A\. Alkharman, S\. A\. A\. Drawsheh, M\. M\. Al\-Khataybeh, Z\. B\. BaniYounes, N\. A\. Hamid Darawsheh, and H\. Alrashdan\(2024\)Cyber attacks and its implication to national security: the need for international law enforcement\.\.Pakistan Journal of Criminology16\(3\)\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1)\.
- \[3\]R\. S\. Sutton A\. G\. Bartoet al\.\(1998\)Reinforcement learning: an introduction\.Vol\.1,MIT press Cambridge\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1)\.
- \[4\]M\. Lopez\-Martin, B\. Carro, and A\. Sanchez\-Esguevillas\(2020\)Application of deep reinforcement learning to intrusion detection for supervised problems\.Expert Systems with Applications141,pp\. 112963\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1)\.
- \[5\]I\. J\. Goodfellow, J\. Shlens, and C\. Szegedy\(2014\)Explaining and harnessing adversarial examples\.arXiv preprint arXiv:1412\.6572\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1),[§3\.3](https://arxiv.org/html/2609.30605#S3.SS3.p2.1)\.
- \[6\]H\. Zhang and C\. Maple\(2023\)Deep reinforcement learning\-based intrusion detection in iot system: a review\.InInternational Conference on AI and the Digital Economy \(CADE 2023\),Vol\.2023,pp\. 88–97\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1)\.
- \[7\]I\. Debicha, B\. Cochez, T\. Kenaza, T\. Debatty, J\. Dricot, and W\. Mees\(2023\)Adv\-bot: realistic adversarial botnet attacks against network intrusion detection systems\.Computers & Security129,pp\. 103176\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[8\]S\. Moosavi\-Dezfooli, A\. Fawzi, O\. Fawzi, and P\. Frossard\(2017\)Universal adversarial perturbations\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 1765–1773\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p2.1),[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p3.2)\.
- \[9\]S\. Zhang, Y\. Xu, and X\. Xie\(2024\)Universal adversarial perturbations against machine learning\-based intrusion detection systems in industrial internet of things\.IEEE Internet of Things Journal\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[10\]H\. Zhang, L\. Zhang, G\. Epiphaniou, and C\. Maple\(2025\)A novel and practical universal adversarial perturbations against deep reinforcement learning based intrusion detection systems\.arXiv preprint arXiv:2511\.18223\.Cited by:[Appendix C](https://arxiv.org/html/2609.30605#A3.p1.1),[§1](https://arxiv.org/html/2609.30605#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.30605#S3.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.30605#S3.T1.p17.1.3.1.1.1.1)\.
- \[11\]S\. Webb, T\. Rainforth, Y\. W\. Teh, and M\. P\. Kumar\(2018\)A statistical approach to assessing neural network robustness\.arXiv preprint arXiv:1811\.07209\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1)\.
- \[12\]T\. Karim, T\. Furon, and M\. Rousset\(2023\)Gradient\-informed neural network statistical robustness estimation\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 323–334\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p1.2)\.
- \[13\]Y\. Zhang, Y\. Tang, W\. Ruan, X\. Huang, S\. Khastgir, P\. Jennings, and X\. Zhao\(2024\)ProTIP: probabilistic robustness verification on text\-to\-image diffusion models against stochastic perturbation\.InEuropean Conference on Computer Vision,pp\. 455–472\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1)\.
- \[14\]X\. Zhao\(2025\)Probabilistic robustness in deep learning: a concise yet comprehensive guide\.arXiv preprint arXiv:2502\.14833\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p1.2)\.
- \[15\]Y\. Zhang, Y\. Chen, Z\. Chen, W\. Ruan, X\. Huang, S\. Khastgir, and X\. Zhao\(2025\)Adversarial training for probabilistic robustness\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 1675–1685\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p3.1),[Figure 1](https://arxiv.org/html/2609.30605#S2.F1),[Figure 1](https://arxiv.org/html/2609.30605#S2.F1.5),[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p1.2),[§3\.3](https://arxiv.org/html/2609.30605#S3.SS3.p3.1),[§3\.3](https://arxiv.org/html/2609.30605#S3.SS3.p5.1)\.
- \[16\]M\. T\. Ribeiro, S\. Singh, and C\. Guestrin\(2016\)" Why should i trust you?" explaining the predictions of any classifier\.InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining,pp\. 1135–1144\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p4.1)\.
- \[17\]D\. L\. Marino, C\. S\. Wickramasinghe, and M\. Manic\(2018\)An adversarial approach for explainable ai in intrusion detection systems\.InIECON 2018\-44th Annual Conference of the IEEE Industrial Electronics Society,pp\. 3237–3243\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p4.1),[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p1.1),[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p2.1)\.
- \[18\]H\. Zhang, D\. Han, S\. Zhuang, Z\. Wang, J\. Sun, Y\. Liu, J\. Liu, and J\. Dong\(2025\)Explainable and transferable adversarial attack for ml\-based network intrusion detectors\.IEEE Transactions on Dependable and Secure Computing\.Cited by:[§1](https://arxiv.org/html/2609.30605#S1.p4.1),[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p1.1)\.
- \[19\]N\. Moustafa and J\. Slay\(2015\)UNSW\-nb15: a comprehensive data set for network intrusion detection systems \(unsw\-nb15 network data set\)\.In2015 military communications and information systems conference \(MilCIS\),pp\. 1–6\.Cited by:[§F\.2](https://arxiv.org/html/2609.30605#A6.SS2.p1.1),[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[20\]L\. Zhang, S\. Lambotharan, G\. Zheng, G\. Liao, B\. AsSadhan, and F\. Roli\(2022\)Attention\-based adversarial robust distillation in radio signal classifications for low\-power iot devices\.IEEE Internet of Things Journal10\(3\),pp\. 2646–2657\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[21\]L\. Zhang, S\. Lambotharan, G\. Zheng, G\. Liao, X\. Liu, F\. Roli, and C\. Maple\(2025\)Vision transformer with adversarial indicator token against adversarial attacks in radio signal classifications\.IEEE Internet of Things Journal\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[22\]Y\. Pacheco and W\. Sun\(2021\)Adversarial machine learning: a comparative study on contemporary intrusion detection datasets\.\.InICISSP,pp\. 160–171\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[23\]U\. A\. Ali, K\. Dogra, and S\. Sharma\(2025\)White\-box adversarial exploitation of nids: insights from fgsm, pgd, and c&w\.In2025 2nd International Conference on Computational Intelligence, Communication Technology and Networking \(CICTN\),pp\. 668–673\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[24\]R\. Sheatsley, N\. Papernot, M\. J\. Weisman, G\. Verma, and P\. McDaniel\(2022\)Adversarial examples for network intrusion detection systems\.Journal of Computer Security30\(5\),pp\. 727–752\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[25\]M\. Teuffenbach, E\. Piatkowska, and P\. Smith\(2020\)Subverting network intrusion detection: crafting adversarial examples accounting for domain\-specific constraints\.InInternational Cross\-Domain Conference for Machine Learning and Knowledge Extraction,pp\. 301–320\.Cited by:[§2\.1](https://arxiv.org/html/2609.30605#S2.SS1.p1.1)\.
- \[26\]T\. Zhang, W\. Ruan, and J\. E\. Fieldsend\(2022\)Proa: a probabilistic robustness assessment against functional perturbations\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 154–170\.Cited by:[§2\.2](https://arxiv.org/html/2609.30605#S2.SS2.p1.1)\.
- \[27\]I\. Sharafaldin, A\. H\. Lashkari, A\. A\. Ghorbani,et al\.\(2018\)Toward generating a new intrusion detection dataset and intrusion traffic characterization\.\.ICISSp1\(2018\),pp\. 108–116\.Cited by:[§3\.1](https://arxiv.org/html/2609.30605#S3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.30605#S4.SS1.p1.1)\.
- \[28\]A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu\(2017\)Towards deep learning models resistant to adversarial attacks\.arXiv preprint arXiv:1706\.06083\.Cited by:[§3\.2](https://arxiv.org/html/2609.30605#S3.SS2.p1.2)\.
- \[29\]F\. Doshi\-Velez and B\. Kim\(2017\)Towards a rigorous science of interpretable machine learning\.arXiv preprint arXiv:1702\.08608\.Cited by:[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p1.1)\.
- \[30\]R\. Guidotti, A\. Monreale, S\. Ruggieri, F\. Turini, F\. Giannotti, and D\. Pedreschi\(2018\)A survey of methods for explaining black box models\.ACM computing surveys \(CSUR\)51\(5\),pp\. 1–42\.Cited by:[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p1.1)\.
- \[31\]S\. Okada, H\. Jmila, K\. Akashi, T\. Mitsunaga, Y\. Sekiya, H\. Takase, G\. Blanc, and H\. Nakamura\(2025\)Xai\-driven black\-box adversarial attacks on network intrusion detectors\.International Journal of Information Security24\(3\),pp\. 1–15\.Cited by:[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p2.1)\.
- \[32\]M\. Sundararajan, A\. Taly, and Q\. Yan\(2017\)Axiomatic attribution for deep networks\.InInternational conference on machine learning,pp\. 3319–3328\.Cited by:[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p3.1)\.
- \[33\]M\. T\. Hosain, J\. R\. Jim, M\. Mridha, and M\. M\. Kabir\(2024\)Explainable ai approaches in deep learning: advancements, applications and challenges\.Computers and electrical engineering117,pp\. 109246\.Cited by:[§3\.4](https://arxiv.org/html/2609.30605#S3.SS4.p3.2)\.
- \[34\]K\. R\. Mopuri, U\. Garg, and R\. V\. Babu\(2017\)Fast feature fool: a data independent approach to universal adversarial perturbations\.arXiv preprint arXiv:1707\.05572\.Cited by:[Table 1](https://arxiv.org/html/2609.30605#S3.T1.p5.1.3.1.1.1.1)\.
- \[35\]K\. R\. Mopuri, A\. Ganeshan, and R\. V\. Babu\(2018\)Generalizable data\-free objective for crafting universal adversarial perturbations\.IEEE transactions on pattern analysis and machine intelligence41\(10\),pp\. 2452–2465\.Cited by:[Table 1](https://arxiv.org/html/2609.30605#S3.T1.p8.1.3.1.1.1.1)\.
- \[36\]Z\. Ye, X\. Cheng, and X\. Huang\(2023\)Fg\-uap: feature\-gathering universal adversarial perturbation\.In2023 International Joint Conference on Neural Networks \(IJCNN\),pp\. 1–8\.Cited by:[Table 1](https://arxiv.org/html/2609.30605#S3.T1.p11.1.3.1.1.1.1),[Table 1](https://arxiv.org/html/2609.30605#S3.T1.p14.1.3.1.1.1.1)\.
- \[37\]A\. Lesne\(2014\)Shannon entropy: a rigorous notion at the crossroads between probability, information theory, dynamical systems and statistical physics\.Mathematical Structures in Computer Science24\(3\),pp\. e240311\.Cited by:[§3\.5](https://arxiv.org/html/2609.30605#S3.SS5.p8.1)\.
- \[38\]V\. Mnih, K\. Kavukcuoglu, D\. Silver, A\. A\. Rusu, J\. Veness, M\. G\. Bellemare, A\. Graves, M\. Riedmiller, A\. K\. Fidjeland, G\. Ostrovski,et al\.\(2015\)Human\-level control through deep reinforcement learning\.nature518\(7540\),pp\. 529–533\.Cited by:[§4\.1](https://arxiv.org/html/2609.30605#S4.SS1.p1.1)\.
- \[39\]A\. Jentzen and P\. E\. Kloeden\(2011\)Taylor approximations for stochastic partial differential equations\.SIAM\.Cited by:[Appendix A](https://arxiv.org/html/2609.30605#A1.p2.2)\.
- \[40\]M\. A\. Merzouk, F\. Cuppens, N\. Boulahia\-Cuppens, and R\. Yaich\(2022\)Investigating the practicality of adversarial evasion attacks on network intrusion detection\.Annals of Telecommunications77\(11\),pp\. 763–775\.Cited by:[Appendix C](https://arxiv.org/html/2609.30605#A3.p2.1)\.
- \[41\]M\. Tavallaee, E\. Bagheri, W\. Lu, and A\. A\. Ghorbani\(2009\)A detailed analysis of the kdd cup 99 data set\.In2009 IEEE symposium on computational intelligence for security and defense applications,pp\. 1–6\.Cited by:[§F\.2](https://arxiv.org/html/2609.30605#A6.SS2.p1.1)\.
- \[42\]J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov\(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§F\.2](https://arxiv.org/html/2609.30605#A6.SS2.p4.1)\.
- \[43\]V\. Mnih, A\. P\. Badia, M\. Mirza, A\. Graves, T\. Lillicrap, T\. Harley, D\. Silver, and K\. Kavukcuoglu\(2016\)Asynchronous methods for deep reinforcement learning\.InInternational conference on machine learning,pp\. 1928–1937\.Cited by:[§F\.2](https://arxiv.org/html/2609.30605#A6.SS2.p4.1)\.

## Appendix ASupplementary Materials for Theoretical Analysis

We provide \(iv\) a first\-order improvement guarantee for the resulting weighted perturbations under standard local assumptions\. The previous theoretical analysis is placed in Sec\.[3\.5](https://arxiv.org/html/2609.30605#S3.SS5)\.

First\-order improvement guarantee under weighted perturbations\.We now connect the entropy\-allocation view to a*first\-order*targeted\-evasion surrogate\. Letv≜P​R​u​a​pv\\triangleq PRuapbe the feasibility\-enforced PR\-based UAP, and define the family of reweighted perturbations

u⁡\(w\)≜v⊙w,δ⁡\(w\)≜ϵ⋅u⁡\(w\)‖u⁡\(w\)‖2,w∈Δmask\.u\(w\)\\triangleq v\\odot w,\\qquad\\delta\(w\)\\triangleq\\epsilon\\cdot\\frac\{u\(w\)\}\{\\\|u\(w\)\\\|\_\{2\}\},\\qquad w\\in\\Delta\_\{\\mathrm\{mask\}\}\.\(10\)Consider the targeted logit functionF​\(x\)≜Clogit​\(x\)F\(x\)\\triangleq C\_\{\\mathrm\{logit\}\}\(x\)foryty\_\{t\}and its local first\-order Taylor approximation:F⁡\(x\+δ\)≈F⁡\(x\)\+∇xF​\(x\)⊤​δF\(x\+\\delta\)\\approx F\(x\)\+\\nabla\_\{x\}F\(x\)^\{\\top\}\\deltafor small‖δ‖2\\\|\\delta\\\|\_\{2\}\[[39](https://arxiv.org/html/2609.30605#bib.bib39)\]\. Averaging over the batchℬ\\mathcal\{B\}in Alg\.[3](https://arxiv.org/html/2609.30605#alg3), letg¯≜1\|ℬ\|​∑x∈ℬ∇xF​\(x\)\\bar\{g\}\\triangleq\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{x\\in\\mathcal\{B\}\}\\nabla\_\{x\}F\(x\)\. Ignoring the \(direction\-preserving\) normalization in Eq\.[10](https://arxiv.org/html/2609.30605#A1.E10)yields the standard linear surrogate gain:

𝒢⁡\(w\)≜⟨g¯,u⁡\(w\)⟩=∑i=1d\(g¯i​vi\)⏟si​wi=⟨s,w⟩,\\mathcal\{G\}\(w\)\\triangleq\\left\\langle\\bar\{g\},\\,u\(w\)\\right\\rangle=\\sum\_\{i=1\}^\{d\}\\underbrace\{\\big\(\\bar\{g\}\_\{i\}v\_\{i\}\\big\)\}\_\{s\_\{i\}\}\\,w\_\{i\}=\\langle s,w\\rangle,\(11\)wheres≜g¯⊙vs\\triangleq\\bar\{g\}\\odot vis a per\-feature score that combines \(i\) how much increasing featureiiraises the target logit viag¯i\\bar\{g\}\_\{i\}, and \(ii\) the PR\-based universal direction already discovered under PR optimization viaviv\_\{i\}\.

###### Corollary 1\(Connection to PX\-UAP with IG scores\)

Supposei​m​pimp\(Eq\.[5](https://arxiv.org/html/2609.30605#S3.E5)\) is used as a proxy score forssin[11](https://arxiv.org/html/2609.30605#A1.E11), which is justified by IG completeness \(Prop\.[1](https://arxiv.org/html/2609.30605#Thmproposition1)\) since IG decomposes the target\-logit difference across features along a path and thus preserves feature ranking for targeted logit increase\. Then the PX\-UAP weighting in Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7)implements the maximizer of an entropy\-regularized first\-order setting overΔmask\\Delta\_\{\\mathrm\{mask\}\}, and reallocating the fixedℓ2\\ell\_\{2\}budgetϵ\\epsilontoward the most influential feasible features while maintaining stable \(non\-degenerate\) allocation through the entropy term\.

Overall, this analysis explains whyP​X​u​a​pPXuapcan outperformP​R​u​a​pPRuapunder feasibility constraints:P​R​u​a​pPRuapidentifies a strong universal direction via PR/kkoptimization, and the IG→\\toSoftmax step then performs a principled budget reallocation toward decision\-critical feasible dimensions\. It can improve the expected targeted\-evasion gain under a fixedℓ2\\ell\_\{2\}budget\.

## Appendix BAddition Algorithims

In this section, we show the two addition algorithms that not shown in the content of main paper, inclduing the UAP generation framework in Alg\.[2](https://arxiv.org/html/2609.30605#alg2)and the PX\-UAP attack generation in Alg\.[3](https://arxiv.org/html/2609.30605#alg3)\.

Algorithm 2UAP Generation for ClassifierCCInput:TrainsetT​rTr, ClassifierCC, Modified feature group maskm​a​s​kmask, seedset sizes​i​z​esize, max iteration numberm​a​x​\_​i​t​e​rmax\\\_iter, perturbation budgetϵ\\epsilon

Output:Universal adversarial perturbationu​a​puap

1:

S​e​e​d​s​e​t←Sample⁡\(T​r,s​i​z​e⋅\|T​r\|\)Seedset\\leftarrow\\mathrm\{Sample\}\(Tr,\\;size\\cdot\|Tr\|\),

u​a​p←𝟎,i​t​e​r​\_​n​u​m←0uap\\leftarrow\\mathbf\{0\},iter\\\_num\\leftarrow 0
2:while

i​t​e​r​\_​n​u​m<m​a​x​\_​i​t​e​riter\\\_num<max\\\_iterdo

3:Randomly shuffle

S​e​e​d​s​e​tSeedset
4:foreach

x∈S​e​e​d​s​e​tx\\in Seedsetdo

5:

L1,L2←C⁡\(x\),C⁡\(x\+u​a​p\)L\_\{1\},L\_\{2\}\\leftarrow C\(x\),\\,C\(x\+uap\)
6:if

L1=L2L\_\{1\}=L\_\{2\}then

7:

g←gradient from Alg\.[1](https://arxiv.org/html/2609.30605#alg1)g\\leftarrow\\text\{gradient from Alg\.~\\ref\{alg:k\.gred\}\},

p​t​b←ϵ⋅m​a​s​k⊙g‖m​a​s​k⊙g‖2ptb\\leftarrow\\epsilon\\cdot\\dfrac\{mask\\odot g\}\{\\\|mask\\odot g\\\|\_\{2\}\},

u​a​p←Projectℓ2​\(u​a​p\+p​t​b\)uap\\leftarrow\\mathrm\{Project\}\_\{\\ell\_\{2\}\}\\\!\\left\(uap\+ptb\\right\)
8:endif

9:endfor

10:

i​t​e​r​\_​n​u​m←i​t​e​r​\_​n​u​m\+1iter\\\_num\\leftarrow iter\\\_num\+1
11:endwhile

12:return

u​a​puap

Algorithm 3PX\-UAP attackInput:TrainsetT​rTr, target modelCC, epsilonϵ\\epsilon, batch sizeBB, XAI feature selection maskmask\\mathrm\{mask\} Output:XAI enhanced UAPP​X​u​a​pPXuap

1:

u​a​p←PR−based​UAP⁡\(C,T​r,ϵ\)uap\\leftarrow\\operatorname\{PR\-based\\ UAP\}\(C,Tr,\\epsilon\),

P​R​u​a​p←recalculate⁡\(u​a​p,ϵ,ℓ2\)PRuap\\leftarrow\\operatorname\{recalculate\}\(uap,\\epsilon,\\ell\_\{2\}\)
2:

FNR←ComputeFNR⁡\(C⁡\(T​r\),C⁡\(T​r\+P​R​u​a​p\)\)\\mathrm\{FNR\}\\leftarrow\\operatorname\{ComputeFNR\}\(C\(Tr\),\\,C\(Tr\+PRuap\)\)
3:Sample a batch

ℬ⊂T​r\\mathcal\{B\}\\subset Trwith

\|ℬ\|=B\|\\mathcal\{B\}\|=B
4:

i​m​pbatch←XAImtd⁡\(ℬ,t​a​r​g​e​t​\_​c​l​a​s​s\)imp\_\{\\text\{batch\}\}\\leftarrow\\operatorname\{XAImtd\}\(\\mathcal\{B\},\\,target\\\_class\)⊳\\trianglerightB×dB\\times d

5:

i​m​p←1\|ℬ\|​∑x∈ℬi​m​pbatch​\(x\)imp\\leftarrow\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{x\\in\\mathcal\{B\}\}imp\_\{\\text\{batch\}\}\(x\)⊳\\trianglerightdd

6:Compute

P​X​u​a​pPXuapaccording to Eq\.[7](https://arxiv.org/html/2609.30605#S3.E7)

7:

P​X​u​a​p←recalculate⁡\(P​X​u​a​p,ϵ,ℓ2\)PXuap\\leftarrow\\operatorname\{recalculate\}\(PXuap,\\epsilon,\\ell\_\{2\}\),

FNR′←ComputeFNR⁡\(C⁡\(T​r\),C⁡\(T​r\+P​X​u​a​p\)\)\\mathrm\{FNR^\{\\prime\}\}\\leftarrow\\operatorname\{ComputeFNR\}\(C\(Tr\),C\(Tr\+PXuap\)\)
8:if

FNR′<FNR\\mathrm\{FNR^\{\\prime\}\}<\\mathrm\{FNR\}then

9:

P​X​u​a​p←P​R​u​a​pPXuap\\leftarrow PRuap
10:endif

11:return

P​X​u​a​pPXuap

## Appendix CNetwork Realistic Domain Constraints

In this section, we present the feature categorization for the CICIDS2018 dataset adopted in this study, together with brief descriptions and recalculation formulations, as shown in Tab\.[2](https://arxiv.org/html/2609.30605#A3.T2)\. We partition the input features into three groups:Modified Features \(MF\),Related Features \(RF\), andUnmodified Features \(UF\)\[[10](https://arxiv.org/html/2609.30605#bib.bib1)\]\.

Specifically, for any input flow feature vectorx∈𝒳x\\in\\mathcal\{X\}\(wherexix\_\{i\}denotes theii\-th feature ofxx\), the adversarial perturbationδ\\deltais applied only to theModified Features\(MF\) first, with the perturbed value\(x\+δ\)i\(x\+\\delta\)\_\{i\}must remain within its valid range \(e\.g\., protocol compliance and admissible feature bounds\[[40](https://arxiv.org/html/2609.30605#bib.bib26)\]\), as formalized in Eq\.[12](https://arxiv.org/html/2609.30605#A3.E12):

∀x∈𝒳,∀i∈M​F,\(x\+δ\)i∈ℛi​\(𝒳\)\\forall x\\in\\mathcal\{X\},\\quad\\forall i\\in MF,\\quad\(x\+\\delta\)\_\{i\}\\in\\mathcal\{R\}\_\{i\}\(\\mathcal\{X\}\)\(12\)
Regarding to RF, whenever the adversary modifies any feature inM​FMF\(denoted as theii\-th feature\), the corresponding features inR​FRF\(denoted as thejj\-th feature\) must be recalculated to preserve inter\-feature mathematical dependencies, shown in Eq\.[13](https://arxiv.org/html/2609.30605#A3.E13):

∀x∈𝒳,∀j∈R​F,\(x\+δ\)j=Recalculate​\(x\+δ\)i\\forall x\\in\\mathcal\{X\},\\quad\\forall j\\in RF,\\quad\(x\+\\delta\)\_\{j\}=\\text\{Recalculate\}\(x\+\\delta\)\_\{i\}\\\(13\)
The remaining features are assigned toUnmodified Features\(UF\) and remain fixed throughout the attack, since their values cannot be altered via feasible manipulation actions\. After these steps, the resulting perturbation is projected onto theL2L\_\{2\}ball of radiusϵ\\epsilonto enforce the prescribed attack budget\.

Table 2:Feature categorization, descriptions, and formulationsFeature NameDescriptionRecalculate FormulationGroupTot Fwd PktsTotal packets in the forward direction–Modified FeaturesTot Bwd PktsTotal packets in the backward direction–TotLen Fwd PktsTotal size of packet in forward direction–TotLen Bwd PktsTotal size of packet in backward direction–Flow DurationDuration of the flow in Microsecond–Fwd Pkts/sNumber of forward packets per secondTot Fwd Pkts×106/Flow Duration\{\\text\{Tot Fwd Pkts\}\\times 10^\{6\}\}\\;/\\;\{\\text\{Flow Duration\}\}Bwd Pkts/sNumber of backward packets per secondTot Bwd Pkts×106/Flow Duration\{\\text\{Tot Bwd Pkts\}\\times 10^\{6\}\}\\;/\\;\{\\text\{Flow Duration\}\}Flow Pkts/sNumber of flow packets per secondFwd Pkts/s\+Bwd Pkts/s\\displaystyle\\text\{Fwd Pkts/s\}\+\\text\{Bwd Pkts/s\}Flow Byts/sNumber of flow bytes per second\(TotLen Fwd Pkts\+TotLen Bwd Pkts\)×106Flow Duration\\displaystyle\\frac\{\(\\text\{TotLen Fwd Pkts\}\+\\text\{TotLen Bwd Pkts\}\)\\times 10^\{6\}\}\{\\text\{Flow Duration\}\}Related FeaturesPkt Size AvgAverage size of packetTotLen Fwd Pkts\+TotLen Bwd PktsTot Fwd Pkts\+Tot Bwd Pkts\\displaystyle\\frac\{\\text\{TotLen Fwd Pkts\}\+\\text\{TotLen Bwd Pkts\}\}\{\\text\{Tot Fwd Pkts\}\+\\text\{Tot Bwd Pkts\}\}Fwd Seg Size AvgAverage size in the forward directionTotLen Fwd Pkts/Tot Fwd Pkts\\text\{TotLen Fwd Pkts\}\\;/\\;\{\\text\{Tot Fwd Pkts\}\}Bwd Seg Size AvgAverage size in the backward directionTotLen Bwd Pkts/Tot Bwd Pkts\{\\text\{TotLen Bwd Pkts\}\}\\;/\\;\{\\text\{Tot Bwd Pkts\}\}Down/Up RatioDownload and upload ratioint​\(Tot Bwd Pkts/Tot Fwd Pkts\)\\text\{int\}\(\{\\text\{Tot Bwd Pkts\}\}\\;/\\;\{\\text\{Tot Fwd Pkts\}\}\)Other FeaturesRemaining features not in above groups–Unmodified Features
## Appendix DHyperparameter Setting for Algorithms

In this section, we report the hyperparameter settings for three algorithms in the main paper, shown in Tab\.[3](https://arxiv.org/html/2609.30605#A4.T3)\. Specifically, to balance computational efficiency with sufficient coverage of the training distribution, we generate the UAP using a small portion of the training set and set the seedset size to1×10−41\\times 10^\{\-4\}\(corresponding to 600 network data samples as the seedset for UAP generation\)\. For the same reason, we set the XAI batch sizeBB=10,000 to improve sample coverage for more representative XAI estimation, while keeping the computational overhead manageable\. All the attack methods and their variants shown in Sec\.[4](https://arxiv.org/html/2609.30605#S4)are evaluated under identical iteration settings \(e\.g\.m​a​x​\_​i​t​e​rmax\\\_iter,NN,nstepn\_\{\\mathrm\{step\}\},S​G​D​\_​n​u​mSGD\\\_num,α\\alpha\)\.

Table 3:Hyperparameter Setting for algorithmsFeature NameFeature DescriptionFeature Valueϵ\\epsilonAdversarial Budget\[0\.04,0\.15\]\[0\.04,\\,0\.15\]s​i​z​esizeSeedset size1×10−41\\times 10^\{\-4\}m​a​x​\_​i​t​e​rmax\\\_iterMax Iteration Number20NNPGD Attack Number10nstepn\_\{\\mathrm\{step\}\}PGD Iteration Step50α\\alphaPGD Step Sizeϵ/nstep\\epsilon/n\_\{\\mathrm\{step\}\}S​G​D​\_​n​u​mSGD\\\_numSGD Max Iteration Number50BBXAI Batch Size10000mmIG Iteration Step500
## Appendix EVariants detail for Ablation Study

To verify that the observed attack effectiveness stems from the combination of the components rather than any single factor, we list the ablation variants in which specific components are removed or modified\. Reviewing the three key components described in Alg\.[1](https://arxiv.org/html/2609.30605#alg1): \(i\) PGD\-based attack initialization, \(ii\)kk\-based selection of AE candidates, and \(iii\)k∗k^\{\*\}\-driven optimization\. We abbreviate these three components asAE Initia\.,K\-Sel\., andOpt\., respectively, in the following Table 4\.

Table 4:Deatils of PR\-based UAP attack and ablation variants\.Index NameAE Initia\.K\-Sel\.Optz\.\(1\): FUAPFGM\(1\)×\\times×\\times\(2\): KFUAPFGM\(N\)√\\surd×\\times\(3\): PUAPPGD\(1\)×\\times×\\times\(4\): KPUAPPGD\(N\)√\\surd×\\times\(5\): PR\-based UAPPGD\(N\)√\\surd√\\surdFor the attack\-initialization component, in addition to the PGD\-based initializer in Alg\.[1](https://arxiv.org/html/2609.30605#alg1), we replace it with a random\-start FGM initializer to form an alternative strategy\. The resulting variants are denoted as FUAP and KFUAP in Tab\.[4](https://arxiv.org/html/2609.30605#A5.T4)\. For variants that include AE candidate selection \(K\-Sel\.\), the initialization attack is repeated multiple times, denoted by ’\(N\)’ after the corresponding attack name\. In parallel, variants that select the AE candidate with the largestkkare prefixed with ’K’\. Otherwise, the sole available AE candidate is used asthe selected result\. Finally, to validate the effectiveness of the optimization, only the proposed PR\-based UAP updates the perturbation using∇xk∗\\nabla\_\{x\}k^\{\*\}, whereas the other variants directly adopt the AE candidateselected in the previous stepas the generated perturbation\.

## Appendix FAdditional Experimental Result

In this section, in addition to FNR and Accuracy results shown in main paper, we provide additional experimental results with a different evaluation metric for the two experiments and the ablation study mentioned in Sec\.[4](https://arxiv.org/html/2609.30605#S4)\. Besides, we also report attack performance across multiple intrusion detection datasets and different DRL architectures, transferability evaluation under the black\-box setting, and ablation study of epsilon step\.

### F\.1F1 score of the results in main paper

The classifier’s F1 score over the displayedϵ\\epsilonrange are provided for a more comprehensive assessment of overall performance, which is calculated according to Eq\.[14](https://arxiv.org/html/2609.30605#A6.E14):

F1\-score=2⋅True Positives2⋅True Positives\+False Positives\+False Negatives​\.\\text\{F1\-score\}=\\frac\{2\\cdot\\text\{True Positives\}\}\{2\\cdot\\text\{True Positives\}\+\\text\{False Positives\}\+\\text\{False Negatives\}\}\\text\{\.\}\(14\)The F1 score complements accuracy and FNR by offering a more balanced assessment by capturing both false alarms and missed detections in a single metric\. This offers an additional evaluation perspective and enables a more comprehensive view of overall detection performance\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure4.png)Figure 4:F1 score of the proposed PR\-based UAP and other UAP baselinesFig\.[4](https://arxiv.org/html/2609.30605#A6.F4)presents the F1 score of the proposed PR\-based UAP and other UAP baselines\. Consistent with the findings delivered in Sec\.[4](https://arxiv.org/html/2609.30605#S4), the proposed PR\-based UAP achieves the lowest F1 score among all the other baselines reported in Tab\.[1](https://arxiv.org/html/2609.30605#S3.T1)over the consideredϵ\\epsilonrange, indicating the strongest degradation of detection performance,i\.e\., strongest attack effectiveness\. Notably, even atϵ=0\.09\\epsilon=0\.09, where the FNR of the PR\-based UAP is the same as that of PCC\_UAP previously introduced in Fig\.[3](https://arxiv.org/html/2609.30605#S4.F3)\(b\), the PR\-based UAP still attains a lower F1 score, specifically it is the only candidate whose F1 score drops below 0\.3 forϵ\>0\.13\\epsilon\>0\.13, thereby achiving a stronger attack\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure5.png)Figure 5:F1 score of \(a\): PX\-UAP with differentTopn\\mathrm\{Topn\}values\. \(b\): The ablation studyWe next report the F1 score of PX\-UAP with differentTopn\\mathrm\{Topn\}values \(Fig\.[5](https://arxiv.org/html/2609.30605#A6.F5)\(a\)\) and the ablation study \(Fig\.[5](https://arxiv.org/html/2609.30605#A6.F5)\(b\)\)\. Overall, the trends in both subfigures are consistent with the results shown in main paper\. Specifically, PX\-UAP yields additional performance degradation in a wide range of perturbation, i\.e\.,ϵ<0\.07\\epsilon<0\.07andϵ≥0\.10\\epsilon\\geq 0\.10\. Among all settings,Topn=3\\mathrm\{Topn\}=3produces the largest reduction in the F1 score\. Turning to the ablation study, the PR\-based UAP achieves the lowest F1 score relative to all ablation variants, indicating the most pronounced degradation in detection performance and the most effective attack\.

In summary, the reported F1\-score results provide a complementary evaluation perspective and further substantiate the effectiveness of our proposed UAP methods\.

### F\.2Generalizability across Datasets and DRL Architectures

In the main paper, we report the results of the proposed PR\-based UAP on the CICIDS2018 dataset\. To further evaluate its generalizability across different scenarios, we additionally launch the proposed PR\-based UAP attacks against multiple classifiers with different DRL architectures and on two commonly used network intrusion detection datasets, namely NSL\-KDD\[[41](https://arxiv.org/html/2609.30605#bib.bib41)\]and UNSW\-NB15\[[19](https://arxiv.org/html/2609.30605#bib.bib19)\]\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure6.png)Figure 6:FNR of proposed PR\-based UAP on DRL\-based IDS \(DQN\-4\-64\) on \(a\): NSL\-KDD overϵ∈\[0,0\.45\]\\epsilon\\in\[0,0\.45\], \(b\) : UNSW\-NB15 overϵ∈\[0,0\.15\]\\epsilon\\in\[0,0\.15\]\.Specifically, since the two datasets differ in dimensionality and feature structure, we adopt dataset\-specificϵ\\epsilonranges to better demonstrate attack performance, withϵ∈\[0,0\.45\]\\epsilon\\in\[0,0\.45\]for NSL\-KDD andϵ∈\[0,0\.15\]\\epsilon\\in\[0,0\.15\]for UNSW\-NB15, resceptively\. The data preprocessing follows the same procedure described in Section[4\.1](https://arxiv.org/html/2609.30605#S4.SS1)\. It can be shown in Fig\.[6](https://arxiv.org/html/2609.30605#A6.F6)that the proposed PR\-based UAP attack and PX\-UAP attack consistently outperform the state\-of\-the\-art PCC\_UAP on the other two datasets with FNR≈\\approx0\.43 and 0\.78, respectively, demonstrating the generalizability of our proposed method in different scenarios in the intrusion detection area\.

Furthermore, in addition to the DRL\-based IDS architecture as the target classifier as used in the main paper, we further assess the attack performance on the same CICIDS2018 dataset but under different DRL agents and network architectures to assess whether the proposed attack can consistently compromise diverse DRL\-based IDS models\. It should be noted that the notation in parentheses indicates the DRL agent and network architecture used by the target IDS\. For example, DQN\-6\-128 denotes a DQN\-based target model with six hidden layers, each containing 128 units\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure7.png)Figure 7:FNR of proposed PR\-based UAP on different DRL\-based IDS on CICIDS2018 dateset with the taget model of \(a\): DQN\-6\-128, \(b\): PPO\-4\-64, \(c\): A2C\-4\-64\.As shown in Fig\.[7](https://arxiv.org/html/2609.30605#A6.F7), in addition to DQN, we further consider two widely used DRL architectures \(agents\): Proximal Policy Optimization \(PPO\)\[[42](https://arxiv.org/html/2609.30605#bib.bib42)\]and Advantage Actor\-Critic \(A2C\)\[[43](https://arxiv.org/html/2609.30605#bib.bib43)\]\. Specifically, the three target models used for evaluation are DQN\-6\-128, PPO\-4\-64, and A2C\-4\-64\. It can be observed that the proposed PR\-based UAP consistently outperforms PCC\_UAP across all three target models and PX\-UAP further improves the attack performance based on that, further demonstrating the generalizability of our method under different intrusion detection scenarios and the effectiveness of our algorithmic design\.

### F\.3Transferability Evaluation under the Black\-Box Setting

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure8.png)Figure 8:FNR for Cross\-model Transferability of AE \(PR\-based UAP\) among different DRL\-based IDS model on CICIDS2018 atϵ=0\.15\\epsilon=0\.15In some real\-world cybersecurity scenarios, attackers may not have access to the internal structure or parameters of the target model, corresponding to a black\-box attack setting\. Therefore, this setting should also be considered for a comprehensive evaluation of attack effectiveness\. We conduct additional transferability experiments, where UAPs are crafted on a surrogate model with white\-box access and then applied to other target models to evaluate their transferability\. The four models used as the surrogate and target models are those considered in the previous CICIDS2018 evaluation in the above section, together with the model used in the main paper, i\.e\., DQN\-4\-64\.

The results shown in Fig\.[8](https://arxiv.org/html/2609.30605#A6.F8)show the performance of the generated AEs under different surrogate\-target model pairs\. It can be observed that, across different transfer scenarios, i\.e\., the off\-diagonal entries, the proposed PR\-based UAP still achieves decent attack performance, which provides evidence of the proposed attack’s effectiveness even beyond the original white\-box setting, as well as the transferability of the generated AEs across different models and architectures\.

### F\.4Ablation study of epsilon step

In this section, we conduct an ablation study on the setting ofϵ\\epsilon\-steps to verify that the strong effectiveness of our proposed attack stems from the attack design itself, rather than being attributable to favorableϵ\\epsilon\-step selection\. The results demonstrate that our proposed PR\-based UAP consistently outperforms the two baselines, both under the fixed unifiedϵ\\epsilon\-step setting and whenϵ\\epsilon\-steps are tuned for each method\.

Specifically, in the experimental setting of Alg\.[2](https://arxiv.org/html/2609.30605#alg2)in the main paper, we set the number ofϵ\\epsilon\-steps to 1, which means each update performs a full\-ϵ\\epsilonstep at line 7\. This design choice is intended to ensure a unified fairness criterion, as different UAP methods may benefit from different small or adaptive step\-size schedules\. If each candidate were allowed to use its own step\-size schedule, the performance comparison would be affected by step\-size tuning as well, rather than reflecting only the difference in UAP design\. Therefore, we adopt the same full\-ϵ\\epsilonupdate rule across all methods in the experiments shown in Chapter[4](https://arxiv.org/html/2609.30605#S4)to ensure a fair comparison\.

Despite this setting, in order to provide a more comprehensive evaluation of our proposed attack method and rule out the influence ofϵ\\epsilon\-step selection, we further conduct an additional step\-size sensitivity analysis for all attack candidates\. Namely, we compare the attack performance of the proposed PR\-based UAP, PCC\_UAP, and UAP under different step sizes, identify the best\-performing step size for each candidate method, and finally report a best\-step\-size comparison\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure9.png)Figure 9:FNR of \(a\): PR\-based UAP \(b\): PCC\_UAP \(c\): UAP with different stepsize on CICIDS2018 dateset\. The applied stepsize are show in parentheses\.Fig\.[9](https://arxiv.org/html/2609.30605#A6.F9)show attack performance of PR\-based UAP, PCC\_UAP, UAP under differentϵ\\epsilonstepsize\. We set the step size toϵ\\epsilon\(the original setting\),ϵ\\epsilon/5,ϵ\\epsilon/10,ϵ\\epsilon/20, and an adaptive step\-size scheme\. In the adaptive setting, the step size is gradually reduced over the 30 iterations \(Alg\.[2](https://arxiv.org/html/2609.30605#alg2)line 2\):ϵ\\epsilon/5 for iterations 1–10,ϵ\\epsilon/10 for iterations 11–20, andϵ\\epsilon/20 for iterations 21–30\. As discussed above, different attack methods may benefit from different step\-size schedules\. Therefore, according to the results in Fig\.[9](https://arxiv.org/html/2609.30605#A6.F9)\(a\)–\(c\), we select the best\-performing step\-size setting for each attack method, namely UAP \(adaptive\) and PCC\_UAP \(ϵ/10\\epsilon/10\) and then compare them with two comparably good settings of the proposed PR\-based UAP \(Fig\.[9](https://arxiv.org/html/2609.30605#A6.F9)\(a\)\), includingϵ/5\\epsilon/5and the original full\-ϵ\\epsilonstep setting\.

![Refer to caption](https://arxiv.org/html/2609.30605v1/Figures/New_Figure10.png)Figure 10:FNR of best\-step\-size comparison for PR\-based UAP, PCC\_UAP and UAP attack\.As shown in Fig\.[10](https://arxiv.org/html/2609.30605#A6.F10), the two PR\-based UAP variants \(black solid and dashed lines\) still outperform the other two baselines under their respective best\-performing settings\. Together with the original comparison under a unified full\-ϵ\\epsilonstep size, this supports a more comprehensive conclusion that the proposed PR\-based UAP significantly outperforms the state\-of\-the\-art PCC\_UAP under both a fairness setting with the same step size and a fairness setting with the best step size, which further demonstrates the effectiveness of our proposed attack method\.

## Appendix GLimitation and future work

In our opinion, the limitation of this work and the corresponding directions for the future work are as follows: Although this work evaluates the transferability of the proposed attacks to provide evidence of their effectiveness under black\-box settings, our main threat model remains white\-box\. Future work can further extend this study to more restrictive black\-box scenarios\. This would enable a closer approximation of practical deployment environments\.

## Appendix HExperiment Device

All experiments in this study were conducted on a laptop workstation equipped with an AMD Ryzen 9 5900HX processor, an NVIDIA GeForce RTX 3080 Laptop GPU with 16 GB VRAM, 32 GB DDR4 RAM, and 2 TB PCIe Gen3 SSD storage\. The operating system used throughout the experiments was Windows 10 64\-bit\.

## Appendix IBroader Impact

This work may have both positive and negative societal impacts\. Positively, by exposing the vulnerability of DRL\-based intrusion detection systems to universal adversarial perturbations, this study can help the research community and security practitioners better understand potential weaknesses in current IDS models and motivate the development of more robust and reliable cybersecurity defenses\. This is particularly important for improving the trustworthiness of intelligent security systems in practical deployment\. Negatively, since the work studies adversarial attack methods, the proposed techniques could potentially be misused to evade real\-world intrusion detection systems\. Therefore, the study is conducted in controlled offline experimental settings and is intended to support academic research and defensive security development\.

相似文章

测试对未知对手的鲁棒性

OpenAI Blog

# 测试对未知对手的鲁棒性 来源:[https://openai.com/index/testing-robustness/](https://openai.com/index/testing-robustness/) OpenAI 我们开发了一种方法来评估神经网络分类器是否能可靠地抵御训练期间未见过的对抗性攻击。我们的方法产生了一个新的指标 UAR(未知攻击鲁棒性),它评估单个模型对意外攻击的鲁棒性,并强调了需要在更多样化的未知攻击范围内测量性能

不同扰动类型之间对抗鲁棒性的迁移

OpenAI Blog

# 不同扰动类型之间对抗鲁棒性的迁移 来源: [https://openai.com/index/transfer-of-adversarial-robustness-between-perturbation-types/](https://openai.com/index/transfer-of-adversarial-robustness-between-perturbation-types/) OpenAI## 摘要 我们研究深度神经网络在不同扰动类型之间的对抗鲁棒性迁移。虽然大多数关于对抗样本的工作专注于L∞L\_∞和L2L\_2有界扰动,但这些并不能捕捉所有t