Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Summary
This paper proposes computation-efficient strategies for latent adversarial training (LAT) of LLMs, using low-rank representation fine-tuning and circuit-guided surrogate models to reduce per-step FLOPs by 48.1% while requiring only 0.0118% trainable parameters.
View Cached Full Text
Cached at: 08/03/26, 07:34 AM
# Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Source: [https://arxiv.org/html/2607.28959](https://arxiv.org/html/2607.28959)
Weiyi He Yuping Lin Jiliang Tang Yue Xing
Michigan State University
\{heweiyi, linyupin, tangjili, xingyue1\}@msu\.edu
###### Abstract
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models \(LLMs\)\. While existing mitigation strategies, e\.g\., latent adversarial training \(LAT\), have been developed, they still incur a high computational cost\. In this work, we comprehensively investigate computation\-efficient strategies to speed up LAT from two complementary perspectives: \(1\) Defense\-side optimization: We explore the representation fine\-tuning \(ReFT\) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack\. \(2\) Attack\-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward\-backward passes through the full model during the attack generation\. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies\. Ultimately, compared to standard LAT with full fine\-tuning, our method on average reduces per\-step adversarial\-training FLOPs by 48\.1% while requiring only 0\.0118% trainable parameters\.
## 1Introduction
Large language models \(LLMs\) have achieved remarkable performance across diverse tasks, yet remain critically vulnerable to adversarial attacks that manipulate model outputs via crafted prompts\(Zouet al\.,[2023b](https://arxiv.org/html/2607.28959#bib.bib1); Yuet al\.,[2023](https://arxiv.org/html/2607.28959#bib.bib2); Andriushchenkoet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib3)\)\. These vulnerabilities pose serious risks for real\-world deployments, making robust defense mechanisms essential\. For example, recent reports have shown that indirect prompt injection can hijack tool\-using LLM systems via poisoned external content, such as malicious calendar invites targeting Gemini, and similar flaws have also been exploited to induce unsafe downstream actions in AI coding agents\(Burgess,[2025](https://arxiv.org/html/2607.28959#bib.bib59); Hart,[2026](https://arxiv.org/html/2607.28959#bib.bib63)\)\.
To ensure the robustness of LLMs, one of the most effective defenses is adversarial training\(Xinget al\.,[2021](https://arxiv.org/html/2607.28959#bib.bib12); Zhaoet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib11)\), which trains LLMs on adversarial examples to help them better defend against adversarial inputs\. However, scaling adversarial training to modern LLMs is computationally infeasible\. The cost arises from two sources: \(i\) updating a large number of model parameters during training, and \(ii\) repeatedly running the full model to generate adversarial examples via iterative optimization in the token space\(Zouet al\.,[2023b](https://arxiv.org/html/2607.28959#bib.bib1); Howeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib7)\)\. Recent works have proposed latent adversarial training \(LAT\)\(Xhonneuxet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib6); Sheshadriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib4); Casperet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib5)\), which perturbs continuous hidden representations rather than optimizing directly in the discrete token space\. This makes the optimization substantially easier than token\-level adversarial example generation, but the computational bottleneck remains significant: running the inner attack with Projected Gradient Descent \(PGD\)\(Madryet al\.,[2017](https://arxiv.org/html/2607.28959#bib.bib13)\)for 8 steps alone accounts for approximately 80% of the computation in LAT under the Kaplan FLOPs estimate\(Kaplanet al\.,[2020](https://arxiv.org/html/2607.28959#bib.bib56)\)\.
Figure 1:A naive latent defense that protects only the last token at the last layer is ineffective\. Robustness improves substantially when the defense covers a suffix window and is moved away from the end of the model\.While parameter\-efficient training strategies such as LoRA and representation fine\-tuning \(ReFT\)\(Wuet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib8); Renet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib9); Zenget al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib10)\)have been widely adopted, significant gaps persist when applying them in the context of LAT\.\(G1\)On the defense side \(i\.e\., updating model parameters to enhance robustness\), ReFT requires substantially fewer trainable parameters than LoRA, making it an attractive candidate for efficient adversarial training\. However, standard ReFT often applies interventions on task\-relevant positions, sometimes only the last token at a specific hidden layer \(usually not the final layer\)\. Since adversarial attacks can perturb multiple tokens, it remains unclear whether such position\-limited interventions can effectively defend multi\-token attacks\.\(G2\)On the attack side, while ReFT and other parameter\-efficient methods reduce the cost of model parameter updates, they do not address the computational burden of attack generation itself: computing adversarial perturbations still requires iterative forward\-backward passes through the full model at each training step\.
To address these gaps, we study adversarial robustness in classification settings with a focus on token\-level suffix attacks, followingHoweet al\.\([2024](https://arxiv.org/html/2607.28959#bib.bib7)\)\. We systematically explore strategies for reducing LAT’s computation cost from both defense and attack perspectives, both theoretically and empirically\. Our contributions are as follows:
- •Defense side:We introduce LAT\-ReFT, a parameter\-efficient latent defense that integrates ReFT into LAT\. To analyze\(G1\), as in Figure[1](https://arxiv.org/html/2607.28959#S1.F1), when the interventions are added on a suffix window of the attacked tokens, the attack success rate is much lower\. We further theoretically prove that single\-token defense is insufficient: attention mechanisms allow attacks from earlier suffix positions to leak into the defended output \(Theorem[1](https://arxiv.org/html/2607.28959#Thmtheorem1)\)\. We then derive conditions under which multi\-token suffix window ReFT can provably suppress this leakage \(Theorem[2](https://arxiv.org/html/2607.28959#Thmtheorem2)\)\. Empirical results demonstrate the effectiveness of the proposed algorithm in reducing computation costs\.
- •Attack side:Inspired by mechanistic interpretability research showing that model behaviors are determined by sparse circuits\(Wanget al\.,[2022](https://arxiv.org/html/2607.28959#bib.bib18); Elhageet al\.,[2022](https://arxiv.org/html/2607.28959#bib.bib17); Hannaet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib15); Chenet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib16)\), we develop a circuit\-guided surrogate attack strategy to overcome\(G2\): generating attacks on a pruned surrogate model and transferring them to the full model\. The key challenge is balancing efficiency and transferability: naive pruning accelerates attack generation but severely degrades attack effectiveness\. Inspired by previous idea\(Molchanovet al\.,[2016](https://arxiv.org/html/2607.28959#bib.bib60),[2019](https://arxiv.org/html/2607.28959#bib.bib61); Syedet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib21)\), we develop an activation\-gradient importance scoring, ActGrad, to identify attack\-critical neurons, and validate it theoretically in Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3)\. Empirical results show that pruning 25% of MLP neurons preserves attack\-relevant geometry while reducing inner\-loop cost by about 20%\. Combined with our defense\-side strategy, we achieve an overall computational reduction of 48\.1% on average compared to standard LAT with full fine\-tuning\.
- •Unified insight:Despite operating through different mechanisms, the defense and attack components reveal a common layerwise pattern: effective ReFT interventions should be placed in early or middle layers rather than very late layers, and the attack surrogate likewise retains most of its important neurons from early and middle layers\. This observation is broadly consistent with recent interpretability evidence suggesting that late layers in Transformers often have limited functional complexity\(Skeanet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib20); Queipo\-de\-Llanoet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib55)\), and it provides practical guidance for efficient adversarial training design\.
## 2Related Work
#### Adversarial Training for LLMs\.
Adversarial training \(AT\) is a widely used approach for improving model robustness by optimizing performance under worst\-case perturbations\(Madryet al\.,[2017](https://arxiv.org/html/2607.28959#bib.bib13); Shafahiet al\.,[2019](https://arxiv.org/html/2607.28959#bib.bib26); Wonget al\.,[2020](https://arxiv.org/html/2607.28959#bib.bib27); Andriushchenko and Flammarion,[2020](https://arxiv.org/html/2607.28959#bib.bib28)\)\. In the context of LLMs, AT typically involves generating adversarial prompts in the discrete token space, which incurs high computational cost\(Ebrahimiet al\.,[2018](https://arxiv.org/html/2607.28959#bib.bib29); Zouet al\.,[2023b](https://arxiv.org/html/2607.28959#bib.bib1); Howeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib7)\)\. To address this challenge, recent work has proposed Continuous Adversarial Training \(CAT\), where adversarial perturbations are applied in the embedding space\(Xhonneuxet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib6); Dékányet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib23)\)\. This allows efficient gradient\-based optimization, reducing the cost of adversarial example generation\. Subsequent work further extends this framework to attacking the latent representations in the intermediate LLM layers\(Sheshadriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib4); Casperet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib5)\)\. Recent theoretical analyses also provide justification for the effectiveness of such continuous perturbations\(Fu and Wang,[2026](https://arxiv.org/html/2607.28959#bib.bib24)\)\.
#### Representation Engineering\.
Representation engineering studies how model behavior can be controlled by operating directly on internal activations, rather than by updating all model parameters\(Zouet al\.,[2023a](https://arxiv.org/html/2607.28959#bib.bib30)\)\. Prior work has shown that hidden representations in LLMs encode rich semantic and behavioral information, and modifying these representations can effectively steer model outputs\(Turneret al\.,[2023](https://arxiv.org/html/2607.28959#bib.bib31); Zhenget al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib32); Linet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib33); Arditiet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib34)\)\. Within this line of work, Representation finetuning \(ReFT\) provides a parameter\-efficient adaptation mechanism by learning localized low\-rank interventions in activation space\(Wuet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib8)\)\. Subsequent studies further show that such representation\-level updates can serve as an effective alternative to full\-parameter finetuning\(Renet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib9); Zenget al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib10)\)\. This is closely related to our setting, since latent adversarial defense also operates directly on internal representations and therefore benefits naturally from parameter\-efficient activation\-level interventions\.
#### Mechanistic Interpretability and Circuits\.
Mechanistic interpretability aims to identify the internal components and computational pathways responsible for specific model behaviors\(Chenet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib16); Leeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib36); Sharkeyet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib35)\)\. A recurring finding in this literature is that many behaviors in LLMs are mediated by relatively sparse circuits or localized subnetworks, rather than being uniformly distributed across all parameters\(Wanget al\.,[2022](https://arxiv.org/html/2607.28959#bib.bib18); Hannaet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib15); Patelet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib37)\)\. To study such locality, prior work has developed methods for localizing and validating important components\(Viget al\.,[2020](https://arxiv.org/html/2607.28959#bib.bib38); Frantar and Alistarh,[2023](https://arxiv.org/html/2607.28959#bib.bib39)\)\. In our work, these findings motivate reducing the computation used in adversarial optimization by focusing on the model components that are most relevant to the expression of adversarial behaviors\.
## 3Preliminaries
#### Transformers\.
A decoder\-only Transformer\(Vaswaniet al\.,[2017](https://arxiv.org/html/2607.28959#bib.bib22)\)maps an input token sequencex=\[x1,…,xT\]x=\[x\_\{1\},\\dots,x\_\{T\}\]to output logits\. Let𝐡tl∈ℝd\\mathbf\{h\}\_\{t\}^\{l\}\\in\\mathbb\{R\}^\{d\}denote the hidden representation of tokenttat layerl∈\[Lall\]l\\in\[L\_\{\\mathrm\{all\}\}\]\. Each layer contains a self\-attention block and an MLP block with residual connections\. For notational simplicity, we write the layer update as𝐡tl=𝐡tl−1\+𝐚tl\+𝐦tl,\\mathbf\{h\}\_\{t\}^\{l\}=\\mathbf\{h\}\_\{t\}^\{l\-1\}\+\\mathbf\{a\}\_\{t\}^\{l\}\+\\mathbf\{m\}\_\{t\}^\{l\},where𝐚tl\\mathbf\{a\}\_\{t\}^\{l\}and𝐦tl\\mathbf\{m\}\_\{t\}^\{l\}denote the attention and MLP updates, respectively\.
#### Latent Adversarial Training \(LAT\)\.
Letfθ\(x\)f\_\{\\theta\}\(x\)denote the model logits for inputxx, parameterized byθ\\theta, and letℓ\(fθ\(x\),y\)\\ell\(f\_\{\\theta\}\(x\),y\)be the classification loss for labelyy\. Standard adversarial training seeks parameters that are robust to worst\-case perturbations:minθ𝔼\(x,y\)\[maxx′∈ℬ\(x,ε\)ℓ\(fθ\(x′\),y\)\],\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{\(x,y\)\}\\left\[\\max\_\{x^\{\\prime\}\\in\\mathcal\{B\}\(x,\\varepsilon\)\}\\ell\(f\_\{\\theta\}\(x^\{\\prime\}\),y\)\\right\],whereℬ\(x,ε\)\\mathcal\{B\}\(x,\\varepsilon\)denotes the set of perturbations aroundxxunder a perturbation budgetε\\varepsilon\.
To save the cost of calculating the discrete attacked tokens, LAT\(Sheshadriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib4); Casperet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib5)\)instead performs the inner maximization in a continuous hidden space\. Let𝐇l\(x;θ\)∈ℝT×d\\mathbf\{H\}\_\{l\}\(x;\\theta\)\\in\\mathbb\{R\}^\{T\\times d\}denote the hidden\-state tensor at layerllfor the inputxx\. Rather than perturbing the discrete input directly, LAT optimizes over hidden representations in a neighborhood of𝐇l\(x;θ\)\\mathbf\{H\}\_\{l\}\(x;\\theta\):
minθ𝔼\(x,y\)\[ℓ\(fθ\(x\),y\)\+λadvmax𝐇~∈ℬ\(𝐇l\(x;θ\),ε\)ℓ\(fθ\(x;𝐇~\),y\)\]\.\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{\(x,y\)\}\\left\[\\ell\(f\_\{\\theta\}\(x\),y\)\+\\lambda\_\{\\mathrm\{adv\}\}\\max\_\{\\widetilde\{\\mathbf\{H\}\}\\in\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\}\(x;\\theta\),\\varepsilon\)\}\\ell\\bigl\(f\_\{\\theta\}\(x;\\widetilde\{\\mathbf\{H\}\}\),y\\bigr\)\\right\]\.\(1\)Here,ℬ\(𝐇l\(x;θ\),ε\)\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\}\(x;\\theta\),\\varepsilon\)denotes the neighborhood in hidden space, andfθ\(x;𝐇~\)f\_\{\\theta\}\(x;\\widetilde\{\\mathbf\{H\}\}\)denotes the model output when the hidden representation at layerllis replaced by𝐇~\\widetilde\{\\mathbf\{H\}\}\. Since the inner optimization is carried out in continuous representation space, it can be efficiently approximated using PGD\.
#### Representation Finetuning \(ReFT\)\.
ReFT\(Wuet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib8)\)restricts model updates to a low\-rank subspace of the hidden representation\. Let𝐑∈ℝr×d\\mathbf\{R\}\\in\\mathbb\{R\}^\{r\\times d\}be a matrix with orthonormal rows, wherer≤dr\\leq dis the subspace rank\. Given trainable parameters𝐖∈ℝr×d\\mathbf\{W\}\\in\\mathbb\{R\}^\{r\\times d\}and𝐛∈ℝr\\mathbf\{b\}\\in\\mathbb\{R\}^\{r\}, the ReFT operator applied to a hidden state𝐡∈ℝd\\mathbf\{h\}\\in\\mathbb\{R\}^\{d\}isΦ𝐑\(𝐡\)=𝐡\+𝐑⊤\(𝐖𝐡\+𝐛−𝐑𝐡\)\.\\Phi\_\{\\mathbf\{R\}\}\(\\mathbf\{h\}\)=\\mathbf\{h\}\+\\mathbf\{R\}^\{\\top\}\(\\mathbf\{W\}\\mathbf\{h\}\+\\mathbf\{b\}\-\\mathbf\{R\}\\mathbf\{h\}\)\.
## 4Method
Our goal is to make adversarial training for LLMs both effective and efficient\. On the defense side, we introduce LAT\-ReFT, a parameter\-efficient defense that restricts adaptation to a low\-rank subspace of the hidden representation\. On the attack side, we introduce a circuit\-guided surrogate that reduces the cost of inner\-loop adversarial optimization while being transferable to the full defended model\.
### 4\.1Defense: LAT\-ReFT
We integrate ReFT into LAT to obtain a parameter\-efficient latent defense by freezing the backbone and training only a low\-rank intervention in representation space\. Letlal\_\{\\mathrm\{a\}\}denote the attack layer andlrl\_\{\\mathrm\{r\}\}the defense layer, withlr≥lal\_\{\\mathrm\{r\}\}\\geq l\_\{\\mathrm\{a\}\}\. Let𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}denote the token positions where latent perturbations are allowed\. Consider perturbationsΔ𝒯∈ℝT×d\\Delta\_\{\\mathcal\{T\}\}\\in\\mathbb\{R\}^\{T\\times d\}, supported only on𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}, i\.e\.,\(Δ𝒯\)t,:=0\(\\Delta\_\{\\mathcal\{T\}\}\)\_\{t,:\}=0for allt∉𝒯LATt\\notin\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\. Thus, the attack is restricted to the selected token positions at layerlal\_\{\\mathrm\{a\}\}, while ReFT is applied at layerlrl\_\{\\mathrm\{r\}\}to correct the resulting adversarial features\. Letf𝐑,𝐖,𝐛\(x;𝐇~\)f\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\(x;\\widetilde\{\\mathbf\{H\}\}\)denote the model output when the hidden representation at layerlal\_\{\\mathrm\{a\}\}is replaced by𝐇~\\widetilde\{\\mathbf\{H\}\}and the ReFT module parameterized by\(𝐑,𝐖,𝐛\)\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\)is applied at layerlrl\_\{\\mathrm\{r\}\}\. The defense objective is
min𝐑,𝐖,𝐛𝔼\(x,y\)\[ℓ\(f𝐑,𝐖,𝐛\(x\),y\)\+λadvmax𝐇~∈ℬ\(𝐇la\(x;θ\),ε;𝒯LAT\)ℓ\(f𝐑,𝐖,𝐛\(x;𝐇~\),y\)\]\.\\min\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\\;\\mathbb\{E\}\_\{\(x,y\)\}\\left\[\\ell\\\!\\left\(f\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\(x\),y\\right\)\+\\lambda\_\{\\mathrm\{adv\}\}\\max\_\{\\widetilde\{\\mathbf\{H\}\}\\in\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\),\\varepsilon;\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\)\}\\ell\\\!\\left\(f\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\(x;\\widetilde\{\\mathbf\{H\}\}\),y\\right\)\\right\]\.\(2\)Here,ℬ\(𝐇la\(x;θ\),ε;𝒯LAT\)\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\),\\varepsilon;\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\)denotes the hidden\-space neighborhood obtained by restricting perturbations to the token positions in𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\.
A central design issue for ReFT\-based latent defense is*where*to place the intervention\. As illustrated in Figure[1](https://arxiv.org/html/2607.28959#S1.F1), this choice has two parts: \(1\) selecting the defended suffix window and \(2\) selecting the intervention layer\. We discuss these two design choices in turn below\.
#### Suffix\-window defense\.
A ReFT intervention applied only at the last token is not sufficient in our setting\. Empirically, restricting the defense to a single position leaves the model vulnerable to suffix\-style attacks, indicating that the intervention is too local\. We therefore defend a contiguous suffix window rather than a single token:𝒯LAT=\{T−L\+1,…,T\}\.\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}=\\\{T\-L\+1,\\dots,T\\\}\.We provide theoretical support for this design in Section[5](https://arxiv.org/html/2607.28959#S5)\. To supplement, we also provide results for prefix attacks and present that a suffix window can still help defend against prefix attacks; see Appendix[B\.3](https://arxiv.org/html/2607.28959#A2.SS3)\.
#### Defense layer\.
Prior latent\-defense work typically treats layer choice empirically\(Sheshadriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib4); Casperet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib5)\), while existing analyses suggest that Transformer layers can be roughly grouped into early, middle, and late stages with different roles\(Skeanet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib20)\)\. Guided by both perspectives, we compare intervention layers from these regions in Section[6\.3](https://arxiv.org/html/2607.28959#S6.SS3)\. Our results show that late\-layer intervention is generally ineffective, and ReFT should be away from the end of the model\.
### 4\.2Circuit\-Guided Surrogate for Attack Generation
Another problem in LAT is the cost of repeatedly solving the inner maximization in Eq\. \([2](https://arxiv.org/html/2607.28959#S4.E2)\)\. Even though latent optimization is cheaper than discrete token\-space search, generating adversarial perturbations still requires many PGD steps with repeated forward–backward passes through a large model\. To reduce this cost, we introduce a circuit\-guided surrogate\.
Specifically, we use a surrogate model for inner\-loop attack generation: the latent perturbation is optimized on the surrogate and then applied to the full defended model for training\. We construct the surrogate by compressing the MLP blocks of the full model\. MLP blocks are a natural target because they contain about two\-thirds of a Transformer’s parameters and account for much computation\. They are also straightforward to compress in a structured way, as their expanded hidden layer is organized neuron\-wise, so we can retain only a subset of neurons and prune the rest\(Gevaet al\.,[2021](https://arxiv.org/html/2607.28959#bib.bib58)\)\. To select neurons to retain, we assign each MLP neuron an importance score inspired by attribution analyses of internal model components, which use activation\- and gradient\-based signals to estimate component importance\(Syedet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib21)\)\. Concretely, for neuronjjin layerll, we define
sl,j=∑\(x,y\)∈𝒟∑t\|at,j\(l\)\(x\)∂ℓ\(fθ\(x\),y\)∂at,j\(l\)\|,s\_\{l,j\}=\\sum\_\{\(x,y\)\\in\\mathcal\{D\}\}\\sum\_\{t\}\\left\|a\_\{t,j\}^\{\(l\)\}\(x\)\\,\\frac\{\\partial\\ell\(f\_\{\\theta\}\(x\),y\)\}\{\\partial a\_\{t,j\}^\{\(l\)\}\}\\right\|,\(3\)whereat,j\(l\)\(x\)a\_\{t,j\}^\{\(l\)\}\(x\)denotes the activation of neuronjjat tokenttin the expanded hidden layer of the MLP block at layerll, and𝒟\\mathcal\{D\}is a small calibration set drawn from the clean training data\. As will be discussed in Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3), this score favors neurons that are not only active, but also influential for changing the adversarial loss\. We rank neurons globally by this score, retain only the top\-scoring fraction, and prune the rest when running the inner\-loop attack\.
Letfsurrf\_\{\\mathrm\{surr\}\}denote the surrogate; we run PGD on it to obtain an adversarial hidden\-state perturbation
𝐇~surr⋆=argmax𝐇~∈ℬ\(𝐇la\(x;θ\),ε;𝒯LAT\)ℓ\(fsurr\(x;𝐇~\),y\),\\widetilde\{\\mathbf\{H\}\}\_\{\\mathrm\{surr\}\}^\{\\star\}=\\operatorname\*\{arg\\,max\}\_\{\\widetilde\{\\mathbf\{H\}\}\\in\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\),\\varepsilon;\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\)\}\\ell\\bigl\(f\_\{\\mathrm\{surr\}\}\(x;\\widetilde\{\\mathbf\{H\}\}\),y\\bigr\),\(4\)and transfer the resulting perturbation back to the full defended model\. In Section[6\.4](https://arxiv.org/html/2607.28959#S6.SS4), we show that this strategy reduces attack\-time computation while transferred well to the full defended model\.
## 5Theoretical Analysis
In this section, we provide theoretical analysis for the suffix\-window defense and the surrogate model\. We write𝐡tlr\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}for the hidden vector at tokentt, i\.e\., thett\-th row of𝐇lr\(x;θ\)\\mathbf\{H\}\_\{l\_\{\\mathrm\{r\}\}\}\(x;\\theta\)\. LetαT,tl\\alpha\_\{T,t\}^\{l\}andα^T,tl\\hat\{\\alpha\}\_\{T,t\}^\{l\}denote the attention weights from query positionTTto tokenttat layerllin the clean and defended forward passes, respectively, where the clean pass is unperturbed, and the defended pass includes latent perturbation together with ReFT intervention\. Here,𝐡~\\tilde\{\\mathbf\{h\}\}denotes the perturbed hidden state before ReFT, and𝐡^\\hat\{\\mathbf\{h\}\}denotes the defended hidden state after ReFT\. We also define the clean and defended value vectors by𝐯tl=𝐖Vl𝐡tl−1\\mathbf\{v\}\_\{t\}^\{l\}=\\mathbf\{W\}\_\{V\}^\{l\}\\mathbf\{h\}\_\{t\}^\{\\,l\-1\}and𝐯^tl=𝐖Vl𝐡^tl−1\.\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}=\\mathbf\{W\}\_\{V\}^\{l\}\\hat\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\-1\}\.Here𝐖Ql,𝐖Kl,𝐖Vl\\mathbf\{W\}\_\{Q\}^\{l\},\\mathbf\{W\}\_\{K\}^\{l\},\\mathbf\{W\}\_\{V\}^\{l\}are the query, key, and value matrices of the attention head at layerll, with‖𝐖Ql‖≤BQ,‖𝐖Kl‖≤BK,‖𝐖Vl‖≤BV\\\|\\mathbf\{W\}\_\{Q\}^\{l\}\\\|\\leq B\_\{Q\},\\;\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\leq B\_\{K\},\\;\\\|\\mathbf\{W\}\_\{V\}^\{l\}\\\|\\leq B\_\{V\}for some finite constantsBQ,BK,BVB\_\{Q\},B\_\{K\},B\_\{V\}\(Edelmanet al\.,[2022](https://arxiv.org/html/2607.28959#bib.bib41); He and Xing,[2025](https://arxiv.org/html/2607.28959#bib.bib40)\)\. Unless otherwise specified,∥⋅∥\\\|\\cdot\\\|denotes theℓ2\\ell\_\{2\}norm for vectors and the spectral norm for matrices\. Additional technical details are deferred to Appendix[D\.1](https://arxiv.org/html/2607.28959#A4.SS1)\. Besides, we also impose the following assumption:
###### Assumption 1\(Local Lipschitz continuity\)\.
Let𝐨Tl\\mathbf\{o\}\_\{T\}^\{l\}denote the attention output at token T in layer l\. Assume that the map from the hidden\-state tuple\(𝐡1lr,…,𝐡Tlr\)\(\\mathbf\{h\}\_\{1\}^\{\\,l\_\{\\mathrm\{r\}\}\},\\dots,\\mathbf\{h\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\}\)to𝐨Tl\\mathbf\{o\}\_\{T\}^\{l\}is locally Lipschitz in a neighborhood of the clean trajectory, and letLαL\_\{\\alpha\}denote the Lipschitz constant\.
Assumption[1](https://arxiv.org/html/2607.28959#Thmassumption1)ensures that small perturbations in the hidden states induce controlled changes in the downstream attention output\. It is standard for Transformer blocks composed of linear maps and softmax\-based attention in a bounded neighborhood of the clean trajectory\(Kimet al\.,[2021](https://arxiv.org/html/2607.28959#bib.bib62)\)\.
### 5\.1Suffix\-Window Defense
We first analyze the limitation of defending only positionTT\. By investigating the layer immediately following the defense layer, namelyl=lr\+1l=l\_\{\\mathrm\{r\}\}\+1, we study how residual perturbations atlrl\_\{\\mathrm\{r\}\}propagate to the attention output𝐨Tl\\mathbf\{o\}\_\{T\}^\{l\}at the final positionTTthrough attention aggregation\. Let𝒯atk=\{T−k\+1,…,T\}\\mathcal\{T\}\_\{\\mathrm\{atk\}\}=\\\{T\-k\+1,\\dots,T\\\}denote an attacked suffix index set withk≥1k\\geq 1such that at layerlrl\_\{\\mathrm\{r\}\}the perturbed hidden states satisfy𝐡~tlr=𝐡tlr\+Δt,\\tilde\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}=\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\+\\Delta\_\{t\},for allt∈𝒯atkt\\in\\mathcal\{T\}\_\{\\mathrm\{atk\}\}, andΔt=𝟎\\Delta\_\{t\}=\\mathbf\{0\}fort∉𝒯atkt\\notin\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\.
###### Theorem 1\(Last\-token defense is insufficient\)\.
Under Assumption[1](https://arxiv.org/html/2607.28959#Thmassumption1), further assume that the ReFT operatorΦ𝐑\\Phi\_\{\\mathbf\{R\}\}is applied only at positionTTand achieves‖𝐡^Tlr−𝐡Tlr‖≤εT\\\|\\hat\{\\mathbf\{h\}\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\}\-\\mathbf\{h\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\}\\\|\\leq\\varepsilon\_\{T\}, whereεT\\varepsilon\_\{T\}measures how close the defended hidden state at T is to the clean hidden state after ReFT intervention\. Since ReFT is applied only at positionTT, the remaining attacked suffix positions are left uncorrected:𝐡^tlr=𝐡tlr\+Δt,t∈𝒯atk∖\{T\}\.\\hat\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}=\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\+\\Delta\_\{t\},t\\in\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\\setminus\\\{T\\\}\.Let𝐨^Tl\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}denote the defended attention outputs atTTin layerll\. LetℰT:=αT,Tl‖𝐖Vl‖εT\\mathcal\{E\}\_\{T\}:=\\alpha\_\{T,T\}^\{l\}\\\|\\mathbf\{W\}\_\{V\}^\{l\}\\\|\\,\\varepsilon\_\{T\}and letℰattn=∑t=1T\|α^T,tl−αT,tl\|⋅‖𝐯^tl‖\.\\mathcal\{E\}\_\{\\mathrm\{attn\}\}=\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|\\cdot\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\.Then
‖𝐨^Tl−𝐨Tl‖≥‖∑t=T−k\+1T−1αT,tl𝐖VlΔt‖−ℰattn−ℰT\.\\bigl\\\|\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}\\bigr\\\|\\;\\geq\\;\\left\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\alpha\_\{T,t\}^\{l\}\\,\\mathbf\{W\}\_\{V\}^\{l\}\\Delta\_\{t\}\\right\\\|\-\\mathcal\{E\}\_\{\\mathrm\{attn\}\}\-\\mathcal\{E\}\_\{T\}\.\(5\)
The first term on the right\-hand side of Eq\. \([5](https://arxiv.org/html/2607.28959#S5.E5)\) captures the contribution of upstream perturbations through the value path\. The termℰT\\mathcal\{E\}\_\{T\}reflects imperfect ReFT intervention at positionTT, whileℰattn\\mathcal\{E\}\_\{\\mathrm\{attn\}\}captures perturbation\-induced changes in the attention weights through the query/key path\. Theorem[1](https://arxiv.org/html/2607.28959#Thmtheorem1)shows that defending only positionTTdoes not generally block adversarial information from earlier suffix positions\. Beyond this attention\-level leakage, a more refined analysis also shows that a ReFT module trained only atTTdoes not generally transfer to earlier positions\. See Appendix[D\.2](https://arxiv.org/html/2607.28959#A4.SS2)and[D\.5](https://arxiv.org/html/2607.28959#A4.SS5)\.
In contrast to the above limitation, we next consider a sequence\-shared ReFT operatorΦshared\\Phi\_\{\\mathrm\{shared\}\}trained jointly over the defended suffix window𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\.
###### Theorem 2\(Sequence\-shared defense over a suffix window\)\.
Under Assumption[1](https://arxiv.org/html/2607.28959#Thmassumption1), supposeΦshared\\Phi\_\{\\mathrm\{shared\}\}satisfies‖Φshared\(𝐡~tlr\)−𝐡tlr‖≤εLAT,\\\|\\Phi\_\{\\mathrm\{shared\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\)\-\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\\\|\\leq\\varepsilon\_\{\\mathrm\{LAT\}\},for allt∈𝒯LATt\\in\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\. HereεLAT\\varepsilon\_\{\\mathrm\{LAT\}\}quantifies how accurately the shared ReFT operator restores the clean hidden states across the defended suffix window\. Further assume that the defended value vectors satisfy‖𝐯^tl‖≤MV\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\\leq M\_\{V\}for someMVM\_\{V\}for anyt∈\[T\]t\\in\[T\]\. For any token\-level attack supported on a suffix set𝒯atk\\mathcal\{T\}\_\{\\mathrm\{atk\}\}such that𝒯atk⊆𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\\subseteq\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}, the defended attention output at positionTTin layerllsatisfies
‖𝐨^Tl−𝐨Tl‖≤εLAT\(BV\+MVLα\(BQ\+kBK\)\)\.\\\|\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}\\\|\\leq\\varepsilon\_\{\\mathrm\{LAT\}\}\\bigl\(B\_\{V\}\+M\_\{V\}L\_\{\\alpha\}\(B\_\{Q\}\+kB\_\{K\}\)\\bigr\)\.\(6\)
Theorem[2](https://arxiv.org/html/2607.28959#Thmtheorem2)shows that sequence\-shared ReFT over a suffix window controls perturbation propagation through both the value path and the attention\-weight path, providing theoretical support for window\-based defense rather than single\-token defense\. The full proof appears in Appendix[D\.3](https://arxiv.org/html/2607.28959#A4.SS3)\.
### 5\.2Effectiveness of the Surrogate Model
While the above presents the effectiveness of our design on the defense side, in the following, we further give an informal theorem explaining why a pruned surrogate can remain effective for inner\-loop attack generation\. The formal version and proof are presented in Appendix[D\.4](https://arxiv.org/html/2607.28959#A4.SS4)\.
###### Theorem 3\(Surrogate approximation\)\.
LetSSdenote the set of neurons retained by the surrogate\. Under local regularity conditions, with probability at least1−2Kp1\-2Kpforp∈\(0,1/\(2K\)\)p\\in\(0,1/\(2K\)\), the PGD optimization gap between the full defended model and the surrogate overKKsteps satisfies
∑s=1KGs≤Csurrη∑s=1KMs\(S\)\+o\(Kη\),\\sum\_\{s=1\}^\{K\}G\_\{s\}\\;\\leq\\;C\_\{\\mathrm\{surr\}\}\\,\\eta\\sum\_\{s=1\}^\{K\}M\_\{s\}\(S\)\\;\+\\;o\(K\\eta\),\(7\)whereGsG\_\{s\}is the difference between the one\-step improvement in the full inner objective obtained by the full PGD direction and that obtained by the surrogate PGD direction,CsurrC\_\{\\mathrm\{surr\}\}is a constant depending on the Jacobian bound from the attacked hidden state to the MLP expansion layer,η\\etais the PGD step size, andws,j:=\|as,j\(la\)\(x\)rs,j\(la\)\|2,Ms\(S\):=\(∑j∉Sws,j\)1/2,w\_\{s,j\}:=\\left\|a\_\{s,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(x\)\\,r\_\{s,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\\right\|^\{2\},M\_\{s\}\(S\):=\\left\(\\sum\_\{j\\notin S\}w\_\{s,j\}\\right\)^\{1/2\},wherews,jw\_\{s,j\}measures the act×\\timesgrad contribution of neuronjjat stepss,Ms\(S\)M\_\{s\}\(S\)is the omitted act×\\timesgrad mass of the neurons pruned by the surrogate at stepss, andrs,j\(la\)r\_\{s,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}denotes the corresponding MLP\-path gradient component\.
Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3)shows that surrogate quality is controlled by the omitted act×\\timesgrad massMs\(S\)M\_\{s\}\(S\)\. At a fixed pruning ratioρ\\rho, an effective surrogate should therefore retain neurons with the largestws,jw\_\{s,j\}\. In contrast, if random pruning discards a fractionρ\\rhoof neurons uniformly at random, then for the retained setSrS\_\{\\mathrm\{r\}\},𝔼\[Ms\(Sr\)2\]=ρ∑jws,j\.\\mathbb\{E\}\\\!\\left\[M\_\{s\}\(S\_\{\\mathrm\{r\}\}\)^\{2\}\\right\]=\\rho\\sum\_\{j\}w\_\{s,j\}\.Thus, random pruning removes a constant fraction of the total act×\\timesgrad mass in expectation, whereas ActGrad pruning preferentially keeps the largest\-contributing neurons and therefore yields a smaller omitted mass at the same sparsity level\.
## 6Experiments
We evaluate the robustness–efficiency tradeoff of LAT\-ReFT and test two key design choices: suffix\-window placement and ActGrad\-based surrogate pruning\.
Table 1:Main results on IMDB across three representative models\. We report clean accuracy \(Acc\) and attack success rate \(ASR\) under RandomToken and GCG\. We also report training\-efficiency statistics: trainable parameters as a percentage of total model parameters and per\-step adversarial\-training compute \(FLOPs/step\)\. See more results and FLOPs details in Appendix[B\.1](https://arxiv.org/html/2607.28959#A2.SS1)and[E](https://arxiv.org/html/2607.28959#A5)\.### 6\.1Experimental Setups
#### Datasets and Models\.
FollowingHoweet al\.\([2024](https://arxiv.org/html/2607.28959#bib.bib7)\), we use three classification tasks: IMDB, EnronSpam and PasswordMatch\. IMDB and EnronSpam are standard natural\-language classification tasks for sentiment analysis and spam detection, respectively, and serve as realistic benchmarks in language understanding settings\(Maaset al\.,[2011](https://arxiv.org/html/2607.28959#bib.bib47); Metsiset al\.,[2006](https://arxiv.org/html/2607.28959#bib.bib46)\)\. PasswordMatch is procedurally constructed tasks inspired by TensorTrust\(Toyeret al\.,[2023](https://arxiv.org/html/2607.28959#bib.bib49)\), and provides a more controlled setting for studying adversarial robustness\. We evaluate three pre\-trained LLMs: Llama\-3\.1\-8B\(Grattafioriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib42)\), Qwen\-2\.5\-3B\(Yanget al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib43)\)and Pythia\-1\.4B\(Bidermanet al\.,[2023](https://arxiv.org/html/2607.28959#bib.bib45)\)\. For classification tasks, we attach a task\-specific linear classification head to the last\-token hidden state and fine\-tune each model on clean data to obtain a task\-adapted classifier for subsequent adversarial training\(Howeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib7)\)\. Detailed settings are provided in Appendix[A\.1](https://arxiv.org/html/2607.28959#A1.SS1)\.
#### Defenses and attack methods\.
We compare three baselines: R2D2, CAT\(Xhonneuxet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib6)\), and LAT\(Sheshadriet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib4); Casperet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib5)\)\. For R2D2, we follow the version ofHoweet al\.\([2024](https://arxiv.org/html/2607.28959#bib.bib7)\), which maintains and continually refreshes a pool of previously generated adversarial examples during training rather than generating all attacks from scratch at every step\(Mazeikaet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib52)\)\. CAT performs adversarial training in continuous embedding space by applying PGD\-based perturbations to input token embeddings, whereas LAT perturbs the model’s latent representations in intermediate layers during training\. All baselines are implemented in our classification setting; additional details are provided in Appendix[A\.4](https://arxiv.org/html/2607.28959#A1.SS4)\. For our method, we implement LAT\-ReFT with the ReFT intervention over a defended suffix window and use the ActGrad\-pruned surrogate for inner\-loop attack generation\. The exact details are provided in Appendix[A\.3](https://arxiv.org/html/2607.28959#A1.SS3)\. We evaluate robustness primarily under token\-level suffix attacks, which match the attack setting assumed by our method and theory\. Specifically, we consider attacks that append adversarial suffixes: RandomToken\(Howeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib7)\)and Greedy Coordinate Gradient \(GCG\)\(Zouet al\.,[2023b](https://arxiv.org/html/2607.28959#bib.bib1)\)as representative for random sampling and gradient\-based optimization\. Additional details are provided in Appendix[A\.5](https://arxiv.org/html/2607.28959#A1.SS5)\.
### 6\.2Main Results
Table[1](https://arxiv.org/html/2607.28959#S6.T1)summarizes the main robustness–efficiency results\. Full\-parameter baselines such as LAT and R2D2 provide strong robustness, but require full\-model updates and substantial attack\-generation compute; R2D2 may further improve with more training rounds at higher offline cost\. CAT reduces the number of trainable parameters while still performing full\-model adversarial search\. In contrast, our method uses only0\.0066%0\.0066\\%–0\.0203%0\.0203\\%trainable parameters and reduces per\-step FLOPs by48\.1%48\.1\\%on average compared with LAT\. As expected, this efficiency gain comes with a loss in absolute robustness compared with full\-parameter LAT\. For example, on Llama\-3\.1\-8B, LAT achieves the lowest GCG ASR of0\.020\.02, while our method obtains0\.150\.15with much lower FLOPs and only0\.0066%0\.0066\\%trainable parameters\. RandomToken is more model\-dependent\. On Qwen\-2\.5\-3B, our method reduces ASR from0\.980\.98to0\.380\.38, but remains less effective than under GCG\. One possible reason is that RandomToken repeatedly samples perturbations over the suffix, inducing a broader and less gradient\-aligned attack distribution than the perturbations used during training\. These random suffixes appear to exploit a negative\-to\-positive class bias not fully exposed by GCG\. Overall, LAT\-ReFT with a circuit\-guided surrogate provides a lightweight alternative to full\-parameter adversarial training with a favorable robustness–efficiency tradeoff\.
### 6\.3Design Analysis of the Defense
Figure 2:Defense\-placement ablations on IMDB with Pythia\-1\.4B under GCG and RandomToken attacks\. \(a\) Suffix\-window length ablation with the defense layer fixed atl=12l=12\. \(b\) Defense\-layer ablation with the suffix length fixed atL=20L=20\. Defending a suffix window and avoiding very late intervention layers both improve robustness, especially under GCG\.We study the two factors of LAT\-ReFT mentioned in Section[4\.1](https://arxiv.org/html/2607.28959#S4.SS1): the suffix window and the defense layer\. Figure[2](https://arxiv.org/html/2607.28959#S6.F2)summarizes the corresponding ablations\.
#### Suffix\-window size\.
Figure[2](https://arxiv.org/html/2607.28959#S6.F2)\(a\) shows that defending only the final suffix token is insufficient, which is consistent with Theorem[1](https://arxiv.org/html/2607.28959#Thmtheorem1)\. Specifically, as the defended suffix lengthLLincreases from 1 to 20, the green bars \(GCG ASR\) decrease substantially, from 61% to 26%\. This indicates that broader suffix coverage is important for suppressing adversarial leakage from nearby attacked positions, which aligns with Theorem[2](https://arxiv.org/html/2607.28959#Thmtheorem2)\. By contrast, yellow bars \(RandomToken ASR\) remain low and vary little withLL, likely because its random perturbations are less consistently aligned with the suffix leakage path characterized in Theorem[1](https://arxiv.org/html/2607.28959#Thmtheorem1)\. The blue bars \(clean accuracy\) also remain unchanged, showing that the robustness gains under GCG are not driven by sacrificing clean performance\.
#### Defense layer\.
Figure[2](https://arxiv.org/html/2607.28959#S6.F2)\(b\) shows that placing the defense too late in the model is ineffective\. Specifically, the green bars \(GCG ASR\) are much higher for late\-layer intervention than for early or middle placement, while the yellow bars \(RandomToken ASR\) show the same qualitative trend at a lower overall level\. In contrast, the blue bars \(clean accuracy\) remain consistently high across all three placements\. To explain the worse performance of late layers, the adversarial perturbation has already propagated through many layers by the time it reaches the final layer, leaving less room for a low\-rank intervention to remove it reliably\. This is broadly consistent with a recent work where decoder\-only Transformers organize computation into distinct depth\-wise phases, with later layers acting more as selective refinement stages\(Queipo\-de\-Llanoet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib55)\)\.
Overall, these results support both parts of our placement strategy: defending a suffix window rather than a single token, and avoiding very late\-layer intervention when applying the latent defense\.
### 6\.4Analysis of Circuit\-Guided Surrogate
We next analyze whether the circuit\-guided surrogate can reduce inner\-loop attack cost while preserving attack transferability\. Figure[3](https://arxiv.org/html/2607.28959#S6.F3)evaluates this design from three aspects: \(1\) the pruning\-ratio tradeoff between training cost and attack success rate, \(2\) the impact of different neuron\-selection rules, and \(3\) the resulting layer\-wise pruning patterns\.
#### Tradeoff between pruning and ASR\.
We first study the tradeoff between the pruning and ASR on IMDB with Pythia\-1\.4B under GCG attack as a representative setting\. Figure[3](https://arxiv.org/html/2607.28959#S6.F3)\(a\) reports the resulting ASR together with the total FLOPs per adversarial training step\. The results show that moderate pruning remains effective: at 25% pruning, the surrogate reduces FLOPs to0\.804×0\.804\\timesthat of the unpruned model, while ASR increases only slightly from0\.260\.26to0\.330\.33\. In contrast, more aggressive pruning sharply degrades transfer: at 50% pruning, ASR rises to0\.890\.89, and at 75% pruning it further increases to0\.950\.95\. We therefore use a pruning ratio of25%25\\%in all remaining experiments\. This degradation is expected: the surrogate is useful only if its PGD directions transfer to the full defended model\. Moderate pruning preserves these adversarial search directions, whereas aggressive pruning distorts them and weakens transfer\.
#### Selection rule comparison\.
We next compare three neuron\-selection rules at the same pruning ratio of 25%: random selection \(serves as an unstructured pruning baseline\), activation\-difference scoring \(ActDiff\), and our activation\-times\-gradient \(ActGrad\) rule\. This tests whether surrogate quality depends on which neurons are retained, rather than only on the pruning ratio\. ActDiff is motivated by the difference\-in\-means style of representation analysis widely used in mechanistic interpretability and steering\-vector methods\(Marks and Tegmark,[2023](https://arxiv.org/html/2607.28959#bib.bib57)\)\. Figure[3](https://arxiv.org/html/2607.28959#S6.F3)\(b\) shows that the distinction is crucial\. While the clean accuracies remain high for all methods, at the same pruning ratio, random pruning gives ASR0\.950\.95, and ActDiff performs similarly poorly at0\.970\.97\. In contrast, ActGrad reduces ASR to0\.330\.33\. This supports the effectiveness of ActGrad and aligns with Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3)\.
Figure 3:Analysis of the circuit\-guided surrogate on IMDB with Pythia\-1\.4B under GCG attack\. \(a\) Tradeoff between ASR and normalized training FLOPs across pruning ratios\. \(b\) Selection\-rule comparison at25%25\\%pruning\. ActGrad strongly outperforms random pruning and activation\-difference scoring in ASR\. \(c\) Layer\-by\-neuron patterns at25%25\\%pruning for ActGrad and ActDiff\.
#### Neuron pruning patterns\.
Figure[3](https://arxiv.org/html/2607.28959#S6.F3)\(c\) provides a more detailed view of the resulting patterns at the default setting25%25\\%pruning, using layer\-by\-neuron heatmaps for ActGrad and ActDiff\. The two methods produce different structures: ActGrad preserves more capacity in earlier and middle layers, with a relatively sharp transition toward heavier pruning only in later layers\. By contrast, ActDiff allocates its budget more diffusely and places comparatively more mass in later\-layer regions\. Together with the results in Figure[3](https://arxiv.org/html/2607.28959#S6.F3)\(b\), this suggests that, in our setting, preserving earlier\- and middle\-layer MLP computation is more important for maintaining transferable adversarial directions\. This may explain why ActDiff performs slightly worse than random pruning: its retained computation is less concentrated in earlier and middle layers\. See more details in Appendix[C](https://arxiv.org/html/2607.28959#A3)\.
## 7Conclusion and Limitation
We proposed an efficient latent adversarial training framework for LLMs that improves the robustness–efficiency tradeoff from both defense and attack sides\. On the defense side, LAT\-ReFT combines LAT with ReFT and shows that effective intervention should cover a suffix window and avoid very late layers\. On the attack side, our circuit\-guided surrogate retains high\-importance MLP neurons using activation\-gradient scores\. Together, these components provide a lightweight alternative to full\-parameter adversarial training, supported by theoretical analysis and empirical evidence\. There are still some limitations in our work\. First, we mainly study classification settings and suffix\-style attacks, with preliminary prefix\-attack results in Appendix[B\.3](https://arxiv.org/html/2607.28959#A2.SS3); future work should extend to generation tasks and adaptive selection of defended token positions under broader jailbreak attacks\. Second, our surrogate uses a fixed MLP pruning ratio, and adaptive pruning or broader circuit extraction may further improve the speed–transferability tradeoff\.
## References
- Jailbreaking leading safety\-aligned llms with simple adaptive attacks\.arXiv preprint arXiv:2404\.02151\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p1.1)\.
- M\. Andriushchenko and N\. Flammarion \(2020\)Understanding and improving fast adversarial training\.Advances in Neural Information Processing Systems33,pp\. 16048–16059\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. Nanda \(2024\)Refusal in language models is mediated by a single direction\.Advances in Neural Information Processing Systems37,pp\. 136037–136083\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff,et al\.\(2023\)Pythia: a suite for analyzing large language models across training and scaling\.InInternational Conference on Machine Learning,pp\. 2397–2430\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- M\. Burgess \(2025\)Hackers hijacked google’s gemini ai with a poisoned calendar invite to take over a smart home\.Note:[https://www\.wired\.com/story/google\-gemini\-calendar\-invite\-hijack\-smart\-home/](https://www.wired.com/story/google-gemini-calendar-invite-hijack-smart-home/)WIRED, accessed April 19, 2026Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p1.1)\.
- S\. Casper, L\. Schulze, O\. Patel, and D\. Hadfield\-Menell \(2024\)Defending against unforeseen failure modes with latent adversarial training\.arXiv preprint arXiv:2403\.05030\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.28959#S3.SS0.SSS0.Px2.p2.4),[§4\.1](https://arxiv.org/html/2607.28959#S4.SS1.SSS0.Px2.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
- P\. Chao, A\. Robey, E\. Dobriban, H\. Hassani, G\. J\. Pappas, and E\. Wong \(2025\)Jailbreaking black box large language models in twenty queries\.In2025 IEEE Conference on Secure and Trustworthy Machine Learning \(SaTML\),pp\. 23–42\.Cited by:[§A\.5](https://arxiv.org/html/2607.28959#A1.SS5.SSS0.Px3.p1.1)\.
- J\. Chen, X\. Wang, Z\. Yao, Y\. Bai, L\. Hou, and J\. Li \(2024\)Towards understanding safety alignment: a mechanistic perspective from safety neurons\.arXiv preprint arXiv:2406\.14144\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Dékány, S\. Balauca, R\. Staab, D\. I\. Dimitrov, and M\. Vechev \(2025\)Mixat: combining continuous and discrete adversarial training for llms\.arXiv preprint arXiv:2505\.16947\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Ebrahimi, A\. Rao, D\. Lowd, and D\. Dou \(2018\)Hotflip: white\-box adversarial examples for text classification\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 31–36\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- B\. L\. Edelman, S\. Goel, S\. Kakade, and C\. Zhang \(2022\)Inductive biases and variable creation in self\-attention mechanisms\.InInternational Conference on Machine Learning,pp\. 5793–5831\.Cited by:[§5](https://arxiv.org/html/2607.28959#S5.p1.19)\.
- N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen,et al\.\(2022\)Toy models of superposition\.arXiv preprint arXiv:2209\.10652\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1)\.
- E\. Frantar and D\. Alistarh \(2023\)Sparsegpt: massive language models can be accurately pruned in one\-shot\.InInternational conference on machine learning,pp\. 10323–10337\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Fu and D\. Wang \(2026\)Understanding and improving continuous llm adversarial training via in\-context learning theory\.International Conference on Learning Representations\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Geva, R\. Schuster, J\. Berant, and O\. Levy \(2021\)Transformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5484–5495\.Cited by:[§4\.2](https://arxiv.org/html/2607.28959#S4.SS2.p2.2)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- M\. Hanna, S\. Pezzelle, and Y\. Belinkov \(2024\)Have faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.arXiv preprint arXiv:2403\.17806\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- R\. Hart \(2026\)The ai security nightmare is here and it looks suspiciously like lobster\.Note:[https://www\.theverge\.com/ai\-artificial\-intelligence/881574/cline\-openclaw\-prompt\-injection\-hack](https://www.theverge.com/ai-artificial-intelligence/881574/cline-openclaw-prompt-injection-hack)The Verge, accessed April 19, 2026Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p1.1)\.
- W\. He and Y\. Xing \(2025\)Impact of positional encoding: clean and adversarial rademacher complexity for transformers under in\-context regression\.arXiv preprint arXiv:2512\.09275\.Cited by:[§5](https://arxiv.org/html/2607.28959#S5.p1.19)\.
- N\. Howe, I\. McKenzie, O\. Hollinsworth, M\. Zajac, T\. Tseng, A\. Tucker, P\. Bacon, and A\. Gleave \(2024\)Scaling trends in language model robustness\.arXiv preprint arXiv:2407\.18213\.Cited by:[§A\.1](https://arxiv.org/html/2607.28959#A1.SS1.p1.1),[§A\.4](https://arxiv.org/html/2607.28959#A1.SS4.SSS0.Px1.p1.27),[§A\.5](https://arxiv.org/html/2607.28959#A1.SS5.SSS0.Px2.p1.2),[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§1](https://arxiv.org/html/2607.28959#S1.p4.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
- J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. Amodei \(2020\)Scaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.Cited by:[Appendix E](https://arxiv.org/html/2607.28959#A5.p1.2),[§1](https://arxiv.org/html/2607.28959#S1.p2.1)\.
- H\. Kim, G\. Papamakarios, and A\. Mnih \(2021\)The lipschitz constant of self\-attention\.InInternational Conference on Machine Learning,pp\. 5562–5571\.Cited by:[§5](https://arxiv.org/html/2607.28959#S5.p2.1)\.
- A\. Lee, X\. Bai, I\. Pres, M\. Wattenberg, J\. K\. Kummerfeld, and R\. Mihalcea \(2024\)A mechanistic understanding of alignment algorithms: a case study on dpo and toxicity\.arXiv preprint arXiv:2401\.01967\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Lin, P\. He, H\. Xu, Y\. Xing, M\. Yamada, H\. Liu, and J\. Tang \(2024\)Towards understanding jailbreak attacks in llms: a representation space analysis\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 7067–7085\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Maas, R\. E\. Daly, P\. T\. Pham, D\. Huang, A\. Y\. Ng, and C\. Potts \(2011\)Learning word vectors for sentiment analysis\.InProceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies,pp\. 142–150\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu \(2017\)Towards deep learning models resistant to adversarial attacks\.arXiv preprint arXiv:1706\.06083\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Marks and M\. Tegmark \(2023\)The geometry of truth: emergent linear structure in large language model representations of true/false datasets\.arXiv preprint arXiv:2310\.06824\.Cited by:[§6\.4](https://arxiv.org/html/2607.28959#S6.SS4.SSS0.Px2.p1.3)\.
- M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang, N\. Mu, E\. Sakhaee, N\. Li, S\. Basart, B\. Li,et al\.\(2024\)Harmbench: a standardized evaluation framework for automated red teaming and robust refusal\.arXiv preprint arXiv:2402\.04249\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
- V\. Metsis, I\. Androutsopoulos, and G\. Paliouras \(2006\)Spam filtering with naive bayes\-which naive bayes?\.InCEAS,Vol\.17,pp\. 28–69\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- P\. Molchanov, A\. Mallya, S\. Tyree, I\. Frosio, and J\. Kautz \(2019\)Importance estimation for neural network pruning\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 11264–11272\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1)\.
- P\. Molchanov, S\. Tyree, T\. Karras, T\. Aila, and J\. Kautz \(2016\)Pruning convolutional neural networks for resource efficient inference\.arXiv preprint arXiv:1611\.06440\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1)\.
- P\. Nair \(2025\)Softmax is1/21/2\-lipschitz: a tight bound across allℓp\\ell\_\{p\}norms\.arXiv preprint arXiv:2510\.23012\.Cited by:[§D\.4](https://arxiv.org/html/2607.28959#A4.SS4.3.p3.5)\.
- D\. Patel, G\. Gervacio, D\. Raimi, K\. Zhu, R\. Lagasse, G\. Grand, A\. Panda, and M\. Chaudhary \(2025\)Alignment\-constrained dynamic pruning for llms: identifying and preserving alignment\-critical circuits\.arXiv preprint arXiv:2511\.07482\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- E\. Queipo\-de\-Llano, Á\. Arroyo, F\. Barbero, X\. Dong, M\. Bronstein, Y\. LeCun, and R\. Shwartz\-Ziv \(2025\)Attention sinks and compression valleys in llms are two sides of the same coin\.arXiv preprint arXiv:2510\.06477\.Cited by:[3rd item](https://arxiv.org/html/2607.28959#S1.I1.i3.p1.1),[§6\.3](https://arxiv.org/html/2607.28959#S6.SS3.SSS0.Px2.p1.1)\.
- J\. Ren, Z\. Dai, X\. Tang, H\. Liu, J\. Zeng, Z\. Li, R\. Goutam, S\. Wang, Y\. Xing, and Q\. He \(2025\)A general framework to enhance fine\-tuning\-based llm unlearning\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18464–18476\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p3.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Shafahi, M\. Najibi, M\. A\. Ghiasi, Z\. Xu, J\. Dickerson, C\. Studer, L\. S\. Davis, G\. Taylor, and T\. Goldstein \(2019\)Adversarial training for free\!\.Advances in neural information processing systems32\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- L\. Sharkey, B\. Chughtai, J\. Batson, J\. Lindsey, J\. Wu, L\. Bushnaq, N\. Goldowsky\-Dill, S\. Heimersheim, A\. Ortega, J\. Bloom,et al\.\(2025\)Open problems in mechanistic interpretability\.arXiv preprint arXiv:2501\.16496\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Sheshadri, A\. Ewart, P\. Guo, A\. Lynch, C\. Wu, V\. Hebbar, H\. Sleight, A\. C\. Stickland, E\. Perez, D\. Hadfield\-Menell,et al\.\(2024\)Latent adversarial training improves robustness to persistent harmful behaviors in llms\.arXiv preprint arXiv:2407\.15549\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.28959#S3.SS0.SSS0.Px2.p2.4),[§4\.1](https://arxiv.org/html/2607.28959#S4.SS1.SSS0.Px2.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
- O\. Skean, M\. R\. Arefin, D\. Zhao, N\. Patel, J\. Naghiyev, Y\. LeCun, and R\. Shwartz\-Ziv \(2025\)Layer by layer: uncovering hidden representations in language models\.arXiv preprint arXiv:2502\.02013\.Cited by:[3rd item](https://arxiv.org/html/2607.28959#S1.I1.i3.p1.1),[§4\.1](https://arxiv.org/html/2607.28959#S4.SS1.SSS0.Px2.p1.1)\.
- A\. Syed, C\. Rager, and A\. Conmy \(2024\)Attribution patching outperforms automated circuit discovery\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 407–416\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1),[§4\.2](https://arxiv.org/html/2607.28959#S4.SS2.p2.2)\.
- S\. Toyer, O\. Watkins, E\. A\. Mendes, J\. Svegliato, L\. Bailey, T\. Wang, I\. Ong, K\. Elmaaroufi, P\. Abbeel, T\. Darrell,et al\.\(2023\)Tensor trust: interpretable prompt injection attacks from an online game\.arXiv preprint arXiv:2311\.01011\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid \(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin \(2017\)Attention is all you need\.Advances in neural information processing systems30\.Cited by:[§3](https://arxiv.org/html/2607.28959#S3.SS0.SSS0.Px1.p1.7)\.
- J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber \(2020\)Investigating gender bias in language models using causal mediation analysis\.Advances in neural information processing systems33,pp\. 12388–12401\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- K\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt \(2022\)Interpretability in the wild: a circuit for indirect object identification in gpt\-2 small\.arXiv preprint arXiv:2211\.00593\.Cited by:[2nd item](https://arxiv.org/html/2607.28959#S1.I1.i2.p1.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px3.p1.1)\.
- E\. Wong, L\. Rice, and J\. Z\. Kolter \(2020\)Fast is better than free: revisiting adversarial training\.arXiv preprint arXiv:2001\.03994\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Wu, A\. Arora, Z\. Wang, A\. Geiger, D\. Jurafsky, C\. D\. Manning, and C\. Potts \(2024\)Reft: representation finetuning for language models\.Advances in Neural Information Processing Systems37,pp\. 63908–63962\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p3.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2607.28959#S3.SS0.SSS0.Px3.p1.6)\.
- S\. Xhonneux, A\. Sordoni, S\. Günnemann, G\. Gidel, and L\. Schwinn \(2024\)Efficient adversarial training in llms with continuous attacks\.Advances in Neural Information Processing Systems37,pp\. 1502–1530\.Cited by:[§A\.4](https://arxiv.org/html/2607.28959#A1.SS4.SSS0.Px2.p1.12),[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
- Y\. Xing, Q\. Song, and G\. Cheng \(2021\)On the algorithmic stability of adversarial training\.Advances in neural information processing systems34,pp\. 26523–26535\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p2.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu \(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px1.p1.1)\.
- J\. Yu, X\. Lin, Z\. Yu, and X\. Xing \(2023\)Gptfuzzer: red teaming large language models with auto\-generated jailbreak prompts\.arXiv preprint arXiv:2309\.10253\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p1.1)\.
- S\. Zeng, P\. He, K\. Guo, T\. Zheng, H\. Lu, Y\. Xing, and H\. Liu \(2025\)Towards context\-robust llms: a gated representation fine\-tuning approach\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 10262–10276\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p3.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Zhao, L\. Zhang, J\. Ye, H\. Lu, B\. Yin, and X\. Wang \(2024\)Adversarial training: a survey\.arXiv preprint arXiv:2410\.15042\.Cited by:[§1](https://arxiv.org/html/2607.28959#S1.p2.1)\.
- C\. Zheng, F\. Yin, H\. Zhou, F\. Meng, J\. Zhou, K\. Chang, M\. Huang, and N\. Peng \(2024\)Prompt\-driven llm safeguarding via directed representation optimization\.arXiv preprint arXiv:2401\.180183\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.\(2023a\)Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson \(2023b\)Universal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[§A\.5](https://arxiv.org/html/2607.28959#A1.SS5.SSS0.Px1.p1.4),[§1](https://arxiv.org/html/2607.28959#S1.p1.1),[§1](https://arxiv.org/html/2607.28959#S1.p2.1),[§2](https://arxiv.org/html/2607.28959#S2.SS0.SSS0.Px1.p1.1),[§6\.1](https://arxiv.org/html/2607.28959#S6.SS1.SSS0.Px2.p1.1)\.
## Appendix
###### Contents
1. [1Introduction](https://arxiv.org/html/2607.28959#S1)
2. [2Related Work](https://arxiv.org/html/2607.28959#S2)
3. [3Preliminaries](https://arxiv.org/html/2607.28959#S3)
4. [4Method](https://arxiv.org/html/2607.28959#S4)1. [4\.1Defense: LAT\-ReFT](https://arxiv.org/html/2607.28959#S4.SS1) 2. [4\.2Circuit\-Guided Surrogate for Attack Generation](https://arxiv.org/html/2607.28959#S4.SS2)
5. [5Theoretical Analysis](https://arxiv.org/html/2607.28959#S5)1. [5\.1Suffix\-Window Defense](https://arxiv.org/html/2607.28959#S5.SS1) 2. [5\.2Effectiveness of the Surrogate Model](https://arxiv.org/html/2607.28959#S5.SS2)
6. [6Experiments](https://arxiv.org/html/2607.28959#S6)1. [6\.1Experimental Setups](https://arxiv.org/html/2607.28959#S6.SS1) 2. [6\.2Main Results](https://arxiv.org/html/2607.28959#S6.SS2) 3. [6\.3Design Analysis of the Defense](https://arxiv.org/html/2607.28959#S6.SS3) 4. [6\.4Analysis of Circuit\-Guided Surrogate](https://arxiv.org/html/2607.28959#S6.SS4)
7. [7Conclusion and Limitation](https://arxiv.org/html/2607.28959#S7)
8. [References](https://arxiv.org/html/2607.28959#bib)
9. [AAdditional Experimental Setup](https://arxiv.org/html/2607.28959#A1)1. [A\.1Datasets and Fine\-tuning](https://arxiv.org/html/2607.28959#A1.SS1) 2. [A\.2Evaluation Metric](https://arxiv.org/html/2607.28959#A1.SS2) 3. [A\.3LAT\-ReFT Training Setup](https://arxiv.org/html/2607.28959#A1.SS3) 4. [A\.4Baseline Defense Configurations](https://arxiv.org/html/2607.28959#A1.SS4) 5. [A\.5Adversarial Attacks Setup](https://arxiv.org/html/2607.28959#A1.SS5) 6. [A\.6Computation Cost](https://arxiv.org/html/2607.28959#A1.SS6)
10. [BAdditional Results](https://arxiv.org/html/2607.28959#A2)1. [B\.1More Results](https://arxiv.org/html/2607.28959#A2.SS1) 2. [B\.2Additional Surrogate Analysis](https://arxiv.org/html/2607.28959#A2.SS2) 3. [B\.3Prefix\-attack result](https://arxiv.org/html/2607.28959#A2.SS3)
11. [CCircuit\-Guided Surrogate Details](https://arxiv.org/html/2607.28959#A3)1. [C\.1ActDiff Method\.](https://arxiv.org/html/2607.28959#A3.SS1) 2. [C\.2Surrogate pruning heatmaps](https://arxiv.org/html/2607.28959#A3.SS2)
12. [DAdditional Theory and Proofs](https://arxiv.org/html/2607.28959#A4)1. [D\.1Complete Mathematical Setup](https://arxiv.org/html/2607.28959#A4.SS1) 2. [D\.2Proof of Theorem1](https://arxiv.org/html/2607.28959#A4.SS2) 3. [D\.3Proof of Theorem2](https://arxiv.org/html/2607.28959#A4.SS3) 4. [D\.4Proof of Theorem3](https://arxiv.org/html/2607.28959#A4.SS4) 5. [D\.5A Spatial Transfer Limitation of Last\-Token ReFT](https://arxiv.org/html/2607.28959#A4.SS5)
13. [EEstimated Compute Calculations](https://arxiv.org/html/2607.28959#A5)
## Appendix AAdditional Experimental Setup
### A\.1Datasets and Fine\-tuning
We use 3 binary classification datasets\[Howeet al\.,[2024](https://arxiv.org/html/2607.28959#bib.bib7)\]\. The two natural\-language datasets are filtered to examples with 100 to 1,000 GPT\-2 tokens\. The short\-text datasets are filtered to examples with 5 to 50 tokens\. From the length\-filtered training split of each dataset, we sample up to 20,000 examples with a fixed random seed \(s=42s=42\) to form the fine\-tuning training set \(ft\_train\)\. From the remaining training examples not selected intoft\_train, we further sample up to 100 examples as the attack set, which is disjoint fromft\_trainby construction\. For evaluation, we apply the same length filter to the validation split \(100–1,000 tokens for natural\-language datasets; 5–50 tokens for short\-text datasets\) and retain*all*examples that pass the filter without further subsampling\. Table[2](https://arxiv.org/html/2607.28959#A1.T2)summarizes the resulting split sizes\.
Table 2:Dataset statistics after length filtering\.We fine\-tune each model by attaching a two\-class linear classification head on top of the final hidden state of the last token and training end\-to\-end\. All models are trained for 3 epochs using AdamW with a learning rate of10−510^\{\-5\}and linear decay to zero \(no warmup\)\. The effective batch size is 16 for all models \(per\-device batch size 4 with 4 gradient accumulation steps for 7B models\)\. Input sequences are truncated to 512 tokens for the natural\-language datasets \(IMDB, Enron\) and to 64 tokens for the short\-text datasets \(PasswordMatch\)\. 7B\-scale models are trained in bfloat16\. Table[3](https://arxiv.org/html/2607.28959#A1.T3)reports validation accuracy of the resulting classifiers\.
Table 3:Fine\-tuned classifier validation accuracy \(%\)\.
### A\.2Evaluation Metric
We evaluate each defended classifier using clean accuracy and attack success rate \(ASR\)\. Clean accuracy is computed on the unperturbed test set\. For adversarial evaluation, we report ASR over the full evaluation set: the fraction of all evaluated examples that are correctly classified before attack and misclassified after attack\.
### A\.3LAT\-ReFT Training Setup
Algorithm[1](https://arxiv.org/html/2607.28959#alg1)summarizes one training iteration of LAT\-ReFT with the circuit\-guided surrogate\. For the main IMDB experiments, we train LoReFT interventions on top of task\-specific fine\-tuned classifiers initialized from the corresponding clean baseline checkpoint\. Table[4](https://arxiv.org/html/2607.28959#A1.T4)lists the hyperparameters used for the reported IMDB results with ReFT rank=64 for all models\.
Table 4:LAT\-ReFT hyperparameters used for reported IMDB results\.Algorithm 1One iteration of LAT\-ReFT with circuit\-guided surrogate0:Pretrained model backbone
fθf\_\{\\theta\}, surrogate
fsurrf\_\{\\mathrm\{surr\}\}, attack layer
lal\_\{\\mathrm\{a\}\}, defense layer
lrl\_\{\\mathrm\{r\}\}, defended token set
𝒯LAT\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}, PGD steps
KK, perturbation budget
ε\\varepsilon, adversarial weight
λadv\\lambda\_\{\\mathrm\{adv\}\}, ReFT parameters
\(𝐑,𝐖,𝐛\)\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\), learning rate
η\\eta
1:Freeze the backbone parameters
θ\\theta; optimize only the ReFT parameters
\(𝐑,𝐖,𝐛\)\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\)at layer
lrl\_\{\\mathrm\{r\}\}
2:foreach minibatch
\(x,y\)\(x,y\)do
3:Compute clean hidden states
𝐇la\(x;θ\)\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\)
4:Compute clean loss
ℒclean←ℓ\(f𝐑,𝐖,𝐛\(x\),y\)\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\\leftarrow\\ell\\\!\\left\(f\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\(x\),y\\right\)
5:Construct the surrogate with the current ReFT parameters at layer
lrl\_\{\\mathrm\{r\}\}
6:Initialize
𝐇~\(0\)∈ℬ\(𝐇la\(x;θ\),ε;𝒯LAT\)\\widetilde\{\\mathbf\{H\}\}^\{\(0\)\}\\in\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\),\\varepsilon;\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\)
7:for
s=0,…,K−1s=0,\\dots,K\-1do
8:Compute surrogate inner\-loss gradient
g\(s\)←∇𝐇~ℓ\(fsurr\(x;𝐇~\(s\)\),y\)g^\{\(s\)\}\\leftarrow\\nabla\_\{\\widetilde\{\\mathbf\{H\}\}\}\\ell\\\!\\left\(f\_\{\\mathrm\{surr\}\}\(x;\\widetilde\{\\mathbf\{H\}\}^\{\(s\)\}\),y\\right\)
9:Take one PGD step and project back to the latent attack set
𝐇~\(s\+1\)←Πℬ\(𝐇la\(x;θ\),ε;𝒯LAT\)\(𝐇~\(s\)\+αsign\(g\(s\)\)\)\\widetilde\{\\mathbf\{H\}\}^\{\(s\+1\)\}\\leftarrow\\Pi\_\{\\mathcal\{B\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x;\\theta\),\\varepsilon;\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\)\}\\Bigl\(\\widetilde\{\\mathbf\{H\}\}^\{\(s\)\}\+\\alpha\\,\\mathrm\{sign\}\(g^\{\(s\)\}\)\\Bigr\)
10:endfor
11:Set
𝐇~⋆←𝐇~\(K\)\\widetilde\{\\mathbf\{H\}\}^\{\\star\}\\leftarrow\\widetilde\{\\mathbf\{H\}\}^\{\(K\)\}
12:Compute adversarial loss on the full defended model
ℒadv←ℓ\(f𝐑,𝐖,𝐛\(x;𝐇~⋆\),y\)\\mathcal\{L\}\_\{\\mathrm\{adv\}\}\\leftarrow\\ell\\\!\\left\(f\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\(x;\\widetilde\{\\mathbf\{H\}\}^\{\\star\}\),y\\right\)
13:Form the outer objective
ℒtotal←ℒclean\+λadvℒadv\\mathcal\{L\}\_\{\\mathrm\{total\}\}\\leftarrow\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\+\\lambda\_\{\\mathrm\{adv\}\}\\,\\mathcal\{L\}\_\{\\mathrm\{adv\}\}
14:Update only the ReFT parameters
\(𝐑,𝐖,𝐛\)←\(𝐑,𝐖,𝐛\)−η∇\(𝐑,𝐖,𝐛\)ℒtotal\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\)\\leftarrow\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\)\-\\eta\\,\\nabla\_\{\(\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\)\}\\mathcal\{L\}\_\{\\mathrm\{total\}\}
15:endfor
### A\.4Baseline Defense Configurations
#### R2D2\.
We follow the R2D2 adversarial training setup ofHoweet al\.\[[2024](https://arxiv.org/html/2607.28959#bib.bib7)\]\(Algorithm 1\), including its dynamic adversarial pool construction and mixed clean/adversarial minibatch sampling strategy\. Each experiment runs forRRadversarial training rounds, whereRRis chosen per model based on the available compute budget; specificallyR=5R\\\!=\\\!5for Llama\-3\.1\-8B,R=4R\\\!=\\\!4for Pythia\-1\.4B, andR=7R\\\!=\\\!7for Qwen\-2\.5\-3B\. In each round we attacknadv=200n\_\{\\text\{adv\}\}\{=\}200examples drawn from the attack split using GCG with a suffix ofN=10N\{=\}10tokens,B=128B\{=\}128candidates per iteration, beam widthk=256k\{=\}256, and a linearly increasing iteration budgetk\(r\)=round\(kstart\+rRmax\(kend−kstart\)\)k\(r\)=\\mathrm\{round\}\\\!\\left\(k\_\{\\text\{start\}\}\+\\tfrac\{r\}\{R\_\{\\max\}\}\(k\_\{\\text\{end\}\}\-k\_\{\\text\{start\}\}\)\\right\)withkstart=8k\_\{\\text\{start\}\}\{=\}8,kend=32k\_\{\\text\{end\}\}\{=\}32,Rmax=8R\_\{\\max\}\{=\}8, yieldingk∈\{11,14,17,20,23,26,29,32\}k\\in\\\{11,14,17,20,23,26,29,32\\\}across rounds\. Newly found adversarial examples are added to a persistent pool and resampled via exponential rank\-weighting \(λ=0\.005\\lambda\{=\}0\.005\) that jointly prioritises high\-loss and recently generated examples\. Each round’s fine\-tuning minibatch containsnaug=1000n\_\{\\text\{aug\}\}\{=\}1000items \(80%80\\%adversarial,20%20\\%clean\), trained for200200gradient steps with AdamW \(lr=2×10−5=2\{\\times\}10^\{\-5\}, batch size88, no weight decay; for Llama\-3\.1\-8B we use lr=5×10−6=5\{\\times\}10^\{\-6\}and batch size44to prevent clean\-accuracy collapse\)\. Total FLOPs reported in Table[1](https://arxiv.org/html/2607.28959#S6.T1)are computed asCadv=Csearch\+CtrainC\_\{\\text\{adv\}\}=C\_\{\\text\{search\}\}\+C\_\{\\text\{train\}\}using the Kaplan formulaC=6NDC=6NDwith model parameter countsN∈\{1\.31B,3\.09B,7\.50B\}N\\in\\\{1\.31\\text\{B\},\\,3\.09\\text\{B\},\\,7\.50\\text\{B\}\\\}for Pythia\-1\.4B, Qwen\-2\.5\-3B, and Llama\-3\.1\-8B respectively; training FLOPs account for less than3%3\\%of the total in all cases\.
#### CAT\.
Continuous Adversarial Training \(CAT\) applies projected gradient descent \(PGD\)*in the embedding space*at every training step\. Given a mini\-batch with token embeddings𝐄∈ℝB×T×d\\mathbf\{E\}\\in\\mathbb\{R\}^\{B\\times T\\times d\}, the inner maximization runskkPGD steps within anℓ2\\ell\_\{2\}ball of radiusε\\varepsilon:
𝜹∗=argmax‖𝜹‖2≤εℒ\(fθ\(𝐄\+𝜹\),y\),\\boldsymbol\{\\delta\}^\{\*\}=\\operatorname\*\{arg\\,max\}\_\{\\\|\\boldsymbol\{\\delta\}\\\|\_\{2\}\\leq\\varepsilon\}\\;\\mathcal\{L\}\\\!\\left\(f\_\{\\theta\}\(\\mathbf\{E\}\+\\boldsymbol\{\\delta\}\),\\,y\\right\),and the outer minimization updates the model via
minθℒclean\+λℒadv\.\\min\_\{\\theta\}\\;\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{adv\}\}\.In our implemented CAT baseline, we use LoRA adapters on top of a 4\-bit quantized backbone rather than full\-model fine\-tuning, followingXhonneuxet al\.\[[2024](https://arxiv.org/html/2607.28959#bib.bib6)\]\. For the final runs, we use LoRA rankr=16r=16, LoRA scalingα=32\\alpha=32,ε=0\.05\\varepsilon=0\.05, PGD step size0\.0050\.005,k=10k=10, adversarial weightλ=0\.5\\lambda=0\.5, learning rate10−410^\{\-4\}, and batch size88\.
#### LAT\.
Latent Adversarial Training \(LAT\) performs adversarial training by injecting perturbations into an intermediate hidden representation rather than the input token embeddings\. For an input\-label pair\(x,y\)\(x,y\), let𝐇la\(x\)\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x\)denote the hidden states at transformer layerlal\_\{\\mathrm\{a\}\}\. The inner maximization searches for a latent perturbation within anℓ2\\ell\_\{2\}ball:
𝜹∗=argmax‖𝜹‖2≤εℒ\(fθ\(la\)\(𝐇la\(x\)\+𝜹\),y\),\\boldsymbol\{\\delta\}^\{\*\}=\\operatorname\*\{arg\\,max\}\_\{\\\|\\boldsymbol\{\\delta\}\\\|\_\{2\}\\leq\\varepsilon\}\\mathcal\{L\}\\\!\\left\(f\_\{\\theta\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(\\mathbf\{H\}\_\{l\_\{\\mathrm\{a\}\}\}\(x\)\+\\boldsymbol\{\\delta\}\),\\,y\\right\),wherefθ\(la\)f\_\{\\theta\}^\{\(l\_\{\\mathrm\{a\}\}\)\}denotes the remainder of the model after injecting the perturbed activation at layerlal\_\{\\mathrm\{a\}\}\. We solve this inner problem with projected gradient descent \(PGD\) using random initialization:
𝜹←Π‖𝜹‖2≤ε\(𝜹\+α∇𝜹ℒadv‖∇𝜹ℒadv‖2\)\.\\boldsymbol\{\\delta\}\\leftarrow\\Pi\_\{\\\|\\boldsymbol\{\\delta\}\\\|\_\{2\}\\leq\\varepsilon\}\\left\(\\boldsymbol\{\\delta\}\+\\alpha\\frac\{\\nabla\_\{\\boldsymbol\{\\delta\}\}\\mathcal\{L\}\_\{\\mathrm\{adv\}\}\}\{\\\|\\nabla\_\{\\boldsymbol\{\\delta\}\}\\mathcal\{L\}\_\{\\mathrm\{adv\}\}\\\|\_\{2\}\}\\right\)\.The outer objective is
minθℒclean\+λℒadv\.\\min\_\{\\theta\}\\;\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\+\\lambda\\,\\mathcal\{L\}\_\{\\mathrm\{adv\}\}\.
In our LAT baseline, we fine\-tune all model parameters and restrict the latent perturbation to a fixed subset of positions: the final 20 non\-padding tokens at the selected transformer layer, matching the suffix\-focused attack setting used in our discrete attacks\. For the final IMDB runs, we usek=8k=8PGD steps,α=ε/5\\alpha=\\varepsilon/5, andλ=1\.0\\lambda=1\.0\. The learning rate is2×10−52\\times 10^\{\-5\}for Pythia\-1\.4B and Qwen\-2\.5\-3B, and1×10−51\\times 10^\{\-5\}for Llama\-3\.1\-8B\. We report the best layer–radius setting from the LAT sweep: layer 4 withε=1\.0\\varepsilon=1\.0for Llama\-3\.1\-8B, layer 8 withε=1\.5\\varepsilon=1\.5for Qwen\-2\.5\-3B, and layer 4 withε=0\.5\\varepsilon=0\.5for Pythia\-1\.4B\.
### A\.5Adversarial Attacks Setup
We evaluate robustness using three adversarial attacks under different suffix\-based search strategies\. All attacks are evaluated on 100 samples drawn from the designated attack split of the dataset\. We report clean accuracy \(ACC\) and Attack Success Rate \(ASR\)\.
#### GCG \(Greedy Coordinate Gradient\)\.
We use token\-levelGCG\[Zouet al\.,[2023b](https://arxiv.org/html/2607.28959#bib.bib1)\]in suffix mode:N=10N\{=\}10adversarial tokens are appended to the original input\. Each round computes the gradient of the cross\-entropy loss with respect to the input embeddings, selects the top\-k=256k\{=\}256candidate replacement tokens per attack position, randomly samples128128of these candidates for forward\-pass evaluation, and applies the single token substitution that maximises the loss\. The attack runs forT=10T\{=\}10rounds\.
#### RandomToken\.
RandomTokenis a gradient\-free baseline designed to be comparable toGCGin the number of adversarial tokens while replacing gradient\-guided search with uniform random sampling, following Algorithm 2 ofHoweet al\.\[[2024](https://arxiv.org/html/2607.28959#bib.bib7)\]\. At each iteration, allN=10N\{=\}10adversarial suffix tokens are replaced simultaneously by tokens drawn uniformly at random from the non\-special vocabulary\. The attack evaluates the resulting prompt with one forward pass and stops immediately upon a successful flip; otherwise it retains the best\-loss candidate seen across all iterations\. Given that an iteration ofRandomTokenis much cheaper than an iteration ofGCG, we useT=500T=500iterations\.
#### Threat model scope\.
Our main evaluation focuses on token\-level suffix attacks, which are directly aligned with both the adversarial training setting studied in this paper and the perturbation model used in our theoretical analysis\. We do not include semantic rewriting attacks such as PAIR\[Chaoet al\.,[2025](https://arxiv.org/html/2607.28959#bib.bib53)\]in the main evaluation, since they operate under a broader threat model: they use an external attacker model to iteratively rewrite the input itself\. In our classification setting, this would allow the attacker to freely reformulate the original example, making the attack no longer comparable to suffix\-based perturbations\.
### A\.6Computation Cost
Experiments were run on GPU clusters using NVIDIA H200, L40S, and RTX A6000 GPUs\. Most jobs used a single GPU worker; CPU workers were used for data loading, preprocessing, and job orchestration\. The exact GPU type varied across runs due to shared\-cluster scheduling\. For the main experiments, we report method\-level compute using FLOPs in Table[1](https://arxiv.org/html/2607.28959#S6.T1)\. Overall, we estimate that the full project required on the order of several hundred GPU\-hours, approximately 800 GPU\-hours in total\.
## Appendix BAdditional Results
### B\.1More Results
Table 5:Additional cross\-dataset robustness results on Pythia\-1\.4B\. We report clean accuracy \(Acc, higher is better\) and attack success rate \(ASR, lower is better\) under RandomToken and GCG suffix attacks\.
### B\.2Additional Surrogate Analysis
Table 6:Clean accuracy and GCG attack success rate at different pruning ratios of the surrogate MLP neurons on Pythia\-1\.4B\. The surrogate is constructed by pruning neurons ranked by the proposed activation\-gradient importance score\.0%0\\%corresponds to the full unpruned model\. Lower GCG ASR is better; higher clean accuracy is better\.Table[6](https://arxiv.org/html/2607.28959#A2.T6)shows that the optimal pruning ratio can vary across datasets\. Moderate pruning sometimes improves transferability over the full surrogate, as in EnronSpam, suggesting that pruning may remove task\-irrelevant or noisy neurons and produce a cleaner adversarial search direction\. This is consistent with the circuit\-sparsity motivation behind our surrogate construction\. At the same time, aggressive pruning substantially weakens transfer, indicating that too much pruning destroys attack\-relevant geometry\.
### B\.3Prefix\-attack result
The suffix\-window defense is motivated by the classifier readout structure rather than by suffix attacks alone\. In decoder\-only sequence classification, the final prediction is primarily read from the last hidden state\. We therefore expect defending a suffix window near this readout position to remain useful even against prefix attacks\.
To verify this, we additionally evaluate prefix\-GCG on Pythia\-1\.4B\. Our method reduces the prefix\-attack ASR to0\.240\.24at the same strength as[A\.5](https://arxiv.org/html/2607.28959#A1.SS5)\.
## Appendix CCircuit\-Guided Surrogate Details
### C\.1ActDiff Method\.
ActDiff is motivated by the difference\-in\-means style of representation analysis widely used in mechanistic interpretability and steering\-vector methods\. Concretely, it ranks neurons by the absolute class\-conditional gap in mean activation magnitude:
sl,jActDiff=\|1\|𝒟\+\|∑x∈𝒟\+∑t\|at,j\(l\)\(x\)\|−1\|𝒟−\|∑x∈𝒟−∑t\|at,j\(l\)\(x\)\|\|,s^\{\\mathrm\{ActDiff\}\}\_\{l,j\}=\\left\|\\frac\{1\}\{\|\\mathcal\{D\}\_\{\+\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{\+\}\}\\sum\_\{t\}\\left\|a^\{\(l\)\}\_\{t,j\}\(x\)\\right\|\-\\frac\{1\}\{\|\\mathcal\{D\}\_\{\-\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{\-\}\}\\sum\_\{t\}\\left\|a^\{\(l\)\}\_\{t,j\}\(x\)\\right\|\\right\|,\(8\)where𝒟\+,𝒟−\\mathcal\{D\}\_\{\+\},\\mathcal\{D\}\_\{\-\}are the positive and negative calibration subsets\. Unlike ActGrad, however, this rule depends only on clean activation statistics and does not measure how much a neuron actually affects the attack objective\.
### C\.2Surrogate pruning heatmaps
We further visualize the layerwise pruning patterns for two representative datasets\. These heatmaps illustrate how the pruning MLP neurons are distributed across layers for different model families\.
Figure 4:Heatmaps of layerwise surrogate pruning patterns on EnronSpam\. 75% means pruning 25% of MLP\.Figure 5:Heatmaps of layerwise surrogate pruning patterns on PasswordMatch\.
## Appendix DAdditional Theory and Proofs
### D\.1Complete Mathematical Setup
We provide the full attention\-level setup used in the theoretical analysis\. As in the main text, we analyze a single attention head at a fixed layerl=lr\+1l=l\_\{\\mathrm\{r\}\}\+1, since the multi\-head case only changes constants\. For the final token at positionTT, define
𝐪Tl=𝐖Ql𝐡Tlr,𝐤tl=𝐖Kl𝐡tlr,𝐯tl=𝐖Vl𝐡tlr,\\mathbf\{q\}\_\{T\}^\{l\}=\\mathbf\{W\}\_\{Q\}^\{l\}\\mathbf\{h\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\},\\quad\\mathbf\{k\}\_\{t\}^\{l\}=\\mathbf\{W\}\_\{K\}^\{l\}\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\},\\quad\\mathbf\{v\}\_\{t\}^\{l\}=\\mathbf\{W\}\_\{V\}^\{l\}\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\},and let the corresponding attention output be
𝐨Tl=∑t=1TαT,tl𝐯tl,αT,tl=exp\(\(𝐪Tl\)⊤𝐤tl/d\)∑j=1Texp\(\(𝐪Tl\)⊤𝐤jl/d\)\.\\mathbf\{o\}\_\{T\}^\{l\}=\\sum\_\{t=1\}^\{T\}\\alpha\_\{T,t\}^\{l\}\\mathbf\{v\}\_\{t\}^\{l\},\\quad\\alpha\_\{T,t\}^\{l\}=\\frac\{\\exp\\\!\\left\(\(\\mathbf\{q\}\_\{T\}^\{l\}\)^\{\\top\}\\mathbf\{k\}\_\{t\}^\{l\}/\\sqrt\{d\}\\right\)\}\{\\sum\_\{j=1\}^\{T\}\\exp\\\!\\left\(\(\\mathbf\{q\}\_\{T\}^\{l\}\)^\{\\top\}\\mathbf\{k\}\_\{j\}^\{l\}/\\sqrt\{d\}\\right\)\}\.
For completeness, we also restate the ReFT operator used throughout the proofs\. Let𝐑∈ℝr×d\\mathbf\{R\}\\in\\mathbb\{R\}^\{r\\times d\}be row\-orthonormal, i\.e\.,
𝐑𝐑⊤=𝐈r,\\mathbf\{R\}\\mathbf\{R\}^\{\\top\}=\\mathbf\{I\}\_\{r\},and let𝐖∈ℝr×d\\mathbf\{W\}\\in\\mathbb\{R\}^\{r\\times d\}and𝐛∈ℝr\\mathbf\{b\}\\in\\mathbb\{R\}^\{r\}\. For a hidden state𝐡∈ℝd\\mathbf\{h\}\\in\\mathbb\{R\}^\{d\}, the ReFT intervention is
Φ𝐑\(𝐡\)=𝐡\+𝐑⊤\(𝐖𝐡\+𝐛−𝐑𝐡\)=\(𝐈−𝐑⊤𝐑\)𝐡\+𝐑⊤\(𝐖𝐡\+𝐛\)\.\\Phi\_\{\\mathbf\{R\}\}\(\\mathbf\{h\}\)=\\mathbf\{h\}\+\\mathbf\{R\}^\{\\top\}\(\\mathbf\{W\}\\mathbf\{h\}\+\\mathbf\{b\}\-\\mathbf\{R\}\\mathbf\{h\}\)=\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\\top\}\\mathbf\{R\}\)\\mathbf\{h\}\+\\mathbf\{R\}^\{\\top\}\(\\mathbf\{W\}\\mathbf\{h\}\+\\mathbf\{b\}\)\.Since𝐑\\mathbf\{R\}has orthonormal rows,𝐑⊤𝐑\\mathbf\{R\}^\{\\top\}\\mathbf\{R\}is the orthogonal projector onto the edited subspacespan\(𝐑\)\\mathrm\{span\}\(\\mathbf\{R\}\), and
Πspan\(𝐑\)⟂:=𝐈−𝐑⊤𝐑\\Pi\_\{\\mathrm\{span\}\(\\mathbf\{R\}\)^\{\\perp\}\}:=\\mathbf\{I\}\-\\mathbf\{R\}^\{\\top\}\\mathbf\{R\}is the orthogonal projector onto its orthogonal complement\.
### D\.2Proof of Theorem[1](https://arxiv.org/html/2607.28959#Thmtheorem1)
###### Proof\.
We analyze the attention layerl=lr\+1l=l\_\{\\mathrm\{r\}\}\+1immediately after the defense layer\. Let\(𝐪Tl,𝐤tl,𝐯tl,𝐨Tl\)\(\\mathbf\{q\}\_\{T\}^\{l\},\\mathbf\{k\}\_\{t\}^\{l\},\\mathbf\{v\}\_\{t\}^\{l\},\\mathbf\{o\}\_\{T\}^\{l\}\)and\(𝐪^Tl,𝐤^tl,𝐯^tl,𝐨^Tl\)\(\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\},\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\},\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\},\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\)denote the clean and defended query, key, value, and attention output, respectively\.
For attacked suffix positionst∈𝒯atk∖\{T\}t\\in\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\\setminus\\\{T\\\}, no correction is applied, so
𝐡^tlr=𝐡tlr\+Δt,\\hat\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}=\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\+\\Delta\_\{t\},which implies
𝐤^tl=𝐤tl\+𝐖KlΔt,𝐯^tl=𝐯tl\+𝐖VlΔt\.\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}=\\mathbf\{k\}\_\{t\}^\{l\}\+\\mathbf\{W\}\_\{K\}^\{l\}\\Delta\_\{t\},\\quad\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}=\\mathbf\{v\}\_\{t\}^\{l\}\+\\mathbf\{W\}\_\{V\}^\{l\}\\Delta\_\{t\}\.At the defended tokenTT, the ReFT correction gives
‖𝐡^Tlr−𝐡Tlr‖≤εT,\\bigl\\\|\\hat\{\\mathbf\{h\}\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\}\-\\mathbf\{h\}\_\{T\}^\{\\,l\_\{\\mathrm\{r\}\}\}\\bigr\\\|\\leq\\varepsilon\_\{T\},and therefore
‖𝐪^Tl−𝐪Tl‖≤‖𝐖Ql‖εT,‖𝐤^Tl−𝐤Tl‖≤‖𝐖Kl‖εT,‖𝐯^Tl−𝐯Tl‖≤‖𝐖Vl‖εT\.\\\|\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\}\-\\mathbf\{q\}\_\{T\}^\{l\}\\\|\\leq\\\|\\mathbf\{W\}\_\{Q\}^\{l\}\\\|\\varepsilon\_\{T\},\\quad\\\|\\hat\{\\mathbf\{k\}\}\_\{T\}^\{l\}\-\\mathbf\{k\}\_\{T\}^\{l\}\\\|\\leq\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\varepsilon\_\{T\},\\quad\\\|\\hat\{\\mathbf\{v\}\}\_\{T\}^\{l\}\-\\mathbf\{v\}\_\{T\}^\{l\}\\\|\\leq\\\|\\mathbf\{W\}\_\{V\}^\{l\}\\\|\\varepsilon\_\{T\}\.
We decompose the output difference as
𝐨^Tl−𝐨Tl\\displaystyle\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}=∑t=1TαT,tl\(𝐯^tl−𝐯tl\)\+∑t=1T\(α^T,tl−αT,tl\)𝐯^tl\\displaystyle=\\sum\_\{t=1\}^\{T\}\\alpha\_\{T,t\}^\{l\}\(\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\-\\mathbf\{v\}\_\{t\}^\{l\}\)\+\\sum\_\{t=1\}^\{T\}\(\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\)\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}=∑t=T−k\+1T−1αT,tl𝐖VlΔt\+αT,Tl\(𝐯^Tl−𝐯Tl\)\+∑t=1T\(α^T,tl−αT,tl\)𝐯^tl\.\\displaystyle=\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\alpha\_\{T,t\}^\{l\}\\mathbf\{W\}\_\{V\}^\{l\}\\Delta\_\{t\}\+\\alpha\_\{T,T\}^\{l\}\(\\hat\{\\mathbf\{v\}\}\_\{T\}^\{l\}\-\\mathbf\{v\}\_\{T\}^\{l\}\)\+\\sum\_\{t=1\}^\{T\}\(\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\)\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\.Applying the reverse triangle inequality gives
‖𝐨^Tl−𝐨Tl‖\\displaystyle\\bigl\\\|\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}\\bigr\\\|≥‖∑t=T−k\+1T−1αT,tl𝐖VlΔt‖−αT,Tl‖𝐯^Tl−𝐯Tl‖−‖∑t=1T\(α^T,tl−αT,tl\)𝐯^tl‖\\displaystyle\\geq\\left\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\alpha\_\{T,t\}^\{l\}\\mathbf\{W\}\_\{V\}^\{l\}\\Delta\_\{t\}\\right\\\|\-\\alpha\_\{T,T\}^\{l\}\\\|\\hat\{\\mathbf\{v\}\}\_\{T\}^\{l\}\-\\mathbf\{v\}\_\{T\}^\{l\}\\\|\-\\left\\\|\\sum\_\{t=1\}^\{T\}\(\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\)\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\right\\\|≥‖∑t=T−k\+1T−1αT,tl𝐖VlΔt‖−αT,Tl‖𝐖Vl‖εT−∑t=1T\|α^T,tl−αT,tl\|⋅‖𝐯^tl‖\.\\displaystyle\\geq\\left\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\alpha\_\{T,t\}^\{l\}\\mathbf\{W\}\_\{V\}^\{l\}\\Delta\_\{t\}\\right\\\|\-\\alpha\_\{T,T\}^\{l\}\\\|\\mathbf\{W\}\_\{V\}^\{l\}\\\|\\,\\varepsilon\_\{T\}\-\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|\\cdot\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\.This proves \([5](https://arxiv.org/html/2607.28959#S5.E5)\) with
ℰT=αT,Tl‖𝐖Vl‖εT,andℰattn=∑t=1T\|α^T,tl−αT,tl\|⋅‖𝐯^tl‖\.\\mathcal\{E\}\_\{T\}=\\alpha\_\{T,T\}^\{l\}\\\|\\mathbf\{W\}\_\{V\}^\{l\}\\\|\\,\\varepsilon\_\{T\},\\quad\\text\{and\}\\quad\\mathcal\{E\}\_\{\\mathrm\{attn\}\}=\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|\\cdot\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\.
It remains to boundℰattn\\mathcal\{E\}\_\{\\mathrm\{attn\}\}\. The attention weights depend on the query𝐪Tl\\mathbf\{q\}\_\{T\}^\{l\}and the key collection\{𝐤tl\}t=1T\\\{\\mathbf\{k\}\_\{t\}^\{l\}\\\}\_\{t=1\}^\{T\}through the softmax map\. Under Assumption[1](https://arxiv.org/html/2607.28959#Thmassumption1), there exists a local Lipschitz constantLα\>0L\_\{\\alpha\}\>0such that
∑t=1T\|α^T,tl−αT,tl\|≤Lα\(‖𝐪^Tl−𝐪Tl‖\+∑t=1T‖𝐤^tl−𝐤tl‖\)\.\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|\\leq L\_\{\\alpha\}\\left\(\\\|\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\}\-\\mathbf\{q\}\_\{T\}^\{l\}\\\|\+\\sum\_\{t=1\}^\{T\}\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\right\)\.Using the bounds established above,
‖𝐪^Tl−𝐪Tl‖≤‖𝐖Ql‖εT,\\\|\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\}\-\\mathbf\{q\}\_\{T\}^\{l\}\\\|\\leq\\\|\\mathbf\{W\}\_\{Q\}^\{l\}\\\|\\,\\varepsilon\_\{T\},and
∑t=1T‖𝐤^tl−𝐤tl‖≤‖𝐖Kl‖εT\+‖𝐖Kl‖∑t=T−k\+1T−1‖Δt‖\.\\sum\_\{t=1\}^\{T\}\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\leq\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\,\\varepsilon\_\{T\}\+\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\\|\\Delta\_\{t\}\\\|\.Therefore,
∑t=1T\|α^T,tl−αT,tl\|≤Lα\(‖𝐖Ql‖εT\+‖𝐖Kl‖εT\+‖𝐖Kl‖∑t=T−k\+1T−1‖Δt‖\)\.\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|\\leq L\_\{\\alpha\}\\left\(\\\|\\mathbf\{W\}\_\{Q\}^\{l\}\\\|\\,\\varepsilon\_\{T\}\+\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\,\\varepsilon\_\{T\}\+\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\\|\\Delta\_\{t\}\\\|\\right\)\.
If the defended trajectory remains in a bounded neighborhood of the clean one, then there exists a constantMV\>0M\_\{V\}\>0such that
sup1≤t≤T‖𝐯^tl‖≤MV\.\\sup\_\{1\\leq t\\leq T\}\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\\leq M\_\{V\}\.Hence
ℰattn\\displaystyle\\mathcal\{E\}\_\{\\mathrm\{attn\}\}≤MV∑t=1T\|α^T,tl−αT,tl\|\\displaystyle\\leq M\_\{V\}\\sum\_\{t=1\}^\{T\}\\bigl\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\\bigr\|≤MVLα\(\(‖𝐖Ql‖\+‖𝐖Kl‖\)εT\+‖𝐖Kl‖∑t=T−k\+1T−1‖Δt‖\)\.\\displaystyle\\leq M\_\{V\}L\_\{\\alpha\}\\left\(\(\\\|\\mathbf\{W\}\_\{Q\}^\{l\}\\\|\+\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\)\\varepsilon\_\{T\}\+\\\|\\mathbf\{W\}\_\{K\}^\{l\}\\\|\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\\|\\Delta\_\{t\}\\\|\\right\)\.Thus there exists a constantCattn\>0C\_\{\\mathrm\{attn\}\}\>0such that
ℰattn≤Cattn\(εT\+∑t=T−k\+1T−1‖Δt‖\),\\mathcal\{E\}\_\{\\mathrm\{attn\}\}\\leq C\_\{\\mathrm\{attn\}\}\\left\(\\varepsilon\_\{T\}\+\\sum\_\{t=T\-k\+1\}^\{T\-1\}\\\|\\Delta\_\{t\}\\\|\\right\),which proves the theorem\.
∎
### D\.3Proof of Theorem[2](https://arxiv.org/html/2607.28959#Thmtheorem2)
###### Proof\.
Becausek≤Lk\\leq L, the attack window is contained in the defended suffix window, i\.e\.,
𝒯atk⊆𝒯LAT\.\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\\subseteq\\mathcal\{T\}\_\{\\mathrm\{LAT\}\}\.Hence, for every attacked positiont∈𝒯atkt\\in\\mathcal\{T\}\_\{\\mathrm\{atk\}\}, the sequence\-shared operator satisfies
‖Φshared\(𝐡~tlr\)−𝐡tlr‖≤εLAT\.\\\|\\Phi\_\{\\mathrm\{shared\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\)\-\\mathbf\{h\}\_\{t\}^\{\\,l\_\{\\mathrm\{r\}\}\}\\\|\\leq\\varepsilon\_\{\\mathrm\{LAT\}\}\.Therefore, the corresponding query, key, and value perturbations are bounded by
‖𝐪^Tl−𝐪Tl‖≤BQεLAT,‖𝐤^tl−𝐤tl‖≤BKεLAT,‖𝐯^tl−𝐯tl‖≤BVεLAT\.\\\|\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\}\-\\mathbf\{q\}\_\{T\}^\{l\}\\\|\\leq B\_\{Q\}\\varepsilon\_\{\\mathrm\{LAT\}\},\\quad\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\leq B\_\{K\}\\varepsilon\_\{\\mathrm\{LAT\}\},\\quad\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\-\\mathbf\{v\}\_\{t\}^\{l\}\\\|\\leq B\_\{V\}\\varepsilon\_\{\\mathrm\{LAT\}\}\.
We decompose the attention distortion into a value term and an attention\-weight term:
‖𝐨^Tl−𝐨Tl‖≤∑t=1TαT,tl‖𝐯^tl−𝐯tl‖\+∑t=1T\|α^T,tl−αT,tl\|‖𝐯^tl‖\.\\\|\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}\\\|\\leq\\sum\_\{t=1\}^\{T\}\\alpha\_\{T,t\}^\{l\}\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\-\\mathbf\{v\}\_\{t\}^\{l\}\\\|\+\\sum\_\{t=1\}^\{T\}\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\|\\,\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\.
For the value term, only attacked positions contribute, and∑tαT,tl=1\\sum\_\{t\}\\alpha\_\{T,t\}^\{l\}=1, so
∑t=1TαT,tl‖𝐯^tl−𝐯tl‖≤BVεLAT\.\\sum\_\{t=1\}^\{T\}\\alpha\_\{T,t\}^\{l\}\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\-\\mathbf\{v\}\_\{t\}^\{l\}\\\|\\leq B\_\{V\}\\varepsilon\_\{\\mathrm\{LAT\}\}\.
For the attention\-weight term, Assumption[1](https://arxiv.org/html/2607.28959#Thmassumption1)gives
∑t=1T\|α^T,tl−αT,tl\|≤Lα\(‖𝐪^Tl−𝐪Tl‖\+∑t=1T‖𝐤^tl−𝐤tl‖\)\.\\sum\_\{t=1\}^\{T\}\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\|\\leq L\_\{\\alpha\}\\left\(\\\|\\hat\{\\mathbf\{q\}\}\_\{T\}^\{l\}\-\\mathbf\{q\}\_\{T\}^\{l\}\\\|\+\\sum\_\{t=1\}^\{T\}\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\right\)\.Only the attacked positions contribute to the key perturbation\. Fort∉𝒯atkt\\notin\\mathcal\{T\}\_\{\\mathrm\{atk\}\}, we have𝐤^tl=𝐤tl\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}=\\mathbf\{k\}\_\{t\}^\{l\}\. Fort∈𝒯atkt\\in\\mathcal\{T\}\_\{\\mathrm\{atk\}\}, the sequence\-shared defense ensures
‖𝐤^tl−𝐤tl‖≤BKεLAT\.\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\leq B\_\{K\}\\varepsilon\_\{\\mathrm\{LAT\}\}\.Since\|𝒯atk\|=k\|\\mathcal\{T\}\_\{\\mathrm\{atk\}\}\|=k, it follows that
∑t=1T‖𝐤^tl−𝐤tl‖≤kBKεLAT\.\\sum\_\{t=1\}^\{T\}\\\|\\hat\{\\mathbf\{k\}\}\_\{t\}^\{l\}\-\\mathbf\{k\}\_\{t\}^\{l\}\\\|\\leq kB\_\{K\}\\varepsilon\_\{\\mathrm\{LAT\}\}\.Therefore,
∑t=1T\|α^T,tl−αT,tl\|≤LαεLAT\(BQ\+kBK\)\.\\sum\_\{t=1\}^\{T\}\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\|\\leq L\_\{\\alpha\}\\varepsilon\_\{\\mathrm\{LAT\}\}\(B\_\{Q\}\+kB\_\{K\}\)\.Assuming‖𝐯^tl‖≤MV\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\\leq M\_\{V\}, we obtain
∑t=1T\|α^T,tl−αT,tl\|‖𝐯^tl‖≤MVLαεLAT\(BQ\+kBK\)\.\\sum\_\{t=1\}^\{T\}\|\\hat\{\\alpha\}\_\{T,t\}^\{l\}\-\\alpha\_\{T,t\}^\{l\}\|\\,\\\|\\hat\{\\mathbf\{v\}\}\_\{t\}^\{l\}\\\|\\leq M\_\{V\}L\_\{\\alpha\}\\varepsilon\_\{\\mathrm\{LAT\}\}\(B\_\{Q\}\+kB\_\{K\}\)\.
Combining the two bounds yields
‖𝐨^Tl−𝐨Tl‖≤εLAT\(BV\+MVLα\(BQ\+kBK\)\),\\\|\\hat\{\\mathbf\{o\}\}\_\{T\}^\{l\}\-\\mathbf\{o\}\_\{T\}^\{l\}\\\|\\leq\\varepsilon\_\{\\mathrm\{LAT\}\}\\bigl\(B\_\{V\}\+M\_\{V\}L\_\{\\alpha\}\(B\_\{Q\}\+kB\_\{K\}\)\\bigr\),which proves the theorem\. ∎
### D\.4Proof of Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3)
###### Lemma 1\.
To simplify notation, write the MLP expansion layer atlal\_\{\\mathrm\{a\}\}asZ=ϕ\(HattnW1\),Z=\\phi\(H^\{\\mathrm\{attn\}\}W\_\{1\}\),where
Hattn=σ\(RowSoftmax\(𝐇~WQK𝐇~⊤dm\)𝐇~WV\),WQK=WQWK⊤,H^\{\\mathrm\{attn\}\}=\\sigma\\\!\\left\(\\operatorname\{RowSoftmax\}\\\!\\left\(\\frac\{\\widetilde\{\\mathbf\{H\}\}W\_\{QK\}\\widetilde\{\\mathbf\{H\}\}^\{\\top\}\}\{\\sqrt\{d\_\{m\}\}\}\\right\)\\widetilde\{\\mathbf\{H\}\}W\_\{V\}\\right\),\\quad W\_\{QK\}=W\_\{Q\}W\_\{K\}^\{\\top\},andσ\\sigmaandϕ\\phiareLσL\_\{\\sigma\}\- andLϕL\_\{\\phi\}\-Lipschitz, respectively\. Assume that
‖WQ‖≤BQ,‖WK‖≤BK,‖WV‖≤BV\.\\\|W\_\{Q\}\\\|\\leq B\_\{Q\},\\quad\\\|W\_\{K\}\\\|\\leq B\_\{K\},\\quad\\\|W\_\{V\}\\\|\\leq B\_\{V\}\.Then there existsCattn=O\(T\)C\_\{\\mathrm\{attn\}\}=O\(T\)such that
‖∂Z∂𝐇~‖≤LϕLσ‖W1‖\(CattnBQBKBVdm\+BV\)\.\\left\\\|\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\leq L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\.
###### Proof of Lemma[1](https://arxiv.org/html/2607.28959#Thmlemma1)\.
By the chain rule,
∂Z∂𝐇~=∂Z∂Hattn∂Hattn∂𝐇~\.\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}=\\frac\{\\partial Z\}\{\\partial H^\{\\mathrm\{attn\}\}\}\\frac\{\\partial H^\{\\mathrm\{attn\}\}\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\.Sinceϕ\\phiisLϕL\_\{\\phi\}\-Lipschitz,
‖∂Z∂Hattn‖≤Lϕ‖W1‖\.\\left\\\|\\frac\{\\partial Z\}\{\\partial H^\{\\mathrm\{attn\}\}\}\\right\\\|\\leq L\_\{\\phi\}\\\|W\_\{1\}\\\|\.
LetA\(𝐇~\)=RowSoftmax\(𝐇~WQK𝐇~⊤dm\),A\(\\widetilde\{\\mathbf\{H\}\}\)=\\operatorname\{RowSoftmax\}\\\!\\left\(\\frac\{\\widetilde\{\\mathbf\{H\}\}W\_\{QK\}\\widetilde\{\\mathbf\{H\}\}^\{\\top\}\}\{\\sqrt\{d\_\{m\}\}\}\\right\),so thatHattn=σ\(A\(𝐇~\)𝐇~WV\)\.H^\{\\mathrm\{attn\}\}=\\sigma\\\!\\left\(A\(\\widetilde\{\\mathbf\{H\}\}\)\\,\\widetilde\{\\mathbf\{H\}\}W\_\{V\}\\right\)\.Sinceσ\\sigmaisLσL\_\{\\sigma\}\-Lipschitz,
‖∂Hattn∂𝐇~‖≤Lσ‖∂\(A\(𝐇~\)𝐇~WV\)∂𝐇~‖\.\\left\\\|\\frac\{\\partial H^\{\\mathrm\{attn\}\}\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\leq L\_\{\\sigma\}\\left\\\|\\frac\{\\partial\\bigl\(A\(\\widetilde\{\\mathbf\{H\}\}\)\\widetilde\{\\mathbf\{H\}\}W\_\{V\}\\bigr\)\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\.
For a single\-row softmaxp=softmax\(z\)p=\\mathrm\{softmax\}\(z\), its Jacobian is
Dsoftmax\(z\)=diag\(p\)−pp⊤,D\\,\\mathrm\{softmax\}\(z\)=\\operatorname\{diag\}\(p\)\-pp^\{\\top\},and satisfies the uniform spectral\-norm bound‖Dsoftmax\(z\)‖≤12\\\|D\\,\\mathrm\{softmax\}\(z\)\\\|\\leq\\tfrac\{1\}\{2\}\[Nair,[2025](https://arxiv.org/html/2607.28959#bib.bib64)\]\. Therefore the RowSoftmax map is uniformly Lipschitz row\-wise\. Then the derivative ofA\(𝐇~\)𝐇~WVA\(\\widetilde\{\\mathbf\{H\}\}\)\\,\\widetilde\{\\mathbf\{H\}\}W\_\{V\}with respect to𝐇~\\widetilde\{\\mathbf\{H\}\}consists of a value\-path term and an attention\-weight term\. The value\-path term is bounded byBVB\_\{V\}\. For the attention\-weight term, the softmax Jacobian bound together with
‖WQK‖≤‖WQ‖‖WK‖≤BQBK,‖WV‖≤BV,\\\|W\_\{QK\}\\\|\\leq\\\|W\_\{Q\}\\\|\\,\\\|W\_\{K\}\\\|\\leq B\_\{Q\}B\_\{K\},\\quad\\\|W\_\{V\}\\\|\\leq B\_\{V\},gives a contribution proportional to‖WQK‖‖WV‖dm,\\frac\{\\\|W\_\{QK\}\\\|\\\|W\_\{V\}\\\|\}\{\\sqrt\{d\_\{m\}\}\},multiplied by a sequence\-length accumulation factor\. Absorbing this factor intoCattn=O\(T\)C\_\{\\mathrm\{attn\}\}=O\(T\), we obtain
‖∂Hattn∂𝐇~‖≤Lσ\(CattnBQBKBVdm\+BV\)\.\\left\\\|\\frac\{\\partial H^\{\\mathrm\{attn\}\}\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\leq L\_\{\\sigma\}\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\.
Combining the above displays gives
‖∂Z∂𝐇~‖≤LϕLσ‖W1‖\(CattnBQBKBVdm\+BV\)\.\\left\\\|\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\leq L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\.∎
###### Proof of Theorem[3](https://arxiv.org/html/2607.28959#Thmtheorem3)\.
We work on the following high\-probability event\.
###### Assumption 2\(Non\-degeneracy on pruned active coordinates\)\.
There existm0\>0m\_\{0\}\>0andp∈\(0,1\)p\\in\(0,1\)such that, with probability at least1−p1\-p,
minj∉S\|at,j\(la\)\(x\)\|≥m0,\\min\_\{j\\notin S\}\\left\|a\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(x\)\\right\|\\geq m\_\{0\},for allt∈\[K\]t\\in\[K\]\.
###### Assumption 3\(Local sign\-stability\)\.
There exists a local constantc\>0c\>0such that, on the same event,
⟨gt,ut⟩−⟨gt,uts⟩≤c‖gt−gts‖,\\langle g\_\{t\},u\_\{t\}\\rangle\-\\langle g\_\{t\},u\_\{t\}^\{s\}\\rangle\\leq c\\,\\\|g\_\{t\}\-g\_\{t\}^\{s\}\\\|,for allt∈\[K\]t\\in\[K\], where
gt=∇𝐇~Lf\(𝐇~\(t\)\),gts=∇𝐇~Ls\(𝐇~\(t\)\),g\_\{t\}=\\nabla\_\{\\widetilde\{\\mathbf\{H\}\}\}L\_\{f\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\\right\),\\quad g\_\{t\}^\{s\}=\\nabla\_\{\\widetilde\{\\mathbf\{H\}\}\}L\_\{s\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\\right\),andut:=sign\(gt\),uts:=sign\(gts\)u\_\{t\}:=\\operatorname\{sign\}\(g\_\{t\}\),u\_\{t\}^\{s\}:=\\operatorname\{sign\}\(g\_\{t\}^\{s\}\)be the one\-stepℓ∞\\ell\_\{\\infty\}\-PGD ascent directions\.
Define
Gt:=Lf\(𝐇~\(t\)\+ηut\)−Lf\(𝐇~\(t\)\+ηuts\)\.G\_\{t\}:=L\_\{f\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\+\\eta u\_\{t\}\\right\)\-L\_\{f\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\+\\eta u\_\{t\}^\{s\}\\right\)\.
LetZZdenote the MLP expansion layer atlal\_\{\\mathrm\{a\}\}, and letrt:=∇ZLf\(𝐇~\(t\)\)\.r\_\{t\}:=\\nabla\_\{Z\}L\_\{f\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\\right\)\.WritePSP\_\{S\}for the projection onto the retained neuron setSS\. Define
st,j:=\|at,j\(la\)\(x\)rt,j\(la\)\|,Mt:=\(∑j∉S\|at,j\(la\)\(x\)rt,j\(la\)\|2\)1/2\.s\_\{t,j\}:=\\left\|a\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(x\)\\,r\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\\right\|,\\quad M\_\{t\}:=\\left\(\\sum\_\{j\\notin S\}\|a\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(x\)\\,r\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\|^\{2\}\\right\)^\{1/2\}\.
By the first\-order expansion ofLfL\_\{f\},
Lf\(𝐇~\(t\)\+ηu\)=Lf\(𝐇~\(t\)\)\+η⟨gt,u⟩\+o\(η\),L\_\{f\}\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\+\\eta u\)=L\_\{f\}\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\)\+\\eta\\langle g\_\{t\},u\\rangle\+o\(\\eta\),hence
Gt=η\(⟨gt,ut⟩−⟨gt,uts⟩\)\+o\(η\)≤cη‖gt−gts‖\+o\(η\)\.G\_\{t\}=\\eta\\bigl\(\\langle g\_\{t\},u\_\{t\}\\rangle\-\\langle g\_\{t\},u\_\{t\}^\{s\}\\rangle\\bigr\)\+o\(\\eta\)\\leq c\\,\\eta\\,\\\|g\_\{t\}\-g\_\{t\}^\{s\}\\\|\+o\(\\eta\)\.
Since the surrogate replaces the full MLP expansionZZby its projected versionZs=PSZ,Z\_\{s\}=P\_\{S\}Z,its Jacobian with respect to the attacked hidden state is correspondingly projected:∂Zs∂𝐇~=PS∂Z∂𝐇~\.\\frac\{\\partial Z\_\{s\}\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}=P\_\{S\}\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\.Let
rt=∇ZLf\(𝐇~\(t\)\),rts=∇ZsLs\(𝐇~\(t\)\)\.r\_\{t\}=\\nabla\_\{Z\}L\_\{f\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\\right\),\\qquad r\_\{t\}^\{s\}=\\nabla\_\{Z\_\{s\}\}L\_\{s\}\\\!\\left\(\\widetilde\{\\mathbf\{H\}\}^\{\(t\)\}\\right\)\.Under the surrogate approximation, we identify the surrogate MLP\-path gradient with the projected full gradient, i\.e\.rts≈PSrt\.r\_\{t\}^\{s\}\\approx P\_\{S\}r\_\{t\}\.Hence
gts=\(∂Zs∂𝐇~\)⊤rts≈\(∂Z∂𝐇~\)⊤PSrt,g\_\{t\}^\{s\}=\\left\(\\frac\{\\partial Z\_\{s\}\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\)^\{\\top\}r\_\{t\}^\{s\}\\approx\\left\(\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\)^\{\\top\}P\_\{S\}r\_\{t\},and therefore
gt−gts≈\(∂Z∂𝐇~\)⊤\(I−PS\)rt\.g\_\{t\}\-g\_\{t\}^\{s\}\\approx\\left\(\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\)^\{\\top\}\(I\-P\_\{S\}\)r\_\{t\}\.Accordingly, we use the bound
‖gt−gts‖≤‖∂Z∂𝐇~‖⋅‖\(I−PS\)rt‖\+εt,\\\|g\_\{t\}\-g\_\{t\}^\{s\}\\\|\\leq\\left\\\|\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\cdot\\\|\(I\-P\_\{S\}\)r\_\{t\}\\\|\+\\varepsilon\_\{t\},whereεt\\varepsilon\_\{t\}collects the surrogate mismatch induced by replacingrtsr\_\{t\}^\{s\}withPSrtP\_\{S\}r\_\{t\}\. In the ideal projected\-gradient case,εt=0\\varepsilon\_\{t\}=0\.
Forj∉Sj\\notin S,
\|rt,j\(la\)\|=st,j\|at,j\(la\)\(x\)\|≤st,jm0,\|r\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\|=\\frac\{s\_\{t,j\}\}\{\|a\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\(x\)\|\}\\leq\\frac\{s\_\{t,j\}\}\{m\_\{0\}\},and thus
‖\(I−PS\)rt‖2=∑j∉S\(rt,j\(la\)\)2≤1m02∑j∉Sst,j2=1m02Mt2\.\\\|\(I\-P\_\{S\}\)r\_\{t\}\\\|^\{2\}=\\sum\_\{j\\notin S\}\\bigl\(r\_\{t,j\}^\{\(l\_\{\\mathrm\{a\}\}\)\}\\bigr\)^\{2\}\\leq\\frac\{1\}\{m\_\{0\}^\{2\}\}\\sum\_\{j\\notin S\}s\_\{t,j\}^\{2\}=\\frac\{1\}\{m\_\{0\}^\{2\}\}M\_\{t\}^\{2\}\.Hence
‖\(I−PS\)rt‖≤1m0Mt\.\\\|\(I\-P\_\{S\}\)r\_\{t\}\\\|\\leq\\frac\{1\}\{m\_\{0\}\}M\_\{t\}\.
Applying Lemma[1](https://arxiv.org/html/2607.28959#Thmlemma1),
‖∂Z∂𝐇~‖≤LϕLσ‖W1‖\(CattnBQBKBVdm\+BV\)\.\\left\\\|\\frac\{\\partial Z\}\{\\partial\\widetilde\{\\mathbf\{H\}\}\}\\right\\\|\\leq L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\.Therefore
Gt≤cLϕLσ‖W1‖m0\(CattnBQBKBVdm\+BV\)ηMt\+o\(η\)\.G\_\{t\}\\leq\\frac\{c\\,L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\}\{m\_\{0\}\}\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\\eta\\,M\_\{t\}\+o\(\\eta\)\.
Summing overt∈\[K\]t\\in\[K\], we obtain
∑t=1KGt≤cLϕLσ‖W1‖m0\(CattnBQBKBVdm\+BV\)η∑t=1KMt\+o\(Kη\)\.\\sum\_\{t=1\}^\{K\}G\_\{t\}\\leq\\frac\{c\\,L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\}\{m\_\{0\}\}\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)\\eta\\sum\_\{t=1\}^\{K\}M\_\{t\}\+o\(K\\eta\)\.
By a union bound, the estimate holds with probability at least1−2Kp1\-2Kp\. Denote
Ccurr=cLϕLσ‖W1‖m0\(CattnBQBKBVdm\+BV\)C\_\{\\mathrm\{curr\}\}=\\frac\{c\\,L\_\{\\phi\}L\_\{\\sigma\}\\\|W\_\{1\}\\\|\}\{m\_\{0\}\}\\left\(C\_\{\\mathrm\{attn\}\}\\frac\{B\_\{Q\}B\_\{K\}B\_\{V\}\}\{\\sqrt\{d\_\{m\}\}\}\+B\_\{V\}\\right\)gives the final result\. ∎
### D\.5A Spatial Transfer Limitation of Last\-Token ReFT
We further explain why a ReFT module trained only at the final token need not transfer optimally to earlier positions\. The key point is that, at an earlier position, the reconstruction error naturally decomposes into two parts: one outside the learned edited subspace, and one inside that subspace\.
###### Proposition 1\(Orthogonal decomposition of transfer error under single\-token LAT\)\.
Let\(𝐑∗,𝐖∗,𝐛∗\)\(\\mathbf\{R\}^\{\*\},\\mathbf\{W\}^\{\*\},\\mathbf\{b\}^\{\*\}\)be obtained by training ReFT only at the final token positionTT:
\(𝐑∗,𝐖∗,𝐛∗\)∈argmin𝐑,𝐖,𝐛𝔼𝐡T\[max‖δ‖≤ε‖Φ𝐑\(𝐡T\+δ\)−𝐡T‖2\]\.\(\\mathbf\{R\}^\{\*\},\\mathbf\{W\}^\{\*\},\\mathbf\{b\}^\{\*\}\)\\in\\arg\\min\_\{\\mathbf\{R\},\\mathbf\{W\},\\mathbf\{b\}\}\\;\\mathbb\{E\}\_\{\\mathbf\{h\}\_\{T\}\}\\left\[\\max\_\{\\\|\\delta\\\|\\leq\\varepsilon\}\\\|\\Phi\_\{\\mathbf\{R\}\}\(\\mathbf\{h\}\_\{T\}\+\\delta\)\-\\mathbf\{h\}\_\{T\}\\\|^\{2\}\\right\]\.For an earlier positiont<Tt<T, let𝐡~t=𝐡t\+Δt\\tilde\{\\mathbf\{h\}\}\_\{t\}=\\mathbf\{h\}\_\{t\}\+\\Delta\_\{t\}\. Then
𝔼\[‖Φ𝐑∗\(𝐡~t\)−𝐡t‖2\]=𝔼\[‖\(𝐈−𝐑∗⊤𝐑∗\)Δt‖2\]\+𝔼\[‖𝐑∗⊤\(𝐖∗𝐡~t\+𝐛∗−𝐑∗𝐡t\)‖2\]\.\\mathbb\{E\}\\Big\[\\\|\\Phi\_\{\\mathbf\{R\}^\{\*\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}\)\-\\mathbf\{h\}\_\{t\}\\\|^\{2\}\\Big\]=\\mathbb\{E\}\\Big\[\\\|\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\)\\Delta\_\{t\}\\\|^\{2\}\\Big\]\+\\mathbb\{E\}\\Big\[\\\|\\mathbf\{R\}^\{\*\\top\}\(\\mathbf\{W\}^\{\*\}\\tilde\{\\mathbf\{h\}\}\_\{t\}\+\\mathbf\{b\}^\{\*\}\-\\mathbf\{R\}^\{\*\}\\mathbf\{h\}\_\{t\}\)\\\|^\{2\}\\Big\]\.
###### Proof\.
By definition of the ReFT operator,
Φ𝐑∗\(𝐡~t\)=\(𝐈−𝐑∗⊤𝐑∗\)\(𝐡t\+Δt\)\+𝐑∗⊤\(𝐖∗𝐡~t\+𝐛∗\)\.\\Phi\_\{\\mathbf\{R\}^\{\*\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}\)=\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\)\(\\mathbf\{h\}\_\{t\}\+\\Delta\_\{t\}\)\+\\mathbf\{R\}^\{\*\\top\}\(\\mathbf\{W\}^\{\*\}\\tilde\{\\mathbf\{h\}\}\_\{t\}\+\\mathbf\{b\}^\{\*\}\)\.Subtracting𝐡t\\mathbf\{h\}\_\{t\}and using
𝐡t=\(𝐈−𝐑∗⊤𝐑∗\)𝐡t\+𝐑∗⊤𝐑∗𝐡t,\\mathbf\{h\}\_\{t\}=\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\)\\mathbf\{h\}\_\{t\}\+\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\\mathbf\{h\}\_\{t\},we obtain
Φ𝐑∗\(𝐡~t\)−𝐡t=\(𝐈−𝐑∗⊤𝐑∗\)Δt\+𝐑∗⊤\(𝐖∗𝐡~t\+𝐛∗−𝐑∗𝐡t\)\.\\Phi\_\{\\mathbf\{R\}^\{\*\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}\)\-\\mathbf\{h\}\_\{t\}=\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\)\\Delta\_\{t\}\+\\mathbf\{R\}^\{\*\\top\}\(\\mathbf\{W\}^\{\*\}\\tilde\{\\mathbf\{h\}\}\_\{t\}\+\\mathbf\{b\}^\{\*\}\-\\mathbf\{R\}^\{\*\}\\mathbf\{h\}\_\{t\}\)\.The first term lies in the orthogonal complement of the edited subspace, while the second lies inside the edited subspace\. Since these two subspaces are orthogonal,
‖Φ𝐑∗\(𝐡~t\)−𝐡t‖2=‖\(𝐈−𝐑∗⊤𝐑∗\)Δt‖2\+‖𝐑∗⊤\(𝐖∗𝐡~t\+𝐛∗−𝐑∗𝐡t\)‖2\.\\\|\\Phi\_\{\\mathbf\{R\}^\{\*\}\}\(\\tilde\{\\mathbf\{h\}\}\_\{t\}\)\-\\mathbf\{h\}\_\{t\}\\\|^\{2\}=\\\|\(\\mathbf\{I\}\-\\mathbf\{R\}^\{\*\\top\}\\mathbf\{R\}^\{\*\}\)\\Delta\_\{t\}\\\|^\{2\}\+\\\|\\mathbf\{R\}^\{\*\\top\}\(\\mathbf\{W\}^\{\*\}\\tilde\{\\mathbf\{h\}\}\_\{t\}\+\\mathbf\{b\}^\{\*\}\-\\mathbf\{R\}^\{\*\}\\mathbf\{h\}\_\{t\}\)\\\|^\{2\}\.Taking expectations yields the stated decomposition\. ∎
The proposition highlights two distinct sources of transfer error when a ReFT module trained only at positionTTis reused at an earlier positiont<Tt<T\. The first term on the right hand side measures perturbation components that lie outside the learned edited subspace and therefore cannot be removed by the reused low\-rank intervention\. The second term measures reconstruction mismatch within the edited subspace itself\. Together, these terms show that even if the final\-token ReFT module is effective at positionTT, its zero\-shot reuse at earlier positions need not remain optimal\. This provides additional support for applying the defense over a suffix window rather than only at the final token\.
## Appendix EEstimated Compute Calculations
We estimate adversarial\-training compute using dense\-equivalent Kaplan\-style FLOPs\. LetNNdenote the dense parameter count of the full model andDDdenote the number of processed tokens\. Following standard scaling\-law approximations\[Kaplanet al\.,[2020](https://arxiv.org/html/2607.28959#bib.bib56)\], we estimate one dense forward pass as
Cfwd≈2ND,C\_\{\\mathrm\{fwd\}\}\\approx 2ND,and one dense forward\-backward training pass with parameter gradient accumulation as
Cfwd\+bwd≈6ND\.C\_\{\\mathrm\{fwd\+bwd\}\}\\approx 6ND\.When a backward pass is used only to obtain gradients with respect to an input or latent perturbation, and no parameter gradients are accumulated, we charge
Cinput\-bwd≈2ND\.C\_\{\\mathrm\{input\\text\{\-\}bwd\}\}\\approx 2ND\.Thus, one forward pass plus one input\-gradient backward pass costs approximately4ND4ND\.
In the main table, we report normalized FLOPs per 512\-token training example\. These estimates are intended to compare training\-time computation across methods\. Unless otherwise stated, we use dense model parameter counts even when the implementation uses quantization or parameter\-efficient wrappers\.
#### R2D2\.
R2D2 first generates an offline adversarial pool and then fine\-tunes the full model on the resulting data\. We therefore decompose the cost into a full\-model fine\-tuning term and an offline search term:
CR2D2=Ctrain\+Csearch\.C\_\{\\mathrm\{R2D2\}\}=C\_\{\\mathrm\{train\}\}\+C\_\{\\mathrm\{search\}\}\.The training term uses full\-model fine\-tuning,
Ctrain≈6NDtrain\.C\_\{\\mathrm\{train\}\}\\approx 6ND\_\{\\mathrm\{train\}\}\.For the offline attack\-generation term, if generating an attacked example requiresnfwdn\_\{\\mathrm\{fwd\}\}forward passes andnbwdn\_\{\\mathrm\{bwd\}\}input\-gradient backward passes overDsearchD\_\{\\mathrm\{search\}\}tokens, we charge
Csearch=\(2nfwd\+2nbwd\)NDsearch,C\_\{\\mathrm\{search\}\}=\(2n\_\{\\mathrm\{fwd\}\}\+2n\_\{\\mathrm\{bwd\}\}\)ND\_\{\\mathrm\{search\}\},summed over all generated adversarial examples\. Since this search is performed offline, we amortize it over normalized 512\-token training examples:
CR2D2step=Ctrain\+CsearchDtrain/512\.C\_\{\\mathrm\{R2D2\\ step\}\}=\\frac\{C\_\{\\mathrm\{train\}\}\+C\_\{\\mathrm\{search\}\}\}\{D\_\{\\mathrm\{train\}\}/512\}\.
#### LAT\.
For LAT, we use the recorded forward/backward pass counts from the training compute summaries\. If one adversarial\-training update usesnfwdn\_\{\\mathrm\{fwd\}\}full\-model forward passes andnbwdn\_\{\\mathrm\{bwd\}\}full\-model backward passes with parameter\-gradient accumulation, the normalized per\-step cost at sequence lengthT=512T=512is
CLAT=\(2nfwd\+4nbwd\)NT\.C\_\{\\mathrm\{LAT\}\}=\(2n\_\{\\mathrm\{fwd\}\}\+4n\_\{\\mathrm\{bwd\}\}\)NT\.
#### CAT\.
CAT performs akk\-step continuous PGD inner attack and updates a LoRA adapter\. Each PGD step requires one dense forward pass and one input\-gradient backward pass through the dense model, so the inner\-search cost is
Csearch=4kNT\.C\_\{\\mathrm\{search\}\}=4kNT\.The subsequent training update uses one clean forward pass and one adversarial forward pass through the dense model\. Although only the LoRA adapter parameters are updated, gradients must still be backpropagated through the frozen dense network to reach the adapter modules\. We therefore estimate the training\-update cost as
Ctrain≈4NT\+2NT\+2NLoRAT=\(6\+2pLoRA\)NT,C\_\{\\mathrm\{train\}\}\\approx 4NT\+2NT\+2N\_\{\\mathrm\{LoRA\}\}T=\(6\+2p\_\{\\mathrm\{LoRA\}\}\)NT,whereNLoRAN\_\{\\mathrm\{LoRA\}\}is the number of trainable LoRA parameters andpLoRA=NLoRA/Np\_\{\\mathrm\{LoRA\}\}=N\_\{\\mathrm\{LoRA\}\}/N\. In our runs,pLoRA<1%p\_\{\\mathrm\{LoRA\}\}<1\\%, so this term is numerically close to6NT6NT\. Thus,
CCAT≈4kNT\+\(6\+2pLoRA\)NT≈\(4k\+6\)NT\.C\_\{\\mathrm\{CAT\}\}\\approx 4kNT\+\(6\+2p\_\{\\mathrm\{LoRA\}\}\)NT\\approx\(4k\+6\)NT\.
#### Ours\.
Our method uses a pruned surrogate for the inner attack search and updates only a small ReFT intervention on the target model\. LetNsN\_\{\\mathrm\{s\}\}denote the dense\-equivalent surrogate parameter count andNtrN\_\{\\mathrm\{tr\}\}denote the number of trainable intervention parameters\. If the ReFT intervention is applied at layerlrl\_\{\\mathrm\{r\}\}, the latent attack perturbation is applied at layerlal\_\{\\mathrm\{a\}\}, and the model hasLallL\_\{\\mathrm\{all\}\}layers, define
ρlr=Lall−lr−1Lall,ρla=Lall−la−1Lall\.\\rho\_\{l\_\{\\mathrm\{r\}\}\}=\\frac\{L\_\{\\mathrm\{all\}\}\-l\_\{\\mathrm\{r\}\}\-1\}\{L\_\{\\mathrm\{all\}\}\},\\qquad\\rho\_\{l\_\{\\mathrm\{a\}\}\}=\\frac\{L\_\{\\mathrm\{all\}\}\-l\_\{\\mathrm\{a\}\}\-1\}\{L\_\{\\mathrm\{all\}\}\}\.Hereρlr\\rho\_\{l\_\{\\mathrm\{r\}\}\}andρla\\rho\_\{l\_\{\\mathrm\{a\}\}\}approximate the fraction of layers traversed by backward computation after the intervention or attack layer\.
The target\-model update requires a clean forward pass, an adversarial forward pass, and a backward pass from the loss to the ReFT intervention\. We estimate this cost as
Ctarget≈4NT\+4ρlrNT\+4NtrT\.C\_\{\\mathrm\{target\}\}\\approx 4NT\+4\\rho\_\{l\_\{\\mathrm\{r\}\}\}NT\+4N\_\{\\mathrm\{tr\}\}T\.The final term accounts for gradient accumulation on the trainable intervention parameters and is typically negligible relative to dense model computation\.
For the surrogate inner search, each PGD step uses a surrogate forward pass and a backward pass from the output to the latent attack layer\. Therefore,
Csurrogate≈2k\(1\+ρla\)NsT\.C\_\{\\mathrm\{surrogate\}\}\\approx 2k\(1\+\\rho\_\{l\_\{\\mathrm\{a\}\}\}\)N\_\{\\mathrm\{s\}\}T\.The total per\-step cost of our method is
Cours=Ctarget\+Csurrogate\.C\_\{\\mathrm\{ours\}\}=C\_\{\\mathrm\{target\}\}\+C\_\{\\mathrm\{surrogate\}\}\.
For surrogate runs where the exact dense\-equivalent surrogate parameter count is not directly available, we approximate
Ns≈N\(1−fMLPrp\),N\_\{\\mathrm\{s\}\}\\approx N\\left\(1\-f\_\{\\mathrm\{MLP\}\}r\_\{\\mathrm\{p\}\}\\right\),wherefMLPf\_\{\\mathrm\{MLP\}\}is the fraction of dense model parameters in MLP blocks andrpr\_\{\\mathrm\{p\}\}is the pruning ratio of MLP neurons\.Similar Articles
Principled Thoughts for Latent Recursive LLM Systems
The paper presents REST, a novel training objective for latent recursive LLM systems that enhances accuracy by up to 7.5 percentage points across benchmarks by incorporating properties like causality and minimality into differentiable losses.
Layer-wise Curriculum Learning for Efficient LLM Compression
The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Introduces LoCA, a two-stage backpropagation-free method for small-shift adaptation of LLMs, using one-shot calibration to fit local credit assignment maps and closed-form ridge solves for low-rank adapters, achieving lower memory and time than LoRA with competitive cross-entropy on multiple benchmarks.
Hybrid Adversarial Defence for Natural Language Understanding Tasks
Researchers from Southampton and Manchester propose a hybrid adversarial defence framework for LLMs that combines entropy-based, uncertainty-based, and geometric-based models to simultaneously address hallucination and adversarial vulnerability in NLU tasks, achieving up to 64.92% improvement in adversarial robustness and 62.27% reduction in attack success rate.
Learning, Fast and Slow: Towards LLMs That Adapt Continually [R]
This paper introduces a Fast-Slow Training framework for LLMs that combines parameter updates with optimized context to improve sample efficiency and reduce catastrophic forgetting during continual learning.