表征简单性与回路规模以阈值依赖的方式解耦:通过对抗训练进行的对照检验
摘要
该预印本以对抗训练作为对照工具,在 GPT-2 Small 上检验表征简单性(SAE 可分解性、集中化归因)是否蕴含回路在因果上的简单性。研究发现,鲁棒模型具有更强的 SAE 可分解性,而回路规模则取决于机制区间:在高忠实度水平(90-95%)下,鲁棒性有所帮助,但在 85% 以下则无此效果。
arXiv:2609.35890v1 Announce Type: new
Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level. Starting from the same pretrained GPT-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms. We then compare the models along three complementary axes: sparse-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph. The robust model is more SAE-decomposable and engages fewer SAE features in task attribution. Circuit size is regime-dependent: on competence-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness (90%, 95%), a pattern established on the primary pair while representational trends generalize across a seven-point sweep and a second corpus.
查看缓存全文
缓存时间: 2026/09/30 09:37
# Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
Source: [https://arxiv.org/html/2609.35890](https://arxiv.org/html/2609.35890)
###### Abstract
Sparse\-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model’s computation is easier to reverse\-engineer\. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question\. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size\. We investigate this question through*reverse\-engineering complexity*: the causal structure required to recover a model’s behavior at a fixed level of faithfulness\. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level\. Starting from the same pretrained GPT\-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms\. We then compare the models along three complementary axes: sparse\-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph\. The robust model is more SAE\-decomposable and engages fewer SAE features in task attribution\. Circuit size is regime\-dependent: on competence\-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness \(90%, 95%\), a pattern established on the primary pair while representational trends generalize across a seven\-point sweep and a second corpus\.
††footnotetext:Preprint\. Under review\.## 1Introduction
Mechanistic interpretability seeks to reverse\-engineer the computations implemented by neural networks\. Prior work has identified circuits for specific behaviors, including indirect object identification \(IOI\) and Greater\-Than in GPT\-2 Small\([Wang et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib12);[Hanna et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib19)\)\. Yet models that exhibit the same behavior need not implement it through mechanisms of comparable complexity\. Their solutions may differ in how distributed, redundant, or causally entangled they are, and therefore in how difficult they are to recover\.
Adversarial training provides a natural setting in which to study this variation\. It can substantially alter the features and representations learned by a model\([Tsipras et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib13);[Engstrom et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib4)\), and has been associated with more interpretable gradients and saliency maps\([Kim et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib5);[Etmann et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib6)\)\. Recent work also links adversarial vulnerability to superposition and interference between overlapping features\([Gorton and Lewis, 2025](https://arxiv.org/html/2609.35890#bib.bib1);[Bereska et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib2);[Stevinson et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib3)\)\. If robust training reduces this entanglement, it may make learned computations easier to isolate\. Conversely, it may induce representational drift, redundancy, or compensatory pathways that make circuits harder to recover\. Moreover, representational simplicity need not imply causal simplicity\([Xu, 2026](https://arxiv.org/html/2609.35890#bib.bib15)\)\. Whether this representational simplicity translates into causal simplicity – a smaller or more tractable circuit – therefore remains an empirical question, and adversarial training offers a natural instrument to test it: a manipulation that reliably alters representations while task behavior can be held fixed\.
We study this question through*reverse\-engineering complexity*: the causal structure required to recover a model’s behavior at a fixed level of faithfulness\. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level, rather than an inference from representational or attributional proxies alone\. We define a reduction in reverse\-engineering complexity as a smaller faithful circuit for a retained behavior, not as easier SAE reconstruction or more concentrated attribution\. We test this with matched continual training from a shared checkpoint, competence and robustness qualification, and three complementary axes of simplicity, detailed in Section[2](https://arxiv.org/html/2609.35890#S2); this design separates changes in representation geometry from changes in causal organization and tests whether apparent simplicity is instead explained by redundant or compensatory computation\. None of the literatures motivating this prediction test it at the circuit level\.
#### Contributions\.
- •Representational simplicity and circuit recoverability give different answers depending on the fidelity target, in one competence\-matched GPT\-2 Small pair: adversarial training increases SAE decomposability and reduces the number of active attribution features, though not their concentration once normalized by pool size; on IOI, standard training recovers more behavior at low edge budgets, while the robust model reaches very high recovery at substantially fewer top\-kkedges \(90%, 95%\)\.
- •We corroborate this withα\\alpha\-ReQ, the decay rate of the residual\-stream covariance spectrum, an SAE\-free measure of representational simplicity: it rises with adversarial strength across a seven\-point sweep, confirming that the representational trend is not an artifact of the SAE training procedure\.
## 2Experimental Design
### 2\.1Overview
We test whether adversarial training changes how readily a learned computation can be reverse engineered while controlling for initialization, data, training budget, and task performance\. We continually train a standard and a family of adversarially trained GPT\-2 Small models from the same pretrained checkpoint and on the same token stream\. We first qualify candidate pairs by checking task retention and robustness under a multi\-step attack\. We then compare the selected standard and robust models along three progressively stronger axes: \(i\) how accurately matched sparse autoencoders \(SAEs\) reconstruct their internal activations, \(ii\) how many SAE features engage in task attribution, and \(iii\) how large a subgraph must be to reach a given faithfulness level\. The first two axes measure decomposability and attribution structure; only the third directly measures reverse\-engineering complexity\.
Our primary matched\-competence task is indirect object identification \(IOI\), for which all trained models retain near\-ceiling performance\. We additionally evaluate the Greater\-Than \(GT\) task, which is not competence\-matched between models\.
### 2\.2Matched continual training
All models initialize from the official 124M\-parameter GPT\-2 Small checkpoint\. Each model receives the same 1B\-token stream in the same order and updates all parameters for 1,908 optimization steps\. The primary comparison uses OpenWebText; Section[3\.5](https://arxiv.org/html/2609.35890#S3.SS5)reports an identical standard\-robust pair trained on FineWeb instead, to test whether the corpus drives the result\. The standard arm minimizes next\-token cross\-entropy\. The adversarial arms optimize
ℒ=\(1−α\)ℒclean\+αℒadv,\\mathcal\{L\}=\(1\-\\alpha\)\\mathcal\{L\}\_\{\\mathrm\{clean\}\}\+\\alpha\\mathcal\{L\}\_\{\\mathrm\{adv\}\},\(1\)whereℒadv\\mathcal\{L\}\_\{\\mathrm\{adv\}\}is evaluated after a ten\-step projected\-gradient attack on token\-embedding activations\. Perturbations are applied after the token lookup and before positional embeddings\. For every non\-special token, the attack is constrained to anℓ2\\ell\_\{2\}ball with radiusϵrel\\epsilon\_\{\\mathrm\{rel\}\}times the mean non\-special\-token embedding norm in the micro\-batch\. We use random initialization, per\-token normalized gradient ascent, and projection after every step; the special token is never perturbed\. This continuous embedding\-space construction follows the general setup of[Xhonneux et al\. \(2024\)](https://arxiv.org/html/2609.35890#bib.bib42)\.
We sweepϵrel∈\{0\.05,0\.075,0\.10\}\\epsilon\_\{\\mathrm\{rel\}\}\\in\\\{0\.05,0\.075,0\.10\\\}andα∈\{0\.2,0\.5\}\\alpha\\in\\\{0\.2,0\.5\\\}, producing six adversarial candidates and one standard control\. All runs use AdamW, a peak learning rate of5×10−55\\times 10^\{\-5\}with 100 warmup steps and cosine decay, a global batch of 524,288 tokens, gradient clipping at 1\.0, and bfloat16 computation\. Appendix[B](https://arxiv.org/html/2609.35890#A2)gives the complete optimization and attack specification\.
### 2\.3Tasks and qualification criteria
For IOI, we evaluate 1,000 fixed prompts balanced between ABBA and BABA templates\([Wang et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib12)\)\. The behavioral score is the logit difference between the correct and incorrect names at the final position, and accuracy is the fraction of prompts with positive logit difference\. The standard model must achieve at least 90% accuracy and each adversarial candidate at least 85%, floors that allow for ordinary continual\-training drift while remaining far above the level indicating circuit absence \(55\.4% for a from\-scratch control that failed to develop the behavior, Appendix[B](https://arxiv.org/html/2609.35890#A2)\)\.
For robustness, we evaluate held\-out OpenWebText under both a single\-step attack and the ten\-step attack used during training\. The masking sanity check requires
\|ℒPGD10−ℒsingle\|≤0\.10nats,\\left\|\\mathcal\{L\}\_\{\\mathrm\{PGD10\}\}\-\\mathcal\{L\}\_\{\\mathrm\{single\}\}\\right\|\\leq 0\.10\\text\{ nats\},\(2\)which guards against selecting a model that appears robust only under a weak attack\([Athalye et al\., 2018](https://arxiv.org/html/2609.35890#bib.bib14)\)\. Among candidates passing this check and the IOI floor, we rank robustness by attack susceptibility,
Δadv=ℒPGD10−ℒclean,\\Delta\_\{\\mathrm\{adv\}\}=\\mathcal\{L\}\_\{\\mathrm\{PGD10\}\}\-\\mathcal\{L\}\_\{\\mathrm\{clean\}\},\(3\)where lower values indicate a smaller attack\-induced loss increase\. We select the qualified candidate with the lowestΔadv\\Delta\_\{\\mathrm\{adv\}\}\. The single\-step versus PGD\-10 gap is a masking diagnostic, not a robustness score; the non\-robust standard control is therefore not required to pass it\.
We also evaluate 1,000 GT prompts generated with the official YearDataset\([Hanna et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib19)\)\. For a two\-digit thresholdYYYY, the probability\-difference score is
PD=∑y\>YYp\(y\)−∑y≤YYp\(y\),\\operatorname\{PD\}=\\sum\_\{y\>YY\}p\(y\)\-\\sum\_\{y\\leq YY\}p\(y\),\(4\)where the sums use the full\-softmax probabilities of the 100 two\-digit year tokens\. We report paired differences from the standard model on identical prompts\. GT is not a qualification gate because no established performance floor analogous to the IOI accuracy threshold exists for this task\.
### 2\.4Axis 1: SAE decomposability
We train matched BatchTopK SAEs\([Bussmann et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib41)\)on the residual stream entering transformer block 8, following standard practice of choosing a layer near the end of the network that contains diverse features without being specialized for final\-layer output computation\([Gao et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib26)\)\. Each SAE has 12,288 features, a 16\-fold expansion over the 768\-dimensional residual stream, and an averageL0L\_\{0\}of 32 selected features per token\. Training uses 300M post\-training OpenWebText tokens disjoint from model training; evaluation uses a subsequent 10M\-token partition\. The standard and robust SAEs share the same architecture, sparsity target, data volume, optimizer, number of updates, and evaluation protocol\. We train three seeds per condition for the principal reconstruction comparison\.
The primary endpoint is mean squared reconstruction error in the original residual space on the held\-out evaluation partition\. We also report normalized\-space training loss, fixed\-threshold inference error, dead\-feature counts, and the between\-seed dispersion\. The robust SAE’s lower error at fixed dictionary size and sparsity indicates that its activation distribution is more compressible by this SAE family\.
### 2\.5Axis 2: SAE feature engagement and selectivity
For each SAE \(seed 42 primary; seeds 43 and 44 replicated\), we attribute the IOI and GT scores to its features using node\-level attribution patching\([Syed et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib17);[Marks et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib34)\)\. We splice the reconstructionx^\\hat\{x\}into the residual stream as
x~=x\+x^−sg\(x^\),\\tilde\{x\}=x\+\\hat\{x\}\-\\operatorname\{sg\}\(\\hat\{x\}\),\(5\)such that the forward pass is exactly the unmodified model while gradients flow through the SAE features\. For featureii, target attribution is
Ai=𝔼x\[∑t∈𝒯∂Y∂fi,t\(fi,tclean−fi,tcorrupt\)\],A\_\{i\}=\\mathbb\{E\}\_\{x\}\\left\[\\sum\_\{t\\in\\mathcal\{T\}\}\\frac\{\\partial Y\}\{\\partial f\_\{i,t\}\}\\left\(f\_\{i,t\}^\{\\mathrm\{clean\}\}\-f\_\{i,t\}^\{\\mathrm\{corrupt\}\}\\right\)\\right\],\(6\)whereYYis IOI logit difference or GT probability difference\. We define the intervention position set𝒯\\mathcal\{T\}per task:𝒯IOI=\{S1,IO,S2,END\}\\mathcal\{T\}\_\{\\mathrm\{IOI\}\}=\\\{S1,IO,S2,\\mathrm\{END\}\\\}, the three name positions and the final position;𝒯GT=\{start\-year span,END\}\\mathcal\{T\}\_\{\\mathrm\{GT\}\}=\\\{\\text\{start\-year span\},\\mathrm\{END\}\\\}, the token span stating the prompt’s start year and the final position\. Corrupted prompts swap the two IOI names or use the official GT counterfactual construction\.
Our primary cross\-model statistic is candidate pool size: the number of features active at task\-relevant positions under an identical existence rule and identical prompts, measured before any ranking or attribution computation\. As a secondary statistic, after sorting existing features by\|Ai\|\|A\_\{i\}\|, we report the smallest number required to account for 50%, 80%, and 90% of total absolute attribution; this concentration statistic is not scale\-free with respect to pool size \(Section[3\.3](https://arxiv.org/html/2609.35890#S3.SS3)\), so we report it alongside pool size rather than as a stand\-alone measure of task simplicity\. As a complementary check on whether a feature’s causal role is specific to the task or also shapes the model’s broader output, we zero each pool feature on the same prompts and measure the shift in the full next\-token distribution rather than the task readout alone \(mean ablation\-KL, Appendix[F](https://arxiv.org/html/2609.35890#A6)\)\.
### 2\.6Axis 3: faithful circuit recovery
We recover task circuits with edge attribution patching \(EAP\-IG\); a pilot of EAP\-GP, a more recent variant designed to reduce gradient saturation, is reported in Appendix[G](https://arxiv.org/html/2609.35890#A7)\([Syed et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib17);[Zhang et al\., 2025a](https://arxiv.org/html/2609.35890#bib.bib33)\)\. For each task and model, we define edge\-native faithfulness atkkretained edges as
F\(k\)=Y\(circuitk\)−Y\(corrupt\)Y\(clean\)−Y\(corrupt\),F\(k\)=\\frac\{Y\(\\mathrm\{circuit\}\_\{k\}\)\-Y\(\\mathrm\{corrupt\}\)\}\{Y\(\\mathrm\{clean\}\)\-Y\(\\mathrm\{corrupt\}\)\},\(7\)wherecircuitk\\mathrm\{circuit\}\_\{k\}patches in the top\-kkedges by\|score\|\|\\mathrm\{score\}\|at their clean values and corrupt\-patches every other edge to its value under the task’s counterfactual distribution \(name\-swap for IOI, the official counterfactual construction for GT, Section[2\.3](https://arxiv.org/html/2609.35890#S2.SS3)\);Y\(clean\)Y\(\\mathrm\{clean\}\)andY\(corrupt\)Y\(\\mathrm\{corrupt\}\)are the task metric \(IOI logit difference or GT probability difference\) on the fully clean and fully corrupted graph, soF=0F=0reproduces the fully corrupted baseline andF=1F=1fully recovers clean performance\. We distinguish two estimands on the sameF\(k\)F\(k\)curve\. Partial\-subgraph recoverability is the area underF\(k\)F\(k\)across the full edge grid, includingkkwhereF<85%F<85\\%\. Faithful\-circuit size is the first gridkkat whichF\(k\)F\(k\)reaches a target in the circuit regimeF≥85%F\\geq 85\\%; we report that crossing at several targets, not at a single canonical cutoff\. The causal claim in this paper is about this estimand\. AUC is reported as a descriptor of partial recovery, not as a circuit\-size statistic\. Conditional co\-ablation tests whether apparently weak components become necessary after substitutes are removed, guarding against redundancy and self\-repair\([Gong et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib45);[Hanna et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib18);[Rushing and Nanda, 2024](https://arxiv.org/html/2609.35890#bib.bib36)\)\.
## 3Results
### 3\.1Adversarial continual training yields a qualified robust comparison
All seven continually trained models retain IOI performance: accuracy ranges from 99\.4% to 99\.6%, compared with 99\.7% for the original checkpoint\. Four adversarial candidates also pass the masking sanity check\. Among them,ft\-e100\-a05has the lowest attack susceptibility, reducingΔadv\\Delta\_\{\\mathrm\{adv\}\}from 1\.642 nats in the standard control to 0\.362 while retaining 99\.4% IOI accuracy\. We therefore useft\-cleanandft\-e100\-a05as the principal standard\-robust pair\.
The near\-constant IOI accuracy does not extend to GT\. GT probability difference decreases smoothly as attack susceptibility falls across the sweep \(Table[1](https://arxiv.org/html/2609.35890#S3.T1)\)\. The selected robust model scores 0\.442, compared with 0\.671 for the standard model, with a paired prompt\-level difference of−0\.229±0\.010\-0\.229\\pm 0\.010\. By construction, PD lies in\[−1,1\]\[\-1,1\], with values near zero indicating no informative signal; both models remain well above that floor, confirming the behavior persists substantially intact despite the competence gap between models on this task\.
Table 1:Qualification sweep and secondary GT performance\. The gap is\|ℒPGD10−ℒsingle\|\|\\mathcal\{L\}\_\{\\mathrm\{PGD10\}\}\-\\mathcal\{L\}\_\{\\mathrm\{single\}\}\|;Δadv\\Delta\_\{\\mathrm\{adv\}\}is the attack\-induced loss increase, so lower is more robust\. Asterisks mark adversarial candidates that pass both qualification gates\.
### 3\.2Robust activations are more sparsely reconstructable
At equal dictionary size and sparsity, the SAE trained on robust\-model activations reconstructs the held\-out residual stream more accurately\. Across three SAE seeds, original\-space MSE is1\.1293±0\.00091\.1293\\pm 0\.0009for the robust model and1\.1601±0\.00221\.1601\\pm 0\.0022for the standard model, a 2\.65% reduction with non\-overlapping seed ranges\. On position\-matched evaluation blocks, the standard\-minus\-robust difference is0\.0323±0\.00070\.0323\\pm 0\.0007\. The same ordering holds under fixed\-threshold inference and in normalized activation space \(Table[2](https://arxiv.org/html/2609.35890#S3.T2)\), and at most one of the three seeds per condition has a single dead feature \(Appendix[E](https://arxiv.org/html/2609.35890#A5)\)\.
Table 2:Matched SAE reconstruction at the residual stream entering block 8\. Lower is better\. The original\-space entry reports mean±\\pmstandard deviation across three SAE seeds; the remaining entries use seed 42\.
### 3\.3Robust task attribution engages fewer SAE features, not more concentrated ones
Two effects distinguish the models, both consistent across all three SAE seeds and both tasks\. First, under an identical existence rule and identical prompts, fewer robust\-SAE features are active at task\-relevant positions: 1,196 versus 1,612 candidates for IOI \(a 26% reduction\) and 360 versus 563 for GT \(36%; Table[3](https://arxiv.org/html/2609.35890#S3.T3)\)\. This measurement precedes any ranking or attribution computation\.
Second, among active features, the robust SAE reaches each attribution\-mass threshold with fewer of them in absolute terms: at 90% attribution mass, IOI requires 276 features versus 343, and GT requires 63 versus 90 \(Table[3](https://arxiv.org/html/2609.35890#S3.T3), Figure[1](https://arxiv.org/html/2609.35890#S3.F1)\)\. Normalized by candidate pool size, this reverses: robust needs a larger share of its \(smaller\) pool to reach 90% mass on both tasks \(IOI 23\.1% versus 21\.3%; GT 17\.5% versus 16\.0%\), so the effect is that fewer features are active at all, not that attribution is more concentrated among them once pool size is accounted for\. GT probability difference is lower under adversarial training \(Table[1](https://arxiv.org/html/2609.35890#S3.T1)\), so fewer GT features can reflect a weaker behavior rather than cleaner features\. We read the GT counts as consistent with IOI, not as an independent matched comparison\. Zeroing each pool feature and measuring the resulting shift in the model’s general output distribution \(mean ablation\-KL\) shows the same directional pattern \(Appendix[F](https://arxiv.org/html/2609.35890#A6)\)\.
Table 3:Number of SAE features required to explain cumulative absolute target attribution\. Reported values use the seed\-42 SAE for each condition; the ordering holds at seeds 43 and 44 as well\.Figure 1:Cumulative attribution share by feature rank \(descending\|Ai\|\|A\_\{i\}\|\), IOI \(left\) and GT \(right\)\. The robust curve reaches each threshold with fewer features in absolute terms; normalized by candidate pool size this reverses \(Section[3\.3](https://arxiv.org/html/2609.35890#S3.SS3)\)\.
### 3\.4Circuit size on IOI depends on the faithfulness target
We recover faithful IOI circuits from the raw computational graph \(156 nodes: 144 attention heads and 12 MLPs; 32,491 edges\) using EAP\-IG \(m=5m=5interpolation steps, Section[2\.6](https://arxiv.org/html/2609.35890#S2.SS6)\), on 200 prompts \(100 ABBA / 100 BABA, name\-swap counterfactuals\)\. Table[4](https://arxiv.org/html/2609.35890#S3.T4)reports the first edge count reaching five faithfulness levels, rather than a single arbitrary threshold: at 70% standard leads; at 80–85% the models tie; at 90% robust needs far fewer edges \(1,000 versus 4,000\); at 95% the gap widens further \(2,000 versus 16,000\)\. Standard’s curve plateaus after 85–90%, requiring a much larger circuit to close the remaining gap to near\-complete faithfulness; robust’s curve continues climbing steeply over the same range \(Figure[2](https://arxiv.org/html/2609.35890#S3.F2)\)\. Below 85% faithfulness, neither subgraph reproduces the behavior well enough to constitute a circuit; the comparison there is between two non\-faithful partial subgraphs, not between circuit sizes \(raw values in Appendix[G](https://arxiv.org/html/2609.35890#A7), Table[12](https://arxiv.org/html/2609.35890#A7.T12)\)\. At and above 85%, robust never needs more edges than standard, and needs substantially fewer at 90% and 95%\.
Table 4:First edge count reaching each faithfulness threshold, IOI, EAP\-IG\. Both models are evaluated on the same geometric grid; a tie reflects grid resolution, not necessarily identical true crossing points\. Left of the rule \(F<85%F<85\\%\): partial subgraphs, not faithful circuits\. Right of the rule \(F≥85%F\\geq 85\\%\): circuit\-size comparison; 85% is a tie\.AUC \(0\.87 standard versus 0\.60 robust\) is the partial\-subgraph recoverability estimand, as defined in Section[2\.6](https://arxiv.org/html/2609.35890#S2.SS6); it does not speak to the edge count required to reach a faithful circuit at 90% or 95%\. Conditional co\-ablation, recovering backup components to account for self\-repair at the node layer \(Section[2\.6](https://arxiv.org/html/2609.35890#S2.SS6)\), preserves the AUC ordering favoring standard at a much smaller scale \(10–14 nodes, both models far below full faithfulness\); it does not test the 90/95% edge\-level crossings above\.
\(a\)  \(b\) 
Figure 2:Two readings of the same IOI edge\-native faithfulness curve, EAP\-IG\.\(a\)F\(k\)F\(k\)against edges retained, log\-kkscale: standard and robust cross 85% faithfulness at the same edge count; standard leads by the largest margin in the steep rising region \(k≈64k\\approx 64–512\); robust is locally ahead neark=32k=32and acrossk≈2000k\\approx 2000–16000; net area under the curve favors standard \(0\.87 versus 0\.60\)\.\(b\)First grid point reaching each faithfulness threshold, evaluated on the geometric gridk∈\{1,2,4,8,16,32,64,128,256,512,1000,2000,4000,8000,16000,32491\}k\\in\\\{1,2,4,8,16,32,64,128,256,512,1000,2000,4000,8000,16000,32491\\\}; a reported crossing atkkmeans the true crossing lies in\(k′,k\]\(k^\{\\prime\},k\]for the preceding grid pointk′k^\{\\prime\}, not atkkexactly\. The dotted line marks 85%, the threshold used for the single\-crossing comparison above; no threshold is a literature standard for this metric, which is why we report the full set of crossings rather than one value\.Greater\-Than moves the same direction at 95% \(2,000 versus 8,000 edges\)\. Because GT PD differs between models, we do not use that curve as a test of the main claim\. Full detail, and the EAP\-GP pilot, are in Appendix[G](https://arxiv.org/html/2609.35890#A7)\.
### 3\.5Representational simplicity tracks robustness across the sweep
The qualification sweep \(Section[3\.1](https://arxiv.org/html/2609.35890#S3.SS1)\) produced seven continually trained models spanningϵrel∈\{0\.05,0\.075,0\.10\}\\epsilon\_\{\\mathrm\{rel\}\}\\in\\\{0\.05,0\.075,0\.10\\\}andα∈\{0\.2,0\.5\}\\alpha\\in\\\{0\.2,0\.5\\\}, not just the single selected robust model used above\. We extend all three axes to all seven, and additionally measureα\\alpha\-ReQ, the power\-law decay exponent of the residual\-stream covariance spectrum \(Appendix[H](https://arxiv.org/html/2609.35890#A8)\), a representational\-simplicity measure that does not depend on a trained SAE\.
Across the sweep and FineWeb, SAE/α\\alpha\-ReQ simplicity increases with robustness while the 85% IOI crossing stays at or above the same grid point; 90/95% crossings were not computed for these models\. The high\-faithfulness gap in Section[3\.4](https://arxiv.org/html/2609.35890#S3.SS4)is reported for the primary pair only\.
Table 5:Representational and causal metrics across the full qualification sweep\. Reconstruction MSE \(inference\-threshold mode\) and IOI candidates are three\-seed means \(SAE training seeds 42/43/44\);α\\alpha\-ReQ, participation ratio, and edges\-to\-85% are single deterministic measurements, since neither theα\\alpha\-ReQ probe \(SAE\-free, computed directly from the activation covariance\) nor circuit recovery \(run on the raw graph\) depends on SAE training stochasticity\. The edges\-to\-85% column is the single 85% gate used for qualification\. Representational metrics move in the predicted direction as adversarial strength increases, with occasional non\-monotonic steps between adjacent configurations of similar measured robustness \(Table[1](https://arxiv.org/html/2609.35890#S3.T1)\); circuit size does not track them at this threshold\.Figure 3:Representational metrics \(panels 1–5\) move in the predicted direction with adversarial strength across the seven\-model sweep, ordered by measuredΔadv\\Delta\_\{\\mathrm\{adv\}\}\(1\.642→\\to0\.362\); two adjacent configurations with near\-identical measured robustness \(e100\-a02,e050\-a05\) produce small non\-monotonic steps in reconstruction MSE and participation ratio\. Panel 6 shows the single 85% qualification gate, flat at 1,000 edges except for one outlier \(e050\-a02, 2,000 edges – larger, not smaller\)\. Reconstruction MSE \(inference\-threshold mode\) and IOI candidates are three\-seed means with±1\\pm 1SD error bars \(seeds 42/43/44\);α\\alpha\-ReQ, participation ratio, and edges\-to\-85% are single deterministic measurements per model \(Appendix[H](https://arxiv.org/html/2609.35890#A8)\)\.We separately test whether the representational trend depends on OpenWebText specifically, using one additional standard\-robust pair continually trained identically to the primary comparison except for the pretraining corpus \(FineWeb in place of OpenWebText; Appendix[I](https://arxiv.org/html/2609.35890#A9)\)\. Every representational metric moves in the same direction as the primary result, and the 85% edge count is again flat \(Table[6](https://arxiv.org/html/2609.35890#S3.T6)\)\.
Table 6:Corpus ablation \(inference\-threshold mode MSE\): standard vs\. robust GPT\-2 Small continually trained on FineWeb in place of OpenWebText, otherwise identical to the primary comparison\.
## 4Conclusion
In this paper, we test whether adversarial training reduces reverse\-engineering complexity, using it as a controlled instrument that reshapes representations while task competence is held fixed\. To this end, we compare matched standard and adversarially trained GPT\-2 Small models along three complementary axes: SAE decomposability, SAE feature engagement, and the size of faithful circuits recovered from the raw computational graph\.
Our results reveal three key findings: \(1\) the robust model is more SAE\-decomposable and engages fewer active SAE features, though not more concentrated ones once normalized by pool size, corroborated byα\\alpha\-ReQ, an SAE\-free measure that rules out an SAE\-specific artifact; \(2\) faithful circuit size is regime\-dependent – standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness \(90%, 95%\); and \(3\) the representational trend generalizes across a seven\-point adversarial\-strength sweep and a second corpus, while the causal result is established only on the primary pair, making this a starting point for testing generalization across scale, architecture, and robustification method\. Our broader motivation is to test whether robustness can improve model auditing and reshape how interpretability\-oriented training is prioritized; this work is a first step toward that goal\.
## 5Related Work
Adversarial examples can exploit predictive but non\-robust features, and adversarial training changes which features a model uses\([Tsipras et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib13);[Ilyas et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib20)\), with robustness linked to more human\-aligned gradients, saliency maps, and representations\([Ross and Doshi\-Velez, 2018](https://arxiv.org/html/2609.35890#bib.bib7);[Etmann et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib6);[Engstrom et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib4);[Wang et al\., 2022](https://arxiv.org/html/2609.35890#bib.bib21)\); these results concern input sensitivity or representational geometry, not causal mechanism size\. A closer precedent comes from vision: adversarially trained networks admit sparser attribution vectors\([Chalasani et al\., 2020](https://arxiv.org/html/2609.35890#bib.bib43)\)and exhibit feature purification\([Allen\-Zhu and Li, 2020](https://arxiv.org/html/2609.35890#bib.bib44)\), with gradients and saliency maps aligning more closely with human\-interpretable structure\([Kim et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib5)\)– predicting simpler features and sparser attributions, not a smaller causal graph, the prediction we test\.
Superposition offers a sharper hypothesis, that models represent more features than dimensions by accepting interference\([Elhage et al\., 2022](https://arxiv.org/html/2609.35890#bib.bib23)\), recently linked to adversarial vulnerability\([Gorton and Lewis, 2025](https://arxiv.org/html/2609.35890#bib.bib1);[Stevinson et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib3);[Gao et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib24);[Zhang et al\., 2025b](https://arxiv.org/html/2609.35890#bib.bib9);[Bereska et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib2)\); SAEs measure this decomposition at scale\([Huben et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib25);[Gao et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib26)\), though reconstruction and sparsity alone do not guarantee causal features\([Karvonen et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib27);[Paulo and Belrose, 2026](https://arxiv.org/html/2609.35890#bib.bib29);[Li et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib10)\), motivating our separate evaluation of decomposability and causal structure\. On the causal side, manual and automated methods have identified and scaled circuit recovery for GPT\-2 Small behaviors\([Wang et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib12);[Hanna et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib19);[Conmy et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib32);[Syed et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib17);[Hanna et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib18)\), though recovered circuits remain conditional on the attribution method, threshold, and pruning procedure\([Hanna et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib18)\), and ablations can trigger self\-repair that understates causal dependence\([McGrath et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib38);[Rushing and Nanda, 2024](https://arxiv.org/html/2609.35890#bib.bib36)\)\.
The closest empirical studies apply mechanistic interpretability to adversarial vulnerability:[García\-Carrasco et al\. \(2024\)](https://arxiv.org/html/2609.35890#bib.bib39)localize components affected by adversarial examples in a fixed GPT\-2 Small circuit, and[Walkowiak et al\. \(2025\)](https://arxiv.org/html/2609.35890#bib.bib11)combine adversarial evaluation with EAP\-IG circuits across languages;[Bereska \(2024\)](https://arxiv.org/html/2609.35890#bib.bib40)lists reverse\-engineering robust models in a self\-published proposal, not executed empirically\. To our knowledge, we are the first to compare matched standard and adversarially trained language models through SAE decomposability, SAE feature engagement, and the size of established task circuits at fixed faithfulness\. Appendix[A](https://arxiv.org/html/2609.35890#A1)discusses SAE evaluation, circuit\-discovery variants, and self\-repair in greater detail\.
## AI Use Disclosure
We used Claude \(Anthropic\) to assist with manuscript preparation\.
Writing and polishing\.Claude assisted with proofreading, polishing, and structuring the paper’s sections\.
Citations\.Claude assisted with verifying citations against primary sources and identifying missing bibliography entries\.
LaTeX and formatting\.Claude assisted with LaTeX formatting, table and figure layout, and troubleshooting page\-length and float\-placement issues\.
All experimental design, model training, data analysis, and interpretation of results were carried out by the authors\. All claims, figures, and interpretations were reviewed and approved by the authors, who take full responsibility for the submission’s content and accuracy\.
## References
- Allen\-Zhu and Li \(2020\)Z\. Allen\-Zhu and Y\. LiFeature purification: how adversarial training performs robust deep learning\.External Links:2005\.10190,[Link](https://arxiv.org/abs/2005.10190)Cited by:[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Athalyeet al\.\(2018\)A\. Athalye, N\. Carlini, and D\. WagnerObfuscated gradients give a false sense of security: circumventing defenses to adversarial examples\.InProceedings of the 35th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.80,pp\. 274–283\.External Links:[Link](https://proceedings.mlr.press/v80/athalye18a.html)Cited by:[§2\.3](https://arxiv.org/html/2609.35890#S2.SS3.p2.2)\.
- Bereskaet al\.\(2025\)L\. F\. Bereska, Z\. Tzifa\-Kratira, R\. Samavi, and S\. GavvesSuperposition as lossy compression: measure with sparse autoencoders and connect to adversarial vulnerability\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=qaNP6o5qvJ)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Bereska \(2024\)L\. BereskaMechanistic interpretability for adversarial robustness: a proposal\.Self\-published\.External Links:[Link](https://leonardbereska.github.io/blog/2024/mechrobustproposal/)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p3.1)\.
- Boopathyet al\.\(2020\)A\. Boopathy, S\. Liu, G\. Zhang, C\. Liu, P\. Chen, S\. Chang, and L\. DanielProper network interpretability helps adversarial robustness in classification\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 1014–1023\.External Links:[Link](https://proceedings.mlr.press/v119/boopathy20a.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1)\.
- Bussmannet al\.\(2024\)B\. Bussmann, P\. Leask, and N\. NandaBatchTopK sparse autoencoders\.InNeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning,External Links:2412\.06410,[Link](https://openreview.net/forum?id=d4dpOCqybL)Cited by:[§2\.4](https://arxiv.org/html/2609.35890#S2.SS4.p1.1)\.
- Chalasaniet al\.\(2020\)P\. Chalasani, J\. Chen, A\. R\. Chowdhury, X\. Wu, and S\. JhaConcise explanations of neural networks using adversarial training\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 1383–1391\.External Links:[Link](https://proceedings.mlr.press/v119/chalasani20a.html)Cited by:[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Chanin \(2026\)D\. ChaninAre sparse autoencoder benchmarks reliable?\.External Links:2605\.18229,[Link](https://arxiv.org/abs/2605.18229)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1)\.
- Conmyet al\.\(2023\)A\. Conmy, A\. N\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-AlonsoTowards automated circuit discovery for mechanistic interpretability\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 16318–16352\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/34e1dbe95d34d7ebaf99b9bcaeb5b2be-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Elhageet al\.\(2022\)N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. OlahToy models of superposition\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2022/toy_model/index.html)Cited by:[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Engstromet al\.\(2019\)L\. Engstrom, A\. Ilyas, S\. Santurkar, D\. Tsipras, B\. Tran, and A\. MadryAdversarial robustness as a prior for learned representations\.External Links:1906\.00945,[Link](https://arxiv.org/abs/1906.00945)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Etmannet al\.\(2019\)C\. Etmann, S\. Lunz, P\. Maass, and C\. SchönliebOn the connection between adversarial robustness and saliency map interpretability\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 1823–1832\.External Links:[Link](https://proceedings.mlr.press/v97/etmann19a.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Gaoet al\.\(2026\)J\. Gao, Z\. Lu, R\. Mudumbai, X\. Wu, J\. Yi, M\. Cho, C\. Xu, H\. Xie, and W\. XuFeature compression is the root cause of adversarial fragility in neural networks\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/cbcce87f745072c819204529be843d16-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Gaoet al\.\(2025\)L\. Gao, T\. Dupré la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. WuScaling and evaluating sparse autoencoders\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=tcsZt9ZNKD)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1),[§2\.4](https://arxiv.org/html/2609.35890#S2.SS4.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- García\-Carrascoet al\.\(2024\)J\. García\-Carrasco, A\. Maté, and J\. TrujilloDetecting and understanding vulnerabilities in language models via mechanistic interpretability\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 385–393\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2024/43),[Link](https://www.ijcai.org/proceedings/2024/0043)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p3.1)\.
- Gonget al\.\(2026\)Z\. Gong, Z\. Zeng, C\. Yuen, and W\. Y\. B\. LimConditional co\-ablation: recovering self\-repair backups in transformer circuits\.External Links:2607\.01940,[Link](https://arxiv.org/abs/2607.01940)Cited by:[Appendix G](https://arxiv.org/html/2609.35890#A7.p6.1),[§2\.6](https://arxiv.org/html/2609.35890#S2.SS6.p1.2)\.
- Gorton and Lewis \(2025\)L\. Gorton and O\. LewisAdversarial examples are not bugs, they are superposition\.External Links:2508\.17456,[Link](https://arxiv.org/abs/2508.17456)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Hadadet al\.\(2026\)I\. Hadad, G\. Katz, and S\. BassanFormal mechanistic interpretability: automated circuit discovery with provable guarantees\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/b5afe13494c825089b1e3944fdaba212-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1)\.
- Hannaet al\.\(2023\)M\. Hanna, O\. Liu, and A\. VariengienHow does GPT\-2 compute greater\-than?: interpreting mathematical abilities in a pre\-trained language model\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://openreview.net/forum?id=p4PckNQR8k)Cited by:[§1](https://arxiv.org/html/2609.35890#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.35890#S2.SS3.p3.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Hannaet al\.\(2024\)M\. Hanna, S\. Pezzelle, and Y\. BelinkovHave faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=grXgesr5dT)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1),[§2\.6](https://arxiv.org/html/2609.35890#S2.SS6.p1.2),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Hubenet al\.\(2024\)R\. Huben, H\. Cunningham, L\. Riggs Smith, A\. Ewart, and L\. SharkeySparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Ilyaset al\.\(2019\)A\. Ilyas, S\. Santurkar, D\. Tsipras, L\. Engstrom, B\. Tran, and A\. MadryAdversarial examples are not bugs, they are features\.InAdvances in Neural Information Processing Systems,Vol\.32\.External Links:[Link](https://proceedings.neurips.cc/paper/2019/hash/e2c420d928d4bf8ce0ff2ec19b371514-Abstract.html)Cited by:[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Karvonenet al\.\(2025\)A\. Karvonen, C\. Rager, J\. Lin, C\. Tigges, J\. I\. Bloom, D\. Chanin, Y\. Lau, E\. Farrell, C\. McDougall, K\. Ayonrinde, M\. Wearden, A\. Conmy, S\. Marks, and N\. NandaSAEBench: a comprehensive benchmark for sparse autoencoders in language model interpretability\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 29223–29264\.External Links:[Link](https://proceedings.mlr.press/v267/karvonen25a.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Kimet al\.\(2019\)B\. Kim, J\. Seo, and T\. JeonBridging adversarial robustness and gradient interpretability\.InICLR Workshop on Safe Machine Learning: Specification, Robustness, and Assurance,External Links:[Link](https://arxiv.org/abs/1903.11626)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Liet al\.\(2026\)A\. J\. Li, S\. Srinivas, U\. Bhalla, and H\. LakkarajuEvaluating adversarial robustness of concept representations in sparse autoencoders\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5940–5957\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.279),[Link](https://aclanthology.org/2026.eacl-long.279/)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Markset al\.\(2025\)S\. Marks, C\. Rager, E\. J\. Michaud, Y\. Belinkov, D\. Bau, and A\. MuellerSparse feature circuits: discovering and editing interpretable causal graphs in language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=I4e82CIDxv)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1),[§2\.5](https://arxiv.org/html/2609.35890#S2.SS5.p1.1)\.
- McDougallet al\.\(2024\)C\. S\. McDougall, A\. Conmy, C\. Rushing, T\. McGrath, and N\. NandaCopy suppression: comprehensively understanding a motif in language model attention heads\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 337–363\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.22),[Link](https://aclanthology.org/2024.blackboxnlp-1.22/)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px4.p1.1)\.
- McGrathet al\.\(2023\)T\. McGrath, M\. Rahtz, J\. Kramár, V\. Mikulik, and S\. LeggThe hydra effect: emergent self\-repair in language model computations\.External Links:2307\.15771,[Link](https://arxiv.org/abs/2307.15771)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px4.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Minegishiet al\.\(2025\)G\. Minegishi, H\. Furuta, Y\. Iwasawa, and Y\. MatsuoRethinking evaluation of sparse autoencoders through the representation of polysemous words\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/380a0b16a7e6f8c5010f798c9f2d3c61-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1)\.
- Noacket al\.\(2021\)A\. Noack, I\. Ahern, D\. Dou, and B\. LiAn empirical study on the relation between network interpretability and adversarial robustness\.SN Computer Science2\(1\),pp\. 32\.External Links:[Document](https://dx.doi.org/10.1007/s42979-020-00390-x),[Link](https://doi.org/10.1007/s42979-020-00390-x)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1)\.
- Paulo and Belrose \(2026\)G\. Paulo and N\. BelroseSparse autoencoders trained on the same data learn different features\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/3c1fe56b043848b211030c202764c6a7-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Ross and Doshi\-Velez \(2018\)A\. S\. Ross and F\. Doshi\-VelezImproving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.32\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v32i1.11504),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/11504)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Rushing and Nanda \(2024\)C\. Rushing and N\. NandaExplorations of self\-repair in language models\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 42836–42855\.External Links:[Link](https://proceedings.mlr.press/v235/rushing24a.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px4.p1.1),[§2\.6](https://arxiv.org/html/2609.35890#S2.SS6.p1.2),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Songet al\.\(2026\)X\. Song, A\. Muhamed, Y\. Zheng, L\. Kong, Z\. Tang, M\. T\. Diab, V\. Smith, and K\. ZhangMechanistic interpretability should prioritize feature consistency in sparse autoencoders\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2172–2210\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.99),[Link](https://aclanthology.org/2026.acl-long.99/)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px2.p1.1)\.
- Stevinsonet al\.\(2026\)E\. Stevinson, L\. Prieto, M\. Barsbey, and T\. BirdalAdversarial vulnerability from interference between features in superposition\.InProceedings of the 43rd International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=fD6c4bD7sR)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Syedet al\.\(2024\)A\. Syed, C\. Rager, and A\. ConmyAttribution patching outperforms automated circuit discovery\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 407–416\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.25),[Link](https://aclanthology.org/2024.blackboxnlp-1.25/)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1),[§2\.5](https://arxiv.org/html/2609.35890#S2.SS5.p1.1),[§2\.6](https://arxiv.org/html/2609.35890#S2.SS6.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Tiggeset al\.\(2024\)C\. Tigges, M\. Hanna, Q\. Yu, and S\. BidermanLLM circuit analyses are consistent across training and scale\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Document](https://dx.doi.org/10.52202/079017-1287),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/47c7edadfee365b394b2a3bd416048da-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px4.p1.1)\.
- Tsipraset al\.\(2019\)D\. Tsipras, S\. Santurkar, L\. Engstrom, A\. Turner, and A\. MadryRobustness may be at odds with accuracy\.InThe Seventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SyxAb30cY7)Cited by:[§1](https://arxiv.org/html/2609.35890#S1.p2.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Walkowiaket al\.\(2025\)P\. Walkowiak, M\. Klonowski, M\. Oleksy, and A\. JanzUnpacking robustness in inflectional languages: adversarial evaluation and mechanistic insights\.External Links:2505\.07856,[Link](https://arxiv.org/abs/2505.07856)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p3.1)\.
- Wanget al\.\(2023\)K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. SteinhardtInterpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NpsVSN6o4ul)Cited by:[§1](https://arxiv.org/html/2609.35890#S1.p1.1),[§2\.3](https://arxiv.org/html/2609.35890#S2.SS3.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
- Wanget al\.\(2022\)Z\. Wang, M\. Fredrikson, and A\. DattaRobust models are more interpretable because attributions look normal\.InProceedings of the 39th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 22625–22651\.External Links:[Link](https://proceedings.mlr.press/v162/wang22e.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.35890#S5.p1.1)\.
- Xhonneuxet al\.\(2024\)S\. Xhonneux, A\. Sordoni, S\. Günnemann, G\. Gidel, and L\. SchwinnEfficient adversarial training in LLMs with continuous attacks\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/0302fb83c62991efbccf0a003e4f5a92-Abstract-Conference.html)Cited by:[§2\.2](https://arxiv.org/html/2609.35890#S2.SS2.p1.2)\.
- Xu \(2026\)Y\. XuPattern selectivity is not task\-causal structure\.External Links:2606\.05378,[Link](https://arxiv.org/abs/2606.05378)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.35890#S1.p2.1)\.
- Zhanget al\.\(2025a\)L\. Zhang, W\. Dong, Z\. Zhang, S\. Yang, L\. Hu, N\. Liu, P\. Zhou, and D\. WangEAP\-GP: mitigating saturation effect in gradient\-based automated circuit identification\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.52202/085713-4748),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/d029c97ee0db162c60f2ebc9cb93387e-Abstract-Conference.html)Cited by:[Appendix A](https://arxiv.org/html/2609.35890#A1.SS0.SSS0.Px3.p1.1),[§2\.6](https://arxiv.org/html/2609.35890#S2.SS6.p1.1)\.
- Zhanget al\.\(2025b\)Q\. Zhang, Y\. Wang, J\. Cui, X\. Pan, Q\. Lei, S\. Jegelka, and Y\. WangBeyond interpretability: the gains of feature monosemanticity on model robustness\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/11822e84689e631615199db3b75cd0e4-Abstract-Conference.html)Cited by:[§5](https://arxiv.org/html/2609.35890#S5.p2.1)\.
## Appendix AExtended Related Work and Methodological Context
#### Conventional interpretations of robust models\.
The connection between robustness and interpretability has most often been studied through gradients and visual representations\. Input\-gradient regularization can improve both resistance to transferred adversarial examples and the human legibility of gradients\([Ross and Doshi\-Velez, 2018](https://arxiv.org/html/2609.35890#bib.bib7)\)\. Robustness has also been connected to the alignment of saliency maps with perceptually relevant directions\([Etmann et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib6);[Kim et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib5)\), while robust image classifiers learn representations and attributions that are often more human\-aligned\([Engstrom et al\., 2019](https://arxiv.org/html/2609.35890#bib.bib4);[Wang et al\., 2022](https://arxiv.org/html/2609.35890#bib.bib21)\)\. Conversely, training for robust interpretation can support robust classification\([Boopathy et al\., 2020](https://arxiv.org/html/2609.35890#bib.bib22);[Noack et al\., 2021](https://arxiv.org/html/2609.35890#bib.bib8)\)\. These findings provide evidence for a relationship between robustness and conventional interpretability, but each operationalizes interpretability through input\-output sensitivity, attribution, or representation quality rather than through the causal organization of a learned computation\.
#### Evaluating sparse autoencoders\.
SAEs can recover fine\-grained language\-model features and scale to large overcomplete dictionaries\([Huben et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib25);[Gao et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib26)\)\. Their evaluation nevertheless remains unsettled\. SAEBench finds that improvements on reconstruction and sparsity proxies do not reliably transfer to practical interpretability tasks\([Karvonen et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib27)\), while semantic evaluations based on polysemous words expose failures hidden by the usual reconstruction\-sparsity frontier\([Minegishi et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib28)\)\. Independent SAEs trained on the same model and data can learn substantially different dictionaries\([Paulo and Belrose, 2026](https://arxiv.org/html/2609.35890#bib.bib29)\), motivating explicit measurement of feature consistency across runs\([Song et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib30)\)\. A recent preprint further reports high noise or weak discriminability in several benchmark metrics\([Chanin, 2026](https://arxiv.org/html/2609.35890#bib.bib31)\)\. SAE representations can themselves be adversarially fragile: small input perturbations can manipulate concept interpretations without materially changing the underlying language\-model activations\([Li et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib10)\)\. These limitations do not make SAEs uninformative, but they prevent reconstruction, sparsity, or automated feature labels from serving as stand\-alone evidence of simpler computation\.
#### Automated circuit discovery\.
Automated Circuit Discovery \(ACDC\) recursively removes edges using activation\-patching effects and recovers compact subgraphs for established tasks\([Conmy et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib32)\)\. Edge Attribution Patching \(EAP\) replaces repeated interventions with a gradient\-based approximation, reducing circuit discovery to two forward passes and one backward pass\([Syed et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib17)\)\. EAP with integrated gradients improves faithfulness relative to local\-gradient EAP and demonstrates that overlap with a reference circuit can remain high even when the recovered circuit is behaviorally unfaithful\([Hanna et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib18)\)\. EAP\-GP modifies the integration path to mitigate saturation and further improve gradient\-based circuit identification\([Zhang et al\., 2025a](https://arxiv.org/html/2609.35890#bib.bib33)\)\. Sparse feature circuits extend causal graph discovery from attention heads and MLPs to SAE features\([Marks et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib34)\), while recent formal work defines circuit guarantees relative to explicit input domains and intervention semantics\([Hadad et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib35)\)\. Together, these methods make circuit recovery scalable, but their outputs remain conditional on the graph granularity, patching baseline, task metric, attribution estimator, threshold, and search procedure\. Our circuit\-size comparison therefore concerns the smallest graph recovered by a fixed algorithm at a specified faithfulness threshold, not a global minimum over every possible circuit\.
#### Circuit stability, selectivity, and self\-repair\.
Circuit analyses can identify similar functional algorithms across training checkpoints and model scales even when individual components change\([Tigges et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib16)\)\. Conversely, attention\-pattern selectivity need not identify the components that causally implement a task\([Xu, 2026](https://arxiv.org/html/2609.35890#bib.bib15)\)\. These results reinforce the distinction between observational structure and causal dependence\. Causal interventions introduce their own complication: ablating a component can change the behavior of downstream components, producing Hydra\-like compensation\([McGrath et al\., 2023](https://arxiv.org/html/2609.35890#bib.bib38)\)\. Self\-repair occurs across language\-model families and scales and can arise through several mechanisms, including LayerNorm scaling and reduced erasure of upstream contributions\([Rushing and Nanda, 2024](https://arxiv.org/html/2609.35890#bib.bib36)\)\. Copy suppression provides a detailed example in which an attention\-head motif explains both a component’s normal function and part of the compensation observed after ablation\([McDougall et al\., 2024](https://arxiv.org/html/2609.35890#bib.bib37)\)\. Because one\-at\-a\-time ablations can therefore underestimate distributed dependence, our causal analysis includes conditional co\-ablation to test whether apparently weak components become important when potential substitutes are simultaneously removed\.
#### Relation to adversarial mechanistic interpretability\.
[García\-Carrasco et al\. \(2024\)](https://arxiv.org/html/2609.35890#bib.bib39)propose a pipeline that first identifies a task circuit in a fixed GPT\-2 Small model, then generates adversarial samples and localizes the affected components\.[Walkowiak et al\. \(2025\)](https://arxiv.org/html/2609.35890#bib.bib11)combine adversarial attacks with EAP\-IG circuits to compare language and inflectional settings\. The feature\-level literature instead studies whether superposition or representation compression creates adversarial vulnerability\([Gorton and Lewis, 2025](https://arxiv.org/html/2609.35890#bib.bib1);[Bereska et al\., 2025](https://arxiv.org/html/2609.35890#bib.bib2);[Stevinson et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib3);[Gao et al\., 2026](https://arxiv.org/html/2609.35890#bib.bib24)\)\. Finally,[Bereska \(2024\)](https://arxiv.org/html/2609.35890#bib.bib40)lists reverse\-engineering robust models as an explicit research objective in a self\-published proposal, alongside investigating feature superposition and designing training methods that improve both robustness and interpretability; the proposal is not executed empirically\. Our work carries out that objective: we experimentally compare matched standard and adversarially trained language models across three non\-equivalent levels: sparse feature decomposition, SAE feature engagement, and fixed\-faithfulness circuit recovery\.
## Appendix BContinual\-training and attack details
### B\.1Initialization and optimization
Every run initializes fromopenai\-community/gpt2, revision607a30d783dfa663caf39e06633721c8d4cfcd7e\. The checkpoint has 124M parameters, a 50,257\-token vocabulary, and learned affine biases\. We update all parameters rather than using parameter\-efficient adapters\. The training stream is drawn from the Hugging Face parquet conversion of OpenWebText at revision433fe0f44ed7894fea29c08b3202aa348ccc6369\. The standard and adversarial runs receive the same examples in the same order\.
Table 7:Continual\-training hyperparameters shared by all seven runs\.The attack operates on the output of the token\-embedding lookup before positional embeddings are added\. In each micro\-batch, its absolute radius is
ϵabs=ϵrel1\|ℐ\|∑t∈ℐ‖E\(xt\)‖2,\\epsilon\_\{\\mathrm\{abs\}\}=\\epsilon\_\{\\mathrm\{rel\}\}\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{t\\in\\mathcal\{I\}\}\\\|E\(x\_\{t\}\)\\\|\_\{2\},\(8\)whereℐ\\mathcal\{I\}contains non\-special\-token positions\. Each token perturbation is initialized at a random point inside itsℓ2\\ell\_\{2\}ball\. We then perform ten normalized\-gradient ascent steps of size2\.5ϵabs/102\.5\\epsilon\_\{\\mathrm\{abs\}\}/10, projecting every token independently after each step\. Token 50256 is excluded from perturbation\. Training and qualification use the same attack implementation and parameterization\.
### B\.2Why continual training is used
We initially attempted a matched comparison trained from scratch on 9\.9B FineWeb tokens\. The standard model reached only 55\.4% IOI accuracy, compared with 99\.7% for the official GPT\-2 checkpoint on the same 1,000 prompts\. Thus, the target behavior did not reliably emerge within the available from\-scratch budget\. A one\-step embedding\-adversarial variant also showed catastrophic overfitting: its loss was 1\.20 under its training\-time attack but 24\.70 under PGD\-10, a 23\.5\-nat discrepancy\. These failures motivated continual training from a checkpoint that already possesses the target computations and the use of multi\-step PGD during both training and verification\.
## Appendix CEvaluation data and provenance
Table[8](https://arxiv.org/html/2609.35890#A3.T8)summarizes the OpenWebText partitions\. Model\-training data, robustness evaluation, SAE training, SAE evaluation, and the generic\-language\-modeling attribution control are positionally separated in the token stream\. Because this OpenWebText conversion has no native document identifiers, we synthesize identifiers from parquet file and row indices and record hashes of the resulting lists\. One document crosses each of two token\-level partition boundaries; the token streams themselves remain disjoint\.
Table 8:OpenWebText partitions used by the study\.The IOI set contains 1,000 prompts, split evenly between ABBA and BABA templates, generated with the official IOIDataset implementation at seed 42\. Names are restricted to single GPT\-2 tokens\. The GT set contains 1,000 balanced prompts generated with the official YearDataset implementation at seed 42\. To construct GT counterfactuals, we use the official “01\-dataset” bad\-sentence branch and require matching start\-year token counts; 973 prompts pass this parity filter\. The unused bad\-sentence branch of the original generator was patched to avoid a batching crash when a sampled counterfactual year is represented by one BPE token\. The sampling procedure, templates, and random\-number stream are otherwise unchanged\.
## Appendix DQualification sweep and behavioral trade\-off
Qualification metrics are computed for every sweep cell rather than stopping after the first passing candidate\. All adversarial runs clear the IOI retention floor, while four satisfy the masking sanity check\. The standard model’s large single\-step versus PGD\-10 gap is reported as a non\-robust reference and is not interpreted as a failed adversarial\-training run\.
The GT results reveal a graded behavioral trade\-off that IOI accuracy does not expose\. Paired prompt\-level differences from the standard model are−0\.063±0\.005\-0\.063\\pm 0\.005,−0\.087±0\.006\-0\.087\\pm 0\.006,−0\.117±0\.007\-0\.117\\pm 0\.007,−0\.112±0\.006\-0\.112\\pm 0\.006,−0\.171±0\.008\-0\.171\\pm 0\.008, and−0\.229±0\.010\-0\.229\\pm 0\.010fore050\-a02,e075\-a02,e100\-a02,e050\-a05,e075\-a05, ande100\-a05, respectively\. The original checkpoint scores 0\.706 on this prompt set, the continually trained standard model scores 0\.671, and the selected robust model scores 0\.442\. Because prompt sampling differs from the setup used to obtain previously published headline values, the paired within\-set difference is the primary GT comparison\.
## Appendix ESparse\-autoencoder implementation and statistics
We cache block\-8 residual activations once per model as bfloat16 shards, recording a hash of the corresponding model weights\. Each SAE is trained for 73,242 updates, one pass over 300M tokens, with batches of 4,096 tokens and Adam at learning rate4×10−44\\times 10^\{\-4\}and\(β1,β2\)=\(0\.9,0\.99\)\(\\beta\_\{1\},\\beta\_\{2\}\)=\(0\.9,0\.99\)\. Training uses a roughly 500,000\-token activation buffer, cross\-shard shuffling, and refill when half the buffer has been consumed\. No early stopping is used\.
Inputs are normalized per token bys=d/‖x‖2s=\\sqrt\{d\}/\\\|x\\\|\_\{2\}before encoding and returned to the original scale after decoding\. BatchTopK selects3232times the number of batch tokens activations across the batch\. Decoder columns are constrained to unit norm by radially projecting their gradients before each update and renormalizing after it\. AuxK uses 512 features, coefficient1/321/32, and defines a feature as dead if it has not been selected for 10M tokens\. At most one dead feature per seed \(Table[9](https://arxiv.org/html/2609.35890#A5.T9)\)\.
For downstream inference, batch\-relative selection is replaced by a fixed threshold\. The threshold is the mean, across held\-out evaluation batches, of the smallest positive activation selected by BatchTopK\. At seed 42, the thresholds are 0\.785 for the standard SAE and 0\.741 for the robust SAE\. This prevents the activation state of one prompt from depending on other prompts in the evaluation batch\.
We quantify measurement uncertainty by pairing residual positions between models and clustering observations into 1,024\-token sequence blocks\. Across 9,756 blocks, standard\-minus\-robust reconstruction MSE is0\.0323±0\.00070\.0323\\pm 0\.0007in original space and0\.0112±0\.00010\.0112\\pm 0\.0001in normalized space\. We separately retrain each SAE with seeds 42, 43, and 44\. The original\-space MSE ranges do not overlap:1\.1601±0\.00221\.1601\\pm 0\.0022for the standard activations and1\.1293±0\.00091\.1293\\pm 0\.0009for the robust activations\. Reconstruction statistics, SAE feature engagement \(Section[3\.3](https://arxiv.org/html/2609.35890#S3.SS3)\), and the ablation\-KL population comparison \(Appendix[F](https://arxiv.org/html/2609.35890#A6)\) use all three seeds\.
Table[9](https://arxiv.org/html/2609.35890#A5.T9)reports the per\-seed values underlying the aggregate statistics above\.
Table 9:Per\-seed SAE statistics\. Bold values indicate the robust condition\.Metricclean s42clean s43clean s44robust s42robust s43robust s44MSE, original residual space \(train\-mode\)1\.16241\.15791\.16011\.13011\.12951\.1284MSE, inference\-threshold mode1\.15461\.15111\.15261\.12621\.12651\.1254Final training loss \(normalized space\)0\.08770\.09020\.08690\.07680\.07700\.0788Dead latents010001Inference thresholdTT0\.7850\.7830\.7870\.7410\.7420\.743Table[10](https://arxiv.org/html/2609.35890#A5.T10)reports per\-seed candidate counts and feature counts at the 90% attribution\-mass threshold underlying Table[3](https://arxiv.org/html/2609.35890#S3.T3)and Section[3\.3](https://arxiv.org/html/2609.35890#S3.SS3)\. The robust SAE requires fewer candidates and fewer features at every seed, both tasks, with no exceptions\.
Table 10:Per\-seed SAE feature engagement\. Values are candidate count / features for 90% attribution mass, per seed\.
## Appendix FAttribution implementation and selectivity controls
For IOI, target attribution is summed over the first subject, indirect object, second subject, and final readout positions\. For GT, it is summed over the start\-year span and final readout\. A feature enters the candidate set if it activates at any intervention position on a clean or corrupted prompt\. Each model\-task analysis ranks its own candidate features; rankings are never merged across SAEs\. The ablation pool includes all features tied at the top\-100 boundary\.
### F\.1Non\-discriminative generic\-language\-modeling control
On held\-out OpenWebText, generic next\-token attributions are∼10−8\\sim 10^\{\-8\}against task attributions∼10−2\\sim 10^\{\-2\}\. A selectivity ratio against that control is uninformative: the denominator is negligible in both models\.
### F\.2Ablation\-KL population comparison
We zero\-ablate each of the top\-100 most\-attributed features and measureDKL\(p∥piabl\)D\_\{\\mathrm\{KL\}\}\(p\\\|p\_\{i\}^\{\\mathrm\{abl\}\}\)from the unspliced to the ablated output distribution, using a real intervention rather than a gradient\-based estimate: KL divergence has zero first derivative at the unablated distribution, so a linearized \(attribution\-patching\) estimate of this quantity is identically zero regardless of the true effect\. For GT, ablation is applied at the readout position, where pool features are already reliably active\. For IOI, an earlier readout\-only version left 98 of the top 100 features permanently inactive, since IOI features act predominantly at the name positions; we instead ablate at the union of positions in𝒯IOI\\mathcal\{T\}\_\{\\mathrm\{IOI\}\}where a given feature is active on that prompt, which resolves this and leaves all 100 features active for both models\.
Because the two SAEs are trained independently, a feature index carries no correspondence across models, so we do not pair features by index or form a per\-feature ratio against target attribution\. Instead, for each model and task we treat the resulting 100 ablation\-KL values as an unpaired population and compare the two models’ populations directly, across all three SAE seeds\. Table[11](https://arxiv.org/html/2609.35890#A6.T11)reports the full per\-seed detail\. Values exclude control\-set\-dead features \(zero effect by construction, since ablating an already\-inactive feature is a no\-op\); dead counts are reported separately per cell\. Medians are close between models on both tasks; the separation is concentrated in the 90th\-percentile tail\. We report this as a directional pattern only: with three seeds, per\-seed means are not separated by standard deviation, and one of three IOI seeds reverses direction\.
Table 11:Ablation\-KL effect, robust vs\. standard, excluding control\-set\-dead features\. Mean/median in units of10−310^\{\-3\}\.
## Appendix GCircuit recovery protocol
For each model\-task pair, we score all 32,491 edges of the 156\-node raw computational graph \(144 attention heads, 12 MLPs\) using EAP\-IG \(m=5m=5interpolation steps\) and, as a pilot, EAP\-GP, on identical prompt sets and counterfactuals: 200 IOI prompts \(100 ABBA / 100 BABA, name\-swap counterfactuals\) and 200 of the first 207 GT prompts scanned, surviving the start\-year parity filter \(collection stops at 200 passes, so this count is smaller than the full\-set 973\-of\-1000 in Appendix[C](https://arxiv.org/html/2609.35890#A3)\), ”01\-dataset” counterfactuals\. Circuit size is the smallest edge count, from a geometric gridk∈\{1,2,4,8,16,32,64,128,256,512,1000,2000,4000,8000,16000,32491\}k\\in\\\{1,2,4,8,16,32,64,128,256,512,1000,2000,4000,8000,16000,32491\\\}\(a denser tail past 512, not a continued power\-of\-two sequence, to avoid skipping the 85% crossing\), that reaches 85% of full\-graph faithfulness, together with the trapezoidal area under the full faithfulness curve on a log\-kkaxis as a finer\-grained discriminator\.
Table[12](https://arxiv.org/html/2609.35890#A7.T12)reports the rawF\(k\)F\(k\)values underlying Figure[2](https://arxiv.org/html/2609.35890#S3.F2), IOI, EAP\-IG, per Equation[7](https://arxiv.org/html/2609.35890#S2.E7)\. Belowk=1000k=1000, several entries are negative, meaning that subgraph performs worse than the fully corrupted baseline; atk=128k=128standard leads by its largest margin in the curve \(F=0\.13F=0\.13versus−0\.09\-0\.09robust\), still well short of the 85% threshold\. Negative values at smallkklikely reflect\|score\|\|\\mathrm\{score\}\|ranking admitting edges with negative effect\.
Table 12:RawF\(k\)F\(k\), EAP\-IG\. IOI underlies Figure[2](https://arxiv.org/html/2609.35890#S3.F2)\. GT underlies the 95% sentence in Section[3\.4](https://arxiv.org/html/2609.35890#S3.SS4); GT is not competence\-matched\.
GT\-standard’s first grid point reaching 95% isk=8000k=8000:k=2000k=2000\(0\.94\) andk=4000k=4000\(0\.87, non\-monotonic\) both fall short\.
EAP\-GP is a recently proposed variant using integrated gradient paths intended to mitigate the saturation effect known to affect EAP\-IG\. No public reference implementation existed at the time of this work; our implementation follows the method description, with one identified approximation – a single shared gradient path computed at the input activations and reused for all downstream edges, rather than one path per node, matching the roughly 5×\\timesper\-step cost the source paper reports relative to EAP\-IG\. EAP\-GP did not outperform EAP\-IG on our graph for either task or model: area under the faithfulness curve was negative for EAP\-GP on IOI for both models \(standard:−0\.97\-0\.97; robust:−0\.70\-0\.70\), against positive AUC for EAP\-IG \(standard: 0\.87; robust: 0\.60\)\. We select the primary method per task by comparing AUC between EAP\-IG and EAP\-GP before inspecting the standard\-versus\-robust comparison, and report EAP\-IG throughout the main text on this basis\.
Figure 4:Greater\-Than edge\-native faithfulness,F\(k\)F\(k\), EAP\-IG \(normalized\-faithfulness\-score, or NFS, normalization\)\. AUC is an unnormalized trapezoid ofF\(k\)F\(k\)overlnk\\ln k\(maximum≈10\.4\\approx 10\.4ifF=1F=1throughout the grid\); GT’s curve stays near zero or above and rises early, while IOI’s curve dips toF≈−1\.08F\\approx\-1\.08before climbing \(Table[12](https://arxiv.org/html/2609.35890#A7.T12)\), which is why GT and IOI AUCs sit on different numeric scales and are not directly comparable\. Not competence\-matched between models; not used as evidence for the main result \(Section[3\.4](https://arxiv.org/html/2609.35890#S3.SS4)\)\.Figure 5:EAP\-GP edge\-native faithfulness, both tasks\. Negative AUC on IOI for both models; EAP\-IG is used for the headline result on this basis\.Top\-kkedge selection can yield a disconnected subgraph, which could in principle make a circuit look less faithful for reasons unrelated to the underlying attribution method’s quality\. As a check, we rebuilt circuits with a connectivity\-preserving greedy edge selector in place of plain top\-kkand recomputed both methods’ faithfulness curves under it\. The verdict does not change: EAP\-IG beats EAP\-GP under both builders, on both tasks and models\. The greedy builder improves both methods’ AUC, since connected circuits are generally more faithful, but it improves EAP\-IG more than EAP\-GP on IOI \(\+0\.81 versus \+0\.48\), so EAP\-IG’s advantage widens rather than narrows under the connectivity\-preserving builder\. The same holds on Greater\-Than, our secondary check: EAP\-IG’s margin over EAP\-GP moves from\+0\.10\+0\.10/\+0\.56\+0\.56\(standard/robust, top\-kk\) to\+0\.13\+0\.13/\+0\.54\+0\.54\(greedy\), essentially unchanged\. Top\-kkdisconnection is therefore not the source of EAP\-GP’s underperformance on either task\.
Conditional co\-ablation \(CoAx\) recovers backup components after primary\-circuit ablation, following the scoring formula of[Gong et al\. \(2026\)](https://arxiv.org/html/2609.35890#bib.bib45)111Propositions 2–3\., and tests whether apparently weak components become necessary once substitutes are removed, guarding against self\-repair understating causal dependence\. This step costs a fixed295295forward passes for backup recovery plus2222for the backup\-validity gate, per model, task, and attribution method – constant by construction \(a function of the candidate set size, not of which circuit is being recovered\), so it does not itself discriminate between the standard and robust conditions; the resulting circuit’s faithfulness curve does\. Node\-level faithfulness on the primary\-plus\-backup set remains far below 85% for both models \(10–14 nodes\)\. It is not a second 90/95% edge\-level test\.
## Appendix Hα\\alpha\-ReQ: representation eigenspectrum decay
α\\alpha\-ReQ is a representational\-simplicity measure independent of any trained SAE, used to test whether the dose\-response in Section[3\.5](https://arxiv.org/html/2609.35890#S3.SS5)is an artifact of the SAE family rather than a property of the underlying representation\. For each residual\-stream layer \(every block’sresid\_post, plus the layer entering block 8 used for the SAE\), we accumulate the centered activation covariance over a fixed evaluation token set, take its eigenvaluesλ1≥…≥λd\\lambda\_\{1\}\\geq\\ldots\\geq\\lambda\_\{d\}\(d=768d=768\), and fitλi∼i−α\\lambda\_\{i\}\\sim i^\{\-\\alpha\}by OLS on a log\-log scale over ranks 6–691 \(dropping the top\-5 spikes, which reflect a small number of dominant directions rather than the bulk spectrum, and the bottom 10% floor, where eigenvalues approach numerical noise\), applied identically to every model and layer\. A more compressed representation decays faster, giving a largerα\\alpha; we report this exponent at the model level \(mean across layers\) and at the SAE\-hook layer specifically\. Participation ratio,\(∑iλi\)2/∑iλi2\(\\sum\_\{i\}\\lambda\_\{i\}\)^\{2\}/\\sum\_\{i\}\\lambda\_\{i\}^\{2\}, is a second summary of the same eigenvalue spectrum: unlikeα\\alpha, it requires no fitting or rank\-range choice, so agreement between the two – both moving the same direction across the sweep – is evidence theα\\alpha\-ReQ trend is not an artifact of the fit range\. HT\-SR weight\-spectrum alpha \(WeightWatcher\) was also measured and found flat across the full sweep \(2\.923–2\.934\) – the continual\-training update does not move the weight spectrum measurably, so this metric is not used further\.
## Appendix ICorpus ablation: FineWeb
The corpus\-ablation pair uses the same continual\-training recipe, budget, and qualification gate as the primary standard\-robust pair \(Section[2\.2](https://arxiv.org/html/2609.35890#S2.SS2)–[2\.3](https://arxiv.org/html/2609.35890#S2.SS3)\), with FineWeb substituted for OpenWebText as the training and SAE corpus\. SAE architecture, training budget, and evaluation protocol match Axis 1 exactly\. Both arms pass the IOI qualification floor and the masking sanity check before comparison\.
## Appendix JReproducibility checks and limitations
All experiments use seed 42 for model initialization where applicable, data order, PGD initialization, and primary SAE training; SAE reconstruction and attribution analyses are additionally replicated at seeds 43 and 44\. The implementation is covered by 80 CPU tests on small models, including token accounting, attack constraints, value identity of the straight\-through splice, attribution\-patching linearity against finite differences, BatchTopK selection semantics, KL\-control computation, and provenance checks\.
The main limitations are the use of one base architecture, one adversarial training regime, and one analyzed residual\-stream site\. SAE reconstruction and SAE feature engagement use three seeds; the ablation\-KL population comparison \(Appendix[F](https://arxiv.org/html/2609.35890#A6)\) also uses three seeds; circuit recovery \(Section[3\.4](https://arxiv.org/html/2609.35890#S3.SS4)\) uses a 200\-prompt subsample, smaller than the 1,000\-prompt sets used elsewhere, reflecting its higher per\-prompt cost\. On Greater\-Than the robust model scores lower than the standard model on the same prompts \(PD 0\.442 vs\. 0\.671; paired difference−0\.229±0\.010\-0\.229\\pm 0\.010\)\. We therefore do not treat GT circuit size as a competence\-matched test\. OpenWebText identifiers are synthesized rather than native\. Our EAP\-GP implementation follows the method description in the absence of a public reference implementation\.相似文章
通过稀疏电路理解神经网络
OpenAI 研究人员提出了一种训练稀疏神经网络的方法,通过强制大部分权重为零使其更易于解释,从而发现能够解释模型行为的小型解耦电路,同时保持性能。这项工作旨在推进机制可解释性,作为对稠密网络事后分析的补充,并支持 AI 安全目标。
架构而非规模:大语言模型中的电路局部化
本文挑战了“随着模型规模扩大,机制可解释性变得愈发困难”的假设,表明架构(特别是分组查询注意力与多头注意力之间的差异)对电路局部化和稳定性的影响比参数量更为关键。
@dair_ai: 值得一读的新论文。GPT-5.4 nano 加上 critic-comparator 编排循环在 SWE-bench Verified 上达到 76.4%,匹配…
一篇新论文表明,使用一个弱模型,通过 k=8 个提议和 critic-comparator 选择循环,可以在 SWE-bench Verified 上匹配前沿模型的性能,达到 76.4% 的准确率。关键见解是,正确的补丁通常已经存在于弱模型的前 k 个候选补丁中,挑战在于如何利用执行验证进行有效选择。
可解释却脆弱?概念瓶颈模型在几何-语义扰动下的鲁棒性
本文提出一种基于生成器的评估框架,用于研究概念瓶颈模型和基于原型的模型在几何-隐变量扰动与语义-概念扰动下的响应差异,论证了可解释性本身并不天然带来鲁棒性,而只是将敏感性重新分配到了不同的扰动空间中。
灾难性遗忘的机制起源:为什么RL比SFT更好地保留电路?
本文研究了LLM中灾难性遗忘的机制起源,发现强化学习比监督微调更好地保留了内部计算电路,从而减少了对先前能力的遗忘。