The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

arXiv cs.LG Papers

Summary

This paper re-derives activation patching from causal mediation analysis, revealing that the natural indirect effect (NIE) captures not only a component's causal effect but also interaction effects with other components. It demonstrates these hidden interactions in the GPT-2 IOI circuit and argues that they are a diagnostic tool rather than a nuisance.

arXiv:2606.27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.
Original Article
View Cached Full Text

Cached at: 06/29/26, 05:23 AM

# The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
Source: [https://arxiv.org/html/2606.27510](https://arxiv.org/html/2606.27510)
Sankaran Vaidyanathan University of Massachusetts Amherst sankaranv@cs\.umass\.edu &David Arbour Adobe Research arbour@adobe\.com &Aaron Mueller Boston University amueller@bu\.edu &Scott Niekum University of Massachusetts Amherst sniekum@cs\.umass\.edu &David Jensen University of Massachusetts Amherst jensen@cs\.umass\.edu

###### Abstract

Activation patching is the primary tool in mechanistic interpretability\. It attributes causal responsibility for a model behavior to each of its individual components by estimating itsnatural indirect effect \(NIE\)\. Re\-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component\. It also containsinteraction effects \(INT\)that measure how much the component’s causal effect itself depends on the state of other components in the model\. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes\. We demonstrate these failure modes in the GPT\-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores\. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher\-order group interactions\. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies\. Its individual and group\-level magnitude and sign signal when causal conclusions are prompt\-dependent, and when greedy NIE\-based component ranking will miss mechanisms only discoverable through combinatorial search\.

## 1Introduction

Mechanistic interpretability aims to identify which internal components in a neural network implement a given behavior or capability, and to characterize the causal role of each identified component\(Olahet al\.,[2020](https://arxiv.org/html/2606.27510#bib.bib38); Elhageet al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib39)\)\. Component localization is grounded incausal mediation analysis, particularly in isolating how much of the total causal effect flows through each component by computing itsnatural indirect effect \(NIE\)\(Pearl,[2001](https://arxiv.org/html/2606.27510#bib.bib28)\)\. In neural networks,activation patching\(Viget al\.,[2020](https://arxiv.org/html/2606.27510#bib.bib25); Geigeret al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib30)\)operationalizes the NIE by substituting a component’s activations with counterfactual values, interpreting the resulting change in output as that component’s causal contribution\. Activation patching and its derivatives\(Goldowsky\-Dillet al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib40); Syedet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib35); Hannaet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib47)\)are the field’s de facto standard instruments for causal attribution, ranking components, and evaluating circuits\.

However, documented failures of activation patching appear to contradict its claim as a measure of causal importance, most notably by missing components that are functionally important for a given task\(Mueller,[2024](https://arxiv.org/html/2606.27510#bib.bib48)\)\. The circuit for indirect object identification \(IOI\)\(Wanget al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib34)\)is the canonical example of such false negatives; it contains backup name mover heads that only show significant causal effects when ablating all other components that performed the same function\. These were only found through combinatorial search and pre\-registered hypotheses about the components’ function—an approach that does not scale and is no longer standard practice\. Circuit evaluations have also encountered other counterintuitive issues, including high variability across ablation methods and prompt distributions\(Milleret al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib27); Mélouxet al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib5)\), non\-monotonicity\(Shiet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib4)\), and low\-scoring circuits that overlap highly with ground truth\(Hannaet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib47)\)\. While some of these issues have been explained by appealing to the complexity of the underlying model, the assumptions of the measurement procedure themselves have rarely been questioned\.

SincePearl \([2001](https://arxiv.org/html/2606.27510#bib.bib28)\)proposed the NIE, the causal inference literature has identified failure cases where the NIE does not correctly isolate the effect through a specific component\(Andrews and Didelez,[2021](https://arxiv.org/html/2606.27510#bib.bib50); VanderWeele,[2014](https://arxiv.org/html/2606.27510#bib.bib36)\)\. A key structural condition that guarantees such failure is the presence of recanting witnesses\(Avinet al\.,[2005](https://arxiv.org/html/2606.27510#bib.bib51)\): variables that lie on both the target path and a bypass route around the mediator\. The transformer residual stream guarantees this condition for every mediator by construction: the skip connection at each layer routes information directly from earlier components to later ones, around whatever activation is being patched\(Elhageet al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib39)\)\. This gives us principled reason to suspect that the NIE may not isolate component\-specific contributions in transformers\.

We prove that the NIE in a transformer containsinteraction effects \(INT\)\(VanderWeele,[2013](https://arxiv.org/html/2606.27510#bib.bib41)\), which measure how much a component’s causal effect itself depends on the state of the rest of the model\. INT arises in transformers because activation patching holds one component fixed at its clean\-input values while the rest of the model runs on a counterfactual input; the output is then a nonlinear function of both, and the portion of the NIE that cannot be attributed to either alone is INT\. Activation patching computesNIE=PIE\+INT\\operatorname\{NIE\}=\\operatorname\{PIE\}\+\\mathrm\{INT\}, where PIE is the pure indirect effect: the causal effect measured solely through paths that contain the component, which the NIE was assumed to be recovering\. INT has been present in every activation patching study\.

In this paper, we study the theoretical properties and empirical consequences of interaction effects in transformer activation patching\. Our key findings include the following:

1. \(a\)NIE and PIE induce substantially different rankings over component importance across models and datasets \([Figure˜1](https://arxiv.org/html/2606.27510#S3.F1)\), with rank correlation as low asρ=0\.51\\rho=0\.51on the GPT\-2 IOI task\.
2. \(b\)INT scales with the distance between clean and patched component activations, and is negligible when the downstream map is locally affine \([Theorem˜3\.2](https://arxiv.org/html/2606.27510#S3.Thmtheorem2)\)\. Empirically, INT is very small in tasks where clean and counterfactual prompts are highly similar, and in compiled Tracr\(Lindneret al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib8)\)models used in interpretability benchmarks\(Guptaet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib23); Muelleret al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib53)\)\.
3. \(c\)In multi\-component patching estimands such as faithfulness scores, INT decomposes combinatorially into individual, pairwise and group\-level interactions \([Theorem˜4\.1](https://arxiv.org/html/2606.27510#S4.Thmtheorem1)\)\. The variance of these terms provides a mechanistic explanation for the prompt\-level instability of faithfulness scores thatMilleret al\.\([2024](https://arxiv.org/html/2606.27510#bib.bib27)\)documented \([Figure˜2](https://arxiv.org/html/2606.27510#S4.F2)\)\.
4. \(d\)Interaction effects explain known failure cases of ranking by activation patching scores in the GPT\-2 IOI circuit \([Section˜5\.1](https://arxiv.org/html/2606.27510#S5.SS1)\)\. Backup name mover heads are known to have low NIE scores since other heads perform the same function; this suppression is represented as negative INT\.
5. \(e\)Removing INT and ranking by PIE introduces ranking failures of its own \([Section˜5\.2](https://arxiv.org/html/2606.27510#S5.SS2)\)\. On the IOI task in GPT\-2, PIE is near\-zero for context\-specific components such as Duplicate Token and Previous Token heads; their importance depends on other heads being simultaneously active\.

INT is a fundamental feature of causal mediation, growing in complexity with the size and nonlinearity of modern transformers\. Its presence in the NIE propagates downstream to ranking and evaluation procedures, and it provides a unifying theory for several known pathologies in the field\. While it does not fully eliminate the need for combinatorial search to resolve these pathologies, INT is measurable at many levels of granularity and enables practitioners to identify when their conclusions can be trusted without that search\. That is not a limitation of activation patching; it is a verifiable statement of its assumptions, and a new lens on model mechanisms that single\-component effects cannot see\.

††footnotetext:Code available at:[https://github\.com/sankaranv/mech\-interp\-mediation\-analysis](https://github.com/sankaranv/mech-interp-mediation-analysis)
## 2Background

Causal mediation analysis \(CMA\) decomposes the total causal effect of a treatmentAAon an outcomeYYinto contributions from specific causal pathways\. For treatmentaaand baselinea′\{a^\{\\prime\}\}, the total causal effect isTE⁡\(a,a′\)=𝔼​\[Y​\(a\)\]−𝔼​\[Y​\(a′\)\]\\operatorname\{TE\}\(a,\{a^\{\\prime\}\}\)=\\mathbb\{E\}\[Y\(a\)\]\-\\mathbb\{E\}\[Y\(\{a^\{\\prime\}\}\)\]\. When amediatorMMlies on some but not all paths fromAAtoYY, the natural indirect effect \(NIE\) measures how much of the total effect passes throughMM\.

###### Definition 2\.1\(Natural Indirect Effect\)\.

For treatmentaaand baselinea′\{a^\{\\prime\}\}, the natural indirect effect\(Pearl,[2001](https://arxiv.org/html/2606.27510#bib.bib28)\)through mediatorMMis

NIE⁡\(a,a′\)=𝔼​\[Y​\(a,M​\(a\)\)\]−𝔼​\[Y​\(a,M​\(a′\)\)\]\\operatorname\{NIE\}\(a,\{a^\{\\prime\}\}\)=\\mathbb\{E\}\[Y\(a,M\(a\)\)\]\-\\mathbb\{E\}\[Y\(a,M\(\{a^\{\\prime\}\}\)\)\]whereY​\(a,M​\(a′\)\)Y\(a,M\(\{a^\{\\prime\}\}\)\)is the outcome under treatmentaawithMMfixed at its baseline valueM​\(a′\)M\(\{a^\{\\prime\}\}\)\.

The NIE is intended to isolate the effect throughMMand block all paths that bypass it\.Avinet al\.\([2005](https://arxiv.org/html/2606.27510#bib.bib51)\)formalizes this target as thepath\-specific effect \(PSE\): the causal effect propagated along a designated subgraph of the model\.

In transformers, mediators are typically activations of model components such as attention heads, MLPs, or sparse autoencoder latents\(Muelleret al\.,[2026](https://arxiv.org/html/2606.27510#bib.bib6)\)\. Because the residual stream introduces a skip connectionxℓ−1→xℓx\_\{\\ell\-1\}\\to x\_\{\\ell\}at each layer, every mediator has a bypass path\. The computation of NIE in transformers goes by various names in the literature such as activation patching\(Heimersheim and Nanda,[2024](https://arxiv.org/html/2606.27510#bib.bib24)\), causal tracing\(Menget al\.,[2022](https://arxiv.org/html/2606.27510#bib.bib31)\), and interchange intervention\(Geigeret al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib30)\); we usepatchingthroughout this paper\.

The treatmentaaand baselinea′\{a^\{\\prime\}\}are conventionally called thecleanandcorruptedinputs\. This distinction gives rise to two patching estimators\.

###### Definition 2\.2\(Patching Estimators\)\.

LetMMbe a set of nodes with treatment \(clean\) activationM​\(a\)M\(a\)and baseline \(corrupted\) activationM​\(a′\)M\(\{a^\{\\prime\}\}\)\. Activation patching yields the following estimators:

1. \(a\)*Denoising*: run the model on inputa′\{a^\{\\prime\}\}, then replaceM​\(a′\)M\(\{a^\{\\prime\}\}\)withM​\(a\)M\(a\)\.
2. \(b\)*Noising*: run the model on inputaa, then replaceM​\(a\)M\(a\)withM​\(a′\)M\(\{a^\{\\prime\}\}\)\.

The two operations differ in which input provides the baseline: denoising restores components to their clean values within a forward pass under the corrupted input; noising corrupts the components within a forward pass under the clean input\. Noising is currently the default approach for measuring component importance in transformers\(Conmyet al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib52); Hannaet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib47); Markset al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib1)\)\.

## 3Interaction Effects in Transformers

Activation patching measures each component’s causal contribution by computing its natural indirect effect, but the NIE is not a function of that component’s activation alone\. Every component operates alongside a bypass: the residual skip connection and same\-layer components route information around it to downstream layers by construction\(Elhageet al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib39)\)\. While patching the component’s activation changes what it contributes to the residual stream, it leaves the bypass on its original input, and downstream nonlinearities couple both contributions in a way that is difficult to attribute to either alone\(Avinet al\.,[2005](https://arxiv.org/html/2606.27510#bib.bib51); Andrews and Didelez,[2021](https://arxiv.org/html/2606.27510#bib.bib50)\)\. Proposition[3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)identifies how the bypass factors into the causal effect that the NIE was assumed to isolate, and what it actually measures\.

###### Proposition 3\.1\(Single\-Mediator Interaction\)\.

Consider a single componentMMat layerℓ\\ell, and letBMB\_\{M\}be the set of edges that bypassMMin the computation graph\.

1. \(a\)Denoising computesPIE=Y​\(a′,M​\(a\)\)−Y​\(a′\)\\operatorname\{PIE\}=Y\(a^\{\\prime\},M\(a\)\)\-Y\(a^\{\\prime\}\), the pure indirect effect, which is the path\-specific effect through the subgraph in which the bypass edgesBMB\_\{M\}are blocked\.
2. \(b\)Noising computesNIE=PIE\+INT\\operatorname\{NIE\}=\\operatorname\{PIE\}\+\\mathrm\{INT\}, where INT=Y​\(a\)−Y​\(a′,M​\(a\)\)−Y​\(a,M​\(a′\)\)\+Y​\(a′\)\\mathrm\{INT\}=Y\(a\)\-Y\(a^\{\\prime\},M\(a\)\)\-Y\(a,M\(a^\{\\prime\}\)\)\+Y\(a^\{\\prime\}\)\(1\)is themediated interactionbetweenMMand its bypassBMB\_\{M\}\.
3. \(c\)The total effect decomposes asTE=PIE\+PDE\+INT\\mathrm\{TE\}=\\operatorname\{PIE\}\+\\mathrm\{PDE\}\+\\mathrm\{INT\}, wherePDE=Y​\(a,M​\(a′\)\)−Y​\(a′\)\\mathrm\{PDE\}=Y\(a,M\(a^\{\\prime\}\)\)\-Y\(a^\{\\prime\}\)is the pure direct effect, i\.e\., the path\-specific effect through the subgraph in which the edge fromMMto the residual stream is blocked\.

Proposition[3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)establishes that noising and denoising are not symmetrical operations, and actually estimate distinct causal quantities, NIE and PIE\. The PIE is referred to as “pure” because it corresponds to the surgical removal of bypass edges; it is the causal effect that exclusively flows through the given component\. The NIE additionally contains aninteractionterm\(VanderWeele,[2013](https://arxiv.org/html/2606.27510#bib.bib41)\)that captures how much the component’s contribution to the outcome depends on what the bypass is simultaneously carrying\. For instance, a positive interaction indicates that the rest of the model amplifies the effect of the component\. Ignoring INT does not eliminate it\. Activation patching absorbs it entirely into the estimate it computes, thereby inflating or deflating it by the full magnitude of INT\.

Interaction is a recognized primary quantity in causal mediation analysis\(VanderWeele,[2015](https://arxiv.org/html/2606.27510#bib.bib37)\), and expected to arise whenever treatment and mediator share downstream causal pathways\. For example, in a study of the interaction between age and alcohol consumption in their association withH\. pyloriinfection, alcohol consumption was found to be causative for younger individuals but preventative for older adults\(Paunioet al\.,[1994](https://arxiv.org/html/2606.27510#bib.bib17)\): the direction of the causal effect depended entirely on the state of a second variable\. The causal inference literature documents a rich taxonomy of such interaction patterns\(VanderWeele,[2019](https://arxiv.org/html/2606.27510#bib.bib16)\), of which sign reversal is only the most dramatic\.Heimersheim and Nanda \([2024](https://arxiv.org/html/2606.27510#bib.bib24)\)identified one such pattern in transformers, attributing the noising\-denoising gap to AND/OR circuit topology\. Interaction is the underlying phenomenon; in[Section˜5](https://arxiv.org/html/2606.27510#S5)we show that the sign and magnitude of INT indicate a wider range of functional roles than AND/OR circuits alone\.

#### INT and Ranking Divergence\.

The practical consequence ofINT≠0\\mathrm\{INT\}\\neq 0is that noising and denoising rank components differently, as shown in[Figure˜1](https://arxiv.org/html/2606.27510#S3.F1)\(left\)\. We measure the Spearman rank correlationρ​\(NIE¯,PIE¯\)\\rho\(\\overline\{\\operatorname\{NIE\}\},\\overline\{\\operatorname\{PIE\}\}\)across all attention heads in GPT\-2 small, Pythia\-70m, and Qwen2\.5\-0\.5B for five tasks and corruption schemes; all datasets are taken from the repository ofHannaet al\.\([2024](https://arxiv.org/html/2606.27510#bib.bib47)\)\. Models were selected to match the original circuit papers for each task\. We use exact activation patching rather than fast approximations such as attribution patching or integrated gradients, since our goal is to study the fundamental properties of NIE and PIE directly\.

Rank agreement between NIE and PIE tracks how semantically similar the clean and counterfactual prompts are\. At the high end,ρ=0\.989\\rho=0\.989for SVA\(Finlaysonet al\.,[2021](https://arxiv.org/html/2606.27510#bib.bib18)\),0\.8260\.826for Gender Bias\(Viget al\.,[2020](https://arxiv.org/html/2606.27510#bib.bib25)\), and0\.9190\.919for IOI\(Wanget al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib34)\)with symmetric corruption where the name tokens are swapped\. These are datasets where the clean and counterfactual prompts are extremely close to each other; for example, an SVA counterfactual changes a single word from singular to plural\. On the other hand,ρ=0\.509\\rho=0\.509for IOI with pABC corruption and0\.5170\.517for Greater Than\(Hannaet al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib45)\), which use counterfactuals that are more substantially different from the clean prompt\.

INT as defined in Equation[1](https://arxiv.org/html/2606.27510#S3.E1)is a discrete quantity requiring no additional assumptions about the downstream function\. In modern transformers, INT admits a second\-order approximation that reveals its relationship to the prompt distance: it scales with the magnitude of the perturbations to the component and its bypass, so when clean and counterfactual prompts are close, INT is suppressed\.

![Refer to caption](https://arxiv.org/html/2606.27510v1/x1.png)

![Refer to caption](https://arxiv.org/html/2606.27510v1/x2.png)

Figure 1:Left\. Spearman rank correlation between mean NIE and mean PIE per head across 144 heads in GPT\-2 small, Pythia\-70M, and Qwen2\.5\-0\.5B\. Datasets where INT is large relative to NIE show lower rank agreement between the two estimators\.Right\. INT as a function of the L2 distance between the mediator’s patched activation and its clean value for 6 heads from the IOI circuit with different roles\. INT grows linearly with patch distance, consistent with[Theorem˜3\.2](https://arxiv.org/html/2606.27510#S3.Thmtheorem2)\.###### Theorem 3\.2\(Second\-Order Approximation of INT\)\.

LetδM=M​\(a\)−M​\(a′\)\\delta\_\{M\}=M\(a\)\-M\(a^\{\\prime\}\)andδB=BM​\(a\)−BM​\(a′\)\\delta\_\{B\}=B\_\{M\}\(a\)\-B\_\{M\}\(a^\{\\prime\}\)be the changes inMMand its bypass induced by changing the input fromaatoa′\{a^\{\\prime\}\}\. Letzℓ​\(a′\)z\_\{\\ell\}\(\{a^\{\\prime\}\}\)be the residual stream at layerℓ\\ellundera′a^\{\\prime\}, and letΨ\\Psibe the downstream function fromzℓ​\(a′\)z\_\{\\ell\}\(\{a^\{\\prime\}\}\)to the output metricYY\. WhenΨ\\Psiis twice differentiable, the mediated interaction is the mixed second directional derivative ofΨ\\Psiatzℓ​\(a′\)z\_\{\\ell\}\(\{a^\{\\prime\}\}\)in the directions of the perturbationsδM\\delta\_\{M\}andδB\\delta\_\{B\},

INT≈D2​Ψ​\(zℓ​\(a′\)\)​\[δM,δB\]\\mathrm\{INT\}\\;\\approx\\;D^\{2\}\\Psi\\big\(z\_\{\\ell\}\(\{a^\{\\prime\}\}\)\\big\)\[\\delta\_\{M\},\\,\\delta\_\{B\}\]whereD2​Ψ​\(zℓ​\(a′\)\)​\[⋅,⋅\]D^\{2\}\\Psi\(z\_\{\\ell\}\(\{a^\{\\prime\}\}\)\)\[\\cdot,\\cdot\]is the Hessian bilinear form\. The approximation is zero under the following sufficient conditions:

1. \(i\)δM=0\\delta\_\{M\}=0orδB=0\\delta\_\{B\}=0, i\.e\., the mediator or the bypass is unaffected by the change in input\.
2. \(ii\)∇2Ψ​\(zℓ​\(a′\)\)=0\\nabla^\{2\}\\Psi\\big\(z\_\{\\ell\}\(\{a^\{\\prime\}\}\)\\big\)=0, i\.e\., the downstream map is locally affine\.
3. \(iii\)D2​Ψ​\(zℓ​\(a′\)\)​\[δM,δB\]=0D^\{2\}\\Psi\\big\(z\_\{\\ell\}\(\{a^\{\\prime\}\}\)\\big\)\[\\delta\_\{M\},\\,\\delta\_\{B\}\]=0withD2​Ψ≠0D^\{2\}\\Psi\\neq 0, i\.e\. the perturbations either have no shared components in the curved directions ofΨ\\Psi, or their contributions in shared directions cancel\.

Figure[1](https://arxiv.org/html/2606.27510#S3.F1)\(right\) verifies the approximation’s prediction on different heads from the IOI circuit with different ground\-truth functional roles\. As the mediator activation is interpolated toward its counterfactual value, INT grows linearly with perturbation distance‖δM‖\\\|\\delta\_\{M\}\\\|across all circuit heads\. The slope sign encodes functional role, negative for L9H9 \(Name Mover\) and positive for L8H6 \(S\-Inhibition\); we expand on this pattern in Section[5](https://arxiv.org/html/2606.27510#S5)\.Viget al\.\([2020](https://arxiv.org/html/2606.27510#bib.bib25)\)reported an approximately linear decomposition of activation patching for Gender Bias and used it to argue for a no\-interaction assumption\. Our analysis explains why: their profession\-swap corruption produces semantically close clean\-corrupt pairs with small‖δM‖\\\|\\delta\_\{M\}\\\|, satisfying condition \(i\) of[Theorem˜3\.2](https://arxiv.org/html/2606.27510#S3.Thmtheorem2)\. More generally,‖δM‖\\\|\\delta\_\{M\}\\\|is small and INT is suppressed when clean and counterfactual prompts are semantically close\.

Condition \(ii\) holds exactly for Tracr\(Lindneret al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib8)\)compiled transformers, which are piecewise affine by construction and used in interpretability benchmarks\(Guptaet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib23); Muelleret al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib53)\)\. INT vanishes in such settings, meaning these benchmarks likely do not reflect the INT conditions present in real transformers\. The magnitude of INT is determined by the evaluation conditions—the model class and counterfactual prompt design—not by the circuit alone\.

## 4Multi\-Component Interactions

Single\-component patching aims to measure causal attribution for each head in isolation, but nearly every mechanistic interpretability study then uses those scores to select or ablate*sets*of components\. Circuit faithfulness evaluation, grouping nodes with a common function, and combinatorial knockouts are all examples of multi\-component patching operations in common practice\. We show that in this setting, interaction effects appear not just at every patched component individually, but also among every pair and subset of components as well — meaning that individual patching scores cannot predict what happens when those components are patched together\.

###### Theorem 4\.1\(Multi\-Component Interaction\)\.

For a subset of model components𝒞⊆ℳ\\mathcal\{C\}\\subseteq\\mathcal\{M\}, letY​\(𝒞\)Y\(\\mathcal\{C\}\)denote the output of the multi\-component noising estimator that runs the model on treatmentaaand patchesHi↦Hi​\(a′\)H\_\{i\}\\mapsto H\_\{i\}\(a^\{\\prime\}\)for allHi∈𝒞H\_\{i\}\\in\\mathcal\{C\}, withY​\(∅\)=Y​\(a\)Y\(\\emptyset\)=Y\(a\)\. The causal effect computed by the multi\-component noising estimator decomposes as:

Y​\(∅\)−Y​\(𝒞\)=∑M∈𝒞\[Y​\(∅\)−Y​\(M\)\]⏟NIE​\(M\)\+∑𝒯⊆𝒞\|𝒯\|≥2xINT​\(𝒯\),Y\(\\emptyset\)\-Y\(\\mathcal\{C\}\)\\;=\\;\\sum\_\{M\\in\\mathcal\{C\}\}\\underbrace\{\\bigl\[Y\(\\emptyset\)\-Y\(M\)\\bigr\]\}\_\{\\mathrm\{NIE\}\(M\)\}\\;\+\\;\\sum\_\{\\begin\{subarray\}\{c\}\\mathcal\{T\}\\subseteq\\mathcal\{C\}\\\\ \|\\mathcal\{T\}\|\\geq 2\\end\{subarray\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\),\(2\)wherexINT​\(𝒯\)=−∑ℛ⊆𝒯\(−1\)\|𝒯\|−\|ℛ\|​Y​\(ℛ\)\\mathrm\{xINT\}\(\\mathcal\{T\}\)=\-\\sum\_\{\\mathcal\{R\}\\subseteq\\mathcal\{T\}\}\(\-1\)^\{\|\\mathcal\{T\}\|\-\|\\mathcal\{R\}\|\}Y\(\\mathcal\{R\}\)is the cross\-component interaction among each subset of nodes𝒯\\mathcal\{T\}\. EachNIE\\operatorname\{NIE\}term can be computed with the single\-component noising estimator, or further decomposed as

Y​\(∅\)−Y​\(𝒞\)=∑M∈𝒞PIE⁡\(M\)⏟path\-specificeffects\+∑M∈𝒞INT​\(M\)⏟individualinteractions\+∑𝒯⊆𝒞\|𝒯\|≥2xINT​\(𝒯\)⏟cross\-componentinteractions\.Y\(\\emptyset\)\-Y\(\\mathcal\{C\}\)\\;=\\;\\underbrace\{\\sum\_\{M\\in\\mathcal\{C\}\}\\operatorname\{PIE\}\(M\)\}\_\{\\begin\{subarray\}\{c\}\\text\{path\-specific\}\\\\ \\text\{effects\}\\end\{subarray\}\}\\;\+\\;\\underbrace\{\\sum\_\{M\\in\\mathcal\{C\}\}\\mathrm\{INT\}\(M\)\}\_\{\\begin\{subarray\}\{c\}\\text\{individual\}\\\\ \\text\{interactions\}\\end\{subarray\}\}\\;\+\\;\\underbrace\{\\sum\_\{\\begin\{subarray\}\{c\}\\mathcal\{T\}\\subseteq\\mathcal\{C\}\\\\ \|\\mathcal\{T\}\|\\geq 2\\end\{subarray\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{cross\-component\}\\\\ \\text\{interactions\}\\end\{subarray\}\}\.\(3\)

Analogous to the single\-mediator decompositionNIE=PIE\+INT\\operatorname\{NIE\}=\\operatorname\{PIE\}\+\\mathrm\{INT\}, the multi\-mediator decomposition reads asNIE=∑PIE\+∑sINT\+∑xINT\\operatorname\{NIE\}=\\sum\\operatorname\{PIE\}\+\\sum\\mathrm\{sINT\}\+\\sum\\mathrm\{xINT\}, wheresINT\\mathrm\{sINT\}andxINT\\mathrm\{xINT\}are the individual interactions and cross\-interactions respectively\. The full proof is in Appendix[A](https://arxiv.org/html/2606.27510#A1)\.

Of the three terms, xINT is invisible to per\-head measurement and arises specifically when multiple nodes are patched, as it measures how components interact with each other when patched together\. Since each component’s output is added to the residual stream, the counterfactual activations that are patched at different points in the model compound through shared nonlinearities in a way that individual patches do not\. As a result, measuring each head’s NIE, PIE, and sINT individually during the localization and ranking process is still not enough to predict what a circuit does collectively\.

![Refer to caption](https://arxiv.org/html/2606.27510v1/x3.png)
![Refer to caption](https://arxiv.org/html/2606.27510v1/x4.png)

Figure 2:Faithfulness decompositionF=∑PIE\+∑sINT\+∑xINTF=\\sum\\text\{PIE\}\+\\sum\\text\{sINT\}\+\\sum\\text\{xINT\}on the IOI task in GPT\-2 small\. Split violins show per\-pair distributions for ABBA and BABA prompt templates\. Dotted lines mark quartiles\.Left \(symmetric\):Both INT components are more dispersed than the PIE term despite near\-zero mean\.Right \(pABC\):∑PIE\\sum\\text\{PIE\}is large and positive \(mean\+5\.78\+5\.78\) while∑xINT\\sum\\text\{xINT\}is large and negative \(mean−4\.05\-4\.05\); the heads collectively have a lower causal effect than the sum of their individual path\-specific contributions\. In all cases, BABA shows higher variability than ABBA\.#### Explaining Faithfulness Anomalies

Faithfulness is the most common metric used to evaluate circuit discovery procedures\. It aims to measure how much of the full model’s performance is recovered by the circuit, usually expressed as a ratio\. Since performance is usually expressed as a total causal effect such as logit difference, we express faithfulness as a difference instead — the gap between the total effect and the multi\-component noising effect where all out\-of\-circuit nodes are patched\. This allows us to use the decomposition of[Theorem˜4\.1](https://arxiv.org/html/2606.27510#S4.Thmtheorem1)to analyze how pure indirect effects, individual interactions, and cross interactions contribute to faithfulness and its variability\.

Cross\-interaction accounts for nearly the entire gap between what individual heads contribute and what the circuit produces jointly, as shown in Figure[2](https://arxiv.org/html/2606.27510#S4.F2)\. Under pABC mean ablation, individual path\-specific contributions sum to∑PIE¯=\+5\.775\\overline\{\\sum\\mathrm\{PIE\}\}=\+5\.775logit\-diff\. The circuit’s jointly measured faithfulness is onlyF¯=1\.523\\bar\{F\}=1\.523, with∑xINT¯=−4\.045\\overline\{\\sum\\mathrm\{xINT\}\}=\-4\.045explaining most of the difference\. The circuit collectively produces roughly a quarter of what summed PIEs predict\.

Milleret al\.\([2024](https://arxiv.org/html/2606.27510#bib.bib27)\)documented a systematic gap between faithfulness scores across IOI prompt templates \(ABBA and BABA\), as well as high per\-prompt variability\. We show in Figure[2](https://arxiv.org/html/2606.27510#S4.F2)that the gap across prompt templates is primarily driven by cross\-interactions\. Under pABC corruption,∑xINT\\sum\\mathrm\{xINT\}and∑sINT\\sum\\mathrm\{sINT\}are affected by the template difference in opposite directions:Δ​∑xINT=\+3\.098\\Delta\\sum\\mathrm\{xINT\}=\+3\.098andΔ​∑sINT=−1\.116\\Delta\\sum\\mathrm\{sINT\}=\-1\.116\. The ABBA/BABA asymmetry is not a property of how individual heads process the two templates, but of how they interact when patched together\.

Both INT terms also carry higher per\-pair variance than PIE\. Under pABC mean ablation,SD​\(∑sINT\)=2\.69\\mathrm\{SD\}\(\\sum\\mathrm\{sINT\}\)=2\.69andSD​\(∑xINT\)=3\.07\\mathrm\{SD\}\(\\sum\\mathrm\{xINT\}\)=3\.07both exceedSD​\(∑PIE\)=2\.21\\mathrm\{SD\}\(\\sum\\mathrm\{PIE\}\)=2\.21\. Interaction terms fluctuate more than path\-specific contributions because they depend on how clean and corrupted contexts jointly activate the circuit, not on any single head’s contribution in isolation\.

## 5When INT is Interpretable: A Case Study on the GPT\-2 IOI Circuit

The Indirect Object Identification circuit in GPT\-2 small\(Wanget al\.,[2023](https://arxiv.org/html/2606.27510#bib.bib34)\)is the only published circuit whose characterization included combinatorial ablation over groups of attention heads, not just single\-head NIE rankings\. This was required in order to discover backup mechanisms that only activated when other groups of heads were ablated\. This methodology predates the field’s widespread adoption of greedy activation patching methods, and it is what makes the circuit’s functional labels available as a pre\-registered baseline against which hypotheses about interaction effects can be tested\. Through our study of interactions in the IOI circuit, we demonstrate the following:

1. \(a\)INT values across heads carry genuine mechanistic information\.By clustering heads based on the signs and magnitudes of INT and PIE, we recover heads with backup compensation, context\-specific mechanisms, and harmful mechanisms\.
2. \(b\)NIE systematically under\-ranks backup mechanisms\.This includes the Backup Name Mover heads as well as the primary Name Mover head with the highest PIE\.
3. \(c\)Removing INT introduces completely different ranking failures\.Ranking by PIE excludes context\-specific mechanisms that are required for the Name Mover heads to work effectively, including Previous Token, Duplicate Token, and Induction heads\.

All data is obtained using the IOI dataset ofN=1000N=1000samples provided byHannaet al\.\([2024](https://arxiv.org/html/2606.27510#bib.bib47)\)\. Mean ablation with prompts from thepA​B​Cp\_\{ABC\}distribution is used as the baseline value for NIE computations, as inWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)\. Further details can be found in[Appendix˜B](https://arxiv.org/html/2606.27510#A2)\.

### 5\.1INT Patterns Across Functional Head Groups

![Refer to caption](https://arxiv.org/html/2606.27510v1/x5.png)Figure 3:Left:Mean NIE vs\. mean PIE per head, colored by their functional group fromWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)\. NMH and BNMH sit below the diagonal; S\-Inh, DTH, PTH, and Induction cluster near theyy\-axis \(PIE≈0\\approx 0,INT≈NIE\\mathrm\{INT\}\\approx\\operatorname\{NIE\}\); NNMH occupies the negative\-PIE region above the diagonal\.Right:MeanINT\\mathrm\{INT\}by functional group; each point is one head, with median INT marked with black lines\. NMH and most BNMH are negative; S\-Inh, DTH, PTH, and Induction are positive; NNMH is moderately positive; the remaining heads form a median\-zero cluster\.We describe three mechanistically distinct interaction patterns that emerge across the attention heads in GPT\-2 small on the IOI task\. The NIE, PIE, and INT values are shown in[Figure˜3](https://arxiv.org/html/2606.27510#S5.F3)\.

#### Backup compensation\.

These are heads withINT<0\\mathrm\{INT\}<0andNIE<PIE\\operatorname\{NIE\}<\\operatorname\{PIE\}; their marginal contribution shrinks when the rest of the circuit sees the clean input, since other heads are active that perform the same function\. This pattern is mainly seen in the Name Mover Heads \(NMH\) and Backup Name Mover Heads \(BNMH\)\. Most notably, the name mover head L9H9 \(name mover\) has the highestPIE=\+2\.575\\operatorname\{PIE\}=\+2\.575but nearly equal and oppositeINT=−2\.338\\mathrm\{INT\}=\-2\.338, resulting in a much lower rankingNIE=\+0\.237\\operatorname\{NIE\}=\+0\.237\.

#### Context\-specific mechanisms\.

These are heads withINT\>0\\mathrm\{INT\}\>0andPIE≈0\\operatorname\{PIE\}\\approx 0; their contribution only materializes when specific downstream components are active\. This pattern is mainly seen in the Duplicate Token \(DTH\), Induction, and Previous Token \(PTH\) heads;[Figure˜3](https://arxiv.org/html/2606.27510#S5.F3)shows that their PIE values are near\-zero and solely contribute through INT\. S\-Inhibition heads also fit this pattern, though they have slightly higherPIE\\operatorname\{PIE\}\(−0\.02\-0\.02to\+0\.25\+0\.25\)\.

#### Harmful mechanisms\.

These are heads withPIE<0\\operatorname\{PIE\}<0andINT\>0\\mathrm\{INT\}\>0; they directly hurt task performance but are partially opposed by other heads when they are active\. This pattern is seen in the two Negative Name Mover heads \(NNMH\), which have PIEs of−0\.8\-0\.8and−1\.6\-1\.6and write the wrong answer to the output\.

As shown in[Figure˜3](https://arxiv.org/html/2606.27510#S5.F3)\(right\), these patterns are largely stable within head groups and structured by sign and magnitude\. They also sharpen the AND/OR circuit picture ofHeimersheim and Nanda \([2024](https://arxiv.org/html/2606.27510#bib.bib24)\): the same INT sign can result from qualitatively different functional relationships with the other circuit heads\. Together, they demonstrate that INT carries information about functional roles that neither NIE nor PIE recover on their own\.

### 5\.2Diagnosing Activation Patching Results

We showed in[Section˜3](https://arxiv.org/html/2606.27510#S3)that NIE and PIE rankings on the IOI task in GPT\-2 IOI disagree substantially\. Here we show that this disagreement is not arbitrary; each estimand systematically under\-ranks a different functional group of heads in the circuit\. The rankings are shown in Figure[4](https://arxiv.org/html/2606.27510#S5.F4)\.

![Refer to caption](https://arxiv.org/html/2606.27510v1/x6.png)Figure 4:All 144 heads ranked by signed mean \(descending\); dashed vertical line marks the Wang et al\. circuit boundary \(k=26k=26\)\. Outlined circles identify the heads each estimand misranks\.\(A\)NIE ranking: backup heads \(BNMH, outlined\) are displaced to ranks 11–116 byINT<0\\mathrm\{INT\}<0suppression, falling in the lower half of the top\-26 circuit despite their functional importance\.\(B\)PIE ranking: heads whose contribution is entirely pathway\-mediated \(DTH, Induction, PTH, outlined\) are pushed to ranks 30–130, below every head with positive PIE, even though NIE ranks them in the top 12\.#### NIE underranks the backup tier\.

INT<0\\mathrm\{INT\}<0results in lower NIE scores relative to their corresponding PIE, which reduces the head’s position in the ranking\. The PIE rankings for seven of the eight backup name mover heads fall within the top 26 nodes, where 26 is the size of theWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)IOI circuit\. Under NIE, the same nodes are spread across ranks 11 to 116\.

Pairwise interaction alone does not reveal the backup mechanism either\. For example, pairwise INT between L9H9 and any individual BNMH head reaches at most7\.1%7\.1\\%of L9H9’s individual INT\. However, cross\-interaction between each backup name mover head and the full NMH group is up to3\.8×3\.8\\timeshigher\. This is consistent withWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\), who selected the backup heads based on whether there was at least a 2% effect on logit difference after ablating the full NMH group\.

#### PIE misses context\-specific mechanisms\.

Atk=26k=26, signedNIEmean\\operatorname\{NIE\}\_\{\\text\{mean\}\}places three Induction heads at ranks 2, 6, and 7; two DTH heads at ranks 5 and 9; and PTH L4H11 at rank 12 — this group occupies the top half of the circuit\. However, PIE rankings exclude all of them except Induction L5H5 \(PIE=\+0\.020\\operatorname\{PIE\}=\+0\.020\), which barely makes it in at rank 19\. These heads support the S\-Inhibition mechanism by identifying and routing positional information about the subject token\. Under pABC corruption, however, there is no repeated name in the prompt, so the duplicate\-token detection that DTH, Induction, and PTH implement has nothing to operate on\.

#### INT reveals a potential competing mechanism

We observe significant negative pairwise interactions between S\-Inhibition and Negative Name Mover heads\. For example, NNMH L10H7 has pairwise INTs of−0\.480\-0\.480and−0\.606\-0\.606with S\-Inhibition heads L8H10 and L8H6 respectively\. Although these S\-Inhibition heads have positive individual INTs \(\+0\.717\+0\.717and\+1\.198\+1\.198respectively\), their negative pairwise INTs with NNMH suggest a different mechanism at play\. S\-Inhibition suppresses subject\-name output while NNMH promotes it, and the negative interaction is consistent with mutual opposition between these head functions\.

While INT values alone do not confirm this type of competing mechanism, they generate a falsifiable hypothesis that could be confirmed through targeted patching, visualization, and ablation experiments of the kindWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)used to characterize the circuit’s other backup mechanisms\.

### 5\.3The IOI Circuit as Ground Truth

The IOI circuit is the closest approximation to ground truth in a realistic language model and the most common benchmark for evaluating localization methods\. Its credibility rests on manual discovery rather than automated procedures, and on a complete characterization covering token position\-level structure, functional roles, and head groups\. Several of the design decisions that produced this characterization are no longer in common practice; we describe below why each was essential for grounding INT values in known functional structure\.

#### Counterfactual prompt design\.

The pABC prompt was designed to isolate mechanisms in the model related to duplicate token detection and processing\.[Figure˜1](https://arxiv.org/html/2606.27510#S3.F1)shows that under symmetric token replacement corruption, INT collapses and NIE rankings reduce to PIE rankings; these would have also left out Duplicate Token, Previous Token, and Induction heads\. The pABC counterfactual is what makes those heads visible, since their contributions exist almost entirely through INT\.

#### Genuine functional modularity\.

GPT\-2’s IOI circuit implements functionally distinct head groups, and INT signs and magnitudes are consistent within those groups\. When the underlying functionality is diffuse across components, INTs between those components reflect mixtures of contributions rather than relationships between distinct mechanisms, making them difficult to interpret\. The modularity of the IOI circuit was established independently, and cannot be assumed for arbitrary models and tasks\.

#### Prior independent characterization through combinatorial search\.

Wanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)identified backup mechanisms through combinatorial ablation over groups of heads\. This is what made it possible to confirm that INT values track known functional distinctions rather than simply disagreeing with NIE\. Measuring INT is not an escape hatch from combinatorial search, but clustering and interpreting interaction values is a step toward constraining and guiding that search\.

## 6Discussion and Future Work

We have demonstrated that non\-zero interaction terms \(INTs\) are likely widespread in interpretability settings \(§[3](https://arxiv.org/html/2606.27510#S3)\), and can quantitatively characterize and explain many known failure modes in circuit analysis \(§[5](https://arxiv.org/html/2606.27510#S5)\)\. Interpretability research thus far has implicitly assumed that components can be explained in a modular fashion, and that counterfactual interventions tend to yield direct estimates of component importance\. Nonetheless, if one desires a truly modular description of a component’s role, then one must explain and control for INTs\. Given how ubiquitous redundancy is in neural networks, overdetermination and preemption \(and therefore INTs\) seem likely to be frequent\(Mueller,[2024](https://arxiv.org/html/2606.27510#bib.bib48)\)\. The only methods guaranteed to control for INTs will require full combinatorial or recursive search over mediators, rendering exact methods intractable for contemporary models of interest\.

This raises a natural question:*should*one control for INTs? While one’s initial instinct might reasonably be to do so, we instead followIkram and VanderWeele \([2015](https://arxiv.org/html/2606.27510#bib.bib10)\)in advocating that INTs be embraced as a useful analysis tool in its own right\. Indeed, interaction terms have been directly studied and have yielded notable insights in epidemiology settings\(VanderWeele and Knol,[2014](https://arxiv.org/html/2606.27510#bib.bib9); VanderWeele,[2019](https://arxiv.org/html/2606.27510#bib.bib16)\)\. In interpretability, natural analogies could include quantifying to what extent certain mediators’ effects are prompt\-dependent \(i\.e\., interact with the treatment\), or to what extent two mechanisms with no component overlap functionally intersect, and how\.

To address the intractability of combinatorial search, future work could explore heuristics\-based methods that reduce the search space\. For example, much of the cleanliness of analyses of the IOI circuit ofWanget al\.\([2023](https://arxiv.org/html/2606.27510#bib.bib34)\)derives from the clustering of components by functional role\. Can the discovery of functionally similar components be automated, then? Interpretability agents\(Rott Shahamet al\.,[2024](https://arxiv.org/html/2606.27510#bib.bib13); Kimet al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib11); Haklayet al\.,[2026](https://arxiv.org/html/2606.27510#bib.bib12)\)and more fine\-grained taxonomies of model representations\(e\.g\., Aradet al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib15)\)seem likely to be informative\. Additionally, advances in concept geometry have led to more selective steerability with fewer side effects\(e\.g\., Markset al\.,[2025](https://arxiv.org/html/2606.27510#bib.bib1); Sarfatiet al\.,[2026](https://arxiv.org/html/2606.27510#bib.bib14)\); while not guaranteed to fully control for INTs, this represents a tractable way to reduce their impact on our interpretation of particular model representations\.

## References

- Insights into the Cross\-world independence assumption of causal mediation analysis\.Epidemiology32\(2\),pp\. 209–219\.External Links:[Document](https://dx.doi.org/10.1097/EDE.0000000000001313),[Link](https://pubmed.ncbi.nlm.nih.gov/33512846/)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p3.1),[§3](https://arxiv.org/html/2606.27510#S3.p1.1)\.
- D\. Arad, A\. Mueller, and Y\. Belinkov \(2025\)SAEs are good for steering – if you select the right features\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 10241–10259\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.519/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.519),ISBN 979\-8\-89176\-332\-6Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- C\. Avin, I\. Shpitser, and J\. Pearl \(2005\)Identifiability of path\-specific effects\.InProceedings of the 19th International Joint Conference on Artificial Intelligence,IJCAI’05,San Francisco, CA, USA,pp\. 357–363\.Cited by:[Appendix A](https://arxiv.org/html/2606.27510#A1.p1.1),[§1](https://arxiv.org/html/2606.27510#S1.p3.1),[§2](https://arxiv.org/html/2606.27510#S2.p2.1),[§3](https://arxiv.org/html/2606.27510#S3.p1.1)\.
- A\. Conmy, A\. N\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-Alonso \(2023\)Towards automated circuit discovery for mechanistic interpretability\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 16318–16352\.External Links:[Link](https://openreview.net/forum?id=89lM2hjkdW)Cited by:[§2](https://arxiv.org/html/2606.27510#S2.p5.1)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah \(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.Note:https://transformer\-circuits\.pub/2021/framework/index\.htmlCited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1),[§1](https://arxiv.org/html/2606.27510#S1.p3.1),[§3](https://arxiv.org/html/2606.27510#S3.p1.1)\.
- M\. Finlayson, A\. Mueller, S\. Gehrmann, S\. Shieber, T\. Linzen, and Y\. Belinkov \(2021\)Causal analysis of syntactic agreement mechanisms in neural language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 1828–1843\.External Links:[Link](https://aclanthology.org/2021.acl-long.144/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.144)Cited by:[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p2.5)\.
- A\. Geiger, H\. Lu, T\. F\. Icard, and C\. Potts \(2021\)Causal abstractions of neural networks\.InAdvances in Neural Information Processing Systems,A\. Beygelzimer, Y\. Dauphin, P\. Liang, and J\. W\. Vaughan \(Eds\.\),External Links:[Link](https://openreview.net/forum?id=RmuXDtjDhG)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1),[§2](https://arxiv.org/html/2606.27510#S2.p3.1)\.
- N\. Goldowsky\-Dill, C\. MacLeod, L\. Sato, and A\. Arora \(2023\)Localizing model behavior with path patching\.External Links:2304\.05969,[Link](https://arxiv.org/abs/2304.05969)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1)\.
- R\. Gupta, I\. Arcuschin, T\. Kwa, and A\. Garriga\-Alonso \(2024\)InterpBench: semi\-synthetic transformers for evaluating mechanistic interpretability techniques\.InThe Thirty\-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=R9gR9MPuD5)Cited by:[item \(b\)](https://arxiv.org/html/2606.27510#S1.I1.i2.p1.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p5.1)\.
- T\. Haklay, N\. Prakash, S\. Pandey, A\. Torralba, A\. Mueller, J\. Andreas, T\. R\. Shaham, and Y\. Belinkov \(2026\)Pitfalls in evaluating interpretability agents\.External Links:2603\.20101,[Link](https://arxiv.org/abs/2603.20101)Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- M\. Hanna, O\. Liu, and A\. Variengien \(2023\)How does GPT\-2 compute greater\-than?: interpreting mathematical abilities in a pre\-trained language model\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=p4PckNQR8k)Cited by:[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p2.5)\.
- M\. Hanna, S\. Pezzelle, and Y\. Belinkov \(2024\)Have faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=TZ0CCGDcuT)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1),[§1](https://arxiv.org/html/2606.27510#S1.p2.1),[§2](https://arxiv.org/html/2606.27510#S2.p5.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p1.2),[§5](https://arxiv.org/html/2606.27510#S5.p3.2)\.
- S\. Heimersheim and N\. Nanda \(2024\)How to use and interpret activation patching\.External Links:2404\.15255,[Link](https://arxiv.org/abs/2404.15255)Cited by:[§2](https://arxiv.org/html/2606.27510#S2.p3.1),[§3](https://arxiv.org/html/2606.27510#S3.p3.1),[§5\.1](https://arxiv.org/html/2606.27510#S5.SS1.SSS0.Px3.p2.1)\.
- M\. A\. Ikram and T\. J\. VanderWeele \(2015\)A proposed clinical and biological interpretation of mediated interaction\.European journal of epidemiology30\(10\),pp\. 1115–1118\.Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p2.1)\.
- B\. Kim, J\. Hewitt, N\. Nanda, N\. Fiedel, and O\. Tafjord \(2025\)Because we have llms, we can and should pursue agentic interpretability\.External Links:2506\.12152,[Link](https://arxiv.org/abs/2506.12152)Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- D\. Lindner, J\. Kramar, S\. Farquhar, M\. Rahtz, T\. McGrath, and V\. Mikulik \(2023\)Tracr: compiled transformers as a laboratory for interpretability\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=tbbId8u7nP)Cited by:[item \(b\)](https://arxiv.org/html/2606.27510#S1.I1.i2.p1.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p5.1)\.
- S\. Marks, C\. Rager, E\. J\. Michaud, Y\. Belinkov, D\. Bau, and A\. Mueller \(2025\)Sparse feature circuits: discovering and editing interpretable causal graphs in language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=I4e82CIDxv)Cited by:[§2](https://arxiv.org/html/2606.27510#S2.p5.1),[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- M\. Méloux, F\. Portet, and M\. Peyrard \(2025\)Mechanistic interpretability as statistical estimation: a variance analysis of EAP\-IG\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=UbOAXViKsU)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p2.1)\.
- K\. Meng, D\. Bau, A\. J\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in GPT\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),External Links:[Link](https://openreview.net/forum?id=-h6WAS6eE4)Cited by:[§2](https://arxiv.org/html/2606.27510#S2.p3.1)\.
- J\. Miller, B\. Chughtai, and W\. Saunders \(2024\)Transformer circuit evaluation metrics are not robust\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=zSf8PJyQb2)Cited by:[item \(c\)](https://arxiv.org/html/2606.27510#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2606.27510#S1.p2.1),[§4](https://arxiv.org/html/2606.27510#S4.SS0.SSS0.Px1.p3.4)\.
- A\. Mueller, J\. Brinkmann, M\. Li, S\. Marks, K\. Pal, N\. Prakash, C\. Rager, A\. Sankaranarayanan, A\. S\. Sharma, J\. Sun, E\. Todd, D\. Bau, and Y\. Belinkov \(2026\)The quest for the right mediator: surveying mechanistic interpretability for nlp through the lens of causal mediation analysis\.Computational Linguistics52\(1\),pp\. 331–378\.External Links:ISSN 0891\-2017,[Document](https://dx.doi.org/10.1162/COLI.a.572),[Link](https://doi.org/10.1162/COLI.a.572),https://direct\.mit\.edu/coli/article\-pdf/52/1/331/2554934/coli\.a\.572\.pdfCited by:[§2](https://arxiv.org/html/2606.27510#S2.p3.1)\.
- A\. Mueller, A\. Geiger, S\. Wiegreffe, D\. Arad, I\. Arcuschin, A\. Belfki, Y\. S\. Chan, J\. F\. Fiotto\-Kaufman, T\. Haklay, M\. Hanna, J\. Huang, R\. Gupta, Y\. Nikankin, H\. Orgad, N\. Prakash, A\. Reusch, A\. Sankaranarayanan, S\. Shao, A\. Stolfo, M\. Tutek, A\. Zur, D\. Bau, and Y\. Belinkov \(2025\)MIB: a mechanistic interpretability benchmark\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=sSrOwve6vb)Cited by:[item \(b\)](https://arxiv.org/html/2606.27510#S1.I1.i2.p1.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p5.1)\.
- A\. Mueller \(2024\)Missed causes and ambiguous effects: counterfactuals pose challenges for interpreting neural networks\.InICML 2024 Workshop on Mechanistic Interpretability,External Links:[Link](https://openreview.net/forum?id=pJs3ZiKBM5)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p2.1),[§6](https://arxiv.org/html/2606.27510#S6.p1.1)\.
- C\. Olah, N\. Cammarata, L\. Schubert, G\. Goh, M\. Petrov, and S\. Carter \(2020\)Zoom in: an introduction to circuits\.Distill\.Note:https://distill\.pub/2020/circuits/zoom\-inExternal Links:[Document](https://dx.doi.org/10.23915/distill.00024.001)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1)\.
- M\. Paunio, J\. Höök\-Nikanne, T\. U\. Kosunen, U\. Vainio, M\. Salaspuro, J\. Mäkinen, and O\. P\. Heinonen \(1994\)Association of alcohol consumption and helicobacter pylori infection in young adulthood and early middle age among patients with gastric complaints: a case\-control study on finnish conscripts, officers and other military personnel\.European Journal of Epidemiology10\(2\),pp\. 205–209\.External Links:[Document](https://dx.doi.org/10.1007/BF01730371)Cited by:[§3](https://arxiv.org/html/2606.27510#S3.p3.1)\.
- J\. Pearl \(2001\)Direct and indirect effects\.InProceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence,UAI’01,San Francisco, CA, USA,pp\. 411–420\.External Links:ISBN 1558608001Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1),[§1](https://arxiv.org/html/2606.27510#S1.p3.1),[Definition 2\.1](https://arxiv.org/html/2606.27510#S2.Thmtheorem1.p1.3)\.
- T\. Rott Shaham, S\. Schwettmann, F\. Wang, A\. Rajaram, E\. Hernandez, J\. Andreas, and A\. Torralba \(2024\)A multimodal automated interpretability agent\.InForty\-first International Conference on Machine Learning,Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- R\. Sarfati, E\. Bigelow, D\. Wurgaft, J\. Merullo, A\. Geiger, O\. Lewis, T\. McGrath, and E\. S\. Lubana \(2026\)The shape of beliefs: geometry, dynamics, and interventions along representation manifolds of language models’ posteriors\.External Links:2602\.02315,[Link](https://arxiv.org/abs/2602.02315)Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.
- C\. Shi, N\. Beltran\-Velez, A\. Nazaret, C\. Zheng, A\. Garriga\-Alonso, A\. Jesson, M\. Makar, and D\. Blei \(2024\)Hypothesis testing the circuit hypothesis in LLMs\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=5ai2YFAXV7)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p2.1)\.
- A\. Syed, C\. Rager, and A\. Conmy \(2024\)Attribution patching outperforms automated circuit discovery\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 407–416\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.25/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.25)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1)\.
- T\. J\. VanderWeele and M\. J\. Knol \(2014\)A tutorial on interaction\.Epidemiologic methods3\(1\),pp\. 33–72\.Cited by:[§6](https://arxiv.org/html/2606.27510#S6.p2.1)\.
- T\. J\. VanderWeele \(2013\)A Three\-way Decomposition of a Total Effect into Direct, Indirect, and Interactive Effects\.Epidemiology24\(2\),pp\. 224–232\.External Links:23487823,ISSN 1044\-3983Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p4.1),[§3](https://arxiv.org/html/2606.27510#S3.p2.1)\.
- T\. J\. VanderWeele \(2014\)A unification of mediation and interaction: a 4\-way decomposition\.Epidemiology25\(5\),pp\. 749–761\.External Links:[Document](https://dx.doi.org/10.1097/EDE.0000000000000121)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p3.1)\.
- T\. J\. VanderWeele \(2015\)Explanation in causal inference: methods for mediation and interaction\.Oxford University Press,New York\.External Links:ISBN 9780199325870Cited by:[§3](https://arxiv.org/html/2606.27510#S3.p3.1)\.
- T\. J\. VanderWeele \(2019\)The interaction continuum\.Epidemiology30\(5\),pp\. 648–658\.Cited by:[§3](https://arxiv.org/html/2606.27510#S3.p3.1),[§6](https://arxiv.org/html/2606.27510#S6.p2.1)\.
- J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber \(2020\)Investigating gender bias in language models using causal mediation analysis\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 12388–12401\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/92650b2e92217715fe312e6fa7b90d82-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2606.27510#S1.p1.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p2.5),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p4.3)\.
- K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt \(2023\)Interpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NpsVSN6o4ul)Cited by:[§B\.1](https://arxiv.org/html/2606.27510#A2.SS1.p1.3),[§B\.2](https://arxiv.org/html/2606.27510#A2.SS2.SSS0.Px1.p1.1),[§B\.2](https://arxiv.org/html/2606.27510#A2.SS2.SSS0.Px1.p2.5),[Table 1](https://arxiv.org/html/2606.27510#A2.T1),[Table 1](https://arxiv.org/html/2606.27510#A2.T1.4.2),[Table 2](https://arxiv.org/html/2606.27510#A2.T2),[Table 2](https://arxiv.org/html/2606.27510#A2.T2.4.2),[Table 2](https://arxiv.org/html/2606.27510#A2.T2.8.4.4.8.1.1),[§1](https://arxiv.org/html/2606.27510#S1.p2.1),[§3](https://arxiv.org/html/2606.27510#S3.SS0.SSS0.Px1.p2.5),[Figure 3](https://arxiv.org/html/2606.27510#S5.F3),[Figure 3](https://arxiv.org/html/2606.27510#S5.F3.7.3.3),[§5\.2](https://arxiv.org/html/2606.27510#S5.SS2.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2606.27510#S5.SS2.SSS0.Px1.p2.2),[§5\.2](https://arxiv.org/html/2606.27510#S5.SS2.SSS0.Px3.p2.1),[§5\.3](https://arxiv.org/html/2606.27510#S5.SS3.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2606.27510#S5.p1.1),[§5](https://arxiv.org/html/2606.27510#S5.p3.2),[§6](https://arxiv.org/html/2606.27510#S6.p3.1)\.

## Appendix AProofs

We reproduce the definition of the path\-specific effect\[Avinet al\.,[2005](https://arxiv.org/html/2606.27510#bib.bib51)\]for reference\.

###### Definition A\.1\(Path\-Specific Effect\)\.

Fix a DAGGGand letg⊆Gg\\subseteq Gbe an*effect subgraph*with complementg¯=G∖g\\bar\{g\}=G\\setminus g\. The*modified model*ℳg\\mathcal\{M\}\_\{g\}is obtained fromGGby replacing eachg¯\\bar\{g\}\-parent of every nodeVVwith its baseline valueV​\(a′\)V\(\{a^\{\\prime\}\}\)\. The*path\-specific effect*alongggis

PSEg​\(a,a′\)=𝔼​\[Ya\|ℳg\]−𝔼​\[Ya′\|ℳg\]\\mathrm\{PSE\}\_\{g\}\(a,\{a^\{\\prime\}\}\)=\\mathbb\{E\}\\bigl\[Y\_\{a\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}\\bigr\]\-\\mathbb\{E\}\\bigl\[Y\_\{\{a^\{\\prime\}\}\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}\\bigr\]

See[3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)

###### Proof\.

Consider a transformer where each layerℓ\\ellcontains multi\-head attention followed by an MLP, such that each layerℓ\\ellcomputes:

zℓ=xℓ−1\+∑i=1nheadhℓi​\(xℓ−1\)xℓ=zℓ\+MLPℓ​\(zℓ\)z\_\{\\ell\}=x\_\{\\ell\-1\}\+\\sum\_\{i=1\}^\{n\_\{\\mathrm\{head\}\}\}h^\{i\}\_\{\\ell\}\(x\_\{\\ell\-1\}\)\\qquad x\_\{\\ell\}=z\_\{\\ell\}\+\\mathrm\{MLP\}\_\{\\ell\}\(z\_\{\\ell\}\)where eachhℓi​\(⋅\)h^\{i\}\_\{\\ell\}\(\\cdot\)andMLPℓ​\(⋅\)\\mathrm\{MLP\}\_\{\\ell\}\(\\cdot\)absorb their respective LayerNorms\. LetMMbe a single attention headHℓiH\_\{\\ell\}^\{i\}or the MLP, such that the edges that bypassMMare given by:

BM=\{\{xℓ−1→zℓ\}∪\{Hℓj→zℓ:j≠i\}if​M=Hℓi\{zℓ→xℓ\}if​M=MLPℓB\_\{M\}=\\begin\{cases\}\\\{x\_\{\\ell\-1\}\\to z\_\{\\ell\}\\\}\\cup\\\{H\_\{\\ell\}^\{j\}\\to z\_\{\\ell\}:j\\neq i\\\}&\\text\{if \}M=H\_\{\\ell\}^\{i\}\\\\ \\\{z\_\{\\ell\}\\to x\_\{\\ell\}\\\}&\\text\{if \}M=\\mathrm\{MLP\}\_\{\\ell\}\\end\{cases\}
We treat the attention head and MLP cases in parallel\. ForM=HℓiM=H\_\{\\ell\}^\{i\}, let

r=xℓ−1​\(a\)\+∑j≠iHℓj​\(a\),r′=xℓ−1​\(a′\)\+∑j≠iHℓj​\(a′\)r=x\_\{\\ell\-1\}\(a\)\+\\sum\_\{j\\neq i\}H\_\{\\ell\}^\{j\}\(a\),\\qquad r^\{\\prime\}=x\_\{\\ell\-1\}\(a^\{\\prime\}\)\+\\sum\_\{j\\neq i\}H\_\{\\ell\}^\{j\}\(a^\{\\prime\}\)denote the total bypass contributions tozℓz\_\{\\ell\}under treatment and baseline respectively, so thatzℓ​\(a\)=r\+hz\_\{\\ell\}\(a\)=r\+handzℓ​\(a′\)=r′\+h′z\_\{\\ell\}\(a^\{\\prime\}\)=r^\{\\prime\}\+h^\{\\prime\}\.

ForM=MLPℓM=\\mathrm\{MLP\}\_\{\\ell\}, setr=zℓ​\(a\)r=z\_\{\\ell\}\(a\)andr′=zℓ​\(a′\)r^\{\\prime\}=z\_\{\\ell\}\(a^\{\\prime\}\), so thatxℓ​\(a\)=r\+hx\_\{\\ell\}\(a\)=r\+handxℓ​\(a′\)=r′\+h′x\_\{\\ell\}\(a^\{\\prime\}\)=r^\{\\prime\}\+h^\{\\prime\}\.

In both cases, letΦ​\(r,h\)\\Phi\(r,h\)denote the function mapping bypass contributionrrand component outputhhto the outcomeYY, and writeh=M​\(a\)h=M\(a\),h′=M​\(a′\)h^\{\\prime\}=M\(a^\{\\prime\}\)\.

The four canonical outcomes are:

Φ​\(r,h\)=Y​\(a\),Φ​\(r′,h\)=Y​\(a′,M​\(a\)\),Φ​\(r,h′\)=Y​\(a,M​\(a′\)\),Φ​\(r′,h′\)=Y​\(a′\)\.\\Phi\(r,h\)=Y\(a\),\\quad\\Phi\(r^\{\\prime\},h\)=Y\(a^\{\\prime\},M\(a\)\),\\quad\\Phi\(r,h^\{\\prime\}\)=Y\(a,M\(a^\{\\prime\}\)\),\\quad\\Phi\(r^\{\\prime\},h^\{\\prime\}\)=Y\(a^\{\\prime\}\)\.

#### Part \(a\)\.

Denoising runs the model ona′a^\{\\prime\}and patchesM​\(a′\)↦M​\(a\)M\(a^\{\\prime\}\)\\mapsto M\(a\), producing the counterfactual outputY​\(a′,M​\(a\)\)Y\(a^\{\\prime\},M\(a\)\)by definition\. The denoising effect is therefore

Y​\(a′,M​\(a\)\)−Y​\(a′\)=Φ​\(r′,h\)−Φ​\(r′,h′\)=PIE\.Y\(a^\{\\prime\},M\(a\)\)\-Y\(a^\{\\prime\}\)\\;=\\;\\Phi\(r^\{\\prime\},h\)\-\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\;=\\;\\mathrm\{PIE\}\.
To identify the corresponding PSE, consider the effect subgraphggwithg¯=BM\\bar\{g\}=B\_\{M\}and all other edges ingg\. Inℳg\\mathcal\{M\}\_\{g\}, all bypass edges are blocked at their baseline values\. ForM=HℓiM=H\_\{\\ell\}^\{i\}, this meansxℓ−1x\_\{\\ell\-1\}is frozen atxℓ−1​\(a′\)x\_\{\\ell\-1\}\(a^\{\\prime\}\)and every other headHℓjH\_\{\\ell\}^\{j\}is frozen atHℓj​\(a′\)H\_\{\\ell\}^\{j\}\(a^\{\\prime\}\), giving bypass contributionr′r^\{\\prime\}\. ForM=MLPℓM=\\mathrm\{MLP\}\_\{\\ell\},zℓz\_\{\\ell\}is frozen atzℓ​\(a′\)=r′z\_\{\\ell\}\(a^\{\\prime\}\)=r^\{\\prime\}directly\. In both cases, the edge fromMMinto the residual stream is ingg, soM\|ℳg=hM\\big\|\_\{\\mathcal\{M\}\_\{g\}\}=h\.

The residual stream at the layer output is thereforeΦ​\(r′,h\)\\Phi\(r^\{\\prime\},h\), and since all downstream edges are ingg,Ya\|ℳg=Φ​\(r′,h\)=Y​\(a′,M​\(a\)\)Y\_\{a\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}=\\Phi\(r^\{\\prime\},h\)=Y\(a^\{\\prime\},M\(a\)\)\. Under baselinea′a^\{\\prime\}, every variable takes its natural value regardless ofgg, soYa′\|ℳg=Y​\(a′\)Y\_\{a^\{\\prime\}\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}=Y\(a^\{\\prime\}\)\. HencePIE=PSEg​\(a,a′\)\\mathrm\{PIE\}=\\mathrm\{PSE\}\_\{g\}\(a,a^\{\\prime\}\)for this subgraph\.

#### Part \(b\)\.

Noising runs the model onaaand patchesM​\(a\)↦M​\(a′\)M\(a\)\\mapsto M\(a^\{\\prime\}\), producing the counterfactual outputY​\(a,M​\(a′\)\)Y\(a,M\(a^\{\\prime\}\)\)by definition\. The noising effect is therefore

NIE=Y​\(a\)−Y​\(a,M​\(a′\)\)=Φ​\(r,h\)−Φ​\(r,h′\)\.\\mathrm\{NIE\}\\;=\\;Y\(a\)\-Y\(a,M\(a^\{\\prime\}\)\)\\;=\\;\\Phi\(r,h\)\-\\Phi\(r,h^\{\\prime\}\)\.
The identityNIE=PIE\+INT\\mathrm\{NIE\}=\\mathrm\{PIE\}\+\\mathrm\{INT\}follows by direct cancellation:

PIE\+INT\\displaystyle\\mathrm\{PIE\}\+\\mathrm\{INT\}=\[Φ​\(r′,h\)−Φ​\(r′,h′\)\]\+\[Φ​\(r,h\)−Φ​\(r′,h\)−Φ​\(r,h′\)\+Φ​\(r′,h′\)\]\\displaystyle=\\bigl\[\\Phi\(r^\{\\prime\},h\)\-\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\bigr\]\+\\bigl\[\\Phi\(r,h\)\-\\Phi\(r^\{\\prime\},h\)\-\\Phi\(r,h^\{\\prime\}\)\+\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\bigr\]=Φ​\(r,h\)−Φ​\(r,h′\)=NIE\\displaystyle=\\Phi\(r,h\)\-\\Phi\(r,h^\{\\prime\}\)\\;=\\;\\mathrm\{NIE\}

#### Part \(c\)\.

The three\-way decomposition follows by direct cancellation from the four\-outcome table:

PIE\+PDE\+INT\\displaystyle\\mathrm\{PIE\}\+\\mathrm\{PDE\}\+\\mathrm\{INT\}=\[Φ​\(r′,h\)−Φ​\(r′,h′\)\]\+\[Φ​\(r,h′\)−Φ​\(r′,h′\)\]\\displaystyle=\\bigl\[\\Phi\(r^\{\\prime\},h\)\-\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\bigr\]\+\\bigl\[\\Phi\(r,h^\{\\prime\}\)\-\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\bigr\]\+\[Φ​\(r,h\)−Φ​\(r′,h\)−Φ​\(r,h′\)\+Φ​\(r′,h′\)\]\\displaystyle\\quad\+\\bigl\[\\Phi\(r,h\)\-\\Phi\(r^\{\\prime\},h\)\-\\Phi\(r,h^\{\\prime\}\)\+\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\bigr\]=Φ​\(r,h\)−Φ​\(r′,h′\)=Y​\(a\)−Y​\(a′\)=TE\\displaystyle=\\Phi\(r,h\)\-\\Phi\(r^\{\\prime\},h^\{\\prime\}\)\\;=\\;Y\(a\)\-Y\(a^\{\\prime\}\)\\;=\\;\\mathrm\{TE\}To identify PDE as a PSE, consider the effect subgraphggwithg¯=\{Hℓi→zℓ\}\\bar\{g\}=\\\{H\_\{\\ell\}^\{i\}\\to z\_\{\\ell\}\\\}for an attention head andg¯=\{MLPℓ→xℓ\}\\bar\{g\}=\\\{\\mathrm\{MLP\}\_\{\\ell\}\\to x\_\{\\ell\}\\\}for the MLP, with all other edges ingg\. Inℳg\\mathcal\{M\}\_\{g\}, all bypass edges are inggso treatment propagates normally through the bypass, givingr~=r\\tilde\{r\}=r\. The component edge is ing¯\\bar\{g\}, soMMis frozen ath′h^\{\\prime\}\. HenceYa\|ℳg=Φ​\(r,h′\)=Y​\(a,M​\(a′\)\)Y\_\{a\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}=\\Phi\(r,h^\{\\prime\}\)=Y\(a,M\(a^\{\\prime\}\)\), andYa′\|ℳg=Y​\(a′\)Y\_\{a^\{\\prime\}\}\\big\|\_\{\\mathcal\{M\}\_\{g\}\}=Y\(a^\{\\prime\}\)\. HencePDE=PSEg​\(a,a′\)\\mathrm\{PDE\}=\\mathrm\{PSE\}\_\{g\}\(a,a^\{\\prime\}\)for this subgraph\. ∎

See[3\.2](https://arxiv.org/html/2606.27510#S3.Thmtheorem2)

###### Proof\.

We first express INT in terms ofΨ\\Psi, then apply a Taylor expansion to obtain the bilinear form, and finally verify each sufficient condition\. Writeh=M​\(a\)h=M\(a\),h′=M​\(a′\)h^\{\\prime\}=M\(a^\{\\prime\}\),r=BM​\(a\)r=B\_\{M\}\(a\), andr′=BM​\(a′\)r^\{\\prime\}=B\_\{M\}\(a^\{\\prime\}\), so thatδM=h−h′\\delta\_\{M\}=h\-h^\{\\prime\}andδB=r−r′\\delta\_\{B\}=r\-r^\{\\prime\}\. Since the downstream computation receivesr\+hr\+has the combined residual stream at layerℓ\\ell, the four counterfactual outputs from[Proposition˜3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)are evaluations ofΨ\\Psiat four points:

Y​\(a\)\\displaystyle Y\(a\)=Ψ​\(r′\+h′\+δM\+δB\),\\displaystyle=\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\+\\delta\_\{B\}\),Y​\(a′,M​\(a\)\)\\displaystyle Y\(a^\{\\prime\},M\(a\)\)=Ψ​\(r′\+h′\+δM\),\\displaystyle=\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\),Y​\(a,M​\(a′\)\)\\displaystyle Y\(a,M\(a^\{\\prime\}\)\)=Ψ​\(r′\+h′\+δB\),\\displaystyle=\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{B\}\),Y​\(a′\)\\displaystyle Y\(a^\{\\prime\}\)=Ψ​\(r′\+h′\)\.\\displaystyle=\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\)\.Substituting into the definition of INT from[Proposition˜3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)\(b\):

INT=Ψ​\(r′\+h′\+δM\+δB\)−Ψ​\(r′\+h′\+δM\)−Ψ​\(r′\+h′\+δB\)\+Ψ​\(r′\+h′\)\.\\mathrm\{INT\}=\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\+\\delta\_\{B\}\)\-\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\)\-\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{B\}\)\+\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\)\.\(4\)

#### Real\-analyticity ofΨ\\Psi\.

Softmax is a composition of the exponential function and division by a strictly positive sum, hence real\-analytic\. GeLU is a product of the identity with either the Gaussian CDF or a sigmoid, both of which are real\-analytic\. LayerNorm is real\-analytic because the denominatorσ​\(z\)\+ϵ≥ϵ\>0\\sigma\(z\)\+\\epsilon\\geq\\epsilon\>0is bounded away from zero\. SinceΨ\\Psiis a composition of these operations with linear maps, it is real\-analytic onℝd\\mathbb\{R\}^\{d\}\.

#### Taylor expansion\.

WriteH=∇2Ψ​\(zℓ​\(a′\)\)H=\\nabla^\{2\}\\Psi\(z\_\{\\ell\}\(a^\{\\prime\}\)\)and expand each term in[Equation˜4](https://arxiv.org/html/2606.27510#A1.E4)aroundzℓ​\(a′\)=r′\+h′z\_\{\\ell\}\(a^\{\\prime\}\)=r^\{\\prime\}\+h^\{\\prime\}:

Ψ​\(r′\+h′\+δM\+δB\)\\displaystyle\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\+\\delta\_\{B\}\)=Ψ​\(r′\+h′\)\+∇Ψ⊤​\(δM\+δB\)\+12​\(δM\+δB\)⊤​H​\(δM\+δB\)\+O​\(‖δ‖3\),\\displaystyle=\\Psi\(r^\{\\prime\}\{\+\}h^\{\\prime\}\)\+\\nabla\\Psi^\{\\top\}\(\\delta\_\{M\}\{\+\}\\delta\_\{B\}\)\+\\tfrac\{1\}\{2\}\(\\delta\_\{M\}\{\+\}\\delta\_\{B\}\)^\{\\top\}H\(\\delta\_\{M\}\{\+\}\\delta\_\{B\}\)\+O\(\\\|\\delta\\\|^\{3\}\),Ψ​\(r′\+h′\+δM\)\\displaystyle\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{M\}\)=Ψ​\(r′\+h′\)\+∇Ψ⊤​δM\+12​δM⊤​H​δM\+O​\(‖δ‖3\),\\displaystyle=\\Psi\(r^\{\\prime\}\{\+\}h^\{\\prime\}\)\+\\nabla\\Psi^\{\\top\}\\delta\_\{M\}\+\\tfrac\{1\}\{2\}\\delta\_\{M\}^\{\\top\}H\\delta\_\{M\}\+O\(\\\|\\delta\\\|^\{3\}\),Ψ​\(r′\+h′\+δB\)\\displaystyle\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\+\\delta\_\{B\}\)=Ψ​\(r′\+h′\)\+∇Ψ⊤​δB\+12​δB⊤​H​δB\+O​\(‖δ‖3\)\.\\displaystyle=\\Psi\(r^\{\\prime\}\{\+\}h^\{\\prime\}\)\+\\nabla\\Psi^\{\\top\}\\delta\_\{B\}\+\\tfrac\{1\}\{2\}\\delta\_\{B\}^\{\\top\}H\\delta\_\{B\}\+O\(\\\|\\delta\\\|^\{3\}\)\.Substituting into[Equation˜4](https://arxiv.org/html/2606.27510#A1.E4), the fourΨ​\(r′\+h′\)\\Psi\(r^\{\\prime\}\+h^\{\\prime\}\)terms cancel and the three first\-order terms cancel, leaving only the second\-order terms:

INT\\displaystyle\\mathrm\{INT\}=12​\(δM\+δB\)⊤​H​\(δM\+δB\)−12​δM⊤​H​δM−12​δB⊤​H​δB\+O​\(‖δ‖3\)\.\\displaystyle=\\tfrac\{1\}\{2\}\(\\delta\_\{M\}\{\+\}\\delta\_\{B\}\)^\{\\top\}H\(\\delta\_\{M\}\{\+\}\\delta\_\{B\}\)\-\\tfrac\{1\}\{2\}\\delta\_\{M\}^\{\\top\}H\\delta\_\{M\}\-\\tfrac\{1\}\{2\}\\delta\_\{B\}^\{\\top\}H\\delta\_\{B\}\+O\(\\\|\\delta\\\|^\{3\}\)\.Expanding\(δM\+δB\)⊤​H​\(δM\+δB\)=δM⊤​H​δM\+2​δM⊤​H​δB\+δB⊤​H​δB\(\\delta\_\{M\}\+\\delta\_\{B\}\)^\{\\top\}H\(\\delta\_\{M\}\+\\delta\_\{B\}\)=\\delta\_\{M\}^\{\\top\}H\\delta\_\{M\}\+2\\delta\_\{M\}^\{\\top\}H\\delta\_\{B\}\+\\delta\_\{B\}^\{\\top\}H\\delta\_\{B\}, the diagonal termsδM⊤​H​δM\\delta\_\{M\}^\{\\top\}H\\delta\_\{M\}andδB⊤​H​δB\\delta\_\{B\}^\{\\top\}H\\delta\_\{B\}cancel in pairs, leaving:

INT\\displaystyle\\mathrm\{INT\}=δM⊤​H​δB\+O​\(‖δ‖3\)\\displaystyle=\\delta\_\{M\}^\{\\top\}H\\delta\_\{B\}\+O\(\\\|\\delta\\\|^\{3\}\)
SinceδM⊤​H​δB=D2​Ψ​\(zℓ​\(a′\)\)​\[δM,δB\]\\delta\_\{M\}^\{\\top\}H\\delta\_\{B\}=D^\{2\}\\Psi\(z\_\{\\ell\}\(a^\{\\prime\}\)\)\[\\delta\_\{M\},\\delta\_\{B\}\]by definition of the Hessian bilinear form, dropping theO​\(‖δ‖3\)O\(\\\|\\delta\\\|^\{3\}\)remainder givesINT≈D2​Ψ​\(zℓ​\(a′\)\)​\[δM,δB\]\\mathrm\{INT\}\\approx D^\{2\}\\Psi\(z\_\{\\ell\}\(a^\{\\prime\}\)\)\[\\delta\_\{M\},\\delta\_\{B\}\]\.

#### Sufficient conditions\.

Each condition ensuresδM⊤​H​δB=0\\delta\_\{M\}^\{\\top\}H\\delta\_\{B\}=0\. Condition \(i\) setsδM=0\\delta\_\{M\}=0orδB=0\\delta\_\{B\}=0, condition \(ii\) setsH=0H=0, and condition \(iii\) setsD2​Ψ​\(zℓ​\(a′\)\)​\[δM,δB\]=0D^\{2\}\\Psi\(z\_\{\\ell\}\(a^\{\\prime\}\)\)\[\\delta\_\{M\},\\delta\_\{B\}\]=0by assumption\. The qualifierD2​Ψ≠0D^\{2\}\\Psi\\neq 0distinguishes condition \(iii\) from condition \(ii\): the Hessian is nonzero, meaning the downstream map has genuine curvature at the operating point, but the perturbationsδM\\delta\_\{M\}andδB\\delta\_\{B\}have no shared components in the directions along which that curvature acts\. ∎

See[4\.1](https://arxiv.org/html/2606.27510#S4.Thmtheorem1)

###### Proof\.

Define the centered set functionf​\(𝒯\)=Y​\(𝒯\)−Y​\(∅\)f\(\\mathcal\{T\}\)=Y\(\\mathcal\{T\}\)\-Y\(\\emptyset\)for𝒯⊆𝒞\\mathcal\{T\}\\subseteq\\mathcal\{C\}, withf​\(∅\)=0f\(\\emptyset\)=0\. We claim:

f​\(𝒞\)=−∑∅≠𝒯⊆𝒞xINT​\(𝒯\)f\(\\mathcal\{C\}\)=\-\\sum\_\{\\emptyset\\neq\\mathcal\{T\}\\subseteq\\mathcal\{C\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\)\(5\)
Substituting the definitionxINT​\(𝒯\)=−∑ℛ⊆𝒯\(−1\)\|𝒯\|−\|ℛ\|​Y​\(ℛ\)\\mathrm\{xINT\}\(\\mathcal\{T\}\)=\-\\sum\_\{\\mathcal\{R\}\\subseteq\\mathcal\{T\}\}\(\-1\)^\{\|\\mathcal\{T\}\|\-\|\\mathcal\{R\}\|\}Y\(\\mathcal\{R\}\)and exchanging the order of summation:

∑∅≠𝒯⊆𝒞xINT​\(𝒯\)\\displaystyle\\sum\_\{\\emptyset\\neq\\mathcal\{T\}\\subseteq\\mathcal\{C\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\)=−∑∅≠𝒯⊆𝒞∑ℛ⊆𝒯\(−1\)\|𝒯\|−\|ℛ\|​Y​\(ℛ\)\\displaystyle=\-\\sum\_\{\\emptyset\\neq\\mathcal\{T\}\\subseteq\\mathcal\{C\}\}\\;\\sum\_\{\\mathcal\{R\}\\subseteq\\mathcal\{T\}\}\(\-1\)^\{\|\\mathcal\{T\}\|\-\|\\mathcal\{R\}\|\}Y\(\\mathcal\{R\}\)=−∑ℛ⊆𝒞Y​\(ℛ\)​∑𝒯:ℛ⊆𝒯⊆𝒞𝒯≠∅\(−1\)\|𝒯\|−\|ℛ\|\\displaystyle=\-\\sum\_\{\\mathcal\{R\}\\subseteq\\mathcal\{C\}\}Y\(\\mathcal\{R\}\)\\sum\_\{\\begin\{subarray\}\{c\}\\mathcal\{T\}:\\mathcal\{R\}\\subseteq\\mathcal\{T\}\\subseteq\\mathcal\{C\}\\\\ \\mathcal\{T\}\\neq\\emptyset\\end\{subarray\}\}\(\-1\)^\{\|\\mathcal\{T\}\|\-\|\\mathcal\{R\}\|\}
We evaluate the inner sum by cases\. Forℛ=𝒞\\mathcal\{R\}=\\mathcal\{C\}, the only superset of𝒞\\mathcal\{C\}in𝒞\\mathcal\{C\}is𝒞\\mathcal\{C\}itself, so the inner sum equals\(−1\)0=1\(\-1\)^\{0\}=1, contributingY​\(𝒞\)Y\(\\mathcal\{C\}\)to the overall sum\. For∅≠ℛ⊊𝒞\\emptyset\\neq\\mathcal\{R\}\\subsetneq\\mathcal\{C\}, we can drop the constraint𝒯≠∅\\mathcal\{T\}\\neq\\emptysetsince every superset of a nonempty set is nonempty\. Lettingk=\|𝒞∖ℛ\|≥1k=\|\\mathcal\{C\}\\setminus\\mathcal\{R\}\|\\geq 1andj=\|𝒯\|−\|ℛ\|j=\|\\mathcal\{T\}\|\-\|\\mathcal\{R\}\|:

∑j=0k\(kj\)​\(−1\)j=\(1−1\)k=0,\\sum\_\{j=0\}^\{k\}\\binom\{k\}\{j\}\(\-1\)^\{j\}=\(1\-1\)^\{k\}=0,where\(kj\)\\binom\{k\}\{j\}counts the supersets ofℛ\\mathcal\{R\}in𝒞\\mathcal\{C\}of size\|ℛ\|\+j\|\\mathcal\{R\}\|\+j\. Forℛ=∅\\mathcal\{R\}=\\emptyset, the𝒯=∅\\mathcal\{T\}=\\emptysetterm is excluded by the constraint, so the inner sum is:

∑j=1k\(kj\)​\(−1\)j=\(1−1\)k−1=−1,\\sum\_\{j=1\}^\{k\}\\binom\{k\}\{j\}\(\-1\)^\{j\}=\(1\-1\)^\{k\}\-1=\-1,contributing−Y​\(∅\)\-Y\(\\emptyset\)to the overall sum\. Collecting all three cases:

−∑∅≠𝒯⊆𝒞xINT​\(𝒯\)=Y​\(𝒞\)−Y​\(∅\)=f​\(𝒞\)\-\\sum\_\{\\emptyset\\neq\\mathcal\{T\}\\subseteq\\mathcal\{C\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\)=Y\(\\mathcal\{C\}\)\-Y\(\\emptyset\)=f\(\\mathcal\{C\}\)which establishes the claim in[Equation˜5](https://arxiv.org/html/2606.27510#A1.E5)\.

For singletons𝒯=\{M\}\\mathcal\{T\}=\\\{M\\\}, direct substitution gives\(−1\)0​Y​\(\{M\}\)\+\(−1\)1​Y​\(∅\)=Y​\(\{M\}\)−Y​\(∅\)\(\-1\)^\{0\}\\,Y\(\\\{M\\\}\)\+\(\-1\)^\{1\}\\,Y\(\\emptyset\)=Y\(\\\{M\\\}\)\-Y\(\\emptyset\)\. Hence we get:

Y​\(𝒞\)−Y​\(∅\)=∑M∈𝒞\[Y​\(\{M\}\)−Y​\(∅\)\]−∑𝒯⊆𝒞\|𝒯\|≥2xINT​\(𝒯\)\.Y\(\\mathcal\{C\}\)\-Y\(\\emptyset\)=\\sum\_\{M\\in\\mathcal\{C\}\}\\bigl\[Y\(\\\{M\\\}\)\-Y\(\\emptyset\)\\bigr\]\-\\sum\_\{\\begin\{subarray\}\{c\}\\mathcal\{T\}\\subseteq\\mathcal\{C\}\\\\ \|\\mathcal\{T\}\|\\geq 2\\end\{subarray\}\}\\mathrm\{xINT\}\(\\mathcal\{T\}\)\.
Negating both sides yields[Equation˜2](https://arxiv.org/html/2606.27510#S4.E2)\.

To recover[Equation˜3](https://arxiv.org/html/2606.27510#S4.E3), apply[Proposition˜3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)\(b\) to each individual termY​\(∅\)−Y​\(\{M\}\)=NIE​\(M\)=PIEM\+INTMY\(\\emptyset\)\-Y\(\\\{M\\\}\)=\\mathrm\{NIE\}\(M\)=\\mathrm\{PIE\}\_\{M\}\+\\mathrm\{INT\}\_\{M\}\. The decomposition in[Proposition˜3\.1](https://arxiv.org/html/2606.27510#S3.Thmtheorem1)applies to eachM∈𝒞M\\in\\mathcal\{C\}individually regardless of the other components in𝒞\\mathcal\{C\}, sinceY​\(∅\)−Y​\(\{M\}\)Y\(\\emptyset\)\-Y\(\\\{M\\\}\)is a single\-component noising effect\. Substituting into the first term of[Equation˜2](https://arxiv.org/html/2606.27510#S4.E2)yields[Equation˜3](https://arxiv.org/html/2606.27510#S4.E3)directly\.

∎

## Appendix BIOI Case Study

All values are from the same dataset used to produce the figures and numerical claims in Section[5](https://arxiv.org/html/2606.27510#S5): 1000 IOI prompt pairs under pABC mean\-ablation corruption in GPT\-2 small\. CIs are mean±1\.96×SD/1000\\pm 1\.96\\times\\mathrm\{SD\}/\\\!\\sqrt\{1000\}across prompt pairs\. Signed ranks are over all 144 heads, descending\.

### B\.1All Circuit Heads: NIE, PIE, and INT Rankings

[Table˜1](https://arxiv.org/html/2606.27510#A2.T1)includes the NIE, PIE, and INT values for all the in\-circuit nodes inWanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]\. Of these numbers,[Section˜5](https://arxiv.org/html/2606.27510#S5)highlights the the L9H9 \(Name Mover\) values, the NIE rank range 11–116 for the seven BNMH heads in the signed PIE top\-26, the near\-zero PIE values for context\-specific heads \(DTH, Induction, PTH\), and the PIE range−0\.02\-0\.02to\+0\.25\+0\.25for S\-Inhibition heads\. For 22 of the 26 heads, the per\-pair SD of PIE exceeds the absolute mean, reflecting large prompt\-to\-prompt variability\. The CIs on means are narrow becauseN=1000N\{=\}1000\.

Table 1:Mean NIE, PIE, and INT \(±\\pm95% CI, logit\-diff units\) for all 26Wanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]IOI circuit heads\. PIE rank / NIE rank = signed rank by descending mean over all 144 heads\. Fuzzy heads:†\\dagger\.
### B\.2Backup Name Mover Heads

The maximum pairwise cross\-interaction between L9H9 and any individual BNMH head is0\.1650\.165\(L10H2\), equal to 7\.1% of L9H9’s individual INT; for L11H9, cross\-interaction with the full NMH group \(−0\.065\-0\.065\) is3\.8×3\.8\\timesthe pairwise value \(−0\.017\-0\.017\)\.

Table[2](https://arxiv.org/html/2606.27510#A2.T2)gives the full set of values of NIE, PIE, and xINT\(L9H9\), which is the cross\-component interaction when L9H9 alone is patched simultaneously with the BNMH head\. xINT\(NMH\) is the same quantity when all NMH heads are patched jointly\.

#### BNMH heterogeneity\.

Wanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]observe heterogeneous functionality within the backup name mover group: of the eight heads they label BNMH, four resemble Name Mover Heads, two attend equally to IO and S, one attends more to S1, and one copies S2 negatively\. They select these heads by an arbitrary top\-eight cutoff on direct effect after knocking out the primary Name Mover Heads, retaining only heads outside other groups whose effect exceeds a low threshold: 2% of the model’s clean baseline logit difference on the IOI task\.

Two BNMH heads do not exhibit the backup compensation pattern \(INT<0<\\penalty 10000\\ 0, NIE<<PIE\) described in Section[5\.1](https://arxiv.org/html/2606.27510#S5.SS1)\. L11H2 has mean PIE=−0\.361=\-0\.361: negative because, asWanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]document, it attends more to the subject S1 and copies it rather than the IO token\. Under pABC corruption there is no repeated name, so its backup function cannot engage; instead it contributes negatively\. L11H2 is therefore the single BNMH excluded from the signed PIE top\-26\. L9H7 is a second exception: its INT is positive \(\+0\.186\+0\.186\), soNIE\>PIE\\operatorname\{NIE\}\>\\operatorname\{PIE\}, the opposite of the backup compensation pattern\.Wanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]describe L9H7 as attending to S2 and writing negatively, suppressing the subject name rather than copying IO\. This places it closer to the context\-specific pattern of S\-Inhibition heads, whose contributions are similarly amplified when the surrounding circuit is intact\. The four heads most clearly displaying backup compensation are L10H1, L10H2, L10H6, and L10H10\.

Table 2:BNMH heads: estimand values \(±\\pm95% CI\), signed PIE/NIE ranks, pairwise cross\-interaction with L9H9, cross\-interaction with the full NMH group, andWanget al\.\[[2023](https://arxiv.org/html/2606.27510#bib.bib34)\]behavioral profile\. All logit\-diff units;N=1000N=1000, pABC mean\-ablation\.

### B\.3Pairwise Cross\-Interaction between NNMH and S\-Inhibition Heads

All eight pairwise xINT values in[Table˜3](https://arxiv.org/html/2606.27510#A2.T3)are negative: patching any NNMH and S\-Inh pair jointly suppresses their combined contribution\. This holds across both NNMH heads and all four S\-Inh heads, not only the two pairs cited in the text, supporting the competing mechanism hypothesis in Section[5\.2](https://arxiv.org/html/2606.27510#S5.SS2)\. L10H7 shows the larger magnitudes \(up to−0\.606±0\.028\-0\.606\\pm 0\.028with L8H6\), consistent with its larger individual INT \(\+0\.313±0\.028\+0\.313\\pm 0\.028\); L11H10’s individual INT is near zero \(−0\.011±0\.015\-0\.011\\pm 0\.015\) and its pairwise interactions are correspondingly smaller but still negative throughout\.

Table 3:Pairwise cross\-interaction xINT between each NNMH and each S\-Inhibition head \(logit\-diff units, mean±\\pm95% CI,N=1000N\{=\}1000pairs, pABC mean\-ablation\)\. Individual INT values for NNMH and S\-Inh heads are in Table[1](https://arxiv.org/html/2606.27510#A2.T1)\.

Similar Articles

When Attribution Patching Lies: Diagnosis and a Second-Order Correction

arXiv cs.LG

This paper diagnoses systematic errors in attribution patching, a gradient-based approximation used for causal localization in language models, and proposes a second-order correction using Hessian-vector products that improves reliability with minimal additional computational cost.

Optimal Experiments for Partial Causal Effect Identification

arXiv cs.AI

This paper introduces the 'max-potency problem' for selecting cost-constrained experiments to maximize the tightening of bounds on partial causal effects. The authors propose graphical pruning criteria to reduce the search space and demonstrate the method on NHANES health data.