Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit

arXiv cs.LG Papers

Summary

This paper introduces a site-asymmetry audit to separate baseline effects from interaction in activation-space order-swap interventions, showing that single-intervention baselines explain most variance in language models and proposing a corrected residual for detecting geometric structure.

arXiv:2608.25315v1 Announce Type: new Abstract: Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network. We introduce a no-fit site-asymmetry audit. For a twice-differentiable readout, the open-path order-swap decomposes into a canonical additive response measured by single interventions and an antisymmetrized second difference free of first-order and pure self-curvature terms to second order. Across six open-weight language-model families, the single-intervention baseline explains 84.3-97.7 percent of the bracket norm (mean 93.7 percent), while the no-interaction self-curvature term is 1.8-5.2 times larger than the corrected residual in the two families with the plus/minus injection split. The corrected residual clears a generic-interaction null in three of six families under a confound-free prompt split and two of six after configuration robustness. A known-positive surrogate recovers planted mixed interaction, while a matched site-separation test changes the baseline share and a random architecture reproduces the first-order regime. The same estimator transfers to released non-language references: trained residual fractions fall below a fixed Gaussian-direction null in 11/12 contrasts (5/6 ViT-B/16, 6/6 ResNet-50), a portability check rather than pooled evidence. The contribution is a reusable measurement criterion: run the single-intervention baseline before reading an order-swap vector as interaction or geometric structure; if it explains the vector, form the second difference instead. All claims are scoped to activation-space interventions at distinct sites; we do not claim that representation geometry is globally Abelian.
Original Article
View Cached Full Text

Cached at: 08/27/26, 09:39 AM

# Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit
Source: [https://arxiv.org/html/2608.25315](https://arxiv.org/html/2608.25315)
###### Abstract

Order\-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network\. We introduce a no\-fit site\-asymmetry audit\. For a twice\-differentiable readout, the open\-path order\-swap decomposes into a canonical additive response measured by single interventions and an antisymmetrized second difference free of first\-order and pure self\-curvature terms to second order\. Across six open\-weight language\-model families, the single\-intervention baseline explains 84\.3\-97\.7 percent of the bracket norm \(mean 93\.7 percent\), while the no\-interaction self\-curvature term is 1\.8\-5\.2 times larger than the corrected residual in the two families with the plus/minus injection split\. The corrected residual clears a generic\-interaction null in three of six families under a confound\-free prompt split and two of six after configuration robustness\. A known\-positive surrogate recovers planted mixed interaction, while a matched site\-separation test changes the baseline share and a random architecture reproduces the first\-order regime\. The same estimator transfers to released non\-language references: trained residual fractions fall below a fixed Gaussian\-direction null in 11/12 contrasts \(5/6 ViT\-B/16, 6/6 ResNet\-50\), a portability check rather than pooled evidence\. The contribution is a reusable measurement criterion: run the single\-intervention baseline before reading an order\-swap vector as interaction or geometric structure; if it explains the vector, form the second difference instead\. All claims are scoped to activation\-space interventions at distinct sites; we do not claim that representation geometry is globally Abelian\.

††proceedings:arXiv: arXiv preprint††year:2026††workshop:Symmetry and Geometry in Neural Representations## 1Introduction

Steering vectors added to a transformer’s residual stream compose, and the composition is order\-dependent: injecting directionviv\_\{i\}beforevjv\_\{j\}does not give the same result as the reverse\. The algebraic reading of that order\-dependence is available and inviting\. The measured quantity looks like a bracket: it is antisymmetric, it vanishes when the two directions coincide, and it is non\-zero in practice\. In*weight*space the reading is already established, with the commutator of two fine\-tuning updates treated as the governing order\-dependent quantity\([Sweeney, 2026](https://arxiv.org/html/2608.25315#bib.bib2)\); in activation space the analogous closed\-loop statistic is being computed and read geometrically\([Richards, 2026](https://arxiv.org/html/2608.25315#bib.bib3);[Sevetlidis and Pavlidis, 2026](https://arxiv.org/html/2608.25315#bib.bib4)\)\. Composable\-intervention and task\-arithmetic pipelines provide related composition settings rather than this bracket statistic\([Kolbeinsson et al\., 2025](https://arxiv.org/html/2608.25315#bib.bib21);[Ilharco et al\., 2023](https://arxiv.org/html/2608.25315#bib.bib20);[van der Weij et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib22)\)\. Closest to our statistic,[Mudarisov et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib30)do form an activation\-space Lie bracket and describe it as “measuring whether their order matters”; our criterion returns*exempt*there on the merits rather than by failing to apply\. We are not aware of a published paper that computes the open\-path activation\-space order\-swap statistic across*different*injection depths and reads it as evidence of non\-Abelian structure, and we do not attribute that error to anyone\. What we supply is a null model for that statistic and a cheap test of it: a reusable site\-asymmetry audit that separates what the construction forces from what the model contributes\. The nearest published statistic of our own form is in weight space and outside what our activation\-space criterion tests \([Schessl, 2026](https://arxiv.org/html/2608.25315#bib.bib34); discussed with our own failed attempt to use it in Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. The criterion is also not binary underneath: contamination is driven byJ1−J2J\_\{1\}\-J\_\{2\}, so it predicts a gradient, and widening site separation with the readout held fixed raises the first\-order term’s share of the bracket in66of66families and91\.7%91\.7\\%of720720matched pairs, where raw depth does not and an untrained network of the same shape reaches only5454–58%58\\%\(Appendix[D](https://arxiv.org/html/2608.25315#A4); we report this as a count, not app\-value, because the six families share one trait inventory and are not six independent draws\)\. A loop through four distinct depths is first\-order dominated while the degenerate loop is inert \(Figure[2](https://arxiv.org/html/2608.25315#S1.F2)\) — a mechanism\-specific test rather than another binary verdict\. The audit is organized by a single principle:order\-swap geometry has a canonical site\-asymmetry baseline that must be measured before it is interpreted\.Define a quantity by subtracting two orderings and you get antisymmetry for free, vanishing on repeated arguments for free, and — when the site mismatch acts non\-trivially — a measurable baseline for free\. The Taylor identity is elementary; the positive claim is that this is the correct audit object for the class: it is measured by four single\-intervention passes, falsifiable by a matched site\-separation perturbation, and validated by known\-positive controls\. The six\-family result then shows that the baseline accounts for most of the measured bracket in this setting\.

Concretely, the difference expands to first order as\(J1−J2\)​\(vi−vj\)\(J\_\{1\}\-J\_\{2\}\)\(v\_\{i\}\-v\_\{j\}\), linear in each direction separately, carrying no interaction, and generically non\-zero when the site mismatch acts on the direction difference\. The same pattern recurs at every level we look: the second\-order term contains a self\-curvature piece that is also interaction\-free and is1\.81\.8–5\.2×5\.2\\times*larger*than the zero\-parameter residual at the primary configuration; closed transport loops, widely assumed exempt, carry the same first\-order term; and the natural diagnostic for conjunctive structure returns near\-zero for*any*antisymmetric form at any rank, so its collapse in real data is entailed rather than observed\. Four apparent findings, one cause\.

### Contribution\.

The central contribution is a site\-asymmetry audit for order\-dependent activation statistics, demonstrable without fitting anything\.The zero\-parameter predictionℒ^i​j=\[s1​\(vi\)\+s2​\(vj\)\]−\[s1​\(vj\)\+s2​\(vi\)\]\\widehat\{\\mathcal\{L\}\}\_\{ij\}=\[s\_\{1\}\(v\_\{i\}\)\{\+\}s\_\{2\}\(v\_\{j\}\)\]\-\[s\_\{1\}\(v\_\{j\}\)\{\+\}s\_\{2\}\(v\_\{i\}\)\], built from four single\-injection forward passes and never shown the bracket it predicts, leaves55–8%8\\%of the bracket’s norm unexplained at the primary configuration in all six families, and22–16%16\\%across all eighteen family\-by\-configuration cells \(mean cosine0\.99730\.9973, never below0\.98690\.9869\)\. The cosine and the residual fraction are one measurement in two units, not two agreeing measurements \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\.

The consequence for practice is a partition, and it is what carries past our own models\. Two statistics are in circulation for “do these two interventions interact”: the order\-swap bracketBri​j=hi​j−hj​i\\mathrm\{Br\}\_\{ij\}=h\_\{ij\}\-h\_\{ji\}, and the*second difference*Di​j=hi​j−s1​\(vi\)−s2​\(vj\)D\_\{ij\}=h\_\{ij\}\-s\_\{1\}\(v\_\{i\}\)\-s\_\{2\}\(v\_\{j\}\)\. They differ by exactly the artifact above,Bri​j=\(Di​j−Dj​i\)\+ℒ^i​j\\mathrm\{Br\}\_\{ij\}=\(D\_\{ij\}\-D\_\{ji\}\)\+\\widehat\{\\mathcal\{L\}\}\_\{ij\}\. The identity is a tautology; the magnitude is not\. Across our eighteen cells the first\-order term accounts for84\.384\.3–97\.7%97\.7\\%of the bracket \(mean93\.7%93\.7\\%\), against0%0\\%of the second difference\. Neither half of that contrast is an independent measurement: the first is11minus the residual reported in Section[4](https://arxiv.org/html/2608.25315#S4), the same quantity in complementary units, and the second is exact by the definition ofDD, surviving even a constant map\. This gives a practical stop condition: run the single\-injection prediction before interpretingBr\\mathrm\{Br\}geometrically; if it explains the bracket, the statistic has not identified interaction\. A reader who wants the interaction should formDDrather than correct a bracket to recover it\. This also fixes the status of our own residue, which*is*the antisymmetrized second difference \(Section[2](https://arxiv.org/html/2608.25315#S2)\)\.

Figure 1:The two central results\.\(a\)Norm of the order\-swap bracket against injection coefficientα\\alpha, normalised atα=1\\alpha=1\. All six families trackα1\\alpha^\{1\}\(dashed\), the scaling of a term with no interaction, rather thanα2\\alpha^\{2\}\(dotted\); fitted exponentsγ∈\[0\.886,1\.043\]\\gamma\\in\[0\.886,1\.043\]\.\(b\)The unsigned shared\-argument ratio for the corrected residual under the*confound\-free*prompt\-split design \(solid\), against the generic\-interaction null \(dashed red\), a calibrated trait\-varying first\-order arm, and an uncalibrated constant\-Jacobian floor \(the linear slack, which has no free parameter to calibrate; Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. Stars mark the33of66families that clear the null — Llama1\.5991\.599, OLMo2\.0592\.059, Qwen1\.6421\.642against the1\.3241\.324bar; DeepSeek, Gemma and Mistral fall to linear\-slack level\. The dotted outline is the same statistic on the full sample, where all pairs share one prompt set: the gap is what the shared\-prompt confound was worth\. We plot the number we claim, not the more favourable full\-sample one\. The trait\-varying arm is clipped at the axis top and annotated with its true value\.Figure 2:Three\-dimensional schematic of the measurement, not model data\.\(a\)Opposite site orders give different Jacobian\-weighted paths and an endpoint chordBri​j\\mathrm\{Br\}\_\{ij\}\.\(b\)The tangent\-plane predictionℒ^\\widehat\{\\mathcal\{L\}\}is separated from the second\-difference curvatureDi​jD\_\{ij\}\.\(c\)Four legs can have zero net injected vector but a nonzero first\-order readout residual when site Jacobians differ\. Endpoint order and loop closure are diagnostics to measure, not evidence of interaction by themselves\.Second, the correction people would apply does not isolate an interaction either\.Expanding one order further splits theα2\\alpha^\{2\}coefficient in two: a genuinely mixed termℳ\\mathcal\{M\}, and a self\-curvature term𝒮\\mathcal\{S\}that is quadratic in each direction*separately*and carries no coupling\. We measure𝒮\\mathcal\{S\}with no fit\. It is1\.81\.8–5\.2×5\.2\\times*larger*than the zero\-parameter residual at the primary configuration, so‖Q‖/‖ℒ‖\\\|Q\\\|/\\\|\\mathcal\{L\}\\\|from a two\-term fit is not an interaction share: the same artifact recurring one order up\.[Piontkovskaia and Nikolenko \(2026\)](https://arxiv.org/html/2608.25315#bib.bib25)separate these second\-order objects at a*shared*base point; what is new is that when the sites differ, self\-curvature contaminates the order\-swap coefficient itself throughH11−H22H\_\{11\}\-H\_\{22\}\.

*Supporting measurements\.*Two further routes agree by different apparatus that the statistic is first\-order dominated \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\), and a randomly initialised network reproduces the same agreement, fixing what the six\-family measurement establishes: not that trained models have a property, but that they do not escape a generic one\. What survives correction clears a generic\-interaction null in33of66families under the confound\-free design, which we report as a negative result\.

## 2Related Work

[Vaidyanathan et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib1)prove the second\-difference interaction equals the Hessian bilinear form, vanishing for locally affine maps; that object is our𝒬\\mathcal\{Q\}, and we take the theory as given\. Note what it entails: a bilinear form applies to argument pairs it was never fitted on, so transfer to unseen arguments is a property of the model class, not a discovery\.[Skifstad et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib24)and[Piontkovskaia and Nikolenko \(2026\)](https://arxiv.org/html/2608.25315#bib.bib25)both support the premise that a first\-order term dominates here, the first finding layer\-wise LLM dynamics well approximated by locally\-linear models, the second finding single perturbations first\-order predictable across nine transformers while pairwise composition has no stable radius\.[Wu et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib27)make the same move for a different statistic: apparent scale\-dependent steerability across 17 models turns out to be produced by an uncalibrated pipeline;[Heap et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib32)do the same for SAE auto\-interpretability with a randomised baseline, the closest neighbour to our untrained null \(Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. The closest methodological neighbour is[Zhang and Wang \(2026\)](https://arxiv.org/html/2608.25315#bib.bib5), who find attribution patching’s first\-order approximation unreliable, trace the error to downstream non\-linearity, and supply a correction: the template is ours; the object is not\.[VanderWeele \(2014\)](https://arxiv.org/html/2608.25315#bib.bib19)provides a causal\-inference analogue of theℒ\+Q\\mathcal\{L\}\+Qsplit, and[Dooms et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib26)independently find low\-rank quadratic structure prevalent in LLM activations — which would be the natural home for the interaction we are looking for, though we do not find it there: the operator we recover is not low\-rank once the estimator’s own rank compression is accounted for \(the full operator comparison is in the artifact\)\.

### Weight\-space and neighbouring statistics\.

[Piontkovskaia and Nikolenko \(2026\)](https://arxiv.org/html/2608.25315#bib.bib25)measure the analogous Lie bracket for sequential task\-gradient steps, while[Sweeney \(2026\)](https://arxiv.org/html/2608.25315#bib.bib2)use a weight\-space bracket as an ordering statistic; both are outside our activation\-space criterion\. The two spaces admit a first\-order correspondence under the conditions analyzed by[Adila et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib29), so agreement across them is closer to entailment than to an independent replication\. Our untrained null and self\-term split are the additional controls\. Among activation\-space neighbours,[Richards \(2026\)](https://arxiv.org/html/2608.25315#bib.bib3)use one fixed layer window,[Sevetlidis and Pavlidis \(2026\)](https://arxiv.org/html/2608.25315#bib.bib4)prove an input\-space affine null, and[Mudarisov et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib30)use a same\-site bracket where the first\-order term cancels \(the full comparison is in the artifact\)\.[Vaidyanathan et al\. \(2026\)](https://arxiv.org/html/2608.25315#bib.bib1)and[Khemais \(2026\)](https://arxiv.org/html/2608.25315#bib.bib33)instead form the second difference, which is first\-order\-free\. Composable interventions expose order effects without forming this statistic \([Kolbeinsson et al\. \(2025\)](https://arxiv.org/html/2608.25315#bib.bib21); task arithmetic\([Ilharco et al\., 2023](https://arxiv.org/html/2608.25315#bib.bib20)\)\);[Ortiz\-Jimenez et al\. \(2023\)](https://arxiv.org/html/2608.25315#bib.bib31)study the closest weight\-space first\-order question\. Exemption is conditional\. The second difference cancels first\-order content by construction; a raw bracket at a nominally shared site requires exactJ1=J2J\_\{1\}=J\_\{2\}, which is not stable\. Because contamination is first order in the edit while interaction is second order, even a1%1\\%Jacobian mismatch leaves the former at99\.6%99\.6\\%of the bracket atα=0\.01\\alpha=0\.01\. Closure is not the criterion; site structure is \(Section[3](https://arxiv.org/html/2608.25315#S3)\), and the artifact records the neighbouring statistics treated outside its scope\.

## 3Method: the estimator and its first\-order term

### Where the artifact bites\.

A measurement is affected when the two interventions enter at sites with different downstream maps: order\-swap brackets, sequential\-edit differences, and anyA→BA\\\!\\to\\\!BversusB→AB\\\!\\to\\\!Acomparison at distinct depths\. Closure is not itself a defence: the first\-order terms cancel pairwise only when each return leg re\-enters at its outgoing site, and that composition is the identity map\.

LetAi\(ℓ\)A\_\{i\}^\{\(\\ell\)\}add directionviv\_\{i\}to the residual stream at layerℓ\\ell, and leth\(ℓout\)h^\{\(\\ell\_\{\\mathrm\{out\}\}\)\}read out at a later layer\. Forℓ1<ℓ2\\ell\_\{1\}<\\ell\_\{2\}the order\-swap bracket isBri​j=𝔼p​\[h\(ℓout\)​\(Aj\(ℓ2\)​Ai\(ℓ1\)​p\)−h\(ℓout\)​\(Ai\(ℓ2\)​Aj\(ℓ1\)​p\)\]\\mathrm\{Br\}\_\{ij\}=\\mathbb\{E\}\_\{p\}\[\\,h^\{\(\\ell\_\{\\mathrm\{out\}\}\)\}\(A\_\{j\}^\{\(\\ell\_\{2\}\)\}A\_\{i\}^\{\(\\ell\_\{1\}\)\}p\)\-h^\{\(\\ell\_\{\\mathrm\{out\}\}\)\}\(A\_\{i\}^\{\(\\ell\_\{2\}\)\}A\_\{j\}^\{\(\\ell\_\{1\}\)\}p\)\]\. Expanding about the unsteered activations and carrying the second order out in full,

Bri​j=\(J1−J2\)​\(vi−vj\)⏟ℒ,first order, no interaction\+12​\[H11​\(vi,vi\)−H11​\(vj,vj\)\+H22​\(vj,vj\)−H22​\(vi,vi\)\]⏟𝒮,second order,*no interaction*\+H12​\(vi,vj\)−H12​\(vj,vi\)⏟ℳ,the interaction\+O⁡\(α3\),\\begin\{split\}\\mathrm\{Br\}\_\{ij\}\\;=\\;&\\underbrace\{\(J\_\{1\}\-J\_\{2\}\)\(v\_\{i\}\-v\_\{j\}\)\}\_\{\\mathcal\{L\},\\ \\text\{first order, no interaction\}\}\\;\+\\;\\underbrace\{\\tfrac\{1\}\{2\}\\big\[H\_\{11\}\(v\_\{i\},v\_\{i\}\)\-H\_\{11\}\(v\_\{j\},v\_\{j\}\)\+H\_\{22\}\(v\_\{j\},v\_\{j\}\)\-H\_\{22\}\(v\_\{i\},v\_\{i\}\)\\big\]\}\_\{\\mathcal\{S\},\\ \\text\{second order, \\emph\{no interaction\}\}\}\\\\\[4\.0pt\] &\+\\;\\underbrace\{H\_\{12\}\(v\_\{i\},v\_\{j\}\)\-H\_\{12\}\(v\_\{j\},v\_\{i\}\)\}\_\{\\mathcal\{M\},\\ \\text\{the interaction\}\}\\;\+\\;O\(\\alpha^\{3\}\),\\end\{split\}\(1\)whereHk​kH\_\{kk\}is the curvature of the readout in the perturbation entering at sitekkandH12H\_\{12\}the mixed second derivative\.

###### Observation 1\(site\-asymmetry audit identity\)\.

For matched single\-site responses, the canonical additive baseline isℒ^i​j=\[s1​\(vi\)\+s2​\(vj\)\]−\[s1​\(vj\)\+s2​\(vi\)\]\\widehat\{\\mathcal\{L\}\}\_\{ij\}=\[s\_\{1\}\(v\_\{i\}\)\+s\_\{2\}\(v\_\{j\}\)\]\-\[s\_\{1\}\(v\_\{j\}\)\+s\_\{2\}\(v\_\{i\}\)\], andBri​j−ℒ^i​j=Di​j−Dj​i\\mathrm\{Br\}\_\{ij\}\-\\widehat\{\\mathcal\{L\}\}\_\{ij\}=D\_\{ij\}\-D\_\{ji\}\. Consequently, to second order the residual removes both the site\-Jacobian and pure self\-curvature terms; what remains is the antisymmetrized mixed derivative plusO⁡\(α3\)O\(\\alpha^\{3\}\)\.

ForC3C^\{3\}readouts, Appendix[C](https://arxiv.org/html/2608.25315#A3)gives an explicit third\-derivative norm certificate for theO⁡\(α3\)O\(\\alpha^\{3\}\)remainder\. We do not claim empirical bounds on those derivatives for the six LLMs, so their residuals remain finite\-scale mixtures unless such bounds are supplied\.

### The second\-order term is not all interaction\.

𝒮\\mathcal\{S\}is antisymmetric under exchange, vanishes whenvi=vjv\_\{i\}=v\_\{j\}, and is non\-zero exactly when the two sites have different curvature — every surface property that makes the statistic look like a bracket\. But it is quadratic in each argument*separately*, with no interaction between them: the first\-order artifact one order up\. Onlyℳ\\mathcal\{M\}couples the two directions\. Writing theα2\\alpha^\{2\}coefficient as a single “bilinear” object𝒬⁡\(vi,vj\)\\mathcal\{Q\}\(v\_\{i\},v\_\{j\}\), which is the natural thing to write, silently merges𝒮\\mathcal\{S\}into the interaction and repeats at second order the exact conflation the paper exists to correct\. Two consequences: a two\-term fitα​ℒ\+α2​𝒬\\alpha\\mathcal\{L\}\+\\alpha^\{2\}\\mathcal\{Q\}estimates𝒮\+ℳ\\mathcal\{S\}\+\\mathcal\{M\}together, so‖𝒬‖/‖ℒ‖\\\|\\mathcal\{Q\}\\\|/\\\|\\mathcal\{L\}\\\|in the injection\-scale fit is*not*an interaction share; and Observation[2](https://arxiv.org/html/2608.25315#Thmobservation2)guarantees only that the fittedW⁡\(vi−vj\)W\(v\_\{i\}\-v\_\{j\}\)removes*linear*structure, which is weaker, because𝒮\\mathcal\{S\}is nonlinear too\.

###### Observation 2\(elementary\)\.

Bri​j=−Brj​i\\mathrm\{Br\}\_\{ij\}=\-\\mathrm\{Br\}\_\{ji\}by construction\. Any linear modelM1​vi\+M2​vjM\_\{1\}v\_\{i\}\+M\_\{2\}v\_\{j\}that is antisymmetric under exchange requiresM1=−M2M\_\{1\}=\-M\_\{2\}, hence has the formW⁡\(vi−vj\)W\(v\_\{i\}\-v\_\{j\}\)\.

Fitting and subtracting the bestW⁡\(vi−vj\)W\(v\_\{i\}\-v\_\{j\}\)therefore removes all linear structure in the model class\. This is the elementary symmetric/antisymmetric splitting of theS2S\_\{2\}action on𝒯⊕𝒯\\mathcal\{T\}\\oplus\\mathcal\{T\}, stated because it dictates the correct control rather than as a result\. It is not exact for the*estimate*: withW^\\widehat\{W\}fitted from finite data the residual retains\(W−W^\)​\(vi−vj\)\(W\-\\widehat\{W\}\)\(v\_\{i\}\-v\_\{j\}\), and Section[4](https://arxiv.org/html/2608.25315#S4)tests that this slack does not manufacture our result\.

Three natural controls do*not*removeℒ\\mathcal\{L\}\.\(C1\)Projecting offspan⁡\(vi,vj\)\\mathrm\{span\}\(v\_\{i\},v\_\{j\}\):ℒ\\mathcal\{L\}lands outside that span at the readout\.\(C2\)ℒ\\mathcal\{L\}is antisymmetric, so brackets sharing a direction in*swapped*argument slots are anti\-correlated; a statistic that pools argument positions averages\+\+against−\-\.\(C3\)ℒ\\mathcal\{L\}alone reproduces which pair a bracket came from, so pair\-identifiability is no evidence of interaction either\.

### Measuring the first\-order term instead of fitting it\.

Both arguments above are indirect, and the fitted correction of Observation[2](https://arxiv.org/html/2608.25315#Thmobservation2)carries16×Dout16\\times D\_\{\\mathrm\{out\}\}free parameters, so a high explained variance is partly a statement about capacity\. We therefore measureJ1J\_\{1\}andJ2J\_\{2\}rather than fitting them\. The single\-injection responsesk​\(v\)=𝔼p​\[h\(ℓout\)​\(Av\(ℓk\)​p\)−h\(ℓout\)​\(p\)\]=Jk​v\+12​Hk​k​\(v,v\)\+O⁡\(α3\)s\_\{k\}\(v\)=\\mathbb\{E\}\_\{p\}\[h^\{\(\\ell\_\{\\mathrm\{out\}\}\)\}\(A\_\{v\}^\{\(\\ell\_\{k\}\)\}p\)\-h^\{\(\\ell\_\{\\mathrm\{out\}\}\)\}\(p\)\]=J\_\{k\}v\+\\tfrac\{1\}\{2\}H\_\{kk\}\(v,v\)\+O\(\\alpha^\{3\}\)costs one forward pass, andℒ^i​j=\[s1​\(vi\)\+s2​\(vj\)\]−\[s1​\(vj\)\+s2​\(vi\)\]\\widehat\{\\mathcal\{L\}\}\_\{ij\}=\[s\_\{1\}\(v\_\{i\}\)\+s\_\{2\}\(v\_\{j\}\)\]\-\[s\_\{1\}\(v\_\{j\}\)\+s\_\{2\}\(v\_\{i\}\)\]predicts the bracket with*no fitted parameters*; it never seesBri​j\\mathrm\{Br\}\_\{ij\}\. Expanding to second order, every Jacobian and every pure self\-term cancels, leavingBri​j−ℒ^i​j=H12​\(vi,vj\)−H12​\(vj,vi\)\+O⁡\(α3\)\\mathrm\{Br\}\_\{ij\}\-\\widehat\{\\mathcal\{L\}\}\_\{ij\}=H\_\{12\}\(v\_\{i\},v\_\{j\}\)\-H\_\{12\}\(v\_\{j\},v\_\{i\}\)\+O\(\\alpha^\{3\}\): the antisymmetrized*mixed*second derivative at second order, plus an uncontrolled higher\-order remainder at finite injection scale\. Thus this route targets the same interaction term as the fitted correction, but its finite\-scale residual is not an exact interaction estimate\. The estimator is falsifiable: the symmetric combination should not predict an antisymmetric object, and a bracket paired with a*different*pair’s responses should not be predicted at all\. \(Exchanging the two site labels negatesℒ^\\widehat\{\\mathcal\{L\}\}identically, so that check tests our arithmetic, not the model\.\) Appendix[C](https://arxiv.org/html/2608.25315#A3)validates it on surrogates with knownJ1,J2,H11,H22,H12J\_\{1\},J\_\{2\},H\_\{11\},H\_\{22\},H\_\{12\}\.

### Known\-positive control\.

On known\-Hessian surrogates, the no\-interaction arm recovers the bracket exactly; a pure mixed arm withJ1=J2J\_\{1\}=J\_\{2\}givesℒ^=0\\widehat\{\\mathcal\{L\}\}=0and residual fraction11; and mixed arms with self\-curvature recover the planted interaction share to machine precision at three strengths \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\. The model\-side residual remains a finite\-scale mixture, not an interaction estimate\.

### Closed loops: site structure, not closure\.

A holonomy statistic is often assumed exempt because its net injection is zero\. That argument is wrong: it uses two Jacobians for four injections, so the loop is the identity map and the statistic is identically zero, leaving nothing to exempt\. A genuine loop through four distinct depths gives\(J1−J3\)​vi\+\(J2−J4\)​vj\(J\_\{1\}\-J\_\{3\}\)v\_\{i\}\+\(J\_\{2\}\-J\_\{4\}\)v\_\{j\}, which does not vanish and carries no interaction \(the loop checks are in the artifact\)\. Exemption is earned from site structure, not assumed from closure\.

### Models, directions, and protocol\.

Six open\-weight base families at 7–9B: DeepSeek\-LLM\-7B, Gemma\-2\-9B, Llama\-3\.1\-8B, Mistral\-7B\-v0\.3, OLMo\-7B\-0724 and Qwen2\.5\-7B\. Steering directions are contrastive activation additions\([Rimsky et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib8);[Turner et al\., 2023](https://arxiv.org/html/2608.25315#bib.bib7)\)over 16 trait contrasts, giving 120 unordered pairs per family, with two extraction seeds from disjoint prompt samples so that cross\-seed agreement is a replication rather than a re\-read of one fit\. Since steering\-vector reliability is itself contested\([Tan et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib9)\), we measured it: the same trait’s direction reproduces across seeds at mean cosine0\.7840\.784\(0\.6670\.667–0\.8330\.833\), which bounds any*cross\-seed*statistic here; within\-run comparisons such ascos⁡\(ℒ^,Br\)\\cos\(\\widehat\{\\mathcal\{L\}\},\\mathrm\{Br\}\)use the same directions on both sides, so direction noise is common\-mode and is not bounded by it\. Unless stated otherwise we report the primary layer pair at fractional depths\(0\.25,0\.50\)\(0\.25,0\.50\), read out three layers downstream, over 48 held\-out prompts, identically across families with no per\-family tuning of the layer pairs \(Appendix[A](https://arxiv.org/html/2608.25315#A1)\)\. The accompanying reviewer artifact is available at[the anonymous repository](https://anonymous.4open.science/r/non-abelian-persona-composition-artifact-94F0/)and contains the source, figures, per\-experiment records, validation scripts, and an independent CPU reference with raw held\-out outputs; every numeric claim is bound to a named record field by a checking script that fails on any value it cannot trace, is itself decoy\-tested, and the single\-injection responses behind Appendix[C](https://arxiv.org/html/2608.25315#A3)are released as scalar summaries rather than tensors \(Appendix[A](https://arxiv.org/html/2608.25315#A1)\)\.

## 4Results: validating the audit

### Validating the additive baseline\.

Buildingℒ^i​j\\widehat\{\\mathcal\{L\}\}\_\{ij\}from single\-injection responses \(Section[3](https://arxiv.org/html/2608.25315#S3)\) and comparing it against the measured bracket givescos⁡\(ℒ^,Br\)∈\[0\.9955,0\.9986\]\\cos\(\\widehat\{\\mathcal\{L\}\},\\mathrm\{Br\}\)\\in\[0\.9955,0\.9986\]at the primary layer pair in all six families, leaving a residual of5\.25\.2–7\.6%7\.6\\%of the bracket’s norm; over all eighteen family\-by\-configuration cells the cosine never falls below0\.98690\.9869\. The informative control behaves as Section[3](https://arxiv.org/html/2608.25315#S3)requires: the symmetric combination does not predict the bracket, signed per\-cell mean\|cossym\|=0\.124\|\\cos\_\{\\mathrm\{sym\}\}\|=0\.124\(per\-pair magnitudes are larger and we do not claim otherwise, Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\. A second control answers the circularity objection: sinceℒ^\\widehat\{\\mathcal\{L\}\}is built from the same four responses as the bracket, a high cosine might be entailed\. It is not, and the strongest form of the objection — drawing the bracket inside the span of those same four responses — reaches the observed agreement in0\.06%0\.06\\%of20,00020\{,\}000draws \(Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. This separates capacity from mechanism, because the fitted correction leaves2\.92\.9–9\.9%9\.9\\%of the bracket while the zero\-parameter route leaves the5\.25\.2–7\.6%7\.6\\%above\. Those are*not*the same statistic — the first is pooled over pairs, the second a mean of per\-pair ratios — so we do not present their overlap as agreement\. On the matched pooled statistic the two routes bracket the first\-order share from opposite directions and land within15%15\\%of each other \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\), which is not the same as agreeing\.

### Measuring the no\-interaction second\-order term\.

𝒮\\mathcal\{S\}is measurable with no fit, from the same apparatus: injecting−v\-vas well as\+v\+vseparates the two orders by parity\. The split is exact on surrogates with knownH11,H22H\_\{11\},H\_\{22\}, and the closure ofℒ^=ℒ\+𝒮\\widehat\{\\mathcal\{L\}\}=\\mathcal\{L\}\+\\mathcal\{S\}is algebra rather than evidence \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\.

The result is unfavourable to the operator claim\.Measured on Llama and OLMo across three injection configurations,‖𝒮‖/‖Br‖\\\|\\mathcal\{S\}\\\|/\\\|\\mathrm\{Br\}\\\|runs0\.0920\.092–0\.3380\.338\. Against the zero\-parameter residual on the same cells, the same statistic on both sides of the ratio,*the no\-interaction second\-order term is1\.81\.8–5\.2×5\.2\\timeslarger than the operator extracted from underneath it*at the primary configuration, and1\.81\.8–11\.2×11\.2\\timesacross all three \(Table[1](https://arxiv.org/html/2608.25315#A3.T1)\)\. A quantity carrying no coupling between the two directions is the larger part of what a two\-term fit reports as “bilinear”\. We do not divide it by the fitted\-correction range2\.92\.9–9\.9%9\.9\\%, which is pooled over pairs rather than a mean of per\-pair ratios: the caution above applies to our own second result too\.

This does not invalidate the operator analysis, because𝒮\\mathcal\{S\}cancels identically inBr−ℒ^\\mathrm\{Br\}\-\\widehat\{\\mathcal\{L\}\}: the zero\-parameter prediction is built from single\-injection responses that carry the self\-curvature at both sites\. It does invalidate reading theα2\\alpha^\{2\}coefficient of a two\-term fit as an interaction share: the fittedW⁡\(vi−vj\)W\(v\_\{i\}\-v\_\{j\}\)removes linear structure only, so a residual that is “purely nonlinear” can be mostly a term with no interaction in it\.

### The untrained null, and what the measurement is evidence for\.

A random\-init, untrained residual network with1616layers andd=768d=768givescos⁡\(ℒ^,Br\)∈\[0\.9956,0\.9988\]\\cos\(\\widehat\{\\mathcal\{L\}\},\\mathrm\{Br\}\)\\in\[0\.9956,0\.9988\]as the injection\-to\-stream ratio sweepsε=0\.01\\varepsilon=0\.01–44\(worst atε=1\\varepsilon=1\), and passes the symmetric control \(\|cossym\|=0\.033\|\\cos\_\{\\mathrm\{sym\}\}\|=0\.033\)\. This validates the architecture\-level baseline: first\-order dominance follows from differing downstream Jacobians, not training\. The trained norm ratio was not recorded, so this is a range, not a matched comparison\.

An independent CPU residual MLP, ViT\-B/16, and ResNet\-50 reference is released with raw held\-out outputs; the vision residual falls below a fixed Gaussian\-direction null in11/1211/12contrasts \(5/65/6ViT,6/66/6ResNet\)\. Unlike the random\-init network null above, this keeps the trained network fixed and randomises directions; it is portability evidence, not pooled evidence \(Appendix[A](https://arxiv.org/html/2608.25315#A1);experiments/external/vision\_summary\.json\)\.

### Mechanism\-specific generalization\.

The audit predicts a graded effect, not merely a yes/no verdict: widening one site while holding the other site and readout fixed should increase the site\-asymmetry share\. In the strictly matched contrast, the residual fraction falls in66of66families, and91\.7%91\.7\\%of720720matched trait pairs move in the predicted direction; matched random residual networks reach only5454–58%58\\%\. Depth alone and the confounded contrast do not reproduce this pattern \(Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. This is a positive mechanism test, not another restatement of the cosine\. A separate injection sweep givesγ∈\[0\.886,1\.043\]\\gamma\\in\[0\.886,1\.043\]forBr⁡\(α\)=α​ℒ\+α2​Q\\mathrm\{Br\}\(\\alpha\)=\\alpha\\mathcal\{L\}\+\\alpha^\{2\}Q, excluding theγ=2\\gamma=2scaling of a pure bilinear term; the random architecture also givesγ≈1\\gamma\\approx 1, so this route validates the audit rather than learned composition \(Figure[1](https://arxiv.org/html/2608.25315#S1.F1)a; the full sweep is in the artifact\)\.

### Delimiting the claim\.

The test asks whether the corrected residual behaves like a pair\-specific interaction or like leftover first\-order structure\. We score it by the*shared\-argument ratio*: the mean\|cos\|\|\\cos\|between residuals of pairs sharing one trait, over the same quantity for pairs sharing none\. If the residual couples its two arguments, pairs sharing one should resemble each other more\. The*generic\-interaction null*sets the bar: an antisymmetric bilinear form with no model in it, on the family’s own empirical directions, gives1\.3241\.324\(Table[3](https://arxiv.org/html/2608.25315#A4.T3); Appendices[B](https://arxiv.org/html/2608.25315#A2)and[D](https://arxiv.org/html/2608.25315#A4)\)\.Under a prompt\-split design that removes the shared\-prompt confound, the corrected residual clears that null in 3 of 6 families— Llama, OLMo and Qwen — while DeepSeek, Gemma and Mistral fall to1\.0101\.010\(DeepSeek\) to1\.0171\.017\(Gemma\), within0\.020\.02of pure linear slack\. The weaker cross\-seed design gives5/65/6, but it does not break the confound, because seeds differ in direction extraction and not in prompts\. We claim the3/63/6\. One diagnostic must be discounted outright: the signed shared\-argument statistic collapses to near zero after correction, and that collapse is*entailed*by antisymmetry rather than observed, since a maximally overlap\-dependent rank\-one form reproduces it\. Appendix[B](https://arxiv.org/html/2608.25315#A2)gives the stratification and the seed construction, the artifact per\-family records, and Appendix[D](https://arxiv.org/html/2608.25315#A4)the three nulls and the interval estimates\.

### What the audit leaves\.

The audit returns a residual rather than forcing a verdict: it retains2\.92\.9–9\.9%9\.9\\%of the raw bracket’s norm, smaller than the no\-interaction second\-order term measured beside it \(1\.81\.8–5\.2×5\.2\\timesat the primary configuration, Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\. It is an estimable bilinear form and not an echo of trait geometry; against a behavioural readout that cancels additive contributions it beats the raw bracket in4 of 6families\. The behavioural direction survives multiplicity correction in22of66, so the audit identifies a candidate interaction without overstating its stability \(Appendix[D](https://arxiv.org/html/2608.25315#A4); per\-family behavioural records are in the artifact\)\. We scope it to distinct\-site activation interventions, not a universal theory of representation geometry\.

## Appendix AProtocol details

### Exact checkpoints\.

All are base \(non\-instruct\) checkpoints, one per line:

deepseek\-ai/deepseek\-llm\-7b\-base google/gemma\-2\-9b meta\-llama/Llama\-3\.1\-8B mistralai/Mistral\-7B\-v0\.3 allenai/OLMo\-7B\-0724\-hf Qwen/Qwen2\.5\-7B

### What is and is not reproducible from the release\.

Most of the analysis is CPU\-only and runs from the committed per\-pair tensors\. The single\-injection responses behind Appendix[C](https://arxiv.org/html/2608.25315#A3)are released as scalar summaries and not as tensors, so Table[2](https://arxiv.org/html/2608.25315#A3.T2)is reproducible only by re\-running the forward passes\. The independent CPU reference is fully rerunnable from its released raw held\-out outputs withscripts/check\_external\_audit\.py; it is a portability and artifact check, not an additional pooled LLM result\.

The three\-family extension is fully specified in the external protocolnonlm\_protocol\.md\. It uses the CPU residual MLP, torchvision ViT\-B/16 and torchvision ResNet\-50 with fixed public checkpoints, class\-balanced direction extraction, disjoint held\-out readout images, all 45 unordered direction pairs, and a fixed Gaussian direction null\. ViT sites are post\-block residual streams \(including direction site 0\); ResNet sites are the five post\-block stage\-3 residual streams\. The vision records include the complete raw bracket, linear\-estimator and residual arrays, model URLs, data provenance, and GPU telemetry\. Exact aggregate values are recomputed byscripts/summarize\_external\_references\.py, whilescripts/check\_external\_vit\_audit\.pyverifies each vision record against its raw arrays\.

Residual width isDout∈\{3584,4096\}D\_\{\\mathrm\{out\}\}\\in\\\{3584,4096\\\}across the six families\. The 16 trait contrasts give\(162\)=120\\binom\{16\}\{2\}=120unordered pairs per family\. Extraction seeds are drawn from disjoint prompt samples: seed agreement therefore reflects a re\-extraction of the direction, not a re\-read of a single fit, which is what makes the cross\-seed comparison a replication\. The measured across\-seed direction reliability \(mean cosine0\.7840\.784, range0\.6670\.667–0\.8330\.833\) is a baseline for downstream agreement, not a formal upper bound, since downstream statistics can transform or normalize direction errors\. Layer pairs are specified as fractional depths so that they are comparable across families of different depth; the primary pair is\(0\.25,0\.50\)\(0\.25,0\.50\)with readout three layers below the deeper injection site, and the two secondary configurations are\(0\.25,0\.75\)\(0\.25,0\.75\)and\(0\.50,0\.75\)\(0\.50,0\.75\)\. All statistics use 48 held\-out prompts, disjoint from the prompts used for direction extraction\.

### Numeric provenance\.

Every number printed in this paper is checked against the committed result records by a script that binds each printed value to a*named*record field, and fails on any value it cannot trace\. The script also runs a decoy test on itself each time it is invoked: it perturbs every number in the paper and confirms the perturbed versions are rejected\. This matters because an earlier version matched on value proximity alone and certified roughly half of all deliberately corrupted numbers\. The current false\-certification rate is under10%10\\%\. We report the counts of traced, unbound and untraced numbers in the released log rather than summarizing them as a pass\. We also verified that the injection\-scale exponents recompute from the raw sweeps to four decimals\.

## Appendix BDelimiting the claim: full treatment

We form a cross\-seed cosine matrix stratified by overlap and argument position:same,sh1\-same,sh1\-crossed, andshare0\. Leave\-one\-pair\-out residuals are positive scalar multiples of in\-sample residuals, so this procedure is not an out\-of\-sample safeguard\. The signed corrected statistic is near zero \(−0\.0100\-0\.0100mean; coherence0\.0700\.070\), but that collapse is entailed by antisymmetry and is not evidence\.

The discriminating statistic isRsh=mean⁡\|cos\|​\(sh1\-same\)/mean⁡\|cos\|​\(share0\)R\_\{\\rm sh\}=\\operatorname\{mean\}\|\\cos\|\(\\textsc\{sh1\-same\}\)/\\operatorname\{mean\}\|\\cos\|\(\\textsc\{share0\}\)\. The observed value is2\.1322\.132, versus0\.9990\.999for constant\-Jacobian slack,1\.3241\.324for a calibrated generic antisymmetric interaction, and6\.0586\.058for a trait\-varying first\-order Jacobian\. Trait bootstrap and delete\-one\-trait jackknife intervals exclude the generic null in five of six families; the disjoint\-trait and third\-seed checks are conditional on the three families they cover\.

The shared\-prompt confound is addressed by recomputing brackets on disjoint prompt halves\. The mean falls2\.132→1\.602→1\.3912\.132\\to 1\.602\\to 1\.391as prompts are halved and then disjointed; only Llama, OLMo, and Qwen clear the recalibrated null \(1\.3251\.325\)\. We therefore claim the effect in3/63/6families under the confound\-free design, with5/65/6retained only as the weaker cross\-seed result\. Intersecting configuration, prompt, trait, and seed controls leaves Llama and OLMo\. The full records and per\-family strata are in the reviewer artifact\.

## Appendix CThe zero\-parameter first\-order measurement

###### Proposition 1\(finite\-scale remainder certificate\)\.

LetF⁡\(a,b\)F\(a,b\)be the averaged two\-site readout and letfk​\(a\)f\_\{k\}\(a\)be its single\-site restriction\. IfF,f1,f2F,f\_\{1\},f\_\{2\}areC3C^\{3\}on the line segments used by the injections, with third\-derivative operator\-norm boundsM,M1,M2M,M\_\{1\},M\_\{2\}, then

Bri​j​\(α\)−ℒ^i​j​\(α\)\\displaystyle\\mathrm\{Br\}\_\{ij\}\(\\alpha\)\-\\widehat\{\\mathcal\{L\}\}\_\{ij\}\(\\alpha\)=α2​ℳi​j\+ℛi​j​\(α\),\\displaystyle=\\alpha^\{2\}\\mathcal\{M\}\_\{ij\}\+\\mathcal\{R\}\_\{ij\}\(\\alpha\),\(2\)‖ℛi​j​\(α\)‖\\displaystyle\\\|\\mathcal\{R\}\_\{ij\}\(\\alpha\)\\\|≤α36​\[2​M​\(‖vi‖2\+‖vj‖2\)3/2\+\(M1\+M2\)​\(‖vi‖3\+‖vj‖3\)\]\.\\displaystyle\\leq\\frac\{\\alpha^\{3\}\}\{6\}\\\!\\left\[2M\(\\\|v\_\{i\}\\\|^\{2\}\+\\\|v\_\{j\}\\\|^\{2\}\)^\{3/2\}\+\(M\_\{1\}\+M\_\{2\}\)\(\\\|v\_\{i\}\\\|^\{3\}\+\\\|v\_\{j\}\\\|^\{3\}\)\\right\]\.

The certificate follows by applying the third\-order Taylor remainder to the two orderings and the four single\-site paths, then using the triangle inequality\. It makes the finite\-scale caveat quantitative; without empirical upper bounds onM,M1,M2M,M\_\{1\},M\_\{2\}, the model\-side residual is a mixture rather than a certified interaction estimate\.

### The second\-order split on models\.

Table[1](https://arxiv.org/html/2608.25315#A3.T1)gives the±\\pm\-split decomposition for the two families where we ran it\. The identityℒ^=ℒ\+𝒮\\widehat\{\\mathcal\{L\}\}=\\mathcal\{L\}\+\\mathcal\{S\}is exact by construction — substituting the parity definitionsℒk=\[sk​\(v\)−sk​\(−v\)\]/2\\mathcal\{L\}\_\{k\}=\[s\_\{k\}\(v\)\{\-\}s\_\{k\}\(\-v\)\]/2and𝒮k=\[sk​\(v\)\+sk​\(−v\)\]/2\\mathcal\{S\}\_\{k\}=\[s\_\{k\}\(v\)\{\+\}s\_\{k\}\(\-v\)\]/2givesℒk\+𝒮k=sk​\(v\)\\mathcal\{L\}\_\{k\}\+\\mathcal\{S\}\_\{k\}=s\_\{k\}\(v\)termwise — so it holds for any numbers whatsoever and is*not*evidence that the split is clean\. We record the closure only as an arithmetic self\-check on the implementation: it sits at1\.9×10−71\.9\\times 10^\{\-7\}relative, which is float32 rounding and nothing more \(the same computation on random float32 input gives6×10−86\\times 10^\{\-8\}\)\. The no\-interaction self term is9\.29\.2–33\.8%33\.8\\%of the bracket’s norm, against a corrected residual of2\.92\.9–9\.9%9\.9\\%: it is the larger object\. Removing only the first\-order part leaves1010–34%34\\%of the bracket \(column*resid,ℒ\\mathcal\{L\}only*\), and it is the additional subtraction of𝒮\\mathcal\{S\}that takes the residual down to a few percent\. Reading theα2\\alpha^\{2\}coefficient as an interaction share would therefore attribute a quantity with no coupling in it to the operator\.

Table 1:Second\-order split from the±\\pminjection design, two families, 120 pairs, three injection configurations\.‖𝒮‖/‖Br‖\\\|\\mathcal\{S\}\\\|/\\\|\\mathrm\{Br\}\\\|is the no\-interaction second\-order term12​\(H11−H22\)​\[\(vi,vi\)−\(vj,vj\)\]\\tfrac\{1\}\{2\}\(H\_\{11\}\-H\_\{22\}\)\[\(v\_\{i\},v\_\{i\}\)\-\(v\_\{j\},v\_\{j\}\)\]as a fraction of the bracket\.*resid,ℒ\\mathcal\{L\}only*removes the first\-order term alone;*resid, full*removesℒ^=ℒ\+𝒮\\widehat\{\\mathcal\{L\}\}=\\mathcal\{L\}\+\\mathcal\{S\}\. The gap between them is what𝒮\\mathcal\{S\}contributes\.
### Results\.

Table[2](https://arxiv.org/html/2608.25315#A3.T2)gives the per\-family values\. The prediction tracks the bracket in every family and every injection configuration, and the two controls behave as the theory requires throughout\. The symmetric combination does not predict the bracket: the signed per\-cell meancossym\\cos\_\{\\mathrm\{sym\}\}averages0\.1240\.124in absolute value, so there is no consistent alignment, though we note the per\-*pair*mean\|cossym\|\|\\cos\_\{\\mathrm\{sym\}\}\|is0\.3020\.302with27%27\\%of pairs above0\.4560\.456, so the control works at the aggregate and we do not claim individual pairs are uninformative\. Pairing a bracket with a*different*pair’s response destroys the agreement: under a fixed derangement the relative norm error rises from0\.3%0\.3\\%to53%53\\%, a factor of167167across all eighteen cells; the prediction is pair\-specific, not a generic property of any antisymmetric object of the right size\. We no longer count the site\-swap check among these:cosswapped=−cos\\cos\_\{\\mathrm\{swapped\}\}=\-\\cosis an algebraic identity ofℒ^\\widehat\{\\mathcal\{L\}\}’s definition rather than a property of the network \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\)\. The residual left by this zero\-parameter route,5\.25\.2–7\.6%7\.6\\%at the primary configuration, is a mean of per\-pair ratios and is not comparable to the2\.92\.9–9\.9%9\.9\\%the fitted correction leaves, which is pooled over pairs\. Matched pooled against pooled, the two routes differ by1313–14%14\\%in the families where both exist \(Appendix[C](https://arxiv.org/html/2608.25315#A3)\), with the zero\-parameter figure the larger, as its noise accumulation and the fitted route’s in\-sample bias both predict\.

The one weaker cell is DeepSeek at\(0\.5,0\.75\)\(0\.5,0\.75\), where the cosine falls to0\.98690\.9869and the residual rises to15\.7%15\.7\\%— the largest in the table\. DeepSeek is also the family whose trait\-clustered interval fails to exclude the generic null\. We note the co\-location without claiming the two are the same effect\.

Table 2:Zero\-parameter first\-order prediction against the measured bracket, per family and injection configuration\.cos\\cosis the agreement betweenℒ^\\widehat\{\\mathcal\{L\}\}\(Section[3](https://arxiv.org/html/2608.25315#S3), no fitted parameters\) and the measured bracket;*resid*is the fraction of the bracket’s norm it leaves, which by that section’s expansion is the antisymmetrized mixed second derivative*plus*an unboundedO⁡\(α3\)O\(\\alpha^\{3\}\)remainder\. We do not measure the cubic share on models, so*resid*is a finite\-scale mixture, not an interaction estimate or a one\-sided bound without explicit remainder control\.cossym\\cos\_\{\\mathrm\{sym\}\}is the symmetric\-combination control, which should carry no bracket signal;cosswap\\cos\_\{\\mathrm\{swap\}\}is the site\-swap control, which should equal−cos\-\\cosexactly\. Both hold in every cell\.

## Appendix DNulls, robustness, and residual geometry

### Two statistics, one exact identity\.

The identityBri​j=\(Di​j−Dj​i\)\+ℒ^i​j\\mathrm\{Br\}\_\{ij\}=\(D\_\{ij\}\-D\_\{ji\}\)\+\\widehat\{\\mathcal\{L\}\}\_\{ij\}is a tautology: the single\-intervention terms cancel by definition\. Its empirical content is the separation of objects: the raw bracket retains the first\-order site\-asymmetry term, whereas the second difference removes it\. The full scale sweep, constant\-map check, and comparison with published second\-difference statistics are released in the artifact; they are not additional model evidence\. The corrected residual is precisely the antisymmetrized second difference, so readers seeking an interaction should measureDDdirectly rather than interpret the raw bracket\.

### The three nulls\.

Each is generated on the family’s own empirical directions\.*Linear slack*:Br=W​δi​j\\mathrm\{Br\}=W\\delta\_\{ij\}with a constant Jacobian\.*Trait\-varying Jacobian*:Br=\(W0\+Δ​Wi\+Δ​Wj\)​δi​j\\mathrm\{Br\}=\(W\_\{0\}\+\\Delta W\_\{i\}\+\\Delta W\_\{j\}\)\\delta\_\{ij\}, still purely first order and still exactly linear in the injection coefficient, but with the Jacobian depending on which traits are injected — this is the alternative a first\-order account of the residual must invoke\.*Generic interaction*: a rank\-8 antisymmetric form on top of a linear term\. All three generate brackets from the*true*directions while the fit sees independently*estimated*ones, so first\-order slack is present wherever it can be\.

The two interaction\-bearing arms carry a free parameter set to the family’s observed residual fraction\. The linear\-slack arm does not and cannot: under a constant Jacobian the residual is annihilated*exactly*at any estimation error \(Appendix[D](https://arxiv.org/html/2608.25315#A4)\), so that arm reports only its own measurement noise,0\.0070\.007–0\.0230\.023of the bracket against an observed0\.0290\.029–0\.0990\.099\. It is a floor, not a matched competitor, and we do not present it as one\.

Table 3:Unsigned shared\-argument ratio against three nulls, 40 draws per family\. The generic\-interaction and trait\-varying\-Jacobian arms are calibrated so their residual is the same fraction of the bracket as the observed residual; the*linear slack*arm is not calibrated, and cannot be — its residual is0\.0070\.007–0\.0230\.023of the bracket against an observed0\.0290\.029–0\.0990\.099, because a constant\-Jacobian bracket is annihilated exactly \(Appendix[D](https://arxiv.org/html/2608.25315#A4)\)\. It is a floor, not a matched competitor\. Brackets are generated from the true directions and fitted on independently estimated ones\. The observation is bracketed on both sides: above a generic interaction in 5 of 6 families \(its own9595th percentile, in parentheses\), and far below a first\-order model whose Jacobian varies with the injected traits, in 6 of 6\.
### Null construction\.

Each null generates brackets from the true directionsvvand fits on independently perturbed estimatesv^\\hat\{v\}; additive noise is scaled by total norm, not per component\. The trait\-varying and generic arms each have one free parameter — the Jacobian\-variation scaleΔ​W\\Delta W, and the interaction scale — set so that the null’s residual is the same fraction of its bracket as the family’s observed residual \(0\.029–0\.099\)\. Without that calibration the arms sit at different signal\-to\-noise and the comparison is not meaningful\.

### The constant\-Jacobian arm is an exact floor, not a competitor\.

It has no free parameter\. Write the stack of pair differences asP​𝖣P\\mathsf\{D\}\. For invertible𝖣\\mathsf\{D\}and𝖣^\\widehat\{\\mathsf\{D\}\},col⁡\(P​𝖣\)=col⁡\(P\)=col⁡\(P​𝖣^\)\\mathrm\{col\}\(P\\mathsf\{D\}\)=\\mathrm\{col\}\(P\)=\\mathrm\{col\}\(P\\widehat\{\\mathsf\{D\}\}\), so the least\-squares projector is unchanged by direction perturbation and a constant\-Jacobian bracket is annihilated exactly\. The arm therefore reports only additive measurement noise, not a matched null\. The invariance is entailed by the shared column space, not a stuck knob: atσ≤2\\sigma\\leq 2the maximum principal angle is1\.9×10−61\.9\\times 10^\{\-6\}degrees and the relative residual is at most1\.9×10−151\.9\\times 10^\{\-15\}\. The trait\-varying arm is the discriminating first\-order account \(6\.0586\.058versus the observed2\.1322\.132\);400400\-draw reruns move each9595th percentile by at most0\.0480\.048and change no verdict\.

### Prompt\-split replication\.

Cross\-seed replication does not break prompt\-level fluctuations because the prompts are shared\. Table[4](https://arxiv.org/html/2608.25315#A4.T4)separates the cost of halving the sample \(within\-half mean1\.6021\.602versus full\-sample2\.1322\.132\) from the cost of disjointness \(a further0\.2110\.211\)\. The cross\-half mean is1\.3911\.391: Llama, OLMo and Qwen clear the recalibrated null1\.3251\.325, while DeepSeek, Gemma and Mistral remain within0\.020\.02of linear slack; Mistral falls2\.578→1\.0872\.578\\to 1\.087already from halving prompts\.

Table 4:Prompt\-split replication on disjoint prompt halves, all six families, layer pair\(0\.25,0\.50\)\(0\.25,0\.50\), 48 prompts split into two disjoint halves of 24\.*Cross\-half*is the confound\-free statistic: brackets computed on half A are compared against brackets computed on half B, so no prompt\-level fluctuation is shared between the two sides\. Three of six clear the generic antisymmetric null of1\.3241\.324\. The three that do not sit within0\.020\.02of the1\.001\.00of pure linear slack\.*The null is recomputed under the design it gates\.*The1\.3241\.324bar is calibrated on the full4848\-prompt sample, while the cross\-half statistic it gates is computed on2424\-prompt halves\. Because each null arm’s knob is chosen so its residual fraction matches the observed one, and halving the prompts raises that observed fraction, the bar is not automatically transportable across the two designs\. We therefore recomputed it: the halved\-sample observed residual fractions \(0\.0350\.035to0\.3370\.337, against0\.0290\.029to0\.0990\.099at full sample\) give a recalibrated bar of1\.3251\.325, and the same three families clear it\. The bar is close to unmoved because the generic\-interaction arm’s shared\-argument ratio is insensitive to the noise level over this range, not because the recomputation was skipped\. Llama1\.5991\.599, OLMo2\.0592\.059and Qwen1\.6421\.642therefore clear a bar calibrated on their own design\.scripts/halved\_null\.py; recordexperiments/controls/halved\_null\.json\. The script reproduces the committed full\-sample null bit\-for\-bit \(1\.32393129254072961\.3239312925407296, to2×10−162\\times 10^\{\-16\}\) before computing anything new, so it is answering with the same instrument that produced the published bar\. This does not touch the paper’s central negative claim, which never uses this null\.
### A disjoint trait inventory\.

Because the six families share one 16\-trait inventory, we repeated the three configuration\-robust families on 14 disjoint traits \(91 pairs, two extraction seeds\)\. Ratios were Llama2\.5602\.560vs2\.7182\.718, Mistral2\.5812\.581vs2\.5782\.578, and OLMo3\.2983\.298vs2\.7982\.798\(mean retention1\.04±0\.071\.04\\pm 0\.07\)\. Size\-matched subsampling slightly deflates rather than inflates the result \(design effect0\.9690\.969\); the disjoint\-inventory null is1\.2841\.284and all three families clear it\. This is evidence for those three families, not all six\.

### Scope: the operator is layer\-specific\.

Withlp0=\(\.25,\.50\)\\mathrm\{lp\}\_\{0\}=\(\.25,\.50\),lp1=\(\.50,\.75\)\\mathrm\{lp\}\_\{1\}=\(\.50,\.75\), andlp2=\(\.25,\.75\)\\mathrm\{lp\}\_\{2\}=\(\.25,\.75\), the effect collapses to≈1\\approx 1when the compared pairs do not share a site, while the shared deeper\-site comparison retains2\.7332\.733,2\.1012\.101, and1\.9461\.946for Llama, Mistral, and OLMo\. Within\-pair analysis finds the effect in only three families at all configurations; Mistral then fails the prompt split, leaving Llama and OLMo as the families surviving every control\. The opposite\-slot ratio is2\.1402\.140versus2\.1322\.132, so the elevation follows argument overlap, not slot choice\. This layer\-specific scope is consistent with independent evidence that single\-layer steering does not sustain effects across layers\([Saadatinia et al\., 2026](https://arxiv.org/html/2608.25315#bib.bib23)\)\.

### Adjacent literature and reusable checklist\.

Continuous causal and low\-rank or multi\-behaviour diagnostics\([Mahadevan, 2026](https://arxiv.org/html/2608.25315#bib.bib28);[Sharma et al\., 2026](https://arxiv.org/html/2608.25315#bib.bib6);[van der Weij et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib22)\), contrastive steering\([Rimsky et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib8);[Turner et al\., 2023](https://arxiv.org/html/2608.25315#bib.bib7);[Tan et al\., 2024](https://arxiv.org/html/2608.25315#bib.bib9)\), and classical bilinear or tabular interaction methods\([Tenenbaum and Freeman, 2000](https://arxiv.org/html/2608.25315#bib.bib10);[Rendle, 2010](https://arxiv.org/html/2608.25315#bib.bib11);[Smolensky, 1990](https://arxiv.org/html/2608.25315#bib.bib12);[Memisevic and Hinton, 2010](https://arxiv.org/html/2608.25315#bib.bib13);[Friedman and Popescu, 2008](https://arxiv.org/html/2608.25315#bib.bib14);[Tsang et al\., 2018](https://arxiv.org/html/2608.25315#bib.bib15);[Janizek et al\., 2021](https://arxiv.org/html/2608.25315#bib.bib16);[Sundararajan et al\., 2020](https://arxiv.org/html/2608.25315#bib.bib17);[Tsai et al\., 2023](https://arxiv.org/html/2608.25315#bib.bib18)\)address adjacent additive\-versus\-coupling questions but do not instantiate this distinct\-site activation\-space order swap\. For reuse: \(i\) sweep injection scale and state truncation order; \(ii\) subtract both the linear and self\-curvature terms; \(iii\) preserve argument slots and use true directions with independently fitted estimates; \(iv\) calibrate each null to the observed residual; \(v\) test site separation with the readout fixed; and \(vi\) replicate on disjoint prompt halves, reporting within\-half and cross\-half values separately\.

## References

- Adilaet al\.\(2026\)D\. Adila, J\. Cooper, A\. Yun, A\. Trost, and F\. SalaWeight updates as activation shifts: a principled framework for steering\.arXiv preprint arXiv:2603\.00425\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Doomset al\.\(2026\)T\. Dooms, W\. Gauderis, G\. Wiggins, and J\. OramasBilinear autoencoders find interpretable manifolds\.arXiv preprint arXiv:2605\.08891\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Friedman and Popescu \(2008\)J\. H\. Friedman and B\. E\. PopescuPredictive learning via rule ensembles\.The Annals of Applied Statistics2\(3\),pp\. 916–954\.External Links:[Document](https://dx.doi.org/10.1214/07-AOAS148)Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Heapet al\.\(2026\)T\. Heap, T\. Lawson, L\. Farnik, and L\. AitchisonAutomated interpretability metrics do not distinguish trained and random transformers\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2501\.17727Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Ilharcoet al\.\(2023\)G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, S\. Gururangan, L\. Schmidt, H\. Hajishirzi, and A\. FarhadiEditing models with task arithmetic\.arXiv preprint arXiv:2212\.04089\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Janizeket al\.\(2021\)J\. D\. Janizek, P\. Sturmfels, and S\. LeeExplaining explanations: axiomatic feature interactions for deep networks\.Journal of Machine Learning Research22\(104\),pp\. 1–54\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Khemais \(2026\)A\. KhemaisCross\-layer interaction under weight\-space ablation: a closed\-form attention jacobian bound and a test on a real pretrained model\.arXiv preprint arXiv:2608\.03629\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Kolbeinssonet al\.\(2025\)A\. Kolbeinsson, K\. O’Brien, T\. Huang, S\. Gao, S\. Liu, J\. R\. Schwarz, A\. Vaidya, F\. Mahmood, M\. Žitnik, T\. Chen, and T\. HartvigsenComposable interventions for language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Mahadevan \(2026\)S\. MahadevanInfinitesimal causality\.arXiv preprint arXiv:2606\.24621\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Memisevic and Hinton \(2010\)R\. Memisevic and G\. E\. HintonLearning to represent spatial transformations with factored higher\-order Boltzmann machines\.Neural Computation22\(6\),pp\. 1473–1492\.External Links:[Document](https://dx.doi.org/10.1162/neco.2010.01-09-953)Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Mudarisovet al\.\(2026\)T\. Mudarisov, M\. Burtsev, and R\. StateFeed\-forward steering in transformer residual dynamics\.arXiv preprint arXiv:2608\.02071\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Ortiz\-Jimenezet al\.\(2023\)G\. Ortiz\-Jimenez, A\. Favero, and P\. FrossardTask arithmetic in the tangent space: improved editing of pre\-trained models\.InAdvances in Neural Information Processing Systems,Vol\.36\.Note:arXiv:2305\.12827External Links:[Document](https://dx.doi.org/10.52202/075280-2913)Cited by:[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Piontkovskaia and Nikolenko \(2026\)I\. Piontkovskaia and S\. NikolenkoFirst\-order predictable but pairwise fragile: local task adaptation in trained transformers\.arXiv preprint arXiv:2607\.16821\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.SS0.SSS0.Px1.p3.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Rendle \(2010\)S\. RendleFactorization machines\.In2010 IEEE International Conference on Data Mining,pp\. 995–1000\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Richards \(2026\)L\. RichardsDo active SAE feature planes carry more holonomy? a preregistered reversal in gemma\.arXiv preprint arXiv:2607\.20522\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15504–15522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828)Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1),[§3](https://arxiv.org/html/2608.25315#S3.SS0.SSS0.Px6.p1.1)\.
- Saadatiniaet al\.\(2026\)M\. Saadatinia, P\. Razmara, A\. Aryashad, A\. Abbasi, and S\. AziziCircuitSteer: geometrically aligned multi\-layer steering via sparse autoencoder circuits\.arXiv preprint arXiv:2608\.05732\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px7.p1.1)\.
- Schessl \(2026\)F\. M\. SchesslForgetting is not a fix: path dependence in sequential engram editing\.arXiv preprint arXiv:2607\.24805\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1)\.
- Sevetlidis and Pavlidis \(2026\)V\. Sevetlidis and G\. PavlidisGauge\-invariant representation holonomy\.arXiv preprint arXiv:2601\.21653\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Sharmaet al\.\(2026\)A\. Sharma, C\. Schroeder de Witt, P\. Torr, A\. Calinescu, and J\. YuA low\-rank subspace analysis of LLM interventions\.arXiv preprint arXiv:2606\.14388\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Skifstadet al\.\(2026\)J\. Skifstad, X\. A\. Yang, and G\. ChouLocal linearity of LLMs enables activation steering via model\-based linear optimal control\.arXiv preprint arXiv:2604\.19018\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Smolensky \(1990\)P\. SmolenskyTensor product variable binding and the representation of symbolic structures in connectionist systems\.Artificial Intelligence46\(1–2\),pp\. 159–216\.External Links:[Document](https://dx.doi.org/10.1016/0004-3702%2890%2990007-M)Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Sundararajanet al\.\(2020\)M\. Sundararajan, K\. Dhamdhere, and A\. AgarwalThe Shapley Taylor interaction index\.InProceedings of the 37th International Conference on Machine Learning,Vol\.119,pp\. 9259–9268\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Sweeney \(2026\)J\. SweeneyThe geometry of sequential learning: lie\-bracket prediction of transfer order\.arXiv preprint arXiv:2606\.24993\.Cited by:[§1](https://arxiv.org/html/2608.25315#S1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1)\.
- Tanet al\.\(2024\)D\. Tan, D\. Chanin, A\. Lynch, B\. Paige, D\. Kanoulas, A\. Garriga\-Alonso, and R\. KirkAnalysing the generalisation and reliability of steering vectors\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1),[§3](https://arxiv.org/html/2608.25315#S3.SS0.SSS0.Px6.p1.1)\.
- Tenenbaum and Freeman \(2000\)J\. B\. Tenenbaum and W\. T\. FreemanSeparating style and content with bilinear models\.Neural Computation12\(6\),pp\. 1247–1283\.External Links:[Document](https://dx.doi.org/10.1162/089976600300015349)Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Tsaiet al\.\(2023\)C\. Tsai, C\. Yeh, and P\. RavikumarFaith\-shap: the faithful Shapley interaction index\.Journal of Machine Learning Research24\(94\),pp\. 1–42\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Tsanget al\.\(2018\)M\. Tsang, D\. Cheng, and Y\. LiuDetecting statistical interactions from neural network weights\.InInternational Conference on Learning Representations,Note:arXiv:1705\.04977Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidSteering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1),[§3](https://arxiv.org/html/2608.25315#S3.SS0.SSS0.Px6.p1.1)\.
- Vaidyanathanet al\.\(2026\)S\. Vaidyanathan, D\. Arbour, A\. Mueller, S\. Niekum, and D\. JensenThe curse of multiple mediators: hidden interaction effects in activation patching\.arXiv preprint arXiv:2606\.27510\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- van der Weijet al\.\(2024\)T\. van der Weij, M\. Poesio, and N\. SchootsExtending activation steering to broad skills and multiple behaviours\.arXiv preprint arXiv:2403\.05767\.Cited by:[Appendix D](https://arxiv.org/html/2608.25315#A4.SS0.SSS0.Px8.p1.1),[§1](https://arxiv.org/html/2608.25315#S1.p1.1)\.
- VanderWeele \(2014\)T\. J\. VanderWeeleA unification of mediation and interaction: a 4\-way decomposition\.Epidemiology25\(5\),pp\. 749–761\.External Links:[Document](https://dx.doi.org/10.1097/EDE.0000000000000121)Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Wuet al\.\(2026\)Y\. Wu, S\. Zhao, and J\. ChenWhen is a steerable concept representation real? measurement confounds in a cross\-family audit of neuroscience parallels in LLMs\.arXiv preprint arXiv:2608\.08159\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.
- Zhang and Wang \(2026\)L\. Zhang and J\. WangWhen attribution patching lies: diagnosis and a second\-order correction\.arXiv preprint arXiv:2606\.09899\.Cited by:[§2](https://arxiv.org/html/2608.25315#S2.p1.1)\.

Similar Articles

A Geometric Account of Activation Steering through Angle-Norm Decomposition

arXiv cs.AI

This paper analyzes linear activation steering in language models by decomposing interventions into angular and radial components. It finds that concepts are primarily encoded in angular structure, but norm adjustments are crucial for stability, supporting spherical steering methods while showing that additive coefficients conflate geometry.

Relation Geometry in Semantic Space of Language Models

arXiv cs.CL

This paper explores how semantic relations are encoded in the geometry of language model semantic spaces, finding that asymmetric relations occupy distinct regions and that lexical information matters more for causal models while contextual information matters more for masked and diffusion models.