LLM组合任务中的因果与可解释结构
摘要
研究借助功能ANOVA(方差分析)分解,揭示了LLM在组合任务中跨层的因果与几何表示机制:中间层采用两token的联合关系表示,后期层则转向三token联合表示,而限制模型仅关注因果相关的表示反而提升了下一token的预测准确率。
arXiv:2609.35970v1 Announce Type: new
Abstract: Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
查看缓存全文
缓存时间: 2026/09/30 09:50
# Causal and Interpretable Structures in LLM Compositional Tasks Source: [https://arxiv.org/html/2609.35970](https://arxiv.org/html/2609.35970) Toni J\.B\. LiuAffiliation:Cornell University, USAJiajun BaoAffiliation:Cornell University, USARaphaël SarfatiChristopher J\. EarlsAffiliation:Goodfire AI, USACorrespondence:[g\.arora@cornell\.edu](mailto:[email protected]) ###### Abstract Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them\. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept \(months, hours, weekdays, and musical notes\) to correctly predict the next token\. Across model families \(Llama, Qwen, Gemma, and Mistral\) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task\. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next\-token prediction\. Crucially, when taken together, these geometric and causal investigations reveal the representation\-level mechanism that progressively organizes and composes the relational information to form the answer\. More surprisingly, restricting the models to such causally relevant joint representations improves next\-token prediction accuracy\. ## 1Introduction Many tasks require a large language model \(LLM\) to use relationships between multiple input tokens\. Consider the prompt, “*Hi John, the conference is scheduled between January 13 and March 13\. This is the same time as between May 13 and*”\. The correct next\-token prediction depends crucially on three tokens:*January*,*March*, and*May*\. LetAA,BB, andCCdenote these three tokens, respectively, and defineγ=B−A\(mod12\)\\gamma=B\-A\\;\(\\mathrm\{mod\\;\}12\)\. The correct answer is thenD=C\+γ\(mod12\)D=C\+\\gamma\\;\(\\mathrm\{mod\\;\}12\)\. The individual representations ofAA,BB, andCClie on the low\-dimensional manifold associated with calendar months\([Engels et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib1)\)\. We ask whether computationally relevant quantities arising jointly from these tokens, such asγ\\gammaorDD, also have geometric structure, and whether that structure is used in the next\-token prediction\. We varyA,B,CA,B,Cindependently over calendar months while keeping everything else fixed\. Across this ensemble of prompts, the last\-token activation varies withAA,BB, andCCindividually and with the particular combinations in which they co\-occur\. These sources of variation are superposed in the full activation ensemble\. We use the functional analysis of variance \(ANOVA\) decomposition to separate the resulting activation ensemble into terms associated with individual variables, pairwise interactions, and their three\-way interaction\([Hoeffding, 1948](https://arxiv.org/html/2609.35970#bib.bib18)\)\. We treat the interaction terms as geometric objects and study their organization and causal role across the layers of an LLM\. We study this question across several cyclic concepts: months, hours, weekdays, and musical notes\. The cyclic structure gives us known relational quantities\. In particular,γ=B−A\\gamma=B\-Ais a pairwise relation, while the answerD=C\+γD=C\+\\gammadepends jointly on all three inputs\. Importantly,γ\\gammais never supplied explicitly and must instead be inferred fromAAandBB\. This allows us to ask how the joint dependence associated withγ\\gammaandDDis represented and used across network depth\. Our contributions are:\(a\) Geometric structure\.We find that the pairwise interaction betweenAAandBBis organized byγ\\gammain the middle layers\. Other pairwise interactions are also organized by their corresponding pairwise differences\. The three\-way interaction betweenAA,BB, andCCbecomes organized byDDat a later depth\.\(b\) Causality\.We show that the representation associated withA,BA,Bis causally relevant in the middle layers, while the representation associated withA,B,CA,B,Cbecomes causally relevant at later layers\.\(c\) Improved performance\.Retaining only the causally relevant three\-way interaction improves the performance on the task\.\(d\) Relationship transfer\.We show that the relationship encoded by interaction terms transfers across cyclic domains, such as from hours to months\. ### 1\.1Related work #### Geometry of concepts\. Prior work has identified low\-dimensional geometric structure associated with individual concepts in latent activations\([Engels et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib1);[Modell et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib4);[Park et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib22);[Karkada et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib11);[Hu et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib16);[Prieto et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib21)\)\. Other work has studied how representational geometry evolves during computation and how interventions along that geometry affect model behavior\([Gurnee et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib10);[Wurgaft et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib23);[Sarfati et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib24)\)\. We instead study the geometry and causality of joint dependence between multiple tokens\. Figure 1:Two\-dimensional visualization using the first two principal components for the activation ensembleΦr\\Phi^\{r\}and its decomposed interaction ensemblesHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}for the prompt template in Section[2\.1](https://arxiv.org/html/2609.35970#S2.SS1)using Llama\-3\.1\-8B\. Here,γ=B−Amod12\\gamma=B\-A\\mod 12and the expected answer isD=C\+γmod12D=C\+\\gamma\\mod 12\. Before fitting the principal components, we exclude prompts satisfyingA=BA=B,B=CB=C, orC=AC=A\(see Appendix[B\.3](https://arxiv.org/html/2609.35970#A2.SS3)\)\. The bottom left of each panel shows the coloring according to eitherγ\\gammaorDD\.HABrH\_\{AB\}^\{r\}gets organized byγ\\gammain early layers\. Until layer 17,HABCrH\_\{ABC\}^\{r\}is unstructured but becomes organized byDDat layer 18\. See Section[3](https://arxiv.org/html/2609.35970#S3)for a quantitative analysis\. #### Relational representations and arithmetic computation\. Prior work has shown that latent activations contain structured representations of relations between inputs, including approximately linear transformations, as well as distributed representations that bind entities or variables to relational roles\([Hernandez et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib12);[Merullo et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib13);[Feng and Steinhardt, 2024](https://arxiv.org/html/2609.35970#bib.bib25);[Davies et al\., 2023](https://arxiv.org/html/2609.35970#bib.bib26);[Dai et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib15);[Todd et al\., 2026](https://arxiv.org/html/2609.35970#bib.bib27);[Wang et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib14)\)\. Related work on arithmetic has identified structured numerical representations and computational mechanisms, including Fourier and periodic structure in transformers trained on modular arithmetic and in pretrained language models\([Nanda et al\., 2023](https://arxiv.org/html/2609.35970#bib.bib28);[Furuta et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib29);[Zhou et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib30);[Kantamneni and Tegmark, 2025](https://arxiv.org/html/2609.35970#bib.bib3);[Levy and Geva, 2025](https://arxiv.org/html/2609.35970#bib.bib31)\)\. A closely related work is that of[Feucht et al\. \(2026\)](https://arxiv.org/html/2609.35970#bib.bib17), who study tasks involving an explicit offsetγ\\gammaand a conceptCCfrom a cyclic domain\. They identify a universal base\-10 addition mechanism in Llama\-3\.1\-8B by characterizing MLP neurons associated with the task’s Fourier structure\. In contrast, we work at the representation level\. The relationγ\\gammain our setting is not supplied explicitly, but arises from the joint dependence of two independently varied inputs\. We isolate the component of the residual\-stream representation associated with this joint dependence and study its geometry and causal role across layers\. #### Causal interventions\. The presence of structured or decodable information in an activation does not by itself establish that the model uses that information to perform the task\. Causal interventions and activation patching have been used in several works to test the functional role of such representations\([Vig et al\., 2020](https://arxiv.org/html/2609.35970#bib.bib32);[Geiger et al\., 2021](https://arxiv.org/html/2609.35970#bib.bib33);[Geiger et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib34);[Zhang and Nanda, 2024](https://arxiv.org/html/2609.35970#bib.bib35);[Arora et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib36)\)\. Our experiments similarly intervene directly on joint representations of multiple input tokens\. ## 2Setup In the main text, all results are reported using base Llama\-3\.1\-8B\([Llama Team, 2024b](https://arxiv.org/html/2609.35970#bib.bib9)\), which contains 32 layers indexed byℓ∈\{0,…,31\}\\ell\\in\\\{0,\\ldots,31\\\}\. In the appendices, we reproduce all our results on base models: Llama\-3\.2\-3B\([Llama Team, 2024a](https://arxiv.org/html/2609.35970#bib.bib40)\), Qwen\-2\.5\-7B\([Qwen Team, 2024](https://arxiv.org/html/2609.35970#bib.bib5)\), Qwen\-3\-8B\([Qwen Team, 2025](https://arxiv.org/html/2609.35970#bib.bib6)\), Gemma\-2\-9B\([Gemma Team, 2024](https://arxiv.org/html/2609.35970#bib.bib7)\), Gemma\-3\-12B\([Gemma Team, 2025](https://arxiv.org/html/2609.35970#bib.bib8)\), Mistral\-small\-24B\-Base\-2501\([Mistral AI Team, 2025](https://arxiv.org/html/2609.35970#bib.bib20)\)\. In each model, we analyze residual\-stream activations after every transformer block\. ### 2\.1Construction of prompts Several other works have used controlled prompt variations to reveal manifolds associated with individual concepts\([Engels et al\., 2025](https://arxiv.org/html/2609.35970#bib.bib1);[Kantamneni and Tegmark, 2025](https://arxiv.org/html/2609.35970#bib.bib3)\)\. We propose a similar experimental design, but we aim to elicit a relationship between concepts\. We consider several cyclic concepts—months, weekdays, hours, musical notes—and let𝒳\\mathcal\{X\}be the associated vocabulary\. In the main text, we discuss the months domain,𝒳:=\{January,…,December\}\\mathcal\{X\}:=\\\{\\text\{January\},\\ldots,\\text\{December\}\\\}\. Consider the variables\(A,B,C\)∈𝒳3\(A,B,C\)\\in\\mathcal\{X\}^\{3\}along with the prompt ``` Hi {name}, the {noun} is scheduled between {A} {dates} and {B} {dates}. This is the same time as between {C} {dates} and ``` Here,`name`is chosen from the most common names of the past century\([Social Security Administration, n\.d\.](https://arxiv.org/html/2609.35970#bib.bib2)\),`noun`varies over synonyms of`conference`, while`dates`∈\{1,…,28\}\\in\\\{1,\\ldots,28\\\}\(see Appendix[A\.1](https://arxiv.org/html/2609.35970#A1.SS1)for a list\)\. We call each combination of\(𝚗𝚊𝚖𝚎,𝚗𝚘𝚞𝚗,𝚍𝚊𝚝𝚎𝚜\)\(\\verb\|name\|,\\verb\|noun\|,\\verb\|dates\|\)areplicateand label it byrr\. These replicates define prompt variations that leave the underlying task unchanged\. We further require that all replicate variables tokenize to the same length so that positional offsets do not introduce a confound\. We randomly subsample 10 such replicates\. Unless stated otherwise, we average over the replicates when we report any quantity\. We design prompts associated with each cyclic concept such that the answer to the task is the next token\. We measure the performance by computing the top\-11and top\-33accuracy, where top\-kkaccuracy is the fraction of prompts for which the correct answer features in thekklargest logits\. The top\-11accuracy of Llama\-3\.1\-8B is62\.2%±5\.1%62\.2\\%\\pm 5\.1\\%, and the top\-33accuracy is85\.4%±3\.3%85\.4\\%\\pm 3\.3\\%, where the variation is reported as the standard deviation across replicatesrr111The accuracies are computed over all prompts with distinct variables\(A,B,C\)\(A,B,C\)\. See Appendix[B\.3](https://arxiv.org/html/2609.35970#A2.SS3)\.\. Because top\-33accuracy varies less with larger offsets ofγ=B−A≳6\\gamma=B\-A\\gtrsim 6than top\-11accuracy \(Appendix[A](https://arxiv.org/html/2609.35970#A1)\), we use top\-33accuracy as the primary performance metric in the main text\. The analogous results using top\-11accuracy are reported in the appendices\. ### 2\.2Ensemble decomposition Letϕℓr\(A,B,C\)∈ℝdmodel\\bm\{\\phi\}^\{r\}\_\{\\ell\}\(A,B,C\)\\in\\mathbb\{R\}^\{d\_\{\\text\{model\}\}\}denote the last\-token activation at layerℓ\\ellfor replicaterr\. We define the corresponding activation ensemble as Φℓr:=\(ϕℓr\(A,B,C\)\)\(A,B,C\)∈𝒳3\.\\Phi^\{r\}\_\{\\ell\}:=\\left\(\\bm\{\\phi\}^\{r\}\_\{\\ell\}\(A,B,C\)\\right\)\_\{\(A,B,C\)\\in\\mathcal\{X\}^\{3\}\}\.\(1\)Since we report all quantities as a function of layer, we drop the labelℓ\\ellfor brevity\. We decompose each activation as ϕr\(A,B,C\)=𝝁r⏟mean\+𝒉Ar\+𝒉Br\+𝒉Cr⏟additive terms\+𝒉ABr\+𝒉BCr\+𝒉CAr\+𝒉ABCr⏟interaction terms,\\bm\{\\phi\}^\{r\}\(A,B,C\)=\\underbrace\{\\bm\{\\mu\}^\{r\}\}\_\{\\text\{mean\}\}\+\\underbrace\{\\bm\{h\}\_\{A\}^\{r\}\+\\bm\{h\}\_\{B\}^\{r\}\+\\bm\{h\}\_\{C\}^\{r\}\}\_\{\\text\{additive terms\}\}\+\\underbrace\{\\bm\{h\}\_\{AB\}^\{r\}\+\\bm\{h\}\_\{BC\}^\{r\}\+\\bm\{h\}\_\{CA\}^\{r\}\+\\bm\{h\}\_\{ABC\}^\{r\}\}\_\{\\text\{interaction terms\}\},\(2\) where every term is defined as an average with respect to the ensembleΦr\\Phi^\{r\}: 𝝁r\\displaystyle\\bm\{\\mu\}^\{r\}:=𝔼ABC\(ϕr\),\\displaystyle:=\\mathbb\{E\}\_\{ABC\}\(\\bm\{\\phi\}^\{r\}\),\(3\)𝒉Ar\\displaystyle\\bm\{h\}\_\{A\}^\{r\}:=𝔼BC\(ϕr\)−𝝁r,\\displaystyle:=\\mathbb\{E\}\_\{BC\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\},𝒉ABr\\displaystyle\\bm\{h\}\_\{AB\}^\{r\}:=𝔼C\(ϕr\)−𝝁r−𝒉Ar−𝒉Br,\\displaystyle:=\\mathbb\{E\}\_\{C\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\},𝒉ABCr\\displaystyle\\bm\{h\}\_\{ABC\}^\{r\}:=ϕr−𝝁r−𝒉Ar−𝒉Br−𝒉Cr−𝒉ABr−𝒉BCr−𝒉CAr\.\\displaystyle:=\\bm\{\\phi\}^\{r\}\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\}\-\\bm\{h\}\_\{C\}^\{r\}\-\\bm\{h\}\_\{AB\}^\{r\}\-\\bm\{h\}\_\{BC\}^\{r\}\-\\bm\{h\}\_\{CA\}^\{r\}\. Analogously to Equation[1](https://arxiv.org/html/2609.35970#S2.E1), we define the ensemble associated with decomposed terms as HAr=\(𝒉Ar\)A∈𝒳,HABr=\(𝒉ABr\)\(A,B\)∈𝒳2,andHABCr=\(𝒉ABCr\)\(A,B,C\)∈𝒳3\.H\_\{A\}^\{r\}=\\left\(\\bm\{h\}\_\{A\}^\{r\}\\right\)\_\{A\\in\\mathcal\{X\}\},\\qquad H\_\{AB\}^\{r\}=\\left\(\\bm\{h\}\_\{AB\}^\{r\}\\right\)\_\{\(A,B\)\\in\\mathcal\{X\}^\{2\}\},\\quad\\text\{and\}\\quad H\_\{ABC\}^\{r\}=\\left\(\\bm\{h\}\_\{ABC\}^\{r\}\\right\)\_\{\(A,B,C\)\\in\\mathcal\{X\}^\{3\}\}\.The remaining terms are defined by cyclic permutation of the expressions above\. Intuitively,𝝁r\\bm\{\\mu\}^\{r\}is the center of the activation ensemble,HArH\_\{A\}^\{r\}captures the dependence onAAalone,HABrH\_\{AB\}^\{r\}captures the joint dependence onAAandBBthat remains after averaging overCCand removing their individual effects, andHABCrH\_\{ABC\}^\{r\}contains the residual three\-way dependence after accounting for all lower\-order terms\. We refer toHABrH\_\{AB\}^\{r\},HBCrH\_\{BC\}^\{r\},HCArH\_\{CA\}^\{r\}, andHABCrH\_\{ABC\}^\{r\}as the interaction ensembles\. We emphasize that Equation[2](https://arxiv.org/html/2609.35970#S2.E2)is an exact identity rather than an approximation\. It is known as the functional ANOVA decomposition\([Hoeffding, 1948](https://arxiv.org/html/2609.35970#bib.bib18);[Owen, 2013](https://arxiv.org/html/2609.35970#bib.bib19)\)\. The decomposed components are pairwise orthogonal under the ensemble\-averaged inner product \(Appendix[B](https://arxiv.org/html/2609.35970#A2)\)\. We do not apply the inferential machinery of ANOVA, since the associated questions of statistical significance are orthogonal to the goal of this paper\. Figure[1](https://arxiv.org/html/2609.35970#S1.F1)shows the first two principal components ofΦr\\Phi^\{r\},HABrH\_\{AB\}^\{r\}, andHABCrH\_\{ABC\}^\{r\}\. The remainder of the paper analyzes the interaction ensembles as geometric objects to study their structure and causal role across network depth\. ## 3Geometric structure of interaction terms The ANOVA decomposition described in Equation[2](https://arxiv.org/html/2609.35970#S2.E2)breaks the ensemble into its constituent parts\. Figure[1](https://arxiv.org/html/2609.35970#S1.F1)visually suggests thatHABrH\_\{AB\}^\{r\}can be parametrized byγ=B−A\\gamma=B\-Aand similarlyHABCrH\_\{ABC\}^\{r\}can be parametrized byD=C\+γD=C\+\\gamma\. In this section, we first rigorously quantify the dependence of interaction terms onγ\\gammaorDDby exploiting the cyclic nature ofA,B,CA,B,Cusing the discrete Fourier transform in Section[3\.1](https://arxiv.org/html/2609.35970#S3.SS1)\. Second, we analyze the norm of vectors in each interaction ensemble in Section[3\.2](https://arxiv.org/html/2609.35970#S3.SS2)to check for ensembles that are highly organized but have negligible norm overall\. ### 3\.1Fourier Transform We perform a discrete Fourier transform \(DFT\) of every interaction term for every replicaterr\. ConsiderHABrH\_\{AB\}^\{r\}as an example\. Its DFT is given by 𝒉^kAkBr=∑A,B𝒉rABe−2πi\(kAA\+kBB\)/12,\\widehat\{\\bm\{h\}\}\_\{k\_\{A\}k\_\{B\}\}^\{r\}=\\sum\_\{A,B\}\\bm\{h\}^\{r\}\_\{AB\}e^\{\-2\\pi i\(k\_\{A\}A\+k\_\{B\}B\)/12\},\(4\)where we identify the months withA,B∈\{0,…,11\}A,B\\in\\\{0,\\ldots,11\\\}andkA,kBk\_\{A\},k\_\{B\}are the associated Fourier modes\. Letγ=B−A\\gamma=B\-A\. If𝒉ABr\\bm\{h\}\_\{AB\}^\{r\}depends only onγ\\gamma, that is,𝒉ABr≡𝒇\(γ\)\\bm\{h\}\_\{AB\}^\{r\}\\equiv\\bm\{f\}\(\\gamma\), then its DFT 𝒉^kAkBr=∑γ𝒇\(γ\)e−2πikBγ/12∑Ae−2πi\(kA\+kB\)A/12\\widehat\{\\bm\{h\}\}\_\{k\_\{A\}k\_\{B\}\}^\{r\}=\\sum\_\{\\gamma\}\\bm\{f\}\(\\gamma\)e^\{\-2\\pi ik\_\{B\}\\gamma/12\}\\sum\_\{A\}e^\{\-2\\pi i\(k\_\{A\}\+k\_\{B\}\)A/12\}vanishes unlesskA\+kB=0mod12k\_\{A\}\+k\_\{B\}=0\\mod\{12\}\. Thus concentration on modes satisfyingkA=−kBk\_\{A\}=\-k\_\{B\}222Strictly, the condition iskA\+kB=0mod12k\_\{A\}\+k\_\{B\}=0\\;\\mathrm\{mod\}\\;\{12\}; all frequency equalities are understood modulo 12\.quantifies the extent to whichHABrH\_\{AB\}^\{r\}is organized according toγ\\gamma\. Therefore we definemode fractionas m\(HABr\):=∑kA=−kB‖𝒉^kAkBr‖22∑kA,kB‖𝒉^kAkBr‖22\.m\(H\_\{AB\}^\{r\}\):=\\frac\{\\sum\_\{k\_\{A\}=\-k\_\{B\}\}\\left\\lVert\\widehat\{\\bm\{h\}\}^\{r\}\_\{k\_\{A\}k\_\{B\}\}\\right\\rVert\_\{2\}^\{2\}\}\{\\sum\_\{k\_\{A\},k\_\{B\}\}\\left\\lVert\\widehat\{\\bm\{h\}\}^\{r\}\_\{k\_\{A\}k\_\{B\}\}\\right\\rVert\_\{2\}^\{2\}\}\.\(5\)The mode fractionm\(HABr\)∈\[0,1\]m\(H^\{r\}\_\{AB\}\)\\in\[0,1\]measures the degree of organization ofHABrH\_\{AB\}^\{r\}byγ\\gamma, withm=1m=1when the ensemble depends onγ\\gammaalone\. We define an analogous mode fraction forHBCrH\_\{BC\}^\{r\},HCArH\_\{CA\}^\{r\}, andHABCrH\_\{ABC\}^\{r\}to probe for their dependence onγ′=C−B\\gamma^\{\\prime\}=C\-B,γ′′=A−C\\gamma^\{\\prime\\prime\}=A\-C, andD=C\+B−AD=C\+B\-Arespectively\. For the latter, the modes under investigation are\(kA,kB,kC\)=\(−k,k,k\)\(k\_\{A\},k\_\{B\},k\_\{C\}\)=\(\-k,k,k\)\. Dependence onγ\\gammacannot occur in terms lower\-order thanHABrH\_\{AB\}^\{r\}, and dependence onDDcannot occur in terms lower\-order thanHABCrH\_\{ABC\}^\{r\}\. Thus, if the residual stream contains structure organized byγ\\gammaorDD, that structure must reside inHABrH\_\{AB\}^\{r\}orHABCrH\_\{ABC\}^\{r\}, respectively\. Consistently, we find large mode fractions for both interaction ensembles but at different depths\. Importantly, we also find that all second\-order interaction terms are strongly organized by their corresponding pairwise differences, and thatHABCrH\_\{ABC\}^\{r\}undergoes a sharp increase in organization byDDaround layers 17–18 \(Figure[2](https://arxiv.org/html/2609.35970#S3.F2)\(a\)\)\. At the same depth, the mode fractions ofHBCrH\_\{BC\}^\{r\}andHCArH\_\{CA\}^\{r\}begin to decrease, whilem\(HABr\)m\(H\_\{AB\}^\{r\}\)remains large throughout the depth of the transformer\. Results across models and domains appear in Appendix[C](https://arxiv.org/html/2609.35970#A3)\. Control experiments detailed in Appendix[E](https://arxiv.org/html/2609.35970#A5)suggest that the pairwise\-difference organization of the second\-order interaction ensembles can arise from the co\-occurrence of the corresponding tokens\. This is in agreement with[Karkada et al\. \(2026\)](https://arxiv.org/html/2609.35970#bib.bib11)\. In contrast,HABCrH\_\{ABC\}^\{r\}is not organized byDDwhen the task structure is removed from the prompt, even when all three tokensA,B,CA,B,Care present\. Figure 2:Mode fraction and energy fraction of the interaction ensembles across layers of Llama\-3\.1\-8B\. We zero out vectors for whichAA,BB, andCCare not pairwise distinct before computing both fractions \(Appendix[B\.3](https://arxiv.org/html/2609.35970#A2.SS3)\)\.\(a\)Mode fraction measures the organization ofHABrH\_\{AB\}^\{r\},HBCrH\_\{BC\}^\{r\}, andHCArH\_\{CA\}^\{r\}by their corresponding pairwise differences, and ofHABCrH\_\{ABC\}^\{r\}byD=C\+B−AD=C\+B\-A\. Under a random null model, the expected mode fraction for second\-order interactions is0\.090\.09and that for third\-order interaction is0\.0080\.008\.\(b\)Energy fraction measures the fraction of mean\-subtracted activation energy attributable to each interaction ensemble\. Throughout this work, colored bands show the central 80% interval across 10 replicates, and vertical gray bands highlight the layers discussed in the text\. ### 3\.2Energy fraction An ensemble can have a high mode fraction even if all of its vectors are close to zero\. Directly comparing the norms of interaction terms across layers is misleading becauseϕr\\bm\{\\phi\}^\{r\}grows substantially with depth, with most of this growth coming from the mean𝝁r\\bm\{\\mu\}^\{r\}\. Therefore we need a measure of relative magnitude for each ensemble across layers\. ForS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}, we define energy fractione\(HSr\)e\(H^\{r\}\_\{S\}\)as333See Appendix[D](https://arxiv.org/html/2609.35970#A4)for a relation between ANOVA decomposition and our definition of energy fraction\. e\(HSr\)=𝔼S\(‖𝒉Sr‖22\)𝔼ABC\(‖ϕr\(A,B,C\)−𝝁r‖22\)e\(H\_\{S\}^\{r\}\)=\\frac\{\\mathbb\{E\}\_\{S\}\\Big\(\\left\\lVert\\bm\{h\}^\{r\}\_\{S\}\\right\\rVert\_\{2\}^\{2\}\\Big\)\}\{\\mathbb\{E\}\_\{ABC\}\\Big\(\\left\\lVert\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{\\mu\}^\{r\}\\right\\rVert\_\{2\}^\{2\}\\Big\)\}\(6\)Figure[2](https://arxiv.org/html/2609.35970#S3.F2)\(b\) shows the energy fraction of each interaction ensemble\. The energy fraction ofHABrH\_\{AB\}^\{r\}peaks at layer 15 and starts decreasing around layers 17–18\. Around the same depth, the energy fraction ofHABCrH\_\{ABC\}^\{r\}increases dramatically\. This coincides with the sharp increase in organization ofHABCrH\_\{ABC\}^\{r\}byDD, as found in the previous Section[3\.1](https://arxiv.org/html/2609.35970#S3.SS1)\. Together, these observations hint at a transition from a pairwise representation of the input relationB−AB\-Ato a three\-way representation aligned with the answerC\+B−AC\+B\-A, which we test causally in the following section\. ## 4Causal use of interaction terms The previous section establishes that the interaction ensembles are highly organized by pairwise differences or byDD, depending on the interaction ensemble and the layer\. To test whether these ensembles are used in downstream computation, we intervene directly on the activation of the last token\([Zhang and Nanda, 2024](https://arxiv.org/html/2609.35970#bib.bib35)\)using its ANOVA decomposition\. In the first set of experiments in Section[4\.1](https://arxiv.org/html/2609.35970#S4.SS1), we intervene on the last token at a particular layer to remove an interaction term\. In Section[4\.2](https://arxiv.org/html/2609.35970#S4.SS2), we instead remove all terms except for the mean and an interaction term\. After each intervention, we let the model run without any further modification and score the performance\. Figure 3:Top row shows the result of ensemble ablation experiments, while the bottom row shows corresponding ensemble replacement experiments\. The horizontal dashed line shows the baseline performance\. Dotted line at 25% is the chance of getting the correct answer under a uniform prior\.\(a\)AblatingHABrH\_\{AB\}^\{r\}produces the largest decrease in top\-3 accuracy around layers 15–17, whereas ablatingHABCrH\_\{ABC\}^\{r\}sharply decreases accuracy after layer 17\.\(b\)AblatingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}has little effect on performance across model depth\.\(c\)Replacement experiments show a complementary transition where retainingHABrH\_\{AB\}^\{r\}preserves performance in the middle layers, while retainingHABCrH\_\{ABC\}^\{r\}preserves performance in later layers\.\(d\)RetainingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}sharply decreases performance from layer 15 onward\.### 4\.1Ensemble Ablations We defineensemble ablationof an interaction ensembleHSrH\_\{S\}^\{r\}by subtracting the corresponding interaction vector𝒉Sr\\bm\{h\}\_\{S\}^\{r\}from the activation of the last token: ϕ~r\(A,B,C\)=ϕr\(A,B,C\)−𝒉Sr,\\widetilde\{\\bm\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{h\}^\{r\}\_\{S\},\(7\)for allA,B,C∈𝒳A,B,C\\in\\mathcal\{X\}andS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}\. Figure[3](https://arxiv.org/html/2609.35970#S4.F3)shows the results of ensemble ablation experiments for every interaction term\. We see that ablatingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}barely changes the top\-33accuracy relative to the baseline \(Figure[3](https://arxiv.org/html/2609.35970#S4.F3)\(b\)\)\. In contrast, ablatingHABrH\_\{AB\}^\{r\}orHABCrH\_\{ABC\}^\{r\}causes a drastic change in the model’s performance, but at layers 15 and 18 respectively \(Figure[3](https://arxiv.org/html/2609.35970#S4.F3)\(a\)\)\. ### 4\.2Ensemble Replacement In a complementary set of experiments, we performensemble replacementby retaining a single interaction ensemble together with the mean\. Concretely, we replace the activation of the last token with ϕ~r\(A,B,C\)=𝝁r\+𝒉Sr,\\widetilde\{\\bm\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\mu\}^\{r\}\+\\bm\{h\}^\{r\}\_\{S\},\(8\)for allA,B,C∈𝒳A,B,C\\in\\mathcal\{X\}andS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}\. Figure[3](https://arxiv.org/html/2609.35970#S4.F3)\(d\) shows that retainingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}alone induces a sharp decrease in top\-3 accuracy around layers 14–15\. However, Figure[3](https://arxiv.org/html/2609.35970#S4.F3)\(c\) shows that retainingHABrH\_\{AB\}^\{r\}preserves performance relative to the unmodified baseline until around layer 17, after which performance decreases sharply\. Because we intervene only at the last\-token position and at a particular layer, the model may recompute task\-relevant information downstream of the intervention\. We therefore also include a𝝁r\\bm\{\\mu\}^\{r\}\-only baseline to measure this recovery\. Before layer 17, retainingHABCrH\_\{ABC\}^\{r\}gives performance similar to the𝝁r\\bm\{\\mu\}^\{r\}\-only baseline\. After layer 17, however, retainingHABCrH\_\{ABC\}^\{r\}substantially improves performance over the𝝁r\\bm\{\\mu\}^\{r\}\-only baseline and even exceeds the unmodified baseline\. Together, ensemble ablation and replacement experiments show a transition in causal relevance fromHABrH\_\{AB\}^\{r\}toHABCrH\_\{ABC\}^\{r\}around layers 17–18\. Between layers 15–17, ablatingHABrH\_\{AB\}^\{r\}causes a sharp decrease in performance while retainingHABrH\_\{AB\}^\{r\}preserves it\. The same aforementioned properties hold forHABCrH\_\{ABC\}^\{r\}starting around layers 17–18\. This transition occurs at the same depth where Section[3](https://arxiv.org/html/2609.35970#S3)finds thatHABCrH\_\{ABC\}^\{r\}becomes organized according to the answerDD, the energy fraction ofHABrH\_\{AB\}^\{r\}starts decreasing, and the energy fraction ofHABCrH\_\{ABC\}^\{r\}rises sharply\. AlthoughHBCrH\_\{BC\}^\{r\}andHCArH\_\{CA\}^\{r\}are also strongly organized by their corresponding pairwise differences \(Section[3](https://arxiv.org/html/2609.35970#S3)\), ablating them has little effect on the performance, and retaining them does not reproduce the layerwise replacement behavior ofHABrH\_\{AB\}^\{r\}\. Additional ablation and replacement experiments for different models and domains are in Appendices[G](https://arxiv.org/html/2609.35970#A7)and[H](https://arxiv.org/html/2609.35970#A8), respectively\. Appendix[F](https://arxiv.org/html/2609.35970#A6)further shows that ablatingHABrH\_\{AB\}^\{r\}at layer 15 prevents the later emergence ofHABCrH\_\{ABC\}^\{r\}\. ## 5Steering So far, we have shown that the ensemblesHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}are causally relevant, but this does not tell us if the model uses their organization according toγ\\gammaorDD\. If the model indeed uses this organization, then changingγ→γ\+δ\\gamma\\to\\gamma\+\\deltashould shift the predicted answer byδ\\delta, that is,D→D\+δD\\to D\+\\delta\. In this section, we test this by steering alongγ\\gammaandDD\. There are multiple ways to steer usingHABrH\_\{AB\}^\{r\}: shiftingB→B\+δB\\to B\+\\deltasuch that𝒉ABr→𝒉A,B\+δr\\bm\{h\}\_\{AB\}^\{r\}\\to\\bm\{h\}\_\{A,B\+\\delta\}^\{r\}, shiftingA→A−δA\\to A\-\\deltasuch that𝒉ABr→𝒉A−δ,Br\\bm\{h\}\_\{AB\}^\{r\}\\to\\bm\{h\}\_\{A\-\\delta,B\}^\{r\}, or using an averaged vector that depends only onγ\\gamma\. Results using a combination of the first two strategies are reported in Appendix[I](https://arxiv.org/html/2609.35970#A9)\. Here, we describe the last strategy\. We define an averaged vector𝒉¯γr\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}by444The chosen steering vector𝒉¯γr\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}is closely related to the definition of the mode fraction in Equation[5](https://arxiv.org/html/2609.35970#S3.E5)\. We show this correspondence in Appendix[B\.1](https://arxiv.org/html/2609.35970#A2.SS1)\. 𝒉¯γr:=𝔼ABB−A=γAB𝒉ABr\.\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}:=\\mathop\{\\mathbb\{E\}\_\{AB\}\}\_\{B\-A=\\gamma\}\\bm\{h\}\_\{AB\}^\{r\}\.\(9\) To steer, we replace the activationϕr\\bm\{\\phi\}^\{r\}withϕ~r\\bm\{\\widetilde\{\\phi\}\}^\{r\}according to ϕ~r\(A,B,C\)=ϕr\(A,B,C\)\+α\(𝒉¯γ\+δr−𝒉¯γr\)\.\\bm\{\\widetilde\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\phi\}^\{r\}\(A,B,C\)\+\\alpha\\left\(\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\+\\delta\}\-\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}\\right\)\.\(10\) Figure 4:Steering using different vectors as a function of layer atα=1\\alpha=1\. The median is evaluated among all the prompts for each layer andδ\\delta\.\(a\)We see that steering using𝒉¯γ\\bm\{\\bar\{h\}\}\_\{\\gamma\}is most effective between layers 15–17\.\(b\)In contrast, steering using𝒉¯D\\bm\{\\bar\{h\}\}\_\{D\}becomes effective from layer 18 onward\.\(c\)Top\-3 accuracy forδ=4\\delta=4after steering\. The dashed line shows the unmodified model’s top\-3 accuracy with respect to the target answerD\+δD\+\\delta\.Similarly, we define the steering vector𝒉¯D\\bm\{\\bar\{h\}\}\_\{D\}for the ensembleHABCrH\_\{ABC\}^\{r\}as 𝒉¯Dr:=𝔼ABCC\+B−A=DABC𝒉ABCr\\bm\{\\bar\{h\}\}^\{r\}\_\{D\}:=\\mathop\{\\mathbb\{E\}\_\{ABC\}\}\_\{C\+B\-A=D\}\\bm\{h\}\_\{ABC\}^\{r\}\(11\)and steer using Equation[10](https://arxiv.org/html/2609.35970#S5.E10)after replacing𝒉¯γr\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}with𝒉¯Dr\\bm\{\\bar\{h\}\}^\{r\}\_\{D\}\. The steering method is equivalent to the difference\-of\-means steering, analogous to that of[Rimsky et al\. \(2024\)](https://arxiv.org/html/2609.35970#bib.bib38)and[Subramani et al\. \(2022\)](https://arxiv.org/html/2609.35970#bib.bib37)\(see Appendix[B\.2](https://arxiv.org/html/2609.35970#A2.SS2)\)\. LetD∗=D\+δD^\{\*\}=D\+\\deltabe the target answer\. To measure the effect of steering, we compute the top\-3 logit marginm3m\_\{3\}, defined as the difference between the target logitzD∗z\_\{D^\{\*\}\}and the third\-largest non\-target logitz3z\_\{3\}: m3=zD∗−z3\.m\_\{3\}=z\_\{D^\{\*\}\}\-z\_\{3\}\.\(12\)Thusm3≥0m\_\{3\}\\geq 0iff the target logit is in top\-3\. Finally, we take the difference between steered and baselinem3m\_\{3\}to get Δm3=m3steered−m3baseline\.\\Delta m\_\{3\}=m\_\{3\}^\{\\text\{steered\}\}\-m\_\{3\}^\{\\text\{baseline\}\}\.\(13\) In Figures[4](https://arxiv.org/html/2609.35970#S5.F4)\(a–b\), we reportmedianΔm3\\mathrm\{median\\;\}\\Delta m\_\{3\}for steering using𝒉¯γr\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\}and𝒉¯Dr\\bm\{\\bar\{h\}\}\_\{D\}^\{r\}atα=1\\alpha=1, and Figure[4](https://arxiv.org/html/2609.35970#S5.F4)\(c\) shows the top\-3 accuracy after steering forδ=4\\delta=4\. Steering using𝒉¯γr\\bm\{\\bar\{h\}\}^\{r\}\_\{\\gamma\}becomes effective around layers 14–15, but around layers 17–18, its effect decreases and steering using𝒉¯Dr\\bm\{\\bar\{h\}\}^\{r\}\_\{D\}becomes effective\. ## 6Cross\-domain transfer Our preceding experiments show thatHABrH\_\{AB\}^\{r\}is causally relevant in the middle layers whileHABCrH\_\{ABC\}^\{r\}is relevant in the later layers\. We now want to understand whether one interaction vector extracted from one context can substitute for the corresponding vector in another context performing the same underlying task\. Prior work has shown that relational and task\-level representations can be reused across contexts\([Wang et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib14);[Todd et al\., 2024](https://arxiv.org/html/2609.35970#bib.bib39)\)\. We focus on the interaction ensembleHABrH\_\{AB\}^\{r\}here; the performance of transplantingHABCrH\_\{ABC\}^\{r\}varies across model and domain combinations\. Appendix[J](https://arxiv.org/html/2609.35970#A10)shows thatHABCrH\_\{ABC\}^\{r\}can be transplanted across domains when their subspaces are aligned\. We will usedomain\-1to refer to the original set of prompts that we have discussed so far and we usedomain\-2for the following set of prompts: ``` Hi {name}, the {noun} is scheduled between {day} {A} o’clock and {day} {B} o’clock. This is the same time as between {day} {C} o’clock and {day} ``` whereA,B,C∈𝒳\(2\)=\{1,…,12\}A,B,C\\in\\mathcal\{X\}^\{\(2\)\}=\\\{1,\\ldots,12\\\}and \(`name`,`noun`,`day`\) are replicate variables \(Appendix[A\.1](https://arxiv.org/html/2609.35970#A1.SS1)\)\. We then decompose the domain\-2 ensemble according to Section[2\.2](https://arxiv.org/html/2609.35970#S2.SS2)\. We transplant an interaction ensemble by replacing each domain\-1 interaction vector with the corresponding domain\-2 interaction vector indexed by the same values of the underlying cyclic variables\. We consider ϕ~r1,\(1\)\(A,B,C\)=ϕr1,\(1\)\(A,B,C\)−𝒉ABr1,\(1\)\+𝒉ABr2,\(2\),\\widetilde\{\\bm\{\\phi\}\}^\{r\_\{1\},\(1\)\}\(A,B,C\)=\\bm\{\\phi\}^\{r\_\{1\},\(1\)\}\(A,B,C\)\-\\bm\{h\}^\{r\_\{1\},\(1\)\}\_\{AB\}\+\\bm\{h\}^\{r\_\{2\},\(2\)\}\_\{AB\},\(14\)where superscripts\(1\),\(2\)\(1\),\(2\)refer to the domain andr1,r2r\_\{1\},r\_\{2\}refer to the corresponding replicate\. Analogous to the ensemble replacement experiments in Section[4\.2](https://arxiv.org/html/2609.35970#S4.SS2), we also consider ϕ~r1,\(1\)\(A,B,C\)=𝝁r1,\(1\)\+𝒉ABr2,\(2\)\.\\widetilde\{\\bm\{\\phi\}\}^\{r\_\{1\},\(1\)\}\(A,B,C\)=\\bm\{\\mu\}^\{r\_\{1\},\(1\)\}\+\\bm\{h\}\_\{AB\}^\{r\_\{2\},\(2\)\}\.\(15\)In both cases, the model runs unmodified after the intervention and we score relative to domain\-1\. Figure[5](https://arxiv.org/html/2609.35970#S6.F5)\(a\) shows that replacing an ablated domain\-1 interaction ensemble with the corresponding domain\-2 ensemble restores the performance lost under ablation\. Figure[5](https://arxiv.org/html/2609.35970#S6.F5)\(b\) shows that the domain\-2 interaction ensemble also reproduces the layerwise trend observed in the replacement experiment in Section[4\.2](https://arxiv.org/html/2609.35970#S4.SS2)\. Figure 5:Shaded bands show the central 80% interval across 10 randomly sampled replicate pairs\(r1,r2\)\(r\_\{1\},r\_\{2\}\)\.\(a\)Transplanting the interaction ensembleHAB\(2\)H\_\{AB\}^\{\(2\)\}from domain\-2 to domain\-1 preserves performance\.\(b\)Transplanting and retaining only the domain\-1 mean𝝁\(1\)\\bm\{\\mu\}^\{\(1\)\}and domain\-2 interaction ensembleHAB\(2\)H^\{\(2\)\}\_\{AB\}also preserves performance compared to the dashed blue within\-domain replacement experiments\.\(c\)Cross\-domain steering using𝒉¯γ\+δr2,\(2\)−𝒉¯γr1,\(1\)\\bm\{\\bar\{h\}\}\_\{\\gamma\+\\delta\}^\{r\_\{2\},\(2\)\}\-\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\_\{1\},\(1\)\}atα=1\\alpha=1\.We next ask whether the organization according toγ\\gammaalso transfers across domains\. Using the steering vector defined in Equation[9](https://arxiv.org/html/2609.35970#S5.E9), we steer using the difference between the domain\-1 vector forγ\\gammaand the domain\-2 vector forγ\+δ\\gamma\+\\delta: ϕ~r1,\(1\)\(A,B,C\)=ϕr1,\(1\)\(A,B,C\)\+α\(𝒉¯γ\+δr2,\(2\)−𝒉¯γr1,\(1\)\)\.\\widetilde\{\\bm\{\\phi\}\}^\{r\_\{1\},\(1\)\}\(A,B,C\)=\\bm\{\\phi\}^\{r\_\{1\},\(1\)\}\(A,B,C\)\+\\alpha\\left\(\\bm\{\\bar\{h\}\}\_\{\\gamma\+\\delta\}^\{r\_\{2\},\(2\)\}\-\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\_\{1\},\(1\)\}\\right\)\.\(16\)If the organization byγ\\gammatransfers between two domains, then the answer should shift accordingly\. Figure[5](https://arxiv.org/html/2609.35970#S6.F5)\(c\) shows that steering vectors constructed across the two domains produce a steering pattern similar to that in Section[5](https://arxiv.org/html/2609.35970#S5)\. ## 7Discussion We started by asking whether it is possible to separate variation due to individual concepts from variation arising through their joint dependence in an activation ensemble\. We showed that using the functional ANOVA decomposition is a simple yet powerful technique to accomplish this task\. The resulting interaction ensembles exhibit clear geometric structure, but not all geometrically organized ensembles are causally necessary for the task\. For the ensembles that are causally relevant, their organization can be exploited to predictably steer the model\. We also found two results that we did not anticipate\. First, retaining only the mean and the third\-order interaction can outperform the unmodified model\. Our interventions show thatHABCrH\_\{ABC\}^\{r\}is used by the LLM to output the correct answer, while other interaction ensembles either are unused or perform a different function\. Therefore, it may happen that those other interaction ensembles interfere destructively in the computation in the last few layers\. Second, transplantingHABH\_\{AB\}between domains that perform the same underlying task recovers much of the performance lost under ablation\. This suggests that the intermediate representation of the inferred relation is sufficiently compatible across domains to support downstream computation\. Neither observation would have been accessible in the full activation ensemble without first isolating the joint dependence from the remaining components\. Our results also suggest a possible connection to the neuron\-level analysis of[Feucht et al\. \(2026\)](https://arxiv.org/html/2609.35970#bib.bib17)\. They identify a small set of MLP neurons at layer 18 that are associated with Fourier components of the arithmetic representation in Llama\-3\.1\-8B\. In our analysis, the same depth is where the three\-way interactionHABCrH\_\{ABC\}^\{r\}sharply increases in energy fraction, becomes organized byDD, and becomes causally relevant\. While we do not analyze individual neurons here, the interaction decomposition may provide a way to identify the layers in which such structure emerges and may also provide a model\-agnostic starting point for neuron\-level analyses\. ### Limitations Scope\. We have studied a controlled setting with three variablesA,B,CA,B,Cthat are drawn from the*same cyclic*concept whose composition has a*definitive*arithmetic answer\. A natural extension is to allowA,B,CA,B,Cto vary over different, potentially non\-cyclic concepts whose composition may not have a well\-defined answer\. Exponential scaling\. As we increase the number of variablespp, the number of prompts required to furnish the ANOVA decomposition scales exponentially∼\|𝒳\|p\\sim\|\\mathcal\{X\}\|^\{p\}\. ### Acknowledgments This research is funded in part by the Gordon and Betty Moore Foundation through Grant GBMF13901 to Cornell to support the work of G\.A\. ### AI use statement We have used generative AI \(GPT\-5\.6 and GPT\-6\) to code, proofread the paper to improve readability, correct typos and grammatical errors\. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\. ## References - Aroraet al\.\(2024\)A\. Arora, D\. Jurafsky, and C\. PottsCausalGym: benchmarking causal interpretability methods on linguistic tasks\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 14638–14663\.External Links:[Link](https://aclanthology.org/2024.acl-long.785/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.785)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px3.p1.1)\. - Daiet al\.\(2026\)Q\. Dai, B\. Heinzerling, and K\. InuiCell\-based representation of relational binding in language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 47464–47524\.External Links:[Link](https://aclanthology.org/2026.acl-long.2194/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2194),ISBN 979\-8\-89176\-390\-6Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Davieset al\.\(2023\)X\. Davies, M\. Nadeau, N\. Prakash, T\. R\. Shaham, and D\. BauDiscovering variable binding circuitry with desiderata\.External Links:2307\.03637,[Link](https://arxiv.org/abs/2307.03637)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Engelset al\.\(2025\)J\. Engels, E\. J\. Michaud, I\. Liao, W\. Gurnee, and M\. TegmarkNot all language model features are one\-dimensionally linear\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d63a4AM4hb)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.35970#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.35970#S2.SS1.p1.1)\. - Feng and Steinhardt \(2024\)J\. Feng and J\. SteinhardtHow do language models bind entities in context?\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zb3b6oKO77)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Feuchtet al\.\(2026\)S\. Feucht, T\. Haklay, U\. Bhalla, D\. Wurgaft, C\. Rager, R\. Sarfati, J\. Merullo, T\. McGrath, O\. Lewis, E\. S\. Lubana, T\. Fel, and A\. GeigerArithmetic in the wild: llama uses base\-10 addition to reason about cyclic concepts\.External Links:2605\.01148,[Link](https://arxiv.org/abs/2605.01148)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1),[§7](https://arxiv.org/html/2609.35970#S7.p3.1)\. - Furutaet al\.\(2024\)H\. Furuta, G\. Minegishi, Y\. Iwasawa, and Y\. MatsuoTowards empirical interpretation of internal circuits and properties in grokked transformers on modular polynomials\.External Links:2402\.16726,[Link](https://arxiv.org/abs/2402.16726)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Geigeret al\.\(2021\)A\. Geiger, H\. Lu, T\. Icard, and C\. PottsCausal abstractions of neural networks\.InAdvances in Neural Information Processing Systems,M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\.S\. Liang, and J\. W\. Vaughan \(Eds\.\),Vol\.34,pp\. 9574–9586\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2021/file/4f5c422f4d49a5a807eda27434231040-Paper.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px3.p1.1)\. - Geigeret al\.\(2024\)A\. Geiger, Z\. Wu, C\. Potts, T\. Icard, and N\. D\. GoodmanFinding alignments between interpretable causal variables and distributed neural representations\.External Links:2303\.02536,[Link](https://arxiv.org/abs/2303.02536)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px3.p1.1)\. - Gemma Team \(2024\)Gemma TeamGemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Link](https://arxiv.org/abs/2408.00118)Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Gemma Team \(2025\)Gemma TeamGemma 3 technical report\.Kaggle\.External Links:[Link](https://goo.gle/Gemma3Report)Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Gurneeet al\.\(2025\)W\. Gurnee, E\. Ameisen, I\. Kauvar, J\. Tarng, A\. Pearce, C\. Olah, and J\. BatsonWhen models manipulate manifolds: the geometry of a counting task\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2025/linebreaks/index.html)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Hernandezet al\.\(2024\)E\. Hernandez, A\. S\. Sharma, T\. Haklay, K\. Meng, M\. Wattenberg, J\. Andreas, Y\. Belinkov, and D\. BauLinearity of relation decoding in transformer language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=w7LU2s14kE)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Hoeffding \(1948\)W\. HoeffdingA Class of Statistics with Asymptotically Normal Distribution\.The Annals of Mathematical Statistics19\(3\),pp\. 293 – 325\.External Links:[Document](https://dx.doi.org/10.1214/aoms/1177730196),[Link](https://doi.org/10.1214/aoms/1177730196)Cited by:[§1](https://arxiv.org/html/2609.35970#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.35970#S2.SS2.p4.1)\. - Huet al\.\(2026\)Z\. Hu, L\. Niu, and S\. VarmaLanguage models represent and transform concepts with shared geometry\.InMechanistic Interpretability Workshop at ICML 2026,External Links:[Link](https://openreview.net/forum?id=wg6Q5fPL1m)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Kantamneni and Tegmark \(2025\)S\. Kantamneni and M\. TegmarkLanguage models use trigonometry to do addition\.External Links:2502\.00873,[Link](https://arxiv.org/abs/2502.00873)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1),[§2\.1](https://arxiv.org/html/2609.35970#S2.SS1.p1.1)\. - Karkadaet al\.\(2026\)D\. Karkada, D\. J\. Korchinski, A\. Nava, M\. Wyart, and Y\. BahriSymmetries in language statistics shape the geometry of model representations\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=XQeMPEkfdd)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.35970#S3.SS1.p3.1)\. - Levy and Geva \(2025\)A\. A\. Levy and M\. GevaLanguage models encode numbers using digit representations in base 10\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 2: Short Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 385–395\.External Links:[Link](https://aclanthology.org/2025.naacl-short.33/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-short.33),ISBN 979\-8\-89176\-190\-2Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Llama Team \(2024a\)Llama TeamLlama\-3\.2\-3B\.Note:Hugging Face model cardExternal Links:[Link](https://huggingface.co/meta-llama/Llama-3.2-3B)Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Llama Team \(2024b\)Llama TeamThe llama 3 herd of models\.CoRRabs/2407\.21783\.External Links:[Link](https://doi.org/10.48550/arXiv.2407.21783),[Document](https://dx.doi.org/10.48550/ARXIV.2407.21783),2407\.21783Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Merulloet al\.\(2024\)J\. Merullo, C\. Eickhoff, and E\. PavlickLanguage models implement simple Word2Vec\-style vector arithmetic\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 5030–5047\.External Links:[Link](https://aclanthology.org/2024.naacl-long.281/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.281)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Mistral AI Team \(2025\)Mistral AI TeamMistral\-Small\-24B\-Base\-2501\.Note:[https://huggingface\.co/mistralai/Mistral\-Small\-24B\-Base\-2501](https://huggingface.co/mistralai/Mistral-Small-24B-Base-2501)Model cardCited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Modellet al\.\(2025\)A\. Modell, P\. Rubin\-Delanchy, and N\. WhiteleyThe origins of representation manifolds in large language models\.External Links:2505\.18235,[Link](https://arxiv.org/abs/2505.18235)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Nandaet al\.\(2023\)N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. SteinhardtProgress measures for grokking via mechanistic interpretability\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=9XFSbDPmdW)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Owen \(2013\)A\. B\. OwenMonte carlo theory, methods and examples\.[https://artowen\.su\.domains/mc/](https://artowen.su.domains/mc/)\.Cited by:[§2\.2](https://arxiv.org/html/2609.35970#S2.SS2.p4.1)\. - Parket al\.\(2025\)K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. VeitchThe geometry of categorical and hierarchical concepts in large language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=bVTM2QKYuA)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Prietoet al\.\(2026\)L\. Prieto, E\. Stevinson, M\. Barsbey, T\. Birdal, and P\. A\. M\. MedianoFrom data statistics to feature geometry: how correlations shape superposition\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=7akSRQS5Xh)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Qwen Team \(2024\)Qwen TeamQwen2\.5: a party of foundation models\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5/)Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Qwen Team \(2025\)Qwen TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§2](https://arxiv.org/html/2609.35970#S2.p1.1)\. - Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15504–15522\.External Links:[Link](https://aclanthology.org/2024.acl-long.828/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828)Cited by:[§5](https://arxiv.org/html/2609.35970#S5.p3.2)\. - Sarfatiet al\.\(2026\)R\. Sarfati, E\. Bigelow, D\. Wurgaft, S\. Boppana, J\. Merullo, A\. Geiger, O\. Lewis, T\. McGrath, and E\. S\. LubanaThe shape of beliefs: geometry, dynamics, and interventions along representation manifolds of language models’ posteriors\.External Links:2602\.02315,[Link](https://arxiv.org/abs/2602.02315)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Social Security Administration \(n\.d\.\)Social Security AdministrationTop names over the last 100 years\.Note:Accessed in 2026External Links:[Link](https://www.ssa.gov/oact/babynames/decades/century.html)Cited by:[Table 2](https://arxiv.org/html/2609.35970#A2.T2.2.2.2.1.1),[§2\.1](https://arxiv.org/html/2609.35970#S2.SS1.p1.3)\. - Subramaniet al\.\(2022\)N\. Subramani, N\. Suresh, and M\. PetersExtracting latent steering vectors from pretrained language models\.InFindings of the Association for Computational Linguistics: ACL 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 566–581\.External Links:[Link](https://aclanthology.org/2022.findings-acl.48/),[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.48)Cited by:[§5](https://arxiv.org/html/2609.35970#S5.p3.2)\. - Toddet al\.\(2026\)E\. Todd, J\. Brinkmann, R\. Gandikota, and D\. BauIn\-context algebra\.InInternational Conference on Learning Representations,C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(Eds\.\),Vol\.2026,pp\. 80695–80729\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/file/82aec8518602748540a42b783468c94d-Paper-Conference.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. - Toddet al\.\(2024\)E\. Todd, M\. Li, A\. S\. Sharma, A\. Mueller, B\. C\. Wallace, and D\. BauFunction vectors in large language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=AwyxtyMwaG)Cited by:[§6](https://arxiv.org/html/2609.35970#S6.p1.1)\. - Viget al\.\(2020\)J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. ShieberInvestigating gender bias in language models using causal mediation analysis\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 12388–12401\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/92650b2e92217715fe312e6fa7b90d82-Paper.pdf)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px3.p1.1)\. - Wanget al\.\(2024\)Z\. Wang, B\. Whyte, and C\. XuLocating and extracting relational concepts in large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 4818–4832\.External Links:[Link](https://aclanthology.org/2024.findings-acl.287/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.287)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.35970#S6.p1.1)\. - Wurgaftet al\.\(2026\)D\. Wurgaft, C\. Rager, M\. Kowal, S\. Feucht, U\. Bhalla, T\. Haklay, E\. Bigelow, R\. Sarfati, J\. Merullo, N\. Goodman, T\. Fel, A\. Geiger, and E\. S\. LubanaManifold steering reveals the shared geometry of neural network representation and behavior\.InThird Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=pf3HKAGT2M)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px1.p1.1)\. - Zhang and Nanda \(2024\)F\. Zhang and N\. NandaTowards best practices of activation patching in language models: metrics and methods\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Hf17y6u9BC)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.35970#S4.p1.1)\. - Zhouet al\.\(2024\)T\. Zhou, D\. Fu, V\. Sharan, and R\. JiaPre\-trained large language models use fourier features to compute addition\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=i4MutM2TZb)Cited by:[§1\.1](https://arxiv.org/html/2609.35970#S1.SS1.SSS0.Px2.p1.1)\. ## Appendix AModels, Prompts, and Capabilities In this work, we evaluate the robustness of our results across seven base models ranging from 3B to 24B parameters: Llama\-3\.2\-3B, Llama\-3\.1\-8B, Qwen\-2\.5\-7B, Qwen\-3\-8B, Gemma\-2\-9B, Gemma\-3\-12B, Mistral\-small\-24B\-Base\-2501\. Table[1](https://arxiv.org/html/2609.35970#A1.T1)lists all the prompts for each model, while Table[4](https://arxiv.org/html/2609.35970#A2.T4)shows their corresponding performances using top\-1 and top\-3 accuracy as metrics\. Here, we report performance using all prompts, including cases with coincident variables, that is, whenA=BA=B,B=CB=C, orC=AC=A\. We use two conditions to fix the prompt template:\(a\)The model should be capable of answering the problem via the next token\.\(b\)All variations overA,B,CA,B,Cmust tokenize to the same width\. Therefore we exclude evaluation of the hours domain for the Qwen, Gemma, and Mistral families since all of them use a single digit tokenizer\. ### A\.1List of variables Tables[2](https://arxiv.org/html/2609.35970#A2.T2)and[3](https://arxiv.org/html/2609.35970#A2.T3)show the values of variables used for replicate variables andA,B,CA,B,Crespectively\. ### A\.2Capabilities Figure[6](https://arxiv.org/html/2609.35970#A1.F6)shows the top\-1 and top\-3 accuracy of each model indexed byγ\\gammaover prompt templates in Table[1](https://arxiv.org/html/2609.35970#A1.T1)\. Table[4](https://arxiv.org/html/2609.35970#A2.T4)lists the aggregate top\-1 and top\-3 accuracy over replicates\. Table 1:Prompt templates used for each model and cyclic domain\.A,B,CA,B,Cand the replicate variables vary as described in Tables[2](https://arxiv.org/html/2609.35970#A2.T2)and[3](https://arxiv.org/html/2609.35970#A2.T3)\.N/Aindicates domains excluded because the cyclic values do not tokenize to a common width\.ModelDomainPrompt templateLlama\-3\.2\-3BmonthsHi \{name\}, the \{noun\} is scheduled from \{A\} \{dates\} to \{B\} \{dates\}\. This is the same as time from \{C\} \{dates\} tohoursHi \{name\}, the \{noun\} is scheduled from \{day\} \{A\} o’clock to \{day\} \{B\} o’clock\. This is the same time as from \{day\} \{C\} o’clock to \{day\}weekdaysHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same as time from \{C\} tomusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteLlama\-3\.1\-8BmonthsHi \{name\}, the \{noun\} is scheduled between \{A\} \{dates\} and \{B\} \{dates\}\. This is the same time as between \{C\} \{dates\} andhoursHi \{name\}, the \{noun\} is scheduled between \{day\} \{A\} o’clock and \{day\} \{B\} o’clock\. This is the same time as between \{day\} \{C\} o’clock and \{day\}weekdaysHi \{name\}, the \{noun\} is scheduled between \{A\} and \{B\}\. This is the same as time between \{C\} andmusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteQwen\-2\.5\-7BmonthsHello \{name\}, the \{noun\} is scheduled from \{A\} \{datesDouble\} to \{B\} \{datesDouble\}\. This is the same as duration from \{C\} \{datesDouble\} tohoursN/AweekdaysHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same time as from \{C\} tomusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteQwen\-3\-8BmonthsHello \{name\}, the \{noun\} is scheduled from \{A\} \{datesDouble\} to \{B\} \{datesDouble\}\. This is the same time as from \{C\} \{datesDouble\} tohoursN/AweekdaysHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same time as from \{C\} tomusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteGemma\-2\-9BmonthsHello \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same as duration from \{C\} tohoursN/AweekdaysHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same time as from \{C\} tomusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteGemma\-3\-12BmonthsHello \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same time as from \{C\} tohoursN/AweekdaysHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same time as from \{C\} tomusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteMistral\-small\-24B\-Base\-2501monthsHi \{name\}, the \{noun\} is scheduled from \{A\} to \{B\}\. This is the same as time from \{C\} tohoursN/AweekdaysHi \{name\}, the \{noun\} is scheduled between \{A\} and \{B\}\. This is the same as time between \{C\} andmusicHi \{name\}, on the musical scale, the interval from the note \{A\} to the note \{B\} is the same as the interval from the note \{C\} to the noteFigure 6:Top\-1 and top\-3 accuracy as a function ofγ=B−A\\gamma=B\-Aacross models and cyclic concepts\. Columns correspond to models and rows to concepts\. Curves show accuracy averaged across the 10 sampled replicates, with shaded bands showing the central 80% interval\. Hours are omitted for Qwen, Gemma, and Mistral because the relevant numerical values do not tokenize to a common width\. Top\-3 accuracy is generally more stable acrossγ\\gammathan top\-1 accuracy\. ## Appendix BTheoretical background This appendix collects properties of the functional ANOVA decomposition used throughout the paper and establishes the identities underlying the mode fraction, the energy fraction and steering analyses\. Let us revisit the functional ANOVA decomposition defined in the main text, ϕr\(A,B,C\)=𝝁r\+𝒉Ar\+𝒉Br\+𝒉Cr\+𝒉ABr\+𝒉BCr\+𝒉CAr\+𝒉ABCr,\\bm\{\\phi\}^\{r\}\(A,B,C\)=\\bm\{\\mu\}^\{r\}\+\\bm\{h\}\_\{A\}^\{r\}\+\\bm\{h\}\_\{B\}^\{r\}\+\\bm\{h\}\_\{C\}^\{r\}\+\\bm\{h\}\_\{AB\}^\{r\}\+\\bm\{h\}\_\{BC\}^\{r\}\+\\bm\{h\}\_\{CA\}^\{r\}\+\\bm\{h\}\_\{ABC\}^\{r\},where 𝝁r\\displaystyle\\bm\{\\mu\}^\{r\}:=𝔼ABC\(ϕr\),\\displaystyle:=\\mathbb\{E\}\_\{ABC\}\(\\bm\{\\phi\}^\{r\}\),𝒉Ar\\displaystyle\\bm\{h\}\_\{A\}^\{r\}:=𝔼BC\(ϕr\)−𝝁r,\\displaystyle:=\\mathbb\{E\}\_\{BC\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\},𝒉ABr\\displaystyle\\bm\{h\}\_\{AB\}^\{r\}:=𝔼C\(ϕr\)−𝝁r−𝒉Ar−𝒉Br,\\displaystyle:=\\mathbb\{E\}\_\{C\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\},𝒉ABCr\\displaystyle\\bm\{h\}\_\{ABC\}^\{r\}:=ϕr−𝝁r−𝒉Ar−𝒉Br−𝒉Cr−𝒉ABr−𝒉BCr−𝒉CAr\.\\displaystyle:=\\bm\{\\phi\}^\{r\}\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\}\-\\bm\{h\}\_\{C\}^\{r\}\-\\bm\{h\}\_\{AB\}^\{r\}\-\\bm\{h\}\_\{BC\}^\{r\}\-\\bm\{h\}\_\{CA\}^\{r\}\.To standardize notation, we use𝒉∅r=𝝁r\\bm\{h\}^\{r\}\_\{\\emptyset\}=\\bm\{\\mu\}^\{r\}\. Using these definitions, it immediately follows that each non\-constant decomposed term has zero mean with respect to each variable on which it depends\. ###### Lemma 1\(Zero mean of ANOVA terms\)\. For every nonemptyS⊆\{A,B,C\}S\\subseteq\\\{A,B,C\\\}and everyX∈SX\\in S, 𝔼X\[𝒉Sr\]=𝟎\.\\mathbb\{E\}\_\{X\}\\\!\\left\[\\bm\{h\}\_\{S\}^\{r\}\\right\]=\\bm\{0\}\.\(17\) ###### Proof\. Consider𝒉Ar\\bm\{h\}\_\{A\}^\{r\}: 𝔼A\(𝒉Ar\)\\displaystyle\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{A\}^\{r\}\)=𝔼A\(𝔼BC\(ϕr\)−𝝁r\)\\displaystyle=\\mathbb\{E\}\_\{A\}\\left\(\\mathbb\{E\}\_\{BC\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\\right\)=𝔼ABC\(ϕr\)−𝝁r\\displaystyle=\\mathbb\{E\}\_\{ABC\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}=𝝁r−𝝁r\\displaystyle=\\bm\{\\mu\}^\{r\}\-\\bm\{\\mu\}^\{r\}=0\.\\displaystyle=0\.Similarly, for𝒉ABr\\bm\{h\}\_\{AB\}^\{r\} 𝔼A\(𝒉ABr\)\\displaystyle\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{AB\}^\{r\}\)=𝔼A\(𝔼C\(ϕr\)−𝝁r−𝒉Ar−𝒉Br\)\\displaystyle=\\mathbb\{E\}\_\{A\}\\left\(\\mathbb\{E\}\_\{C\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\}\\right\)=𝔼AC\(ϕr\)−𝝁r−𝔼A\(𝒉Ar\)−𝒉Br\\displaystyle=\\mathbb\{E\}\_\{AC\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\-\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{A\}^\{r\}\)\-\\bm\{h\}\_\{B\}^\{r\}=\(𝝁r\+𝒉Br\)−𝝁r−0−𝒉Br\\displaystyle=\\left\(\\bm\{\\mu\}^\{r\}\+\\bm\{h\}\_\{B\}^\{r\}\\right\)\-\\bm\{\\mu\}^\{r\}\-0\-\\bm\{h\}\_\{B\}^\{r\}=0\.\\displaystyle=0\.The same argument holds for𝔼B\(𝒉AB\)=0\\mathbb\{E\}\_\{B\}\(\\bm\{h\}\_\{AB\}\)=0\. Finally, consider𝒉ABCr\\bm\{h\}\_\{ABC\}^\{r\} 𝔼A\(𝒉ABCr\)\\displaystyle\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{ABC\}^\{r\}\)=𝔼A\(ϕr−𝝁r−𝒉Ar−𝒉Br−𝒉Cr−𝒉ABr−𝒉BCr−𝒉CAr\)\\displaystyle=\\mathbb\{E\}\_\{A\}\\left\(\\bm\{\\phi\}^\{r\}\-\\bm\{\\mu\}^\{r\}\-\\bm\{h\}\_\{A\}^\{r\}\-\\bm\{h\}\_\{B\}^\{r\}\-\\bm\{h\}\_\{C\}^\{r\}\-\\bm\{h\}\_\{AB\}^\{r\}\-\\bm\{h\}\_\{BC\}^\{r\}\-\\bm\{h\}\_\{CA\}^\{r\}\\right\)=𝔼A\(ϕr\)−𝝁r−𝔼A\(𝒉Ar\)−𝒉Br−𝒉Cr−𝔼A\(𝒉ABr\)−𝒉BCr−𝔼A\(𝒉CAr\)\\displaystyle=\\mathbb\{E\}\_\{A\}\(\\bm\{\\phi\}^\{r\}\)\-\\bm\{\\mu\}^\{r\}\-\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{A\}^\{r\}\)\-\\bm\{h\}\_\{B\}^\{r\}\-\\bm\{h\}\_\{C\}^\{r\}\-\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{AB\}^\{r\}\)\-\\bm\{h\}\_\{BC\}^\{r\}\-\\mathbb\{E\}\_\{A\}\(\\bm\{h\}\_\{CA\}^\{r\}\)=\(𝝁r\+𝒉Br\+𝒉Cr\+𝒉BCr\)−𝝁r−0−𝒉Br−𝒉Cr−0−𝒉BCr−0\\displaystyle=\\left\(\\bm\{\\mu\}^\{r\}\+\\bm\{h\}\_\{B\}^\{r\}\+\\bm\{h\}\_\{C\}^\{r\}\+\\bm\{h\}\_\{BC\}^\{r\}\\right\)\-\\bm\{\\mu\}^\{r\}\-0\-\\bm\{h\}\_\{B\}^\{r\}\-\\bm\{h\}\_\{C\}^\{r\}\-0\-\\bm\{h\}\_\{BC\}^\{r\}\-0=0\.\\displaystyle=0\.Similarly,𝔼B\(𝒉ABCr\)=𝔼C\(𝒉ABCr\)=0\\mathbb\{E\}\_\{B\}\(\\bm\{h\}\_\{ABC\}^\{r\}\)=\\mathbb\{E\}\_\{C\}\(\\bm\{h\}\_\{ABC\}^\{r\}\)=0\. In general, it can be proved by induction\. ∎ ###### Definition 1\(Ensemble\-averaged form\)\. ForS,T⊆\{A,B,C\}S,T\\subseteq\\\{A,B,C\\\}and two decomposed ensemblesHSrH\_\{S\}^\{r\}andHTrH\_\{T\}^\{r\}with vectors𝐡Sr∈HSr\\bm\{h\}\_\{S\}^\{r\}\\in H\_\{S\}^\{r\}and𝐡Tr∈HTr\\bm\{h\}\_\{T\}^\{r\}\\in H\_\{T\}^\{r\}, define the ensemble\-averaged form as ⟨HSr,HTr⟩ens:=𝔼ABC\[⟨𝒉Sr,𝒉Tr⟩\]\\left\\langle H\_\{S\}^\{r\},H\_\{T\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}:=\\mathbb\{E\}\_\{ABC\}\\left\[\\left\\langle\\bm\{h\}\_\{S\}^\{r\},\\bm\{h\}\_\{T\}^\{r\}\\right\\rangle\\right\]\(18\)where the inner product on the RHS is the standard Euclidean inner product defined overℝdmodel\\mathbb\{R\}^\{\\text\{d\}\_\{\\text\{model\}\}\}\. Table 2:Replicate variables used in the experiments\.Table 3:Values assigned to the cyclic variablesAA,BB, andCCin each domain\. Months and hours have cycle lengthN=12N=12, while weekdays and musical notes haveN=7N=7\.Table 4:Top\-1 and top\-3 accuracy by model and domain\. Values are mean±\\pmstandard deviation across 10 replicates, computed over all prompts, including coincident values ofAA,BB, andCC\.###### Proposition 1\(Inner product\)\. Under conditions of Definition[1](https://arxiv.org/html/2609.35970#Thmdefinition1), Equation[18](https://arxiv.org/html/2609.35970#A2.E18)satisfies the properties of an inner product: 1. 1\.Symmetry\.⟨HSr,HTr⟩ens=⟨HTr,HSr⟩ens\\left\\langle H\_\{S\}^\{r\},H\_\{T\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=\\left\\langle H\_\{T\}^\{r\},H\_\{S\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\. 2. 2\.Linearity\.⟨αHSr\+βHTr,HUr⟩ens=α⟨HSr,HUr⟩ens\+β⟨HTr,HUr⟩ens\\left\\langle\\alpha H\_\{S\}^\{r\}\+\\beta H\_\{T\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=\\alpha\\left\\langle H\_\{S\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\+\\beta\\left\\langle H\_\{T\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}for anyα,β∈ℝ\\alpha,\\beta\\in\\mathbb\{R\}\. 3. 3\.Positive definiteness\.⟨HSr,HSr⟩ens≥0\\left\\langle H\_\{S\}^\{r\},H\_\{S\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\\geq 0and the equality holds iffHSr=0H\_\{S\}^\{r\}=0\. ###### Proof\. Symmetry\.By symmetry of the Euclidean inner product, ⟨HSr,HTr⟩ens=𝔼ABC\[⟨𝒉Sr,𝒉Tr⟩\]=𝔼ABC\[⟨𝒉Tr,𝒉Sr⟩\]=⟨HTr,HSr⟩ens\.\\left\\langle H\_\{S\}^\{r\},H\_\{T\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=\\mathbb\{E\}\_\{ABC\}\\left\[\\left\\langle\\bm\{h\}\_\{S\}^\{r\},\\bm\{h\}\_\{T\}^\{r\}\\right\\rangle\\right\]=\\mathbb\{E\}\_\{ABC\}\\left\[\\left\\langle\\bm\{h\}\_\{T\}^\{r\},\\bm\{h\}\_\{S\}^\{r\}\\right\\rangle\\right\]=\\left\\langle H\_\{T\}^\{r\},H\_\{S\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\. Linearity\.For anyα,β∈ℝ\\alpha,\\beta\\in\\mathbb\{R\}and ensemblesHSr,HTr,HUrH\_\{S\}^\{r\},H\_\{T\}^\{r\},H\_\{U\}^\{r\}, linearity of the Euclidean inner product and of expectation gives ⟨αHSr\+βHTr,HUr⟩ens\\displaystyle\\left\\langle\\alpha H\_\{S\}^\{r\}\+\\beta H\_\{T\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=𝔼ABC\[⟨α𝒉Sr\+β𝒉Tr,𝒉Ur⟩\]\\displaystyle=\\mathbb\{E\}\_\{ABC\}\\left\[\\left\\langle\\alpha\\bm\{h\}\_\{S\}^\{r\}\+\\beta\\bm\{h\}\_\{T\}^\{r\},\\bm\{h\}\_\{U\}^\{r\}\\right\\rangle\\right\]=α⟨HSr,HUr⟩ens\+β⟨HTr,HUr⟩ens\.\\displaystyle=\\alpha\\left\\langle H\_\{S\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\+\\beta\\left\\langle H\_\{T\}^\{r\},H\_\{U\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}\. Positive definiteness\.For every ensembleHSrH\_\{S\}^\{r\}, ⟨HSr,HSr⟩ens=𝔼ABC\[‖𝒉Sr‖22\]≥0\.\\left\\langle H\_\{S\}^\{r\},H\_\{S\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=\\mathbb\{E\}\_\{ABC\}\\left\[\\\|\\bm\{h\}\_\{S\}^\{r\}\\\|\_\{2\}^\{2\}\\right\]\\geq 0\.A nonnegative random variable has expectation zero if and only if it is zero\. Hence ⟨HSr,HSr⟩ens=0⟺𝒉Sr=0\.\\left\\langle H\_\{S\}^\{r\},H\_\{S\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=0\\quad\\Longleftrightarrow\\quad\\bm\{h\}\_\{S\}^\{r\}=0\.This is equivalent toHSr=0H\_\{S\}^\{r\}=0in the ensemble space\. Thus the form satisfies all three inner product axioms\. ∎ ###### Proposition 2\(Orthogonality of the ANOVA terms\)\. Under the full ensemble, any two distinct terms in the decomposition are orthogonal under the inner product in Definition[1](https://arxiv.org/html/2609.35970#Thmdefinition1): ⟨HSr,HTr⟩ens=0,S≠T\.\\left\\langle H\_\{S\}^\{r\},H\_\{T\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=0,\\qquad S\\neq T\.\(19\)Here,S,T⊆\{A,B,C\}S,T\\subseteq\\\{A,B,C\\\} ###### Proof\. LetS≠TS\\neq T\. Without loss of generality, choose a variableX∈S∖TX\\in S\\setminus T\. Since𝒉Tr\\bm\{h\}\_\{T\}^\{r\}does not depend onXX, we may average overXXfirst: ⟨HSr,HTr⟩ens\\displaystyle\\left\\langle H\_\{S\}^\{r\},H\_\{T\}^\{r\}\\right\\rangle\_\{\\mathrm\{ens\}\}=𝔼\{ABC\}∖\{X\}\[⟨𝔼X\[𝒉Sr\],𝒉Tr⟩\]\\displaystyle=\\mathbb\{E\}\_\{\\\{ABC\\\}\\setminus\\\{X\\\}\}\\left\[\\left\\langle\\mathbb\{E\}\_\{X\}\\left\[\\bm\{h\}\_\{S\}^\{r\}\\right\],\\bm\{h\}\_\{T\}^\{r\}\\right\\rangle\\right\]=0,\\displaystyle=0,where the final equality follows from Lemma[1](https://arxiv.org/html/2609.35970#Thmlemma1)\. Hence all distinct terms are pairwise orthogonal\. ∎ ###### Corollary 1\(Energy decomposition\)\. Under the conditions of Proposition[2](https://arxiv.org/html/2609.35970#Thmproposition2), 𝔼ABC‖ϕr\(A,B,C\)−𝒉∅r‖22=∑∅≠S⊆\{A,B,C\}𝔼S‖𝒉Sr‖22\.\\mathbb\{E\}\_\{ABC\}\\left\\\|\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{h\}\_\{\\emptyset\}^\{r\}\\right\\\|\_\{2\}^\{2\}=\\sum\_\{\\emptyset\\neq S\\subseteq\\\{A,B,C\\\}\}\\mathbb\{E\}\_\{S\}\\left\\\|\\bm\{h\}\_\{S\}^\{r\}\\right\\\|\_\{2\}^\{2\}\.\(20\) ###### Proof\. By Equation[2](https://arxiv.org/html/2609.35970#S2.E2), ϕr−𝒉∅r=∑∅≠S⊆\{A,B,C\}𝒉Sr\.\\bm\{\\phi\}^\{r\}\-\\bm\{h\}\_\{\\emptyset\}^\{r\}=\\sum\_\{\\emptyset\\neq S\\subseteq\\\{A,B,C\\\}\}\\bm\{h\}\_\{S\}^\{r\}\.Expanding the squared norm and applying Proposition[2](https://arxiv.org/html/2609.35970#Thmproposition2)eliminates all cross terms: 𝔼ABC‖ϕr−𝒉∅r‖22\\displaystyle\\mathbb\{E\}\_\{ABC\}\\left\\\|\\bm\{\\phi\}^\{r\}\-\\bm\{h\}\_\{\\emptyset\}^\{r\}\\right\\\|\_\{2\}^\{2\}=∑∅≠S⊆\{A,B,C\}𝔼ABC‖𝒉Sr‖22\\displaystyle=\\sum\_\{\\emptyset\\neq S\\subseteq\\\{A,B,C\\\}\}\\mathbb\{E\}\_\{ABC\}\\left\\\|\\bm\{h\}\_\{S\}^\{r\}\\right\\\|\_\{2\}^\{2\}=∑∅≠S⊆\{A,B,C\}𝔼S‖𝒉Sr‖22,\\displaystyle=\\sum\_\{\\emptyset\\neq S\\subseteq\\\{A,B,C\\\}\}\\mathbb\{E\}\_\{S\}\\left\\\|\\bm\{h\}\_\{S\}^\{r\}\\right\\\|\_\{2\}^\{2\},where the second equality follows because𝒉Sr\\bm\{h\}\_\{S\}^\{r\}depends only on the variables contained inSS\. ∎ ### B\.1Mode fraction and cluster analysis In Section[3](https://arxiv.org/html/2609.35970#S3)of the main text, we defined mode fraction using the DFT\. It can instead be defined directly without invoking the DFT\. Here, we show the equivalence\. Figure[1](https://arxiv.org/html/2609.35970#S1.F1)suggests that the interaction ensembles form clusters indexed byγ\\gammaorDD\. ###### Definition 2\(Cluster vectors\)\. ForHABrH\_\{AB\}^\{r\}, we define the centroid of the cluster indexed byγ\\gammaas 𝒉¯γr:=𝔼\[𝒉ABr\|B−A=γ\]\.\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}:=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{AB\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\.\(21\)Similarly, forHABCrH\_\{ABC\}^\{r\}, we define the centroid of the cluster indexed byDDas 𝒉¯Dr:=𝔼\[𝒉ABCr\|C\+B−A=D\]\.\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}:=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{ABC\}^\{r\}\\,\\middle\|\\,C\+B\-A=D\\right\]\.\(22\) Note that these are the same vectors used in the steering experiments in Section[5](https://arxiv.org/html/2609.35970#S5)\. We now define the corresponding projection operators for each ensemble\. ###### Definition 3\(Projection operator\)\. LetPγP\_\{\\gamma\}be the projection operator that operates onHABrH\_\{AB\}^\{r\}to give the corresponding cluster vector: PγHABr:=𝒉¯γrP\_\{\\gamma\}H\_\{AB\}^\{r\}:=\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}\(23\)Similarly, we definePDP\_\{D\}as PDHABCr:=𝒉¯DrP\_\{D\}H\_\{ABC\}^\{r\}:=\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}\(24\) Thus,PγHABrP\_\{\\gamma\}H\_\{AB\}^\{r\}is constant over pairs with the same value ofγ\\gamma, whilePDHABCrP\_\{D\}H\_\{ABC\}^\{r\}is constant over triples with the same value ofDD\. Now we define cluster fraction using these projection operators and show that it is equivalent to the definition of mode fraction in Section[3](https://arxiv.org/html/2609.35970#S3)\. ###### Definition 4\(Cluster fraction\)\. Letc\(Hr\)c\(H^\{r\}\)be the cluster fraction defined by c\(HABr\):=𝔼γ‖PγHABr‖22𝔼AB‖𝒉ABr‖22=𝔼γ‖𝒉¯γr‖22𝔼AB‖𝒉ABr‖22\\displaystyle c\(H\_\{AB\}^\{r\}\):=\\frac\{\\mathbb\{E\}\_\{\\gamma\}\\left\\\|P\_\{\\gamma\}H\_\{AB\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\{\\mathbb\{E\}\_\{AB\}\\left\\\|\\bm\{h\}\_\{AB\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}=\\frac\{\\mathbb\{E\}\_\{\\gamma\}\\left\\\|\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\{\\mathbb\{E\}\_\{AB\}\\left\\\|\\bm\{h\}\_\{AB\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\(25\)and c\(HABCr\):=𝔼D‖PDHABCr‖22𝔼ABC‖𝒉ABCr‖22=𝔼D‖𝒉¯Dr‖22𝔼ABC‖𝒉ABCr‖22\\displaystyle c\(H\_\{ABC\}^\{r\}\):=\\frac\{\\mathbb\{E\}\_\{D\}\\left\\\|P\_\{D\}H\_\{ABC\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\{\\mathbb\{E\}\_\{ABC\}\\left\\\|\\bm\{h\}\_\{ABC\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}=\\frac\{\\mathbb\{E\}\_\{D\}\\left\\\|\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\{\\mathbb\{E\}\_\{ABC\}\\left\\\|\\bm\{h\}\_\{ABC\}^\{r\}\\right\\\|\_\{2\}^\{2\}\}\(26\) ###### Proposition 3\(Equivalence between cluster fraction and mode fraction\)\. Ifc\(Hr\)c\(H^\{r\}\)is the cluster fraction as in Definition[4](https://arxiv.org/html/2609.35970#Thmdefinition4)andm\(Hr\)m\(H^\{r\}\)is the mode fraction as defined in Section[3](https://arxiv.org/html/2609.35970#S3), then c\(HSr\)=m\(HSr\),c\(H\_\{S\}^\{r\}\)=m\(H\_\{S\}^\{r\}\),\(27\)forS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}\. ###### Proof\. We suppress the superscriptrrfor readability\. First considerHABH\_\{AB\}\. Let 𝒑AB:=PγHAB=𝒉¯γ\.\\bm\{p\}\_\{AB\}:=P\_\{\\gamma\}H\_\{AB\}=\\bar\{\\bm\{h\}\}\_\{\\gamma\}\.Here,γ=B−A\\gamma=B\-A, and henceB=A\+γB=A\+\\gamma\. The discrete Fourier transform \(DFT\) of𝒑\\bm\{p\}is 𝒑^kAkB\\displaystyle\\widehat\{\\bm\{p\}\}\_\{k\_\{A\}k\_\{B\}\}=∑A,B𝒑ABe−2πi\(kAA\+kBB\)/N\\displaystyle=\\sum\_\{A,B\}\\bm\{p\}\_\{AB\}e^\{\-2\\pi i\(k\_\{A\}A\+k\_\{B\}B\)/N\}=∑A,γ𝒉¯γe−2πi\[\(kA\+kB\)A\+kBγ\]/N\\displaystyle=\\sum\_\{A,\\gamma\}\\bar\{\\bm\{h\}\}\_\{\\gamma\}e^\{\-2\\pi i\[\(k\_\{A\}\+k\_\{B\}\)A\+k\_\{B\}\\gamma\]/N\}=\(∑Ae−2πi\(kA\+kB\)A/N\)\(∑γ𝒉¯γe−2πikBγ/N\)\.\\displaystyle=\\left\(\\sum\_\{A\}e^\{\-2\\pi i\(k\_\{A\}\+k\_\{B\}\)A/N\}\\right\)\\left\(\\sum\_\{\\gamma\}\\bar\{\\bm\{h\}\}\_\{\\gamma\}e^\{\-2\\pi ik\_\{B\}\\gamma/N\}\\right\)\.Here,NNis the length of the cycle;N=12N=12for months and hours, whileN=7N=7for weekdays and musical notes\. The sum overAAvanishes unlesskA\+kB=0k\_\{A\}\+k\_\{B\}=0\. Therefore, 𝒑^kAkB=𝟎wheneverkA≠−kB\.\\widehat\{\\bm\{p\}\}\_\{k\_\{A\}k\_\{B\}\}=\\bm\{0\}\\qquad\\text\{whenever \}k\_\{A\}\\neq\-k\_\{B\}\. For frequencies\(kA,kB\)=\(−k,k\)\(k\_\{A\},k\_\{B\}\)=\(\-k,k\), 𝒑^−k,k\\displaystyle\\widehat\{\\bm\{p\}\}\_\{\-k,k\}=N∑γ𝒉¯γe−2πikγ/N\\displaystyle=N\\sum\_\{\\gamma\}\\bar\{\\bm\{h\}\}\_\{\\gamma\}e^\{\-2\\pi ik\\gamma/N\}\(28\)=∑A,γ𝒉A,A\+γe−2πikγ/N\\displaystyle=\\sum\_\{A,\\gamma\}\\bm\{h\}\_\{A,A\+\\gamma\}e^\{\-2\\pi ik\\gamma/N\}\(29\)=∑A,B𝒉ABe−2πik\(B−A\)/N\\displaystyle=\\sum\_\{A,B\}\\bm\{h\}\_\{AB\}e^\{\-2\\pi ik\(B\-A\)/N\}\(30\)=𝒉^−k,k,\\displaystyle=\\widehat\{\\bm\{h\}\}\_\{\-k,k\},\(31\)where the second equality uses 𝒉¯γ=1N∑A𝒉A,A\+γ\.\\bar\{\\bm\{h\}\}\_\{\\gamma\}=\\frac\{1\}\{N\}\\sum\_\{A\}\\bm\{h\}\_\{A,A\+\\gamma\}\.Thus, the Fourier transform ofPγHABP\_\{\\gamma\}H\_\{AB\}agrees with that ofHABH\_\{AB\}on the modeskA=−kBk\_\{A\}=\-k\_\{B\}and vanishes on all other modes\. Considerc\(HAB\)c\(H\_\{AB\}\): c\(HAB\)\\displaystyle c\(H\_\{AB\}\)=𝔼AB‖𝒑AB‖22𝔼AB‖𝒉AB‖22\\displaystyle=\\frac\{\\mathbb\{E\}\_\{AB\}\\left\\\|\\bm\{p\}\_\{AB\}\\right\\\|\_\{2\}^\{2\}\}\{\\mathbb\{E\}\_\{AB\}\\left\\\|\\bm\{h\}\_\{AB\}\\right\\\|\_\{2\}^\{2\}\}=∑kA=−kB‖𝒑^kAkB‖22∑kA,kB‖𝒉^kAkB‖22\\displaystyle=\\frac\{\\sum\_\{k\_\{A\}=\-k\_\{B\}\}\\left\\\|\\widehat\{\\bm\{p\}\}\_\{k\_\{A\}k\_\{B\}\}\\right\\\|\_\{2\}^\{2\}\}\{\\sum\_\{k\_\{A\},k\_\{B\}\}\\left\\\|\\widehat\{\\bm\{h\}\}\_\{k\_\{A\}k\_\{B\}\}\\right\\\|\_\{2\}^\{2\}\}=∑kA=−kB‖𝒉^kAkB‖22∑kA,kB‖𝒉^kAkB‖22\\displaystyle=\\frac\{\\sum\_\{k\_\{A\}=\-k\_\{B\}\}\\left\\\|\\widehat\{\\bm\{h\}\}\_\{k\_\{A\}k\_\{B\}\}\\right\\\|\_\{2\}^\{2\}\}\{\\sum\_\{k\_\{A\},k\_\{B\}\}\\left\\\|\\widehat\{\\bm\{h\}\}\_\{k\_\{A\}k\_\{B\}\}\\right\\\|\_\{2\}^\{2\}\}=m\(HAB\),\\displaystyle=m\(H\_\{AB\}\),where we get the second equality using Parseval’s identity and the third equality using Equation[31](https://arxiv.org/html/2609.35970#A2.E31)\. A similar proof holds forc\(HABC\)=m\(HABC\)c\(H\_\{ABC\}\)=m\(H\_\{ABC\}\)\. ∎ ### B\.2Steering vector Finally, we show the correspondence between the choice of steering vector as in Section[5](https://arxiv.org/html/2609.35970#S5)and the conditional mean of the activation ensemble\. ###### Proposition 4\. If𝐡¯γr\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}is the vector according to Definition[2](https://arxiv.org/html/2609.35970#Thmdefinition2), then 𝒉¯γr=𝔼ABC\[ϕr\(A,B,C\)\|B−A=γ\]−𝝁r\.\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}=\\mathbb\{E\}\_\{ABC\}\\left\[\\bm\{\\phi\}^\{r\}\(A,B,C\)\\;\|\\;B\-A=\\gamma\\right\]\-\\bm\{\\mu\}^\{r\}\.\(32\)Similarly, 𝒉¯Dr=𝔼ABC\[ϕr\(A,B,C\)\|C\+B−A=D\]−𝝁r\.\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}=\\mathbb\{E\}\_\{ABC\}\\left\[\\bm\{\\phi\}^\{r\}\(A,B,C\)\\;\|\\;C\+B\-A=D\\right\]\-\\bm\{\\mu\}^\{r\}\.\(33\) ###### Proof\. To see this, start with the conditional average overϕr\\bm\{\\phi\}^\{r\}: 𝔼\[ϕr\|B−A=γ\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=𝝁r\+𝔼\[𝒉Ar\|B−A=γ\]\+𝔼\[𝒉Br\|B−A=γ\]\\displaystyle=\\bm\{\\mu\}^\{r\}\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{A\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{B\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+𝔼\[𝒉Cr\|B−A=γ\]\+𝔼\[𝒉ABr\|B−A=γ\]\\displaystyle\\quad\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{C\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{AB\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+𝔼\[𝒉BCr\|B−A=γ\]\+𝔼\[𝒉CAr\|B−A=γ\]\\displaystyle\\quad\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{BC\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{CA\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\+𝔼\[𝒉ABCr\|B−A=γ\]\.\\displaystyle\\quad\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{ABC\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\. Using Lemma[1](https://arxiv.org/html/2609.35970#Thmlemma1), we get 𝔼\[𝒉Ar\|B−A=γ\]=𝔼\[𝒉Br\|B−A=γ\]=𝔼\[𝒉Cr\|B−A=γ\]=𝟎,\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{A\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{B\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{C\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\bm\{0\},and 𝔼\[𝒉BCr\|B−A=γ\]=𝔼\[𝒉CAr\|B−A=γ\]=𝔼\[𝒉ABCr\|B−A=γ\]=𝟎\.\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{BC\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{CA\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{ABC\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\bm\{0\}\. Hence, 𝔼\[ϕr\|B−A=γ\]=𝝁r\+𝔼\[𝒉ABr\|B−A=γ\]\.\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]=\\bm\{\\mu\}^\{r\}\+\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{AB\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\.By definition, 𝒉¯γr=𝔼\[𝒉ABr\|B−A=γ\],\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}=\\mathbb\{E\}\\\!\\left\[\\bm\{h\}\_\{AB\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\],and therefore 𝒉¯γr=𝔼\[ϕr\|B−A=γ\]−𝝁r\.\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}=\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\]\-\\bm\{\\mu\}^\{r\}\.A similar proof holds for𝒉¯Dr\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}\. ∎ Therefore, we get the following identities using Proposition[4](https://arxiv.org/html/2609.35970#Thmproposition4): 𝒉¯γ\+δr−𝒉¯γr=𝔼\[ϕr\|B−A=γ\+δ\]−𝔼\[ϕr\|B−A=γ\],\\displaystyle\\bar\{\\bm\{h\}\}\_\{\\gamma\+\\delta\}^\{r\}\-\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}=\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\+\\delta\\right\]\-\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,B\-A=\\gamma\\right\],\(34\)𝒉¯D\+δr−𝒉¯Dr=𝔼\[ϕr\|C\+B−A=D\+δ\]−𝔼\[ϕr\|C\+B−A=D\]\.\\displaystyle\\bar\{\\bm\{h\}\}\_\{D\+\\delta\}^\{r\}\-\\bar\{\\bm\{h\}\}\_\{D\}^\{r\}=\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,C\+B\-A=D\+\\delta\\right\]\-\\mathbb\{E\}\\\!\\left\[\\bm\{\\phi\}^\{r\}\\,\\middle\|\\,C\+B\-A=D\\right\]\.\(35\)Thus, the steering using𝒉¯γ\+δr−𝒉¯γr\\bar\{\\bm\{h\}\}\_\{\\gamma\+\\delta\}^\{r\}\-\\bar\{\\bm\{h\}\}\_\{\\gamma\}^\{r\}is equivalent to the difference\-of\-means steering\. ### B\.3Coincident variables We always compute the interaction ensembles using the full set of prompts, including cases withA=BA=B,B=CB=C, orC=AC=A\. Table[5](https://arxiv.org/html/2609.35970#A2.T5)reports the mean norm of the interaction vectors grouped byγ=B−A\\gamma=B\-A,γ′=C−B\\gamma^\{\\prime\}=C\-B, andγ′′=A−C\\gamma^\{\\prime\\prime\}=A\-Cfor a representative ensemble using the months prompt template at layer 15 of Llama\-3\.1\-8B\. The vectors corresponding toγ=0,γ′=0,γ′′=0\\gamma=0,\\;\\gamma^\{\\prime\}=0,\\;\\gamma^\{\\prime\\prime\}=0have substantially larger norms than the rest of the ensemble and can therefore bias quantitative results toward coincident\-variable cases\. We exclude such prompts when reporting the results in the main text\. Unless stated otherwise, we repeat our analyses without this exclusion in the following appendices\. This also serves as a sanity check that our conclusions are not driven by the coincident\-variable cases\. Table 5:Mean norm of each interaction vector categorized byγ=B−A\\gamma=B\-A,γ′=C−B\\gamma^\{\\prime\}=C\-B, andγ′′=A−C\\gamma^\{\\prime\\prime\}=A\-C, respectively, for a single replicate of the months prompt template at layer 15 of Llama\-3\.1\-8B\. The interaction vectors associated with zero difference have substantially larger norms than the remaining classes\. This motivates the exclusion of coincident\-variable prompts from the quantitative analyses in the main text\. ## Appendix CMode fraction Figure 7:Mode fraction of the interaction ensembles across models and cyclic domains\. Columns correspond to models and rows to domains; shaded bands show the central 80% interval across 10 replicates\. The vertical gray band marks the transition in causal relevance fromHABrH\_\{AB\}^\{r\}toHABCrH\_\{ABC\}^\{r\}identified by the ablation and replacement experiments\. Across all models and prompt templates, the pairwise interaction ensembles are strongly organized by their corresponding pairwise differences beforeHABCrH\_\{ABC\}^\{r\}becomes organized byDD\. The transition inHABCrH\_\{ABC\}^\{r\}is more distributed across depth in the Gemma models\.The main text reports the layerwise organization of the interaction ensembles for Llama\-3\.1\-8B on the months domain\. Here, we test whether the same pattern generalizes across models and cyclic domains\. Recall the definition of mode fraction forHABrH\_\{AB\}^\{r\}as defined in the main text: m\(HABr\):=∑kA=−kB‖𝒉^kAkBr‖22∑kA,kB‖𝒉^kAkBr‖22\.m\(H\_\{AB\}^\{r\}\):=\\frac\{\\sum\_\{k\_\{A\}=\-k\_\{B\}\}\\left\\lVert\\widehat\{\\bm\{h\}\}^\{r\}\_\{k\_\{A\}k\_\{B\}\}\\right\\rVert\_\{2\}^\{2\}\}\{\\sum\_\{k\_\{A\},k\_\{B\}\}\\left\\lVert\\widehat\{\\bm\{h\}\}^\{r\}\_\{k\_\{A\}k\_\{B\}\}\\right\\rVert\_\{2\}^\{2\}\}\.\(36\) Similarly, we define the corresponding mode fractions forHBCrH\_\{BC\}^\{r\}andHCArH\_\{CA\}^\{r\}using modes satisfyingkB\+kC=0k\_\{B\}\+k\_\{C\}=0andkC\+kA=0k\_\{C\}\+k\_\{A\}=0respectively\. ForHABCrH\_\{ABC\}^\{r\}, the mode fraction is defined using modes\(kA,kB,kC\)=\(−k,k,k\)\(k\_\{A\},k\_\{B\},k\_\{C\}\)=\(\-k,k,k\)\. Figure[7](https://arxiv.org/html/2609.35970#A3.F7)shows the mode fractions for all interaction ensembles using the prompts and models as detailed in Table[1](https://arxiv.org/html/2609.35970#A1.T1)\. Across all models and prompt templates, the second\-order interaction ensembles become strongly organized by their corresponding pairwise differences before the three\-way interaction becomes organized byDD\. The increase inm\(HABCr\)m\(H\_\{ABC\}^\{r\}\)occurs near the same relative depth within a particular model across domains\. The transition is sharp for most models but is distributed across several layers in the Gemma models\. Notice that we report Llama\-3\.1\-8B on the months domain again: we include prompts with coincident variables and we see that the mode fraction is inflated as a result of their inclusion\. ## Appendix DEnergy Fraction Figure 8:Energy fraction of the interaction ensembles across models and cyclic domains\. Columns correspond to models and rows to domains; shaded bands show the central 80% interval across 10 replicates\. The vertical gray band marks the transition in causal relevance identified by the ablation and replacement experiments\. Across all models and prompt templates, the energy fraction ofHABrH\_\{AB\}^\{r\}decreases near the same depth at which the energy fraction ofHABCrH\_\{ABC\}^\{r\}increases\.Recall Corollary[1](https://arxiv.org/html/2609.35970#Thmcorollary1)from Appendix[B](https://arxiv.org/html/2609.35970#A2): 𝔼A,B,C‖ϕr\(A,B,C\)−𝝁r‖22=∑S𝔼S‖𝒉Sr‖22,\\mathbb\{E\}\_\{A,B,C\}\\left\\\|\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{\\mu\}^\{r\}\\right\\\|\_\{2\}^\{2\}=\\sum\_\{S\}\\mathbb\{E\}\_\{S\}\\left\\\|\\bm\{h\}\_\{S\}^\{r\}\\right\\\|^\{2\}\_\{2\},\(37\)whereSSranges over non\-empty subsets of\{A,B,C\}\\\{A,B,C\\\}\. This motivates us to define theenergy fractionof each ensemble as e\(HSr\)=𝔼S\(‖𝒉Sr‖22\)𝔼ABC\(‖ϕr\(A,B,C\)−𝝁r‖22\),with∑Se\(HSr\)=1\.e\(H\_\{S\}^\{r\}\)=\\frac\{\\mathbb\{E\}\_\{S\}\\Big\(\\left\\lVert\\bm\{h\}^\{r\}\_\{S\}\\right\\rVert\_\{2\}^\{2\}\\Big\)\}\{\\mathbb\{E\}\_\{ABC\}\\Big\(\\left\\lVert\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{\\mu\}^\{r\}\\right\\rVert\_\{2\}^\{2\}\\Big\)\},\\qquad\\text\{with\}\\qquad\\sum\_\{S\}e\(H\_\{S\}^\{r\}\)=1\.\(38\)We next test whether the layerwise redistribution of energy fraction observed in the main text generalizes across models and domains\. Figure[8](https://arxiv.org/html/2609.35970#A4.F8)shows the energy fraction for each interaction ensemble\. Across all models and prompt templates, the energy fraction ofHABrH\_\{AB\}^\{r\}decreases near the same depth at which the energy fraction ofHABCrH\_\{ABC\}^\{r\}increases\. As with mode fraction, this transition is spread over several layers in the Gemma models\. ## Appendix EControls for Llama\-3\.1\-8B The strong organization of the pairwise interaction ensembles by their corresponding pairwise differences does not by itself imply that this structure is task\-specific\. We therefore test whether similar organization arises when the cyclic tokens co\-occur without the relational computation used in the main task\. We design three control prompts as listed in Table[6](https://arxiv.org/html/2609.35970#A5.T6)\. For the*non\-cyclic*control, we construct a template similar to the Llama\-3\.1\-8B months template but useA,B,C∈A,B,C\\in\{berry, apple, orange, banana, peach, pear, fig, mango, lemon, lime, cherry, plum\} instead of the usual calendar months\. The*off\-by\-1*control retainsA,B,CA,B,Cas calendar months, but the answer depends only onCC;AAandBBtherefore co\-occur in the prompt without contributing to the answer\. The*list*control also retains calendar months but removes the arithmetic relation entirely;A,B,CA,B,Care presented as items in a list and the model is asked to returnBB\. Figure 9:Mode and energy fractions for three control prompt templates using Llama\-3\.1\-8B\. The*non\-cyclic*control replaces months with an unordered vocabulary, while the*off\-by\-1*and*list*controls retain calendar months but remove the relational computation of the main task\. In the latter two controls, the pairwise interaction ensembles remain strongly organized by their corresponding pairwise differences despite having a small energy fraction\. In contrast,HABCrH\_\{ABC\}^\{r\}remains negligible in both mode and energy fraction across all controls\. All results reported here exclude vectors corresponding to the coincident\-variable cases\.In Figure[9](https://arxiv.org/html/2609.35970#A5.F9), we have computed the mode and energy fractions for the control prompt templates\. In all three controls, the mode and energy fractions ofHABCrH\_\{ABC\}^\{r\}remain close to zero\. In contrast, when calendar months co\-occur in the*off\-by\-1*and*list*controls,HABrH\_\{AB\}^\{r\},HBCrH\_\{BC\}^\{r\}, andHCArH\_\{CA\}^\{r\}can still have large mode fractions, but retain a small energy fraction\. This suggests that pairwise\-difference geometry can arise from co\-occurrence of cyclic tokens alone, whereas theDD\-organized interaction ensembleHABCrH\_\{ABC\}^\{r\}requires the task structure\. Table 6:Control prompt templates for Llama\-3\.1\-8B\. ## Appendix FCausal handoff Section[4](https://arxiv.org/html/2609.35970#S4)shows thatHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}are causally relevant at different depths, but this does not establish whether the laterHABCrH\_\{ABC\}^\{r\}representation depends on the earlierHABrH\_\{AB\}^\{r\}representation\. To test this, we ablateHABrH\_\{AB\}^\{r\}at layer 15 of Llama\-3\.1\-8B on the months domain, allow the model to run normally thereafter, and recompute the ANOVA decomposition at every subsequent layer\. We then compare the mode and energy fractions of the resultingHABCrH\_\{ABC\}^\{r\}with those of the unmodified model\. Figure[10](https://arxiv.org/html/2609.35970#A6.F10)shows the comparison of mode and energy fractions between the ablated and the unmodified ensemble forHABCrH\_\{ABC\}^\{r\}\. In the unmodified case, both the mode fraction and energy fraction ofHABCrH\_\{ABC\}^\{r\}rise sharply around layer 18\. After ablatingHABrH\_\{AB\}^\{r\}at layer 15, neither transition occurs at that depth\. The mode fraction recovers slightly in later layers, but the energy fraction remains substantially reduced\. Thus, the emergence of the laterHABCrH\_\{ABC\}^\{r\}representation depends on the earlierHABrH\_\{AB\}^\{r\}representation\. Figure 10:DownstreamHABCrH\_\{ABC\}^\{r\}after ablatingHABrH\_\{AB\}^\{r\}at layer 15\. Solid curves show the mode and energy fractions of the recomputedHABCrH\_\{ABC\}^\{r\}after the intervention; dashed curves show the corresponding quantities in the unmodified model\. The vertical gray band marks layer 18, whereHABCrH\_\{ABC\}^\{r\}normally becomes organized byDDand increases sharply in energy fraction\.\(a\)After ablatingHABrH\_\{AB\}^\{r\}, the mode fraction does not undergo its normal increase at layer 18, although some organization recovers later but never regains the unmodified baseline\.\(b\)The energy fraction likewise fails to increase at layer 18 and remains substantially below that of the unmodified ensemble\. These results show that the normal emergence ofHABCrH\_\{ABC\}^\{r\}depends on the earlierHABrH\_\{AB\}^\{r\}representation\. ## Appendix GEnsemble ablation We next test whether the layerwise ablation pattern from the main text generalizes across models and cyclic concepts\. We conduct the ensemble ablation experiments for all prompts and models in Table[1](https://arxiv.org/html/2609.35970#A1.T1)\. Recall that we intervene on the last token position at a particular layer using ϕ~r\(A,B,C\)=ϕr\(A,B,C\)−𝒉Sr,\\widetilde\{\\bm\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\phi\}^\{r\}\(A,B,C\)\-\\bm\{h\}^\{r\}\_\{S\},\(39\)for allA,B,C∈𝒳A,B,C\\in\\mathcal\{X\}andS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}\. Figures[12](https://arxiv.org/html/2609.35970#A10.F12)and[13](https://arxiv.org/html/2609.35970#A10.F13)report the results of the ensemble ablation experiments using top\-1 and top\-3 accuracy as metrics, respectively\. We observe that across every model and domain, the layerwise behavior of interaction ensembles under ablation remains the same\. AblatingHABrH\_\{AB\}^\{r\}always affects the accuracy at an earlier network depth than ablatingHABCrH\_\{ABC\}^\{r\}\. However, ablatingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}has little to no effect on performance\. AblatingHABCrH\_\{ABC\}^\{r\}at later layers decreases the performance dramatically, with performance dropping to near chance when it is ablated near the final layers\. ## Appendix HEnsemble replacement We repeat the ensemble replacement experiments across all models and prompt templates using the following intervention ϕ~r\(A,B,C\)=𝝁r\+𝒉Sr,\\widetilde\{\\bm\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\mu\}^\{r\}\+\\bm\{h\}^\{r\}\_\{S\},\(40\)for allA,B,C∈𝒳A,B,C\\in\\mathcal\{X\}andS∈\{AB,BC,CA,ABC\}S\\in\\\{AB,BC,CA,ABC\\\}\. Figures[14](https://arxiv.org/html/2609.35970#A10.F14)and[15](https://arxiv.org/html/2609.35970#A10.F15)show the corresponding results using top\-1 and top\-3 accuracy respectively\. RetainingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}together with the mean performs similarly to theμr\\mu^\{r\}\-only baseline\. RetainingHABrH\_\{AB\}^\{r\}preserves performance in intermediate layers but becomes insufficient later in depth, while retainingHABCrH\_\{ABC\}^\{r\}becomes effective at later layers and can exceed the unmodified baseline\. Together with Appendix[G](https://arxiv.org/html/2609.35970#A7), these results show that the transition fromHABrH\_\{AB\}^\{r\}toHABCrH\_\{ABC\}^\{r\}is reproduced by both ablation and replacement interventions across models and prompt templates\. ## Appendix ISteering We report steering results across all models and prompt templates using the averaged vectors introduced in Section[5](https://arxiv.org/html/2609.35970#S5)\. ForS∈\{γ,D\}S\\in\\\{\\gamma,D\\\}, we intervene according to ϕ~r\(A,B,C\)=ϕr\(A,B,C\)\+α\(𝒉¯S\+δr−𝒉¯Sr\)\.\\bm\{\\widetilde\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\phi\}^\{r\}\(A,B,C\)\+\\alpha\\left\(\\bm\{\\bar\{h\}\}^\{r\}\_\{S\+\\delta\}\-\\bm\{\\bar\{h\}\}^\{r\}\_\{S\}\\right\)\.\(41\) To compare steering effects across models, domains, and layers, we report thenormalizedmedianΔm3\\mathrm\{normalized\\;median\\;\}\\Delta m\_\{3\}\. We normalize by Δm3max=maxS,δ,r,ℓ\|Δm3r,ℓ\(𝒉¯S\+δr−𝒉¯Sr\)\|\.\\Delta m\_\{3\}^\{\\mathrm\{max\}\}=\\mathop\{\\mathrm\{max\}\}\_\{S,\\delta,r,\\ell\}\\left\|\\Delta m\_\{3\}^\{r,\\ell\}\(\\bm\{\\bar\{h\}\}^\{r\}\_\{S\+\\delta\}\-\\bm\{\\bar\{h\}\}^\{r\}\_\{S\}\)\\right\|\.\(42\)Figure[16](https://arxiv.org/html/2609.35970#A10.F16)shows the resulting steering profiles atα=1\\alpha=1\. Steering with𝒉¯γr\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\}is effective in intermediate layers, whereas steering with𝒉¯Dr\\bm\{\\bar\{h\}\}\_\{D\}^\{r\}becomes effective at later layers\. These regions coincide with the layers whereHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}are found to be causally relevant in Appendices[G](https://arxiv.org/html/2609.35970#A7)and[H](https://arxiv.org/html/2609.35970#A8)\. ### I\.1Sensitivity to steering strength The experiments in the main text useα=1\\alpha=1\. To test whether the observed steering effects depend strongly on this choice, Figure[11](https://arxiv.org/html/2609.35970#A9.F11)shows themedianΔm3\\mathrm\{median\\;\}\\Delta m\_\{3\}as a function ofα\\alphaat representative layers for𝒉¯γr\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\}and𝒉¯Dr\\bm\{\\bar\{h\}\}\_\{D\}^\{r\}\. In both cases, the steering effect varies smoothly withα\\alpha\. Figure 11:MedianΔm3\\Delta m\_\{3\}as a function of steering strengthα\\alphafor the averaged steering vectors\.\(a\)Steering with𝒉¯γr\\bm\{\\bar\{h\}\}\_\{\\gamma\}^\{r\}at layer 15\.\(b\)Steering with𝒉¯Dr\\bm\{\\bar\{h\}\}\_\{D\}^\{r\}at layer 22\. Both panels use the months prompt template with Llama\-3\.1\-8B\. ### I\.2Symmetrized pointwise steering The averaged vectors above depend only on the relation labelγ\\gammaorDD\. As a complementary intervention, we construct prompt\-specific counterfactual interaction vectors by averaging over the coordinate changes that produce the same shift in the corresponding relation\. ForHABrH\_\{AB\}^\{r\}, bothB→B\+δB\\to B\+\\deltaandA→A−δA\\to A\-\\deltachangeγ→γ\+δ\\gamma\\to\\gamma\+\\deltaand thereforeD→D\+δD\\to D\+\\delta\. We define 𝒉~AB\|δr=12\(𝒉A,B\+δr\+𝒉A−δ,Br\)\.\\widetilde\{\\bm\{h\}\}^\{r\}\_\{AB\\mid\\delta\}=\\frac\{1\}\{2\}\\left\(\\bm\{h\}^\{r\}\_\{A,B\+\\delta\}\+\\bm\{h\}^\{r\}\_\{A\-\\delta,B\}\\right\)\.\(43\)Similarly, each ofA→A−δA\\to A\-\\delta,B→B\+δB\\to B\+\\delta, andC→C\+δC\\to C\+\\deltachangesDDtoD\+δD\+\\delta, so we define 𝒉~ABC\|δr=13\(𝒉A−δ,B,Cr\+𝒉A,B\+δ,Cr\+𝒉A,B,C\+δr\)\.\\widetilde\{\\bm\{h\}\}^\{r\}\_\{ABC\\mid\\delta\}=\\frac\{1\}\{3\}\\left\(\\bm\{h\}^\{r\}\_\{A\-\\delta,B,C\}\+\\bm\{h\}^\{r\}\_\{A,B\+\\delta,C\}\+\\bm\{h\}^\{r\}\_\{A,B,C\+\\delta\}\\right\)\.\(44\)Whenδ=0\\delta=0, these expressions reduce to the original interaction vectors\. We steer using ϕ~r\(A,B,C\)=ϕr\(A,B,C\)\+α\(𝒉~S\|δr−𝒉Sr\),\\bm\{\\widetilde\{\\phi\}\}^\{r\}\(A,B,C\)=\\bm\{\\phi\}^\{r\}\(A,B,C\)\+\\alpha\\left\(\\widetilde\{\\bm\{h\}\}^\{r\}\_\{S\\mid\\delta\}\-\\bm\{h\}^\{r\}\_\{S\}\\right\),\(45\)whereS∈\{AB,ABC\}S\\in\\\{AB,ABC\\\}\. Figure[17](https://arxiv.org/html/2609.35970#A10.F17)shows thenormalizedmedianΔm3\\mathrm\{normalized\\;median\\;\}\\Delta m\_\{3\}obtained using this intervention\. The same layerwise pattern appears as with the averaged steering vectors: steering using𝒉~AB\|δr\\widetilde\{\\bm\{h\}\}^\{r\}\_\{AB\\mid\\delta\}is effective in intermediate layers, whereas steering using𝒉~ABC\|δr\\widetilde\{\\bm\{h\}\}^\{r\}\_\{ABC\\mid\\delta\}becomes effective at later layers\. ## Appendix JCross\-domain transfer We report the cross\-domain transfer experiments across models and domains\. Figures[18](https://arxiv.org/html/2609.35970#A10.F18)and[19](https://arxiv.org/html/2609.35970#A10.F19)show top\-1 and top\-3 accuracy after transplantingHABrH\_\{AB\}^\{r\}from one domain into another following ablation of the corresponding domain\-1 interaction\. Figures[20](https://arxiv.org/html/2609.35970#A10.F20)and[21](https://arxiv.org/html/2609.35970#A10.F21)show the corresponding replacement experiments, in which the domain\-1 mean is retained together with the domain\-2 interaction\. Figure[22](https://arxiv.org/html/2609.35970#A10.F22)shows cross\-domain steering using𝒉¯γ\\bar\{\\bm\{h\}\}\_\{\\gamma\}, includingδ=0\\delta=0as a control\. Direct transplantation ofHABrH\_\{AB\}^\{r\}generally recovers part of the performance lost under ablation, although the degree of recovery depends on the source and target domains\. In particular, transfer between the months and hours domains is asymmetric for the Llama models: months→\\tohours decreases performance relative to the unmodified model, whereas hours→\\tomonths improves the performance \(Figures[18](https://arxiv.org/html/2609.35970#A10.F18)and[19](https://arxiv.org/html/2609.35970#A10.F19)\)\. In contrast, direct transfer ofHABCrH\_\{ABC\}^\{r\}is substantially less consistent across domains \(Figures[23](https://arxiv.org/html/2609.35970#A10.F23)and[24](https://arxiv.org/html/2609.35970#A10.F24)\)\. ### J\.1Alignment of interaction subspaces The inconsistent direct transfer ofHABCrH\_\{ABC\}^\{r\}may arise because different domains use different residual\-stream subspaces for the same underlying interaction structure\. We therefore align the interaction subspaces before transplantation\. We perform the alignment in aq=N−1q=N\-1dimensional subspace and divide the replicates into a training set \(70%\), denoted byrtr\_\{t\}, and a held\-out test set \(30%\), denoted byrer\_\{e\}\. For domainc∈\{1,2\}c\\in\\\{1,2\\\}, let WAB\(c\)=\[𝒉ABrt,\(c\)\]∈ℝtN2×dmodel,W\_\{AB\}^\{\(c\)\}=\\left\[\\bm\{h\}\_\{AB\}^\{r\_\{t\},\(c\)\}\\right\]\\in\\mathbb\{R\}^\{tN^\{2\}\\times d\_\{\\mathrm\{model\}\}\},\(46\)and similarly, WABC\(c\)=\[𝒉ABCrt,\(c\)\]∈ℝtN3×dmodel\.W\_\{ABC\}^\{\(c\)\}=\\left\[\\bm\{h\}\_\{ABC\}^\{r\_\{t\},\(c\)\}\\right\]\\in\\mathbb\{R\}^\{tN^\{3\}\\times d\_\{\\mathrm\{model\}\}\}\.\(47\)BothWAB\(c\)W\_\{AB\}^\{\(c\)\}andWABC\(c\)W\_\{ABC\}^\{\(c\)\}have zero mean by Lemma[1](https://arxiv.org/html/2609.35970#Thmlemma1)\. We describe the alignment procedure forWABW\_\{AB\}; the same procedure is applied toWABCW\_\{ABC\}\. Consider transplanting the domain\-2 interaction into domain\-1\. For each domain, we first compute WAB\(c\)=U\(c\)Σ\(c\)\(V\(c\)\)T,W\_\{AB\}^\{\(c\)\}=U^\{\(c\)\}\\Sigma^\{\(c\)\}\\left\(V^\{\(c\)\}\\right\)^\{T\},\(48\)where the columns ofV\(c\)V^\{\(c\)\}are the right singular vectors\. LetVq\(c\)∈ℝdmodel×qV\_\{q\}^\{\(c\)\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\\times q\}contain the leadingq=N−1q=N\-1right singular vectors\. We project the interaction ensembles into these subspaces: ZAB\(c\)=WAB\(c\)Vq\(c\)∈ℝtN2×q\.Z\_\{AB\}^\{\(c\)\}=W\_\{AB\}^\{\(c\)\}V\_\{q\}^\{\(c\)\}\\in\\mathbb\{R\}^\{tN^\{2\}\\times q\}\.\(49\) We then align the domain\-2 coordinates to domain\-1 by solving the orthogonal Procrustes problem RAB∗=argminRTR=𝑰q‖ZAB\(2\)R−ZAB\(1\)‖F2,R\_\{AB\}^\{\*\}=\\underset\{R^\{T\}R=\{\\bm\{I\}\}\_\{q\}\}\{\\operatorname\{arg\\;min\}\}\\left\\lVert Z\_\{AB\}^\{\(2\)\}R\-Z\_\{AB\}^\{\(1\)\}\\right\\rVert\_\{F\}^\{2\},\(50\)whereRAB∗∈ℝq×qR\_\{AB\}^\{\*\}\\in\\mathbb\{R\}^\{q\\times q\}is the orthogonal transformation that best aligns the two coordinate systems\. Transforming a domain\-2 interaction vector into the domain\-1 residual\-stream basis then consists of projection onto the domain\-2 subspace, alignment in theqq\-dimensional coordinate system, and reconstruction in the domain\-1 basis\. We therefore define Q=Vq\(2\)RAB∗\(Vq\(1\)\)T∈ℝdmodel×dmodel\.Q=V\_\{q\}^\{\(2\)\}R\_\{AB\}^\{\*\}\\left\(V\_\{q\}^\{\(1\)\}\\right\)^\{T\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\\times d\_\{\\mathrm\{model\}\}\}\.\(51\) We evaluate the learned alignment only on held\-out replicates\. For the aligned ablation\-transplant experiment, we use ϕ~re,\(1\)\(A,B,C\)=ϕre,\(1\)\(A,B,C\)−𝒉ABre,\(1\)\+𝒉ABre,\(2\)Q\.\\widetilde\{\\bm\{\\phi\}\}^\{r\_\{e\},\(1\)\}\(A,B,C\)=\\bm\{\\phi\}^\{r\_\{e\},\(1\)\}\(A,B,C\)\-\\bm\{h\}\_\{AB\}^\{r\_\{e\},\(1\)\}\+\\bm\{h\}\_\{AB\}^\{r\_\{e\},\(2\)\}Q\.\(52\)For the corresponding aligned replacement experiment, we use ϕ~re,\(1\)\(A,B,C\)=𝝁re,\(1\)\+𝒉ABre,\(2\)Q\.\\widetilde\{\\bm\{\\phi\}\}^\{r\_\{e\},\(1\)\}\(A,B,C\)=\\bm\{\\mu\}^\{r\_\{e\},\(1\)\}\+\\bm\{h\}\_\{AB\}^\{r\_\{e\},\(2\)\}Q\.\(53\) Figures[25](https://arxiv.org/html/2609.35970#A10.F25)and[26](https://arxiv.org/html/2609.35970#A10.F26)show the aligned ablation\-transplant results for bothHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}, while Figures[27](https://arxiv.org/html/2609.35970#A10.F27)and[28](https://arxiv.org/html/2609.35970#A10.F28)show the corresponding replacement experiments\. Alignment improves cross\-domain transfer for both interaction ensembles and produces a particularly large improvement forHABCrH\_\{ABC\}^\{r\}, whose direct transfer is otherwise inconsistent across domains\. This suggests that part of the failure of directHABCrH\_\{ABC\}^\{r\}transplantation can be explained by differences in the residual\-stream subspaces used by different domains\. Figure 12:Top\-1 accuracy after ablating an interaction ensemble at the last\-token position, shown as a function of layer\. The vertical shaded bands mark the model\-specific intermediate layers where ablatingHABrH\_\{AB\}^\{r\}produces its largest drop in performance\. Across models and prompts, ablatingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}has little effect on accuracy, whereas ablatingHABrH\_\{AB\}^\{r\}orHABCrH\_\{ABC\}^\{r\}substantially degrades performance\. The effect of ablatingHABrH\_\{AB\}^\{r\}occurs earlier in depth than the effect of ablatingHABCrH\_\{ABC\}^\{r\}\.Figure 13:Top\-3 accuracy after ablating each interaction ensemble at the last\-token position, shown as a function of layer\. Across models and prompt templates, ablatingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}has little effect on performance, whereas ablatingHABrH\_\{AB\}^\{r\}orHABCrH\_\{ABC\}^\{r\}substantially decreases accuracy\. The effect of ablatingHABrH\_\{AB\}^\{r\}consistently occurs earlier in depth than the effect of ablatingHABCrH\_\{ABC\}^\{r\}\.Figure 14:Top\-1 accuracy after replacing the activation of the last token with only the mean𝝁r\\bm\{\\mu\}^\{r\}and an interaction ensemble, shown as a function of layer\. The vertical shaded bands mark the intermediate layers where ablatingHABrH\_\{AB\}^\{r\}produced the largest drop in performance, as shown in Figures[12](https://arxiv.org/html/2609.35970#A10.F12)and[13](https://arxiv.org/html/2609.35970#A10.F13)\. Replacing withHBCrH\_\{BC\}^\{r\},HCArH\_\{CA\}^\{r\}, orHABCrH\_\{ABC\}^\{r\}substantially degrades performance at the first vertical band, whereas replacing withHABrH\_\{AB\}^\{r\}does not affect performance until the second vertical band\. At later layers, retainingHABCrH\_\{ABC\}^\{r\}preserves performance and can exceed the unmodified baseline\.Figure 15:Top\-3 accuracy after replacing the last\-token activation with the ensemble mean𝝁r\\bm\{\\mu\}^\{r\}and one interaction term\. Across models and prompt templates, retainingHABrH\_\{AB\}^\{r\}preserves performance in intermediate layers but becomes insufficient later in depth, while retainingHABCrH\_\{ABC\}^\{r\}becomes effective at later layers and can outperform the unmodified baseline\. RetainingHBCrH\_\{BC\}^\{r\}orHCArH\_\{CA\}^\{r\}generally degrades the performance beginning from the first vertical band\.Figure 16:Heatmap of steering using𝒉¯γ\\bar\{\\bm\{h\}\}\_\{\\gamma\}and𝒉¯D\\bar\{\\bm\{h\}\}\_\{D\}across every model and prompt\. Color indicatesnormalizedmedianΔm3\\mathrm\{normalized\\;median\\;\}\\Delta m\_\{3\}\. Steering with𝒉¯γ\\bar\{\\bm\{h\}\}\_\{\\gamma\}is most effective in intermediate layers, with boldface layer labels marking the layers whereHABrH\_\{AB\}^\{r\}was found to be causally relevant\. At later layers, steering with𝒉¯D\\bar\{\\bm\{h\}\}\_\{D\}becomes effective\.Figure 17:Heatmap of steering using𝒉~AB\|δ\\widetilde\{\\bm\{h\}\}\_\{AB\|\\delta\}and𝒉~ABC\|δ\\widetilde\{\\bm\{h\}\}\_\{ABC\|\\delta\}across every model and prompt\. Color indicatesnormalizedmedianΔm3\\mathrm\{normalized\\;median\\;\}\\Delta m\_\{3\}\. As with the steering results in Figure[16](https://arxiv.org/html/2609.35970#A10.F16), steering using𝒉~AB\|δ\\widetilde\{\\bm\{h\}\}\_\{AB\|\\delta\}is effective in intermediate layers, while steering using𝒉~ABC\|δ\\widetilde\{\\bm\{h\}\}\_\{ABC\|\\delta\}becomes effective at later layers\.Figure 18:Top\-1 accuracy for cross\-domain ablation transplant ofHABrH\_\{AB\}^\{r\}\. Solid curves show performance after replacing an ablated domain\-1 interaction vector with the corresponding domain\-2 vector; dashed blue curves show ablation without transplantation and dashed black curves show the unmodified baseline\. Successful transfer is indicated by recovery from the dashed blue ablation curve toward the baseline\. Direct transplantation generally recovers performance, although the amount of recovery depends on the model and cross\-domain pair\.Figure 19:Top\-3 accuracy for direct cross\-domain ablation transplant ofHABrH\_\{AB\}^\{r\}\. As in Figure[18](https://arxiv.org/html/2609.35970#A10.F18), transplantation generally recovers part of the performance lost under ablation, with variation across model and cross\-domain pairs\.Figure 20:Top\-1 accuracy for cross\-domain replacement usingHABrH\_\{AB\}^\{r\}\. Solid curves retain the domain\-1 mean together with the domain\-2 interaction vector; dashed blue curves show the corresponding within\-domain replacement experiment\. Solid and dashed trajectories follow each other, which indicates that the domain\-2 interaction can substitute for the domain\-1 interaction in downstream computation\.Figure 21:Top\-3 accuracy for direct cross\-domain replacement usingHABrH\_\{AB\}^\{r\}\. Across all models and cross\-domain pairs, the transplanted domain\-2 interaction reproduces the layerwise behavior of the corresponding within\-domain replacement\.Figure 22:Cross\-domain steering using vectors extracted from domain\-2\. Color indicates normalized medianΔm3\\Delta m\_\{3\}\. Steering is strongest in the same intermediate layers whereHABrH\_\{AB\}^\{r\}is causally relevant\. Theδ=0\\delta=0case provides a non\-trivial control: replacing the domain\-1 relation vector with the corresponding domain\-2 vector should preserve the original answer if the relation representation transfers across domains\.Figure 23:Top\-3 accuracy for cross\-domain ablation transplant ofHABCrH\_\{ABC\}^\{r\}\. UnlikeHABrH\_\{AB\}^\{r\}, direct transfer of the three\-way interaction is not consistently successful across domains\. Recovery in performance is observed for the months↔\\leftrightarrowhours transfers in the Llama models, but most other model–domain pairs show little improvement over ablation\.Figure 24:Top\-3 accuracy for cross\-domain replacement usingHABCrH\_\{ABC\}^\{r\}\. Consistent with Figure[23](https://arxiv.org/html/2609.35970#A10.F23), the ability of a domain\-2HABCrH\_\{ABC\}^\{r\}to substitute for its domain\-1 counterpart depends strongly on the model and domain pair\.Figure 25:Top\-1 accuracy for cross\-domain ablation transplant after aligning the interaction subspaces using training replicates and evaluating on held\-out replicates\. Solid curves show aligned domain\-2 interaction vectors transplanted into domain 1; dashed curves show the corresponding ablation and unmodified baselines\. Alignment improves transfer for bothHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}, although the degree of recovery varies across models and domain pairs\.Figure 26:Top\-3 accuracy for cross\-domain ablation transplant after interaction\-subspace alignment\. Alignment improves recovery over the ablated baseline for bothHABrH\_\{AB\}^\{r\}andHABCrH\_\{ABC\}^\{r\}, with the largest qualitative change occurring forHABCrH\_\{ABC\}^\{r\}, whose direct transfer was often weak\.Figure 27:Top\-1 accuracy for cross\-domain replacement after aligning the interaction subspaces\. Solid curves retain the domain\-1 mean together with an aligned domain\-2 interaction vector; dashed curves show the corresponding within\-domain replacement\. Alignment allows the transplanted interaction to reproduce the within\-domain replacement trajectory more closely, particularly forHABCrH\_\{ABC\}^\{r\}\.Figure 28:Top\-3 accuracy for cross\-domain replacement after interaction\-subspace alignment\. As in the top\-1 results in Figure[27](https://arxiv.org/html/2609.35970#A10.F27), alignment improves agreement between cross\-domain and within\-domain replacement, especially forHABCrH\_\{ABC\}^\{r\}, although the degree of recovery depends on the model and domain\.
相似文章
因果世界模型何时有助于模块化 LLM 智能体
论文提出 FedCausalCompose,一个面向模块化 LLM 智能体的因果世界模型框架,探讨因果结构在何种条件下有助于跨模块工具调用,并指出只有当跨模块接口在统计上可识别且以决策可用的形式呈现时,因果世界模型才真正有效。
可解释的人类与陌生的LLMs:评估响应中潜结构的专家分析
本研究采用探索性因子分析,比较人类与LLMs在评估中响应的潜在结构,揭示LLMs依赖与人类推理不同的统计上不透明的机制。
揭示大语言模型中的数学推理:内部机制的方法学研究
本文通过早期解码分析大语言模型的内部机制,研究其如何执行算术运算。研究发现,能力强的模型在推理任务中,注意力模块和 MLP 模块之间呈现明确的分工。
Cross-LLM推理一致性:来自共享交互的证据
本文利用基于交互的解释方法,研究了不同LLM在预测相同词元时是否共享共同的推理模式。结果表明,先进LLM展现出一致的交互模式,暗示它们隐式地优化到了共享的推理机制。
MLLM 为何失败及其原因:面向能力失败诊断的因果任务分解
该论文提出了一种因果分解框架,用于判定 MLLM 在组合任务上的失败究竟源于其内在能力不足,还是来自上游前置知识缺失所引发的级联错误;同时推出了 CADET 基准,包含 46 项单元任务和超过 33,000 条标注问题,研究发现:补充关键前置条件后,模型表现出的推理能力缺陷在很大程度上会消失。