Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

arXiv cs.CL Papers

Summary

The paper introduces Coherentist Probabilistic Compositionalism (CPC), a framework for interpreting transformer computation through four operator roles, and validates it across 15 models from five architecture families.

arXiv:2608.22034v1 Announce Type: new Abstract: Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
Original Article
View Cached Full Text

Cached at: 08/25/26, 04:24 AM

# Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
Source: [https://arxiv.org/html/2608.22034](https://arxiv.org/html/2608.22034)
\\jvol

vv\\jnumnn 2026\\docheadPreprint

\\affilblock

André Freitas1,2Email:[nura\.aljaafari@manchester\.ac\.uk](mailto:[email protected])Affiliation:University of Manchester, United KingdomEmail:[andre\.freitas@idiap\.ch](mailto:[email protected])Affiliation:Idiap Research Institute, Switzerland

###### Abstract

Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures\. We introduce*Coherentist Probabilistic Compositionalism*\(CPC\), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles\.*Alignment*identifies candidate relations,*unification*integrates supporting information,*suppression*reduces incompatible alternatives, and*routing*carries selected information to the output\. Across 15 models from five architecture families, the suppression, unification, and routing weight\-space signatures correlate with held\-out activation\-level role measures above random baselines\. Suppression is more stable across tasks than unification\. Ablating alignment heads reduces downstream suppressive activity beyond a random\-head control in 10 models, but similar effects on no\-conflict prompts indicate a general upstream dependency, not contradiction\-specific coupling\. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model\. Base and instruction\-tuned variants preserve induction\-head score structure \(r≥0\.98r\{\\geq\}0\.98\) without a consistent shift of operator signatures towards later layers\. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture\-specific\.111Code and data will be made publicly available upon acceptance\.

## 1Introduction

As transformer language models scale, their capabilities extend beyond surface pattern matching to in\-context learning and other emergent behaviours\([51](https://arxiv.org/html/2608.22034#bib.bib1);[49](https://arxiv.org/html/2608.22034#bib.bib3);[70](https://arxiv.org/html/2608.22034#bib.bib4)\)\. Post\-training further refines model behaviour\([52](https://arxiv.org/html/2608.22034#bib.bib21);[4](https://arxiv.org/html/2608.22034#bib.bib22)\)\. Mechanistic interpretability seeks to explain these behaviours through the analysis of internal components and computations, often identifying circuits: sets of interacting components that contribute to a particular computation or behaviour\([68](https://arxiv.org/html/2608.22034#bib.bib6)\)\. Such analyses have identified induction heads, the indirect\-object identification \(IOI\) circuit, copy\-suppression heads, feedforward memories, and instruction\-conditioned conflict resolution\([15](https://arxiv.org/html/2608.22034#bib.bib8);[21](https://arxiv.org/html/2608.22034#bib.bib5);[51](https://arxiv.org/html/2608.22034#bib.bib1);[68](https://arxiv.org/html/2608.22034#bib.bib6);[45](https://arxiv.org/html/2608.22034#bib.bib7);[69](https://arxiv.org/html/2608.22034#bib.bib23);[62](https://arxiv.org/html/2608.22034#bib.bib13);[2](https://arxiv.org/html/2608.22034#bib.bib20)\)\. These mechanisms are well characterised individually, but how their functional roles fit within a broader framework of transformer computation is not\.

Existing frameworks approach this question at different levels of abstraction\. Bayesian and algorithmic analyses treat prompts as evidence for latent task inference or as inputs to implicit estimators\([72](https://arxiv.org/html/2608.22034#bib.bib15);[1](https://arxiv.org/html/2608.22034#bib.bib16);[67](https://arxiv.org/html/2608.22034#bib.bib18);[18](https://arxiv.org/html/2608.22034#bib.bib17)\), abstracting away the components that implement the computation\. Feature\-level analyses decompose activations into interpretable directions\([14](https://arxiv.org/html/2608.22034#bib.bib9);[33](https://arxiv.org/html/2608.22034#bib.bib10);[8](https://arxiv.org/html/2608.22034#bib.bib12)\), primarily characterising representational content and its local causal roles, with less emphasis on how this content is combined and revised across depth\. Circuit analyses identify causal components and reusable computational structure\([73](https://arxiv.org/html/2608.22034#bib.bib47);[28](https://arxiv.org/html/2608.22034#bib.bib67);[3](https://arxiv.org/html/2608.22034#bib.bib19)\), but do not resolve how recurring functions can be described through graded roles that generalise across tasks and architectures\. We address this gap by organising these functions into a shared vocabulary grounded in*coherence refinement*: the view that a model interprets its context by linking representational fragments and progressively resolving inconsistencies among them\. We test how the organisation of these functions varies across tasks, architectures, and post\-training methods\.

Figure 1:The CPC vocabulary on the give\-frame running example\. Partial cues enter, alignment proposes candidate role links, unification integrates the compatible assignments, suppression prunes the incompatible recipient, and routing conveys the consistent interpretation to the readout, with coherence increasing across depth\. The legend names a characteristic mechanism for each operator, and the same operators also cover factual recall, contradiction handling, and instruction tuning\. Individual mechanisms may instantiate one or more operator roles\.Since language is a structured field of inferential and referential relations, plausible continuations are constrained by discourse entities, semantic roles, and coherence relations established earlier in the context, not by surface adjacency alone\([24](https://arxiv.org/html/2608.22034#bib.bib52);[23](https://arxiv.org/html/2608.22034#bib.bib53);[16](https://arxiv.org/html/2608.22034#bib.bib51)\)\. For a token\-based language model, these relations must be realised through representations distributed across token positions\. The representations link mentions, roles, and events, and the model must keep them compatible as the context extends \(see Figure[2](https://arxiv.org/html/2608.22034#S3.F2)\)\. Self\-supervised language models encode aspects of this structure, including syntactic dependencies, semantic roles, and coreference\([63](https://arxiv.org/html/2608.22034#bib.bib55);[43](https://arxiv.org/html/2608.22034#bib.bib54);[32](https://arxiv.org/html/2608.22034#bib.bib56)\)\. Models can draw on this structure by relating context fragments, integrating compatible information, reducing support for contradictions, and carrying the interpretation to the prediction position\([51](https://arxiv.org/html/2608.22034#bib.bib1);[68](https://arxiv.org/html/2608.22034#bib.bib6);[21](https://arxiv.org/html/2608.22034#bib.bib5);[45](https://arxiv.org/html/2608.22034#bib.bib7)\)\.

We propose*coherence refinement*as the organising principle for a vocabulary of these operations and formalise it as*Coherentist Probabilistic Compositionalism*\(CPC\)\. In CPC \(Figure[1](https://arxiv.org/html/2608.22034#S1.F1)\), transformer computation is modelled as the iterative refinement of representational fragments under a*global coherence functional*𝒞\\mathcal\{C\}, and components across layers may instantiate four operator roles that realise these operations:*alignment*,*unification*,*suppression*, and*routing*\. Each describes what a component \(e\.g\., an attention head or a feedforward layer\) contributes to the developing interpretation: a single component may serve several roles, and one role may be spread across many components\. The empirical question is how far weights and activations carry measurable signatures of these roles\. We treat this as a modelling hypothesis, organising known mechanisms and generating testable predictions without assuming that transformers optimise𝒞\\mathcal\{C\}directly\.

The empirical evaluation examines four dimensions: \(i\) the reproducibility and functional validity of operator signatures \(P1\); \(ii\) the dependence of suppression on upstream alignment \(P2\); \(iii\) the effect of contradiction on the coherence proxy near layers with high alignment and suppression activity \(P3\); and \(iv\) the extent to which post\-training reweights operator structure or introduces new structure \(P4\)\. The paper makes three contributions:

- •An interpretive framework that describes transformer computation as coherence refinement carried out by four operator roles, giving mechanistic findings a shared representational language \(Section[3](https://arxiv.org/html/2608.22034#S3)\)\.
- •Operational definitions of attention\-head operator signatures, with activation\-level measures and causal tests of their proposed functional roles \(Sections[3\.4](https://arxiv.org/html/2608.22034#S3.SS4)and[5\.3](https://arxiv.org/html/2608.22034#S5.SS3)\)\.
- •Empirical evaluation across 15 models from five architecture families, identifying which properties of the vocabulary generalise between architectures \(Section[5](https://arxiv.org/html/2608.22034#S5)\)\.

## 2Coherence in Language and Cognition

Coherence has been investigated in both philosophy and linguistics\. In philosophy, theories take an epistemological view, where a belief is justified partly by the degree to which it aligns with the broader system of beliefs held by a reasoner\([7](https://arxiv.org/html/2608.22034#bib.bib57)\)\. In linguistics, coherence is used to characterise the relations between text units that make a discourse interpretable as a structured whole, different from an arbitrary sequence of sentences\([29](https://arxiv.org/html/2608.22034#bib.bib63);[35](https://arxiv.org/html/2608.22034#bib.bib62)\)\. Despite these different objects of study, both perspectives treat coherence as a relational property: an element is evaluated by how it fits with other elements, not in isolation\.

### 2\.1Coherence in Language and Discourse

In linguistic settings, coherence is commonly studied through*coherence relations*: relations that connect linguistic units such as clauses, sentences, or larger discourse spans\([29](https://arxiv.org/html/2608.22034#bib.bib63)\)\. They explain how one part of a discourse is interpreted with respect to another\. For example, two spans may stand in a causal, elaborative, temporal, or contrastive relation\([35](https://arxiv.org/html/2608.22034#bib.bib62)\)\. One of the most widely used theories is the Rhetorical Structure Theory, which formalises this view by representing a text through relations between discourse spans, including relations such asElaboration,Cause, andContrast\([42](https://arxiv.org/html/2608.22034#bib.bib58)\)\.

[35](https://arxiv.org/html/2608.22034#bib.bib62)distinguish between local and global coherence\. Local coherence refers to how nearby parts of a text are connected\. It has three main sources: \(i\) coherence relations between neighbouring text spans, \(ii\) continuity of discourse entities, where successive spans remain centred on the same people, objects, or events, and \(iii\) lexical or semantic continuity, where neighbouring spans remain within a compatible topic or semantic field\. Global coherence concerns the organisation of the discourse over longer stretches of text, including how individual sentences and local relations contribute to the overall discourse structure\.

Several computational approaches formalise these sources separately\. Cohesion models describe surface links created through repetition, reference, and lexical relations\([24](https://arxiv.org/html/2608.22034#bib.bib52)\)\. Entity\-based models describe how discourse entities are introduced and maintained\. For instance, Centring Theory tracks the changing focus of attention across utterances and predicts greater coherence when transitions preserve salient discourse entities\([23](https://arxiv.org/html/2608.22034#bib.bib53)\)\. On the other hand, Semantic approaches measure compatibility between neighbouring textual representations\. More recent computational models learn such compatibility from coherent and reordered text, producing a score that distinguishes well\-formed discourse from less coherent alternatives\([39](https://arxiv.org/html/2608.22034#bib.bib59)\)\. These approaches differ in what constitutes a discourse unit and how compatibility is measured, but all describe coherence in terms of relations among parts of the text\.

An important consequence is that semantic opposition does not itself imply incoherence\. A discourse may contain disagreement, contrast, correction, or concession while remaining coherent if the relation between the opposing contents is represented explicitly\. For example,Contrastis itself a coherence relation in Rhetorical Structure Theory\([42](https://arxiv.org/html/2608.22034#bib.bib58)\)\. In what follows,*incompatibility*refers instead to two assignments that cannot be jointly maintained within a single interpretation\.

### 2\.2Coherence as Constraint Satisfaction

A more general formalisation treats coherence as a constraint\-satisfaction problem\.[64](https://arxiv.org/html/2608.22034#bib.bib60)define a finite set of elementsE=\{e1,…,en\}E=\\\{e\_\{1\},\\ldots,e\_\{n\}\\\}with weighted positive and negative constraints between pairs of elements\. Positive constraints connect elements that support or fit with one another, while negative constraints connect elements that conflict\. A solution partitions the elements into an accepted set𝒜\\mathcal\{A\}and a rejected setℛ\\mathcal\{R\}\. A positive constraint betweeneie\_\{i\}andeje\_\{j\}is satisfied when both elements receive the same status, while a negative constraint is satisfied when one is accepted and the other rejected\. The coherence of a partition can then be written as

W⁡\(𝒜,ℛ\)=∑\(ei,ej\)​satisfiedwi​j,\\displaystyle W\(\\mathcal\{A\},\\mathcal\{R\}\)=\\sum\_\{\(e\_\{i\},e\_\{j\}\)\\,\\mathrm\{satisfied\}\}w\_\{ij\},\(1\)wherewi​jw\_\{ij\}is the weight associated with the constraint betweeneie\_\{i\}andeje\_\{j\}\. The coherence problem is to find the partition that maximisesWW\.

Exact maximisation of this objective is computationally difficult\. A connectionist approximation represents elements as units in a network, positive constraints as excitatory connections, and negative constraints as inhibitory connections\([64](https://arxiv.org/html/2608.22034#bib.bib60)\)\. Mutually supporting elements reinforce one another, while conflicting elements compete\. The network progressively settles into a configuration in which compatible elements tend to be active together and incompatible alternatives are separated\.[65](https://arxiv.org/html/2608.22034#bib.bib61)applies this general formulation to different cognitive problems, including explanation, perception, analogy, and decision making\. The elements and constraints differ between domains, but the underlying problem remains the same: finding a configuration that satisfies as many mutually weighted constraints as possible\.

This formulation provides a useful bridge between epistemological and linguistic perspectives\. In discourse, the elements may correspond to propositions, entities, events, or discourse spans\. Positive constraints can represent relations that support a joint interpretation, such as coreference, causal relations, or compatible semantic roles\. Negative constraints can represent mutually incompatible assignments\. Coherence then depends on the configuration of these relations across the interpretation\.

## 3Coherentist Probabilistic Compositionalism

CPC describes transformer computation through the layerwise construction of an interpretation in the residual stream\. The operator roles specify functional contributions to this process, while analyses of transformer components and circuits provide the empirical basis for identifying them\.

### 3\.1Coherence State and Functional

A*coherence state*A=\{\(m,g\)\}A\{=\}\\\{\(m,g\)\\\}is a set of representational fragmentsmmpaired with groundingsgg\. A fragment captures a partial hypothesis about the input, such as a token span, entity mention, syntactic role, or latent feature\. A grounding maps a fragment into a relational structureGGrepresenting the current interpretation\. Figure[2](https://arxiv.org/html/2608.22034#S3.F2)illustrates a coherence state and its groundings\. We define the global coherence functional as

𝒞⁡\(A\)=\\displaystyle\\mathcal\{C\}\(A\)\{=\}μ​cov​\(A\)⏟coverage reward\(unification\)\+∑\(m,g\)∈Alog⁡𝒮coh​\(g∣m,G\)⏟local fit\(alignment, unification\)\\displaystyle\\underbrace\{\\mu\\,\\mathrm\{cov\}\(A\)\}\_\{\\begin\{subarray\}\{c\}\\text\{coverage reward\}\\\\ \\text\{\(unification\)\}\\end\{subarray\}\}\+\\underbrace\{\\sum\_\{\(m,g\)\\in A\}\\log\\mathcal\{S\}\_\{\\text\{coh\}\}\(g\\mid m,G\)\}\_\{\\begin\{subarray\}\{c\}\\text\{local fit\}\\\\ \\text\{\(alignment, unification\)\}\\end\{subarray\}\}\(2\)−λ2​∑\(m,g\),\(m′,g′\)∈A\(m,g\)≠\(m′,g′\)κ⁡\(m,g,m′,g′\)⏟contradiction penalty\(suppression\),\\displaystyle\-\\underbrace\{\\frac\{\\lambda\}\{2\}\\\!\\\!\\sum\_\{\\begin\{subarray\}\{c\}\(m,g\),\(m^\{\\prime\},g^\{\\prime\}\)\\in A\\\\ \(m,g\)\\neq\(m^\{\\prime\},g^\{\\prime\}\)\\end\{subarray\}\}\\\!\\\!\\kappa\(m,g;m^\{\\prime\},g^\{\\prime\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{contradiction penalty\}\\\\ \\text\{\(suppression\)\}\\end\{subarray\}\},where coverage is defined over grounded positions,cov⁡\(A\)=\|⋃\(m,g\)∈Adom⁡\(g\)\|\\mathrm\{cov\}\(A\)\{=\}\\bigl\|\\bigcup\_\{\(m,g\)\\in A\}\\operatorname\{dom\}\(g\)\\bigr\|,𝒮coh​\(g∣m,G\)∈\(0,1\]\\mathcal\{S\}\_\{\\text\{coh\}\}\(g\\mid m,G\)\{\\in\}\(0,1\]measures how well a grounding fits the relational structureGG, andκ≥0\\kappa\{\\geq\}0is a symmetric contradiction kernel that vanishes on compatible pairs\. Formal definitions of fragments, groundings, admissible pairings, and the contradiction kernel are given in Appendix[Appendix B: Formal Definitions](https://arxiv.org/html/2608.22034#Sx4)\. The functional rewards coverage \(μ\>0\\mu\{\>\}0\) and local compatibility while penalising incompatible assignments \(λ\>0\\lambda\{\>\}0\)\. The coverage term favours interpretations that incorporate more of the available input, the local\-fit term favours groundings compatible with the developing relational structure, and the contradiction term penalises grounded fragments that cannot be jointly maintained\. Figure[3](https://arxiv.org/html/2608.22034#S3.F3)illustrates the resulting view of coherence as a configuration of grounded fragments connected by relations of support and incompatibility\. Appendix[Appendix A: The CPC Functional and Coherentism](https://arxiv.org/html/2608.22034#Sx3)develops the relation to the formalisation of coherence in Section[2](https://arxiv.org/html/2608.22034#S2)\.

MaryandJohnwenttothestore,Johngaveadrinkto\_\_\_m1m\_\{1\}: entity*Mary*m2m\_\{2\}: entity*John*m3m\_\{3\}: location*store*m4m\_\{4\}: event*give*\(agent, recipient, theme\)relational graphGGMaryJohnstoregivedrinkrecipientagentatthemesolid grey lines mark fragment positionsUmU\_\{m\}dashed teal arrows are groundingsg:Um⇀Vg\\colon U\_\{m\}\\rightharpoonup V

Figure 2:A coherence state and its groundings\. Fragmentsm1m\_\{1\}\-m4m\_\{4\}capture partial hypotheses about spans of the input, grey lines mark the positionsUmU\_\{m\}they cover, and dashed teal groundings map fragment positions into the relational graphGG\.m1m\_\{1\}m2m\_\{2\}m3m\_\{3\}m4m\_\{4\}m5m\_\{5\}m6m\_\{6\}m7m\_\{7\}fragmented /competing staterefinementm1m\_\{1\}m2m\_\{2\}m3m\_\{3\}m4m\_\{4\}m5m\_\{5\}m6m\_\{6\}m7m\_\{7\}conflict exposed /compatible structure strengthensrefinementm1m\_\{1\}m2m\_\{2\}m3m\_\{3\}m4m\_\{4\}m5m\_\{5\}m6m\_\{6\}m7m\_\{7\}coherentconfigurationsupport \(compatibility\)conflict \(incompatibility\)fragment support level\(low→\\rightarrowhigh\)coverage↑\\uparrowcompatibility↑\\uparrowconflict↓\\downarrow⇒\\Rightarrow𝒞⁡\(A\)↑\\mathcal\{C\}\(A\)\\uparrow

Figure 3:Coherence as a representational configuration\. Circles are grounded fragments, fill level shows support, teal links mark mutual support, and red dashed links mark unresolved incompatibility, not disagreement\. Refinement increases coverage and compatibility and decreases conflict, raising𝒞⁡\(A\)\\mathcal\{C\}\(A\)\(Section[2](https://arxiv.org/html/2608.22034#S2);[7](https://arxiv.org/html/2608.22034#bib.bib57);[64](https://arxiv.org/html/2608.22034#bib.bib60)\)\. All fragments can be retained provided their relations are admissible, and support stays graded\.
### 3\.2The Four Operators

The four operators describe recurring functional contributions to the refinement of a coherence state\. To connect these roles to transformer computation, we use the residual\-stream and attention decomposition of[15](https://arxiv.org/html/2608.22034#bib.bib8)\. Transformer blocks update an additive residual stream, while attention heads separate into query\-key \(QK\) interactions, which determine which positions attend to one another, and value\-output \(OV\) maps, which determine what information is written from attended positions back to the residual stream\. Figure[4](https://arxiv.org/html/2608.22034#S3.F4)locates the four roles on this architecture\.

x1x\_\{1\}x2x\_\{2\}x3x\_\{3\}x4x\_\{4\}xTx\_\{T\}layerℓ\\ellℓ\+1\\ell\{\+\}1ℓ\+2\\ell\{\+\}2align \(QK scores relations\)unify \(OV/MLP write,\+\+\)suppress \(−\-\)route \(mover head\)readout

Figure 4:The CPC operators on the transformer architecture, with tokens on the left and depth running left to right\. Query\-key structure aligns positions, OV and feedforward writes unify content, subtractive writes suppress incompatible content, and a mover head routes the result to the readout\.#### Alignment

It identifies candidate relations between positions or representational fragments\. In transformers, QK interactions provide input\-dependent compatibility scores that determine which positions exchange information\.

#### Unification

It incorporates compatible content into the developing interpretation\. Once a fragment is aligned with a candidate role or relation, additive component writes can increase support for the corresponding interpretation in the residual stream without removing competing alternatives\.

#### Suppression

It reduces support for continuations, features, or fragments that are incompatible with the developing interpretation\. In transformers, this role can be implemented through subtractive component writes along particular readout directions\.

#### Routing

It carries selected content to positions where it can affect the readout\. It changes where information is available without necessarily introducing new representational content\.

These roles correspond to different changes in the coherence state and functional\. Alignment proposes candidate groundings\. Unification incorporates compatible grounded fragments, increasing coverage and local fit\. Suppression reduces support for alternatives that contribute to the contradiction penalty\. Routing does not directly change𝒞\\mathcal\{C\}and instead controls where the resulting content is available to affect the model’s prediction\.

### 3\.3Layerwise Coherence Refinement

CPC interprets transformer depth as a sequence of approximate coherence\-refinement steps in which the operators act on the current representational state\. Figure[1](https://arxiv.org/html/2608.22034#S1.F1)illustrates this process on the running example\. A stronger conjecture is that the latent coherence functional tends to improve across depth on typical inputs:

𝒞⁡\(A\(ℓ\+1\)\)≳𝒞⁡\(A\(ℓ\)\)\.\\mathcal\{C\}\\\!\\left\(A^\{\(\\ell\+1\)\}\\right\)\\gtrsim\\mathcal\{C\}\\\!\\left\(A^\{\(\\ell\)\}\\right\)\.\(3\)We treat Equation[3](https://arxiv.org/html/2608.22034#S3.E3)as a falsifiable conjecture, not as an assumption of monotonic improvement at every layer\. Operators may overlap, repeat, or interact across layers, and CPC does not impose a fixed sequence in which alignment must precede unification, suppression, or routing\.

### 3\.4Transformer Mechanisms as CPC Operators

Known circuits and transformer mechanisms can be described through the four operator roles, although the correspondence is not one\-to\-one\. Circuit findings are measured through attention patterns, activations, ablations, and logit contributions, while the operator labels describe functional contributions at a higher level of abstraction\. A single mechanism may combine several roles, and one role may be implemented by several components\.

Table 1:CPC operators, their candidate mechanistic signatures, and representative circuits associated with each role; sources are cited in the text\.Table[1](https://arxiv.org/html/2608.22034#S3.T1)summarises the proposed correspondence\. Induction combines alignment and unification: a previous\-token head supplies shifted positional context, and the induction head copies a compatible continuation through its OV map\([51](https://arxiv.org/html/2608.22034#bib.bib1)\)\. The IOI circuit combines all four roles: duplicate\-token heads align repeated names, S\-inhibition heads suppress the duplicated\-subject pathway, and name\-mover heads route the correct name to the prediction position while their OV writes increase support for that name\([68](https://arxiv.org/html/2608.22034#bib.bib6)\)\. Copy\-suppression heads provide a direct example of subtractive writes\([45](https://arxiv.org/html/2608.22034#bib.bib7)\), while feedforward key\-value memories provide an example of content integration\([21](https://arxiv.org/html/2608.22034#bib.bib5)\)\. Instruction\-conditioned suppression connects the vocabulary to post\-training\([69](https://arxiv.org/html/2608.22034#bib.bib23)\)\. Sparse autoencoder features offer a candidate measurement basis for CPC fragments\([33](https://arxiv.org/html/2608.22034#bib.bib10);[62](https://arxiv.org/html/2608.22034#bib.bib13)\), although the correspondence remains partial\.

#### Per\-head operator scores

For each headhh, we derive weight\-space quantities corresponding to the proposed mechanistic roles\. LetWO​V=WO​WVW\_\{OV\}=W\_\{O\}W\_\{V\}denote the effective OV map, with singular value decompositionWO​V=∑iσi​ui​vi⊤W\_\{OV\}=\\sum\_\{i\}\\sigma\_\{i\}u\_\{i\}v\_\{i\}^\{\\top\}\. For residual staterrand token readout directionete\_\{t\}, its contribution to the token logit is

et⊤​WO​V​r=∑iσi​⟨et,ui⟩​⟨vi,r⟩\.e\_\{t\}^\{\\top\}W\_\{OV\}\\,r=\\sum\_\{i\}\\sigma\_\{i\}\\langle e\_\{t\},u\_\{i\}\\rangle\\langle v\_\{i\},r\\rangle\.\(4\)A mode contributes subtractively to tokenttwhen the corresponding term in Equation[4](https://arxiv.org/html/2608.22034#S3.E4)is negative\. This motivates separating additive and subtractive OV structure at the weight level, while task\-specific effects still depend on the residual staterrand readout directionete\_\{t\}\. We compute

salign​\(h\)\\displaystyle s\_\{\\mathrm\{align\}\}\(h\)=\(σh,1Q​K\)2/∑i\(σh,iQ​K\)2,\\displaystyle=\\bigl\(\\sigma^\{QK\}\_\{h,1\}\\bigr\)^\{2\}\\big/\{\\textstyle\\sum\_\{i\}\}\\bigl\(\\sigma^\{QK\}\_\{h,i\}\\bigr\)^\{2\},\(5\)ssup​\(h\)\\displaystyle s\_\{\\mathrm\{sup\}\}\(h\)=∑iσh,iO​V​max⁡\(0,−⟨uh,iO​V,vh,iO​V⟩\),\\displaystyle=\{\\textstyle\\sum\_\{i\}\}\\sigma^\{OV\}\_\{h,i\}\\,\\max\\\!\\left\(0,\-\\langle u^\{OV\}\_\{h,i\},v^\{OV\}\_\{h,i\}\\rangle\\right\),sunify​\(h\)\\displaystyle s\_\{\\mathrm\{unify\}\}\(h\)=∑iσh,iO​V​max⁡\(0,⟨uh,iO​V,vh,iO​V⟩\),\\displaystyle=\{\\textstyle\\sum\_\{i\}\}\\sigma^\{OV\}\_\{h,i\}\\,\\max\\\!\\left\(0,\\langle u^\{OV\}\_\{h,i\},v^\{OV\}\_\{h,i\}\\rangle\\right\),strace​\(h\)\\displaystyle s\_\{\\mathrm\{trace\}\}\(h\)=1d​tr​\(WO​V,h\),\\displaystyle=\\tfrac\{1\}\{d\}\\,\\mathrm\{tr\}\(W\_\{OV,h\}\),whered=dmodeld\{=\}d\_\{\\mathrm\{model\}\},σh,iQ​K\\sigma^\{QK\}\_\{h,i\}are the singular values of the QK map, andσh,iO​V\\sigma^\{OV\}\_\{h,i\},uh,iO​Vu^\{OV\}\_\{h,i\}, andvh,iO​Vv^\{OV\}\_\{h,i\}are the singular values and vectors of the OV map\. The alignment score measures QK spectral concentration, while the suppression and unification scores quantify subtractive and additive OV modes\. Since

tr⁡\(WO​V,h\)=∑iσh,iO​V​⟨uh,iO​V,vh,iO​V⟩,\\mathrm\{tr\}\(W\_\{OV,h\}\)=\\sum\_\{i\}\\sigma^\{OV\}\_\{h,i\}\\langle u^\{OV\}\_\{h,i\},v^\{OV\}\_\{h,i\}\\rangle,the trace descriptor satisfiesstrace=\(sunify−ssup\)/ds\_\{\\mathrm\{trace\}\}=\(s\_\{\\mathrm\{unify\}\}\-s\_\{\\mathrm\{sup\}\}\)/dexactly and is not treated as an independent signature coordinate\.

Routing is instead represented by the embedding\-space copy score of[15](https://arxiv.org/html/2608.22034#bib.bib8)\. For a seeded token sample with effective embeddings collected inEE\(token embedding plus the first\-layer MLP output;[45](https://arxiv.org/html/2608.22034#bib.bib7)\) and unembedding directions inUU, the mapMh=E​WO​V,h​UM\_\{h\}=E\\,W\_\{OV,h\}\\,Umeasures how strongly headhhmaps token representations towards their corresponding readout directions\. We define

scopy​\(h\)=∑imax⁡\{0,λi​\(Mh\)\}∑i\|λi​\(Mh\)\|∈\[0,1\],s\_\{\\mathrm\{copy\}\}\(h\)=\\frac\{\\sum\_\{i\}\\max\\\{0,\\lambda\_\{i\}\(M\_\{h\}\)\\\}\}\{\\sum\_\{i\}\|\\lambda\_\{i\}\(M\_\{h\}\)\|\}\\in\[0,1\],\(6\)whereλi\\lambda\_\{i\}are the eigenvalues ofMhM\_\{h\}\. A head that preserves token identity towards the readout scores near one, while a head dominated by inverted mappings scores near zero\. The copy score depends on the embedding geometry and is not an algebraic function of the other scores\. The resulting signature is\(salign,ssup,sunify,scopy\)\(s\_\{\\mathrm\{align\}\},s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{unify\}\},s\_\{\\mathrm\{copy\}\}\)\.

These signatures are defined for attention heads\. Feedforward contributions to unification \(Table[1](https://arxiv.org/html/2608.22034#S3.T1)\) enter CPC through the functional mapping and are not assigned separate weight\-space scores in the present evaluation\. Position\-dependent routing is also measured separately at the activation level\. The signatures nominate candidate operator roles whose functional validity is tested in Section[5](https://arxiv.org/html/2608.22034#S5)using held\-out activation measures and causal ablations\. CPC does not assume that every component belongs to one operator class, that every task requires all four roles, or that transformers explicitly optimise Equation[2](https://arxiv.org/html/2608.22034#S3.E2)\.

### 3\.5Activation\-Level Coherence Proxy

The functional𝒞\\mathcal\{C\}is latent, so the empirical analysis uses an activation\-level proxy for one possible geometric consequence of coherence:

𝒞^link\(ℓ\)=1\|P\|​∑\(i,j\)∈Psim⁡\(ri\(ℓ\),rj\(ℓ\)\),\\hat\{\\mathcal\{C\}\}\_\{\\mathrm\{link\}\}^\{\(\\ell\)\}=\\frac\{1\}\{\|P\|\}\\sum\_\{\(i,j\)\\in P\}\\mathrm\{sim\}\\\!\\left\(r\_\{i\}^\{\(\\ell\)\},\\,r\_\{j\}^\{\(\\ell\)\}\\right\),\(7\)whereri\(ℓ\)r\_\{i\}^\{\(\\ell\)\}is the residual\-stream vector at positioniiand layerℓ\\ell,PPis a set of position pairs, andsim\\mathrm\{sim\}is a bounded similarity function\. Higher similarity is treated as a candidate signature of representational coherence across the prompt, and an explicit contradiction is predicted to reduce it\.

The proxy captures only one geometric consequence that may accompany changes in𝒞\\mathcal\{C\}and does not uniquely isolate relational coherence\. We compare it against matched non\-contradictory and shuffled position\-pair controls, and evaluate a whitened variant that reduces the contribution of global residual\-stream covariance\. The contradiction experiment compares matched consistent and contradictory prompts across depth \(P3, Section[5\.5](https://arxiv.org/html/2608.22034#S5.SS5)\), while the ascent analysis tests the relation between the proxy and layer depth on mixed prompts \(Section[3\.3](https://arxiv.org/html/2608.22034#S3.SS3)\)\. The experimental choices ofPPandsim\\mathrm\{sim\}are specified in Section[4](https://arxiv.org/html/2608.22034#S4)\.

### 3\.6Post\-Training as Operator Reweighting

Post\-training methods, including supervised instruction tuning and reinforcement learning from human feedback \(RLHF\), modify model behaviour after pretraining\. The mechanistic question is how much these changes alter the organisation of computations already present in the base model\. RLHF provides one formal instance: it is commonly expressed as reward optimisation under a KL constraint to a reference policy\([52](https://arxiv.org/html/2608.22034#bib.bib21);[4](https://arxiv.org/html/2608.22034#bib.bib22)\)\. In CPC terms, the preference objective changes the relative importance of behaviours expressed through the operator roles, while the KL constraint limits divergence from the pretrained computation\. This motivates testing if post\-training primarily reweights existing operator structure\.

## 4Methodology

The experiments test four predictions: operator signatures in weight space \(P1\), suppression\-alignment coupling \(P2\), coherence under contradiction \(P3\), and operator reweighting under post\-training \(P4\)\. Layerwise composition and cross\-task stability serve as secondary analyses\.

#### Models

We evaluate 15 models from five families: GPT\-2\([58](https://arxiv.org/html/2608.22034#bib.bib37), Small, Medium, Large;\), Pythia\([5](https://arxiv.org/html/2608.22034#bib.bib11), 410M, 1\.4B, 2\.8B;\), Qwen 2\.5\([57](https://arxiv.org/html/2608.22034#bib.bib39), 1\.5B, 3B;\), Gemma 2\([61](https://arxiv.org/html/2608.22034#bib.bib40), 2B;\), and LLaMA 3\.2\([12](https://arxiv.org/html/2608.22034#bib.bib41), 1B, 3B;\), plus the available instruction\-tuned variants \(Gemma, LLaMA\-1B, LLaMA\-3B, Qwen\-1\.5B\)\. The P1 weight\-space analyses operate on model weights without inference\. Experiments using prompts cover all models, except P4, which compares the four base/instruction\-tuned pairs\.

#### Tasks and data

We use indirect object identification\([68](https://arxiv.org/html/2608.22034#bib.bib6), IOI;\), greater\-than comparison\([25](https://arxiv.org/html/2608.22034#bib.bib30)\), and factual recall\([46](https://arxiv.org/html/2608.22034#bib.bib32)\), formally defined in Appendix[Appendix C: Task Definitions](https://arxiv.org/html/2608.22034#Sx5)\. For each task, examples are filtered to those the model predicts correctly, with a target of 500 examples per model\-task pair\. Greater\-than requires the two\-digit year continuation to be a single token, which the LLaMA, Qwen, and Gemma tokenisers do not provide; the affected models are excluded from this task only\. Each task set is split 50/50 into discovery and held\-out subsets using a seeded hash fixed at generation time\. Head selection and ranking are performed on the discovery split before held\-out evaluation\. Contradiction experiments use 500 matched consistent/contradictory prompt pairs per run, sampled with a fixed seed from a frozen pool of 3,854 unique pairs \(generation templates and pool composition in Appendix[Appendix E: Contradiction Prompt Templates](https://arxiv.org/html/2608.22034#Sx7)\)\.

#### Operational measures

CPC is evaluated through weight\-space and activation\-level measures\. In weight space, each attention head is represented by the four signature coordinates defined in Section[3\.4](https://arxiv.org/html/2608.22034#S3.SS4),\(salign,ssup,sunify,scopy\)\(s\_\{\\mathrm\{align\}\},s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{unify\}\},s\_\{\\mathrm\{copy\}\}\)\. At the activation level, unification and suppression are measured through positive and negative direct logit attribution\([50](https://arxiv.org/html/2608.22034#bib.bib50), DLA;\)over the model’s top\-10 predicted tokens\. Routing is measured through an attention\-weighted copy score, and alignment through the maximum previous\-token or duplicate\-token attention on repeated random sequences\([51](https://arxiv.org/html/2608.22034#bib.bib1);[68](https://arxiv.org/html/2608.22034#bib.bib6)\)\. Layerwise coherence is measured using𝒞^link\(ℓ\)\\hat\{\\mathcal\{C\}\}\_\{\\mathrm\{link\}\}^\{\(\\ell\)\}\(Equation[7](https://arxiv.org/html/2608.22034#S3.E7)\), instantiated as the mean cosine similarity over all position pairs in the residual stream, excluding the beginning\-of\-sequence token\. Task\-based activation measures are averaged over correctly predicted examples\. For comparison across architectures of different depths, layerwise measures are aggregated into early, middle, and late thirds of the layer stack \(Appendix[Appendix D: Parameter and Threshold Choices](https://arxiv.org/html/2608.22034#Sx6)\)\.

#### P1: operator signatures

We first characterise the geometry of the four\-dimensional weight\-space signature\. Within each model, heads from all layers are pooled and each coordinate is standardised\. To identify recurring signature profiles, we applykk\-means clustering\([41](https://arxiv.org/html/2608.22034#bib.bib42)\)with 50 restarts fork∈\{2,…,7\}k\\in\\\{2,\\ldots,7\\\}\. To characterise continuous axes of variation, we apply principal component analysis\([54](https://arxiv.org/html/2608.22034#bib.bib43), PCA;\)to the same head\-by\-coordinate matrix\. The algebraically derived trace descriptor is excluded from both analyses\. Pooling is used because P1 concerns model\-wide operator structure, and the influence of depth is quantified separately through the depth\-separation ratio and depth residualisation\. Clustering quality is measured by the silhouette coefficient\([36](https://arxiv.org/html/2608.22034#bib.bib48)\)and calibrated against two nulls in the same four\-coordinate space: a column\-shuffle null, which preserves each coordinate’s marginal distribution while removing cross\-score dependence, and an isotropic standard Gaussian null with the same number of heads and dimensions\.

We test functional validity by correlating the suppression, unification, and routing weight\-space scores \(Spearman\) with their corresponding held\-out activation\-level measures\. The correlations are compared against the 95% quantile of 100 random weight projections, weight\-magnitude ranking, and, for suppression and unification, scores computed after shuffling task relations\. Alignment is evaluated separately through its activation\-level detector\. We also correlate the weight\-space scores with held\-out single\-head ablation effects, defined as the absolute change in the correct\-token logit after zero\-ablating a head\. This analysis uses the union of the top\-10 heads under each operator ranking, so a positive correlation indicates that higher\-scoring heads have larger causal effects irrespective of direction\.

For direct causal validation, we ablate the top five heads per signature dimension on IOI, using the same selection rule across models, and measure changes in the correct\-token \(IO\) logit, wrong\-token \(subject\) logit, and logit margin\. We compare these effects against layer\-matched random\-head ablations \(20 control samples per intervention\) and a magnitude\-matched control\. As a task\-conditioned check, suppression and unification heads are additionally ranked by the targeted DLA margin on the discovery split and ablated on the held\-out split\.

#### P2: suppression\-alignment coupling

Ablating upstream alignment heads should reduce the summed suppressive contributionssupact​\(h,x,t\)=max⁡\(0,−et⊤​oi∗\(h\)​\(x\)\)s^\{\\mathrm\{act\}\}\_\{\\mathrm\{sup\}\}\(h,x,t\)=\\max\\bigl\(0,\-e\_\{t\}^\{\\top\}o^\{\(h\)\}\_\{i^\{\\ast\}\}\(x\)\\bigr\)of downstream suppressor headshh, whereoi∗\(h\)​\(x\)o^\{\(h\)\}\_\{i^\{\\ast\}\}\(x\)is the head output at the readout position andete\_\{t\}the unembedding direction of the incompatible candidatett\. The candidate is the top\-1 continuation of the matched consistent prompt, rendered incompatible by the inserted contradiction\. On contradiction prompts, suppression is compared before and after alignment\-head ablation against two controls: layer\-matched random\-head ablation and the same intervention on matched no\-conflict prompts\. Significance is assessed by paired permutation tests \(10,000 permutations\), Holm\-corrected\([30](https://arxiv.org/html/2608.22034#bib.bib49)\)across the 30 model\-by\-contrast comparisons, with percentile bootstrap confidence intervals\([13](https://arxiv.org/html/2608.22034#bib.bib46)\)based on 10,000 resamples\. GPT\-2 Small uses the validated IOI circuit, with duplicate\- and previous\-token heads as alignment heads, and S\-inhibition and negative name\-mover heads as suppressors\. Elsewhere, alignment heads are the five heads scoring highest on the activation\-level alignment measure, which uses repeated random sequences and no task data, and suppressor heads are the six heads with the most\-negative mean DLA ontt, estimated on the discovery split of the contradiction prompts\.

#### P3: coherence under contradiction

The coherence gap is the consistent\-minus\-contradictory difference in𝒞^link\(ℓ\)\\hat\{\\mathcal\{C\}\}\_\{\\mathrm\{link\}\}^\{\(\\ell\)\}on matched prompt pairs\. It is tested per layer third using pairedtt\-tests with Holm correction, and checked for co\-location with the layer third in which the model’s alignment or suppression activity peaks\. Three controls qualify the raw gap\. Lexically matched non\-contradictory pairs isolate the contradiction substitution, while a shuffled position\-pair control, which pairs positions across prompts, separates within\-prompt relational structure from a global state shift\. A per\-layer whitened proxy applies shrinkage\-regularised ZCA whitening\([37](https://arxiv.org/html/2608.22034#bib.bib44)\)to residual vectors before cosine similarity, with covariance shrunk towards the average\-variance identity \(α=0\.1\\alpha\{=\}0\.1\), and is used to test the directional prediction\. The layerwise logit\-lens rank trajectory\([50](https://arxiv.org/html/2608.22034#bib.bib50)\)of the coherent continuation provides a scale\-free convergent measure, and a naturalistic Wikipedia prompt set probes out\-of\-template replication\.

#### P4: post\-training reweighting

The base/instruction\-tuned pairs are compared through per\-head operator signatures in the early and late halves of the network\. We additionally measure the Pearson correlation of induction scores for corresponding heads between variants\. Preservation is assessed using Fisher’szz\-transformation\([17](https://arxiv.org/html/2608.22034#bib.bib45)\)with the one\-sided hypothesesH0:ρ≤0\.9H\_\{0\}\{:\}\\ \\rho\{\\leq\}0\.9andH1:ρ\>0\.9H\_\{1\}\{:\}\\ \\rho\{\>\}0\.9\.

#### Secondary analyses

Layerwise composition averages the per\-head activation\-level measures by layer third on correctly predicted prompts\. A targeted variant contrasts compatible and incompatible tokens through the per\-head DLA margin \(indirect object vs\. subject in IOI, answer vs\. distractor in factual recall\)\. Cross\-task stability correlates per\-head DLA\-based unification and suppression scores between task pairs, subject to the task exclusions above\.

#### Decision rules and statistical testing

The following decision rules were fixed before the runs\. Outcomes satisfying none of the statedsupport,partial, ormixedcriteria are labelledagainst\.*P1 \(signatures\):*supportrequires the mean of the corresponding functional correlations to exceed all applicable baselines and the ablation\-effect correlation to be positive;partialrequires at least one applicable baseline to be exceeded\. The optimalkkis the silhouette peak overk∈\{2,…,7\}k\\in\\\{2,\\ldots,7\\\}\.*P2 \(coupling\):*supportrequires alignment\-head ablation to reduce summed suppressive contribution more than random\-head ablation \(Holm\-correctedα=0\.05\\alpha\{=\}0\.05\) and the reduction to be weaker on matched no\-conflict prompts;partialrequires only the first condition\.*P3 \(contradiction sensitivity\):*supportrequires a significant raw coherence gap in at least one layer third \(Holm\-corrected\), with the largest absolute gap co\-locating with the model’s peak alignment or suppression third \(Section[5\.2](https://arxiv.org/html/2608.22034#S5.SS2)\);mixeddenotes a significant but non\-co\-located gap\. The directional prediction is scored separately on the whitened proxy, wheresupportrequires a significant positive consistent\-minus\-contradictory gap\.*P4 \(base vs\. tuned\):*supportrequires the late suppression\-and\-routing shift to exceed the early alignment shift and induction preservation to pass the one\-sided test with thresholdr=0\.9r\{=\}0\.9;partialrequires one condition\.*Layerwise composition:*supportrequires unification to peak in the middle or late third, routing in the late third, and suppression to increase from early to middle;partialrequires two of three conditions\.*Cross\-task stability:*supportrequires mean cross\-taskr\>0\.5r\{\>\}0\.5for both suppression and unification;partialrequires the threshold for one operator\.

#### Software and implementation

Experiments were run on NVIDIA RTX A6000 GPUs with Python 3\.11\.15\. Core dependencies are NumPy 2\.4\.6\([27](https://arxiv.org/html/2608.22034#bib.bib72)\), PyTorch 2\.7\.1 with CUDA 12\.6\([53](https://arxiv.org/html/2608.22034#bib.bib64)\), TransformerLens 3\.5\.1\([48](https://arxiv.org/html/2608.22034#bib.bib38)\), Transformers 5\.14\.1\([71](https://arxiv.org/html/2608.22034#bib.bib66)\), scikit\-learn 1\.9\.0\([55](https://arxiv.org/html/2608.22034#bib.bib65)\), SciPy 1\.17\.1\([66](https://arxiv.org/html/2608.22034#bib.bib73)\), and h5py 3\.16\.0\.

## 5Results and Discussion

### 5\.1Operator Signatures \(P1\)

CPC predicts reproducible functional information in the operator scores\(salign,ssup,sunify,scopy\)\(s\_\{\\mathrm\{align\}\},s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{unify\}\},s\_\{\\mathrm\{copy\}\}\)\(Section[3\.4](https://arxiv.org/html/2608.22034#S3.SS4)\), without assuming that heads form four disjoint operator classes\.

Table 2:Operator\-signature structure and functional validation \(means over 5 seeds\), computed on the four independent signature coordinates \(alignment, suppression, unification, routing\-copy\)\. Sil\.: silhouette atk=4k\{=\}4; Bestkk: silhouette\-optimal cluster count;pshufp\_\{\\mathrm\{shuf\}\}: maximum column\-shuffle nullpp\-value across seeds; W→\\toActrr: held\-out Spearman correlation between weight scores and activation\-level role measures, against the weight\-magnitude baseline \(Magn\.\); Bl\.: baselines beaten \(random\-projection, shuffled\-relation, weight\-magnitude\); Abl\.: positive rank correlation with held\-out single\-head ablation effects\. The pre\-specified P1 verdict issupportwith 3/3 baselines and a positive ablation correlation \(three models\),partialotherwise \(twelve\)\.#### Signature scores correlate with held\-out functional roles

On held\-out inputs, the suppression, unification, and routing scores correlate with the matching activation\-level role measures at mean Spearmanrrbetween0\.210\.21and0\.580\.58per model, positive throughout \(Table[2](https://arxiv.org/html/2608.22034#S5.T2)\)\. The copy score is the strongest single predictor \(r=0\.59r\{=\}0\.59against the activation\-level copy measure, versus0\.400\.40for the derived trace score\)\. The correlation exceeds the random\-projection and shuffled\-relation baselines in all models, but the weight\-magnitude baseline in only 10: in the Qwen family and two LLaMA variants, generic OV magnitude tracks activation\-level roles as well as the structured scores do,*suggesting that these families concentrate functional roles in their largest\-norm heads*\. However, the alignment weight score does not track the attention\-based alignment measure \(r≈−0\.04r\{\\approx\}\{\-\}0\.04\); evidence for the alignment role accordingly depends on the activation\-level detection score\. Predicting held\-out single\-head ablation effects is harder, with a positive rank correlation in 3 models\. The pre\-specified rule gives 3supportand 12partialverdicts, with no modelagainst:*the signatures carry functional role information beyond chance in every model and beyond naive magnitude in two\-thirds*, while reliable prediction of causal effect sizes remains the open gap\.

#### Structured signatures appear across all families, but the optimal cluster count varies

Silhouette coefficients atk=4k\{=\}4range from0\.360\.36to0\.500\.50\(Table[2](https://arxiv.org/html/2608.22034#S5.T2)\)\. A two\-cluster separation between QK\- and OV\-specialised heads provides a natural baseline, with higherkkindicating finer role\-dominant profiles\. The preferred cluster count remains architecture\-dependent:k=2k\{=\}2across the Qwen family,k=3k\{=\}3in the GPT\-2 family, Pythia\-1\.4B, and LLaMA\-1B,k=4k\{=\}4in the remaining LLaMAs, andk=5k\{=\}5ork=6k\{=\}6in Gemma, Pythia\-410M, and Pythia\-2\.8B\. No single cluster count describes all families, and no family cleanly reproduces a four\-profile organisation\.*This supports the framework’s position that operator roles are graded properties of heads, not disjoint classes*\(Section[3\.4](https://arxiv.org/html/2608.22034#S3.SS4)\): most heads mix additive and subtractive OV modes, and the preferred count tracks how sharply a family separates its few specialised heads from this mixed background\.

#### Cross\-score structure exceeds marginal\-preserving nulls in most models

Two nulls calibrate the silhouettes \(Table[2](https://arxiv.org/html/2608.22034#S5.T2); Section[4](https://arxiv.org/html/2608.22034#S4)\)\. Every model exceeds the Gaussian null \(p<0\.01p\{<\}0\.01against 200 draws; meanΔ=\+0\.28\\Delta\{=\}\+0\.28\), and 11 exceed the column\-shuffle null atp<0\.01p\{<\}0\.01in every seed \(meanΔ=\+0\.09\\Delta\{=\}\+0\.09, exceptions GPT\-2 Small, Pythia\-410M, and Qwen\-1\.5B variants\)\. The joint structure is carried largely by the copy score’s coupling to the OV mode balance \(r⁡\(ssup,scopy\)=−0\.61r\(s\_\{\\mathrm\{sup\}\},s\_\{\\mathrm\{copy\}\}\)\{=\}\{\-\}0\.61,r⁡\(sunify,scopy\)=\+0\.44r\(s\_\{\\mathrm\{unify\}\},s\_\{\\mathrm\{copy\}\}\)\{=\}\{\+\}0\.44\): heads dominated by subtractive OV modes anti\-copy token identity\. Clusters track depth only partially: the depth\-separation ratio \(between\-band share of signature variance\) is0\.020\.02–0\.240\.24\(mean0\.080\.08\), and regressing out layer position leaves substantial clusterability \(best\-kksilhouette0\.350\.35–0\.560\.56\), changing the preferred count in 5 models\. Base and instruction\-tuned variants produce nearly identical silhouettes \(difference≤0\.001\{\\leq\}0\.001; bestkkunchanged except LLaMA\-1B,3→43\{\\to\}4\), suggesting limited weight\-space reorganisation\.

#### The signatures separate QK specialisation from OV mode balance

PCA on the four standardised coordinates gives components of49%49\\%,27%27\\%,18%18\\%, and7%7\\%: no single axis dominates\. Alignment is nearly uncorrelated with every other score \(\|r\|≤0\.07\|r\|\\leq 0\.07\) and suppression and unification are moderately anti\-correlated \(r=−0\.30r\{=\}\{\-\}0\.30\)\. QK specialisation hence varies independently of an OV subspace in which mode balance and copying are coupled\. Functionally, the suppression, unification, and copy scores predict their matching activation\-level measures atr=\+0\.21r\{=\}\{\+\}0\.21,\+0\.28\+0\.28, and\+0\.59\+0\.59\. The geometry separates the signature space into two nearly independent families, a QK axis and an OV subspace, and explains why the alignment role needs its own activation\-level measure:*the additive or subtractive character of an OV map is fixed by the weights, while QK structure acts only in combination with the input; the present weight\-space alignment score does not recover the activation\-level attention pattern*\.

### 5\.2Layerwise Operator Composition

Following Section[3\.3](https://arxiv.org/html/2608.22034#S3.SS3), which interprets depth as coherence refinement, CPC expects a depth\-stratified distribution of operator roles: alignment early, unification and suppression mid\-network, routing late\. The discriminative prediction is that suppression should not follow the generic late\-peaking pattern\.

Suppression peaks in the middle third for Pythia\-1\.4B and Pythia\-2\.8B, and in the early third for Pythia\-410M, GPT\-2 Large, and both Gemma variants; the decomposition distinguishes operator timing beyond the generic late\-layer bias\. The full decision rule \(unification peaking mid\-or\-late, routing late, suppression increasing early\-to\-mid\) holds in 8 models; the remaining 7 satisfy two of three conditions\. Gemma remains the most divergent family: suppression and alignment peak in the early third while routing peaks late, suggesting early suppressive processing followed by conventional late readout\. Operators distribute non\-uniformly across depth and suppression is separable from the late\-concentration baseline, but the three\-stage partition \(align/suppress/route\) does not hold cleanly: unification and routing co\-occur in the later part, in line with content becoming readout\-aligned only late in the model, and suppression timing is architecture\-dependent\.*The three\-stage organisation of depth holds as a tendency, not a strict sequence\.*The targeted DLA variant \(Section[4](https://arxiv.org/html/2608.22034#S4)\) concentrates the correct\-token margin in the late third of both tasks \(GPT\-2 Small IOI: early0\.000\.00, late\+0\.08\+0\.08per head\), with factual margins an order of magnitude smaller and no support for an early factual margin\.

### 5\.3Causal Ablation of Operator\-Ranked Heads

The analyses above measure association, not causation\. We test if operator\-ranked heads cause the predicted behaviour with the ablation design of Section[4](https://arxiv.org/html/2608.22034#S4)\(Figure[5](https://arxiv.org/html/2608.22034#S5.F5)\)\.

Figure 5:Causal ablation of operator\-ranked heads on held\-out IOI \(mean over 5 seeds\)\. Filled dots: change in the correct\-token logit when zero\-ablating the top five heads per operator ranking \(routing ranked by the embedding\-space copy score, Equation[6](https://arxiv.org/html/2608.22034#S3.E6)\); open diamonds: layer\-matched random\-head control\. Dots left of zero indicate removed support for the correct answer\.Ablating the top unification heads reduces the correct\-token logit in 10 models, exceeding the layer\-matched control in 9 and the magnitude\-matched control in 11\. For routing, the copy\-score ranking selects a head set largely distinct from unification \(mean top\-5 overlap0\.190\.19\), and ablating it reduces the correct\-token logit in 14 models, exceeding both controls in 10\. Magnitude and direction vary by family: Gemma\-2B\-It shows an IO\-logit change of−3\.45\-3\.45for routing ablation \(control:\+0\.83\+0\.83\), while GPT\-2 Small’s unification ablation*raises*the logit \(\+1\.27\+1\.27; control:−0\.80\-0\.80\), the clearest counter\-example to the prediction, matching the backup behaviour documented for this circuit, where removing supportive heads recruits compensatory ones\([68](https://arxiv.org/html/2608.22034#bib.bib6)\)\. Its routing ablation behaves as predicted \(−0\.89\-0\.89\)\.

For suppression, the result is weaker: the ablation exceeds the layer\-matched control in only 5 models, and the predicted release of the wrong token appears in 7, with the margin moving oppositely or negligibly in the remaining 8\. The weight score thus does not reliably isolate task\-specific suppression \(see Limitations\): it measures a head’s total subtractive capacity, while which content that capacity acts on depends on the residual state\. Task\-agnostic structure does not guarantee task\-relevant suppression\. Re\-ranking suppression and unification heads by the targeted DLA margin on the discovery split restores causal reliability on the held\-out split: the margin moves in the predicted direction in all models for both rankings \(mean\+2\.07\+2\.07suppression,−1\.86\-1\.86unification\), and DLA\-ranked unification ablation reduces the correct\-token logit in every model \(mean−1\.08\-1\.08\), at the cost of requiring task labels\.*Weight\-based rankings suffice for routing and unification, while identifying task\-specific suppression requires activation evidence\.*

### 5\.4Suppression–Alignment Coupling \(P2\)

CPC treats suppression as a response to conflicts exposed through alignment: ablating upstream alignment heads should reduce the summed suppressive contribution of downstream suppressor heads \(design in Section[4](https://arxiv.org/html/2608.22034#S4)\)\.

Table 3:Suppression\-alignment coupling \(mean over 5 seeds\): change in summed suppressive contribution when ablating alignment heads on contradiction prompts \(∗: exceeds the layer\-matched random\-head control, Holm\-correctedp<0\.05p\{<\}0\.05\), the random\-head control, and the same intervention on matched no\-conflict prompts\. P2: pre\-specified verdict \(supp\./part\./agai\.\)\.Table[3](https://arxiv.org/html/2608.22034#S5.T3)reports the per\-model results\. Ablating alignment heads reduces the summed suppressive contribution on contradiction prompts in 12 models, and*the reduction exceeds the layer\-matched random\-head control with Holm\-corrected significance in 10*\(95% bootstrap CIs exclude zero\)\. The conflict\-specificity control is rarely passed: only LLaMA\-3\.2\-3B shows a reduction that is significantly weaker on matched no\-conflict prompts, and in the remaining significant models the same intervention reduces suppressive activity comparably with and without conflict\. Under the pre\-specified rule, the outcome is 1support, 9partial, and 5against\. Alignment ablation thus propagates to downstream suppressor heads in most architectures, aligned with compositional coupling through the residual stream, but the propagation reflects generic upstream dependence, not conflict\-specific signalling\. The CPC prediction survives only in its weaker, non\-selective form, with suppressor heads appearing to rely on a broad set of upstream writes instead of a dedicated conflict channel\. Conflict\-specific signalling, if present, is not separable by zero\-ablation at this granularity\.

### 5\.5Coherence Under Contradiction \(P3\)

The framework predicts that an explicit contradiction depresses the layerwise coherence proxy𝒞^link\(ℓ\)\\hat\{\\mathcal\{C\}\}\_\{\\mathrm\{link\}\}^\{\(\\ell\)\}\(Equation[7](https://arxiv.org/html/2608.22034#S3.E7)\), with the largest disruption near the layers where the model’s alignment and suppression activity concentrates \(Section[5\.2](https://arxiv.org/html/2608.22034#S5.SS2)\)\.

Figure[6](https://arxiv.org/html/2608.22034#S5.F6)shows the layerwise coherence gap for all 15 models\.*The gap is significant \(pairedtt\-test, Holm\-corrected\) in 14 models, but its sign is architecture\-dependent*: GPT\-2 and Pythia show the predicted positive gap \(magnitudes∼10−3\{\\sim\}10^\{\-3\}to10−210^\{\-2\}\), whereas Gemma, Qwen, and three of the four LLaMA variants show a comparable*negative*gap: contradictory prompts*raise*pairwise similarity\. The exception is LLaMA\-3\.2\-3B, whose gaps are small and do not survive correction\. LLaMA\-1B\-Instruct, the fourth variant, shows a positive gap and follows the GPT\-2 and Pythia pattern\. The pre\-specified co\-location rule gives 10support, 4mixed\(both Gemma variants, GPT\-2 Small, and Qwen\-1\.5B\-Instruct\), and 1against\. The directional prediction is assessed separately below\. Co\-location is exact and seed\-stable in all Pythia models: the gap and the suppression profile peak in the same third at every size and seed\.

Figure 6:Layerwise coherence gap \(consistent minus contradictory\) for all 15 models\. Lines give the mean over 5 seeds, bands the range across seeds, and colour the sign of the mean gap \(navy positive, orange negative\)\. GPT\-2 and Pythia show the predicted positive gap, Gemma, Qwen, and three of the four LLaMA variants show negative gaps, and LLaMA\-1B\-Instruct follows the positive families\.#### Controls and robustness

Three analyses qualify the interpretation \(designs in Section[4](https://arxiv.org/html/2608.22034#S4)\)\. First, lexically matched non\-contradictory pairs produce gaps one to two orders of magnitude smaller in every model: the effect is driven by the contradiction substitution\. Second, the shuffled position\-pair control reproduces the raw gap almost exactly in 14 models: the raw proxy is a global state marker of contradiction, not a measure of within\-prompt relational structure\. Third, on the whitened proxy the directional prediction holds: the gap is positive in all models and significant in 14 \(Holm\-corrected, exception LLaMA\-1B\-Instruct\)\. This attributes the raw sign flip to residual\-stream covariance\. One possible explanation is that contradiction engages a shared response component in these families, raising all pairwise similarities at once and masking the relational effect; removing the shared covariance recovers the predicted decrease in every family\. As with the signature scores, a global measure mainly reflects the family’s overall geometry, and the predicted relational effect appears only once that shared structure is removed\. The logit\-lens measure agrees: contradiction worsens the rank of the coherent continuation in all models, with the largest changes concentrated in the late third\. On the naturalistic Wikipedia set the gap is positive in 10, a weak out\-of\-template replication\. The ascent conjecture \(Equation[3](https://arxiv.org/html/2608.22034#S3.E3)\) behaves consistently: on mixed prompts the proxy’s correlation with depth is non\-negative in 14 models \(Spearmanρ\\rhoup to0\.930\.93in Qwen and LLaMA\), although it is close to zero for GPT\-2 Large and Pythia\-410M and negative for GPT\-2 Medium\.

Table 4:Base vs\. instruction\-tuned reweighting \(mean over 5 seeds\): late\-half suppression\-and\-routing signature shift, early\-half alignment shift, their ratio, the induction\-head correlation between variants, and the outcome of the one\-sided preservation test \(thresholdr=0\.9r\{=\}0\.9\)\. Verdicts follow the pre\-specified P4 rule\.

### 5\.6Base vs\. Instruction\-Tuned Models \(P4\)

CPC treats post\-training as a possible reweighting of operator structure already present in the base model \(Section[3\.6](https://arxiv.org/html/2608.22034#S3.SS6)\)\. Table[4](https://arxiv.org/html/2608.22034#S5.T4)summarises the four pairs\. Induction\-head matching passes the pre\-specified preservation test in all four pairs \(r=0\.98r\{=\}0\.98–1\.001\.00; one\-sided test againstr=0\.9r\{=\}0\.9,p<0\.05p\{<\}0\.05\):*induction\-head score structure is strongly preserved after post\-training, while a generic late concentration of the signature shift is not observed*\. Under the pre\-specified operator\-specific rule \(Section[4](https://arxiv.org/html/2608.22034#S4)\), the late suppression\-and\-routing shift exceeds the early alignment shift in both LLaMA pairs but not in Gemma or Qwen, giving twosupportand twopartialverdicts\. The operator\-specific reweighting prediction thus varies by family, while induction\-head score preservation is consistent across the four pairs\. The preservation result matches Section[3\.6](https://arxiv.org/html/2608.22034#S3.SS6), where the KL constraint keeps post\-training close to the pretrained computation: post\-training adjusts behaviour while leaving the base circuits supporting in\-context prediction intact\. Where the adjustment concentrates appears to follow each family’s depth organisation, echoing the architecture dependence seen throughout\.

### 5\.7Cross\-Task Operator Stability

If CPC operators are functional kinds, the same heads should play similar roles on different tasks \(design in Section[4](https://arxiv.org/html/2608.22034#S4)\)\. Suppression is consistently more task\-stable than unification: mean cross\-task suppressionr=0\.68r\{=\}0\.68versus unificationr=0\.45r\{=\}0\.45, with suppression exceeding0\.50\.5in 12 models and Gemma\-2B the most stable \(meanr=0\.99r\{=\}0\.99\)\. Among the models evaluated on IOI and factual recall, unification stability is higher \(meanr=0\.58r\{=\}0\.58\)\. The three\-task average underestimates cross\-task consistency where greater\-than has low accuracy\. The prediction holds fully in 6 models \(both operatorsr\>0\.5r\{\>\}0\.5\), partially in 6, and fails on both operators in 3\. The pattern suggests*suppression is a consistent, task\-general operation, while unification is more task\-dependent*\. CPC currently treats both roles as equally task\-general; the evidence supports this only for suppression\.

#### Overall pattern

Across architectures, the most stable findings concern the functional operator axes, causal effects of routing and unification, cross\-task suppression, and induction\-score preservation\. Architecture dependence is stronger in signature profile count, operator timing, raw contradiction geometry, and the location of post\-training changes\. Role\-level regularities are consequently more stable across families than their depth and geometric expression\.*Role\-level claims can be stated architecture\-independently, while depth\-level claims should be validated per family\.*

## 6Related Work

#### Mechanistic interpretability

Circuit analysis provides CPC’s empirical basis, including QK/OV decomposition\([15](https://arxiv.org/html/2608.22034#bib.bib8)\), induction heads\([51](https://arxiv.org/html/2608.22034#bib.bib1);[10](https://arxiv.org/html/2608.22034#bib.bib2)\), the IOI circuit\([68](https://arxiv.org/html/2608.22034#bib.bib6)\), copy suppression\([45](https://arxiv.org/html/2608.22034#bib.bib7)\), feedforward memories\([21](https://arxiv.org/html/2608.22034#bib.bib5)\), and analyses of greater\-than, factual recall, and component reuse\([25](https://arxiv.org/html/2608.22034#bib.bib30);[20](https://arxiv.org/html/2608.22034#bib.bib31);[46](https://arxiv.org/html/2608.22034#bib.bib32);[47](https://arxiv.org/html/2608.22034#bib.bib29)\)\. Sparse autoencoders, feature circuits, and attribution graphs link interpretable features to causal subgraphs\([33](https://arxiv.org/html/2608.22034#bib.bib10);[62](https://arxiv.org/html/2608.22034#bib.bib13);[44](https://arxiv.org/html/2608.22034#bib.bib28);[40](https://arxiv.org/html/2608.22034#bib.bib14)\)\. Path patching, automated circuit discovery, and edge\-attribution methods locate the causally relevant components on which such findings rest\([22](https://arxiv.org/html/2608.22034#bib.bib26);[11](https://arxiv.org/html/2608.22034#bib.bib25);[26](https://arxiv.org/html/2608.22034#bib.bib70)\), and the component types they identify are consistent with the operator roles CPC defines\. Closest to CPC’s aim, modular\-circuit approaches seek a global vocabulary of reusable task\-agnostic subgraphs\([28](https://arxiv.org/html/2608.22034#bib.bib67)\); CPC instead defines its vocabulary through graded functional roles tied to a coherence objective, under which heads may mix roles instead of partitioning into discrete modules\. Circuit analyses of logical reasoning find dedicated structure for propositional inference and syllogisms, including a middle\-term suppression circuit\([31](https://arxiv.org/html/2608.22034#bib.bib69);[38](https://arxiv.org/html/2608.22034#bib.bib68)\); CPC treats such contradiction handling as one instance of its suppression role\. The operators could be formalised as causal\-abstraction variables, with interchange interventions or causal scrubbing verifying proposed head\-to\-operator assignments\([19](https://arxiv.org/html/2608.22034#bib.bib27);[9](https://arxiv.org/html/2608.22034#bib.bib71)\), or induced from circuit evidence through inductive\-logic theory construction\([3](https://arxiv.org/html/2608.22034#bib.bib19)\)\.

#### Theoretical frameworks for in\-context learning and post\-training

Bayesian and algorithmic frameworks treat transformers as implicit inference machines or estimators\([72](https://arxiv.org/html/2608.22034#bib.bib15);[1](https://arxiv.org/html/2608.22034#bib.bib16);[67](https://arxiv.org/html/2608.22034#bib.bib18);[18](https://arxiv.org/html/2608.22034#bib.bib17)\); CPC differs by naming the component\-level functions that these frameworks abstract away: alignment, unification, suppression, and routing\. Preference\-training objectives change model behaviour\([52](https://arxiv.org/html/2608.22034#bib.bib21);[4](https://arxiv.org/html/2608.22034#bib.bib22);[59](https://arxiv.org/html/2608.22034#bib.bib33)\), and mechanistic studies suggest that fine\-tuning can enhance or reweight existing circuits\([56](https://arxiv.org/html/2608.22034#bib.bib34);[34](https://arxiv.org/html/2608.22034#bib.bib35);[60](https://arxiv.org/html/2608.22034#bib.bib36)\), with instruction\-conditioned suppression\([69](https://arxiv.org/html/2608.22034#bib.bib23)\)providing a circuit\-level example for CPC’s reweighting interpretation\.[6](https://arxiv.org/html/2608.22034#bib.bib24)combines symbolic\-like and continuous computation, compatible with parts of CPC but without a central role for contradiction suppression\.

## 7Conclusion

We proposed Coherentist Probabilistic Compositionalism \(CPC\), an interpretive framework that describes transformer computation through four operator roles involved in coherence construction and readout, and evaluated its predictions on 15 models from five architecture families\. The vocabulary identifies axes of head variation that hold across architectures, while the number of separable signature profiles, the depth at which suppression peaks, and the way contradictions register all depend on the family\. CPC provides a shared vocabulary for transformer mechanisms and post\-training effects, with predictions that should be stated conditionally on architecture\. Verifying the operator assignments through causal abstraction and encoding the task\-generality asymmetry between suppression and unification are natural next steps\.

## Limitations

CPC is an interpretive framework, not a mechanistic proof: the four operators are not claimed to be exhaustive, the treatment covers decoder\-only English models, and the predicted failure regimes \(ambiguity overload, depth overload, preference distortion\) remain untested\. The main measurement caveat concerns the coherence proxy, which conflates semantic coherence with residual\-stream geometry: the raw gap’s sign reverses between families and is reproduced by cross\-prompt position pairs, so it is a global state marker of contradiction, not a measure of within\-prompt relational structure \(Section[5\.5](https://arxiv.org/html/2608.22034#S5.SS5)\)\. Whitening and the logit\-lens trajectory provide convergent directional support, but a proxy that isolates relational coherence pair\-specifically remains future work, and the ascent conjecture \(Equation[3](https://arxiv.org/html/2608.22034#S3.E3)\) is weak or absent in the GPT\-2 family and Pythia\-410M\.

The suppression evidence is the weakest\. The weight score measures total subtractive OV mode weight, a property of the matrix, and does not reliably isolate task\-specific suppression \(layer\-matched control exceeded in 5/15 models\); task\-targeted DLA rankings restore causal reliability \(Section[5\.3](https://arxiv.org/html/2608.22034#S5.SS3)\) but require task labels\. The evaluation rests on synthetic tasks with a single naturalistic negation set \(gap positive in 10/15 models\), greater\-than covers only the six GPT\-2 and Pythia models for tokeniser reasons, and the base\-vs\-tuned comparison covers four pairs\. The coupling test \(P2\) uses the validated IOI circuit only for GPT\-2 Small, with automatically detected heads elsewhere, and its conflict\-specificity control passed only in LLaMA\-3\.2\-3B, so the P2 verdicts rest on the random\-head control; other operationalisations of “the suppressed alternative” remain untested\.

## Ethics Statement

CPC is a mechanistic interpretive framework\. Its predictions concern internal representations, not system deployment\. We do not anticipate direct misuse risks from this analytical vocabulary; however, improved understanding of suppression and instruction\-following circuits may inform both the design and circumvention of safety mechanisms\.

## Appendix A: The CPC Functional and Coherentism

The CPC functional in Equation[2](https://arxiv.org/html/2608.22034#S3.E2)adapts the constraint\-based view of Section[2\.2](https://arxiv.org/html/2608.22034#S2.SS2)to transformer computation\. Grounded fragments\(m,g\)\(m,g\)play the role of elements in the constraint formulation, while the relational structureGGspecifies how these fragments can be connected\. The local score𝒮coh​\(g∣m,G\)\\mathcal\{S\}\_\{\\text\{coh\}\}\(g\\mid m,G\)rewards a grounding that fits the developing interpretation, corresponding to positive compatibility between elements\. The kernelκ\\kappapenalises incompatible grounded fragments, corresponding to negative constraints\. The additional coverage term rewards interpretations that incorporate more of the available input, preventing coherence from being increased simply by retaining a small set of mutually compatible fragments\.

Three differences separate CPC from the discrete constraint\-satisfaction formulation\. First, support is graded: fragments receive continuous compatibility scores instead of belonging only to accepted or rejected sets\. Second, the problem is incremental\. New linguistic material is introduced token by token, so the interpretation must be updated as the context grows\. Third, the optimisation is implicit\. CPC does not claim that a transformer explicitly represents Equation[2](https://arxiv.org/html/2608.22034#S3.E2)or solves a constraint\-satisfaction problem\. The stronger claim is instead formulated as the coherence\-ascent conjecture of Section[3\.3](https://arxiv.org/html/2608.22034#S3.SS3): computation across depth can be interpreted as repeated refinement of the current representational state towards configurations with greater coherence \(see Figure[3](https://arxiv.org/html/2608.22034#S3.F3)\)\.

The four CPC operators describe functional contributions to this refinement\.*Alignment*identifies candidate relations between fragments\.*Unification*incorporates mutually compatible information into the developing representation\.*Suppression*reduces support for incompatible alternatives, paralleling inhibitory competition in the constraint\-satisfaction formulation\.*Routing*then carries the resulting information to positions where it can affect the readout\. The correspondence is functional, not an identification of linguistic coherence theories with transformer mechanisms\.

This connection also clarifies the status of the empirical coherence proxy in Equation[7](https://arxiv.org/html/2608.22034#S3.E7)\. Similarity between internal representations measures one possible consequence of compatible information becoming more closely represented, and is related to semantic\-similarity approaches to discourse coherence\. It does not measure the complete functional in Equation[2](https://arxiv.org/html/2608.22034#S3.E2)\. In particular, it does not directly identify discourse relations, entity continuity, or incompatibility between specific groundings\. For this reason, CPC treats the proxy as an empirical diagnostic of representational coherence, while the coherence state and functional provide the broader theoretical construct\.

The distinction between measurement and construct in this account parallels the levels of description in physical theory\. Circuit\-level measurements are analogous to instrument readings, which sit close to the observable\. The operator labels and the functional are analogous to constructs such as energy and entropy, which organise many observations at a higher level of description and are not read from any single measurement\.

## Appendix B: Formal Definitions

A fragmentm=\(U,F,C\)m\{=\}\(U,F,C\)specifies a position setU⊆\{1,…,T\}U\\subseteq\\\{1,\\ldots,T\\\}, a content labelFFfrom a finite vocabulary of semantic and syntactic types, and compositional constraintsCC\. A grounding is a partial functiong:U⇀Vg:U\\rightharpoonup Vinto the relational graphG=\(V,E\)G\{=\}\(V,E\), admissible when it satisfiesCC\. Two grounded fragments are compatible when their assignments agree on all shared positions and jointly satisfy both constraint sets\. The contradiction kernelκ≥0\\kappa\\geq 0is symmetric withκ=0\\kappa\{=\}0on compatible pairs\. The local score𝒮coh​\(g∣m,G\)∈\(0,1\]\\mathcal\{S\}\_\{\\text\{coh\}\}\(g\\mid m,G\)\\in\(0,1\]is a content\-compatibility factor times the geometric mean of pairwise relational compatibilitiesψ⁡\(g⁡\(u\),g⁡\(u′\)\)\\psi\\bigl\(g\(u\),g\(u^\{\\prime\}\)\\bigr\); for nodes with activation vectors,ψ=ϵψ\+\(1−ϵψ\)​\(1\+cos⁡\(rv,rv′\)\)/2\\psi=\\epsilon\_\{\\psi\}\+\(1\-\\epsilon\_\{\\psi\}\)\\bigl\(1\+\\cos\(r\_\{v\},r\_\{v^\{\\prime\}\}\)\\bigr\)/2, bounded in\(0,1\]\(0,1\]\.

## Appendix C: Task Definitions

The three tasks of Section[4](https://arxiv.org/html/2608.22034#S4)share a next\-token format: the context is a token sequencec=\{tok1,tok2,…,tokn\}c=\\\{\\mathrm\{tok\}\_\{1\},\\mathrm\{tok\}\_\{2\},\\dots,\\mathrm\{tok\}\_\{n\}\\\}, the model predicts its continuation, and an example enters the experiments only when the prediction is correct under the task’s criterion\. Multi\-token targets are scored on their first token\.

#### Indirect Object Identification \(IOI\)

This task requires the model to predict the indirect object of a transfer clause, given a context that introduces two names and repeats one of them as the subject, as inWhen␣Mary␣and␣John␣went␣to␣the␣store,␣John␣gave␣a␣drink␣to\. Formally, the contextccintroduces namesAAandBBand repeatsBBas the subject of the final clause, and the model seeks to produce the non\-repeated nameAAsuch that

A=arg⁡maxt∈𝒱⁡P⁡\(t∣c\),A=\\arg\\max\_\{t\\in\\mathcal\{V\}\}P\(t\\mid c\),\(C\.1\)where𝒱\\mathcal\{V\}denotes the model’s vocabulary\. Predictions are deemed correct only if the predicted token matches the first token ofAA\. The repeated nameBBprovides the wrong\-token contrast for the logit margin used in the ablation experiments \(Section[4](https://arxiv.org/html/2608.22034#S4)\)\. Prompts instantiate 15 template frames over a pool of roughly 100 names, with places and objects varied\.

#### Greater\-Than Comparison

This task requires the model to complete a year interval consistently, given a context that states a start year and the century prefix of the end year, as inThe␣war␣lasted␣from␣the␣year␣1732␣to␣the␣year␣17\. WithYY\\mathrm\{YY\}the two\-digit suffix of the start year, the model seeks a two\-digit continuationy^\\hat\{y\}such that

y^=arg⁡maxt∈𝒱⁡P⁡\(t∣c\),YY<val⁡\(y^\)≤99,\\hat\{y\}=\\arg\\max\_\{t\\in\\mathcal\{V\}\}P\(t\\mid c\),\\qquad\\mathrm\{YY\}<\\operatorname\{val\}\(\\hat\{y\}\)\\leq 99,\(C\.2\)whereval⁡\(⋅\)\\operatorname\{val\}\(\\cdot\)reads a two\-digit token as a number\. Predictions are deemed correct only if the predicted token is a single\-token two\-digit year satisfying the inequality; this requirement excludes nine models from the task \(Section[4](https://arxiv.org/html/2608.22034#S4)\)\. Examples are drawn from an exhaustive noun\-by\-year grid over the years 1102–1898, excluding suffixes 00 and 99\.

#### Factual Recall

This task requires the model to predict the object of a subject\-relation pair, given a natural\-language prompt that expresses the subject and the relation, as inThe␣Eiffel␣Tower␣is␣located␣in␣the␣city␣of\. Formally, the model seeks to produce the gold objectoosuch that

o=arg⁡maxt∈𝒱⁡P⁡\(t∣c\)\.o=\\arg\\max\_\{t\\in\\mathcal\{V\}\}P\(t\\mid c\)\.\(C\.3\)Prompts and gold objects come from the CounterFact set\([46](https://arxiv.org/html/2608.22034#bib.bib32)\), and predictions are deemed correct only if the predicted token matches the first token ofoo\.

## Appendix D: Parameter and Threshold Choices

The constants below were fixed in the experiment plan before the confirmatory runs, and none was tuned against held\-out outcomes\. Choices fall into three classes: values preregistered with an explicit rationale, values matched to validated reference circuits, and conventional defaults\. Every stochastic component \(clustering, null draws, bootstraps, control\-head samples\) runs at seeds\{0,…,4\}\\\{0,\\ldots,4\\\}, with means and ranges reported across seeds\.

#### Prompt counts and splits

The target of 500 model\-correct examples per task balances stable per\-head averages and paired tests against the cost of running 15 models with 5 seeds under interventions; candidate prompts are oversampled fourfold before correctness filtering\. The 50/50 discovery/held\-out split gives selection and verdict computation equal power, and hashing at generation time fixes the assignment so no example migrates between splits\. The contradiction pool of 3,854 pairs exhausts the unique fills of the generation templates; 500 pairs per run matches the per\-task target\.

#### Weight\-space analysis

The cluster rangek∈\{2,…,7\}k\\in\\\{2,\\ldots,7\\\}brackets the hypothesised four\-role organisation from both sides, withk=2k\{=\}2the minimal nontrivial split, letting the silhouette peak fall below, at, or above four instead of being forced towards it\. Fiftykk\-means restarts guard against local minima, and assignments are stable across the five seeds\. The two silhouette nulls use 200 draws each, giving app\-value resolution of1/2011/201, finer than the reporting level ofp<0\.01p\{<\}0\.01; the 100 random weight projections give the 95% baseline quantile a resolution of one draw\.

#### Activation measures

DLA is computed on the model’s top\-10 predicted tokens so that attribution is restricted to the model’s own candidate set, where logit contributions bear on the prediction; widening the set dilutes the measure over tokens the model never considers\. All layerwise quantities are computed at every layer; the layer third is the unit of summary and testing, not of measurement \(per\-layer profiles appear in Figure[6](https://arxiv.org/html/2608.22034#S5.F6)\)\. Testing at thirds serves three purposes\. The CPC predictions are phase\-level claims \(early alignment, mid\-network unification and suppression, late routing\); the test resolution matches the claim\. Thirds are comparable across architectures whose depths range from 12 to 36 layers, where individual layer indices do not align\. Three tests per model also keep the Holm family small, whereas per\-layer testing would inflate it by an order of magnitude and reduce power\. The whitening shrinkage ofα=0\.1\\alpha\{=\}0\.1addresses the rank\-deficient per\-batch covariance \(positions are far fewer than dimensions\); the value is a fixed conventional intensity, and the eigenvalue floor guards numerical stability\.

#### Head\-set sizes

The detection sizes for P2 mirror the validated GPT\-2 Small IOI sets: five alignment heads \(three duplicate\-token plus two previous\-token\) and six suppressors \(four S\-inhibition plus two negative name\-movers\), so that automatic detection selects sets of the same size as the reference circuit\. The causal ablations remove the top five heads per operator, applied uniformly: the set is comparable to identified circuit sizes and remains a small fraction of the head budget of even the smallest model \(five of GPT\-2 Small’s 144 heads\)\. The ablation\-effect correlation uses the union of the top\-10 heads per operator because a correlation over all heads would be dominated by the large bulk of heads with near\-zero scores and near\-zero effects\.

#### Statistical thresholds

Holm correction atα=0\.05\\alpha\{=\}0\.05within each experiment family follows convention\. Permutation tests and bootstrap intervals use 10,000 draws, placing the Monte Carlo standard error near0\.0020\.002atp≈0\.05p\\approx 0\.05\. The preservation marginr=0\.9r\{=\}0\.9was preregistered and uses a pre\-specified preservation threshold ofr=0\.9r=0\.9: shared variance must exceed0\.810\.81, so preservation is never inferred from a non\-significant difference\. The cross\-task stability thresholdr\>0\.5r\{\>\}0\.5is a conventional boundary for a moderate\-to\-strong correlation, fixed before analysis\.

## Appendix E: Contradiction Prompt Templates

Contradiction prompts are drawn from a pool of 3,854 globally unique consistent/contradictory pairs, generated once \(seed 42\)\. Each pair shares identical surface structure up to a key substitution that introduces or removes a contradiction\. Pool 1 \(3,333 pairs\) instantiates four named\-agent IOI\-style frame pairs over a 100\-name pool with place and object slots\. Pool 2 \(221 pairs\) covers seven contradiction types \(factual, temporal, causal, logical, narrative, spatial, quantitative\) via parameterised generators\. Pool 3 \(300 pairs\) crosses 20 agent\-competence frames with 15 role pairs, without name placeholders\. A representative pool\-1 pair contrasts “When*A*and*B*went to the store,*A*stayed in the store, and*A*gave a drink to” with the same frame containing “*A*immediately left the store”; the other pools follow the same substitution pattern\. A separate naturalistic set pairs Wikipedia lead sentences with pronoun\-coreferent continuations, negating the first sentence’s copula in the contradictory variant; it provides the out\-of\-template replication of Section[5\.5](https://arxiv.org/html/2608.22034#S5.SS5)\.

###### Acknowledgements\.

This is an example of acknowlwdgment\. This is a sample paper presented jost for the coding of different elements of a document inLaTeXusing this package file\. We thank all readers\.

## References

- Akyüreket al\.\(2023\)E\. Akyürek, D\. Schuurmans, J\. Andreas, T\. Ma, and D\. ZhouWhat learning algorithm is in\-context learning? investigations with linear models\.InThe Eleventh International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Aljaafariet al\.\(2026a\)N\. Aljaafari, D\. Carvalho, and A\. FreitasEmergence and localisation of semantic role circuits in LLMs\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 39402–39433\.External Links:[Link](https://aclanthology.org/2026.findings-acl.1964/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1964),ISBN 979\-8\-89176\-395\-1Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1)\.
- Aljaafariet al\.\(2026b\)N\. Aljaafari, D\. S\. Carvalho, and A\. FreitasFrom circuit evidence to mechanistic theory: an inductive logic approach\.arXiv preprint arXiv:2605\.21303\.External Links:2605\.21303Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Baiet al\.\(2022\)Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon, C\. Chen, C\. Olsson, C\. Olah, D\. Hernandez, D\. Drain, D\. Ganguli, D\. Li, E\. Tran\-Johnson, E\. Perez, J\. Kerr, J\. Mueller, J\. Ladish, J\. Landau, K\. Ndousse, K\. Lukosuite, L\. Lovitt, M\. Sellitto, N\. Elhage, N\. Schiefer, N\. Mercado, N\. DasSarma, R\. Lasenby, R\. Larson, S\. Ringer, S\. Johnston, S\. Kravec, S\. E\. Showk, S\. Fort, T\. Lanham, T\. Telleen\-Lawton, T\. Conerly, T\. Henighan, T\. Hume, S\. R\. Bowman, Z\. Hatfield\-Dodds, B\. Mann, D\. Amodei, N\. Joseph, S\. McCandlish, T\. Brown, and J\. KaplanConstitutional ai: harmlessness from ai feedback\.arXiv preprint arXiv:2212\.08073\.External Links:[Document](https://dx.doi.org/2212.08073v1),[Link](https://arxiv.org/abs/2212.08073v1)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§3\.6](https://arxiv.org/html/2608.22034#S3.SS6.p1.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Bidermanet al\.\(2023\)S\. Biderman, H\. Schoelkopf, Q\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. Van Der WalPythia: a suite for analyzing large language models across training and scaling\.InProceedings of the 40th International Conference on Machine Learning,ICML’23\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px1.p1.1)\.
- Boleda \(2025\)G\. BoledaLLMs as a synthesis between symbolic and distributed approaches to language\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 9365–9379\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.498/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.498),ISBN 979\-8\-89176\-335\-7Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- BonJour \(1985\)L\. BonJourThe structure of empirical knowledge\.Harvard University Press,Cambridge, MA\.Cited by:[§2](https://arxiv.org/html/2608.22034#S2.p1.1),[Figure 3](https://arxiv.org/html/2608.22034#S3.F3)\.
- Brickenet al\.\(2023\)T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. L\. Turner, C\. Anil, C\. Denison, A\. Askell,et al\.Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.html](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1)\.
- Chanet al\.\(2022\)L\. Chan, A\. Garriga\-Alonso, N\. Goldowsky\-Dill, R\. Greenblatt, J\. Nitishinskaya, A\. Radhakrishnan, B\. Shlegeris, and N\. ThomasCausal scrubbing: a method for rigorously testing interpretability hypotheses\.Note:AI Alignment Forum[https://www\.alignmentforum\.org/posts/JvZhhzycHu2Yd57RN/causal\-scrubbing\-a\-method\-for\-rigorously\-testing](https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2024\)S\. Chen, H\. Sheen, T\. Wang, and Z\. YangUnveiling induction heads: provable training dynamics and feature learning in transformers\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 66479–66567\.Note:[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/7aae9e3ec211249e05bd07271a6b1441\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/7aae9e3ec211249e05bd07271a6b1441-Paper-Conference.pdf)External Links:[Document](https://dx.doi.org/10.52202/079017-2127)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Conmyet al\.\(2023\)A\. Conmy, A\. N\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-AlonsoTowards automated circuit discovery for mechanistic interpretability\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Dubeyet al\.\(2024\)A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan, and et al\.The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px1.p1.1)\.
- Efron \(1979\)B\. EfronBootstrap methods: another look at the jackknife\.The Annals of Statistics7\(1\),pp\. 1–26\.External Links:ISSN 00905364, 21688966,[Link](http://www.jstor.org/stable/2958830)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px5.p1.1)\.
- Elhageet al\.\(2022\)N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. OlahToy models of superposition\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2022/toy\_model/index\.html](https://transformer-circuits.pub/2022/toy_model/index.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1)\.
- Elhageet al\.\(2021\)N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. OlahA mathematical framework for transformer circuits\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2021/framework/index\.html](https://transformer-circuits.pub/2021/framework/index.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§3\.2](https://arxiv.org/html/2608.22034#S3.SS2.p1.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.SSS0.Px1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Fillmore \(1982\)C\. J\. FillmoreFrame semantics\.InLinguistics in the Morning Calm,The Linguistic Society of Korea \(Ed\.\),pp\. 111–137\.Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1)\.
- Fisher \(1915\)R\. A\. FisherFrequency distribution of the values of the correlation coefficient in samples from an indefinitely large population\.Biometrika10\(4\),pp\. 507–521\.External Links:ISSN 00063444,[Link](http://www.jstor.org/stable/2331838)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px7.p1.1)\.
- Garget al\.\(2022\)S\. Garg, D\. Tsipras, P\. S\. Liang, and G\. ValiantWhat can transformers learn in\-context? a case study of simple function classes\.Advances in neural information processing systems35,pp\. 30583–30598\.Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Geigeret al\.\(2025\)A\. Geiger, D\. Ibeling, A\. Zur, M\. Chaudhary, S\. Chauhan, J\. Huang, A\. Arora, Z\. Wu, N\. Goodman, C\. Potts, and T\. IcardCausal abstraction: a theoretical foundation for mechanistic interpretability\.Journal of Machine Learning Research26\(83\),pp\. 1–54\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Gevaet al\.\(2023\)M\. Geva, J\. Bastings, K\. Filippova, and A\. GlobersonDissecting recall of factual associations in auto\-regressive language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 12216–12235\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.751/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.751)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Gevaet al\.\(2021\)M\. Geva, R\. Schuster, J\. Berant, and O\. LevyTransformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 5484–5495\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.446/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.446)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Goldowsky\-Dillet al\.\(2023\)N\. Goldowsky\-Dill, C\. MacLeod, L\. Sato, and A\. AroraLocalizing model behavior with path patching\.arXiv preprint arXiv:2304\.05969\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Groszet al\.\(1995\)B\. J\. Grosz, A\. K\. Joshi, and S\. WeinsteinCentering: a framework for modeling the local coherence of discourse\.Computational Linguistics21\(2\),pp\. 203–225\.External Links:[Link](https://aclanthology.org/J95-2003/)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p3.1)\.
- Halliday and Hasan \(1976\)M\. A\. K\. Halliday and R\. HasanCohesion in english\.English Language Series,Longman,London\.External Links:ISBN 9780582550414Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p3.1)\.
- Hannaet al\.\(2023\)M\. Hanna, O\. Liu, and A\. VariengienHow does gpt\-2 compute greater\-than? interpreting mathematical abilities in a pre\-trained language model\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Hannaet al\.\(2024\)M\. Hanna, S\. Pezzelle, and Y\. BelinkovHave faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.InFirst Conference on Language Modeling,Note:[https://openreview\.net/forum?id=TZ0CCGDcuT](https://openreview.net/forum?id=TZ0CCGDcuT)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Harriset al\.\(2020\)C\. R\. Harris, K\. J\. Millman, S\. J\. van der Walt, R\. Gommers, P\. Virtanen, D\. Cournapeau, E\. Wieser, J\. Taylor, S\. Berg, N\. J\. Smith, R\. Kern, M\. Picus, S\. Hoyer, M\. H\. van Kerkwijk, M\. Brett, A\. Haldane, J\. F\. del Río, M\. Wiebe, P\. Peterson, P\. Gérard\-Marchant, K\. Sheppard, T\. Reddy, W\. Weckesser, H\. Abbasi, C\. Gohlke, and T\. E\. OliphantArray programming with numpy\.Nature585,pp\. 357–362\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Heet al\.\(2025\)Y\. He, W\. Zheng, Y\. Dong, Y\. Zhu, C\. Chen, and J\. LiTowards global\-level mechanistic interpretability: a perspective of modular circuits of large language models\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 22865–22880\.External Links:[Link](https://proceedings.mlr.press/v267/he25x.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Hobbs \(1979\)J\. R\. HobbsCoherence and coreference\.Cognitive Science3\(1\),pp\. 67–90\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1207/s15516709cog0301%5F4),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1207/s15516709cog0301_4),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1207/s15516709cog0301\_4Cited by:[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p1.1),[§2](https://arxiv.org/html/2608.22034#S2.p1.1)\.
- Holm \(1979\)S\. HolmA simple sequentially rejective multiple test procedure\.Scandinavian Journal of Statistics6\(2\),pp\. 65–70\.External Links:ISSN 03036898, 14679469,[Link](http://www.jstor.org/stable/4615733)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px5.p1.1)\.
- Honget al\.\(2025\)G\. Z\. Hong, N\. Dikkala, E\. Luo, C\. Rashtchian, X\. Wang, and R\. PanigrahyA implies b: circuit analysis in LLMs for propositional logical reasoning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,Note:[https://openreview\.net/forum?id=M0U8wUow8c](https://openreview.net/forum?id=M0U8wUow8c)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Huet al\.\(2020\)J\. Hu, J\. Gauthier, P\. Qian, E\. Wilcox, and R\. P\. LevyA systematic assessment of syntactic generalization in neural language models\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 1725–1744\.External Links:[Link](https://aclanthology.org/2020.acl-main.158/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.158)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1)\.
- Hubenet al\.\(2024\)R\. Huben, H\. Cunningham, L\. Smith, A\. Ewart, and L\. SharkeySparse autoencoders find highly interpretable features in language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 7827–7845\.Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Jainet al\.\(2024\)S\. Jain, R\. Kirk, E\. S\. Lubana, R\. P\. Dick, H\. Tanaka, E\. Grefenstette, T\. Rocktäschel, and D\. KruegerMechanistically analyzing the effects of fine\-tuning on procedurally defined tasks\.InICLR 2024 Workshop on Secure and Trustworthy Large Language Models,External Links:[Link](https://openreview.net/forum?id=bWimc91mtK)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Jurafsky and Martin \(2026\)D\. Jurafsky and J\. H\. MartinSpeech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition with language models\.3rd edition\.Note:Online manuscript released August 19, 2026External Links:[Link](https://web.stanford.edu/%CB%9Cjurafsky/slp3/)Cited by:[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p2.1),[§2](https://arxiv.org/html/2608.22034#S2.p1.1)\.
- Kaufman and Rousseeuw \(1990\)L\. Kaufman and P\. RousseeuwFinding groups in data: an introduction to cluster analysis\.External Links:ISBN 0\-471\-87876\-6,[Document](https://dx.doi.org/10.2307/2532178)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px4.p1.1)\.
- Kessyet al\.\(2018\)A\. Kessy, A\. Lewin, and K\. StrimmerOptimal whitening and decorrelation\.The American Statistician72\(4\),pp\. 309–314\.External Links:ISSN 1537\-2731,[Link](http://dx.doi.org/10.1080/00031305.2016.1277159),[Document](https://dx.doi.org/10.1080/00031305.2016.1277159)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px6.p1.1)\.
- Kimet al\.\(2025\)G\. Kim, M\. Valentino, and A\. FreitasReasoning circuits in language models: a mechanistic interpretation of syllogistic inference\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 10074–10095\.External Links:[Link](https://aclanthology.org/2025.findings-acl.525/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.525),ISBN 979\-8\-89176\-256\-5Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Li and Hovy \(2014\)J\. Li and E\. HovyA model of coherence based on distributed sentence representation\.InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),A\. Moschitti, B\. Pang, and W\. Daelemans \(Eds\.\),Doha, Qatar,pp\. 2039–2048\.External Links:[Link](https://aclanthology.org/D14-1218/),[Document](https://dx.doi.org/10.3115/v1/D14-1218)Cited by:[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p3.1)\.
- Lindseyet al\.\(2025\)J\. Lindsey, W\. Gurnee, E\. Ameisen, B\. Chen, A\. Pearce, N\. L\. Turner, C\. Citro, D\. Abrahams, S\. Carter, B\. Hosmer, J\. Marcus, M\. Sklar, A\. Templeton, T\. Bricken, C\. McDougall, H\. Cunningham, T\. Henighan, A\. Jermyn, A\. Jones, A\. Persic, Z\. Qi, T\. B\. Thompson, S\. Zimmerman, K\. Rivoire, T\. Conerly, C\. Olah, and J\. BatsonOn the biology of a large language model\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2025/attribution\-graphs/biology\.html](https://transformer-circuits.pub/2025/attribution-graphs/biology.html)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- MacQueen \(1967\)J\. B\. MacQueenSome methods for classification and analysis of multivariate observations\.InProceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability,Vol\.1,pp\. 281–297\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px4.p1.1)\.
- Mann and Thompson \(1988\)W\. C\. Mann and S\. A\. ThompsonRhetorical structure theory: toward a functional theory of text organization\.Text \- Interdisciplinary Journal for the Study of Discourse8\(3\),pp\. 243–281\.External Links:[Link](https://doi.org/10.1515/text.1.1988.8.3.243),[Document](https://dx.doi.org/doi%3A10.1515/text.1.1988.8.3.243)Cited by:[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.22034#S2.SS1.p4.1)\.
- Manninget al\.\(2020\)C\. D\. Manning, K\. Clark, J\. Hewitt, U\. Khandelwal, and O\. LevyEmergent linguistic structure in artificial neural networks trained by self\-supervision\.Proceedings of the National Academy of Sciences117\(48\),pp\. 30046–30054\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1907367117),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.1907367117),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.1907367117Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1)\.
- Markset al\.\(2025\)S\. Marks, C\. Rager, E\. J\. Michaud, Y\. Belinkov, D\. Bau, and A\. MuellerSparse feature circuits: discovering and editing interpretable causal graphs in language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=I4e82CIDxv)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- McDougallet al\.\(2024\)C\. S\. McDougall, A\. Conmy, C\. Rushing, T\. McGrath, and N\. NandaCopy suppression: comprehensively understanding a motif in language model attention heads\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 337–363\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.22/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.22)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.SSS0.Px1.p2.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Menget al\.\(2022\)K\. Meng, D\. Bau, A\. Andonian, and Y\. BelinkovLocating and editing factual associations in gpt\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1),[Factual Recall](https://arxiv.org/html/2608.22034#Sx5.SS0.SSS0.Px3.p1.2)\.
- Merulloet al\.\(2024\)J\. Merullo, C\. Eickhoff, and E\. PavlickCircuit component reuse across tasks in transformer language models\.InThe Twelfth International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=fpoAYV6Wsk)Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Nanda and Bloom \(2022\)N\. Nanda and J\. BloomTransformerLens\.Note:[https://github\.com/TransformerLensOrg/TransformerLens](https://github.com/TransformerLensOrg/TransformerLens)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Nandaet al\.\(2023\)N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. SteinhardtProgress measures for grokking via mechanistic interpretability\.InThe Eleventh International Conference on Learning Representations,Note:[https://openreview\.net/forum?id=9XFSbDPmdW](https://openreview.net/forum?id=9XFSbDPmdW)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting GPT: the logit lens\.Note:[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)LessWrongCited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px6.p1.1)\.
- Olssonet al\.\(2022\)C\. Olsson, N\. Elhage, N\. Nanda, N\. Joseph, N\. DasSarma, T\. Henighan, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, S\. Johnston, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. OlahIn\-context learning and induction heads\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2022/in\-context\-learning\-and\-induction\-heads/index\.html](https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§3\.6](https://arxiv.org/html/2608.22034#S3.SS6.p1.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Paszkeet al\.\(2019\)A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga, A\. Desmaison, A\. Kopf, E\. Yang, Z\. DeVito, M\. Raison, A\. Tejani, S\. Chilamkurthy, B\. Steiner, L\. Fang, J\. Bai, and S\. ChintalaPyTorch: an imperative style, high\-performance deep learning library\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Pearson \(1901\)K\. PearsonOn lines and planes of closest fit to systems of points in space\.The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science2\(11\),pp\. 559–572\.External Links:[Document](https://dx.doi.org/10.1080/14786440109462720)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px4.p1.1)\.
- Pedregosaet al\.\(2011\)F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and É\. DuchesnayScikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Prakashet al\.\(2024\)N\. Prakash, T\. Shaham, T\. Haklay, Y\. Belinkov, and D\. BauFine\-tuning enhances existing mechanisms: a case study on entity tracking\.InInternational Conference on Learning Representations \(ICRL\),Vol\.2024,pp\. 8057–8082\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Qwenet al\.\(2025\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px1.p1.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. SutskeverLanguage models are unsupervised multitask learners\.OpenAI Blog\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px1.p1.1)\.
- Rafailovet al\.\(2023\)R\. Rafailov, A\. Sharma, E\. Mitchell, S\. Ermon, C\. D\. Manning, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Ruscioet al\.\(2026\)V\. Ruscio, E\. Khedouri, and K\. ThompsonWhere pretraining writes and alignment reads: the asymmetry of transformer weight space\.arXiv preprint arXiv:2605\.16600\.Cited by:[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Teamet al\.\(2024\)G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé, J\. Ferret, P\. Liu, P\. Tafti, A\. Friesen, M\. Casbon, S\. Ramos, R\. Kumar, C\. L\. Lan, S\. Jerome, A\. Tsitsulin, N\. Vieillard, P\. Stanczyk, S\. Girgin, N\. Momchev, M\. Hoffman, S\. Thakoor, J\. Grill, B\. Neyshabur, O\. Bachem, A\. Walton, A\. Severyn, A\. Parrish, A\. Ahmad, A\. Hutchison, A\. Abdagic, A\. Carl, A\. Shen, A\. Brock, A\. Coenen, A\. Laforge, A\. Paterson, B\. Bastian, B\. Piot, B\. Wu, B\. Royal, C\. Chen, C\. Kumar, C\. Perry, C\. Welty, C\. A\. Choquette\-Choo, D\. Sinopalnikov, D\. Weinberger, D\. Vijaykumar, D\. Rogozińska, D\. Herbison, E\. Bandy, E\. Wang, E\. Noland, E\. Moreira, E\. Senter, E\. Eltyshev, F\. Visin, G\. Rasskin, G\. Wei, G\. Cameron, G\. Martins, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Batra, H\. Dhand, I\. Nardini, J\. Mein, J\. Zhou, J\. Svensson, J\. Stanway, J\. Chan, J\. P\. Zhou, J\. Carrasqueira, J\. Iljazi, J\. Becker, J\. Fernandez, J\. van Amersfoort, J\. Gordon, J\. Lipschultz, J\. Newlan, J\. Ji, K\. Mohamed, K\. Badola, K\. Black, K\. Millican, K\. McDonell, K\. Nguyen, K\. Sodhia, K\. Greene, L\. L\. Sjoesund, L\. Usui, L\. Sifre, L\. Heuermann, L\. Lago, L\. McNealus, L\. B\. Soares, L\. Kilpatrick, L\. Dixon, L\. Martins, M\. Reid, M\. Singh, M\. Iverson, M\. Görner, M\. Velloso, M\. Wirth, M\. Davidow, M\. Miller, M\. Rahtz, M\. Watson, M\. Risdal, M\. Kazemi, M\. Moynihan, M\. Zhang, M\. Kahng, M\. Park, M\. Rahman, M\. Khatwani, N\. Dao, N\. Bardoliwalla, N\. Devanathan, N\. Dumai, N\. Chauhan, O\. Wahltinez, P\. Botarda, P\. Barnes, P\. Barham, P\. Michel, P\. Jin, P\. Georgiev, P\. Culliton, P\. Kuppala, R\. Comanescu, R\. Merhej, R\. Jana, R\. A\. Rokni, R\. Agarwal, R\. Mullins, S\. Saadat, S\. M\. Carthy, S\. Cogan, S\. Perrin, S\. M\. R\. Arnold, S\. Krause, S\. Dai, S\. Garg, S\. Sheth, S\. Ronstrom, S\. Chan, T\. Jordan, T\. Yu, T\. Eccles, T\. Hennigan, T\. Kocisky, T\. Doshi, V\. Jain, V\. Yadav, V\. Meshram, V\. Dharmadhikari, W\. Barkley, W\. Wei, W\. Ye, W\. Han, W\. Kwon, X\. Xu, Z\. Shen, Z\. Gong, Z\. Wei, V\. Cotruta, P\. Kirk, A\. Rao, M\. Giang, L\. Peran, T\. Warkentin, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, D\. Sculley, J\. Banks, A\. Dragan, S\. Petrov, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, S\. Borgeaud, N\. Fiedel, A\. Joulin, K\. Kenealy, R\. Dadashi, and A\. AndreevGemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Link](https://arxiv.org/abs/2408.00118)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px1.p1.1)\.
- Templetonet al\.\(2024\)A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, A\. Tamkin, E\. Durmus, T\. Hume, F\. Mosconi, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. HenighanScaling monosemanticity: extracting interpretable features from Claude 3 Sonnet\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2024/scaling\-monosemanticity/index\.html](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Tenneyet al\.\(2019\)I\. Tenney, D\. Das, and E\. PavlickBERT rediscovers the classical NLP pipeline\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4593–4601\.External Links:[Link](https://aclanthology.org/P19-1452/),[Document](https://dx.doi.org/10.18653/v1/P19-1452)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p3.1)\.
- Thagard and Verbeurgt \(1998\)P\. Thagard and K\. VerbeurgtCoherence as constraint satisfaction\.Cognitive Science22\(1\),pp\. 1–24\.External Links:ISSN 0364\-0213,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/S0364-0213%2899%2980033-0),[Link](https://www.sciencedirect.com/science/article/pii/S0364021399800330)Cited by:[§2\.2](https://arxiv.org/html/2608.22034#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2608.22034#S2.SS2.p2.1),[Figure 3](https://arxiv.org/html/2608.22034#S3.F3)\.
- Thagard \(2000\)P\. ThagardCoherence in thought and action\.The MIT Press\.External Links:ISBN 9780262284820,[Document](https://dx.doi.org/10.7551/mitpress/1900.001.0001),[Link](https://doi.org/10.7551/mitpress/1900.001.0001)Cited by:[§2\.2](https://arxiv.org/html/2608.22034#S2.SS2.p2.1)\.
- Virtanenet al\.\(2020\)P\. Virtanen, R\. Gommers, T\. E\. Oliphant, M\. Haberland, T\. Reddy, D\. Cournapeau, E\. Burovski, P\. Peterson, W\. Weckesser, J\. Bright, S\. J\. van der Walt, M\. Brett, J\. Wilson, K\. J\. Millman, N\. Mayorov, A\. R\. J\. Nelson, E\. Jones, R\. Kern, E\. Larson, C\. J\. Carey, İ\. Polat, Y\. Feng, E\. W\. Moore, J\. VanderPlas, D\. Laxalde, J\. Perktold, R\. Cimrman, I\. Henriksen, E\. A\. Quintero, C\. R\. Harris, A\. M\. Archibald, A\. H\. Ribeiro, F\. Pedregosa, P\. van Mulbregt, and SciPy 1\.0 ContributorsSciPy 1\.0: fundamental algorithms for scientific computing in python\.Nature Methods17,pp\. 261–272\.Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Von Oswaldet al\.\(2023\)J\. Von Oswald, E\. Niklasson, E\. Randazzo, J\. Sacramento, A\. Mordvintsev, A\. Zhmoginov, and M\. VladymyrovTransformers learn in\-context by gradient descent\.InInternational Conference on Machine Learning,pp\. 35151–35174\.Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2023a\)K\. R\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. SteinhardtInterpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NpsVSN6o4ul)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§1](https://arxiv.org/html/2608.22034#S1.p3.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px3.p1.1),[§5\.3](https://arxiv.org/html/2608.22034#S5.SS3.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2023b\)T\. Wang, M\. Kai, K\. Hariharan, and N\. ShavitForbidden facts: an investigation of competing objectives in llama 2\.InSocially Responsible Language Modelling Research,External Links:[Link](https://openreview.net/forum?id=1kavETZX3Y)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1),[§3\.4](https://arxiv.org/html/2608.22034#S3.SS4.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler, E\. H\. Chi, T\. Hashimoto, O\. Vinyals, P\. Liang, J\. Dean, and W\. FedusEmergent abilities of large language models\.Transactions on Machine Learning Research\.Note:[https://openreview\.net/forum?id=yzkSU5zdwD](https://openreview.net/forum?id=yzkSU5zdwD)External Links:ISSN 2835\-8856Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p1.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. RushTransformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Q\. Liu and D\. Schlangen \(Eds\.\),Online,pp\. 38–45\.External Links:[Link](https://aclanthology.org/2020.emnlp-demos.6/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6)Cited by:[§4](https://arxiv.org/html/2608.22034#S4.SS0.SSS0.Px10.p1.1)\.
- Xieet al\.\(2022\)S\. M\. Xie, A\. Raghunathan, P\. Liang, and T\. MaAn explanation of in\-context learning as implicit bayesian inference\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=RdJVFCHjUMI)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1),[§6](https://arxiv.org/html/2608.22034#S6.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2025\)Z\. Zheng, Y\. Wang, Y\. Huang, S\. Song, M\. Yang, B\. Tang, F\. Xiong, and Z\. LiAttention heads of large language models\.Patterns6\(2\),pp\. 101176\.External Links:ISSN 2666\-3899,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.patter.2025.101176),[Link](https://www.sciencedirect.com/science/article/pii/S2666389925000248)Cited by:[§1](https://arxiv.org/html/2608.22034#S1.p2.1)\.

Similar Articles

On the Expressive Power of Transformers

arXiv cs.AI

A survey paper examining the expressive power of transformers as language recognizers, using concepts and methods from circuit complexity to compare them with classical models of computation.

Transformers Are Inherently Succinct

Hacker News Top

This paper argues that transformer architectures are inherently succinct, meaning they can represent certain functions more efficiently than other models. It presents theoretical analysis and proofs.