线性表示假说需要群作用

arXiv cs.LG 论文

摘要

本文通过使用群作用来定义表示等价性,将线性表示假说形式化为一系列声明,澄清了不同分析中的假设。

arXiv:2609.27158v1 Announce Type: new Abstract: To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
查看原文
查看缓存全文

缓存时间: 2026/09/24 09:36

# The Linear Representation Hypothesis Needs a Group Action
Source: [https://arxiv.org/html/2609.27158](https://arxiv.org/html/2609.27158)
\\BiblatexSplitbibDefernumbersWarningOff

###### Abstract

Tomakeclaimsaboutrepresentationsthatgeneralizebeyondaparticulartrainedmodel,weneedtospecifywhentworepresentationsshouldcountasequivalent\.TheLinearRepresentationHypothesisisoftendiscussedwithoutmakingthisequivalenceexplicit\.Differentnotionsofequivalencepreservedifferentstructures,sometrics,probes,andinterventionsthatappeartostudythesamerepresentationmayinfactcorrespondtodifferenthypotheses\.WethereforearguethattheLinearRepresentationHypothesisisnotonehypothesisbutafamilyofclaimsdistinguishedbyrepresentationequivalence\.Weformalizethisideausinggroupactions,specifyingtherepresentationobject,theprocedurethatproducesit,andthepropertyultimatelyasserted,whileaccountingforequivalencesimposedbythemodelarchitecture\.Thisframeworkclarifieshowassumptionscanchangeacrossmetrics,readingpoints,andanalysisstages,andweuseittoauditcommonrepresentationquantitiesandrecentinterpretabilityanalyses\.

## 1Introduction

The Linear Representation Hypothesis underlies much of modern interpretability, whether it is tested explicitly\[[24](https://arxiv.org/html/2609.27158#bib.bib1),[19](https://arxiv.org/html/2609.27158#bib.bib13)\]or assumed when constructing linear probes\[[1](https://arxiv.org/html/2609.27158#bib.bib3),[14](https://arxiv.org/html/2609.27158#bib.bib12),[33](https://arxiv.org/html/2609.27158#bib.bib31)\], steering interventions\[[27](https://arxiv.org/html/2609.27158#bib.bib11),[18](https://arxiv.org/html/2609.27158#bib.bib32),[26](https://arxiv.org/html/2609.27158#bib.bib34)\], learned dictionaries\[[4](https://arxiv.org/html/2609.27158#bib.bib16),[28](https://arxiv.org/html/2609.27158#bib.bib17)\], and methods for comparing representations\[[16](https://arxiv.org/html/2609.27158#bib.bib8),[30](https://arxiv.org/html/2609.27158#bib.bib29),[5](https://arxiv.org/html/2609.27158#bib.bib33)\]\. Such studies usually aim to draw conclusions about representations rather than artifacts of a particular realization, which requires specifying when two realizations count as the same representation\. Calling a representation “linear” does not answer this question: one must also specify which changes are merely changes of description\.

Existing formulations rarely make this equivalence explicit\. Different analyses implicitly choose different answers, so methods that appear to study the same representation may in fact be probing different objects and testing different hypotheses\. The issue can arise even within a single study\.

\[[2](https://arxiv.org/html/2609.27158#bib.bib14)\], for example, estimate a difference\-in\-means vector and use the same array in two ways\. It is added to activations without normalization, where it acts as a displacement inVVand its magnitude affects the intervention, and separately normalized before its span is projected out, where only its projective direction inℙ⁡\(V\)\\mathbb\{P\}\(V\)matters\. Both interventions are effective and are reported as evidence for a single refusal direction, even though the two uses place that direction in different mathematical spaces and require different notions of equivalence\.

Neither operation is invalid, since each is well defined under its assumed structure\. The problem is that these assumptions are often unstated and may change between estimation, comparison, and intervention\. As a result, different stages of an analysis can silently refer to different notions of the representation, while the final conclusion is stated as though they referred to the same object\.

We propose that a representation claim be specified by the spaceMMin which its object lives together with the action of an equivalence groupGG, the procedureFFthat produces the object, and the predicatePPultimately asserted of it\. The procedure must transform equivariantly under changes of representation, while the predicate must remain invariant\. The choice ofGGis also constrained by function\-preserving reparameterizations of the architecture\. If a function\-preserving reparameterization changes the value of a quantity computed from an internal representation, that quantity is a property of the parameterization rather than of the model\.

Using this specification, we show that the Linear Representation Hypothesis is not one hypothesis but a family of claims\. Displacement, decoding, superposition, and subspace formulations place features in different mathematical objects and admit different transformation laws\. The analogous distinction applies to the procedures used to estimate these objects: two procedures may produce objects of the same type while respecting different equivalences\. We further show that admissible symmetry depends on where a representation is read and that the symmetry of a multi\-stage analysis must be checked for the composition as a whole\.

Our position is that this specification is part of stating a representation claim rather than metadata attached to it\. Making it explicit may narrow some conclusions, but it makes clear when two analyses are genuinely studying the same representation hypothesis\.

The paper proceeds as follows\. Sections 2–3 motivate and formalize representation equivalence\. Section 4 develops its architectural and compositional consequences\. Sections 5–6 audit common quantities and recent interpretability analyses\. Section 7 extends the argument to the Platonic Representation Hypothesis, and Section 8 states the reporting principle\.

## 2The Linear Representation Hypothesis Is Not One Hypothesis

The Linear Representation Hypothesis is not a single claim\. Different formulations place features in different mathematical objects and admit different transformation laws under which those objects remain meaningful\. We first review some common formulations below\.

A classical formulation treats a feature or relation as a*displacement*\. In word embeddings, some lexical relations were found to satisfy approximately constant offsets,hb−ha≈vh\_\{b\}\-h\_\{a\}\\approx v\[[20](https://arxiv.org/html/2609.27158#bib.bib2)\]\. Modern methods such as difference\-in\-means\[[19](https://arxiv.org/html/2609.27158#bib.bib13)\]and contrastive activation addition\[[27](https://arxiv.org/html/2609.27158#bib.bib11)\]use the same basic object\. Under an affine transformationh=A​h\+th=Ah\+t, the displacement transforms asv=A​vv=Av, since the translation cancels\. The coordinates of the vector change, but the constant\-offset relation does not\. This formulation therefore places the feature inVV, with the primal affine actionv↦A​vv\\mapsto Av\.

Linear decodability provides a different formulation\. If a property can be recovered from an activationh∈Vh\\in Vby a linear functionalw∈Vw\\in Vthroughw​hwh\[[1](https://arxiv.org/html/2609.27158#bib.bib3)\], then an invertible transformationh=A​hh=Ahis accompanied by the dual actionw=A​ww=Aw, which preserves the readout\. With a bias term,w​h\+bwh\+b, this extends to affine transformations by takingb=b−w​tb=b\-wt\. Thus even within probing, the admissible group depends on whether the probe includes a bias term\. More importantly, the object has changed: a probe lives inVVrather thanVV\. Although probe weights and steering vectors are both stored asdd\-dimensional arrays, they obey different transformation laws and cannot be canonically identified without additional structure\. Related work similarly distinguishes measurement from intervention geometry and linear representation from linear decodability\[[24](https://arxiv.org/html/2609.27158#bib.bib1),[23](https://arxiv.org/html/2609.27158#bib.bib7),[10](https://arxiv.org/html/2609.27158#bib.bib10)\]\.

Superposition gives another meaning to linear representation\. In this formulation, an activation is written ash=b\+∑ixi​Wih=b\+\\sum\_\{i\}x\_\{i\}W\_\{i\}, whereWiW\_\{i\}denotes a feature direction andxix\_\{i\}its coefficient\[[6](https://arxiv.org/html/2609.27158#bib.bib4)\]\. This additive view later motivated dictionary\-learning approaches and sparse autoencoders\[[4](https://arxiv.org/html/2609.27158#bib.bib16),[28](https://arxiv.org/html/2609.27158#bib.bib17)\]\. As a claim about the activation space, the decomposition is affine\-covariant: takingb↦A​b\+tb\\mapsto Ab\+tandWi↦A​WiW\_\{i\}\\mapsto AW\_\{i\}reproduces the transformed activation with the coefficients unchanged\. It also has a distinct ambiguity in its latent coordinates, reduced but generally not eliminated by sparsity or other constraints \(Section\)\. This latent non\-identifiability is distinct from equivalence under changes of activation coordinates\.

These examples show that “linear representation” does not determine a unique object or transformation law\. Moreover, even for a fixed object space, different procedures can impose different symmetry requirements\. We separate these choices before formalizing them in\.

The object itself must be specified\. Linearity also need not imply that a feature is one\-dimensional\[[21](https://arxiv.org/html/2609.27158#bib.bib5)\]\. In our framework, a one\-dimensional feature naturally lives inℙ⁡\(V\)\\mathbb\{P\}\(V\), while akk\-dimensional linear feature lives inGr⁡\(k,V\)\\mathrm\{Gr\}\(k,V\)\. Circular representations of days and months provide an example in which a two\-dimensional structure mediates computation and no single direction suffices\[[8](https://arxiv.org/html/2609.27158#bib.bib6)\]\. This distinction is compatible with linearity because the object space and admissible transformation group are separate parts of the claim: an invertible linear map preserves both lines andkk\-dimensional subspaces\.

The object alone also does not determine the symmetry of the claim\. A difference\-in\-means direction and a leading principal direction can both lie inℙ⁡\(V\)\\mathbb\{P\}\(V\), but the procedures that produce them have different symmetry requirements\. Difference\-in\-means requires no inner product, whereas PCA requires one to order directions by explained variance\. Thus two procedures may return objects in the same space while being equivariant under different groups\.

This distinction has an immediate consequence for empirical evidence\. None of the formulations above intrinsically requires an inner product\. Yet common evidence for them does: cosines and spectral quantities require similarity geometry, while absolute norms and intervention magnitudes require a fixed scale\. The metric therefore enters through the evidence used to support the hypothesis rather than through the hypothesis itself\.

This perspective differs from work that fixes a nuisance group to define representation similarity\. In the generalized shape metrics of\[[30](https://arxiv.org/html/2609.27158#bib.bib29)\], for example, quotienting by a chosen group yields a metric on representation space\. Our position differs\. The group is presupposed by the claim rather than chosen to make a measurement well behaved, so it is a property of the hypothesis\. And it also constrained by the architectural floor and by the analysis pipeline \(Sectionsand\)\.

## 3Specifying a Member of the Family

showed that formulations of the Linear Representation Hypothesis must specify not only what object represents a feature, but also how that object transforms and how it is obtained\. We now formalize these choices as a specification of a representation claim\.

### 3\.1Representation Equivalence

Let𝒳\\mathcal\{X\}be an input domain,VVa finite\-dimensional vector space, and𝒮=\{ϕ:𝒳→V\}\\mathcal\{S\}=\\\{\\phi:\\mathcal\{X\}\\to V\\\}the space of representations\. Suppose a groupGGacts on𝒮\\mathcal\{S\}\. In the settings considered below, this action is pointwise on the representation values inVV: for eachg∈Gg\\in G, there is a transformationTg:V→VT\_\{g\}:V\\to Vsuch that

\(g⋅ϕ\)​\(x\)=Tg​\(ϕ⁡\(x\)\)\.\(g\\cdot\\phi\)\(x\)=T\_\{g\}\(\\phi\(x\)\)\.\(1\)This action specifies which transformations count as changes of description rather than changes in the underlying representation\.

###### Definition 3\.1\(Representation equivalence\)\.

Two representationsϕ1,ϕ2∈𝒮\\phi\_\{1\},\\phi\_\{2\}\\in\\mathcal\{S\}are equivalent underGG, writtenϕ1∼Gϕ2\\phi\_\{1\}\\sim\_\{G\}\\phi\_\{2\}, ifϕ2=g⋅ϕ1\\phi\_\{2\}=g\\cdot\\phi\_\{1\}for someg∈Gg\\in G\. The equivalence class ofϕ\\phiis its orbit\[ϕ\]G=\{g⋅ϕ:g∈G\}\[\\phi\]\_\{G\}=\\\{g\\cdot\\phi:g\\in G\\\}\.

The choice of equivalence group determines the strength of the representation claim\. A larger group leaves fewer quantities invariant, while a smaller group risks promoting coordinate\-specific properties to properties of the representation\. The equivalence relation is therefore part of the substantive claim, and it is not always freely chosen\. A function\-preserving reparameterization may alter the coordinates of an internal representation, so a claim about the model rather than one parameterization must remain valid across such realizations\. The admissible group must therefore contain the transformations the architecture already realizes, a constraint made precise in Section\.

### 3\.2Common Choices of Transformation Group

We now specialize the pointwise transformations into affine mapsTg​\(h\)=A​h\+tT\_\{g\}\(h\)=Ah\+t, whereAAandttmay depend ongg\. Within this class, three common transformation groups form the hierarchy

Giso⊂Gsim⊂Gaff\.G\_\{\\mathrm\{iso\}\}\\subset G\_\{\\mathrm\{sim\}\}\\subset G\_\{\\mathrm\{aff\}\}\.\(2\)The affine groupGaffG\_\{\\mathrm\{aff\}\}allows arbitrary invertibleAAand translationstt\. The similarity groupGsimG\_\{\\mathrm\{sim\}\}restricts the linear part toA=s​QA=sQ, wheres\>0s\>0andQQis orthogonal\. And the isometry groupGisoG\_\{\\mathrm\{iso\}\}further requiress=1s=1\.

These groups preserve progressively stronger structures\. Affine transformations preserve affine relations, collinearity, subspace dimension, and intersection structure, but not angles or lengths\. Similarities additionally preserve angles, orthogonality, and cosine similarity\. Isometries further preserve norms, distances, and fixing absolute scale\. A larger equivalence group therefore imposes stronger invariance requirements, while allowing fewer quantities to be attributed to the representation\.

### 3\.3The Object and the Procedure

The equivalence group does not fully specify a representation claim\. An analysis must also specify what mathematical object is extracted from the representation and how that object is constructed\. In practice, the construction uses finitely many sampled activations\{ϕ⁡\(xi\)\}i=1\\\{\\phi\(x\_\{i\}\)\\\}\_\{i=1\}\. Under the pointwise actions of, these samples transform with the same coordinate change\.

LetMMdenote the space of mathematical objects, equipped with an action ofGG, and letF:𝒮→MF:\\mathcal\{S\}\\to Mdenote the procedure that constructs the object\. A claim is a predicatePPonMM, so the resulting statement aboutϕ\\phiisP⁡\(F⁡\(ϕ\)\)P\(F\(\\phi\)\)\. For the statement to be independent of the chosen representation coordinates, the procedure and predicate must satisfy

F⁡\(g⋅ϕ\)=g⋅F⁡\(ϕ\),P⁡\(g⋅m\)=P⁡\(m\)\.F\(g\\cdot\\phi\)=g\\cdot F\(\\phi\),\\qquad P\(g\\cdot m\)=P\(m\)\.\(3\)It then follows thatP⁡\(F⁡\(g⋅ϕ\)\)=P⁡\(g⋅F⁡\(ϕ\)\)=P⁡\(F⁡\(ϕ\)\)P\(F\(g\\cdot\\phi\)\)=P\(g\\cdot F\(\\phi\)\)=P\(F\(\\phi\)\)\. Notice that the conditionis stronger than invariance of the composite predicateP∘FP\\circ F\. We impose this factorization because the intermediate object is itself part of the representation claim and may be used in subsequent analyses\. Its transformation law is therefore substantive: specifyingMMrequires specifying not only what kind of object it contains, but also howGGacts on that object\.

This distinction already matters for two of the most common objects in representation analysis\. A displacement is naturally an element ofVVand transforms asv↦A​vv\\mapsto Av, whereas a linear probe is naturally an element ofVVand transforms asw↦A​ww\\mapsto Aw\. Although both are vectors in the implementation, they belong to differentGG\-spaces\. The following proposition makes precise what additional structure is required to identify them\.

###### Proposition 3\.2\(Primal and dual objects\)\.

LetdimV≥2\\dim V\\geq 2and letGL⁡\(V\)\\mathrm\{GL\}\(V\)act onVVbyv↦A​vv\\mapsto Avand onVVbyw↦A​ww\\mapsto Aw\. Then:

1. i\.There is no nonzeroGL⁡\(V\)\\mathrm\{GL\}\(V\)\-equivariant mapV→VV\\rightarrow V\.
2. ii\.An inner productgginduces an equivariant map♯g:V→V\\sharp\_\{g\}:V\\to VunderGisoG\_\{\\mathrm\{iso\}\}\.
3. iii\.The induced map\[w\]↦\[♯g​w\]\[w\]\\mapsto\[\\sharp\_\{g\}w\]on projective spaces is equivariant underGsimG\_\{\\mathrm\{sim\}\}\.

A metric thus supplies an identification betweenVVandVVonly under isometries, while passing to projective directions removes sensitivity to uniform scale and enlarges the symmetry to similarities\. Regularized probe fitting provides another example: theℓ2\\ell\_\{2\}penalty breaks equivariance under generalGL⁡\(V\)\\mathrm\{GL\}\(V\)transformations, leaving only isometric equivariance\.

### 3\.4The Specification

The preceding components can now be collected into a complete specification:

Complete specification requires​G,M,F,and​P\.\\text\{Complete specification requires \}G,\\ M,\\ F,\\ \\text\{and \}P\.
###### Definition 3\.3\(Representation claim\)\.

Let groupGGact on space𝒮\\mathcal\{S\}\. A representation claim consists of aGG\-spaceMM, a mapF:𝒮→MF:\\mathcal\{S\}\\rightarrow M, and a predicatePPonMM\. The claim aboutϕ∈𝒮\\phi\\in\\mathcal\{S\}isP⁡\(F⁡\(ϕ\)\)P\(F\(\\phi\)\)\.

TheGG\-space specifies both the mathematical object and how it transforms\. The mapFFis separate because the same object space may be reached by procedures with different symmetry properties\.

###### Definition 3\.4\(Admissibility\)\.

A representation claim\(G,M,F,P\)\(G,M,F,P\)is admissible underGGif

1. \(A1\)FFis equivariant andPPis invariant under the declared actions, as in\.
2. \(A2\)The action ofGGon𝒮\\mathcal\{S\}contains the architecture\-induced equivalences described in\.

Condition \(A1\) is internal to the analysis: it requires the method and conclusion to be well defined on the declared equivalence classes rather than on a selected coordinate realization\. Condition \(A2\) supplies an external floor\. If the architecture realizes a function\-preserving transformation, a claim about the model cannot distinguish representations related by it\. An analysis may therefore choose a larger group, but not a smaller one than the architecture permits\.

###### Definition 3\.5\(Comparability\)\.

Two admissible claims are directly comparable if they use the same representation equivalence and the sameGG\-space\. Claims on differentGG\-spaces require an explicit equivariant map relating those spaces and are otherwise not directly comparable\.

Non\-comparable claims may both be correct without supporting one another: equal numerical shape is insufficient unless the objects belong to spaces carrying compatible actions\. When the object space is shared andG1⊆G2G\_\{1\}\\subseteq G\_\{2\}with theG1G\_\{1\}action obtained by restriction, \(A1\) underG2G\_\{2\}implies \(A1\) underG1G\_\{1\}, although \(A2\) must be rechecked\. Restricting the group can therefore make additional predicates well defined, as when an object estimated under a larger group is later reported through angles or norms requiring a smaller one\.

## 4Two Consequences of the Specification

The specification has two immediate consequences: an external constraint imposed by the architecture \(A2\), and an internal constraint arising from composition \(A1\)\.

### 4\.1The Architectural Floor

The equivalence group is not always freely chosen\. Some transformations arise from the architecture itself at the parameter level and constrain which equivalence groups are admissible\. LetΘ\\Thetabe the parameter space, letfθf\_\{\\theta\}denote the function computed by parametersθ\\theta, and define the group of function\-preserving parameter transformations as

Γ=\{γ:Θ→Θ∣γinvertible andfγ⋅θ\(x\)=fθ\(x\)for allθ,x\}\.\\Gamma=\\\{\\gamma:\\Theta\\to\\Theta\\mid\\gamma\\text\{ invertible and \}f\_\{\\gamma\\cdot\\theta\}\(x\)=f\_\{\\theta\}\(x\)\\text\{ for all \}\\theta,x\\\}\.\(4\)
At a fixed reading point, a parameter symmetry induces a representation action when the transformation descends to the activations\. Writingϕ:𝒳→V\\phi:\\mathcal\{X\}\\to Vfor the representation at a fixed reading point, such an action exists whenϕ⁡\(x\)=ρ⁡\(γ\)​ϕ​\(x\)\\phi\(x\)=\\rho\(\\gamma\)\\phi\(x\)for allx∈𝒳x\\in\\mathcal\{X\}\. The image ofρ\\rhois the architectural symmetry group at that reading point\.

###### Proposition 4\.1\(Architectural invariance\)\.

Supposeγ∈Γ\\gamma\\in\\Gammainducesρ⁡\(γ\)\\rho\(\\gamma\)at a reading point, and let\(G,M,F,P\)\(G,M,F,P\)be a representation claim satisfying \(A1\)\. IfP⁡\(F⁡\(ϕ\)≠P⁡\(F⁡\(ϕ\)\)𝐶𝐿𝑂𝑆𝐸P\(F\(\\phi\)\\neq P\(F\(\\phi\)\)for someθ∈Θ\\theta\\in\\Theta, then the claim is not a property of the model function: it takes different values on parameter settings that compute the same function\.

Thus condition \(A2\) requiresG⊇Im​ρG\\supseteq\\mathrm\{Im\}\\rhoat the reading point\.

###### Corollary 4\.2\(Reading\-point dependence\)\.

The same architecture can induce different symmetries at different reading points\. In dot\-product attention,WQ↦Λ​WQW\_\{Q\}\\mapsto\\Lambda W\_\{Q\}andWK↦Λ​WKW\_\{K\}\\mapsto\\Lambda W\_\{K\}preserve the model function for anyΛ∈GL⁡\(dhead\)\\Lambda\\in\\mathrm\{GL\}\(d\_\{\\mathrm\{head\}\}\)\. At the residual stream this transformation acts trivially, so it excludes no angular claim there\. At a query or key site, however,Im⁡ρ\\operatorname\{Im\}\\rhocontainsGL⁡\(dhead\)\\mathrm\{GL\}\(d\_\{\\mathrm\{head\}\}\), which admits no nonzero invariant bilinear form: takingΛ=c​I\\Lambda=cIgivesb⁡\(c​v,c​w\)=c​b​\(v,w\)=b⁡\(v,w\)b\(cv,cw\)=cb\(v,w\)=b\(v,w\)for allc\>0c\>0, forcingb=0b=0\. Angular and metric predicates defined from a fixed inner product are therefore not invariant at those sites\.

Architecture\-induced symmetries of internal representations have been studied previously\[[12](https://arxiv.org/html/2609.27158#bib.bib18)\], including joint query–key rotations in transformers\[[34](https://arxiv.org/html/2609.27158#bib.bib19)\]\. With RoPE, the induced symmetry is reduced but remains generally anisotropic, so the same conclusion holds\. The derivation is given in Appendix\.

Equal\-dimensional representations at different reading points may carry differentGG\-actions, so transporting a direction between them requires an explicit equivariant map in the sense of Definition, rather than identification by shared coordinates\.

### 4\.2Composition of Analysis Stages

Representation analyses are usually pipelines rather than single maps\. Therefore, admissibility must be established for the composite procedure rather than inferred from individual stages\.

###### Proposition 4\.3\(Composition\)\.

LetF=Fn∘⋯∘F1F=F\_\{n\}\\circ\\cdots\\circ F\_\{1\}withFi:Mi−1→MiF\_\{i\}:M\_\{i\-1\}\\to M\_\{i\}\. IfGGacts on everyMiM\_\{i\}, eachFiF\_\{i\}is equivariant under the corresponding actions, andPPis invariant onMnM\_\{n\}, then\(G,Mn,F,P\)\(G,M\_\{n\},F,P\)satisfies \(A1\) underGG\.

If a stage is not equivariant, admissibility must instead be established for the composite\. Its symmetry is not generally obtained by intersecting groups assigned to the stages in isolation, because a stage may change the object space and hence the action seen by the next stage\. For example, underh↦A​h\+th\\mapsto Ah\+t, centered activations transform ash−h¯↦A⁡\(h−h¯\)h\-\\bar\{h\}\\mapsto A\(h\-\\bar\{h\}\), so cosine after centering can be invariant to translations that would change cosine on the uncentered activations\.

When all stages carry compatible actions of a common group, the pipeline is limited by its most restrictive stage\. A direction may be estimated underGaffG\_\{\\mathrm\{aff\}\}, compared by cosine underGsimG\_\{\\mathrm\{sim\}\}, and calibrated by a norm underGisoG\_\{\\mathrm\{iso\}\}, so the resulting claim is guaranteed admissible only underGisoG\_\{\\mathrm\{iso\}\}unless stronger invariance of the composite is established\.

## 5A Symmetry Audit of Common Representation Quantities

Tablerecords, for quantities in common use, the largest group within the hierarchy considered here under which each is well defined; using a smaller group narrows the claim it can support\.

Table 1:Symmetry audit of common representation quantities\. Rows are grouped by the largest group under which the quantity is preserved, generically in the spectral parameters\. The final column names the structure that fixes the restriction\.QuantityComputed fromWhat fixes the group*Preserved underGaffG\_\{\\mathrm\{aff\}\}*Exact rankCentered activationsLinear dependence onlySpanCentered activationsLinear dependence onlyContainment, intersection dim\.Subspace pairIncidence structureCCARepresentation pairCentered linear relations*Preserved underGsimG\_\{\\mathrm\{sim\}\}*Leading principal subspaceCovariance formVariance orderingRelative\-threshold rankCovariance spectrumSpectral ratiosEffective rank, participation ratioCovariance spectrumSpectral ratiosIntrinsic dimensionNeighbour distancesDistance ratiosCosine, angle, orthogonalityDirection pairInner product up to scalePrincipal anglesSubspace pairInner product up to scaleGrassmann distancesSubspace pairInner product up to scaleLinear CKARepresentation pairNormalized Gram geometryNearest\-neighbour agreementRepresentation pairDistance ordering*Preserved underGisoG\_\{\\mathrm\{iso\}\}*Norm, distanceActivationsFixed Euclidean scaleIntervention magnitudeDisplacement inVVNorm and scaleAbsolute\-threshold rankCovariance spectrumAbsolute thresholdProcrustes distanceRepresentation pairEuclidean alignment lossReconstruction lossResidual inVVEuclidean residual norm*Preserved under signed permutations of the latent coordinates*Coordinatewise sparsity penaltyLatent codesCoordinatewise axesDictionary learning acts on two spaces\.Underh↦A​h\+th\\mapsto Ah\+t, takingD↦A​DD\\mapsto AD,Wenc↦Wenc​AW\_\{\\mathrm\{enc\}\}\\mapsto W\_\{\\mathrm\{enc\}\}A, and transforming the encoder and decoder biases to absorbttleaves every encoder preactivation, and hence the latent codes, unchanged\. The reconstruction residual transforms ash−h^↦A⁡\(h−h^\)h\-\\hat\{h\}\\mapsto A\(h\-\\hat\{h\}\), so the Euclidean loss is preserved for all residuals only whenA​A=IAA=I\. The restriction toGisoG\_\{\\mathrm\{iso\}\}on the activation side therefore comes from the reconstruction loss\[[4](https://arxiv.org/html/2609.27158#bib.bib16),[28](https://arxiv.org/html/2609.27158#bib.bib17)\]\. On the latent side,W↦W​BW\\mapsto WBandx↦B​xx\\mapsto Bxleave the reconstruction unchanged, while a coordinatewise sparsity penalty reduces this mixing symmetry to signed permutations\. These are distinct restrictions on distinct spaces\.

## 6Current Practice Leaves the Specification Implicit

The preceding sections make the specification explicit\. We now examine what happens when its components remain implicit\. The recurring problem is not that strong structural assumptions are necessarily unwarranted, but that objects, actions, and procedures are identified or changed without recording the corresponding change in the claim\.

### 6\.1Identification by Storage Format

A linear probe produces a coefficient vector inVV, while a steering method such as difference\-in\-means produces a displacement inVV\. Because both are stored as length\-ddarrays, they are routinely treated as the same kind of direction\. Probe weights are compared to steering vectors by cosine, used as intervention directions, or combined with vectors of other provenance\. For example,\[[3](https://arxiv.org/html/2609.27158#bib.bib26)\]use both steering vectors and linear\-probe weights as additive intervention directions and compare their directions by cosine similarity\. The same identification appears across reading points\. A direction estimated at one layer is often transported to another by the identity map because both residual streams have dimensiondd, even though they need not carry the same group action\.

Propositionseparates two operations that this practice conflates\. Once a metric is fixed, the induced map fromVVtoVVdescends to projective spaces equivariantly underGsimG\_\{\\mathrm\{sim\}\}\. A cosine comparison between a probe direction and a steering direction can therefore be meaningful under similarity geometry\. Reusing the probe coefficients themselves as an additive displacement is stronger\. The vector\-level identification is equivariant only underGisoG\_\{\\mathrm\{iso\}\}\.

### 6\.2Normalization and Calibration

Hidden group choices also enter through normalization and calibration\. Contrastive Activation Addition estimates a difference\-in\-means vector and applies it as a scaled translation\[[27](https://arxiv.org/html/2609.27158#bib.bib11)\]\. The reported procedure normalises vector magnitudes across behaviours but not across layers, where residual\-stream norms grow over the forward pass\. These choices assign meaning to length and therefore introduce Euclidean scale into an otherwise affine\-covariant construction\. Activation Addition adopts the opposite convention\[[29](https://arxiv.org/html/2609.27158#bib.bib15)\]: its activation\-difference vector is left unnormalised, so the displacement produced by a coefficient depends on the norm supplied by the sampled activations\. The same numerical coefficient therefore does not denote the same displacement under the two conventions\.\[[31](https://arxiv.org/html/2609.27158#bib.bib27)\]make the required calibration explicit by scaling an optimised refusal direction to match the norm of a difference\-in\-means direction, thereby stipulating a common Euclidean magnitude\.

Because directional ablation depends only on the normalised direction,\[[31](https://arxiv.org/html/2609.27158#bib.bib27)\]sample unit directions within refusal cones directly rather than normalising arbitrary convex combinations, which would bias the induced distribution over directions\. In our terminology, the sampling procedure is adapted to the projective object actually consumed by the intervention\.

### 6\.3Unnamed Reading Points

The reading point is itself part of theGG\-space\. Corollaryshows that the same function\-preserving parameter transformation can act trivially at the residual stream and as a nontrivial general linear transformation at a query or key site\.

KeyDiff provides a direct case\[[22](https://arxiv.org/html/2609.27158#bib.bib22)\]\. The method evicts KV\-cache entries using pairwise cosine similarity among keys within a head, and its supporting analyses report key and query cosines, theirLLnorms, PCA of key caches, andlogdet\(KK\)\\log\\det\(KK\), all in per\-head query or key space\. The empirical performance of the eviction rule is not in question here\. What is at issue is the status of the geometric quantities offered in explanation of it\. In ordinary dot\-product attention, the functionally relevant quantity is the bilinear pairingq​kqk, equivalentlyWQ​WKW\_\{Q\}W\_\{K\}at the parameter level\[[7](https://arxiv.org/html/2609.27158#bib.bib30)\]; with RoPE, it is the familyWQ​R​\(τ\)​WKW\_\{Q\}R\(\\tau\)W\_\{K\}indexed by relative position\. By Corollaryand Appendix, this gauge is anisotropic even under RoPE, so none of these quantities is invariant\. The quantities used to explain the computation are therefore not invariants of the computation they explain\.

Related analyses of head\-space quantities inherit the same dependence on chosen head coordinates\. Cross\-Gram singular values between head projections are not invariant under general invertible changes of basis, although the underlying residual\-stream subspaces are, unless the spanning matrices are first orthonormalised\[[9](https://arxiv.org/html/2609.27158#bib.bib23)\]\. Likewise, cosine comparisons of head\-space singular vectors and their calibration against a fixed Euclidean reference distribution depend on the chosen head\-space geometry\[[11](https://arxiv.org/html/2609.27158#bib.bib24)\]\. Appendixgives corresponding transformations\.

By contrast,\[[32](https://arxiv.org/html/2609.27158#bib.bib25)\]compare column spaces of transposed projection matrices, which are subspaces of the residual stream\. UnderWQ↦Λ​WQW\_\{Q\}\\mapsto\\Lambda W\_\{Q\}, for example,WQ↦WQ​ΛW\_\{Q\}\\mapsto W\_\{Q\}\\Lambdaleaves its column span unchanged\. In our framework, the comparison targets an object preserved by the parameter symmetry rather than geometry internal to the gauge\-dependent head coordinates\.

Gauge\-related parameter settings compute the same model function but can cause a cosine\-based key\-space eviction rule to select different tokens\. Naming the reading point is therefore necessary to determine whether a reported geometric property belongs to the model, to an explicitly chosen geometry, or only to the sampled parameterization\.

### 6\.4Inconsistency Within a Single Analysis

The clearest cases arise when the specification changes within one pipeline\. Propositionrequires the conclusion to respect the actions introduced throughout the analysis, not merely the symmetry of its initial estimator\. Three patterns recur\.

One estimate, differentGG\-spaces\.\[[2](https://arxiv.org/html/2609.27158#bib.bib14)\]use one difference\-in\-means estimate both as an unnormalised displacement inVVand as a projective direction inℙ⁡\(V\)\\mathbb\{P\}\(V\)\(Section\)\. The estimate changesGG\-space between uses, and the combined conclusion is stronger than either use establishes\.

Change of object mid\-analysis\.A probe coefficient belongs toVV, while an additive intervention consumes an element ofVV, so moving from probing to intervention requires specifying how the two spaces are related\.\[[17](https://arxiv.org/html/2609.27158#bib.bib28)\]compare probe\-weight, mass\-mean\-shift, and CCS directions as alternatives for intervention after using probes to identify relevant attention heads\. Their intervention partially makes the required calibration explicit: the chosen direction is normalised, and the added displacement is scaled by the empirical standard deviation of activations along that direction\. Under a similarityA=s​QA=sQ, a normalised probe direction transforms asQ​w^Q\\hat\{w\}while the projected standard deviation transforms ass​σs\\sigma, so the calibrated displacementσ​w^\\sigma\\hat\{w\}transforms covariantly asA⁡\(σ​w^\)A\(\\sigma\\hat\{w\}\)\. The complete construction is therefore compatible withGsimG\_\{\\mathrm\{sim\}\}, while identifying a probe direction inVVwith an intervention direction inVVstill depends on the chosen similarity geometry\.

Preprocessing that restricts the group\.SVCCA provides the failure at the level of procedures\[[25](https://arxiv.org/html/2609.27158#bib.bib20)\]\. CCA alone is affine\-invariant on centered representations, but SVCCA first performs singular\-value truncation, a spectral operation not preserved by general invertible linear transformations\. The CCA stage cannot restore the affine invariance lost in preprocessing\.

In each case the problem is not the presence of a strong assumption\. It is that the assumption enters at one stage while the conclusion is stated as though it inherited the weaker assumptions of another\.

## 7The Same Omission Beyond Linear Representation

The same specification issue appears in the Platonic Representation Hypothesis\[[15](https://arxiv.org/html/2609.27158#bib.bib21)\], although the hypothesis itself is unrelated to linear representation\. Representations of models with different widths need not lie in a common representation space, so cross\-model comparison proceeds through structures induced on shared inputs, such as neighbourhoods or similarity matrices\. Because two representations may induce the same structure on one input distribution and differ elsewhere, the resulting claim is also indexed by the distribution on which the comparison is made\. These dependencies are typically left implicit\.

Convergence then amounts to saying that differently realised representations belong to a common equivalence class\. The object, predicate, and comparison procedure are supplied by the hypothesis, while the transformation group remains open, and it is this remaining component on which the truth value turns\. On a finite sample, sufficiently permissive linear equivalence can become degenerate\.\[[16](https://arxiv.org/html/2609.27158#bib.bib8)\]show that when the representation width is at least the number of sampled inputs and both activation matrices have full row rank, a similarity measure invariant to arbitrary invertible linear transformations cannot distinguish them\. Finite\-sample convergence under such an equivalence can therefore become near\-trivial in this regime\. Under an isometric interpretation, by contrast, convergence would require agreement in absolute metric structure that models with different widths, tokenizers, and objectives generally cannot satisfy\. Existing evaluations instead operate closer toGsimG\_\{\\mathrm\{sim\}\}, with the effective equivalence determined by the alignment measures in use rather than specified independently of them\.

This places the Platonic Representation Hypothesis directly inside the audit of Section\. Nearest\-neighbour agreement, kernel\-based alignment, and angular quantities inherit the same geometric commitments already catalogued there\. The truth of the Platonic Representation Hypothesis cannot be assessed until the object being compared and the equivalence under which convergence is asserted have been specified\.

## 8The Specification Principle

One clarification is needed before stating the principle\. Condition \(A1\) can always be satisfied by shrinkingGGor weakeningPP, so admissibility alone carries no content\. What carries content is maximality: the strength of a claim is the largest group under which it survives\. Reporting a predicate under an unnecessarily small group is not an error of validity but a loss of content\. This is why Tableis organized by the largest preserving group rather than by any group under which a quantity happens to be defined\.

We therefore propose a simple reporting principle\. A representation study should state theGG\-space in which its object lives, the procedureFFthat produces the object and why it is equivariant under the declared action, and the predicatePPultimately asserted of it\. It should also name the reading point at which the representation is taken and check the symmetry of the complete analysis when several stages are composed\. These requirements need not be presented in group\-theoretic notation, but the underlying choices must be recoverable from the statement of the claim rather than reconstructed from the implementation\.

This requirement has a real cost\. Some conclusions become narrower once their geometric assumptions are made explicit, some comparisons require an additional map between spaces, and some quantities must be attributed to a representation together with a chosen metric or parameterization rather than to the representation alone\. One might instead declareGisoG\_\{\\mathrm\{iso\}\}throughout, but Condition \(A2\) forbids this when the architecture realizes transformations outsideGisoG\_\{\\mathrm\{iso\}\}\.

A second objection is that training returns one specificθ\\theta, so the geometry of thatθ\\thetais a fact about the artifact one actually possesses, whatever the orbit contains\. This is correct, and the framework does not forbid it\. It fixes what such a statement is about\. A property that varies acrossIm⁡ρ\\operatorname\{Im\}\\rhois a property ofθ\\thetatogether with the procedure that produced it, not offθf\_\{\\theta\}, and asserting it therefore incurs an empirical obligation that is rarely discharged: the reported geometry must be shown stable across seeds and training runs before it can be attributed to anything more general than the run in hand\.

A third objection is that the proposal merely relabels assumptions already present in existing analyses\.\[[24](https://arxiv.org/html/2609.27158#bib.bib1),[23](https://arxiv.org/html/2609.27158#bib.bib7)\], for example, make the geometry relating measurement and intervention explicit by specifying an inner product under which causally separable concepts are orthogonal\.\[[13](https://arxiv.org/html/2609.27158#bib.bib9)\]argue that near\-orthogonality can arise generically in high dimensions and be induced by whitening, so the resulting geometry is not uniquely diagnostic of learned conceptual structure\. Making the geometry explicit therefore does more than relabel an assumption: it turns the disagreement into a precise question about which inner product supports the orthogonality predicate\.

## 9Conclusion

The Linear Representation Hypothesis is not one hypothesis but a family, and its members are distinguished by which representations they treat as the same\. Stating a member requires aGG\-space, an equivariant procedure producing the object, and a predicate asserted of it, subject to the floor the architecture already fixes\. Once these are recorded, two studies can be checked for whether they assert the same claim without reconstructing both analyses\. Nothing in the requirement depends on linearity, and the same omission appears wherever a claim about representations is made at all\.

## Acknowledgment

L\.H\.Y\. thanks Tianyu Jiang for helpful discussions\.

## References

- \[1\]\(2016\)Understanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p3.1)\.
- \[2\]A\. Arditi, O\. B\. Obeso, A\. Syed, D\. Paleka, N\. Rimsky, W\. Gurnee, and N\. Nanda\(2024\)Refusal in language models is mediated by a single direction\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=pH3XAQME6c)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p3.1),[§6\.4](https://arxiv.org/html/2609.27158#S6.SS4.p2.1)\.
- \[3\]U\. Bhalla, S\. Srinivas, A\. Ghandeharioun, and H\. Lakkaraju\(2024\)Towards unifying interpretability and control: evaluation via intervention\.arXiv preprint arXiv:2411\.04430\.Cited by:[§6\.1](https://arxiv.org/html/2609.27158#S6.SS1.p1.1)\.
- \[4\]T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. L\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. Olah\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.html](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p4.1),[§5](https://arxiv.org/html/2609.27158#S5.p2.1)\.
- \[5\]L\. Domenichelli, D\. Brunato, and F\. Dell’Orletta\(2026\)Linguistic profiling of transformer embedding geometry\.InProceedings of the 30th Conference on Computational Natural Language Learning,External Links:[Link](https://aclanthology.org/2026.conll-main.10/)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1)\.
- \[6\]N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen,et al\.\(2022\)Toy models of superposition\.arXiv preprint arXiv:2209\.10652\.Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p4.1)\.
- \[7\]N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly,et al\.\(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2021/framework/index\.html](https://transformer-circuits.pub/2021/framework/index.html)Cited by:[§6\.3](https://arxiv.org/html/2609.27158#S6.SS3.p2.1)\.
- \[8\]J\. Engels, E\. Michaud, I\. Liao, W\. Gurnee, and M\. Tegmark\(2025\)Not all language model features are one\-dimensionally linear\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 84591–84622\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/d3221cdb27e49d9c1cd35ad254feccfe-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p6.1)\.
- \[9\]E\. FokouĂŠ\(2026\)Multi\-head attention as ensemble nadaraya\-watson estimation: variance reduction, decorrelation, and optimal head diversity\.arXiv preprint arXiv:2605\.20271\.Cited by:[§C\.4](https://arxiv.org/html/2609.27158#A3.SS4.p1.1),[§6\.3](https://arxiv.org/html/2609.27158#S6.SS3.p3.1)\.
- \[10\]N\. Garg, J\. Kleinberg, and K\. Peng\(2026\)How many features can a language model store under the linear representation hypothesis?\.arXiv preprint arXiv:2602\.11246\.Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p3.1)\.
- \[11\]Y\. Ge, S\. Liu, Y\. Wang, T\. Liu, B\. Bi, L\. Mei, J\. Yao, J\. Guo, and X\. Cheng\(2026\)Prism\-Δ\\Delta: differential subspace steering for prompt highlighting in large language models\.arXiv preprint arXiv:2603\.10705\.Cited by:[§C\.4](https://arxiv.org/html/2609.27158#A3.SS4.p2.1),[§6\.3](https://arxiv.org/html/2609.27158#S6.SS3.p3.1)\.
- \[12\]C\. Godfrey, D\. Brown, T\. Emerson, and H\. Kvinge\(2022\)On the symmetries of deep learning models and their internal representations\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 11893–11905\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/4df3510ad02a86d69dc32388d91606f8-Paper-Conference.pdf)Cited by:[§4\.1](https://arxiv.org/html/2609.27158#S4.SS1.p4.1)\.
- \[13\]S\. Golechha, L\. Bushnaq, E\. Ong, N\. Kayal, and N\. Schoots\(2025\)Intricacies of feature geometry in large language models\.InThe Fourth Blogpost Track at ICLR 2025,External Links:[Link](https://openreview.net/forum?id=Ut3ml7Hdwx)Cited by:[§8](https://arxiv.org/html/2609.27158#S8.p5.1)\.
- \[14\]W\. Gurnee and M\. Tegmark\(2024\)Language models represent space and time\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jE8xbmvFin)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1)\.
- \[15\]M\. Huh, B\. Cheung, T\. Wang, and P\. Isola\(2024\)The platonic representation hypothesis\.InProceedings of the 41st International Conference on Machine Learning,Cited by:[§7](https://arxiv.org/html/2609.27158#S7.p1.1)\.
- \[16\]S\. Kornblith, M\. Norouzi, H\. Lee, and G\. Hinton\(2019\)Similarity of neural network representations revisited\.InProceedings of the 36th International Conference on Machine Learning,pp\. 3519–3529\.External Links:[Link](https://proceedings.mlr.press/v97/kornblith19a.html)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§7](https://arxiv.org/html/2609.27158#S7.p2.1)\.
- \[17\]K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. Wattenberg\(2023\)Inference\-time intervention: eliciting truthful answers from a language model\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=aLLuYpn83y)Cited by:[§6\.4](https://arxiv.org/html/2609.27158#S6.SS4.p3.1)\.
- \[18\]L\. Liu, T\. Zhan, L\. H\. Yao, S\. Ghosh, and T\. Jiang\(2026\)Cross\-lingual steering for figurative language generation\.arXiv preprint arXiv:2605\.30443\.Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1)\.
- \[19\]S\. Marks and M\. Tegmark\(2024\)The geometry of truth: emergent linear structure in large language model representations of true/false datasets\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=aajyHYjjsk)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p2.1)\.
- \[20\]T\. Mikolov, W\. Yih, and G\. Zweig\(2013\)Linguistic regularities in continuous space word representations\.InProceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 746–751\.External Links:[Link](https://aclanthology.org/N13-1090/)Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p2.1)\.
- \[21\]C\. Olah\(2024\)What is a linear representation? what is a multidimensional feature?\.Note:Transformer Circuits ThreadEdited by Adam JermynExternal Links:[Link](https://transformer-circuits.pub/2024/july-update/index.html)Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p6.1)\.
- \[22\]J\. Park, D\. Jones, M\. J\. Morse, R\. Goel, M\. Lee, and C\. Lott\(2025\)KeyDiff: key similarity\-based KV cache eviction for long\-context LLM inference in resource\-constrained environments\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=uBaFH7aQnC)Cited by:[§6\.3](https://arxiv.org/html/2609.27158#S6.SS3.p2.1)\.
- \[23\]K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. Veitch\(2025\)The geometry of categorical and hierarchical concepts in large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 76441–76463\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/be7430d22a4dae8516894e32f2fcc6db-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.27158#S2.p3.1),[§8](https://arxiv.org/html/2609.27158#S8.p5.1)\.
- \[24\]K\. Park, Y\. J\. Choe, and V\. Veitch\(2024\)The linear representation hypothesis and the geometry of large language models\.InProceedings of the 41st International Conference on Machine Learning,pp\. 39643–39666\.External Links:[Link](https://proceedings.mlr.press/v235/park24c.html)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p3.1),[§8](https://arxiv.org/html/2609.27158#S8.p5.1)\.
- \[25\]M\. Raghu, J\. Gilmer, J\. Yosinski, and J\. Sohl\-Dickstein\(2017\)SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability\.InProceedings of the 31st International Conference on Neural Information Processing Systems,pp\. 6078–6087\.Cited by:[§6\.4](https://arxiv.org/html/2609.27158#S6.SS4.p4.1)\.
- \[26\]S\. Raval, H\. J\. Song, L\. Wu, A\. Harrasse, J\. M\. Phillips, F\. Barez, and A\. Abdullah\(2026\)Curveball steering: the right direction to steer isn’t always linear\.arXiv preprint arXiv:2603\.09313\.Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1)\.
- \[27\]N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. Turner\(2024\)Steering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15504–15522\.External Links:[Link](https://aclanthology.org/2024.acl-long.828/)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p2.1),[§6\.2](https://arxiv.org/html/2609.27158#S6.SS2.p1.1)\.
- \[28\]A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, A\. Tamkin, E\. Durmus, T\. Hume, F\. Mosconi, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. Henighan\(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.Note:[https://transformer\-circuits\.pub/2024/scaling\-monosemanticity/index\.html](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p4.1),[§5](https://arxiv.org/html/2609.27158#S5.p2.1)\.
- \[29\]A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid\(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§6\.2](https://arxiv.org/html/2609.27158#S6.SS2.p1.1)\.
- \[30\]A\. H\. Williams, E\. Kunz, S\. Kornblith, and S\. Linderman\(2021\)Generalized shape metrics on neural representations\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=L9JM-pxQOl)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1),[§2](https://arxiv.org/html/2609.27158#S2.p9.1)\.
- \[31\]T\. Wollschläger, J\. Elstner, S\. Geisler, V\. Cohen\-Addad, S\. Günnemann, and J\. Gasteiger\(2025\)The geometry of refusal in large language models: concept cones and representational independence\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=80IwJqlXs8)Cited by:[§6\.2](https://arxiv.org/html/2609.27158#S6.SS2.p1.1),[§6\.2](https://arxiv.org/html/2609.27158#S6.SS2.p2.1)\.
- \[32\]H\. Yamagiwa, Y\. Takase, and H\. Shimodaira\(2026\)Measuring affinity between attention\-head weight subspaces via the projection kernel\.arXiv preprint arXiv:2601\.10266\.Cited by:[§6\.3](https://arxiv.org/html/2609.27158#S6.SS3.p4.1)\.
- \[33\]L\. H\. Yao, V\. Anand, Y\. Zhuang, and T\. Jiang\(2026\)Rhetorical questions in LLM representations: a linear probing study\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 155–172\.External Links:[Link](https://aclanthology.org/2026.acl-long.5/)Cited by:[§1](https://arxiv.org/html/2609.27158#S1.p1.1)\.
- \[34\]B\. Zhang, Z\. Zheng, Z\. Chen, and J\. Li\(2025\)Beyond the permutation symmetry of transformers: the role of rotation for model fusion\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=wBJIO15pBV)Cited by:[§4\.1](https://arxiv.org/html/2609.27158#S4.SS1.p4.1)\.

## Appendix ANotation and Mathematical Preliminaries

### A\.1Groups and Symmetries

A symmetry is a transformation that preserves a specified structure\. Mathematically, a collection of compatible transformations is described by a group\. The choice of group determines which transformations are regarded as equivalent descriptions of the same object\.

A groupGGact on a set𝒳\\mathcal\{X\}, if each elementg∈Gg\\in Gassigns a transformation

x↦g⋅xx\\mapsto g\\cdot x\(5\)such that the identity transformation acts trivially and composition of transformations follows the group operation\. The orbit of an elementxxunderGGis

\[x\]G=\{g⋅x:g∈G\}\.\[x\]\_\{G\}=\\\{g\\cdot x:g\\in G\\\}\.\(6\)The orbit defines the equivalence relation induced by the group action: two elements are equivalent if they belong to the same orbit\.

### A\.2Linear Spaces

LetVVbe a finite\-dimensional real vector space, and the dual space ofVVis the vector space of linear functionals

V=\{w:V→ℝ∣wis linear\}\.V=\\\{w:V\\rightarrow\\mathbb\{R\}\\mid w\\text\{ is linear\}\\\}\.\(7\)A vectorv∈Vv\\in Vand a dual vectorw∈Vw\\in Vare related through the natural pairing

⟨w,v⟩=w⁡\(v\)\.\\langle w,v\\rangle=w\(v\)\.\(8\)Although bothVVandVVare finite\-dimensional vector spaces with the same dimension, they are distinct spaces without additional structure such as an inner product\.

An affine space is a vector space without a fixed origin, in which the difference between points forms vector, and the addition of a point and a vector results in a new point\. Given a vector spaceVV, an affine transformation takes the form

h↦A​h\+t,h\\mapsto Ah\+t,\(9\)wherehhandttare vectors inVV, andAAis an invertible linear transformation onVV\. Under an affine transformation, displacements transform as

\(h2−h1\)↦A⁡\(h2−h1\),\(h\_\{2\}\-h\_\{1\}\)\\mapsto A\(h\_\{2\}\-h\_\{1\}\),\(10\)so affine relations are preserved while lengths and angles are generally not\.

The projective space ofVVis the space of one\-dimensional linear subspaces ofVV

ℙ\(V\)=\{span\(v\):v∈V,v≠0\}\.\\mathbb\{P\}\(V\)=\\\{\\mathrm\{span\}\(v\):v\\in V,v\\neq 0\\\}\.\(11\)Equivalently, projective space identifies vectors that differ only by a nonzero scalar multiplication\. Therefore, quantities defined onℙ⁡\(V\)\\mathbb\{P\}\(V\)describe directions rather than vector magnitudes\.

The GrassmannianGr⁡\(k,V\)\\mathrm\{Gr\}\(k,V\)denotes the space of allkk\-dimensional linear subspaces ofVV

Gr⁡\(k,V\)=\{U⊆V:dim⁡\(U\)=k\}\.\\mathrm\{Gr\}\(k,V\)=\\\{U\\subseteq V:\\mathrm\{dim\}\(U\)=k\\\}\.\(12\)
The projective spaceℙ⁡\(V\)\\mathbb\{P\}\(V\)is the special caseGr⁡\(1,V\)\\mathrm\{Gr\}\(1,V\), where an element corresponds to a one\-dimensional feature\. More generally, a multidimensional linear feature is naturally represented as an element ofGr⁡\(k,V\)\\mathrm\{Gr\}\(k,V\)\.

For a vector spaceVV, the general linear group is

GL\(V\)=\{A:V→V∣Ais linear and invertible\}\.\\mathrm\{GL\}\(V\)=\\\{A:V\\rightarrow V\\mid A\\text\{ is linear and invertible\}\\\}\.\(13\)It acts naturally on vectors byv↦A​vv\\mapsto Av\. The induced action on the dual space isw↦A​ww\\mapsto Awwhich preserves the vector\-dual pairing\.

An inner product onVVis a positive\-definite bilinear form that induces norms, distances, and angular quantities\. For example,

‖v‖=⟨v,v⟩,\\\|v\\\|=\\sqrt\{\\langle v,v\\rangle\},\(14\)or

cos⁡\(u,v\)=⟨u,v⟩‖u‖​‖v‖\.\\cos\(u,v\)=\\frac\{\\langle u,v\\rangle\}\{\\\|u\\\|\\\|v\\\|\}\.\(15\)Such quantities depend on the chosen inner product and are therefore not intrinsic to a vector space alone\.

## Appendix BProof of Proposition\(i\)

Letf:V→Vf:V\\to Vsatisfyf⁡\(A​w\)=A​f​\(w\)f\(Aw\)=Af\(w\)for allA∈GL⁡\(V\)A\\in\\mathrm\{GL\}\(V\)and allw∈Vw\\in V\. Fixw≠0w\\neq 0and letHw=\{A∈GL⁡\(V\):A​w=w\}H\_\{w\}=\\\{A\\in\\mathrm\{GL\}\(V\):Aw=w\\\}be its stabilizer under the dual action, equivalently\{A:w∘A=w\}\\\{A:w\\circ A=w\\\}\. Equivariance givesf⁡\(w\)=A​f​\(w\)f\(w\)=Af\(w\)for everyA∈HwA\\in H\_\{w\}, sof⁡\(w\)f\(w\)lies in the subspace ofVVfixed pointwise byHwH\_\{w\}\.

WriteK=ker⁡wK=\\ker w, a hyperplane, and picku∉Ku\\notin K\. EveryAApreservingwwacts arbitrarily onKKand fixesuumoduloKK, soHwH\_\{w\}contains all maps of the formu↦u\+κu\\mapsto u\+\\kappa,A\|K∈GL⁡\(K\)\\left\.A\\right\|\_\{K\}\\in\\mathrm\{GL\}\(K\)withκ∈K\\kappa\\in Karbitrary\. Letv∈Vv\\in Vbe fixed by all ofHwH\_\{w\}\. Writingv=α​u\+κ0v=\\alpha u\+\\kappa\_\{0\}withκ0∈K\\kappa\_\{0\}\\in Kand applyingu↦u\+κu\\mapsto u\+\\kappagivesα​κ=0\\alpha\\kappa=0for allκ∈K\\kappa\\in K, henceα=0\\alpha=0wheneverdimV≥2\\dim V\\geq 2\. Thenv=κ0∈Kv=\\kappa\_\{0\}\\in Kis fixed by all ofGL⁡\(K\)\\mathrm\{GL\}\(K\), which forcesκ0=0\\kappa\_\{0\}=0sincedimK≥1\\dim K\\geq 1andGL⁡\(K\)\\mathrm\{GL\}\(K\)acts transitively onK∖\{0\}K\\setminus\\\{0\\\}\. Thereforef⁡\(w\)=0f\(w\)=0for everyw≠0w\\neq 0, andf⁡\(0\)=0f\(0\)=0follows from equivariance underA=c​IA=cIwithc≠1c\\neq 1\. Hencef≡0f\\equiv 0\.

The hypothesisdimV≥2\\dim V\\geq 2cannot be dropped\. OnV=ℝV=\\mathbb\{R\}the dual action isw↦w/aw\\mapsto w/aand the mapf⁡\(w\)=c/wf\(w\)=c/wsatisfiesf⁡\(w/a\)=a​f​\(w\)f\(w/a\)=af\(w\)for everya≠0a\\neq 0, so nonzero equivariant maps exist in dimension one\.

## Appendix CArchitecture\-Induced Representation Symmetries

### C\.1The rotary commutant

Letdhead=2​nd\_\{\\mathrm\{head\}\}=2nand let RoPE act at positionttbyR⁡\(t\)=⨁j=1R2​\(t​θj\)R\(t\)=\\bigoplus\_\{j=1\}R\_\{2\}\(t\\theta\_\{j\}\), whereR2​\(α\)R\_\{2\}\(\\alpha\)is the planar rotation byα\\alphaandθ1,…,θn\\theta\_\{1\},\\dots,\\theta\_\{n\}are the rotary frequencies\. The attention logit between a query at positionttand a key at positionssis

⟨R⁡\(t\)​WQ​x,R⁡\(s\)​WK​y⟩=x​WQ​R​\(s−t\)​WK​y,\\langle R\(t\)W\_\{Q\}x,R\(s\)W\_\{K\}y\\rangle=xW\_\{Q\}R\(s\-t\)W\_\{K\}\\,y,\(16\)usingR⁡\(t\)​R​\(s\)=R⁡\(s−t\)R\(t\)R\(s\)=R\(s\-t\)\.Writingτ=s−t\\tau=s\-t, a reparameterizationWQ↦Λ​WQW\_\{Q\}\\mapsto\\Lambda W\_\{Q\},WK↦Λ​WKW\_\{K\}\\mapsto\\Lambda W\_\{K\}preserves every logit if and only if

Λ​R​\(τ\)​Λ=R⁡\(τ\)for all relative positions​τ\.\\Lambda R\(\\tau\)\\Lambda=R\(\\tau\)\\qquad\\text\{for all relative positions \}\\tau\.\(17\)Takingτ=0\\tau=0givesΛ=Λ\\Lambda=\\Lambda, recovering the gauge of Corollary\. Substituting back, EquationbecomesR⁡\(τ\)​Λ=Λ​R​\(τ\)R\(\\tau\)\\Lambda=\\Lambda R\(\\tau\)for allτ\\tau: the admissibleΛ\\Lambdaare exactly the elements of the commutant of\{R⁡\(τ\)\}τ\\\{R\(\\tau\)\\\}\_\{\\tau\}inGL⁡\(2​n\)\\mathrm\{GL\}\(2n\)\.

Identifyℝ≅ℂ\\mathbb\{R\}\\cong\\mathbb\{C\}by pairing the two coordinates of each rotary plane, under whichR⁡\(τ\)R\(\\tau\)becomes multiplication bydiag⁡\(e,…,e\)\\mathrm\{diag\}\(e,\\dots,e\)\. When the frequencies are pairwise distinct and the relative positionsτ\\taurange over enough values to separate them, theℝ\\mathbb\{R\}\-algebra generated by\{R⁡\(τ\)\}τ\\\{R\(\\tau\)\\\}\_\{\\tau\}is the full diagonal algebraℂ\\mathbb\{C\}, whose commutant inEndℝ​\(ℝ\)\\mathrm\{End\}\_\{\\mathbb\{R\}\}\(\\mathbb\{R\}\)is againℂ\\mathbb\{C\}\. Its invertible elements are

Λ=⨁j=1sj​R2​\(αj\),sj\>0,αj∈\[0,2​π\),\\Lambda=\\bigoplus\_\{j=1\}s\_\{j\}R\_\{2\}\(\\alpha\_\{j\}\),\\qquad s\_\{j\}\>0,\\ \\alpha\_\{j\}\\in\[0,2\\pi\),\(18\)soIm⁡ρ\\operatorname\{Im\}\\rhoat a key site is the group of independently scaled rotations of the rotary planes, of real dimension2​n=dhead2n=d\_\{\\mathrm\{head\}\}\. Degenerate frequency sets enlarge the commutant: ifmmfrequencies coincide, the corresponding block isGLm​\(ℂ\)\\mathrm\{GL\}\_\{m\}\(\\mathbb\{C\}\)rather thanℂ\\mathbb\{C\}, so Equationis the smallest image the architecture realizes\.

### C\.2Consequences at a key site

Equationis a proper subgroup ofGL⁡\(dhead\)\\mathrm\{GL\}\(d\_\{\\mathrm\{head\}\}\), so the argument of Corollarydoes not apply directly\. It nonetheless excludes the metric predicates in question, because the scale factorssjs\_\{j\}vary independently across planes\. Writek=\(k,…,k\)k=\(k,\\dots,k\)in rotary\-plane blocks\. Under Equation,

‖k‖↦∑jsj​‖k‖,\\\|k\\\|\\mapsto\\sum\_\{j\}s\_\{j\}\\\|k\\\|,\(19\)which equals‖k‖\\\|k\\\|for allkkonly when everysj=1s\_\{j\}=1\. Norms are therefore not invariant, and neither are cosines: choosings1≠s2s\_\{1\}\\neq s\_\{2\}andsj=1s\_\{j\}=1otherwise changes the relative weight of the first two planes in⟨k1,k2⟩\\langle k\_\{1\},k\_\{2\}\\ranglewhile changing the norms differently, socos⁡\(k1,k2\)\\cos\(k\_\{1\},k\_\{2\}\)is not preserved\. The same anisotropy moves the eigenvalues ofK​KKKnon\-uniformly, so PCA directions,logdet\(KK\)\\log\\det\(KK\), and any spectral quantity not expressible through ratios fixed by the per\-plane scaling are likewise gauge dependent\. Only a uniform scalings1=⋯=sns\_\{1\}=\\cdots=s\_\{n\}would restrict the action toGsimG\_\{\\mathrm\{sim\}\}and rescue the angular quantities, and the architecture does not impose it\.

By contrast, the criterion of Sectionsurvives, since Equationis a subgroup of the gauge under whichΩ\\Omegaand the key differences were shown to transform oppositely\. The RoPE restriction therefore shrinks the architectural image without restoring any of the quantities Corollaryexcludes\.

### C\.3What an invariant version looks like

The framework does more than reject quantities\. Because the QK gauge preserves realised attention logitsq​kqkexactly, any predicate expressible through those logits is admissible at these sites\. Without RoPE, the corresponding parameter\-level object isWQ​WKW\_\{Q\}W\_\{K\}; with RoPE, it is the familyWQ​R​\(τ\)​WKW\_\{Q\}R\(\\tau\)W\_\{K\}indexed by relative position\. This is enough to restate the eviction criterion invariantly\. What KeyDiff needs is a notion of when two realised keys are interchangeable in their effect on attention, and that effect is mediated throughq⁡\(k1−k2\)q\(k\_\{1\}\-k\_\{2\}\)\. Define the second moment of the realised queries by

Ω=𝔼⁡\[q​q\],d⁡\(k1,k2\)=\(k1−k2\)​Ω​\(k1−k2\)\.\\Omega=\\mathbb\{E\}\\\!\\left\[qq\\right\],\\qquad d\(k\_\{1\},k\_\{2\}\)=\(k\_\{1\}\-k\_\{2\}\)\\,\\Omega\\,\(k\_\{1\}\-k\_\{2\}\)\.\(20\)In the no\-RoPE case,Ω=WQ​Σ​WQ\\Omega=W\_\{Q\}\\Sigma W\_\{Q\}, whereΣ\\Sigmais the second moment of the residual\-stream inputs on the query side\. Under the induced gaugeq↦Λ​qq\\mapsto\\Lambda qandk↦Λ​kk\\mapsto\\Lambda k, we haveΩ↦Λ​Ω​Λ\\Omega\\mapsto\\Lambda\\Omega\\Lambdaandk1−k2↦Λ⁡\(k1−k2\)k\_\{1\}\-k\_\{2\}\\mapsto\\Lambda\(k\_\{1\}\-k\_\{2\}\), so the two transformations cancel andddis unchanged\. The quantity it measures is also the right one:d⁡\(k1,k2\)d\(k\_\{1\},k\_\{2\}\)is the mean squared change in attention logit incurred by substituting one key for the other\. WhenΩ∝I\\Omega\\propto Ithe criterion reduces to Euclidean distance in key space, which orders pairs as cosine does when key norms are equal, so the original rule is recovered as the special case in which the realized query distribution is isotropic in the gauge that happens to have been sampled\. The distinction is not merely formal: gauge\-related parameter settings compute the same model function while producing different eviction sets under the cosine rule and identical eviction sets underdd\.

### C\.4Head\-space quantities

The same architectural gauge constrains geometric comparisons between attention heads\.\[[9](https://arxiv.org/html/2609.27158#bib.bib23)\]compute cross\-Gram matrices between key\-projection matrices and interpret their singular values in terms of principal angles between head subspaces\. IfGh​hG\_\{hh\}denotes such a cross\-Gram matrix, independent invertible changes of basis within the two heads give

Gh​h↦Bh​Gh​h​Bh\.G\_\{hh\}\\mapsto B\_\{h\}G\_\{hh\}B\_\{h\}\.\(21\)Its singular values are therefore not invariant under general invertibleBhB\_\{h\}andBhB\_\{h\}unless the spanning matrices have first been orthonormalised\. This does not imply that the underlying subspaces are gauge dependent: when represented in residual\-stream coordinates, those subspaces are unchanged by invertible changes of basis within the heads\.

\[[11](https://arxiv.org/html/2609.27158#bib.bib24)\]compare leading left singular vectors of different heads by cosine similarity and calibrate the observed statistic against uniformly sampled unit vectors in a fixed Euclideanℝ\\mathbb\{R\}\. Both constructions depend on the chosen head\-space geometry\. An anisotropic change of head coordinates changes the cosine statistic, while the uniform distribution on the Euclidean unit sphere is itself not preserved by such a transformation\.

相似文章

线性表示假说综述

arXiv cs.AI

本文综述了线性表示假说在AI及相关领域的研究,分析了不一致性,并提出更严格的表述以使其成为可证伪的科学主张。

表示对齐基于线性结构

arXiv cs.LG

本文研究了Platonic Representation Hypothesis,提出对齐源于表示中的线性结构,并引入了一个包含信号、偏置和噪声的统计框架。

无理解的趋同:语言模型表征一致但推理分歧

arXiv cs.CL

本文通过考察来自8个家族的16个语言模型在800个推理问题上的表现,探究了Platonic Representation Hypothesis。研究发现,虽然模型在内部表征上趋于一致,但在推理过程中,尤其是决策后阶段,它们出现分歧,而且共享的表征对预测的因果影响极小。

语言模型编码命题的语境真实性

arXiv cs.CL

本文研究了大语言模型(LLMs)如何将情境真值——即真值取决于上下文证据的命题——编码为激活空间中的线性方向,表明这些表征在不同输出策略下持续存在,并可通过对话方的断言而发生偏移,相关证据将表征偏移与谄媚行为联系起来。