线性表示假说综述

arXiv cs.AI 论文

摘要

本文综述了线性表示假说在AI及相关领域的研究,分析了不一致性,并提出更严格的表述以使其成为可证伪的科学主张。

arXiv:2609.22695v1 Announce Type: new Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science. But previous works have not consistently treated the LRH as a falsifiable scientific hypothesis; we analyze these inconsistencies and examine their implications for how prior theoretical and methodological results should be interpreted. Based on this analysis, we argue that claims regarding linear representations become well-defined only through careful examination of the model, representation location, feature definition, and evaluation dataset. We therefore propose a more rigorous formalization of the LRH that makes these dependencies explicit and allows the hypothesis to be evaluated as a falsifiable scientific claim. Finally, we identify some non-trivial open problems that warrant further attention from the research community.
查看原文
查看缓存全文

缓存时间: 2026/09/23 09:12

# A Survey on the Linear Representation Hypothesis
Source: [https://arxiv.org/html/2609.22695](https://arxiv.org/html/2609.22695)
Marc E\. CanbyAffiliation:The Grainger College of EngineeringIkhyun ChoAffiliation:University of Illinois Urbana\-ChampaignJulia HockenmaierAffiliation:\{samuel27, marcec2, ihcho2, juliahmr\}@illinois\.edu

###### Abstract

The term “linear representation hypothesis” \(LRH\) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science\. But previous works have not consistently treated the LRH as a falsifiable scientific hypothesis; we analyze these inconsistencies and examine their implications for how prior theoretical and methodological results should be interpreted\. Based on this analysis, we argue that claims regarding linear representations become well\-defined only through careful examination of the model, representation location, feature definition, and evaluation dataset\. We therefore propose a more rigorous formalization of the LRH that makes these dependencies explicit and allows the hypothesis to be evaluated as a falsifiable scientific claim\. Finally, we identify some non\-trivial open problems that warrant further attention from the research community\.

## 1Introduction

Neural activations enable neural systems to do remarkably complex things\. Yet it is not obvious how such capabilities arise, which is why the underlying mechanisms of these systems have long been a focus of inquiry\. Neural network activations correspond todd\-dimensional hidden embedding vectorsemb⁡\(x\)∈ℝd\\mathrm\{emb\}\(x\)\\in\\mathbb\{R\}^\{d\}for inputsxx\. The discovery of linear regularities in the embeddings of recurrent neural network language models\([Mikolov et al\., 2013](https://arxiv.org/html/2609.22695#bib.bib106)\), e\.g\.emb⁡\("king"\)−emb⁡\("queen"\)≈emb⁡\("man"\)−emb⁡\("woman"\)\\mathrm\{emb\}\(\\text\{"king"\}\)\-\\mathrm\{emb\}\(\\text\{"queen"\}\)\\approx\\mathrm\{emb\}\(\\text\{"man"\}\)\-\\mathrm\{emb\}\(\\text\{"woman"\}\), likely contributed to the prominence of the concept of linear representations\. The presence of such a relation would suggest the existence of a linear direction in the embedding space such that:

emb⁡\("queen"\)≈emb⁡\("king"\)\+vmale→female,\\mathrm\{emb\}\(\\text\{"queen"\}\)\\approx\\mathrm\{emb\}\(\\text\{"king"\}\)\+v\_\{\\text\{male\}\\rightarrow\\text\{female\}\},\(1\)wherevmale→female∈ℝdv\_\{\\text\{male\}\\rightarrow\\text\{female\}\}\\in\\mathbb\{R\}^\{d\}is a gender vector \(or, more formally, acounterfactual intervention vector\), providing a more interpretable picture of previously human\-incomprehensible hidden embedding vectors\.

Table 1:Number of papers explicitly employing the term “linear representation hypothesis” over time, based on Google Scholar search results \(counts as of December 2025\)\. The early history is discussed in Section[2\.1](https://arxiv.org/html/2609.22695#S2.SS1)\. The sharp increase after 2022 was triggered by[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.22695#bib.bib98)\.Whether these early findings generalize across evaluation settings has been subject to debate\. For example, GloVe word embeddings\([Pennington et al\., 2014](https://arxiv.org/html/2609.22695#bib.bib96)\)were reported to achieve 75% accuracy on analogy tests of the form “ais tobascis to?” where the answer is predicted by finding the embedding with the highest cosine similarity toemb⁡\(b\)−emb⁡\(a\)\+emb⁡\(c\)\\mathrm\{emb\}\(\\text\{b\}\)\-\\mathrm\{emb\}\(\\text\{a\}\)\+\\mathrm\{emb\}\(\\text\{c\}\)\. On the larger BATS analogy benchmark introduced by[Gladkova et al\. \(2016\)](https://arxiv.org/html/2609.22695#bib.bib95), however, GloVe’s accuracy fell below 30%\. This discrepancy shows that whether linear representations are observed in machine\-learned embeddings depends on the embedding model, evaluation dataset, and linguistic features under investigation\.

More recently, however, there has been a noticeable shift in terminology: as shown in Table[1](https://arxiv.org/html/2609.22695#S1.T1), the term “linear representation hypothesis \(LRH\)” has come into widespread use\. This naturally leads to a deeper question that has received surprisingly little discussion:is “linear representation”truly something that merits being called a hypothesis?And if so, what would falsifying the LRH imply for research that relies on it as a methodological assumption?

We first categorize the different ways in which the LRH is used in the literature\. Some works treat it as a theoretical premise about the structure of neural representations, from which other results are derived \(Type 7 in Table[3](https://arxiv.org/html/2609.22695#S2.T3)\)\. In others, it is examined empirically, with studies presenting supporting evidence, counterexamples, or analyses of when linear representations appear to hold \(Types 2–5; see Table[3](https://arxiv.org/html/2609.22695#S2.T3)\)\. Still other work relies on the LRH more implicitly, as a methodological assumption that motivates techniques such as representation steering and related interventions \(Type 8 in Table[3](https://arxiv.org/html/2609.22695#S2.T3)\)\. Our examination leads us to consider whether the LRH, as commonly formulated, satisfies the criteria of a falsifiable scientific hypothesis\. Our analysis suggests that claims about linear representations are only well\-defined relative to a specific model, representation location, feature definition, and evaluation dataset; we therefore propose a formalization of the LRH that makes these requirements explicit and places prior empirical findings within the scope of a well\-defined hypothesis\.

Table 2:Summary of studies providing empirical evidence for, or counterexamples to, the linear representation hypothesis \(LRH\)\.TheRepresentationcolumn indicates whether the LRH holds in the study; when it does not hold, the observed non\-linear representation is specified\. In theLayercolumn,∙∘∘\\bullet\\circ\\circdenotes the initial \(non\-contextualized\) embedding layer,∘∙∘\\circ\\bullet\\circintermediate layers, and∘∘∙\\circ\\circ\\bulletthe final layer before unembedding\. Gray circles \(∙\{\\color\[rgb\]\{0\.75,0\.75,0\.75\}\\bullet\}\) indicate that the corresponding layer was also examined but exhibited weaker linear representations\.
## 2Formulation of the LRH

### 2\.1Origins of the LRH

As shown in Table[1](https://arxiv.org/html/2609.22695#S1.T1), the phrase*linear representation hypothesis*has appeared sporadically before the recent rise of its use in machine learning\. These earlier occurrences used the phrase in different contexts:[Pitz et al\. \(1976\)](https://arxiv.org/html/2609.22695#bib.bib89)used it in cognitive science to describe the hypothesis that numerical samples are mentally encoded along an imaginary number line;[Zentall \(2001\)](https://arxiv.org/html/2609.22695#bib.bib88)used it in animal cognition to describe linear orderings over stimuli; and[González et al\. \(2007\)](https://arxiv.org/html/2609.22695#bib.bib87)used the related phrase*non\-linear representation hypothesis*in the context of classification problems that are not linearly separable in input space\.

In contemporary machine learning, the term*linear representation hypothesis*was popularized by[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.22695#bib.bib98)\. The rapid increase in usage shown in Table[1](https://arxiv.org/html/2609.22695#S1.T1)follows their formulation of the following claim:

###### Definition 1\(LRH;[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.22695#bib.bib98)’s\)\.

Neural networks represent input features as directions in activation space\.

However, as[Olah \(2024\)](https://arxiv.org/html/2609.22695#bib.bib86)points out, the notion of a*feature*in this definition is itself non\-trivial\. In this survey, we use*feature*broadly to refer to any concept or relation in an input that prior work seeks to measure in internal representations or manipulate by intervening on those representations; examples are summarized in Table[2](https://arxiv.org/html/2609.22695#S1.T2)\.

### 2\.2Early Empirical Findings

Building on[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.22695#bib.bib98)’s definition,[Nanda et al\. \(2023\)](https://arxiv.org/html/2609.22695#bib.bib108)provide early empirical evidence for the LRH by showing that the board state of an OthelloGPT model trained on Othello move sequences can be recovered from its residual stream using supervised linear probes\. Unsupervised methodological progress has also been made, beginning with the findings by[Huben et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib92)that sparse autoencoders \(SAEs\) with a single\-layer \(linear\) encoder can find interpretable features, such as words beginning with “w”, from large language models like Pythia\-70M\([Biderman et al\., 2023](https://arxiv.org/html/2609.22695#bib.bib91)\)\. Notably,[Templeton et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib93)identify the well\-known “Golden Gate Bridge” feature in Claude 3 Sonnet\([Anthropic, 2024](https://arxiv.org/html/2609.22695#bib.bib90)\)and shows that increasing this feature’s activation causally steers the model’s generation toward Golden\-Gate\-Bridge\-related outputs\.

### 2\.3The Strong, the Weak and the Causal

#### The Strong LRH

Later,[Smith \(2024\)](https://arxiv.org/html/2609.22695#bib.bib101)points out that the prior definition \(Definition[1](https://arxiv.org/html/2609.22695#Thmdefinition1)\) is vague, and shows that it can be understood in two distinct ways:

###### Definition 2\(LRH;[Smith \(2024\)](https://arxiv.org/html/2609.22695#bib.bib101)’s\)\.

In a given representation space,

- \(i\)\(Strong LRH\)Allfeatures used by neural networks are represented by linear directions\.
- \(ii\)\(Weak LRH\)Somefeatures used by neural networks are represented as linear directions\.

This classification is meaningful in that it makes at least the strong LRH more easily falsifiable\. Under such constraints, several recent studies provide concrete counterexamples to the strong LRH by identifying features whose representations are inherently non\-linear\.[Csordás et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib100)demonstrate that GRU models encode certain concepts, such as token position, using concentric representations\. Similarly,[Engels et al\. \(2025a\)](https://arxiv.org/html/2609.22695#bib.bib102)argue that transformer\-based language models represent periodic concepts, such as days of the week, using circular representation\. Together, these results suggest that*not*all features used by neural networks are linearly represented\.

#### The Weak LRH

In Definition[2](https://arxiv.org/html/2609.22695#Thmdefinition2), the weak LRH asserts the existence of some features that are represented as linear directions\. As summarized in Table[2](https://arxiv.org/html/2609.22695#S1.T2), several studies report empirical evidence for linear representations of particular features\. However, the existence of supporting examples does not entail that the weak LRH is generally true\. This leads researchers to habitually refer to it as a hypothesis, while leaving unclear how long we should continue to call it a hypothesis\. More concretely, the difficulty of generalizing evidence for the weak LRH can be understood along two main dimensions\.

First, empirical evidence supporting linear representation for one feature is insufficient to justify the claim that another feature is also linearly represented\. Moreover, because each study defines a feature with respect to a particular dataset or input, features that appear identical on the surface may in fact differ\. For example, for the linguistic feature of gender,[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)examine whether masculine nouns such as “king” can be transformed into feminine counterparts like “queen”, whereas[Templeton et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib93)focus on gender bias using sentence\-level contexts such as “I asked the nurse a question, and …”\. Under this approach, researchers are forced to carry out an endless series of validations for the infinitely many fine\-grained features that may exist\.

Second, a more serious challenge is that the class of embedding spaces under investigation is not fixed: they change as new LLMs are developed\. This makes hypothesis formulation in this field particularly challenging\. For instance, in physics, while researchers gather evidence in support of a hypothesis, the universe has not suddenly changed into a different one overnight\. By contrast, an LLM under study may eventually become outdated, and researchers will then encounter the embedding spaces of new LLMs\. In other words, experiments that support the LRH for one model do not guarantee generalization to another model trained with different data or architectures\.

In sum, although the weak LRH may sound weak by name, it is in fact the most difficult to handle; unlike the strong LRH, it is not easily falsified\. Because its loose formulation can cover nearly all LRH\-related work, references to the LRH in practice typically fall into this category, and we follow this categorization in this paper unless otherwise specified\. At the same time, because of its broad applicability, the merit of calling it a hypothesis is questionable; since this issue lies at the core of this paper, Section[5](https://arxiv.org/html/2609.22695#S5)will focus on this question in greater depth\.

Table 3:Contexts in which the term “linear representation hypothesis \(LRH\)” is used across the literature\.A substantial fraction of prior work employs the LRH primarily as a methodological assumption \(Type 8\) to justify linear analysis tools, such as linear probing, sparse autoencoder \(SAE\), and steering vector\. As argued in Section[5\.3](https://arxiv.org/html/2609.22695#S5.SS3), when such assumptions are invoked across different layers or feature definitions, the resulting methodological justification may not necessarily transfer across studies\.- •†\\daggerTheoretical findings derived under the assumption that the LRH is true\.
- •‡\\ddaggerLRH used to justify linear methods \(e\.g\., linear probing, SAE, steering vector\)\.

#### The Causal LRH

There is also another line of work investigating linear representations of causality\([Park et al\., 2024c](https://arxiv.org/html/2609.22695#bib.bib12);[Jiang et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib105);[Nguyen and Leng, 2025](https://arxiv.org/html/2609.22695#bib.bib99)\)\. Prior work has often cited[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)as a seminal contribution\. Their study starts from the premise that causality can be defined by constructing counterfactual pairs between two opposing factual states, as illustrated in Eq\. \([1](https://arxiv.org/html/2609.22695#S1.E1)\)\. Given such pairs, they characterize the geometry of linear representations through the following three questions:

###### Definition 3\(LRH;[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)’s\)\.

For allcounterfactual pairsof a concept,

- \(i\)\(Subspace LRH\) All pairs belong to a shared subspace\.
- \(ii\)\(Measurement LRH\) The concept can be measured using a linear probe\.
- \(iii\)\(Intervention LRH\) The concept can be modified by adding a steering vector\.

A central theoretical contribution of[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)is to show that these three definitions can be unified under an appropriate choice of inner product, which they term the*causal inner product*\. Empirically, they further demonstrate that a range of linguistic features, such as male→\\rightarrowfemale \(gender\), English→\\rightarrowFrench \(translation\), and country→\\rightarrowcapital \(holonym/meronym\), have linear representations that are mutually orthogonal when the concepts are causally separable \(see Table[2](https://arxiv.org/html/2609.22695#S1.T2)\)\.

However, this counterfactual\-pair formulation does not naturally extend to all features studied under the LRH, such as the Golden Gate Bridge feature in[Templeton et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib93)\. These limitations are partially addressed by[Park et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib15), which extends the framework to general hierarchical111Earlier work on hierarchical semantic relations explored representations designed to capture hierarchical structure, including embeddings that encode asymmetric entailment and hypernymy relations\([Vendrov et al\., 2016](https://arxiv.org/html/2609.22695#bib.bib3)\), probabilistic denotational representations\([Lai and Hockenmaier, 2017](https://arxiv.org/html/2609.22695#bib.bib4)\), box embeddings for richer set relations\([Vilnis et al\., 2018](https://arxiv.org/html/2609.22695#bib.bib1)\), and hyperbolic embeddings that capture tree\-like hierarchies more efficiently than Euclidean embeddings in low dimensions\([Nickel and Kiela, 2017](https://arxiv.org/html/2609.22695#bib.bib2)\)\. These works were motivated by the view that standard Euclidean representations, including simple linear representations, are not well suited to capturing such relations\.is\-arelations by relaxing the reliance on counterfactual pairs, while retaining the constraint that features should be causally separable\.222It is worth noting, however, that empirical claims that causally separable concept vectors are orthogonal under the causal inner product\([Park et al\., 2024c](https://arxiv.org/html/2609.22695#bib.bib12);[Park et al\., 2024a](https://arxiv.org/html/2609.22695#bib.bib13)\)were later weakened by the observation that even random concepts can become nearly orthogonal in high\-dimensional spaces\([Golechha and Nandi, 2024](https://arxiv.org/html/2609.22695#bib.bib17)\)\. In response, for hierarchical concepts with set inclusion, such asmammal⊂animal\\text\{mammal\}\\subset\\text\{animal\},[Park et al\. \(2024b\)](https://arxiv.org/html/2609.22695#bib.bib14)further showed thatcos⁡\(vmammal−vanimal,vanimal\)\\cos\\\!\\left\(v\_\{\\text\{mammal\}\}\-v\_\{\\text\{animal\}\},v\_\{\\text\{animal\}\}\\right\)is close to zero\. However, sincevmammal⊤​vrandom≈0v\_\{\\text\{mammal\}\}^\{\\top\}v\_\{\\text\{random\}\}\\approx 0in high dimensions, this contrast largely reduces to the trivial observation that‖vrandom‖2\\\|v\_\{\\text\{random\}\}\\\|^\{2\}is nonzero;[Golechha et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib16)further showed that whitening can make random vectors nearly orthogonal even without causal separability\.Nevertheless,[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12),[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib105), and[Park et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib15)all share the common limitation that their theoretical analyses apply exclusively to the last layer of the model\. We discuss this layer\-specific limitation in more detail in Section[3](https://arxiv.org/html/2609.22695#S3)\.

## 3Target Embedding Layers

#### Last layer

[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)provide a formalization of the linear representation hypothesis for binary concepts defined with counterfactual word pairs \(e\.g\., male→\\rightarrowfemale\), showing that such concepts are well\-defined linear representations as directions in the representation space\. Further,[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib105)provide a theoretical analysis demonstrating that architectural choices such as softmax and cross\-entropy promote a linear structure in the representations used for next\-token prediction\.

However, it is important to note that, in both works\([Park et al\., 2024c](https://arxiv.org/html/2609.22695#bib.bib12);[Jiang et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib105)\), the model is defined as

P⁡\(y∣x\)=softmax​\(emb​\(x\)⊤​unemb​\(y\)\),P\(y\\mid x\)=\\text\{softmax\}\(\\text\{emb\}\(x\)^\{\\top\}\\text\{unemb\}\(y\)\),wherexxdenotes the input sequence \(e\.g\., “He is the”\) andyydenotes the next token to predict \(e\.g\., “king”\)\. By construction, this formulation applies only to the final layer representation immediately preceding the unembedding linear transformation\. Accordingly, the theoretical guarantees established in these works are restricted to the last layer\. Therefore, using these results to justify claims about intermediate layers1,…,L−11,~\\ldots,~L\-1in a network with a total ofLLextends their conclusions beyond the scope of what is formally established\.

#### Intermediate layers

As an early attempt to support the LRH,[Nanda et al\. \(2023\)](https://arxiv.org/html/2609.22695#bib.bib108)report that, when linear probes are trained on the hidden states of a GPT model trained on Othello game records to distinguish between white and black stones, performance is relatively poor \(62–75%\)\. This might naively be interpreted as evidence for a non\-linear representation\. However, when the same probes are trained to distinguish*my*stone versus*your*stone, linear probing achieves extremely high accuracy \(99%\) at the 4th layer of the 7\-layer model, highlighting that the choice of features is crucial for determining whether a representation is linearly encoded\.

Studies measuring linear probing performance and steering effectiveness for sentiment and detoxification tasks similarly find that these effects peak in the middle layers of the network and then decline toward the final layers\([Turner et al\., 2023](https://arxiv.org/html/2609.22695#bib.bib79);[Tigges et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib58)\)\. Additionally, information about time and position is also linearly represented, with layer\-wise performance plateauing around the halfway point of the network\([Gurnee and Tegmark, 2024](https://arxiv.org/html/2609.22695#bib.bib64)\)\. Linear probing and intervention for political perspectives in LLMs are also observed most effective in middle layers\([Kim et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib60)\), and a broad investigation spanning hallucination, persuasion, pessimism, refusal, sycophancy, and truthfulness reports that, in a 28\-layer model, linear probing and steering are most effective at the 15th layer\([Agarwal et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib56)\)\.

Recently,[Merullo et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib109)provide empirical evidence about when linear relations emerge for\(subject,relation,object\)\(\\text\{subject\},\\text\{relation\},\\text\{object\}\)triplets by using the subject\-side hidden statehsubject\(l\)h^\{\(l\)\}\_\{\\text\{subject\}\}extracted from thell\-th intermediate layer, and modeling the relation as:

hobject\(L\)=W​hsubject\(l\)\+b,where​l∈\{1,…,L−1\}\.h^\{\(L\)\}\_\{\\text\{object\}\}=Wh^\{\(l\)\}\_\{\\text\{subject\}\}\+b,~\\text\{where \}l\\in\\\{1,\\dots,L\-1\\\}\.Building on the findings of[Hernandez et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib82)333Earlier work by[Hernandez et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib82)observed that some relations \(e\.g\., country–capital\) are well captured by faithful linear relational embeddings, whereas others \(e\.g\., company–CEO\) are not, despite both being well learned\.,[Merullo et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib109)further demonstrate that linear relational embeddings tend to emerge when the corresponding triplets have high co\-occurrence frequency in the training data, and fail to do so when the frequency is low\. This work is distinguished from prior studies in that it does not merely add supporting or counterexamples to the LRH, but instead seeks to identify the underlying boundary conditions that determine*when*linear representations arise\.

## 4Gap between Theory and Experiment

The existing literature on the LRH can be organized along several dimensions, including the perspective taken on LRH, the models studied, the layers examined, and the features considered, as summarized in Table[2](https://arxiv.org/html/2609.22695#S1.T2)\. In addition, the extent to which each work provides theoretical versus empirical support for LRH is summarized in Table[3](https://arxiv.org/html/2609.22695#S2.T3)\. Having organized the literature by target embedding type, a natural question arises: If thefinal layeris where linear structure is theoretically formalized and promoted[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12);[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.22695#bib.bib105);[Park et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib15);[Liu et al\. \(2026\)](https://arxiv.org/html/2609.22695#bib.bib73), why do empirical studies consistently report the clearest linear separability\([Marks and Tegmark, 2024](https://arxiv.org/html/2609.22695#bib.bib107)\), the lowest probing error\([Nanda et al\., 2023](https://arxiv.org/html/2609.22695#bib.bib108);[Li et al\., 2023](https://arxiv.org/html/2609.22695#bib.bib94)\), and the most effective linear intervention\([Tigges et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib58);[Li et al\., 2025a](https://arxiv.org/html/2609.22695#bib.bib72);[Shen et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib63);[Kim et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib60);[Agarwal et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib56)\)atintermediate layers? Furthermore, as shown in Table[2](https://arxiv.org/html/2609.22695#S1.T2), the layers that experimentally exhibited non\-linear representations were most often the last layer\([Kantamneni and Tegmark, 2025](https://arxiv.org/html/2609.22695#bib.bib40);[Lei and Cooper, 2025a](https://arxiv.org/html/2609.22695#bib.bib77);[Walker et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib62)\)\.

[Skean et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib81)propose an explanation for why the last layer in LLMs is not always the best for embeddings, and why intermediate layers often provide better representations, based on information theory, geometry, and stability analysis\. However, why*linear*representations are empirically observed more frequently in the intermediate layers still remains an open problem, which we provide further details in Section[6](https://arxiv.org/html/2609.22695#S6)\.

## 5What exactly is hypothetical in LRH?

The term*hypothesis*can be used in at least two distinct senses in the literature: scientific hypotheses, which propose explanatory claims about the world, and statistical hypotheses, which are formal statements used within statistical testing procedures\([Alger, 2022](https://arxiv.org/html/2609.22695#bib.bib24)\)\. However, regarding the LRH, the literature is inconsistent about whether it should be understood as a statistical or a scientific hypothesis\. In this section, we review both possibilities and, in particular, discuss the modifications required for the LRH to function as a falsifiable scientific hypothesis\.

### 5\.1LRH as a Statistical Hypothesis

There is a vein of work that approaches the linear representation hypothesis from the perspective of statistical hypothesis testing\. Treating the LRH as a statistical hypothesis means translating a claim about linear representation into a null hypothesis and an alternative hypothesis, along with a test statistic\.

#### Examples of Statistical Formulations\.

For instance,[Guo et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib28)applied statistical hypothesis testing, finding that embeddings significantly satisfy a linear representation hypothesis, exhibiting stronger linear alignment and additive compositional generalization with respect to interpretable attributes \(e\.g\., gender, age, and occupation\) than expected under randomly permuted attribute–embedding pairings\. Another example is[Reizinger et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib69), who treat binary semantic attributes in ImageNet\-X\([Idrissi et al\., 2023](https://arxiv.org/html/2609.22695#bib.bib27)\)defined by human annotations \(e\.g\., object size or position\) as proxies for latent factors, and test whether these attributes can be predicted from the second\-to\-last layer representations of ResNet\-50\([He et al\., 2016](https://arxiv.org/html/2609.22695#bib.bib26)\)and ViT\-B/16\([Dosovitskiy et al\., 2021](https://arxiv.org/html/2609.22695#bib.bib25)\)using linear decoders at levels significantly above chance\.

#### Limitations of Statistical Formulations\.

Despite their methodological clarity, formulating the LRH as a statistical hypothesis faces two major limitations\. First, specifying an appropriate null hypothesis requires strong assumptions about how embeddings would be structured*in the absence of linear representations*\. In practice, such tests do not allow us to accept the hypothesis that an embedding space is structured by linear representations; for example, they only allow us to reject the hypothesis that the embeddings follow a multivariate Gaussian distribution\([Li et al\., 2025b](https://arxiv.org/html/2609.22695#bib.bib80)\), offering limited insight into whether the LRH itself holds\. Second, these approaches rely on assumptions about the underlying data\-generating process, including the distributions of latent features or embeddings, such as von Mises–Fisher models\([Reizinger et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib69)\)or Monte Carlo simulation\([Guo et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib28)\), which are difficult to justify or not clearly generalizable beyond the given dataset\.

Nevertheless, viewing the LRH through the lens of statistical hypothesis testing has the advantage of yielding concrete and well\-scoped contributions under clearly defined assumptions\. However, work that formulates and evaluates the LRH in hypothesis testing remains relatively rare in the literature\.

### 5\.2LRH as a Scientific Hypothesis

One of the fundamental requirements of a scientific hypothesis is falsifiability\([Popper, 1959](https://arxiv.org/html/2609.22695#bib.bib97)\); an idea that is formulated in a way that does not allow for falsification cannot be regarded as a scientific hypothesis\.

#### Toward Falsifiability\.

In Section[2\.3](https://arxiv.org/html/2609.22695#S2.SS3), we discussed that the definition implicitly used in the majority of the literature corresponds to the weak LRH \(Definition[2](https://arxiv.org/html/2609.22695#Thmdefinition2)\), and that this formulation is inherently challenging to falsify\. Although a number of counterexamples are shown to be represented in non\-linear ways \(Table[2](https://arxiv.org/html/2609.22695#S1.T2)\), these results do not directly refute the weak LRH, since they can always escape refutation through post hoc redefinition of features—“well, that wasn’t one of the*some*features\!” Can the LRH, as it is currently formulated and being*increasingly*used \(Table[1](https://arxiv.org/html/2609.22695#S1.T1)\), be regarded as a scientific hypothesis?

The preceding discussion shows that the question*“Is the linear representation hypothesis true or not?”*is not, by itself, a well\-defined open question\. As Table[2](https://arxiv.org/html/2609.22695#S1.T2)shows, the LRH may hold or fail depending on the modelMM, representation locationll, featureff, and dataset𝒟\\mathcal\{D\}\. The question, therefore, depends on a specific tuple\(M,l,f,𝒟\)\(M,l,f,\\mathcal\{D\}\)\. This echoes a fundamental lesson from[Gladkova et al\. \(2016\)](https://arxiv.org/html/2609.22695#bib.bib95)introduced in Section[1](https://arxiv.org/html/2609.22695#S1): linguistic regularities must be examined together with the data on which they are evaluated\. That is, the LRH must be framed not as a hypothesis about ‘some’ general feature as in previous definitions, but as a hypothesis about ‘one’ specific feature, if it is to function as a falsifiable scientific claim\.

#### Formation vs\. Use of Representations\.

A second source of confusion is that the LRH literature often conflates two separate claims\. The first is a claim about*representation formation*: whether a feature becomes linearly decodable in a model representation, corresponding tox→emb⁡\(x\)x\\to\\mathrm\{emb\}\(x\)\. This is the kind of claim supported by linear probes, as in[Nanda et al\. \(2023\)](https://arxiv.org/html/2609.22695#bib.bib108)\. The second is a claim about*representation use*: whether the downstream computation actually uses that representation when producing the output, corresponding toemb⁡\(x\)→y\\mathrm\{emb\}\(x\)\\to y\. This is the kind of claim targeted by causal approaches such as intervention\([Park et al\., 2024c](https://arxiv.org/html/2609.22695#bib.bib12), e\.g\.,\)\.

These two claims are logically distinct\. As[Geiger et al\. \(2021\)](https://arxiv.org/html/2609.22695#bib.bib19)point out, probing is by itself unable to establish that the model causally uses that information in producing its output\. This distinction also explains why the attempt by[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12), discussed in Section[4](https://arxiv.org/html/2609.22695#S4), to unify different notions of linear representation through the causal inner product is limited as a general definition of the LRH\. Their construction applies a causal inner product to concept vectors obtained from counterfactual token\-pair differences in the unembedding space of the final layer\. This is essentially different from much of the probing\-based literature, where the LRH is often used to determine whether a feature is linearly separable in intermediate\-layer representations \(For more details, see Appendix[A](https://arxiv.org/html/2609.22695#A1)\)\.

Consequently, a falsifiable LRH statement must do two things\. First, it must specify the local setting in which the claim is evaluated: the modelMM, representation locationll, featureff, and evaluation distribution𝒟\\mathcal\{D\}\. Second, it must specify the type of claim being made: whether the claim concerns representation formation\(x→emb⁡\(x\)\)\(x\\to\\mathrm\{emb\}\(x\)\)or representation use\(emb⁡\(x\)→y\)\(\\mathrm\{emb\}\(x\)\\to y\)\. Therefore, we propose that a falsifiable and clear scientific formulation of the LRH should be stated as follows:

###### Definition 4\(LRH; proposed\)\.

Fix a modelMM, representation locationll, evaluation distribution𝒟\\mathcal\{D\}over an input space𝒳\\mathcal\{X\}, and feature specification

ℱ=\(𝒴f,yf,ℓf,ε\),yf:𝒳→𝒴f,\\mathcal\{F\}=\(\\mathcal\{Y\}\_\{f\},y\_\{f\},\\ell\_\{f\},\\varepsilon\),\\qquad y\_\{f\}:\\mathcal\{X\}\\to\\mathcal\{Y\}\_\{f\},whereℓf\\ell\_\{f\}is the evaluation loss andε≥0\\varepsilon\\geq 0its tolerance\. LetembM\(l\)​\(x\)∈ℝd\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\)\\in\\mathbb\{R\}^\{d\}denote the representation ofxxatll\.

- \(i\)Representation formation\.ℱ\\mathcal\{F\}is linearly represented atllinMMover𝒟\\mathcal\{D\}iff ∃W,b:\\displaystyle\\exists\\,W,b:𝔼x∼𝒟​\[ℓf​\(W​embM\(l\)​\(x\)\+b,yf​\(x\)\)\]≤ε\.\\displaystyle\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\\\!\\left\[\\ell\_\{f\}\\\!\\left\(W\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\)\\\!\+\\\!b,\\;y\_\{f\}\(x\)\\right\)\\\!\\right\]\\\!\\leq\\varepsilon\.
- \(ii\)Representation use\.Given a linear representation satisfying \(i\), letℐf\\mathcal\{I\}\_\{f\}be a specified intervention on the representation atll\. The representation is*causally used*byMMover𝒟\\mathcal\{D\}if replacingembM\(l\)​\(x\)\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\)withℐf​\(embM\(l\)​\(x\)\)\\mathcal\{I\}\_\{f\}\(\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\)\)produces the corresponding counterfactual change in the model’s output behavior\.

Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)unifies the different ways the LRH has been studied in the literature, covering both representation formation and causal use\. For representation formation, standard multiclass linear probing accuracy can be expressed within Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)by taking𝒴f=\[K\]\\mathcal\{Y\}\_\{f\}=\[K\]and writingz⁡\(x\)=W​embM\(l\)​\(x\)\+b∈ℝK\.z\(x\)=W\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\)\+b\\in\\mathbb\{R\}^\{K\}\.With the00–11evaluation lossℓf\(z,y\)=𝕀\[argmaxkzk≠y\]\\ell\_\{f\}\(z,y\)=\\mathbb\{I\}\[\\arg\\max\_\{k\}z\_\{k\}\\neq y\],

𝔼x∼𝒟​\[ℓf​\(z⁡\(x\),yf​\(x\)\)\]=1−Accuracy𝒟⁡\(z\)\.\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\}\[\\ell\_\{f\}\(z\(x\),y\_\{f\}\(x\)\)\]=1\-\\operatorname\{Accuracy\}\_\{\\mathcal\{D\}\}\(z\)\.Scalar\-valued targetsyf​\(x\)∈ℝy\_\{f\}\(x\)\\in\\mathbb\{R\}include sparse\-autoencoder feature activations, while vector\-valued targetsyf​\(x\)∈ℝdy\_\{f\}\(x\)\\in\\mathbb\{R\}^\{d\}recover the linear\-relation setting of[Merullo et al\. \(2025\)](https://arxiv.org/html/2609.22695#bib.bib109), whereyf​\(x\)=emb​\(o\)y\_\{f\}\(x\)=\\mathrm\{emb\}\(o\)\.

For representation use, takingℐf\\mathcal\{I\}\_\{f\}to add a concept direction to the context representation recovers the intervention of[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12), where the intervention changes the target concept in the model’s output while leaving causally separable off\-target concepts unchanged\. Similarly, forhb=embM\(l\)​\(xbase\)h\_\{b\}=\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\_\{\\mathrm\{base\}\}\)andhs=embM\(l\)​\(xsource\)h\_\{s\}=\\mathrm\{emb\}^\{\(l\)\}\_\{M\}\(x\_\{\\mathrm\{source\}\}\), the intervention used by DAS\([Geiger et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib11)\)can be written asℐf​\(hb\)=R⊤​\[\(Id−Π\)​R​hb\+Π​R​hs\],\\mathcal\{I\}\_\{f\}\(h\_\{b\}\)=R^\{\\top\}\\\!\\left\[\(I\_\{d\}\-\\Pi\)Rh\_\{b\}\+\\Pi Rh\_\{s\}\\right\],whereRRis the learned orthogonal rotation andΠ\\Piis a fixed orthogonal projector onto a pre\-specified subspace of the rotated representation, whose dimension is chosen in advance as a hyperparameter\.

### 5\.3From Slogan to Statement

Would there be merit in formulating the hypothesis as in Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)? One reason this question matters becomes clear from Table[3](https://arxiv.org/html/2609.22695#S2.T3): the LRH is frequently used as experimental background for employing specific methods \(e\.g\., linear probing or sparse autoencoder with a single\-layer encoder\)\. However, this practice is a double\-edged sword\. On the one hand, building methods on prior work is essential for cumulative scientific progress\. On the other hand, grounding methodological justification in different layers or unrelated linguistic features can create an illusion that a choice of methods is theoretically justified, even when no shared justification actually exists\. In particular, when it is unclear whether different studies are referring to the same feature and layer, using the LRH as a slogan to motivate linear methodologies leads to claims that are unfalsifiable and thus, borrowing Wolfgang Pauli’s famous phrase,*not even wrong*\([Peierls and Peierls, 1960](https://arxiv.org/html/2609.22695#bib.bib18)\)\.

Then what is currently built on top of the LRH? Table[3](https://arxiv.org/html/2609.22695#S2.T3)also sheds light on this issue by indicating which lines of work would be affected if the LRH is falsified for a particular combination of model, feature, and layer\. For studies that invoke the LRH as a theoretical assumption \(Type 7in Table[3](https://arxiv.org/html/2609.22695#S2.T3)\), the implications are straightforward: once the LRH is falsified for a given model, feature, or layer, the corresponding theoretical claims no longer hold in that setting\. For studies that rely on the LRH as a methodological justification \(Type 8in Table[3](https://arxiv.org/html/2609.22695#S2.T3)\), while their empirical findings remain valid as observations about specific models and datasets, the methodological justification becomes weaker if the LRH is assumed without being tested\. In particular, when non\-linear representations are not explored as alternatives, one cannot rule out the possibility that more appropriate non\-linear methods would yield stronger or more informative results\([White et al\., 2021](https://arxiv.org/html/2609.22695#bib.bib10)\)\.

## 6Open problems

Building on the preceding review of the literature, we highlight two open problems that are well\-formulated and warrant further investigation\.

#### The relationship between layer depth and linear representation

As in Section[3](https://arxiv.org/html/2609.22695#S3), existing research has theoretically examined linear representations of various features and has shown that cross\-entropy loss and softmax can promote a linear structure in the final layer of neural networks\([Park et al\., 2024c](https://arxiv.org/html/2609.22695#bib.bib12);[Jiang et al\., 2024](https://arxiv.org/html/2609.22695#bib.bib105);[Park et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib15)\)\. However, it remains unclear why empirical studies consistently find stronger linear representations in intermediate layers than at the final layer \(Section[4](https://arxiv.org/html/2609.22695#S4)\)\.

#### Why do linear representations arise only for certain\(M,l,f,𝒟\)\(M,l,f,\\mathcal\{D\}\)tuples?

One of the earliest proposed explanations is that neural networks may prefer to retrieve information “cheaply”\([Henighan and Olah, 2023](https://arxiv.org/html/2609.22695#bib.bib104)\), though this remains largely speculative\. Another possibility is that linear representations arise from co\-occurrence in pre\-training data\([Merullo et al\., 2025](https://arxiv.org/html/2609.22695#bib.bib109)\)\. However, whether this is the only factor, and whether the relationship extends beyond correlation to causality, remains an open question\. More fundamentally, using Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4), we currently lack an explanation of why a featureffexhibits a linear representation for some tuples\(M,l,f,𝒟\)\(M,l,f,\\mathcal\{D\}\)but not for others\.

## 7Conclusion

In this paper, we have examined the Linear Representation Hypothesis \(LRH\)\. We have shown that definitions of the LRH have evolved across prior works and are not always consistent\. In particular, existing studies often draw on results obtained across different layers or feature types\. We propose that any LRH claim should explicitly specify the model, representation location, target feature, and evaluation distribution, as in Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)\. Based on a systematic categorization of prior work, we have identified several open problems for future research\.

## Limitations

We deliberately focus on literature that explicitly frames its contributions around the linear representation*hypothesis*\. While this choice provides coherence when organizing an increasingly large body of work, it necessarily excludes many earlier and influential studies that investigate linear structure but do not connect their findings to the LRH\. A notable example is[Hewitt and Manning \(2019\)](https://arxiv.org/html/2609.22695#bib.bib78), which demonstrates that distances in syntactic parse trees can be recovered by a linear probe\. In that work, linearity is not introduced to support a hypothesis, but rather as a methodological constraint intended to prevent the probe from learning the syntactic parsing task itself\. Such studies fall outside the scope of our survey, although they have inspired later research in this literature\. For a more detailed discussion of representative work in this broader literature and its methodological connections to the LRH, see Appendix[B](https://arxiv.org/html/2609.22695#A2)\.

## Acknowledgments

We would like to thank Rainer Engelken for introducing us to intriguing issues related to the linear representation hypothesis and for providing inspiration in shaping the direction of this work, and Jinu Lee for helpful feedback on this manuscript\.

## References

- Agarwalet al\.\(2025\)I\. Agarwal, S\. Navani, and F\. BarezContext matters: analyzing the generalizability of linear probing and steering across diverse scenarios\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=H5sbfvEbTh)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.12.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p2.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Alain and Bengio \(2017\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.In5th International Conference on Learning Representations, Workshop Track Proceedings,External Links:[Link](https://openreview.net/forum?id=ryF7rTqgl)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px1.p1.1)\.
- Alger \(2022\)B\. E\. AlgerNeuroscience needs to test both statistical and scientific hypotheses\.Journal of Neuroscience42\(45\),pp\. 8432–8438\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.1134-22.2022),ISSN 0270\-6474,[Link](https://www.jneurosci.org/content/42/45/8432),https://www\.jneurosci\.org/content/42/45/8432\.full\.pdfCited by:[§5](https://arxiv.org/html/2609.22695#S5.p1.1)\.
- Anthropic \(2024\)AnthropicThe Claude 3 model family: opus, sonnet, haiku\.Note:Model cardExternal Links:[Link](https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf)Cited by:[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1)\.
- Ayonrindeet al\.\(2024\)K\. Ayonrinde, M\. T\. Pearce, and L\. SharkeyInterpretability as compression: reconsidering sae explanations of neural activations\.InNeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning,Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Baek and Tegmark \(2025\)D\. D\. Baek and M\. TegmarkTowards understanding distilled reasoning models: a representational approach\.Building Trust Workshop at ICLR 2025\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Bhallaet al\.\(2024\)U\. Bhalla, A\. Oesterling, S\. Srinivas, F\. P\. Calmon, and H\. LakkarajuInterpreting CLIP with sparse linear concept embeddings \(SpLiCE\)\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 84298–84328\.External Links:[Document](https://dx.doi.org/10.52202/079017-2678),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/996bef37d8a638f37bdfcac2789e835d-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Bidermanet al\.\(2023\)S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff, A\. Skowron, L\. Sutawika, and O\. Van Der WalPythia: a suite for analyzing large language models across training and scaling\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 2397–2430\.External Links:[Link](https://proceedings.mlr.press/v202/biderman23a.html)Cited by:[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1)\.
- Boix\-Adsera \(2024\)E\. Boix\-AdseraTowards a theory of model distillation\.arXiv preprint arXiv:2403\.09053\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.8.2.1.1)\.
- Canbyet al\.\(2025\)M\. E\. Canby, A\. Davies, C\. Rastogi, and J\. HockenmaierHow reliable are causal probing interventions?\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,Mumbai, India,pp\. 857–878\.External Links:[Link](https://aclanthology.org/2025.ijcnlp-long.47/),[Document](https://dx.doi.org/10.18653/v1/2025.ijcnlp-long.47),ISBN 979\-8\-89176\-298\-5Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Chaninet al\.\(2024\)D\. Chanin, J\. Wilken\-Smith, T\. Dulka, H\. Bhatnagar, and J\. I\. BloomA is for absorption: studying feature splitting and absorption in sparse autoencoders\.InInterpretable AI: Past, Present and Future,External Links:[Link](https://openreview.net/forum?id=Wzav8fesTL)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Chenet al\.\(2025\)A\. Chen, J\. Merullo, A\. Stolfo, and E\. PavlickTransferring linear features across language models with model stitching\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 54114–54146\.External Links:[Document](https://dx.doi.org/10.52202/085713-1621),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/4569a868e7aa891248832ec08445d071-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Chenet al\.\(2024\)X\. Chen, Z\. Zhu, and A\. PerraultUnderstanding learned representations and action collapse in visual reinforcement learning\.InReinforcement Learning Journal,Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Cho and Hockenmaier \(2025\)I\. Cho and J\. HockenmaierToward efficient sparse autoencoder\-guided steering for improved in\-context learning in large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 28961–28973\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1474/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1474),ISBN 979\-8\-89176\-332\-6Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Costaet al\.\(2025\)V\. Costa, T\. Fel, E\. S\. Lubana, B\. Tolooshams, and D\. BaFrom flat to hierarchical: extracting sparse representations with matching pursuit\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 38625–38668\.External Links:[Document](https://dx.doi.org/10.52202/085713-1154),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/311bd32bcd952f93cf820bacd6955f0a-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Csordáset al\.\(2024\)R\. Csordás, C\. Potts, C\. D\. Manning, and A\. GeigerRecurrent neural networks learn to store and generate sequences using non\-linear representations\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 248–262\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.17/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.17)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.14.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px1.p2.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1)\.
- Dalvaet al\.\(2025\)Y\. Dalva, K\. Venkatesh, and P\. YanardagFluxSpace: disentangled semantic editing in rectified flow models\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 13083–13092\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Dosovitskiyet al\.\(2021\)A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. HoulsbyAn image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=YicbFdNTTy)Cited by:[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px1.p1.1)\.
- Elhageet al\.\(2022\)N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. OlahToy models of superposition\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2022/toy_model/index.html)Cited by:[Table 1](https://arxiv.org/html/2609.22695#S1.T1),[§2\.1](https://arxiv.org/html/2609.22695#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1),[Definition 1](https://arxiv.org/html/2609.22695#Thmdefinition1)\.
- Engelset al\.\(2025a\)J\. Engels, E\. J\. Michaud, I\. Liao, W\. Gurnee, and M\. TegmarkNot all language model features are one\-dimensionally linear\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/d3221cdb27e49d9c1cd35ad254feccfe-Abstract-Conference.html)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.15.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px1.p2.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1)\.
- Engelset al\.\(2025b\)J\. Engels, L\. R\. Smith, and M\. TegmarkDecomposing the dark matter of sparse autoencoders\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=sXq3Wb3vef)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Geigeret al\.\(2025\)A\. Geiger, D\. Ibeling, A\. Zur, M\. Chaudhary, S\. Chauhan, J\. Huang, A\. Arora, Z\. Wu, N\. Goodman, C\. Potts, and T\. IcardCausal abstraction: a theoretical foundation for mechanistic interpretability\.Journal of Machine Learning Research26\(83\),pp\. 1–64\.External Links:[Link](http://jmlr.org/papers/v26/23-0058.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Geigeret al\.\(2021\)A\. Geiger, H\. Lu, T\. Icard, and C\. PottsCausal abstractions of neural networks\.Advances in Neural Information Processing Systems34,pp\. 9574–9586\.Cited by:[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p2.1)\.
- Geigeret al\.\(2024\)A\. Geiger, Z\. Wu, C\. Potts, T\. Icard, and N\. GoodmanFinding alignments between interpretable causal variables and distributed neural representations\.InCausal Learning and Reasoning,pp\. 160–187\.Cited by:[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p5.1)\.
- Gladkovaet al\.\(2016\)A\. Gladkova, A\. Drozd, and S\. MatsuokaAnalogy\-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t\.\.InProceedings of the NAACL Student Research Workshop,San Diego, California,pp\. 8–15\.External Links:[Link](https://aclanthology.org/N16-2002/),[Document](https://dx.doi.org/10.18653/v1/N16-2002)Cited by:[§1](https://arxiv.org/html/2609.22695#S1.p2.1),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px1.p2.1)\.
- Golechhaet al\.\(2024\)S\. Golechha, L\. Bushnaq, and NandiIntricacies of feature geometry in large language models\(Website\)Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/eczwWrmX5XNEo7JsS/intricacies-of-feature-geometry-in-large-language-models)Cited by:[footnote 2](https://arxiv.org/html/2609.22695#footnote2)\.
- Golechha and Nandi \(2024\)S\. Golechha and NandiThe geometry of feelings and nonsense in large language models\(Website\)Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/C8LZ3DW697xcpPaqC/the-geometry-of-feelings-and-nonsense-in-large-language)Cited by:[footnote 2](https://arxiv.org/html/2609.22695#footnote2)\.
- Gonzálezet al\.\(2007\)A\. González, G\. Russel, A\. Márquez, J\. A\. Moreno, C\. Garcia, C\. Dominguez, O\. Colmenares, and J\. J\. MachadoSupervised farm classification from remote sensing images based on kernel adatron algorithm\.In2007 IEEE International Geoscience and Remote Sensing Symposium,pp\. 3345–3348\.Cited by:[§2\.1](https://arxiv.org/html/2609.22695#S2.SS1.p1.1)\.
- Guoet al\.\(2025\)Z\. Guo, C\. Xue, Z\. Xu, H\. Bo, Y\. Ye, J\. B\. Pierrehumbert, and M\. LewisQuantifying compositionality of classic and state\-of\-the\-art embeddings\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 22130–22146\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.6.2.1.1),[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px2.p1.1)\.
- Gurnee and Tegmark \(2024\)W\. Gurnee and M\. TegmarkLanguage models represent space and time\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=jE8xbmvFin)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.8.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p2.1)\.
- Heet al\.\(2016\)K\. He, X\. Zhang, S\. Ren, and J\. SunDeep residual learning for image recognition\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 770–778\.Cited by:[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px1.p1.1)\.
- Heinzerling and Inui \(2024\)B\. Heinzerling and K\. InuiMonotonic representation of numeric attributes in language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 175–195\.External Links:[Link](https://aclanthology.org/2024.acl-short.18/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-short.18)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Henighan and Olah \(2023\)T\. Henighan and C\. OlahDictionary learning worries\.External Links:[Link](https://transformer-circuits.pub/2023/may-update/index.html#dictionary-worries)Cited by:[§6](https://arxiv.org/html/2609.22695#S6.SS0.SSS0.Px2.p1.1)\.
- Hernandezet al\.\(2024\)E\. Hernandez, A\. S\. Sharma, T\. Haklay, K\. Meng, M\. Wattenberg, J\. Andreas, Y\. Belinkov, and D\. BauLinearity of relation decoding in transformer language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=w7LU2s14kE)Cited by:[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p3.2),[footnote 3](https://arxiv.org/html/2609.22695#footnote3)\.
- Hewitt and Liang \(2019\)J\. Hewitt and P\. LiangDesigning and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2733–2743\.External Links:[Link](https://aclanthology.org/D19-1275/),[Document](https://dx.doi.org/10.18653/v1/D19-1275)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px1.p2.1)\.
- Hewitt and Manning \(2019\)J\. Hewitt and C\. D\. ManningA structural probe for finding syntax in word representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4129–4138\.External Links:[Link](https://aclanthology.org/N19-1419/),[Document](https://dx.doi.org/10.18653/v1/N19-1419)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px1.p3.1),[Limitations](https://arxiv.org/html/2609.22695#Sx1.p1.1)\.
- Hubenet al\.\(2024\)R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. SharkeySparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=F76bwRSLeK)Cited by:[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1)\.
- Hübotteret al\.\(2025\)J\. Hübotter, P\. Wolf, A\. Shevchenko, D\. Jüni, A\. Krause, and G\. KurSpecialization after generalization: towards understanding test\-time training in foundation models\.39th Conference on Neural Information Processing Systems \(NeurIPS 2025\) Workshop on Continual and Compatible Foundation Model Updates \(CCFM\)\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.8.2.1.1)\.
- Idrissiet al\.\(2023\)B\. Y\. Idrissi, D\. Bouchacourt, R\. Balestriero, I\. Evtimov, C\. Hazirbas, N\. Ballas, P\. Vincent, M\. Drozdzal, D\. Lopez\-Paz, and M\. IbrahimImageNet\-X: understanding model mistakes with factor of variation annotations\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HXz7Vcm3VgM)Cited by:[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px1.p1.1)\.
- Jianget al\.\(2024\)Y\. Jiang, G\. Rajendran, P\. K\. Ravikumar, B\. Aragam, and V\. VeitchOn the origins of linear representations in large language models\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 21879–21911\.External Links:[Link](https://proceedings.mlr.press/v235/jiang24d.html)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.7.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p3.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.2.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px1.p2.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1),[§6](https://arxiv.org/html/2609.22695#S6.SS0.SSS0.Px1.p1.1)\.
- Kalajdzievski \(2025\)D\. KalajdzievskiThe logical implication steering method for conditional interventions on transformer generation\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 28689–28720\.External Links:[Link](https://proceedings.mlr.press/v267/kalajdzievski25a.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Kantamneni and Tegmark \(2025\)S\. Kantamneni and M\. TegmarkLanguage models use trigonometry to do addition\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,External Links:[Link](https://openreview.net/forum?id=CqViN4dQJk)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.16.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Karvonenet al\.\(2024\)A\. Karvonen, B\. Wright, C\. Rager, R\. Angell, J\. Brinkmann, L\. Smith, C\. Mayrink Verdun, D\. Bau, and S\. MarksMeasuring progress in dictionary learning for language model interpretability with board game models\.Advances in Neural Information Processing Systems37,pp\. 83091–83118\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Kimet al\.\(2025\)J\. Kim, J\. Evans, and A\. ScheinLinear representations of political perspective emerge in large language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/1557cc3e1a67f985dbdb162db9963dd0-Abstract-Conference.html)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.11.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p2.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Kramáret al\.\(2026\)J\. Kramár, J\. Engels, Z\. Wang, B\. Chughtai, R\. Shah, N\. Nanda, and A\. ConmyBuilding production\-ready probes for gemini\.arXiv preprint arXiv:2601\.11516\.Cited by:[§A\.3](https://arxiv.org/html/2609.22695#A1.SS3.p6.1)\.
- Lai and Hockenmaier \(2017\)A\. Lai and J\. HockenmaierLearning to predict denotational probabilities for modeling entailment\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers,M\. Lapata, P\. Blunsom, and A\. Koller \(Eds\.\),Valencia, Spain,pp\. 721–730\.External Links:[Link](https://aclanthology.org/E17-1068/)Cited by:[footnote 1](https://arxiv.org/html/2609.22695#footnote1)\.
- Leeet al\.\(2025\)S\. Lee, A\. Davies, M\. E\. Canby, and J\. HockenmaierEvaluating and designing sparse autoencoders by approximating quasi\-orthogonality\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=XhdNFeMclS)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.8.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Lei and Cooper \(2025a\)G\. Lei and S\. J\. CooperDo llamas understand the periodic table?\.Digital Discovery4\(12\),pp\. 3455–3465\.Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.17.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Lei and Cooper \(2025b\)G\. Lei and S\. J\. CooperThe representation and recall of interwoven structured knowledge in llms: a geometric and layered analysis\.arXiv preprint arXiv:2502\.10871\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Liet al\.\(2023\)K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. WattenbergInference\-time intervention: eliciting truthful answers from a language model\.Advances in Neural Information Processing Systems36,pp\. 41451–41530\.Cited by:[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Liet al\.\(2025a\)Y\. Li, Z\. Fan, R\. Chen, X\. Gai, L\. Gong, Y\. Zhang, and Z\. LiuFairSteer: inference time debiasing for LLMs with dynamic activation steering\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 11293–11312\.External Links:[Link](https://aclanthology.org/2025.findings-acl.589/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.589),ISBN 979\-8\-89176\-256\-5Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Liet al\.\(2025b\)Y\. Li, E\. J\. Michaud, D\. D\. Baek, J\. Engels, X\. Sun, and M\. TegmarkThe geometry of concepts: sparse autoencoder feature structure\.Entropy27\(4\),pp\. 344\.External Links:[Document](https://dx.doi.org/10.3390/e27040344),[Link](https://www.mdpi.com/1099-4300/27/4/344),ISSN 1099\-4300Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1),[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px2.p1.1)\.
- Liuet al\.\(2026\)Y\. Liu, D\. Gong, Y\. Cai, E\. Gao, Z\. Zhang, B\. Huang, M\. Gong, A\. van den Hengel, and J\. Q\. ShiI predict therefore i am: is next token prediction enough to learn human\-interpretable concepts from data?\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/hash/db7c4a8a0c1bf8b405f21c783bbf9eae-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.2.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Luet al\.\(2026\)Y\. Lu, Y\. Liu, and H\. SchützeRelational linearity is a predictor of hallucinations\.arXiv preprint arXiv:2601\.11429\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Makelovet al\.\(2024\)A\. Makelov, G\. Lange, A\. Geiger, and N\. NandaIs this the subspace you are looking for? an interpretability illusion for subspace activation patching\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Ebt7JgMHv1)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Makelovet al\.\(2025\)A\. Makelov, G\. Lange, and N\. NandaTowards principled evaluations of sparse autoencoders for interpretability and control\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1Njl73JKjB)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Marks and Tegmark \(2024\)S\. Marks and M\. TegmarkThe geometry of truth: emergent linear structure in large language model representations of true/false datasets\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=aajyHYjjsk)Cited by:[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Marshallet al\.\(2024\)T\. Marshall, A\. Scherlis, and N\. BelroseRefusal in LLMs is an affine function\.External Links:2411\.09003,[Link](https://arxiv.org/abs/2411.09003)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Mencattiniet al\.\(2026\)T\. Mencattini, R\. Cadei, and F\. LocatelloExploratory causal inference in SAEnce\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Ml8t8kQMUP)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Merulloet al\.\(2025\)J\. Merullo, N\. A\. Smith, S\. Wiegreffe, and Y\. ElazarOn linear representations and pretraining data frequency in language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=EDoD3DgivF)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.13.5.1.1),[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.19.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.5.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p3.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p3.2),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p4.2),[§6](https://arxiv.org/html/2609.22695#S6.SS0.SSS0.Px2.p1.1)\.
- Mikolovet al\.\(2013\)T\. Mikolov, W\. Yih, and G\. ZweigLinguistic regularities in continuous space word representations\.InProceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Atlanta, Georgia,pp\. 746–751\.External Links:[Link](https://aclanthology.org/N13-1090/)Cited by:[§1](https://arxiv.org/html/2609.22695#S1.p1.1)\.
- Minsky and Papert \(1968\)M\. Minsky and S\. A\. PapertLinear separation and learning\.External Links:[Link](http://hdl.handle.net/1721.1/6170)Cited by:[§A\.3](https://arxiv.org/html/2609.22695#A1.SS3.p1.1)\.
- Modellet al\.\(2025\)A\. Modell, P\. Rubin\-Delanchy, and N\. WhiteleyThe origins of representation manifolds in large language models\.arXiv preprint arXiv:2505\.18235\.Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.15.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1)\.
- Muraleedharan \(2025\)A\. MuraleedharanOn the limits of linear representation hypotheses in large language models: a dynamical systems analysis\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=fetTz1Xs2l)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Nandaet al\.\(2023\)N\. Nanda, A\. Lee, and M\. WattenbergEmergent linear representations in world models of self\-supervised sequence models\.InProceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, S\. Hao, J\. Jumelet, N\. Kim, A\. McCarthy, and H\. Mohebbi \(Eds\.\),Singapore,pp\. 16–30\.External Links:[Link](https://aclanthology.org/2023.blackboxnlp-1.2/),[Document](https://dx.doi.org/10.18653/v1/2023.blackboxnlp-1.2)Cited by:[§A\.3](https://arxiv.org/html/2609.22695#A1.SS3.p6.1),[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.2.5.1.1),[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p1.1)\.
- Nguyen and Leng \(2025\)T\. Nguyen and Y\. LengToward a flexible framework for linear representation hypothesis using maximum likelihood estimation\.arXiv preprint arXiv:2502\.16385\.Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.9.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1)\.
- Nickel and Kiela \(2017\)M\. Nickel and D\. KielaPoincaré embeddings for learning hierarchical representations\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 6338–6347\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/hash/59dfa2df42d9e3d41f5b02bfc32229dd-Abstract.html)Cited by:[footnote 1](https://arxiv.org/html/2609.22695#footnote1)\.
- Olah \(2024\)C\. OlahThe next five hurdles\.Note:[https://transformer\-circuits\.pub/2024/july\-update/index\.html\#hurdles](https://transformer-circuits.pub/2024/july-update/index.html#hurdles)Circuits Updates — July 2024, Transformer Circuits ThreadExternal Links:[Link](https://transformer-circuits.pub/2024/july-update/index.html#hurdles)Cited by:[§2\.1](https://arxiv.org/html/2609.22695#S2.SS1.p3.1)\.
- Parket al\.\(2024a\)K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. VeitchThe geometry of categorical and hierarchical concepts in large language models\.InICML 2024 Workshop on Mechanistic Interpretability,External Links:[Link](https://openreview.net/forum?id=KXuYjuBzKo)Cited by:[footnote 2](https://arxiv.org/html/2609.22695#footnote2)\.
- Parket al\.\(2024b\)K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. VeitchThe geometry of categorical and hierarchical concepts in large language models\.External Links:2406\.01506,[Link](https://arxiv.org/pdf/2406.01506v2)Cited by:[footnote 2](https://arxiv.org/html/2609.22695#footnote2)\.
- Parket al\.\(2025\)K\. Park, Y\. J\. Choe, Y\. Jiang, and V\. VeitchThe geometry of categorical and hierarchical concepts in large language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=bVTM2QKYuA)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.10.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p3.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.2.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1),[§6](https://arxiv.org/html/2609.22695#S6.SS0.SSS0.Px1.p1.1)\.
- Parket al\.\(2024c\)K\. Park, Y\. J\. Choe, and V\. VeitchThe linear representation hypothesis and the geometry of large language models\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 39643–39666\.External Links:[Link](https://proceedings.mlr.press/v235/park24c.html)Cited by:[§A\.1](https://arxiv.org/html/2609.22695#A1.SS1),[§A\.1](https://arxiv.org/html/2609.22695#A1.SS1.p1.1),[§A\.1](https://arxiv.org/html/2609.22695#A1.SS1.p6.1),[§A\.1](https://arxiv.org/html/2609.22695#A1.SS1.p7.1),[§A\.3](https://arxiv.org/html/2609.22695#A1.SS3.p2.1),[§A\.3](https://arxiv.org/html/2609.22695#A1.SS3.p5.1),[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.6.5.1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px2.p2.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p2.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p3.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.2.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px1.p2.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p2.1),[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.SSS0.Px2.p5.1),[§6](https://arxiv.org/html/2609.22695#S6.SS0.SSS0.Px1.p1.1),[Definition 3](https://arxiv.org/html/2609.22695#Thmdefinition3),[footnote 2](https://arxiv.org/html/2609.22695#footnote2)\.
- Pauloet al\.\(2025\)G\. S\. Paulo, A\. T\. Mallen, C\. Juang, and N\. BelroseAutomatically interpreting millions of features in large language models\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=EemtbhJOXc)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Peierls and Peierls \(1960\)R\. E\. Peierls and R\. E\. PeierlsWolfgang ernst pauli, 1900\-1958\.Biographical Memoirs of Fellows of the Royal Society5\(1\),pp\. 174–192\.External Links:ISSN 0080\-4606,[Document](https://dx.doi.org/10.1098/rsbm.1960.0014),[Link](https://doi.org/10.1098/rsbm.1960.0014),https://royalsocietypublishing\.org/rsbm/article\-pdf/5/1/174/444938/rsbm\.1960\.0014\.pdfCited by:[§5\.3](https://arxiv.org/html/2609.22695#S5.SS3.p1.1)\.
- Penningtonet al\.\(2014\)J\. Pennington, R\. Socher, and C\. ManningGloVe: global vectors for word representation\.InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Doha, Qatar,pp\. 1532–1543\.External Links:[Link](https://aclanthology.org/D14-1162/),[Document](https://dx.doi.org/10.3115/v1/D14-1162)Cited by:[§1](https://arxiv.org/html/2609.22695#S1.p2.1)\.
- Pimentelet al\.\(2020a\)T\. Pimentel, N\. Saphra, A\. Williams, and R\. CotterellPareto probing: Trading off accuracy for complexity\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 3138–3153\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.254/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.254)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px2.p2.1)\.
- Pimentelet al\.\(2020b\)T\. Pimentel, J\. Valvoda, R\. H\. Maudslay, R\. Zmigrod, A\. Williams, and R\. CotterellInformation\-theoretic probing for linguistic structure\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4609–4622\.External Links:[Link](https://aclanthology.org/2020.acl-main.420/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.420)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px2.p1.1)\.
- Pitzet al\.\(1976\)G\. F\. Pitz, L\. S\. Leung, C\. Hamilos, and W\. TerpeningThe use of probabilistic information in making predictions\.Organizational Behavior and Human Performance17\(1\),pp\. 1–18\.Cited by:[§2\.1](https://arxiv.org/html/2609.22695#S2.SS1.p1.1)\.
- Popper \(1959\)K\. PopperThe logic of scientific discovery\.Julius Springer, Hutchinson & Co\.Cited by:[§5\.2](https://arxiv.org/html/2609.22695#S5.SS2.p1.1)\.
- Qiuet al\.\(2025\)H\. Qiu, Y\. Wu, D\. Li, J\. Guo, and Q\. YaoSuperpose task\-specific features for model merging\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 4200–4214\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.210/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.210),ISBN 979\-8\-89176\-332\-6Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Reizingeret al\.\(2025\)P\. Reizinger, A\. Bizeul, A\. Juhos, J\. E\. Vogt, R\. Balestriero, W\. Brendel, and D\. KlindtCross\-entropy is all you need to invert the data generating process\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=hrqNOxpItr)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.6.2.1.1),[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.22695#S5.SS1.SSS0.Px2.p1.1)\.
- Shenet al\.\(2025\)W\. F\. Shen, X\. Qiu, M\. Kurmanji, A\. Iacob, L\. Sani, Y\. Chen, N\. Cancedda, and N\. D\. LaneLLM unlearning via neural activation redirection\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 49474–49511\.External Links:[Document](https://dx.doi.org/10.52202/085713-1475),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/3ed6a903be20e1478aa6793f14938091-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Skeanet al\.\(2024\)O\. Skean, M\. R\. Arefin, and R\. Shwartz\-ZivDoes representation matter? exploring intermediate layers in large language models\.InWorkshop on Machine Learning and Compression, NeurIPS 2024,External Links:[Link](https://openreview.net/forum?id=FN0tZ9pVLz)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Skeanet al\.\(2025\)O\. Skean, M\. R\. Arefin, D\. Zhao, N\. N\. Patel, J\. Naghiyev, Y\. Lecun, and R\. Shwartz\-ZivLayer by layer: uncovering hidden representations in language models\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 55854–55875\.External Links:[Link](https://proceedings.mlr.press/v267/skean25a.html)Cited by:[§4](https://arxiv.org/html/2609.22695#S4.p2.1)\.
- Smith \(2024\)L\. SmithThe “strong” feature hypothesis could be wrong\.Note:AI Alignment ForumExternal Links:[Link](https://www.alignmentforum.org/posts/tojtPCCRpKLSHBdpn/the-strong-feature-hypothesis-could-be-wrong)Cited by:[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px1.p1.1),[Definition 2](https://arxiv.org/html/2609.22695#Thmdefinition2)\.
- Songet al\.\(2025\)J\. Song, Z\. Xu, and Y\. ZhongOut\-of\-distribution generalization via composition: a lens through induction heads in transformers\.Proceedings of the National Academy of Sciences122\(6\),pp\. e2417182122\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2417182122),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2417182122),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2417182122Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Stevinsonet al\.\(2025\)E\. Stevinson, L\. Prieto, M\. Barsbey, and T\. BirdalAdversarial attacks leverage interference between features in superposition\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=LqI52GG2Ss)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.8.2.1.1)\.
- Sutteret al\.\(2025\)D\. Sutter, J\. Minder, T\. Hofmann, and T\. PimentelThe non\-linear representation dilemma: is causal abstraction enough for mechanistic interpretability?\.InAdvances in Neural Information Processing Systems,Vol\.38,pp\. 166009–166057\.External Links:[Document](https://dx.doi.org/10.52202/085713-5000),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/dbb98528c9870377f3f0d133aae6050b-Abstract-Conference.html)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Tanet al\.\(2024\)D\. Tan, D\. Chanin, A\. Lynch, B\. Paige, D\. Kanoulas, A\. Garriga\-Alonso, and R\. KirkAnalysing the generalisation and reliability of steering vectors\.Advances in Neural Information Processing Systems37,pp\. 139179–139212\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Tas and Wagner \(2024\)O\. S\. Tas and R\. WagnerWords in motion: extracting interpretable control vectors for motion transformers\.InInterpretable AI: Past, Present and Future,External Links:[Link](https://openreview.net/forum?id=jT1WiYTIHR)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Templetonet al\.\(2024\)A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. HenighanScaling monosemanticity: extracting interpretable features from Claude 3 Sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.4.5.1.1),[§2\.2](https://arxiv.org/html/2609.22695#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px2.p2.1),[§2\.3](https://arxiv.org/html/2609.22695#S2.SS3.SSS0.Px3.p3.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Tibliaset al\.\(2025\)F\. Tiblias, I\. Bigoulaeva, J\. Niu, S\. Balloccu, and I\. GurevychShape happens: automatic feature manifold discovery in LLMs via supervised multi\-dimensional scaling\.arXiv preprint arXiv:2510\.01025\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Tiggeset al\.\(2024\)C\. Tigges, O\. J\. Hollinsworth, A\. Geiger, and N\. NandaLanguage models linearly represent sentiment\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 58–87\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.5/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.5)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.5.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p2.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Tolooshamset al\.\(2025\)B\. Tolooshams, A\. Shen, and A\. AnandkumarSparse autoencoder neural operators: model recovery in function spaces\.InUniReps: 3rd Edition of the Workshop on Unifying Representations in Neural Models,External Links:[Link](https://openreview.net/forum?id=jfGBkkaeyY)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Turneret al\.\(2023\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidSteering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.3.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.3.2.1.1),[§3](https://arxiv.org/html/2609.22695#S3.SS0.SSS0.Px2.p2.1)\.
- Valoiset al\.\(2025\)P\. H\. Valois, D\. Satav, R\. A\. de Campos, G\. Q\. Pratamasunu, and K\. FukuiVision language model interpretability with concept guided decoding\.In2025 IEEE International Conference on Image Processing \(ICIP\),pp\. 397–402\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Vendrovet al\.\(2016\)I\. Vendrov, R\. Kiros, S\. Fidler, and R\. UrtasunOrder\-embeddings of images and language\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/1511.06361)Cited by:[footnote 1](https://arxiv.org/html/2609.22695#footnote1)\.
- Vilniset al\.\(2018\)L\. Vilnis, X\. Li, S\. Murty, and A\. McCallumProbabilistic embedding of knowledge graphs with box lattice measures\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 263–272\.External Links:[Document](https://dx.doi.org/10.18653/v1/P18-1025),[Link](https://aclanthology.org/P18-1025/)Cited by:[footnote 1](https://arxiv.org/html/2609.22695#footnote1)\.
- Voita and Titov \(2020\)E\. Voita and I\. TitovInformation\-theoretic probing with minimum description length\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 183–196\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.14/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.14)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px2.p2.1)\.
- Vompaet al\.\(2026\)E\. Vompa, T\. Tammet, and M\. VaishnavBeyond the linear separability ceiling: aligning representations in VLMs\.Transactions on Machine Learning Research\.Note:J2C CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=3uX4p80bN0)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Walkeret al\.\(2025\)T\. Walker, A\. I\. Humayun, R\. Balestriero, and R\. BaraniukCentroid affinity: how deep networks represent features\.InMechanistic Interpretability Workshop at NeurIPS 2025,External Links:[Link](https://openreview.net/forum?id=CLrq72YHgd)Cited by:[Table 2](https://arxiv.org/html/2609.22695#S1.T2.2.18.5.1.1),[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.4.2.1.1),[§4](https://arxiv.org/html/2609.22695#S4.p1.1)\.
- Wang and Xu \(2025\)Z\. Wang and C\. XuThoughtProbe: classifier\-guided LLM thought space exploration via probing representations\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 6018–6039\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.307/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.307),ISBN 979\-8\-89176\-332\-6Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Wattenberg and Viégas \(2024\)M\. Wattenberg and F\. B\. ViégasRelational composition in neural networks: a survey and call to action\.Workshop on Mechanistic Interpretability\.Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.7.2.1.1)\.
- Whiteet al\.\(2021\)J\. C\. White, T\. Pimentel, N\. Saphra, and R\. CotterellA non\-linear structural probe\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,K\. Toutanova, A\. Rumshisky, L\. Zettlemoyer, D\. Hakkani\-Tur, I\. Beltagy, S\. Bethard, R\. Cotterell, T\. Chakraborty, and Y\. Zhou \(Eds\.\),Online,pp\. 132–138\.External Links:[Link](https://aclanthology.org/2021.naacl-main.12/),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.12)Cited by:[Appendix B](https://arxiv.org/html/2609.22695#A2.SS0.SSS0.Px3.p1.1),[§5\.3](https://arxiv.org/html/2609.22695#S5.SS3.p2.1)\.
- Wuet al\.\(2024\)Z\. Wu, A\. Geiger, J\. Huang, A\. Arora, T\. Icard, C\. Potts, and N\. D\. GoodmanA reply to makelov et al\. \(2023\)’s “interpretability illusion” arguments\.External Links:2401\.12631,[Link](https://arxiv.org/abs/2401.12631)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Yaoet al\.\(2026\)Y\. Yao, H\. Zhang, and M\. DuAdaptiveK: complexity\-driven sparse autoencoders for interpretable language model representations\.InFindings of the Association for Computational Linguistics: ACL 2026,San Diego, California, United States,pp\. 23702–23728\.External Links:[Link](https://aclanthology.org/2026.findings-acl.1187/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1187),ISBN 979\-8\-89176\-395\-1Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Zentall \(2001\)T\. R\. ZentallThe case for a cognitive approach to animal learning and behavior\.Behavioural Processes54\(1–3\),pp\. 65–78\.Cited by:[§2\.1](https://arxiv.org/html/2609.22695#S2.SS1.p1.1)\.
- Zhanget al\.\(2025\)S\. Zhang, Y\. Zhai, K\. Guo, H\. Hu, S\. Guo, Z\. Fang, L\. Zhao, C\. Shen, C\. Wang, and Q\. WangJBShield: defending large language models from jailbreak attacks through activated concept analysis and manipulation\.In34th USENIX Security Symposium \(USENIX Security 25\),Seattle, WA,pp\. 8215–8234\.External Links:ISBN 978\-1\-939133\-52\-6,[Link](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-shenyi)Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.9.2.1.1)\.
- Zhuet al\.\(2025\)J\. Zhu, Z\. Li, C\. Squires, Q\. Wang, B\. Han, and P\. RavikumarOn the fragility of latent knowledge: layer\-wise influence under unlearning in large language model\.InICML 2025 Workshop on Machine Unlearning for Generative AI,Cited by:[Table 3](https://arxiv.org/html/2609.22695#S2.T3.2.8.2.1.1)\.

## Appendix AConnections Between Definitions

### A\.1[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)’s Approach

First, the three different notions of linear representation summarized by[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)can be expressed mathematically as follows using the example described in Eq\.[1](https://arxiv.org/html/2609.22695#S1.E1):

Subspace:

vmale→female≈emb⁡\("queen"\)−emb⁡\("king"\)\.v\_\{\\text\{male\}\\to\\text\{female\}\}\\approx\\mathrm\{emb\}\(\\text\{"queen"\}\)\-\\mathrm\{emb\}\(\\text\{"king"\}\)\.
Intervention:

emb⁡\("king"\)\+α​vmale→female≈emb⁡\("queen"\)\.\\mathrm\{emb\}\(\\text\{"king"\}\)\+\\alpha v\_\{\\text\{male\}\\to\\text\{female\}\}\\approx\\mathrm\{emb\}\(\\text\{"queen"\}\)\.
Measurement:444This can be viewed as a formalization of Definition[1](https://arxiv.org/html/2609.22695#Thmdefinition1), which states that “neural networks represent input features as directions in activation space\.” Under this formulation, the feature vector can be defined more flexibly than under the other two definitions, since it does not require the feature vector to correspond to a counterfactual concept\. For example, in Section[2\.2](https://arxiv.org/html/2609.22695#S2.SS2), when using a single\-layer sparse autoencoder, each feature is measured asemb​\(x\)⊤​wenc\\mathrm\{emb\}\(x\)^\{\\top\}w\_\{\\text\{enc\}\}, which implicitly adopts this definition as the underlying assumption for feature detection\. Importantly, “being the Golden Gate Bridge” is not itself a counterfactual concept, yet it can still be represented by a measurement of the formemb​\(x\)⊤​vGoldenGateBridge\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{GoldenGateBridge\}\}\.\.

feature⁡\(x\)=emb​\(x\)⊤​vfeature\.\\mathrm\{feature\}\(x\)=\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\}\.
Here,[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)restrictemb⁡\(x\)\\mathrm\{emb\}\(x\)to the final layer where the softmax is applied in order to derive the following relationship:

P⁡\(Y=y∣x\)\\displaystyle P\(Y=y\\mid x\)=softmax⁡\(emb​\(x\)⊤​unemb​\(y\)\)\\displaystyle=\\mathrm\{softmax\}\\\!\\left\(\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(y\)\\right\)=exp⁡\(emb​\(x\)⊤​unemb​\(y\)\)∑y′exp⁡\(emb​\(x\)⊤​unemb​\(y′\)\),\\displaystyle=\\frac\{\\exp\\\!\\left\(\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(y\)\\right\)\}\{\\sum\_\{y^\{\\prime\}\}\\exp\\\!\\left\(\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(y^\{\\prime\}\)\\right\)\},whereYYis a random variable representing the next token andxxdenotes the input text\. For example,x="He is the "x=\\text\{"He is the "\},Y⁡\(0\)="king"Y\(0\)=\\text\{"king"\}, andY⁡\(1\)="queen"Y\(1\)=\\text\{"queen"\}\. To measure how much more likely the outcomeY⁡\(1\)Y\(1\)is compared toY⁡\(0\)Y\(0\), consider the log\-odds:

log⁡P⁡\(Y⁡\(1\)∣x\)P⁡\(Y⁡\(0\)∣x\)=log⁡exp⁡\(emb​\(x\)⊤​unemb​\(Y⁡\(1\)\)\)exp⁡\(emb​\(x\)⊤​unemb​\(Y⁡\(0\)\)\)\\log\\frac\{P\(Y\(1\)\\mid x\)\}\{P\(Y\(0\)\\mid x\)\}=\\log\\frac\{\\exp\(\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(Y\(1\)\)\)\}\{\\exp\(\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(Y\(0\)\)\)\}=emb​\(x\)⊤​unemb​\(Y⁡\(1\)\)−emb​\(x\)⊤​unemb​\(Y⁡\(0\)\)=\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(Y\(1\)\)\-\\mathrm\{emb\}\(x\)^\{\\top\}\\mathrm\{unemb\}\(Y\(0\)\)=emb​\(x\)⊤​\(unemb⁡\(Y⁡\(1\)\)−unemb⁡\(Y⁡\(0\)\)\)=\\mathrm\{emb\}\(x\)^\{\\top\}\\left\(\\mathrm\{unemb\}\(Y\(1\)\)\-\\mathrm\{unemb\}\(Y\(0\)\)\\right\)=emb​\(x\)⊤​vfeature\.=\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\}\.
While[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)’s derivation is advantageous in thatvfeaturev\_\{\\text\{feature\}\}can be easily obtained from next\-token unembedding vectors and is more theoretically tractable, the three definitions can be unified without restricting the analysis to the final layer, as will be described in the next section\.

### A\.2Beyond the Final Layer

Formally, if the subspace exists, the definition of intervention follows trivially withα=1\\alpha=1\. To see the connection with the measurement definition, suppose we apply an intervention to the embedding of an input textxx, producing a modified embeddingemb​\(x\)′=emb⁡\(x\)\+α​vfeature\.\\mathrm\{emb\}\(x\)^\{\\prime\}=\\mathrm\{emb\}\(x\)\+\\alpha v\_\{\\text\{feature\}\}\.Then,

feature⁡\(emb​\(x\)′\)=\(emb⁡\(x\)\+α​vfeature\)⊤​vfeature\\mathrm\{feature\}\(\\mathrm\{emb\}\(x\)^\{\\prime\}\)=\(\\mathrm\{emb\}\(x\)\+\\alpha v\_\{\\text\{feature\}\}\)^\{\\top\}v\_\{\\text\{feature\}\}=emb​\(x\)⊤​vfeature\+α​vfeature⊤​vfeature=\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\}\+\\alpha v\_\{\\text\{feature\}\}^\{\\top\}v\_\{\\text\{feature\}\}=feature⁡\(emb⁡\(x\)\)\+α​‖vfeature‖2\.=\\mathrm\{feature\}\(\\mathrm\{emb\}\(x\)\)\+\\alpha\\\|v\_\{\\text\{feature\}\}\\\|^\{2\}\.Thus, an intervention defined as above leads to a linear change in the feature value measured by the measurement definition\.

### A\.3The Fourth Definition

Linear classifiers and linear separability are foundational concepts in machine learning that long predate the recent LRH literature\([Minsky and Papert, 1968](https://arxiv.org/html/2609.22695#bib.bib20)\)\. In the context of learned representations, linear separability asks whether examples differing in a target feature can be separated by a linear decision boundary in representation space\. This notion is particularly relevant to the LRH because linear probing is widely used to test empirically whether a feature is linearly represented\.

[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)’s formalization, however, does not cover the standard supervised linear\-probing setting, in which the separating directionwwand thresholdbbare learned from labeled examples\. This gap motivates considering linear separability alongside the three notions above\.

Linear separability:

∃w,bs\.t\.:w⊤emb\("queen"\)\+b\\displaystyle\\exists\\,w,b\\quad\\text\{s\.t\.:\}\\quad w^\{\\top\}\\mathrm\{emb\}\(\\text\{"queen"\}\)\+b\>0\\displaystyle\>0w⊤​emb​\("king"\)\+b\\displaystyle w^\{\\top\}\\mathrm\{emb\}\(\\text\{"king"\}\)\+b≤0\.\\displaystyle\\leq 0\.
Given the measurement definition

feature⁡\(emb⁡\(x\)\)=emb​\(x\)⊤​vfeature,\\mathrm\{feature\}\(\\mathrm\{emb\}\(x\)\)=\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\},introducing a thresholdbbyields the linear decision rule

emb​\(x\)⊤​vfeature\+b\>0\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\}\+b\>0oremb​\(x\)⊤​vfeature\+b≤0\.\\quad\\text\{or\}\\quad\\mathrm\{emb\}\(x\)^\{\\top\}v\_\{\\text\{feature\}\}\+b\\leq 0\.
Thus, adding a bias termbbto[Park et al\. \(2024c\)](https://arxiv.org/html/2609.22695#bib.bib12)’s measurement score gives a linear decision rule, with the classifier weightw=vfeaturew=v\_\{\\text\{feature\}\}\. This establishes a direct connection to linear separability, but only for the special case in which the separating direction isvfeaturev\_\{\\text\{feature\}\}\. Standard linear probing is more general: given labeled examples, it learns the separating directionwwand thresholdbbrather than fixingw=vfeaturew=v\_\{\\text\{feature\}\}\. To capture this probing\-based setting, the proposed Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)can take the same form as a linear probe classifier,y^=σ⁡\(w⊤​emb​\(x\)\+b\)\\hat\{y\}=\\sigma\(w^\{\\top\}\\text\{emb\}\(x\)\+b\), whereσ⁡\(⋅\)\\sigma\(\\cdot\)denotes the sigmoid function\. When the learned probe directionwwaligns withvfeaturev\_\{\\text\{feature\}\}, the probe measures the same feature and applies a threshold to determine its presence\.

This connection yields a simple corollary: linear probes can be used to test whether a feature is linearly represented\. To encompass the diverse probing methodologies in existing literature, Definition 4 explicitly introduces the\(M,l,f,D\)\(M,l,f,D\)tuple\. For instance, the exact definition of the target featureffcan determine whether it appears linearly represented, as shown in recent studies where representations emerge only under specific feature definitions[Nanda et al\. \(2023\)](https://arxiv.org/html/2609.22695#bib.bib108)\. Furthermore, modern probes frequently aggregate activations across multiple token positions rather than relying on a single token[Kramár et al\. \(2026\)](https://arxiv.org/html/2609.22695#bib.bib22)\. Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)accounts for these variations by requiring the probing locationllto be explicitly specified\. By formalizing these elements, Definition[4](https://arxiv.org/html/2609.22695#Thmdefinition4)establishes a mapping between empirical linear probes and the theoretical framework of the linear representation hypothesis\.

## Appendix BProbing Literature Related to the LRH

This paper mainly focuses on work that explicitly discusses “the linear representation hypothesis\.” However, an earlier probing literature studied closely related questions without using the term\. Therefore, this section provides context for interpreting linear decodability as evidence about linear representations\.

#### Linear probes as diagnostic tools\.

An early influential example is[Alain and Bengio \(2017\)](https://arxiv.org/html/2609.22695#bib.bib5), who use linear classifiers, which they call “probes,” to analyze intermediate neural representations\. By training these probes independently of the underlying model, they measure how well task\-relevant information can be linearly decoded at different layers\. This established linear probing as a useful diagnostic tool for studying how representations change across a network\.

[Hewitt and Liang \(2019\)](https://arxiv.org/html/2609.22695#bib.bib6)subsequently questioned how such probe performance should be interpreted\. A high\-capacity probe may learn the probing task itself rather than simply reveal structure already present in the representation\. They therefore introduce*control tasks*and*selectivity*to distinguish predictive performance attributable to the representation from that attributable to probe capacity\.

A complementary use of linearity appears in[Hewitt and Manning \(2019\)](https://arxiv.org/html/2609.22695#bib.bib78)\. Their structural probe learns a linear transformation under which distances and norms in the transformed representation recover dependency\-tree distances and depths\. Here, linearity serves as a restriction on the probe, limiting its ability to learn the parsing task independently of the representation\. This extends probing beyond categorical prediction to the geometric structure of representations\.

#### Information, accessibility, and complexity\.

[Pimentel et al\. \(2020b\)](https://arxiv.org/html/2609.22695#bib.bib7)emphasize that the presence of information in a representation is distinct from its linear decodability\. From an information\-theoretic perspective, a more expressive probe may provide a better estimate of whether information is present at all\. A linear probe therefore addresses the more specific question of whether that information is*linearly decodable*\.

This distinction naturally raises the question of probe complexity\.[Pimentel et al\. \(2020a\)](https://arxiv.org/html/2609.22695#bib.bib8)argue that probe accuracy should be interpreted together with probe complexity: the same predictive performance can provide different evidence depending on how complex a decoder is required to achieve it\. Similarly,[Voita and Titov \(2020\)](https://arxiv.org/html/2609.22695#bib.bib9)use minimum description length \(MDL\) to capture not only final probe performance but also how much effort is required to achieve it, such as the amount of training data needed to reach good predictive performance\. Together, these approaches emphasize that decodability is not simply binary: the same information may be accessible with substantially different amounts of data or decoder complexity\.

#### Beyond linear accessibility\.

Finally,[White et al\. \(2021\)](https://arxiv.org/html/2609.22695#bib.bib10)show that restricting analysis to linear probes can miss structure that is more naturally recovered non\-linearly\. Extending the structural\-probing framework with a non\-linear probe based on an RBF \(radial basis function\) kernel, they recover syntactic structure more accurately than comparable linear probes across multiple languages\. Thus, failure of a linear probe does not necessarily indicate that the target structure is absent; it may simply mean that the structure cannot be recovered well by a linear probe\.

Taken together, this literature progressively refines what can be concluded from probing: one must distinguish whether information is present, whether it is linearly decodable, and how much decoder complexity is required to extract it\. These considerations provide background for interpreting linear probing as evidence for representation formation under the LRH\.

相似文章

线性表示假说需要群作用

arXiv cs.LG

本文通过使用群作用来定义表示等价性,将线性表示假说形式化为一系列声明,澄清了不同分析中的假设。

表示对齐基于线性结构

arXiv cs.LG

本文研究了Platonic Representation Hypothesis,提出对齐源于表示中的线性结构,并引入了一个包含信号、偏置和噪声的统计框架。

真理不是方向:塔斯基对LLM探针的批判

Hacker News Top

本文提出了一个受塔斯基启发的对角线论证,表明在LLM的嵌入空间上,没有线性探针能够可靠地检测真理,并与哥德尔不完备定理和图灵停机问题进行了类比。该文批判了语言模型中关于真理的线性表示假说。

人们到底想从AI得到什么?映射偏好多元性

arXiv cs.CL

本文分析了来自75个国家的1500份开放式回答,揭示了人们对AI的偏好多样且常常相互冲突,其中真实是唯一被广泛需求的价值(49%),但定义方式却互不兼容。研究认为,当前的RLHF方法将这些多元偏好扁平化为通用奖励模型,延续了认知暴力。