Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Summary
This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.
View Cached Full Text
Cached at: 07/24/26, 05:12 AM
# Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Source: [https://arxiv.org/html/2607.20596](https://arxiv.org/html/2607.20596)
Seonglae Cho1,2Zekun Wu1,2Kleyton Da Costa1 Rishi Kalra1Ilham Wicaksono1Adriano Koshiyama1,2 1Holistic AI2University College London
###### Abstract
Sparse autoencoder \(SAE\) features are used to interpret and steer large language models, yet whether a feature’s causal role is stable across SAE families remains untested\. Single\-token features that activate on one vocabulary item provide the diagnostic case where ground truth permits direct comparison\. We analyze 3\.9M features across six models and three SAE families using zero\-ablation at full layer depth\. Single\-token features cluster 4\.7×\\timestighter in decoder space and concentrate in early layers \(Layer 0 in GPT2\-Small; L0–L4 in Gemma\)\. Ablating them yields Benjamini\-Hochberg\-significant logit reductions in 178 of 208 full\-layer conditions, with depth controlling whether damage cascades downstream or shapes the output directly\. Cross\-family causal differences exceed within\-family scale effects: on the same base model, GemmaScope and BatchTopK features remain*causally anchored*, while LlamaScope features are*locally redundant*\. The target token’s rank recovers to within 2×\\timesbaseline 96–98% of the time after the same ablation, and a controlled activation\-function comparison reverses sign within the same model, leaving training recipe as the residual candidate\. Cross\-family interpretability claims are therefore sensitive to training methodology, not just activation function or scale\.
Are Single\-Token Sparse Autoencoder Features Causally Necessary? Layer\-Depth and SAE\-Family Effects
## 1Introduction
Figure 1:Single\-token prevalence and SAE family effects\.\(A\) Principal component analysis \(PCA\) scatter of decoder vectors \(GPT2\-Small L0\): green = single\-token, gray = polysemantic\. \(B\) Prevalence vs model scale: GemmaScope/res\-jb \(blue\) declines 1\.48%→\\to0\.14%; LlamaScope \(red\) near\-zero\. No fit line \(n=3n=3\)\. \(C\) GemmaScope shows 46×\\timeshigher prevalence than LlamaScope at the 8–9B scale \(Gemma\-2\-9B vs Llama\-3\.1\-8B\)\.On the same base model, a feature whose ablation destroys the target token under one sparse autoencoder \(SAE\) family can be replaced almost immediately under another, leaving cross\-family interpretability claims contingent on which SAE was applied\. We studysingle\-token features, those activating on one vocabulary item and its morphological variants, as the SAE analog of grandmother cells\(Gross,[2002](https://arxiv.org/html/2607.20596#bib.bib22)\): the diagnostic endpoint of the monosemantic\-polysemantic spectrum where vocabulary\-level ground truth allows direct cross\-family comparison\.
Three reasons motivate this focus: \(i\) single\-token features are the*diagnostic case*with vocabulary\-level ground truth, enabling direct validation; \(ii\) they form a measurable*bridge*between embedding space and feature space, at 1\.72×\\timeshigher embedding alignment than polysemantic features andp<10−42p<10^\{\-42\}; \(iii\) they have*practical relevance*for token\-level reliability tasks where steering and editing must operate on token\-identity directions\.
We analyze 3\.9M SAE features from Neuronpedia\(Lin and Bloom,[2023](https://arxiv.org/html/2607.20596#bib.bib33)\)across six models: GPT2\-Small \(124M\)\(Radford et al\.,[2019](https://arxiv.org/html/2607.20596#bib.bib44); Bloom,[2024](https://arxiv.org/html/2607.20596#bib.bib5)\), Gemma\-2\-2B and Gemma\-2\-9B\(Team et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib48); Lieberum et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib32)\), Gemma\-3\-1B\(Gemma Team,[2025](https://arxiv.org/html/2607.20596#bib.bib19)\), Llama\-3\.1\-8B\(Grattafiori et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib21)\), and DeepSeek\-R1\-8B\(Guo et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib23)\)\. These span three SAE families: GemmaScope/res\-jb, LlamaScope, and community BatchTopK\. The set spans nearly two orders of magnitude in scale\.
Our analysis yields two main findings\. First, single\-token features are geometrically and causally distinct\. They concentrate in Layer 0 with 91% in GPT2\-Small, cluster 4\.7×\\timestighter in decoder space, and exhibit a sharp L0→\\rightarrowL1 transition where the Grassmannian alignment is 0\.26 versus\>\>0\.9 for later layers\. These features persist 2\.7×\\timesmore strongly across layers than polysemantic features atp<10−73p<10^\{\-73\}\. Causal ablation across the seven full\-depth model×\\timesSAE configurations confirms necessity under zero\-ablation at the measured readouts: 178 of their 208 layers are BH\-significant\. Second, cross\-SAE\-family differences exceed within\-family scale effects\. LlamaScope SAEs show 46×\\timeslower prevalence than GemmaScope at comparable 8–9B scale, and token\-matched comparisons on the same base model extend this gap to causal structure: GemmaScope and BatchTopK features are causally anchored while LlamaScope features are locally redundant, recovering their pre\-ablation rank 96–98% of the time, though activation function alone does not explain this \(Section[5](https://arxiv.org/html/2607.20596#S5)\)\. We additionally observe that causal importance varies by layer depth and semantic category\. Single\-token features are the endpoint where ground truth is unambiguous and cross\-family matching is exact; if causal roles diverge at this simplest matched case, comparisons of more complex features, where establishing correspondence is itself contested, plausibly inherit at least this much instability, a scope argument rather than a demonstrated bound\. The takeaway: a feature’s causal necessity is real but cannot be assumed portable across SAE families; treat the SAE family as an experimental variable to be checked, not a detail to be abstracted away\.
## 2Related Work
#### SAE interpretability foundations\.
Sparse autoencoders decompose neural activations into interpretable features under the superposition hypothesis\(Elhage et al\.,[2022](https://arxiv.org/html/2607.20596#bib.bib14); Bricken et al\.,[2023](https://arxiv.org/html/2607.20596#bib.bib8); Cunningham et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib13); Shu et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib47)\)\. Monosemantic features at scale range from named entities to syntax patterns and abstract concepts\(Templeton et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib49); Gao et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib17)\)\. SAE features feed downstream steering, circuit analysis, and model editing pipelines\(Marks et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib38); Chalnev et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib10); Arad et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib2)\), making cross\-family stability practically consequential\.
#### Cross\-method evaluation and feature universality\.
The features SAEs recover depend on architectural choices and the metrics used to evaluate them\(Leask et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib29); Locatello et al\.,[2019](https://arxiv.org/html/2607.20596#bib.bib35); Karvonen et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib26); Korznikov et al\.,[2026](https://arxiv.org/html/2607.20596#bib.bib27)\)\. Recent work tracks feature evolution across layers\(Balcells et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib4); Balagansky et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib3); Olah et al\.,[2020](https://arxiv.org/html/2607.20596#bib.bib40); Elhage et al\.,[2021](https://arxiv.org/html/2607.20596#bib.bib15)\)and tests cross\-model universality\(Lan et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib28); Paulo and Belrose,[2025](https://arxiv.org/html/2607.20596#bib.bib43); Chanin et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib12)\)\.
## 3Methods
Figure 2:Three\-way validation\.\(A\) Detection schematic: activation metrics and decoder geometry\. \(B\) Gap ratio vs lexical purity separates single\-token \(green\) from polysemantic \(gray\)\. \(C\) Decoder\-geometry clustering on GPT2\-Small Layer 0: single\-token decoder vectors are 7\.5×\\timestighter in mean pairwise cosine than a strict polysemantic set \(gap<0\.2<0\.2, purity<0\.4<0\.4\) read at top\-10\. At the canonical operating point used throughout the paper \(top\-20, full polysemantic complement\) the ratio is 4\.7×\\times\(Table[2](https://arxiv.org/html/2607.20596#S3.T2)\)\.### 3\.1Data and Models
We analyze pre\-trained SAEs hosted on Neuronpedia\(Lin and Bloom,[2023](https://arxiv.org/html/2607.20596#bib.bib33)\)across six models spanning three SAE families \(Table[1](https://arxiv.org/html/2607.20596#S3.T1)\)\. All SAE checkpoints, activation data, and feature explanations are publicly available\. For each feature, we use: \(i\) decoder vectors𝒘dec∈ℝd\{\\bm\{w\}\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{d\}from SAE checkpoints\(Chanin and Bloom,[2024](https://arxiv.org/html/2607.20596#bib.bib11)\), \(ii\) top\-kkactivating tokens and values, and \(iii\) auto\-interpretability explanations\. GPT2\-Small uses res\-jb SAEs\(Bloom,[2024](https://arxiv.org/html/2607.20596#bib.bib5)\)\. Gemma\-2 and Gemma\-3 models use GemmaScope JumpReLU SAEs\(Lieberum et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib32); Rajamanoharan et al\.,[2024b](https://arxiv.org/html/2607.20596#bib.bib46),[a](https://arxiv.org/html/2607.20596#bib.bib45)\)\. Llama\-3\.1 and DeepSeek\-R1 use LlamaScope TopK SAEs\(He et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib24); Gao et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib17)\), following thekk\-sparse autoencoder framework\(Makhzani and Frey,[2014](https://arxiv.org/html/2607.20596#bib.bib37)\)\. For causal experiments, we additionally use community BatchTopK SAEs\(Bussmann et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib9); Chanin and Bloom,[2024](https://arxiv.org/html/2607.20596#bib.bib11)\)on Gemma\-2\-2B and Gemma\-3\-1B to compare competitive versus independent activation functions on the same base models\.
Table 1:Models and SAE configurations analyzed\. SAE Family reflects training methodology\.
### 3\.2Single\-Token Detection
We define asingle\-token featureas one whose activation is dominated by a single vocabulary item and its morphological variants\. We operationalize this concept using three continuous metrics computed from the top\-k=20k=20activating tokens:
Gap Ratiomeasures how sharply the top token dominates:
gap=\(v1−v2\)/v1,\\text\{gap\}=\(v\_\{1\}\-v\_\{2\}\)/v\_\{1\},\(1\)wherev1≥v2v\_\{1\}\\geq v\_\{2\}are the top two activation values among the top\-kkactivating tokens\.
Lexical Puritymeasures the fraction of the top\-kkslots occupied by the most frequent token and its case variants:
purity=maxt\|\{i∈\[k\]:ti=t\}\|/k,\\text\{purity\}=\\max\_\{t\}\|\\\{i\\in\[k\]:t\_\{i\}=t\\\}\|/k,\(2\)after case normalization, wheretit\_\{i\}is the surface string of theii\-th top activating token and the maximum is taken over all unique tokensttappearing in the top\-kkset\.
Complete Wordrequires a word\-boundary prefix \(space for byte pair encoding \(BPE\), U\+2581 for SentencePiece\), excluding subword fragments\.
We select an operating point of gap≥0\.3\\geq 0\.3, purity≥0\.6\\geq 0\.6, and complete word \(bolded row in Table[2](https://arxiv.org/html/2607.20596#S3.T2)\), validated by three independent signals: 64% explanation match, 4\.7×\\timesgeometric clustering, and causal necessity under ablation\. This operating point is conservative: under a Dirichlet null \(α=1\.0\\alpha\\\!=\\\!1\.0\) overk=20k\\\!=\\\!20tokens, the 99th\-percentile gap ratio is 0\.77, placing our threshold well below the null tail; the conjunction of three criteria constrains detection to 1\.48% of features\. All qualitative findings are robust to threshold choice across a 2×\\timesrange \(Table[2](https://arxiv.org/html/2607.20596#S3.T2)\), and our causal experiments \(Section[4\.5](https://arxiv.org/html/2607.20596#S4.SS5)\) use an independent, percentile\-based decoder\-alignment method\. Prevalence is the single\-token count divided by features per layer; for GPT2\-Small this gives 364/24,576 = 1\.48% overall, 91% of which are in Layer 0\.
Table 2:Detection thresholds and validation \(GPT2\-Small, 24,576 features/layer×\\times12 layers\)\.†\\daggerExpl Match: fraction of features whose Neuronpedia auto\-interpretability explanation contains the literal top\-activating token\.
### 3\.3Decoder\-Alignment Detection
Activation\-based detection finds near\-zero LlamaScope features because the TopK→\\toJumpReLU conversion alters the activation distribution\. For causal experiments requiring cross\-family detection, we usedecoder\-alignment detection: cosine similarity between each decoder vector𝒘idec\{\\bm\{w\}\}\_\{i\}^\{\\text\{dec\}\}and the model’s token embedding matrix𝐄\\mathbf\{E\}, applying the same gap and purity thresholds\. This method requires no activation data\. The two methods are consistent where both apply: mean cosine 0\.67 for single\-token features and 89% logit\-lens top\-token match in Section[4](https://arxiv.org/html/2607.20596#S4)\. As a direct LlamaScope check, all 2,872 decoder\-aligned single\-token features across 32 layers activate on their putative token at least once in 4,096 evaluation sequences, with 88\.3% showing negativeΔ\\Deltalogit on ablation and 69\.7% exceeding\|Δlogit\|\>0\.1\|\\Delta\\text\{logit\}\|\>0\.1\.
### 3\.4Geometric Analysis
We characterize feature geometry using three metrics on decoder vectors\{𝒘i\}i∈ℱ\\\{\{\\bm\{w\}\}\_\{i\}\\\}\_\{i\\in\\mathcal\{F\}\}:
Within\-Type Similarity\.For a feature subsetℱ\\mathcal\{F\}with decoder vectors\{𝒘i\}i∈ℱ\\\{\{\\bm\{w\}\}\_\{i\}\\\}\_\{i\\in\\mathcal\{F\}\},
simwithin\(ℱ\)=2\|ℱ\|\(\|ℱ\|−1\)∑i<j∈ℱcos\(𝒘i,𝒘j\)\.\\text\{sim\}\_\{\\text\{within\}\}\(\\mathcal\{F\}\)=\\frac\{2\}\{\|\\mathcal\{F\}\|\(\|\\mathcal\{F\}\|\-1\)\}\\sum\_\{i<j\\in\\mathcal\{F\}\}\\cos\(\{\\bm\{w\}\}\_\{i\},\{\\bm\{w\}\}\_\{j\}\)\.\(3\)Theclustering ratioissimwithin\(ℱST\)/simwithin\(ℱpoly\)\\text\{sim\}\_\{\\text\{within\}\}\(\\mathcal\{F\}\_\{\\text\{ST\}\}\)/\\text\{sim\}\_\{\\text\{within\}\}\(\\mathcal\{F\}\_\{\\text\{poly\}\}\), whereℱpoly\\mathcal\{F\}\_\{\\text\{poly\}\}is the complement ofℱST\\mathcal\{F\}\_\{\\text\{ST\}\}among activation\-detected features at the same layer\. A ratio\>1\>1indicates tighter ST clustering\.
Intrinsic Dimension\.We use the pooled\-average variant of the Levina\-Bickel maximum likelihood estimation \(MLE\) estimator\(Levina and Bickel,[2004](https://arxiv.org/html/2607.20596#bib.bib30)\):
d^=\[1n\(kID−1\)∑i=1n∑j=1kID−1logrkID\(𝒘i\)rj\(𝒘i\)\]−1,\\hat\{d\}=\\left\[\\frac\{1\}\{n\(k\_\{\\text\{ID\}\}\-1\)\}\\sum\_\{i=1\}^\{n\}\\sum\_\{j=1\}^\{k\_\{\\text\{ID\}\}\-1\}\\log\\frac\{r\_\{k\_\{\\text\{ID\}\}\}\(\{\\bm\{w\}\}\_\{i\}\)\}\{r\_\{j\}\(\{\\bm\{w\}\}\_\{i\}\)\}\\right\]^\{\-1\},\(4\)withn=\|ℱ\|n=\|\\mathcal\{F\}\|,kID=10k\_\{\\text\{ID\}\}=10nearest neighbors distinct from the detection top\-k=20k=20,jjindexing nearest\-neighbor ranks11throughkID−1k\_\{\\text\{ID\}\}\-1, andrj\(𝒘i\)r\_\{j\}\(\{\\bm\{w\}\}\_\{i\}\)the Euclidean distance from𝒘i\{\\bm\{w\}\}\_\{i\}to itsjj\-th nearest neighbor inℱ\\mathcal\{F\}\.
Cross\-Layer Grassmannian Alignment\.LetUℓ,Uℓ\+1∈ℝd×dPCAU\_\{\\ell\},U\_\{\\ell\+1\}\\in\\mathbb\{R\}^\{d\\times d\_\{\\text\{PCA\}\}\}have orthonormal columns spanning the top\-dPCA=50d\_\{\\text\{PCA\}\}\{=\}50principal subspaces of the decoder matrices at adjacent layers\(Balagansky et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib3); Lindsey et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib34)\)\. The alignment is
align\(ℓ,ℓ\+1\)=1dPCA∑i=1dPCAcos\(θi\),\\text\{align\}\(\\ell,\\ell\+1\)=\\frac\{1\}\{d\_\{\\text\{PCA\}\}\}\\sum\_\{i=1\}^\{d\_\{\\text\{PCA\}\}\}\\cos\(\\theta\_\{i\}\),\(5\)with\{θi\}i=1dPCA\\\{\\theta\_\{i\}\\\}\_\{i=1\}^\{d\_\{\\text\{PCA\}\}\}the principal angles from the singular value decomposition \(SVD\) ofUℓ⊤Uℓ\+1U\_\{\\ell\}^\{\\top\}U\_\{\\ell\+1\}, equivalently the singular valuesσi\(Uℓ⊤Uℓ\+1\)=cos\(θi\)\\sigma\_\{i\}\(U\_\{\\ell\}^\{\\top\}U\_\{\\ell\+1\}\)=\\cos\(\\theta\_\{i\}\)\. Values approach 1 for identical subspaces and 0 for orthogonal subspaces under a representational shift\.
## 4Results
Table[3](https://arxiv.org/html/2607.20596#S4.T3)maps each major empirical claim to its model set, detector, layer coverage, and evidence type\.
Table 3:Claim→\\toevidence mapping\. Detector: A = activation\-based; D = decoder\-alignment\. Evidence: G = Geometric; Ds = Descriptive; C = Causal\.### 4\.1Scale\-Dependent Prevalence
Single\-token feature prevalence \(the fraction of a layer’s features meeting all three detection criteria; Section[3](https://arxiv.org/html/2607.20596#S3)\) decreases consistently with model scale within the GemmaScope/res\-jb SAE family \(Figure[1](https://arxiv.org/html/2607.20596#S1.F1)\)\. GPT2\-Small \(124M\) yields 1\.48% single\-token features \(364/24,576, with 91% in Layer 0\), declining to 0\.47% in Gemma\-2\-2B \(2B\) and 0\.14% in Gemma\-2\-9B \(9B\)\. The decline is steeper at Layer 0 \(1\.35%→\\rightarrow0\.50%→\\rightarrow0\.01%\), where token identity is primarily encoded\. The pattern holds only for GemmaScope/res\-jb SAEs; both analyzed LlamaScope SAEs exhibit near\-zero prevalence at the 8B scale\. The prevalence gap co\-varies with SAE family and cannot be attributed to methodology alone, since base model, tokenizer, and training data vary simultaneously\.
Wider SAEs allocate more capacity to monosemantic features \(prevalence increases sublinearly with width; Appendix[A\.2](https://arxiv.org/html/2607.20596#A1.SS2)\)\. We report the three\-point trend qualitatively; three data points cannot distinguish functional forms\. A descriptive fit with explicit “n=3n\\\!=\\\!3, descriptive only” caveat is in Appendix[A](https://arxiv.org/html/2607.20596#A1)\. The cleanest scale comparison is within\-family: Gemma\-2\-2B to Gemma\-2\-9B, with the same tokenizer and family, declines from 0\.47% to 0\.14%, while GPT2\-Small extends the range without carrying inferential weight\. The Gemma\-2\-9B/Llama\-3\.1\-8B contrast \(Figure[1](https://arxiv.org/html/2607.20596#S1.F1)C\) is a motivating cross\-model observation; the within\-model comparisons in Table[8](https://arxiv.org/html/2607.20596#S4.T8)are the evidentiary basis for the causal methodology claim\.
### 4\.2Validation
We characterize and validate detection through three signals \(Figure[2](https://arxiv.org/html/2607.20596#S3.F2)\)\.Activation profile:Single\-token features have higher max activations and peaked distributions \(kurtosis 2\.65 vs−\-0\.08,p<10−154p<10^\{\-154\}\), with 92% mass on the top token vs 34% for polysemantic; peakedness is expected given the gap\-ratio criterion and confirms a clean threshold rather than independent evidence\.Geometry, independent of activation detector: the detection criteria use only activation statistics; geometric metrics use only decoder vectors\. Single\-token vectors cluster 4\.7×\\timestighter \(mean cosine 0\.103 inℱST\\mathcal\{F\}\_\{\\text\{ST\}\}vs 0\.022 inℱpoly\\mathcal\{F\}\_\{\\text\{poly\}\}\) with 2×\\timeslower intrinsic dimension \(60\.4 vs 118\.4\); ratios stable across models \(4\.5–5\.0×\\timessimilarity, 1\.7–2×\\timesdimension\)\.Explanation, independent of activation and geometry:*explanation match*is the fraction of features whose Neuronpedia auto\-interpretability explanation contains the literal case\-normalized surface form of the top\-activating token, 64% for single\-token vs 8% for polysemantic, widening at stricter thresholds per Table[2](https://arxiv.org/html/2607.20596#S3.T2)\.
We test whether single\-token decoder vectors align with token embeddings, restricting toactivation\-detectedfeatures to avoid circularity with decoder\-alignment detection\. We define the*mean cosine alignment*of a feature setℱ\\mathcal\{F\}as the average ofcos\(𝒘idec,𝐄ti∗\)\\cos\(\{\\bm\{w\}\}\_\{i\}^\{\\text\{dec\}\},\\mathbf\{E\}\_\{t\_\{i\}^\{\*\}\}\)overi∈ℱi\\in\\mathcal\{F\}, where𝐄ti∗\\mathbf\{E\}\_\{t\_\{i\}^\{\*\}\}is the input\-embedding row for featureii’s top activating tokenti∗t\_\{i\}^\{\*\}\. Single\-token features show higher mean cosine alignment than polysemantic \(0\.67 vs 0\.39,p<10−42p<10^\{\-42\}\), decreasing with layer depth as expected for increasingly abstract representations\. Using the logit lens\(nostalgebraist,[2020](https://arxiv.org/html/2607.20596#bib.bib39); Bloom and Lin,[2024](https://arxiv.org/html/2607.20596#bib.bib6)\), the*logit\-lens match*indicator is 1 ifargmaxt\(𝒘idec⋅𝐔t\)=ti∗\\arg\\max\_\{t\}\(\{\\bm\{w\}\}\_\{i\}^\{\\text\{dec\}\}\\cdot\\mathbf\{U\}\_\{t\}\)=t\_\{i\}^\{\*\}, 0 otherwise; 89% match for single\-token vs 12% for polysemantic\. This bidirectional alignment, from embedding space and toward unembedding, validates that single\-token features function as lexical identity detectors\.
### 4\.3Layer\-0 Concentration and Representational Shift
In GPT2\-Small, the majority of single\-token features concentrate in Layer 0 \(Figure[3](https://arxiv.org/html/2607.20596#S4.F3); per\-layer breakdown in Table[13](https://arxiv.org/html/2607.20596#A1.T13), Appendix\)\. This concentration precedes a sharp representational shift: cross\-layer feature tracking shows 86\.8% of L0 features have no close L1 match at the 0\.3 cosine threshold, 3×\\timesthe mean random\-pair similarity, while L1→\\rightarrowL2 shows only 16\.0% transformation and later transitions maintain 60–70% feature persistence\. Gemma\-2\-2B confirms this pattern: L0→\\rightarrowL1 transformation is 82\.8%, then L1→\\rightarrowL2 drops to 51\.9%, with later layers showing<<5% transformation\. This aligns with Layer 0 encoding token identity before attention enables cross\-position flow\. Grassmannian alignment \(Table[4](https://arxiv.org/html/2607.20596#S4.T4)\) quantifies the shift: GPT2 L0→\\toL1 alignment is 0\.26 while later transitions stabilize at\>\>0\.9\. Gemma L0→\\toL1 alignment is similarly low at 0\.23 under its distributed single\-token pattern\. Intrinsic dimension analysis \(full table in Appendix[A\.4](https://arxiv.org/html/2607.20596#A1.SS4)\): single\-token features in GPT2 L0 occupy a lower\-dimensional manifold than polysemantic features\. Despite the L0→\\toL1 shift, single\-token features show 2\.7×\\timeshigher cross\-layer persistence than polysemantic features at Z\-score 10\.7 vs 4\.0,p<10−73p<10^\{\-73\}\.
Table 4:Grassmannian alignment between adjacent layers\.High\-persistence pairs atZ\>50Z\>50preserve semantic meaning, 0\.68 vs 0\.16 in explanation similarity, replicating in Gemma\-2\-2B at 0\.61 vs 0\.32\. Single\-token features also show 1\.42×\\timessmaller direction change at L1→\\rightarrowL2 and near\-zero co\-occurrence at Jaccard<<0\.001, partitioning vocabulary into non\-overlapping regions\.
Figure 3:Cross\-layer dynamics\.\(A\) Cross\-layer similarity matrix \(GPT2\): L0→\\toL1 shows sharp transition \(highlighted\), then stabilizes\. \(B\) Feature preservation across layers: both models show low L0→\\toL1 similarity, then recover\. \(C\) Persistence Z\-scores: single\-token features show 2\.7×\\timeshigher cross\-layer persistence \(10\.7 vs 4\.0\)\.
### 4\.4Token Characterization
Single\-token features are enriched 2\.5×\\timesfor mid\-frequency tokens at rank 1k–10k, 66\.5% vs 26\.7%\. Content nouns at 52%, proper nouns at 8\.4%, and function words at 4\.6% dominate \(Table[5](https://arxiv.org/html/2607.20596#S4.T5)\)\. Ablation damage varies across categories under Kruskal\-Wallisp=4×10−72p=4\\times 10^\{\-72\}\. The “Goldilocks zone” pattern is tested against three chi\-squared nulls withdf=5df=5\. The uniform null givesχ2=206\.4\\chi^\{2\}=206\.4,p=1\.2×10−42p=1\.2\\times 10^\{\-42\}\. The vocab\-frequency\-proportional null givesχ2=625\.4\\chi^\{2\}=625\.4,p=6\.6×10−133p=6\.6\\times 10^\{\-133\}\. The polysemantic\-matched null givesχ2=297\.5\\chi^\{2\}=297\.5,p=3\.5×10−62p=3\.5\\times 10^\{\-62\}\. All three are rejected, ruling out simple allocation models\.
Despite different tokenizers, 24% of GPT2 single\-token features have a Gemma counterpart with sentence\-embedding cosine\>\>0\.7 over auto\-interpretability explanations \(Table[17](https://arxiv.org/html/2607.20596#A2.T17)\): partial semantic convergence across models, measured at the explanation\-text level rather than the geometric feature level\.
Table 5:Semantic category distribution\.Δlogit\\Delta\\text\{logit\}: mean logit change on ablation \(more negative = more damage\)\. Damaged: fraction with\|Δlogit\|\>0\.1\|\\Delta\\text\{logit\}\|\>0\.1\.Kruskal\-WallisH=496\.5H=496\.5,p=4×10−72p=4\\times 10^\{\-72\}\.N=26,594N=26\{,\}594across 6 models\. Top 5 categories shown \(69% of features\); remaining 31% span minor categories\.
### 4\.5Causal Validation
To establish functional necessity beyond correlational evidence, we perform zero\-ablation interventions across eight model×\\timesSAE combinations \(Table[6](https://arxiv.org/html/2607.20596#S4.T6)\)\. For a featureiiactivating with valuefi\>0f\_\{i\}\>0at a \(sequence, position\) pair, we replace the residual\-stream activation𝐚\\mathbf\{a\}with𝐚−fi⋅𝐰idec\\mathbf\{a\}\-f\_\{i\}\\cdot\\mathbf\{w\}\_\{i\}^\{\\text\{dec\}\}, run the modified forward pass, and record*ablation damage*Δlogiti=logpablated\(ti∗\)−logpclean\(ti∗\)\\Delta\\text\{logit\}\_\{i\}=\\log p\_\{\\text\{ablated\}\}\(t\_\{i\}^\{\*\}\)\-\\log p\_\{\\text\{clean\}\}\(t\_\{i\}^\{\*\}\)for the top activating tokenti∗t\_\{i\}^\{\*\}, averaged over all positions wherefi\>0f\_\{i\}\>0\. Layer\-level significance uses a one\-sided Mann\-WhitneyUUon signedΔlogit\\Delta\\text\{logit\}\(alternative: single\-token ablation shifts the target\-token logit more negatively than matched controls\), BH\-corrected atp<0\.05p<0\.05globally across all 208 layer tests\. Features are identified via decoder\-alignment detection, Section[3](https://arxiv.org/html/2607.20596#S3)\. Detection consistency with the activation\-based detector is supported by 89% logit\-lens match and 0\.67 mean cosine\. Activation positions come from 4,096 sequences of 128 tokens from OpenWebText\(Gokaslan and Cohen,[2019](https://arxiv.org/html/2607.20596#bib.bib20)\); each model tokenizes its own copy of the raw text\. Each single\-token feature is paired with a*size\-matched random control*drawn uniformly from non\-single\-token features of the same SAE, restricted to controls whose mean active activation falls within a 2×\\timesrange of the ST feature’s\.
Necessity\.Single\-token feature ablation reduces target token logits at BH\-significant levels across all eight conditions \(Table[6](https://arxiv.org/html/2607.20596#S4.T6)\)\. Within the four full\-layer experiments on Gemma models, 77/102 tested layers yield significant logit reduction under Mann\-WhitneyUUglobally BH\-corrected atp<0\.05p<0\.05: 50/51 Gemma\-2\-2B layers and 27/51 Gemma\-3\-1B layers, concentrated in later layers\. Extending to two additional full\-depth configurations, Llama\-3\.1\-8B×\\timesLlamaScope yields 31/32 significant layers with peakp=2\.3×10−141p=2\.3\\times 10^\{\-141\}at L1,nST=2,872n\_\{\\text\{ST\}\}=2\{,\}872decoder\-aligned features across 32 layers, and prevalence 0\.27%\. Gemma\-2\-9B×\\timesGemmaScope yields 42/42 with peakp=6\.1×10−26p=6\.1\\times 10^\{\-26\}at L35\. DeepSeek\-R1×\\timesLlamaScope at full 32\-layer depth yields 28/32 BH\-significant \(peakp=2\.0×10−13p=2\.0\\times 10^\{\-13\}at L2\), bringing total full\-layer causal coverage to 208 layers across 7 model×\\timesSAE configurations\. This effect is not explained by activation magnitude: even in the lowest activation quartile, single\-token features cause more damage than magnitude\-matched random controls, atp<0\.0001p<0\.0001and rank\-biserialr=0\.27r=0\.27–0\.430\.43; see Appendix[B\.12](https://arxiv.org/html/2607.20596#A2.SS12)\.
Anchoring and redundancy\.A feature at source layerℓ\\ell*anchors*downstream layerℓ′\>ℓ\\ell^\{\\prime\}\>\\ellif zero\-ablation atℓ\\ellproduces a BH\-significant change in the logit\-lens top\-token readout atℓ′\\ell^\{\\prime\}, using one\-sided Mann\-WhitneyUUagainst magnitude\-matched controls atp<0\.05p<0\.05\. A source layer*anchors≥1\\geq 1downstream layer*if at least oneℓ′\>ℓ\\ell^\{\\prime\}\>\\ellmeets this criterion under global BH correction\. Anchoring is consistent for GemmaScope and BatchTopK at 92–100% of source layers and sparser for LlamaScope, where Llama\-3\.1\-8B anchors 31% and DeepSeek\-R1 34% of source layers at full 32\-layer depth; the residual anchoring on Llama\-3\.1\-8B concentrates in early layers L0–L9\. Same\-layer*recovery*is the fraction of features whose mean rank forti∗t\_\{i\}^\{\*\}under the same\-layer logit lens stays within twice its pre\-ablation value \(floor 5\)*after*ablation\. High recovery indicates*local redundancy*where other features compensate; low recovery indicates critical reliance on the ablated feature\. Recovery mirrors anchoring: GemmaScope 62–71%, LlamaScope 96–98%; Gemma\-3\-1B GemmaScope at 91% is the exception\. This SAE family split in causal structure parallels the prevalence split: SAE training methodology shapes not only which features are detected but their degree of causal importance\.
Layer depth dissociation\.Necessity and anchoring show opposite layer profiles\. Necessity damage \(\|Δlogit\|\|\\Delta\\text\{logit\}\|\) increases monotonically with depth \(Spearmanρ=0\.97\\rho=0\.97for BatchTopK,0\.700\.70for GemmaScope on Gemma\-2\-2B;p<0\.001p<0\.001\), with late layers showing 13–30×\\timesmore damage than early layers\. In contrast, anchor damage \(downstream propagation\) is concentrated in early layers: early\-layer ablations cause 4–16×\\timesmore total downstream disruption than late\-layer ablations \(Table[7](https://arxiv.org/html/2607.20596#S4.T7)\)\. This dissociation reveals complementary roles: early features serve as propagation anchors whose ablation cascades through the network, while late features directly shape the output distribution\. BatchTopK shows a near\-perfect monotonic necessity gradient \(ρ=0\.97\\rho=0\.97\) while GemmaScope shows a noisier profile \(ρ=0\.70\\rho=0\.70\): SAE families distribute causal load across layers\.
Cross\-architecture validation\.Gemma\-3\-1B reproduces the pattern: anchoring holds at 92–96% of layers, necessity is BH\-significant in 27/51 layers concentrated late, and BatchTopK on the same model recovers 18/25 \(Table[6](https://arxiv.org/html/2607.20596#S4.T6)\)\. The weaker Gemma\-3\-1B GemmaScope replication \(9/26 vs 26/26 on Gemma\-2\-2B\) is consistent with smaller effect sizes rather than a methodology breakdown\. Per\-layer mean\|Δlogit\|\|\\Delta\\text\{logit\}\|on Gemma\-3\-1B GemmaScope is order\-of\-magnitude smaller than Gemma\-2\-2B GS at matchednSTn\_\{\\text\{ST\}\}\(Table[7](https://arxiv.org/html/2607.20596#S4.T7)\)\. On Gemma\-3\-1B BatchTopK, where effect sizes recover, 18/25 layers are BH\-significant\.
Table 6:Causal validation across 8 model×\\timesSAE conditions; the seven full\-depth configurations contribute the 208 layers reported in the main text \(GPT2\-Small is a single\-layer condition\)\. Nec\. sig: layers significant under a single global BH correction across all 208 tests\. Peakpp: strongest per\-layer one\-sided Mann\-Whitney result on the signed statistic\. Recovery: fraction of features whose target\-token rank after ablation stays within twice its pre\-ablation rank \(floor 5\); blue cells mark the anchored regime, rust cells the locally redundant LlamaScope regime\.ModelSAE FamilyLayersnSTn\_\{\\text\{ST\}\}Nec\. sigPeakppAnchor \(≥\\geq1\)Avg anc\.RecoveryGPT2\-Smallres\-jb \(ReLU\)1/122201/12\.5e−802\.5\\mathrm\{e\}\{\-80\}1/1100%26\.8%Gemma\-2\-2BGemmaScope \(JR\)26/2651–15726/261\.4e−341\.4\\mathrm\{e\}\{\-34\}25/2650%70\.6%Gemma\-2\-2BBatchTopK25/2564–32224/252\.9e−𝟔𝟕\\mathbf\{2\.9\\mathrm\{e\}\{\-67\}\}25/2552%68\.0%Gemma\-2\-9BGemmaScope \(JR\)42/4242–15742/426\.1e−𝟐𝟔\\mathbf\{6\.1\\mathrm\{e\}\{\-26\}\}41/4250%62\.1%Gemma\-3\-1BGemmaScope \(JR\)26/2656–1509/267\.8e−87\.8\\mathrm\{e\}\{\-8\}24/2625%90\.8%Gemma\-3\-1BBatchTopK25/2525–23618/251\.0e−291\.0\\mathrm\{e\}\{\-29\}24/2546%68\.4%Llama\-3\.1\-8BLlamaScope \(TopK→\\toJR\)32/321–32831/322\.3e−𝟏𝟒𝟏\\mathbf\{2\.3\\mathrm\{e\}\{\-141\}\}10/3223%97\.7%DeepSeek\-R1LlamaScope \(TopK→\\toJR\)32/323–32128/322\.0e−𝟏𝟑\\mathbf\{2\.0\\mathrm\{e\}\{\-13\}\}11/3214%95\.5%
Table 7:Layer depth dissociation: necessity increases with depth while anchor damage concentrates in early layers\. Q1/Q4: first/last quartile of layers\.Table[8](https://arxiv.org/html/2607.20596#S4.T8)\(full breakdowns in Appendix[B\.8](https://arxiv.org/html/2607.20596#A2.SS8),[B\.9](https://arxiv.org/html/2607.20596#A2.SS9)\) jointly indicates that the activation function alone does not explain the cross\-family pattern: token\-matched \(N=627N\\\!=\\\!627\) shows BatchTopK\>\>GemmaScope \(p=1\.2×10−18p=1\.2\\times 10^\{\-18\},r=0\.36r\\\!=\\\!0\.36\), but the activation\-function\-isolated controlled comparison on the same model/layer/width \(N=142N\\\!=\\\!142\) shows the opposite direction \(JumpReLU\>\>TopK,p=0\.036p\\\!=\\\!0\.036\); opposite signs from the same model leave training\-recipe factors as residual candidates \(discussed in Section[5](https://arxiv.org/html/2607.20596#S5)\)\.
Table 8:Within\-model decompositions of the cross\-family causal gap\. Top: token\-matched paired \(BTK vs GS,N=627N\\\!=\\\!627\)\. Bottom: controlled activation\-function on Gemma\-2\-2B L1 \(N=142N\\\!=\\\!142, width 18k\)\.ComparisonMetricppEffect sizeDirectionToken\-matched paired \(N=627N\\\!=\\\!627, Gemma\-2\-2B \+ Gemma\-3\-1B\):BatchTopK vs GemmaScopeAnchor \(Δ\\Deltalogit\-lens\)1\.2×10−181\.2\\times 10^\{\-18\}r=0\.36r\\\!=\\\!0\.36BTK\>\>GSBatchTopK vs GemmaScopeNecessity \(Δ\\Deltarank\)0\.0040\.004r=0\.12r\\\!=\\\!0\.12BTK\>\>GSControlled activation function \(N=142N\\\!=\\\!142, Gemma\-2\-2B L1, width 18k\):TopK vs JumpReLUNecessity \(meanΔ\\Deltarank\)0\.0360\.036r=0\.20r\\\!=\\\!0\.20JR\>\>TopKTopK vs JumpReLUNecessity \(maxΔ\\Deltarank\)0\.0030\.003r=0\.27r\\\!=\\\!0\.27JR\>\>TopK
Figure 4:Causal structure across SAE families\.\(A\) Recovery curves for three SAE types on Gemma\-2\-2B L1: activation function and training recipe each contribute to the recovery gap\. \(B\) Anchoring \(solid\) and recovery \(hatched\) across 17 sampled model×\\timeslayer conditions; per\-layer breakdown for the full 208\-layer coverage is in Table[6](https://arxiv.org/html/2607.20596#S4.T6)\. GemmaScope \(blue\) and BatchTopK \(green\) show high anchoring with moderate recovery; LlamaScope \(gray\) shows near\-zero anchoring with high recovery\.
### 4\.6Factors and Robustness
Extending to all six models \(Figure[1](https://arxiv.org/html/2607.20596#S1.F1)B\), GemmaScope/res\-jb prevalence is 0\.01–1\.35% in Layer 0 while LlamaScope is near\-zero at<<0\.01%\. The 46×\\timesfamily contrast exceeds the within\-family scale trend, with GemmaScope gap ratios up to 0\.76 vs LlamaScope’s 0\.17\. An eight\-architecture comparison on Gemma\-2\-2B L12 \(Appendix[A\.2](https://arxiv.org/html/2607.20596#A1.SS2); Table[14](https://arxiv.org/html/2607.20596#A1.T14)\) shows prevalence insensitive to activation function at mid\-layers \(0\.92–0\.99%\): single\-token features are primarily pre\-compositional\. Threshold ablations \(Table[2](https://arxiv.org/html/2607.20596#S3.T2)\) confirm the geometric signatures are stable across operating points\. The causal experiments \(Section[4\.5](https://arxiv.org/html/2607.20596#S4.SS5)\) use independent percentile\-based decoder\-alignment detection\.
## 5Discussion
#### Geometric Findings\.
The 4\.7×\\timestighter decoder clustering, 2×\\timeslower intrinsic dimension, and 1\.72×\\timeshigher embedding alignment \(p<10−42p<10^\{\-42\}\) provide activation\-independent validation, supporting the linear representation hypothesis\(Park et al\.,[2023](https://arxiv.org/html/2607.20596#bib.bib42),[2024](https://arxiv.org/html/2607.20596#bib.bib41); Li et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib31); Engels et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib16)\)\. Similarity ratios are stable across models \(4\.5–5\.0×\\timesat L0; mechanism figure in Appendix[B](https://arxiv.org/html/2607.20596#A2)\); cross\-model analysis reveals zero surface\-form overlap but 24% semantic correspondence \(Appendix Table[17](https://arxiv.org/html/2607.20596#A2.T17)\), localizing the tokenizer\-invariant boundary at the single\-token endpoint and complementingLan et al\. \([2024](https://arxiv.org/html/2607.20596#bib.bib28)\)’s universal feature\-space findings\(Templeton et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib49); Gao et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib17); Venhoff et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib50)\)\.
#### Layerwise Dynamics\.
GPT2’s Layer\-0 concentration \(91%\) preceding low L0→\\toL1 alignment \(0\.26\) supports Layer 0 encoding token identity before attention enables cross\-position flow\(Elhage et al\.,[2021](https://arxiv.org/html/2607.20596#bib.bib15); Ameisen et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib1)\); the L0 concentration is model\-specific \(Gemma models show distributed L0–L4 patterns\)\. The depth dissociation \(Section[4\.5](https://arxiv.org/html/2607.20596#S4.SS5)\) shows complementary roles: necessity damage scales with depth \(ρ=0\.97\\rho=0\.97BatchTopK,0\.700\.70GemmaScope at 2B;0\.810\.81at 9B\) while anchoring concentrates in early layers \(ρ=−0\.65\\rho=\-0\.65\)\. This refines prior cross\-layer SAE tracking work\(Balcells et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib4); Balagansky et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib3)\), which characterize persistence or subspace motifs but do not separate necessity from anchor profiles\. Gemma\-3\-1B GemmaScope \(91% recovery\) is an exception, but BatchTopK on the same model recovers at 68\.4% so the SAE\-family ordering still holds at 1B\.
#### SAE Family Comparison\.
The 46×\\timescross\-family prevalence contrast exceeds the roughly 10×\\timeswithin\-family scale trend, and the causal contrast follows the same family split; because the cross\-family comparison varies training data, dictionary width, training recipe, and post\-hoc conversion, we treat within\-model evidence as load\-bearing and the 46×\\timesnumber as a magnitude bound, consistent withPaulo and Belrose \([2025](https://arxiv.org/html/2607.20596#bib.bib43)\); Leask et al\. \([2025](https://arxiv.org/html/2607.20596#bib.bib29)\); Chanin et al\. \([2024](https://arxiv.org/html/2607.20596#bib.bib12)\); Locatello et al\. \([2019](https://arxiv.org/html/2607.20596#bib.bib35)\)\. Across 208 full\-layer tests, the anchoring split persists: GemmaScope/BatchTopK anchor 92–100% of source layers while LlamaScope anchors 31–34% across Llama\-3\.1\-8B and DeepSeek\-R1, with category\-dependent effects at the token level: domain\-specific tokens converge at 93%, function words at 29%; Appendix[B\.11](https://arxiv.org/html/2607.20596#A2.SS11)\. The within\-model picture dissociates per Table[8](https://arxiv.org/html/2607.20596#S4.T8): token\-matched BatchTopK\>\>GemmaScope atN=627N\\\!=\\\!627,p=1\.2×10−18p=1\.2\\times 10^\{\-18\}, but the activation\-function\-isolatedN=142N\\\!=\\\!142comparison shows JumpReLU\>\>TopK atp=0\.036p\\\!=\\\!0\.036; opposite signs from the same model leave training corpus, dictionary width, and recipe as residual candidates\(Geiger et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib18); Hindupur et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib25); Braun et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib7); Makelov et al\.,[2024](https://arxiv.org/html/2607.20596#bib.bib36)\)\.
## 6Conclusion
At the single\-token endpoint, where ground truth is unambiguous, the causal role of a feature depends on which SAE produced it\. Ablating a single\-token feature reduces the target token’s logit across six transformer language models and three SAE families, with depth controlling whether damage cascades downstream or shapes the output directly\. The same token can be causally anchored under one SAE family yet locally redundant under another, so cross\-SAE interpretability claims must control for training methodology, not activation function or scale\.
For practitioners this implies three concrete steps\. Re\-verify causal claims when switching SAE families: a feature’s necessity under one family cannot be assumed to transfer, even on the same base model, so steering and editing pipelines should re\-run ablation checks under the family they deploy\. Profile a checkpoint’s causal behavior directly rather than inferring it from family name or activation function: anchoring and recovery statistics are a safer guide than either label\. Add per\-feature causal necessity to SAE evaluation: two SAEs trained on the same base model can differ in causal structure, a dimension that current benchmarks such as SAEBench\(Karvonen et al\.,[2025](https://arxiv.org/html/2607.20596#bib.bib26)\)do not directly measure\.
## Limitations
Cross\-family comparisons co\-vary training data, dictionary width, recipe, and post\-hoc conversion; within\-model pairings on Gemma\-2\-2B and Gemma\-3\-1B and the activation\-function\-controlled comparison partially constrain the recipe factor\. All four full\-depth configurations \(Llama\-3\.1\-8B 32/32, DeepSeek\-R1 32/32, Gemma\-2\-9B 42/42, Gemma\-2\-2B 26/26\) show consistent results; Gemma\-3\-1B at 26 layers shows weaker effects\. Activation\-based detection uses a fixed operating point; conclusions are robust across a 2×\\timesthreshold range \(Table[2](https://arxiv.org/html/2607.20596#S3.T2)\) and the causal experiments use independent percentile\-based decoder\-alignment detection\. The single\-token endpoint is the diagnostic case; multi\-token spans and compositional features are future work via span\-embedding alignment\.
## References
- Ameisen et al\. \(2025\)Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L\. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, and 8 others\. 2025\.[Circuit tracing: Revealing computational graphs in language models](https://transformer-circuits.pub/2025/attribution-graphs/methods.html)\.*Transformer Circuits Thread*\.
- Arad et al\. \(2025\)Dana Arad, Aaron Mueller, and Yonatan Belinkov\. 2025\.[SAEs are good for steering – if you select the right features](https://doi.org/10.18653/v1/2025.emnlp-main.519)\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 10241–10259, Suzhou, China\. Association for Computational Linguistics\.
- Balagansky et al\. \(2025\)Nikita Balagansky, Ian Maksimov, and Daniil Gavrilov\. 2025\.[Mechanistic permutability: Match features across layers](https://openreview.net/forum?id=MDvecs7EvO)\.In*The Thirteenth International Conference on Learning Representations*\.
- Balcells et al\. \(2024\)Daniel Balcells, Benjamin Lerner, Michael Oesterle, Ediz Ucar, and Stefan Heimersheim\. 2024\.[Evolution of sae features across layers in llms](https://arxiv.org/abs/2410.08869)\.*Preprint*, arXiv:2410\.08869\.
- Bloom \(2024\)Joseph Bloom\. 2024\.Open source sparse autoencoders for all residual stream layers of GPT\-2 small\.[https://www\.alignmentforum\.org/posts/f9EgfLSurAiqRJySD](https://www.alignmentforum.org/posts/f9EgfLSurAiqRJySD)\.
- Bloom and Lin \(2024\)Joseph Bloom and Johnny Lin\. 2024\.Understanding sae features with the logit lens\.[https://www\.lesswrong\.com/posts/qykrYY6rXXM7EEs8Q](https://www.lesswrong.com/posts/qykrYY6rXXM7EEs8Q)\.
- Braun et al\. \(2024\)Dan Braun, Jordan Taylor, Nicholas Goldowsky\-Dill, and Lee Sharkey\. 2024\.[Identifying functionally important features with end\-to\-end sparse dictionary learning](https://openreview.net/forum?id=7txPaUpUnc)\.In*The Thirty\-eighth Annual Conference on Neural Information Processing Systems*\.
- Bricken et al\. \(2023\)Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield\-Dodds, Alex Tamkin, Karina Nguyen, and 6 others\. 2023\.Towards monosemanticity: Decomposing language models with dictionary learning\.*Transformer Circuits Thread*\.Https://transformer\-circuits\.pub/2023/monosemantic\-features/index\.html\.
- Bussmann et al\. \(2024\)Bart Bussmann, Patrick Leask, and Neel Nanda\. 2024\.[BatchTopK sparse autoencoders](https://openreview.net/forum?id=d4dpOCqybL)\.In*NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning*\.
- Chalnev et al\. \(2024\)Sviatoslav Chalnev, Matthew Siu, and Arthur Conmy\. 2024\.[Improving steering vectors by targeting sparse autoencoder features](https://arxiv.org/abs/2411.02193)\.*Preprint*, arXiv:2411\.02193\.
- Chanin and Bloom \(2024\)David Chanin and Joseph Bloom\. 2024\.[Saelens: SAE training and analysis library](https://github.com/jbloomAus/SAELens)\.
- Chanin et al\. \(2024\)David Chanin, James Wilken\-Smith, Tomáš Dulka, Hardik Bhatnagar, Satvik Golechha, and Joseph Bloom\. 2024\.[A is for absorption: Studying feature splitting and absorption in sparse autoencoders](https://arxiv.org/abs/2409.14507)\.*Preprint*, arXiv:2409\.14507\.
- Cunningham et al\. \(2024\)Hoagy Cunningham, Aidan Ewart, Logan Riggs Smith, Robert Huben, and Lee Sharkey\. 2024\.[Sparse autoencoders find highly interpretable features in language models](https://openreview.net/forum?id=F76bwRSLeK)\.In*The Twelfth International Conference on Learning Representations*\.
- Elhage et al\. \(2022\)Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield\-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah\. 2022\.[Toy models of superposition](https://transformer-circuits.pub/2022/toy_model/index.html)\.*Transformer Circuits Thread*\.
- Elhage et al\. \(2021\)Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield\-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, and 6 others\. 2021\.[A mathematical framework for transformer circuits](https://transformer-circuits.pub/2021/framework/index.html)\.*Transformer Circuits Thread*\.
- Engels et al\. \(2024\)Joshua Engels, Eric J\. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark\. 2024\.[Not all language model features are linear](https://arxiv.org/abs/2405.14860)\.*Preprint*, arXiv:2405\.14860\.
- Gao et al\. \(2025\)Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu\. 2025\.[Scaling and evaluating TopK sparse autoencoders](https://openreview.net/forum?id=tcsZt9ZNKD)\.In*The Thirteenth International Conference on Learning Representations*\.
- Geiger et al\. \(2025\)Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard\. 2025\.[Causal abstraction: A theoretical foundation for mechanistic interpretability](https://arxiv.org/abs/2301.04709)\.*Journal of Machine Learning Research*, 26\.
- Gemma Team \(2025\)Gemma Team\. 2025\.[Gemma 3 technical report](https://arxiv.org/abs/2503.19786)\.*Preprint*, arXiv:2503\.19786\.
- Gokaslan and Cohen \(2019\)Aaron Gokaslan and Vanya Cohen\. 2019\.Openwebtext corpus\.[https://skylion007\.github\.io/OpenWebTextCorpus/](https://skylion007.github.io/OpenWebTextCorpus/)\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others\. 2024\.[The llama 3 herd of models](https://arxiv.org/abs/2407.21783)\.*Preprint*, arXiv:2407\.21783\.
- Gross \(2002\)Charles G\. Gross\. 2002\.[Genealogy of the “grandmother cell”](https://doi.org/10.1177/107385802237175)\.*The Neuroscientist*, 8\(5\):512–518\.
- Guo et al\. \(2025\)Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z\. F\. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, and 175 others\. 2025\.[DeepSeek\-R1 incentivizes reasoning in LLMs through reinforcement learning](https://doi.org/10.1038/s41586-025-09422-z)\.*Nature*, 645\(8081\):633–638\.
- He et al\. \(2024\)Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, Yu\-Gang Jiang, and Xipeng Qiu\. 2024\.[Llama scope: Extracting millions of features from llama\-3\.1\-8b with sparse autoencoders](https://arxiv.org/abs/2410.20526)\.*Preprint*, arXiv:2410\.20526\.
- Hindupur et al\. \(2025\)Sai Sumedh R\. Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba\. 2025\.[Projecting assumptions: The duality between sparse autoencoders and concept geometry](https://arxiv.org/abs/2503.01822)\.*Preprint*, arXiv:2503\.01822\.
- Karvonen et al\. \(2025\)Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu\-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Matthew Wearden, Arthur Conmy, Samuel Marks, and Neel Nanda\. 2025\.[SAEBench: A comprehensive benchmark for sparse autoencoders in language model interpretability](https://arxiv.org/abs/2503.09532)\.In*Proceedings of the 42nd International Conference on Machine Learning \(ICML\)*\.
- Korznikov et al\. \(2026\)Anton Korznikov, Andrey Galichin, Alexey Dontsov, Oleg Rogov, Ivan Oseledets, and Elena Tutubalina\. 2026\.[Sanity checks for sparse autoencoders: Do SAEs beat random baselines?](https://arxiv.org/abs/2602.14111)*Preprint*, arXiv:2602\.14111\.
- Lan et al\. \(2024\)Michael Lan, Philip Torr, Austin Meek, Ashkan Khakzar, David Krueger, and Fazl Barez\. 2024\.[Sparse autoencoders reveal universal feature spaces across large language models](https://arxiv.org/abs/2410.06981)\.*Preprint*, arXiv:2410\.06981\.
- Leask et al\. \(2025\)Patrick Leask, Bart Bussmann, Michael T Pearce, Joseph Isaac Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda\. 2025\.[Sparse autoencoders do not find canonical units of analysis](https://openreview.net/forum?id=9ca9eHNrdH)\.In*The Thirteenth International Conference on Learning Representations*\.
- Levina and Bickel \(2004\)Elizaveta Levina and Peter Bickel\. 2004\.[Maximum likelihood estimation of intrinsic dimension](https://proceedings.neurips.cc/paper_files/paper/2004/file/74934548253bcab8490ebd74afed7031-Paper.pdf)\.In*Advances in Neural Information Processing Systems*, volume 17, pages 777–784\. MIT Press\.
- Li et al\. \(2024\)Yuxiao Li, Eric J\. Michaud, David D\. Baek, Joshua Engels, Xiaoqing Sun, and Max Tegmark\. 2024\.[The geometry of concepts: Sparse autoencoder feature structure](https://arxiv.org/abs/2410.19750)\.*Preprint*, arXiv:2410\.19750\.
- Lieberum et al\. \(2024\)Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, Janos Kramar, Anca Dragan, Rohin Shah, and Neel Nanda\. 2024\.[Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2](https://doi.org/10.18653/v1/2024.blackboxnlp-1.19)\.In*Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP*, pages 278–300, Miami, Florida, US\. Association for Computational Linguistics\.
- Lin and Bloom \(2023\)Johnny Lin and Joseph Bloom\. 2023\.Neuronpedia: Interactive reference and tooling for analyzing neural networks\.[https://neuronpedia\.org](https://neuronpedia.org/)\.Software\.
- Lindsey et al\. \(2024\)Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah\. 2024\.[Sparse crosscoders for cross\-layer features and model diffing](https://transformer-circuits.pub/2024/crosscoders/index.html)\.Transformer Circuits Thread, Anthropic\. Research update, not peer\-reviewed\.
- Locatello et al\. \(2019\)Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem\. 2019\.[Challenging common assumptions in the unsupervised learning of disentangled representations](https://proceedings.mlr.press/v97/locatello19a.html)\.In*Proceedings of the 36th International Conference on Machine Learning*, volume 97 of*Proceedings of Machine Learning Research*, pages 4114–4124\. PMLR\.
- Makelov et al\. \(2024\)Aleksandar Makelov, George Lange, and Neel Nanda\. 2024\.[Towards principled evaluations of sparse autoencoders for interpretability and control](https://arxiv.org/abs/2405.08366)\.*Preprint*, arXiv:2405\.08366\.
- Makhzani and Frey \(2014\)Alireza Makhzani and Brendan Frey\. 2014\.[k\-sparse autoencoders](https://arxiv.org/abs/1312.5663)\.*Preprint*, arXiv:1312\.5663\.
- Marks et al\. \(2025\)Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller\. 2025\.[Sparse feature circuits: Discovering and editing interpretable causal graphs in language models](https://openreview.net/forum?id=I4e82CIDxv)\.In*The Thirteenth International Conference on Learning Representations*\.
- nostalgebraist \(2020\)nostalgebraist\. 2020\.Interpreting gpt: the logit lens\.[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru)\.
- Olah et al\. \(2020\)Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter\. 2020\.[Zoom in: An introduction to circuits](https://doi.org/10.23915/distill.00024.001)\.*Distill*\.Https://distill\.pub/2020/circuits/zoom\-in\.
- Park et al\. \(2024\)Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch\. 2024\.[The geometry of categorical and hierarchical concepts in large language models](https://openreview.net/forum?id=KXuYjuBzKo)\.In*ICML 2024 Workshop on Mechanistic Interpretability*\.
- Park et al\. \(2023\)Kiho Park, Yo Joong Choe, and Victor Veitch\. 2023\.[The linear representation hypothesis and the geometry of large language models](https://openreview.net/forum?id=T0PoOJg8cK)\.In*Causal Representation Learning Workshop at NeurIPS 2023*\.
- Paulo and Belrose \(2025\)Gonçalo Paulo and Nora Belrose\. 2025\.[Sparse autoencoders trained on the same data learn different features](https://arxiv.org/abs/2501.16615)\.*Preprint*, arXiv:2501\.16615\.
- Radford et al\. \(2019\)Alec Radford, Jeffrey Wu, Rewon Child, David Luan, and Dario Amodei\. 2019\.[Language models are unsupervised multitask learners](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)\.Technical report, OpenAI\.
- Rajamanoharan et al\. \(2024a\)Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda\. 2024a\.[Improving dictionary learning with gated sparse autoencoders](https://arxiv.org/abs/2404.16014)\.In*Advances in Neural Information Processing Systems*\.
- Rajamanoharan et al\. \(2024b\)Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda\. 2024b\.[Jumping ahead: Improving reconstruction fidelity with JumpReLU sparse autoencoders](https://arxiv.org/abs/2407.14435)\.*Preprint*, arXiv:2407\.14435\.
- Shu et al\. \(2025\)Dong Shu, Xuansheng Wu, Haiyan Zhao, Daking Rai, Ziyu Yao, Ninghao Liu, and Mengnan Du\. 2025\.[A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models](https://arxiv.org/abs/2503.05613)\.*Preprint*, arXiv:2503\.05613\.
- Team et al\. \(2024\)Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, and 179 others\. 2024\.[Gemma 2: Improving open language models at a practical size](https://arxiv.org/abs/2408.00118)\.*Preprint*, arXiv:2408\.00118\.
- Templeton et al\. \(2024\)Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C\. Daniel Freeman, Theodore R\. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, and 3 others\. 2024\.[Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)\.*Transformer Circuits Thread*\.
- Venhoff et al\. \(2024\)Constantin Venhoff, Anisoara Calinescu, Philip Torr, and Christian Schroeder de Witt\. 2024\.[Sage: Scalable ground truth evaluations for large sparse autoencoders](https://arxiv.org/abs/2410.07456)\.*Preprint*, arXiv:2410\.07456\.
## Appendix AAppendix
### A\.1Model and SAE Configurations
Table 9:Models and SAE configurations analyzed\. SAE Family reflects training methodology\. Community SAEs \(†\) used in controlled comparison \(Section[4](https://arxiv.org/html/2607.20596#S4)\)\.ModelParamsdmodeld\_\{\\text\{model\}\}LayersFeat/LSAE TypeFamilyTotalGPT2\-Small124M7681224,576ReLUres\-jb294,912Gemma\-2\-2B2B2,3042616,384JumpReLUGemmaScope425,984Gemma\-2\-9B9B3,5844216,384JumpReLUGemmaScope688,128Gemma\-3\-1B1B1,1522616,384JumpReLUGemmaScope425,984Llama\-3\.1\-8B8B4,0963232,768TopKLlamaScope1,048,576DeepSeek\-R18B4,0963232,768TopKLlamaScope1,048,576BatchTopK SAEs for budget pressure comparison:Chanind2B2,3042532,768BatchTopKCommunity819,200Chanind1B1,1522532,768BatchTopKCommunity819,200Community SAEs for controlled comparison \(Gemma\-2\-2B L1\):Chanind†2B2,304L118,432TopK \(k=25k\\\!=\\\!25\)Community18,432Chanind†2B2,304L118,432JumpReLUCommunity18,432GemmaScope2B2,304L116,384JumpReLUGemmaScope16,384SAEBench architectures \(Gemma\-2\-2B L12, Section[4](https://arxiv.org/html/2607.20596#S4)\):SAEBench2B2,304L1216,3848 variantsSAEBench131,072
### A\.2Expansion Factor Analysis
SAE expansion factor affects prevalence\. Using GPT2\-Small SAEs with widths 4×\\times, 8×\\times, 16×\\times, and 32×\\timesthe model dimension, single\-token prevalence increases sublinearly: from 0\.8% at 4×\\timesto 1\.48% at 32×\\times, following approximatelyST%∝W0\.3\\text\{ST\\%\}\\propto W^\{0\.3\}\. This suggests wider SAEs allocate more capacity to monosemantic features, but the rate of increase diminishes\. Both geometric signatures, the similarity ratio and the dimension ratio, remain stable across widths; single\-token features maintain their distinct structure regardless of SAE capacity\.
### A\.3Scaling Trend Fit Details
A descriptive power\-law fitST%∝Nα\\text\{ST\\%\}\\propto N^\{\\alpha\}across three GemmaScope/res\-jb models \(n=3n\\\!=\\\!3\) yieldsα=−0\.51±0\.08\\alpha=\-0\.51\\pm 0\.08for all\-layer analysis \(R2=0\.97R^\{2\}=0\.97\) andα=−1\.33\\alpha=\-1\.33for Layer 0 only \(R2=0\.73R^\{2\}=0\.73\)\. We emphasize that three data points cannot distinguish among functional forms; this fit is primarily descriptive\. The steeper Layer 0 exponent reflects concentrated single\-token encoding in early layers\. The geometric similarity ratio \(4\.5–5\.0×\\timesat L0\) remains stable across model families\. Intrinsic dimension ratios increase with depth \(1\.7–2×\\timesat L0,∼\\sim14×\\timesat mid layers; Table[11](https://arxiv.org/html/2607.20596#A1.T11)\), reflecting greater dimensionality divergence as representations become more compositional\.
We fit the power lawST%=a⋅Nα\\text\{ST\\%\}=a\\cdot N^\{\\alpha\}using weighted least squares on log\-transformed data\. Bootstrap resampling \(1000 iterations\) yields confidence intervals for each feature type\.
Table 10:Scaling trend parameters for different feature types \(n=3n\\\!=\\\!3; primarily descriptive\)\.Interpretation:The exponentα\\alphaindicates how feature prevalence scales with model size\. Negativeα\\alphameans the feature type becomes rarer in larger models\. Single\-token features show the steepest decline \(α=−0\.51\\alpha=\-0\.51\), meaning each 10×\\timesincrease in parameters reduces their prevalence to roughly one\-third \(10−0\.51≈0\.3110^\{\-0\.51\}\\approx 0\.31\)\. Morpheme and concept features decline more slowly, while polysemantic features remain constant \(α≈0\\alpha\\approx 0\), absorbing capacity freed by diminishing monosemantic features\. The intercept represents log\-prevalence atN=1N=1, used only for curve fitting\.
### A\.4Layer Distribution and Alignment Notes
Table 11:Intrinsic dimension by layer and feature type\.Table 12:Representative single\-token features \(GPT2\-Small Layer 0\)\.Tables[4](https://arxiv.org/html/2607.20596#S4.T4)and[11](https://arxiv.org/html/2607.20596#A1.T11)in the main text show Grassmannian alignment and intrinsic dimension\. Table[13](https://arxiv.org/html/2607.20596#A1.T13)shows per\-layer single\-token feature counts\. GPT2\-Small concentrates 91% of single\-token features \(331/364\) in Layer 0, with exponential decay\. Gemma\-2\-2B shows a distributed pattern under GemmaScope \(no layer exceeds 6%\), while BatchTopK concentrates features in later layers\. Gemma\-3\-1B shows intermediate behavior\. Feature counts reflect decoder\-alignment detection for Gemma models and activation\-based detection for GPT2\.
Table 13:Single\-token feature count by layer \(L0–L11\)\.Grassmannian alignment is computed between top\-dPCA=50d\_\{\\text\{PCA\}\}\{=\}50principal subspaces of decoder matrices at adjacent layers usingalign\(U,V\)=1dPCA∑i=1dPCAσi\(U⊤V\)\\text\{align\}\(U,V\)=\\frac\{1\}\{d\_\{\\text\{PCA\}\}\}\\sum\_\{i=1\}^\{d\_\{\\text\{PCA\}\}\}\\sigma\_\{i\}\(U^\{\\top\}V\), where the singular valuesσi\(U⊤V\)=cos\(θi\)\\sigma\_\{i\}\(U^\{\\top\}V\)=\\cos\(\\theta\_\{i\}\)recover the principal\-angle form of eq\.[5](https://arxiv.org/html/2607.20596#S3.E5)\. Values near 0 indicate orthogonal subspaces, i\.e\. major representational shift; values near 1 indicate aligned subspaces with gradual change\. GPT2’s low L0→\\toL1 alignment \(0\.26\) coincides with single\-token feature concentration \(91% in L0\)\. After this shift, alignment stabilizes at\>\>0\.9\. Gemma shows similarly low L0→\\toL1 alignment \(0\.23\)\.
Single\-token features occupy manifolds with intrinsic dimension 60–107, while polysemantic features span 118–180 dimensions, a 1\.7–2×\\timesratio\. MLE dimension estimation usesk=20k\\\!=\\\!20nearest neighbors followingLevina and Bickel \([2004](https://arxiv.org/html/2607.20596#bib.bib30)\)\.
### A\.5Multi\-Architecture Comparison
Table 14:ST prevalence across SAE architectures \(Gemma\-2\-2B, L12\)\.SAEAct\. FnWidthnSTn\_\{\\text\{ST\}\}%STGemmaScopeJumpReLU16k1580\.96SAEBenchJumpReLU16k1570\.96SAEBenchTopK16k1500\.92SAEBenchBatchTopK16k1560\.95SAEBenchMat\. BTK16k1620\.99SAEBenchReLU16k1620\.99SAEBenchGated16k630\.38SAEBenchP\-Anneal16k680\.42GemmaScopeJumpReLU65k6541\.00ChanindBTK→\\toJR32k2740\.84To test whether activation function choice affects prevalence independently of training recipe, we compare eight SAE architectures trained on the same model \(Gemma\-2\-2B\) at Layer 12 using decoder\-alignment detection \(Table[14](https://arxiv.org/html/2607.20596#A1.T14)\)\. All 16k\-width SAEs yield similar prevalence \(0\.92–0\.99%\), with only Gated \(0\.38%\) and P\-Anneal \(0\.42%\) lower\. This uniformity at mid\-layers contrasts with Layer\-0 differences, reinforcing that single\-token features are primarily pre\-compositional\.
### A\.6Top\-kkSensitivity
Detection usesk=20k=20top activating tokens \(Section[3](https://arxiv.org/html/2607.20596#S3)\); this is the cache window exposed by the Neuronpedia public API for the original analysis pipeline\. To characterize sensitivity, we re\-exported the top\-10 cache and computed detection counts under smallerkkon the GPT2\-Small cache, holding the gap and purity thresholds fixed at the canonical operating point: gap≥0\.3\\geq 0\.3, purity≥0\.6\\geq 0\.6, complete word\.
Table 15:Top\-kksensitivity on GPT2\-Small \(gap≥0\.3\\geq 0\.3, purity≥0\.6\\geq 0\.6, complete word\)\.Within the re\-exported cache the count declines and then flattens \(219→\\to171→\\to164, a 4% step fromk=7k=7tok=10k=10\): a feature with one strong primary activation and a long tail of weaker activations can satisfy the purity threshold at smallkk, but the additional cache positions surface enough secondary tokens to fail purity for some borderline features\. Thek=20k=20count comes from the original Neuronpedia export rather than this re\-export, so it is not on a common scale with the sweep; what the sweep bounds is how many borderline features a given cache window admits, and the main\-text results are computed atk=20k=20throughout\.
### A\.7Detection Threshold Ablation
Table 16:Extended ablation on detection thresholds showing precision\-recall tradeoff\.Sim = similarity ratio \(ST/Poly\); ID = intrinsic dimension ratio; Prec/Rec = qualitative precision\-recall\.
## Appendix BAdditional Visualizations
Figure 5:Geometric analysis \(mechanism overview\)\.\(A\) Grassmannian alignment for Gemma\-2\-2B: low L0→\\toL1 at 0\.23, gradual increase through middle layers\. \(B\) Intrinsic dimension: high at L0, compresses through middle layers, and rises again at final layers\. \(C\) Semantic categories: content nouns at 52% and proper nouns at 8\.4% dominate\.### B\.1Activation Shape Analysis
A key distinction between single\-token and polysemantic features lies in their activation distributions across the vocabulary\. Figure[6](https://arxiv.org/html/2607.20596#A2.F6)visualizes three complementary measures of activation shape\. Panel \(A\) shows how normalized activation values decay across the top\-10 activating tokens: single\-token features exhibit a sharp drop\-off \(from 1\.0 to∼\\sim0\.3 by rank 10\), while polysemantic features maintain relatively flat activation profiles \(from 1\.0 to∼\\sim0\.65\)\. Single\-token features respond strongly to one token and weakly to others; polysemantic features respond moderately to many tokens\.
Panel \(B\) shows the distribution of kurtosis \(peakedness\) across feature types\. Single\-token features have mean kurtosis 2\.65, indicating sharply peaked distributions with heavy tails, while polysemantic features have mean kurtosis−\-0\.08, indicating flat, uniform\-like distributions\. This difference is highly significant \(p<10−154p<10^\{\-154\}\) and provides a statistical signature independent of our gap ratio metric\. Panel \(C\) shows entropy distributions: single\-token features have lower entropy \(mean 2\.23\) than polysemantic \(mean 2\.28\), confirming their more concentrated activation patterns\. Together, these metrics validate that our detection method captures features with genuinely distinct activation behavior\.
Figure 6:Activation shape comparison\.\(A\) Mean activation decay across top\-10 tokens: single\-token features show sharp drop\-off while polysemantic remain flat\. \(B\) Kurtosis distribution: single\-token features are sharply peaked \(mean 2\.65,p<10−154p<10^\{\-154\}\) vs polysemantic \(mean−\-0\.08\)\. \(C\) Entropy distribution: single\-token features show lower entropy \(2\.23 vs 2\.28\), confirming concentrated activations\.
### B\.2Token Frequency Distribution
Why do single\-token features encode some tokens but not others? Figure[7](https://arxiv.org/html/2607.20596#A2.F7)shows the encoding pattern\. Single\-token features preferentially encode mid\-frequency tokens \(the “Goldilocks zone”\): in GPT2\-Small Layer 0, 66\.5% of single\-token features correspond to tokens ranked 1k–10k in frequency, compared to 26\.7% for polysemantic features, a 2\.5×\\timesenrichment\. Conversely, high\-frequency function words \(rank<<1k\) comprise only 5\.7% of single\-token features versus 20\.4% of polysemantic, and low\-frequency tokens \(rank\>\>10k\) are also underrepresented \(27\.8% vs 52\.9%\)\.
Mid\-frequency tokens are common enough to warrant dedicated neural representations, unlike rare tokens that must share capacity, yet not so frequent that the model can rely on compositional or contextual features as it does for function words\. Content nouns, proper nouns, and function words dominate the single\-token population \(Table[5](https://arxiv.org/html/2607.20596#S4.T5)\), categories that benefit from stable, context\-independent representations\.
Figure 7:Token frequency distribution by feature type\.Single\-token features show 2\.5×\\timesenrichment for mid\-frequency tokens \(rank 1k–10k\), avoiding both high\-frequency function words that require compositional encoding and rare tokens that must share capacity\. This “Goldilocks zone” pattern suggests single\-token features encode tokens that benefit from dedicated, context\-independent representations\.
### B\.3Cross\-Layer Persistence by Feature Type
Do single\-token features persist more across layers than polysemantic features? Figure[8](https://arxiv.org/html/2607.20596#A2.F8)quantifies this by computing the Z\-score of each feature’s cross\-layer similarity against a null distribution of random vector pairs\. Single\-token features show Z\-score 10\.7 for L0→\\toL1 persistence, meaning their cross\-layer similarity is 10\.7 standard deviations above what would be expected by chance\. Polysemantic features show Z\-score 4\.0, still significant, but 2\.7×\\timeslower than single\-token features \(p<10−73p<10^\{\-73\}\)\.
This difference is consistent with single\-token features maintaining stable directions across layers because they encode fixed token identities, while polysemantic features encode context\-dependent combinations and must transform more as contextual information accumulates\. Both feature types undergo the same L0→\\toL1 shift, but single\-token features are more likely to survive it with their direction intact\.
Figure 8:Cross\-layer persistence by feature type\.Single\-token features show Z\-score 10\.7 \(vs null distribution of random pairs\), 2\.7×\\timeshigher than polysemantic features \(Z=4\.0,p<10−73p<10^\{\-73\}\)\. The dashed line indicates the significance threshold \(Z=3\)\. This confirms single\-token features serve as stable reference points that persist through the L0→\\rightarrowL1 representational shift\.
### B\.4Semantic Preservation in Persistent Features
Persistence alone does not guarantee semantic preservation; a feature could maintain a similar direction while encoding entirely different concepts at each layer\. Figure[9](https://arxiv.org/html/2607.20596#A2.F9)tests whether persistent features actually share semantic meaning by comparing the auto\-generated explanations of matched feature pairs across layers\. We measure semantic similarity using sentence embeddings of the Neuronpedia explanations\.
High\-persistence feature pairs atZ\>50Z\>50, indicating very strong cross\-layer similarity, show mean explanation similarity of 0\.682, while low\-persistence pairs atZ<5Z<5show only 0\.157, a difference of 0\.525\. Manual inspection confirms this pattern: high\-persistence pairs are overwhelmingly single\-token features encoding proper names \(“Cruz,” “Matt,” “Scott”\), where the L5 and L6 explanations both reference the same token\. This result validates that our geometric persistence metric captures genuine semantic continuity, not mere directional accident\. The pattern replicates in Gemma\-2\-2B at 0\.61 vs 0\.32, difference 0\.29: a cross\-model phenomenon\.
Figure 9:Semantic preservation in persistent features\.High\-persistence feature pairs \(Z\>\>50\) show explanation similarity 0\.682, while low\-persistence pairs \(Z<<5\) show only 0\.157, a difference of 0\.525\. This confirms that geometric persistence corresponds to genuine semantic continuity: features that maintain similar directions across layers also maintain similar meanings\.Despite different tokenizers, GPT2 and Gemma single\-token features show semantic correspondence\. Table[17](https://arxiv.org/html/2607.20596#A2.T17)breaks down match rates by semantic category: proper nouns show highest cross\-model correspondence \(33%\) and function words lowest \(14%\)\. Partial semantic convergence despite zero surface\-form overlap\.
Table 17:Cross\-model feature correspondence by semantic category\.
### B\.5Decoder\-Embedding Alignment
If single\-token features truly encode token identity, their decoder vectors should align with the corresponding token embeddings\. Figure[10](https://arxiv.org/html/2607.20596#A2.F10)tests this prediction by computing cosine similarity between each feature’s decoder vector𝐰dec\\mathbf\{w\}\_\{\\text\{dec\}\}and the embedding of its top activating token𝐞tok\\mathbf\{e\}\_\{\\text\{tok\}\}\. Single\-token features show mean alignment 0\.674 versus 0\.392 for polysemantic, a 1\.72×\\timesdifference;t=14\.4t=14\.4,p<10−42p<10^\{\-42\}\.
This alignment is independent of activation patterns: single\-token decoder vectors point toward the token embedding subspace, close to the original input representation\. The alignment decreases with layer depth \(not shown\), consistent with representations becoming more abstract in later layers\. Combined with the logit attribution analysis, where 89% of single\-token features have the top\-token match, this confirms single\-token features function as bidirectional bridges between the embedding and unembedding spaces\.
Figure 10:Decoder\-embedding alignment\.Single\-token decoder vectors show 1\.72×\\timeshigher cosine similarity with corresponding token embeddings \(0\.674 vs 0\.392,p<10−42p<10^\{\-42\}\)\. This mechanistic validation confirms single\-token features recover directions close to the original embedding space, functioning as stable reference points for token identity\.
### B\.6Full\-Layer Causal Ablation Results
Tables[19](https://arxiv.org/html/2607.20596#A2.T19)–[19](https://arxiv.org/html/2607.20596#A2.T19)present layer\-wise necessitypp\-values for the four full\-layer experiments summarized in Table[6](https://arxiv.org/html/2607.20596#S4.T6)\. Allpp\-values are from Mann\-WhitneyUUtests of ST vs size\-matched random controls\.∗\\ast: BH\-correctedp<0\.05p<0\.05\.
Table 18:Gemma\-2\-2B: GemmaScope \(26/26 sig\) vs BatchTopK \(24/25 sig\)\.GemmaScope peaks at late layers \(L20\+\); BatchTopK peaks at mid\-layers \(L6–L9\)\.
Table 19:Gemma\-3\-1B: GemmaScope \(9/26 sig\) vs BatchTopK \(18/25 sig\)\.BatchTopK ST count inverts: peaks at late layers \(L17–L24\) vs GemmaScope early\-mid\.
### B\.7Full\-Layer Extension: Llama\-3\.1\-8B \(G1\) and Gemma\-2\-9B \(G2\)
Tables[21](https://arxiv.org/html/2607.20596#A2.T21)and[21](https://arxiv.org/html/2607.20596#A2.T21)report layer\-wise necessity \(pBHp\_\{\\text\{BH\}\}and meanΔ\\Deltalogit\) for two further full\-depth experiments: Llama\-3\.1\-8B×\\timesLlamaScope \(32 layers\) and Gemma\-2\-9B×\\timesGemmaScope \(42 layers\)\. Allpp\-values are BH\-corrected globally across all layers in each configuration;∗\\ast:pBH<0\.05p\_\{\\text\{BH\}\}<0\.05\.
Table 20:Llama\-3\.1\-8B×\\timesLlamaScope, full 32 layers \(G1\)\. 31/32 layers BH\-significant for necessity; peakpBH=2\.3×10−141p\_\{\\text\{BH\}\}=2\.3\\times 10^\{\-141\}at L1\. Anchor: 10/32 source layers show≥1\\geq 1significantly disrupted downstream layer\.Last layer \(L31\) hasnST=1n\_\{\\text\{ST\}\}=1feature, insufficient for significance\.
Table 21:Gemma\-2\-9B×\\timesGemmaScope, full 42 layers \(G2\)\. 42/42 BH\-significant; peakpBH=6\.1e−26p\_\{\\text\{BH\}\}=6\.1\\mathrm\{e\}\{\-26\}at L35; anchor 41/42; Spearmanρ\(depth,\|Δlogit\|\)=0\.81\\rho\(\\text\{depth\},\|\\Delta\\text\{logit\}\|\)=0\.81\(p<10−9p<10^\{\-9\}\)\.
### B\.8Paired Comparison: BatchTopK vs GemmaScope
To test whether the BatchTopK advantage reflects a genuine budget pressure effect, we perform a token\-ID matched paired comparison between GemmaScope and BatchTopK on two models \(Gemma\-2\-2B and Gemma\-3\-1B\)\. For each layer, we identify tokens detected as single\-token features bybothSAE types, yieldingN=627N\\\!=\\\!627matched pairs: 473 from Gemma\-2\-2B L0–L24 and 154 from Gemma\-3\-1B\.
Table 22:Paired comparison: BatchTopK vs GemmaScope \(token\-ID matched,N=627N\\\!=\\\!627\)\.BatchTopK features show stronger downstream anchoring \(p=1\.2×10−18p=1\.2\\times 10^\{\-18\},r=0\.36r\\\!=\\\!0\.36\) and stronger necessity \(p=0\.004p\\\!=\\\!0\.004\)\. The effect is layer\-dependent: early\-mid layers \(L0–L18\) show stronger necessity \(p=2\.0×10−6p=2\.0\\times 10^\{\-6\}\), while late layers converge, consistent with output\-proximal layers enforcing token identity regardless of activation function\.
### B\.9Controlled Comparison: TopK vs JumpReLU
Table 23:Controlled comparison: TopK vs JumpReLU \(N=142N\\\!=\\\!142token\-matched\)\.To disentangle activation function from training recipe, we compare community SAEs on Gemma\-2\-2B L1 with the same dictionary width \(18k\): TopK \(k∈\{25,50,75,100\}k\\in\\\{25,50,75,100\\\}\) and JumpReLU \(ℓ0∈\{14,26,95\}\\ell\_\{0\}\\in\\\{14,26,95\\\}\), all using tied decoders without norm constraint\. The controlled comparison \(Table[23](https://arxiv.org/html/2607.20596#A2.T23)\) shows theoppositedirection from the cross\-model result \(Table[22](https://arxiv.org/html/2607.20596#A2.T22)\): JumpReLU features cause more damage than TopK \(p=0\.036p\\\!=\\\!0\.036, matched\-pair rank\-biserialr=0\.20r\\\!=\\\!0\.20\)\. This dissociation indicates the activation function alone does not explain the cross\-SAE\-family differences\. Training recipe factors such as decoder norms, training scale, and post\-hoc conversion remain the residual candidates at the recipe level\. Comparing the community JumpReLU SAE with GemmaScope L1, both JumpReLU but differing in recipe, shows lower degradation for GemmaScope at median0\.510\.51vs2\.062\.06log2\\log\_\{2\}withp=3\.1×10−15p=3\.1\\times 10^\{\-15\}, consistent with unit\-norm decoders reducing per\-feature perturbation\.
### B\.10Sparsity Dose\-Response Analysis
Table 24:Dose\-response: meanΔ\\Deltalogit by sparsity level\. Tighter TopK budget monotonically increases necessity; JumpReLU shows no trend\.To test whether competitive budget size monotonically predicts necessity, we compare TopK SAEs at four sparsity levels \(k∈\{25,50,75,100\}k\\in\\\{25,50,75,100\\\}\) and JumpReLU SAEs at three levels \(ℓ0∈\{14,26,95\}\\ell\_\{0\}\\in\\\{14,26,95\\\}\), all on Gemma\-2\-2B L1 with the same dictionary width \(18k\)\. TopK shows a strong monotonic trend \(Table[24](https://arxiv.org/html/2607.20596#A2.T24)\): Pearsonr=0\.98r\\\!=\\\!0\.98\(p=0\.020p\\\!=\\\!0\.020, group\-level\) and Spearmanρ=0\.077\\rho\\\!=\\\!0\.077\(p=0\.04p\\\!=\\\!0\.04, per\-feature\)\.Δ\\Deltarank shows the same monotonic pattern \(r=−0\.928r\\\!=\\\!\-0\.928,p=0\.072p\\\!=\\\!0\.072\)\. JumpReLU shows no dose\-response \(r=−0\.015r\\\!=\\\!\-0\.015, ns\), consistent with its variable budget: changing the threshold does not create competitive allocation pressure\.
### B\.11Cross\-SAE Token\-Matched Analysis
To understand which tokens are robust versus sensitive to SAE methodology, we categorize theN=474N\\\!=\\\!474token\-matched pairs \(Gemma\-2\-2B, GemmaScope vs BatchTopK\) into semantic categories and compare convergence and anchoring dominance \(Table[25](https://arxiv.org/html/2607.20596#A2.T25)\)\.
Table 25:SAE methodology sensitivity by semantic category\. Convergent: both SAEs agree on anchored layer count \(±\\pm2\)\. BTK\>\>GS: fraction where BatchTopK shows stronger anchoring\.Two patterns emerge\. First,domain\-specific tokens are SAE\-robust: code and math tokens \(operatorname,mathbf,createElement\) show 93% convergence across SAE families, with matching anchor layer counts across 8–13 layers\. These tokens occupy a narrow, well\-defined region in the training distribution, leaving little room for SAE methodology to alter their representation\. For example, theoperatornamefeature shows identical anchor counts \(within±\\pm0 layers\) across all 13 layers where both SAEs detect it\.111operatornamesingle\-token features: GemmaScope[L3\#1853](https://neuronpedia.org/gemma-2-2b/3-gemmascope-res-16k/1853)\(Neuronpedia: “mathematical operators and functions”\), BatchTopK[L3\#940](https://neuronpedia.org/gemma-2-2b/3-res-matryoshka-dc/940)\(“the string operatorname”\);mathbfsingle\-token features: GemmaScope[L2\#212](https://neuronpedia.org/gemma-2-2b/2-gemmascope-res-16k/212)\(“mathematical terminologies”\), BatchTopK[L2\#5067](https://neuronpedia.org/gemma-2-2b/2-res-matryoshka-dc/5067)\(“linear algebra expressions”\)\. All identified via decoder\-alignment detection \(Section[3](https://arxiv.org/html/2607.20596#S3)\)\.
Second,function words are SAE\-dependent: only 29% convergence, with the highest BTK dominance \(80%\)\. The same function word can be causally necessary under one SAE but redundant under another\. For instance, “al” at L22 showsΔ\\Deltalogit=−0\.38=\-0\.38under GemmaScope but−2\.72\-2\.72under BatchTopK\.222alsingle\-token features: GemmaScope[L22\#2521](https://neuronpedia.org/gemma-2-2b/22-gemmascope-res-16k/2521)\(Neuronpedia: “the word et”\), BatchTopK[L22\#736](https://neuronpedia.org/gemma-2-2b/22-res-matryoshka-dc/736)\(“et al\. citation abbreviation”\)\. Identified via decoder\-alignment detection\.Function words are frequent enough that multiple SAE features can share their representation, making the allocation of causal importance sensitive to the competitive dynamics imposed by the activation function\.
This category\-dependent sensitivity suggests that SAE methodology comparisons should be stratified by token type: conclusions drawn from domain\-specific tokens \(where SAEs converge\) may not generalize to function words \(where they diverge\)\.
### B\.12Magnitude\-Matched Baseline
To test whether single\-token feature ablation damage reflects activation magnitude rather than feature type, we compare each single\-token feature against size\-matched random controls evaluated on the same input positions\. Across Gemma\-2\-2B GemmaScope \(N=2,626N\\\!=\\\!2\{,\}626pairs, 26 layers\) and BatchTopK \(N=4,574N\\\!=\\\!4\{,\}574pairs, 25 layers\), single\-token features cause more logit damage than controls under a pooled Wilcoxon test,p<10−94p<10^\{\-94\}and rank\-biserialr=0\.56r=0\.56–0\.700\.70\. Stratifying by activation magnitude quartile, the effect holds even in the lowest quartile \(Q1:p<0\.0001p<0\.0001,r=0\.27r=0\.27–0\.430\.43\) and strengthens with magnitude \(Q4:r=0\.69r=0\.69–0\.830\.83\)\. This rules out the alternative explanation that single\-token features are merely high\-activation features whose ablation damage reflects magnitude rather than functional role\.
### B\.13Control\-Population Inertness by Configuration
Every causal experiment pairs each single\-token feature with five magnitude\-matched random controls drawn from non\-single\-token features of the same SAE, measured at the same positions and target\-token readout\. Table[26](https://arxiv.org/html/2607.20596#A2.T26)reports the pooled control population per full\-depth configuration under the pipeline’s stored recovery flag \(ablated same\-layer logit\-lens rank within2×2\\timesof baseline\)\. Control ablations are causally inert on the paired token readouts in every configuration and under both SAE families: recovery is at least 99\.96% and the median relative rank displacement is 1\.000, so the anchored\-versus\-redundant family split does not appear in the non\-single\-token control population and is not a generic artifact of the ablation protocol\.
Table 26:Non\-single\-token control population per full\-depth configuration: counts, median signedΔlogit\\Delta\\text\{logit\}, recovery, and median relative rank displacement \(1\.000 = no effect\)\. Single\-token medians shown for reference\.
### B\.14Alignment\-Matched Null Control
Decoder\-alignment detection selects features whose decoder vector aligns with the target token’s input embedding, so the necessity effect could in principle be an artifact of that selection geometry\. The direct test is a null population matched on decoder–embedding cosine\.
#### Selection procedure and matching tolerance\.
For each single\-token feature we compute the cosine between every decoder vector in the SAE dictionary and the target token’s input embedding, exclude all detected single\-token features, and select controls within±0\.02\\pm 0\.02of the feature’s own cosine; when fewer than five in\-band candidates exist, we take the five nearest non\-single\-token features by cosine distance\. Exact matching is only partially constructible: the median single\-token feature has 0–2 in\-band candidates over the full 16k dictionary on Gemma\-2\-2B GemmaScope and 0 at all four analyzed Llama\-3\.1\-8B layers, where single\-token features at layer 1 have median cosine 0\.65 against a far lower non\-single\-token maximum\. At single\-token alignment levels, decoder–embedding alignment and single\-token behavior nearly coincide as populations\. The achieved matching gap for the nearest\-null controls is median\|Δcos\|=0\.18\|\\Delta\\cos\|=0\.18on GemmaScope and0\.510\.51on LlamaScope\.
#### In\-band subset\.
Pooling the measured controls that fall within±0\.02\\pm 0\.02of their single\-token feature’s cosine across five Gemma\-2\-2B layers \(315 in\-band controls vs\. 590 single\-token features, same measurement protocol\): control median signedΔlogit\\Delta\\text\{logit\}is\+0\.00004\+0\.00004versus−0\.0032\-0\.0032for single\-token features, one\-sided Mann\-WhitneyUUp=6\.1×10−20p=6\.1\\times 10^\{\-20\}, rank\-biserial 0\.37\. Per layer the comparison is significant at 4 of 5 layers \(p=3\.1×10−4p=3\.1\\times 10^\{\-4\},3\.5×10−33\.5\\times 10^\{\-3\},2\.8×10−42\.8\\times 10^\{\-4\},1\.4×10−121\.4\\times 10^\{\-12\}\); layer 6 has only 11 in\-band controls and its point estimate reverses, so we report it as inconclusive\. Alignment also does not predict ablation damage within the control population: Spearman correlation between a control’s cosine and its signedΔlogit\\Delta\\text\{logit\}is−0\.024\-0\.024\(p=0\.19p=0\.19,n=2,950n=2\{,\}950Gemma\-2\-2B control ablations\)\.
#### Nearest\-null comparison\.
Table[27](https://arxiv.org/html/2607.20596#A2.T27)reports single\-token features against the five nearest\-alignment non\-single\-token controls per target on Gemma\-2\-2B \(five layers\) and Llama\-3\.1\-8B \(four layers\)\. Necessity remains significant at eight of nine layer\-configurations under the signed one\-sided test, and the same eight survive BH correction within this nine\-test family\. As a clustering sensitivity check, collapsing each single\-token feature to a single paired comparison \(Wilcoxon signed\-rank on the feature’sΔlogit\\Delta\\text\{logit\}minus the median of its five controls\) leaves the same eight configurations significant\. The nearest\-alignment controls carry a non\-trivial effect of their own, most clearly on Llama\-3\.1\-8B layer 1, so a geometric component of the necessity effect is detectable; it does not, however, account for the single\-token effect, which exceeds the nearest available null in every configuration except Gemma\-2\-2B layer 6, where the single\-token effect itself is smallest\.
Table 27:Single\-token features vs\. nearest\-alignment non\-single\-token controls: mean signedΔlogit\\Delta\\text\{logit\}and one\-sided Mann\-Whitneyppper layer\-configuration\.Similar Articles
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.
From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness
This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.
WriteSAE: Sparse Autoencoders for Recurrent State
WriteSAE introduces the first sparse autoencoder that decomposes matrix cache writes in state-space and hybrid recurrent language models, enabling superior token-level interventions compared to existing methods.
Feature Starvation as Geometric Instability in Sparse Autoencoders
This paper identifies feature starvation in sparse autoencoders as a geometric instability and proposes adaptive elastic net SAEs (AEN-SAEs) to mitigate it without heuristics.
Decompose Sparsely Where You Should, Absorb Densely Where You Should No
The paper hypothesizes that language model activations contain a low-rank dense component that is inefficiently represented by sparse autoencoders (SAEs). By adding a linear bottleneck to absorb dense structure, the authors reduce dense latents and improve sparse probing performance on Gemma-2-2B.