Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

arXiv cs.CL Papers

Summary

This paper uses sparse autoencoders to decompose LLMs into interpretable features and shows that semantic features explain brain alignment with cortical semantic topography, generalizing across English, Chinese, and French.

arXiv:2605.23035v1 Announce Type: new Abstract: Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational neurolinguistics, yet why remains mechanistically unexplained. We address this gap by bridging sparse autoencoders (SAEs) from mechanistic interpretability with neural encoding models, decomposing GPT-2 XL and Llama-3.1-8B into 16K-32K interpretable features per layer. A human-validated taxonomy ($\kappa \geq 0.74$) reveals that semantic features alone recover 94% of peak encoding performance ($r=0.285$), substantially exceeding variance-matched baselines ($p<0.001$, $d=1.31$). Beyond this aggregate dominance, we test a novel cortical topography prediction: five semantic subcategories derived a priori from three independent neuroscience programs should map onto distinct brain regions. A formal convergence test confirms this alignment (Spearman $\rho=0.72$, $p<0.001$; hypergeometric $p=0.007$), demonstrating that SAE-discovered features recapitulate known cortical semantic organization at a granularity inaccessible to prior methods. SAE features further predict human reading times beyond lexical controls ($\Delta\mathrm{logLik}=38.4$, $p<0.001$), and an exploratory prediction-error analysis provides preliminary evidence that the brain additionally encodes unexpected semantic content. Results generalize across English, Chinese, and French.
Original Article
View Cached Full Text

Cached at: 05/25/26, 08:57 AM

# Sparse Autoencoders Map Brain–LLM Alignment onto Cortical Semantic Topography
Source: [https://arxiv.org/html/2605.23035](https://arxiv.org/html/2605.23035)
Dongxin Guo The University of Hong Kong Hong Kong, China bettyguo@connect\.hku\.hk &Jikun Wu Stellaris AI Limited Hong Kong, China hk950014@connect\.hku\.hk &Siu Ming Yiu The University of Hong Kong Hong Kong, China smyiu@cs\.hku\.hk

###### Abstract

Intermediate layers of large language models \(LLMs\) best predict human brain responses to language, one of the most robust findings in computational neurolinguistics, yet why remains mechanistically unexplained\. We address this gap by bridging sparse autoencoders \(SAEs\) from mechanistic interpretability with neural encoding models, decomposing GPT\-2 XL and Llama\-3\.1\-8B into 16K–32K interpretable features per layer\. A human\-validated taxonomy \(κ≥0\.74\\kappa\\geq 0\.74\) reveals that semantic features alone recover 94% of peak encoding performance \(r=0\.285r\{=\}0\.285\), substantially exceeding variance\-matched baselines \(p<0\.001p\{<\}0\.001,d=1\.31d\{=\}1\.31\)\. Beyond this aggregate dominance, we test a novel*cortical topography*prediction: five semantic subcategories derived*a priori*from three independent neuroscience programs should map onto distinct brain regions\. A formal convergence test confirms this alignment \(Spearmanρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001; hypergeometricp=0\.007p\{=\}0\.007\), demonstrating that SAE\-discovered features recapitulate known cortical semantic organization at a granularity inaccessible to prior methods\. SAE features further predict human reading times beyond lexical controls \(Δ\\DeltalogLik=38\.4\{=\}38\.4,p<0\.001p\{<\}0\.001\), and an exploratory prediction\-error analysis provides preliminary evidence that the brain additionally encodes unexpected semantic content\. Results generalize across English, Chinese, and French\.

Sparse Autoencoders Map Brain–LLM Alignment onto Cortical Semantic Topography

Dongxin GuoThe University of Hong KongHong Kong, Chinabettyguo@connect\.hku\.hkJikun WuStellaris AI LimitedHong Kong, Chinahk950014@connect\.hku\.hkSiu Ming YiuThe University of Hong KongHong Kong, Chinasmyiu@cs\.hku\.hk

## 1Introduction

What computational properties of language representations make them predictive of human brain activity? This question, rooted in decades of neurolinguistic theory\(Hale,[2001](https://arxiv.org/html/2605.23035#bib.bib1); Levy,[2008](https://arxiv.org/html/2605.23035#bib.bib23); Huthet al\.,[2016](https://arxiv.org/html/2605.23035#bib.bib2); Mitchellet al\.,[2008](https://arxiv.org/html/2605.23035#bib.bib24)\), has gained new urgency as large language models \(LLMs\) predict neural responses with remarkable accuracy\(Schrimpfet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib25); Goldsteinet al\.,[2022](https://arxiv.org/html/2605.23035#bib.bib26); Pereiraet al\.,[2018](https://arxiv.org/html/2605.23035#bib.bib27)\)\.Tuckuteet al\.\([2024b](https://arxiv.org/html/2605.23035#bib.bib28)\)demonstrated that LLM\-optimized stimuli causally drive the brain’s language network, andFedorenkoet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib29)\)argued that this network constitutes a natural kind whose properties align with LLM\-discovered representations\. The “direct fit” perspective\(Hassonet al\.,[2020](https://arxiv.org/html/2605.23035#bib.bib52)\)suggests this alignment arises because both brains and models are optimized for similar computational objectives\.

Yet a fundamental puzzle remains\. Brain\-predictive power is not uniform across depth:*intermediate*layers consistently outperform early and late layers\(Caucheteux and King,[2022](https://arxiv.org/html/2605.23035#bib.bib30); Toneva and Wehbe,[2019](https://arxiv.org/html/2605.23035#bib.bib3); Tuckuteet al\.,[2024b](https://arxiv.org/html/2605.23035#bib.bib28)\)\. This inverted\-U is among the most robust findings in computational neurolinguistics\(Maoet al\.,[2025](https://arxiv.org/html/2605.23035#bib.bib67)\)\. But*why*do intermediate layers best align with the brain?

Three non\-mutually\-exclusive accounts have been proposed\.\(1\) Predictive processing:under hierarchical predictive coding\(Rao and Ballard,[1999](https://arxiv.org/html/2605.23035#bib.bib31); Friston,[2005](https://arxiv.org/html/2605.23035#bib.bib32); Clark,[2013](https://arxiv.org/html/2605.23035#bib.bib34)\), neural responses reflect prediction errors; intermediate layers may generate the most brain\-like predictions\(Friston,[2010](https://arxiv.org/html/2605.23035#bib.bib33); Heilbronet al\.,[2022](https://arxiv.org/html/2605.23035#bib.bib35)\)\.\(2\) Feature discovery:LLMs discover rich linguistic features as a byproduct of prediction\(Antonello and Huth,[2023](https://arxiv.org/html/2605.23035#bib.bib36)\); intermediate layers contain the richest semantics before late\-layer specialization\.\(3\) Geometric artifact:intermediate layers may simply have higher\-dimensional representations that are more linearly decodable\(Kornblithet al\.,[2019](https://arxiv.org/html/2605.23035#bib.bib4)\)\. Distinguishing*what*the brain represents from*how*it computes requires decomposing representations into interpretable features\. Sparse autoencoders \(SAEs\) provide exactly this capability\(Hubenet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib5); Templetonet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib68)\)\.

We bridge mechanistic interpretability and neural encoding to provide a feature\-level decomposition of the intermediate\-layer advantage\. We use “mechanistic” here in the sense of*mechanistic interpretability*\(decomposing representations into interpretable features\), not as a claim of full causal\-mechanistic explanation\. Our contributions:

1. 1\.Primary empirical \(novel\):We derive five semantic subcategories*a priori*from the independent neuroscience programs of Huth et al\., Binder et al\., and Deniz et al\., then test whether SAE features recapitulate their predicted cortical topography\. A formal convergence test \(Spearman rank correlationρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001; hypergeometric overlapp=0\.007p\{=\}0\.007\) confirms significant alignment between the predicted and observed subcategory×\\timesregion patterns, demonstrating that SAE\-derived features capture neurally meaningful semantic distinctions at a granularity unavailable to prior methods\.
2. 2\.Methodological \(enabling\):We train SAEs on every fourth layer of GPT\-2 XL and Llama\-3\.1\-8B, extracting 16K–32K features per layer with a validated five\-way taxonomy \(κ≥0\.74\\kappa\\geq 0\.74for substantive categories\)\. We empirically demonstrate that SAEs outperform simpler word\-level semantic norm annotations for cortical topography mapping \(§[5\.2](https://arxiv.org/html/2605.23035#S5.SS2)\)\.
3. 3\.Supporting empirical \(confirmatory \+ exploratory\):Three converging analyses \(encoding models, variance partitioning, and activation patching with variance\-matched controls\) confirm at feature\-level resolution that semantic content dominates brain–LLM alignment, extending prior aggregate findings\(Kaufet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib37); Antonello and Huth,[2023](https://arxiv.org/html/2605.23035#bib.bib36)\)\. Behavioral validation shows SAE semantic features predict human reading times \(§[5\.3](https://arxiv.org/html/2605.23035#S5.SS3)\)\. An exploratory prediction\-error analysis provides preliminary evidence that the brain may additionally encode unexpected semantic content\.

Central claim\.SAE\-extracted semantic features simultaneously explain the well\-known intermediate\-layer advantage in brain–LLM alignment*and*a cortical topography that converges with established neuroscientific accounts of semantic organization: a single feature\-level account spanning two previously disconnected empirical regularities\.

Prior work byKaufet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib37)\)andAntonello and Huth \([2023](https://arxiv.org/html/2605.23035#bib.bib36)\)established that semantic content, specifically the lexical\-semantic dimension contrasted with syntactic structure, dominates brain alignment using category\-level probes and aggregate measures\. Our SAE decomposition advances beyond these predecessors in four ways: \(i\)*feature\-level granularity*\(16K–32K individual features per layer vs\.∼\{\\sim\}3–5 broad categories\) enabling the subcategory×\\timesbrain\-region interaction analysis \(§[4\.4](https://arxiv.org/html/2605.23035#S4.SS4)\); \(ii\)*a priori subcategories and formal convergence testing*against independent neuroscience programs \(§[4\.5](https://arxiv.org/html/2605.23035#S4.SS5)\); \(iii\)*behavioral validation*showing these features predict human reading times \(§[5\.3](https://arxiv.org/html/2605.23035#S5.SS3)\); and \(iv\) an exploratory test of semantic prediction errors \(§[5\.4](https://arxiv.org/html/2605.23035#S5.SS4)\)\. Our finer\-grained decomposition further reveals that the brain\-predictive variance attributable to semantic content is captured by*contextual*SAE features rather than by lexical features alone \(Table[2](https://arxiv.org/html/2605.23035#S4.T2)\), refining the lexical\-semantic emphasis ofKaufet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib37)\)\. Together, these contributions transform the question from*whether*semantics drives alignment to*which specific semantic features*drive alignment*where in the brain*\.

## 2Background and Related Work

#### The intermediate\-layer advantage\.

Schrimpfet al\.\([2021](https://arxiv.org/html/2605.23035#bib.bib25)\)systematically documented that intermediate layers best predict brain responses\.Caucheteux and King \([2022](https://arxiv.org/html/2605.23035#bib.bib30)\)confirmed this across 100\+ subjects;Goldsteinet al\.\([2022](https://arxiv.org/html/2605.23035#bib.bib26)\)found shared principles using ECoG\. The earliest layer\-wise alignment work used ELMo and BERT\(Jain and Huth,[2018](https://arxiv.org/html/2605.23035#bib.bib19); Toneva and Wehbe,[2019](https://arxiv.org/html/2605.23035#bib.bib3)\)\.Maoet al\.\([2025](https://arxiv.org/html/2605.23035#bib.bib67)\)showed earlier layers correspond to fast processing signals while later layers align with N400 amplitudes, connecting to work byMichaelovet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib53)\)demonstrating that LLM surprisal predicts single\-trial N400 amplitudes\.Kietzmannet al\.\([2019](https://arxiv.org/html/2605.23035#bib.bib51)\)demonstrated analogous hierarchical correspondence in vision\.

#### What drives alignment?

A central debate concerns which properties of LLM representations are responsible for their brain\-predictive power\. One line of work emphasizes*lexical\-semantic content*:Kaufet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib37)\)showed that lexical\-semantic information dominates LM–brain alignment when contrasted with syntactic structure, building on earlier work disentangling syntax and semantics in brain responses\(Caucheteuxet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib6)\)\. A second line emphasizes*feature discovery and contextualization*:Antonello and Huth \([2023](https://arxiv.org/html/2605.23035#bib.bib36)\)argued that brain\-predictive features emerge as a byproduct of next\-word prediction;Ootaet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib38)\)showed alignment depends on multiple linguistic dimensions across modalities; andVafaeiet al\.\([2026](https://arxiv.org/html/2605.23035#bib.bib74)\)demonstrated that fine\-grained semantic annotations improve brain encoding\. A third strand examines*training and scale*:Hosseiniet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib39)\)found brain alignment in models trained on only∼\{\\sim\}100M words,Pasquiouet al\.\([2022](https://arxiv.org/html/2605.23035#bib.bib7)\)showed training dynamics shape brain fit, andDoeriget al\.\([2025](https://arxiv.org/html/2605.23035#bib.bib69)\)showed geometric alignment in multilingual embeddings\. Beyond raw activations,Mathiset al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib70)\)showed XAI attributions outperform raw activations, andKumaret al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib40)\)deconstructed attention heads into functionally specialized components\. Methodologically, the representational similarity analysis tradition\(Kriegeskorte,[2008](https://arxiv.org/html/2605.23035#bib.bib59); Niliet al\.,[2014](https://arxiv.org/html/2605.23035#bib.bib21)\)offers a complementary view; the THINGS initiative\(Hebartet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib65)\)and Algonauts challenges\(Cichyet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib22)\)have advanced shared evaluation for object semantics, though we focus on naturalistic language stimuli\. We focus on encoding models \(rather than RSA\) for their superior voxelwise predictive resolution and ability to decompose variance by feature type\.Liuet al\.\([2025](https://arxiv.org/html/2605.23035#bib.bib77)\)provides complementary neuroscience\-inspired evaluation criteria for neural network representations\.

#### Cortical semantic organization\.

Huthet al\.\([2016](https://arxiv.org/html/2605.23035#bib.bib2)\)revealed systematic cortical semantic maps;Binderet al\.\([2009](https://arxiv.org/html/2605.23035#bib.bib46)\)established a core semantic network via meta\-analysis;Denizet al\.\([2019](https://arxiv.org/html/2605.23035#bib.bib47)\)showed modality\-invariant maps\.Patterson and Lambon Ralph \([2016](https://arxiv.org/html/2605.23035#bib.bib73)\)proposed the hub\-and\-spoke architecture\.Binderet al\.\([2016](https://arxiv.org/html/2605.23035#bib.bib61)\)developed a 65\-dimensional experiential attribute space, directly motivating our subcategory analysis\. Semantic feature production norms\(McRaeet al\.,[2005](https://arxiv.org/html/2605.23035#bib.bib62)\)provide converging behavioral evidence\.

#### Mechanistic interpretability\.

The residual stream framework\(Elhageet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib71)\)views transformers as iteratively refining representations\. SAEs decompose activations into monosemantic features\(Hubenet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib5)\);Templetonet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib68)\)scaled this to production models;Liet al\.\([2025](https://arxiv.org/html/2605.23035#bib.bib9)\)discovered spatial modules\. Recent work has raised important concerns: feature splitting at larger dictionary sizes\(Makelovet al\.,[2025](https://arxiv.org/html/2605.23035#bib.bib20)\), feature absorption\(Chaninet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib54)\), and questions about whether SAE features represent “true” computational features\(Hubenet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib5)\)\. A small body of concurrent work has begun to apply SAE\-style sparse decompositions to neural\-data analysis in adjacent domains \(visual brain encoding\(Wassermanet al\.,[2025](https://arxiv.org/html/2605.23035#bib.bib78)\); audio\-model interpretability\(Aparinet al\.,[2026](https://arxiv.org/html/2605.23035#bib.bib79)\); and sparse\-dictionary brain encoding\(Zeng and Gallant,[2026](https://arxiv.org/html/2605.23035#bib.bib80)\)\), developed in parallel with the present work\. To our knowledge no prior work has applied SAEs to decompose LLM–brain alignment for*language*stimuli using fMRI, and none has linked SAE features to a formal a priori cortical\-topography prediction; this is the gap we address\.

#### Predictive coding and surprisal\.

Surprisal theory\(Hale,[2001](https://arxiv.org/html/2605.23035#bib.bib1); Levy,[2008](https://arxiv.org/html/2605.23035#bib.bib23)\)predicts processing difficulty scales with negative log\-probability;Shainet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib41)\)andWilcoxet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib12)\)provided cross\-linguistic evidence, whileShainet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib41)\)established logarithmic effects at very large scale\. Hierarchical predictive coding\(Rao and Ballard,[1999](https://arxiv.org/html/2605.23035#bib.bib31); Friston,[2005](https://arxiv.org/html/2605.23035#bib.bib32),[2010](https://arxiv.org/html/2605.23035#bib.bib33)\)predicts neural responses reflect precision\-weighted prediction errors\.Heilbronet al\.\([2022](https://arxiv.org/html/2605.23035#bib.bib35)\)demonstrated a hierarchy of linguistic predictions;Caucheteuxet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib42)\)provided direct evidence of a predictive coding hierarchy in the human brain listening to speech, demonstrating that cortical responses track prediction errors at multiple representational levels\.

#### LLMs as cognitive models\.

Mahowaldet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib43)\)distinguished formal from functional competence\.McCoyet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib44)\)showed autoregressive objectives shape LLM representations in ways that both enable and constrain their cognitive plausibility, what they term “embers of autoregression\.”Pereiraet al\.\([2018](https://arxiv.org/html/2605.23035#bib.bib27)\)developed universal semantic decoders relevant to cross\-linguistic representation\.Misra and Mahowald \([2024](https://arxiv.org/html/2605.23035#bib.bib57)\)showed LLMs learn rare constructions from distributional statistics, whileTänzeret al\.\([2022](https://arxiv.org/html/2605.23035#bib.bib66)\)characterized the memorization–generalization boundary\.

## 3Methodology

### 3\.1Overview

Our approach proceeds in five stages \(Figure[1](https://arxiv.org/html/2605.23035#S3.F1)\)\. In compact form:

- •Stage 1 — Activations\.Run stories through GPT\-2 XL / Llama\-3\.1\-8B; extract residual\-stream activations every 4 layers\.
- •Stage 2 — Decomposition\.Train per\-layer SAEs \(16K–32K features\); label each via GPT\-4 \+ human validation into \{semantic, syntactic, lexical, prediction, other\}\.
- •Stage 3 — Encoding\.Fit voxelwise ridge regression from each feature subset to fMRI; partition unique vs\. shared variance against count\- and variance\-matched baselines\.
- •Stage 4 — Topography\.Derive five semantic subcategories*a priori*from independent neuroscience programs; test whether the predicted subcategory×\\timesregion matrix matches the observed pattern \(rank\-correlation, hypergeometric, Mantel\)\.
- •Stage 5 — Convergent validation\.Activation patching, GLMM\-based reading\-time prediction, and an exploratory semantic\-prediction\-error analysis\.

StoryStimuliLLM\(GPT\-2 XL /Llama\-3\.1\)𝐱ℓ∈ℝd\\mathbf\{x\}\_\{\\ell\}\\in\\mathbb\{R\}^\{d\}per layerSAE \+Categ\.𝐟​\(𝐱\)∈ℝM\\mathbf\{f\}\(\\mathbf\{x\}\)\\in\\mathbb\{R\}^\{M\}16–32K featsfMRIRecordingA PrioriSubcateg\.Encoding \+ConvergencePatching \+Behav\. Val\.Stage 1Stage 2Stages 3–4Stage 5Figure 1:Five\-stage pipeline: extract LLM activations→\\toSAE decomposition→\\toencoding models with variance partitioning→\\toa priori subcategory convergence→\\tocausal/behavioral/prediction\-error validation\. Dashed line indicates features also feed the patching stage\.
### 3\.2Language Models and SAE Training

We analyzeGPT\-2 XL\(1\.5B, 48 layers,d=1600d\{=\}1600\) andLlama\-3\.1\-8B\(8B, 32 layers,d=4096d\{=\}4096\), extracting residual stream activations at every fourth layer\. For each layerℓ\\ell, we train an SAE:

𝐟​\(𝐱\)\\displaystyle\\mathbf\{f\}\(\\mathbf\{x\}\)=ReLU​\(𝐖enc​\(𝐱−𝐛d\)\+𝐛e\)\\displaystyle=\\text\{ReLU\}\(\\mathbf\{W\}\_\{\\text\{enc\}\}\(\\mathbf\{x\}\-\\mathbf\{b\}\_\{d\}\)\+\\mathbf\{b\}\_\{e\}\)\(1\)𝐱^\\displaystyle\\hat\{\\mathbf\{x\}\}=𝐖dec​𝐟​\(𝐱\)\+𝐛d\\displaystyle=\\mathbf\{W\}\_\{\\text\{dec\}\}\\mathbf\{f\}\(\\mathbf\{x\}\)\+\\mathbf\{b\}\_\{d\}\(2\)minimizingℒ=‖𝐱−𝐱^‖22\+λ​‖𝐟​\(𝐱\)‖1\\mathcal\{L\}=\\\|\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\\\|\_\{2\}^\{2\}\+\\lambda\\\|\\mathbf\{f\}\(\\mathbf\{x\}\)\\\|\_\{1\}, withM=16,384M\{=\}16\{,\}384\(GPT\-2 XL\) or32,76832\{,\}768\(Llama\), L0≈\\approx50, trained on 500M tokens\. ReconstructionR2≥0\.95R^\{2\}\\geq 0\.95at all layers; encoding from reconstruction error achieves onlyr=0\.031r\{=\}0\.031, confirming SAEs preserve brain\-relevant information \(Appendix[N](https://arxiv.org/html/2605.23035#A14)\)\.

#### Addressing SAE feature quality concerns\.

Feature splitting\(Makelovet al\.,[2025](https://arxiv.org/html/2605.23035#bib.bib20)\), absorption\(Chaninet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib54)\), and questions about “true” features\(Hubenet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib5)\)motivate three design choices: \(1\) robustness across dictionary sizes \(8K–32K; Appendix[R](https://arxiv.org/html/2605.23035#A18)\); \(2\) soft probabilistic categorization yielding identical rankings \(Appendix[F](https://arxiv.org/html/2605.23035#A6)\); and \(3\) subcategory analysis operating at the level of feature\-type aggregates\.

### 3\.3Feature Categorization and Validation

Within stage 2 of the pipeline, we categorize features into five types using a two\-pass protocol:

Pass 1 \(Automated\):GPT\-4 assigns each feature to*semantic*,*syntactic*,*lexical*,*prediction*\(correlationr\>0\.5r\>0\.5with tuned lens output entropy\), or*other/uninterpretable*\. We acknowledge that using one LLM to categorize another introduces potential shared biases\.

Pass 2 \(Human validation\):Two graduate\-student annotators \(computational linguistics, compensated $25/hour; Appendix[D](https://arxiv.org/html/2605.23035#A4)\) validated 500 features per layer \(100/category\), achievingκ=0\.81\\kappa\{=\}0\.81\. The confusion matrix \(Table[15](https://arxiv.org/html/2605.23035#A27.T15), Appendix\) shows 14% overall disagreement; 11% of GPT\-4\-labeled semantic features were relabeled by annotators\. Per\-categoryκ\\kapparanges from 0\.74 \(prediction\) to 0\.83 \(syntactic\); the residual “other” category has lower agreement \(κ=0\.58\\kappa\{=\}0\.58\)\. Soft probabilistic categorization yields qualitatively identical results \(Appendix[F](https://arxiv.org/html/2605.23035#A6)\); Appendix[E](https://arxiv.org/html/2605.23035#A5)provides qualitative examples including ambiguous cases\. Appendix[AB](https://arxiv.org/html/2605.23035#A28)additionally reports a cross\-LLM relabeling check, a label\-perturbation sensitivity analysis, and a confidence\-thresholded re\-run, all of which preserve the qualitative pattern of results\. As illustrative examples \(full table in Appendix[E](https://arxiv.org/html/2605.23035#A5)\), at GPT\-2 XL L24 a*concrete\-semantic*feature fires on “thedogran across the”; an*affective\-semantic*feature on “terrifiedof the dark”; a*social\-semantic*feature on “shebelievedthat he”; a*syntactic*feature on “the manwhocame”; and a*prediction*feature on “Inconclusion, we find”\.

### 3\.4Neural and Behavioral Data

Primary \(fMRI\):We use the publicly released naturalistic\-language fMRI dataset ofLeBelet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib48)\)\(UTS dataset; 8 native\-English participants, 27 stories,∼\{\\sim\}6 hours/subject; TR=\{=\}2\.0 s\)\. Because this dataset does not include a per\-participant functional language localizer, we restrict analysis to language\-network voxels using the group\-level*anatomical*language\-network parcellation summarized inFedorenkoet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib29)\), intersected with each subject’s native\-space gray matter mask, yielding∼\{\\sim\}5,000 bilateral voxels per subject after motion\-spike censoring \(FD\>\{\>\}0\.5 mm\) and high\-pass filtering \(1/1281/128Hz\)\. LLM activations are extracted per word and downsampled to the fMRI TR by averaging within each TR window after applying a canonical hemodynamic response function \(HRF\) convolution; full preprocessing and feature\-alignment details are reported in Appendix[B](https://arxiv.org/html/2605.23035#A2)\. For the subcategory analysis, we further parcellate the language network into five regions \(posterior temporal, anterior temporal, inferior frontal, angular gyrus, and dorsomedial prefrontal cortex; dmPFC\) followingFedorenkoet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib29)\)\.

Behavioral:Natural Stories Corpus\(Boyce and Levy,[2023](https://arxiv.org/html/2605.23035#bib.bib75)\)\(self\-paced reading times, 181 participants, 10 stories\) and the Provo Corpus\(Luke and Christianson,[2017](https://arxiv.org/html/2605.23035#bib.bib55)\)\(eye\-tracking with first\-fixation, gaze, and total reading time, 84 participants\)\. Both are publicly available\.

Generalization \(fMRI\):For the Llama\-3\.1\-8B generalization check \(§[5\.3](https://arxiv.org/html/2605.23035#S5.SS3), §[8](https://arxiv.org/html/2605.23035#S8)\), we use two publicly available naturalistic\-listening fMRI datasets from the cross\-linguistic “Little Prince” tradition: a Mandarin Chinese subset \(15 native speakers, 2 chapters\) and a French subset \(12 native speakers, 1 chapter\)\. Both are open\-access datasets whose original collection followed approved institutional protocols \(informed consent and ethics approvals documented in the corresponding data\-paper releases\)\. We did*not*collect any new human data for this study\. Dataset identifiers, preprocessing parameters, and language\-network ROI definitions for these subsets are reported in Appendix[B](https://arxiv.org/html/2605.23035#A2)\.

### 3\.5Encoding Models and Variance Partitioning

For each layerℓ\\elland feature subsetSS, we train voxelwise ridge regressiony^v=𝐰v⊤​𝐟S\+bv\\hat\{y\}\_\{v\}=\\mathbf\{w\}\_\{v\}^\{\\top\}\\mathbf\{f\}\_\{S\}\+b\_\{v\}, evaluated via 4\-fold cross\-validation \(Pearsonrr, bootstrapped 95% CIs\)\. Cross\-validation used a leave\-KK\-stories\-out scheme \(K=K\{=\}6–7 stories per fold\), ensuring no temporal leakage\. The regularization parameterλ\\lambdawas selected via nested cross\-validation within each training fold, searching overλ∈\{100,101,…,106\}\\lambda\\in\\\{10^\{0\},10^\{1\},\\ldots,10^\{6\}\\\}, independently for each feature subset\. The features\-to\-data ratio \(∼\{\\sim\}6,700 semantic features,∼\{\\sim\}5,000 voxels,∼\{\\sim\}6 hours of data\) is within the well\-regularized regime for ridge regression; the effective dimensionality of semantic features after accounting for collinearity is∼\{\\sim\}1,200 \(estimated via participation ratio of the feature covariance matrix; Appendix[X](https://arxiv.org/html/2605.23035#A24)\)\.

Unique variance:Δ​Runique2​\(S\)=Rfull2−Rfull∖S2\\Delta R^\{2\}\_\{\\text\{unique\}\}\(S\)=R^\{2\}\_\{\\text\{full\}\}\-R^\{2\}\_\{\\text\{full\}\\setminus S\}\. This additive decomposition assumes independent contributions; shared variance \(∼\{\\sim\}22%\) absorbs interactions\. Shapley values\(Covertet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib13)\)relax this assumption \(Appendix[M](https://arxiv.org/html/2605.23035#A13)\)\.

We construct two baselines: \(i\)count\-matched random: same number of features as semantic subset \(10 seeds\); \(ii\)variance\-matched random: random features matching total L2 activation variance\.

### 3\.6A Priori Subcategory Derivation

To avoid post\-hoc pattern matching, we derive semantic subcategories and their predicted cortical topography*a priori*from three independent neuroscience programs, constructing a predicted subcategory×\\timesregion matrix*before*examining SAE results:

Binder et al\. \(2009\):Meta\-analysis of 120 neuroimaging studies identifies seven core semantic regions\. We extract predicted mappings from their Table 2 and Figure 5: sensorimotor/concrete content→\\toposterior temporal and angular gyrus; affective content→\\toventromedial PFC and anterior temporal; social/interpersonal content→\\toinferior frontal and medial prefrontal; spatial content→\\toangular gyrus and posterior parietal; temporal/causal content→\\tolateral temporal cortex\.

Huth et al\. \(2016\):Cortical semantic atlas from naturalistic speech reveals category\-specific topography \(their Figures 2–3\): concrete objects and tactile properties cluster in lateral temporal cortex; emotional and mental content concentrates in anterior temporal and prefrontal areas; social interaction content maps to inferior frontal and medial frontal cortex; spatial/locational semantics activate angular gyrus and retrosplenial cortex\.

Deniz et al\. \(2019\):Modality\-invariant semantic maps \(their Figure 3\) confirm that the category\-region associations in Huth et al\. generalize across listening and reading modalities, with concreteness in posterior temporal, affect in anterior temporal, and spatial relations in angular gyrus\.

From the intersection of these three programs, we derive five a priori subcategories with predicted region preferences \(Table[16](https://arxiv.org/html/2605.23035#A29.T16)\)\. A cell in the predicted subcategory×\\timesregion matrix is marked as “predicted primary” when at least two of the three source programs explicitly map that subcategory to that region in their primary findings; Appendix[AD](https://arxiv.org/html/2605.23035#A30)reports the source\-paper citation supporting each predicted cell, making the derivation fully transparent and auditable\. We additionally verified robustness against alternative subcategorizations: using ’s\([2016](https://arxiv.org/html/2605.23035#bib.bib61)\)65\-dimensional experiential attribute space collapsed into analogous groups, and using feature production norms fromMcRaeet al\.\([2005](https://arxiv.org/html/2605.23035#bib.bib62)\)\(Appendix[V](https://arxiv.org/html/2605.23035#A22)\)\.

These mappings \(Table[16](https://arxiv.org/html/2605.23035#A29.T16), Appendix\) predict: concreteness/animacy→\\toposterior temporal and angular gyrus; event structure→\\toposterior temporal and inferior frontal; affect/emotion→\\toanterior temporal and dmPFC; social/mental→\\toinferior frontal and anterior temporal; spatial/locational→\\toangular gyrus and posterior temporal\.

### 3\.7Activation Patching

For subsetSSat layerℓ\\ell, we replace activations with corpus\-mean values and measureΔ​rv2​\(S,ℓ\)\\Delta r^\{2\}\_\{v\}\(S,\\ell\)\. We additionally re\-fit ridge regression on ablated activations to control for encoding\-model bias\. Mean\-ablation is a coarse intervention\(Wuet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib14)\); results are*evidence consistent with*a causal role, not definitive proof\. Interchange interventions\(Geigeret al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib15)\)would be stronger but require matched stimulus pairs infeasible with naturalistic stimuli\.

## 4Results

### 4\.1The Intermediate\-Layer Advantage

Both models show inverted\-U profiles \(Figure[2](https://arxiv.org/html/2605.23035#S4.F2)\): GPT\-2 XL peaks at L20–24 \(r=0\.310r\{=\}0\.310, 95% CI\[0\.278,0\.342\]\[0\.278,0\.342\]\); Llama at L12–16 \(r=0\.341r\{=\}0\.341,\[0\.305,0\.377\]\[0\.305,0\.377\]\)\. SAE\-reconstructed activations track raw withinΔ​r≤0\.01\\Delta r\\leq 0\.01\.

LayerMeanrr0\.10\.20\.30\.4012243644GPT\-2 XLLlama\-3\.1SAE recon\.

Figure 2:Brain prediction across layers\. Both models show the inverted\-U; SAE reconstructions \(dashed\) track raw activations withinΔ​r≤0\.01\\Delta r\\leq 0\.01\.
### 4\.2Confirmatory Validation: Semantic Features Dominate

Confirming prior findings\(Kaufet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib37); Antonello and Huth,[2023](https://arxiv.org/html/2605.23035#bib.bib36)\)at feature\-level resolution, Table[1](https://arxiv.org/html/2605.23035#S4.T1)shows semantic features alone achiever=0\.285r\{=\}0\.285at L24, recovering 94% of SAE\-reconstructed encoding \(r=0\.304r\{=\}0\.304\) and substantially outperforming both thecount\-matched random baseline\(r=0\.198r\{=\}0\.198,p<0\.001p\{<\}0\.001,d=1\.54d\{=\}1\.54, 95% CI\[0\.181,0\.215\]\[0\.181,0\.215\]\) and thevariance\-matched random baseline\(r=0\.213r\{=\}0\.213,p<0\.001p\{<\}0\.001,d=1\.31d\{=\}1\.31, 95% CI\[0\.195,0\.231\]\[0\.195,0\.231\]\)\. In variance space, semantic features explain 88% of the model’s captured neural variance\. Sensitivity: reclassifying the 11% false\-positive semantic features reduces encoding tor=0\.274r\{=\}0\.274\(stillp<0\.001p\{<\}0\.001vs\. random\)\.

Table 1:Encoding performance \(meanrr\) by feature type, GPT\-2 XL\. Bootstrapped 95% CIs in Appendix[O](https://arxiv.org/html/2605.23035#A15)\. Rnd\.†: count\-matched random\.
### 4\.3Variance Partitioning

Table[2](https://arxiv.org/html/2605.23035#S4.T2)reports unique variance at L24\. Semantic features contributeΔ​Runique2=0\.048\\Delta R^\{2\}\_\{\\text\{unique\}\}\{=\}0\.048\(52% of total explained variance\)\. The 28% shared variance \(the remaining proportion of total explained variance not captured by any single category’s unique contribution\) likely reflects semantic\-syntactic interactions\(Caucheteuxet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib6)\)\. Shapley decomposition confirms the ranking with semantic share slightly higher \(65% vs\. 52%; Appendix[M](https://arxiv.org/html/2605.23035#A13)\)\.

#### On feature\-count imbalance\.

A natural concern is whether semantic dominance is an artifact of feature\-count imbalance: at L24, 41% of features are categorized as semantic versus 11% as lexical \(Appendix[P](https://arxiv.org/html/2605.23035#A16)\)\. We address this with three complementary controls\. \(i\) The*variance\-matched random baseline*\(§[4\.2](https://arxiv.org/html/2605.23035#S4.SS2)\) selects random features until their total L2 activation variance matches that of the semantic subset; this still falls0\.0720\.072short of semantic encoding \(r=0\.213r\{=\}0\.213vs\.0\.2850\.285,d=1\.31d\{=\}1\.31\)\. \(ii\) Shapley values\(Covertet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib13)\)are by construction insensitive to the number of features per coalition member; semantic contribution remains dominant \(65%; Appendix[M](https://arxiv.org/html/2605.23035#A13)\)\. \(iii\) Lexical features contributeΔ​Runique2=0\.000\\Delta R^\{2\}\_\{\\text\{unique\}\}\{=\}0\.000despite being the second\-largest category at*early*layers \(Appendix[P](https://arxiv.org/html/2605.23035#A16)\), demonstrating that feature count alone does not predict unique variance\. Together these controls indicate that the semantic\-feature advantage is not a counting artifact\.

Table 2:Variance partitioning at GPT\-2 XL L24\. Brackets: bootstrapped 95% CIs\.†% total = proportion of total explained variance \(Rfull2=0\.092R^\{2\}\_\{\\text\{full\}\}\{=\}0\.092\); shared =Rfull2R^\{2\}\_\{\\text\{full\}\}minus sum of unique variances \(0\.066\)\.

### 4\.4Cortical Topography of Semantic Alignment

We map semantic SAE features to the five a priori subcategories derived in §[3\.6](https://arxiv.org/html/2605.23035#S3.SS6)and test region\-specific alignment across five language\-network regions \(Figure[4](https://arxiv.org/html/2605.23035#A26.F4), Appendix; Figure[3](https://arxiv.org/html/2605.23035#S4.F3)\)\.

A significant subcategory×\\timesregion interaction \(F​\(16,112\)=3\.87F\(16,112\)\{=\}3\.87,p<0\.001p\{<\}0\.001, permutation test, 10,000 iterations\) reveals systematic dissociations\. After FDR correction\(Benjamini and Hochberg,[1995](https://arxiv.org/html/2605.23035#bib.bib56)\)atq<0\.05q\{<\}0\.05,7 of 25 cells survive, and these cells form a neuroanatomical pattern consistent with the a priori predictions: concreteness features best predict posterior temporal cortex \(r=0\.141r\{=\}0\.141\) and angular gyrus \(r=0\.131r\{=\}0\.131\); affect features predict anterior temporal regions \(r=0\.128r\{=\}0\.128\) and dmPFC \(r=0\.121r\{=\}0\.121\); social/mental features predict inferior frontal cortex \(r=0\.119r\{=\}0\.119\) and anterior temporal cortex \(r=0\.122r\{=\}0\.122\); and spatial features predict angular gyrus \(r=0\.135r\{=\}0\.135\) and posterior temporal cortex \(r=0\.124r\{=\}0\.124\)\. We quantify this convergence formally in §[4\.5](https://arxiv.org/html/2605.23035#S4.SS5)\.

### 4\.5Formal Convergence Test

We formally test whether the observed subcategory×\\timesregion pattern matches the a priori predictions derived in §[3\.6](https://arxiv.org/html/2605.23035#S3.SS6)using three complementary tests\.

\(1\) Rank correlation\.We flatten the predicted preference matrix \(Table[16](https://arxiv.org/html/2605.23035#A29.T16)\) into an ordinal vector \(1 = predicted primary region, 0 = not predicted\) and the observedrr\-value matrix into a continuous vector\. Spearman rank correlation:ρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001\(permutation null: 10,000 shuffles of row/column labels; 95% CI\[0\.48,0\.88\]\[0\.48,0\.88\]\)\. The observed alignment substantially exceeds chance\.

\(2\) Hypergeometric overlap\.Of the 25 cells, the a priori matrix predictsK=10K\{=\}10cells as “primary region” associations\. Of the 7 FDR\-surviving cells, 6 fall within these 10 predicted cells\. The hypergeometric probability of≥\\geq6 overlapping cells givenK=10K\{=\}10predicted out of 25 total and 7 observed significant isp=0\.007p\{=\}0\.007\(Fisher’s exact test: OR=21\.0=21\.0, 95% CI\[1\.9,229\]\[1\.9,229\]\)\.

\(3\) Mantel test\.We compute the Pearson correlation between the full 5×\\times5 predicted dissimilarity matrix and the observed encodingrr\-value matrix, testing significance via 10,000 row/column permutations:rMantel=0\.64r\_\{\\text\{Mantel\}\}\{=\}0\.64,p=0\.002p\{=\}0\.002\.

Figure[3](https://arxiv.org/html/2605.23035#S4.F3)shows the predicted and observed matrices side by side\. All three tests converge on the same conclusion: the SAE\-derived cortical topography significantly matches independent neuroscience predictions that were formulated without any reference to LLMs or SAEs\.

A priori predictedPTATIFAGdmConEvtAffSocSpaObserved \(SAE\)PTATIFAGdm\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*Figure 3:Predicted vs\. observed subcategory×\\timesregion patterns\. Left: a priori predictions fromBinderet al\.\([2009](https://arxiv.org/html/2605.23035#bib.bib46)\)/Huthet al\.\([2016](https://arxiv.org/html/2605.23035#bib.bib2)\)/Denizet al\.\([2019](https://arxiv.org/html/2605.23035#bib.bib47)\)\(dark = predicted primary association\)\. Right: observed SAE encodingrr\-values with FDR\-significant cells marked\. Formal convergence:ρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001; hypergeometricp=0\.007p\{=\}0\.007; Mantelr=0\.64r\{=\}0\.64,p=0\.002p\{=\}0\.002\.Statistical power\.Post\-hoc power analysis \(f=0\.42f\{=\}0\.42, N=\{=\}8,α=\\alpha\{=\}0\.05\) yields power=0\.81=0\.81\. The 7/25 FDR\-surviving cells are consistent with power constraints; observed effect sizes \(r=0\.07r\{=\}0\.07–0\.140\.14\) match typical voxelwise values in naturalistic neuroimaging\(Huthet al\.,[2016](https://arxiv.org/html/2605.23035#bib.bib2); Schrimpfet al\.,[2021](https://arxiv.org/html/2605.23035#bib.bib25)\)\(Appendix[W](https://arxiv.org/html/2605.23035#A23)\)\.

## 5Further Analysis

### 5\.1Activation Patching

At L24 \(Figure[5](https://arxiv.org/html/2605.23035#A32.F5), Appendix\): ablating semantic features producesΔ​r2=−0\.14\\Delta r^\{2\}\{=\}\{\-\}0\.14\(p<0\.001p\{<\}0\.001,d=1\.82d\{=\}1\.82, 95% CI\[−0\.17,−0\.11\]\[\-0\.17,\-0\.11\]\); syntactic:−0\.02\{\-\}0\.02\(p=0\.04p\{=\}0\.04\); prediction:−0\.01\{\-\}0\.01\(p=0\.12p\{=\}0\.12\)\. Count\-matched random:−0\.05\{\-\}0\.05\.Variance\-matched random\(matching total L2 variance removed\):−0\.06\{\-\}0\.06\(d=0\.74d\{=\}0\.74\), still significantly less than semantic \(p<0\.001p\{<\}0\.001\)\. These results are*evidence consistent with*a causal role for semantic features, though mean\-ablation is a coarse intervention\.

#### Refitted and non\-semantic baselines\.

To address encoding\-model bias\(Wuet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib14)\), we re\-fitted ridge regression on ablated \(non\-semantic\) activations:r=0\.194r\{=\}0\.194\. A model trained*exclusively*on non\-semantic features from scratch \(syntactic \+ prediction \+ lexical \+ other\) achievesr=0\.189r\{=\}0\.189\. Both fall substantially short of the full model \(r=0\.304r\{=\}0\.304\), confirming semantic features carry genuinely unique brain\-predictive information\.

### 5\.2SAE Features vs\. Word\-Level Semantic Norms

A key question is whether the cortical topography requires SAE decomposition or could emerge from simpler word\-level annotations\. We compared SAE features against word\-level norms \(concreteness\(Brysbaertet al\.,[2013](https://arxiv.org/html/2605.23035#bib.bib63)\), valence/arousal\(Warrineret al\.,[2013](https://arxiv.org/html/2605.23035#bib.bib64)\), socialness\(Diveicaet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib76)\)\) and PCA decomposition \(Table[19](https://arxiv.org/html/2605.23035#A34.T19), Appendix\)\. Word\-level norms produce lower overall encoding \(r=0\.194r\{=\}0\.194vs\.r=0\.285r\{=\}0\.285,p<0\.001p\{<\}0\.001\), with only 2/25 cells surviving FDR and convergenceρ=0\.31\\rho\{=\}0\.31\(p=0\.12p\{=\}0\.12\) vs\.ρ=0\.72\\rho\{=\}0\.72\(p<0\.001p\{<\}0\.001\) for SAEs\. The SAE advantage is both quantitative and qualitative: SAE features capture*contextual*semantics, span 16K–32K dimensions vs\. 4 norm dimensions, and uniquely enable activation patching\.

### 5\.3Behavioral Validation

#### Setup\.

For each reading\-time dataset, we fit nested linear mixed\-effects models \(GLMMs;lme4\) to log\-transformed reading times with by\-subject and by\-item random intercepts\. Following standard psycholinguistic practice\(Shainet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib41)\), thebaseline modelincludes per\-word fixed\-effect regressors for log frequency, word length \(characters\), position in sentence, preceding\-word log frequency and length \(spillover\), and unigram and 5\-gram surprisal estimated from the corresponding LM\. TheSAE\-augmented modeladditionally enters the layer\-ℓ\\ellsemantic\-feature vector at the peak layer \(L24 for GPT\-2 XL, L16 for Llama\-3\.1\-8B\), projected onto its top\-50 PCA components to control dimensionality\. Model fit is compared via likelihood\-ratio tests on−2​Δ​log⁡L\-2\\Delta\\log L, distributedχ2\\chi^\{2\}under the null\. Theword\-level\-norm contrastreplaces the SAE vector with concreteness, valence, arousal, and socialness norms\(Brysbaertet al\.,[2013](https://arxiv.org/html/2605.23035#bib.bib63); Warrineret al\.,[2013](https://arxiv.org/html/2605.23035#bib.bib64); Diveicaet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib76)\); therandom baselinereplaces SAE features with the same number of randomly sampled non\-semantic features \(10 seeds, averaged\)\.

#### Results\.

SAE semantic features predict human reading times beyond all baseline regressors in both datasets: Natural Stories\(Boyce and Levy,[2023](https://arxiv.org/html/2605.23035#bib.bib75)\)\(Δ​log⁡L=38\.4\\Delta\\log L\{=\}38\.4,χ2=76\.8\\chi^\{2\}\{=\}76\.8,p<0\.001p\{<\}0\.001; random baseline:Δ​log⁡L=2\.1\\Delta\\log L\{=\}2\.1,p=0\.34p\{=\}0\.34\) and Provo eye\-tracking\(Luke and Christianson,[2017](https://arxiv.org/html/2605.23035#bib.bib55)\)\(gaze duration:Δ​log⁡L=24\.7\\Delta\\log L\{=\}24\.7,p<0\.001p\{<\}0\.001; total reading time:Δ​log⁡L=31\.2\\Delta\\log L\{=\}31\.2,p<0\.001p\{<\}0\.001\)\. Critically, SAE features capture additional variance beyond word\-level norms \(Δ​log⁡L=12\.8\\Delta\\log L\{=\}12\.8,p<0\.001p\{<\}0\.001\), confirming the contextual advantage\.

#### Generalization\.

The semantic advantage generalizes to Chinese \(r=0\.276r\{=\}0\.276\) and French \(r=0\.291r\{=\}0\.291\) fMRI with Llama\-3\.1\-8B \(Table[18](https://arxiv.org/html/2605.23035#A33.T18), Appendix\); we frame this as a generalization check, not cross\-linguistic evidence \(§[8](https://arxiv.org/html/2605.23035#S8)\)\.

### 5\.4Exploratory: Semantic Prediction Errors

Under hierarchical predictive coding\(Friston,[2005](https://arxiv.org/html/2605.23035#bib.bib32); Clark,[2013](https://arxiv.org/html/2605.23035#bib.bib34); Caucheteuxet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib42)\), the brain predicts*semantic*content and responds to*unexpected*semantic information\. We test this with an exploratory analysis\.

For each semantic featurefif\_\{i\}at layerℓ\\ell, we compute thesemantic prediction error: the residual not predicted from the aligned feature at the previous analyzed layer:

ϵi\(ℓ\)=fi\(ℓ\)−\(αi​fπ​\(i\)\(ℓ−Δ\)\+βi\)\\epsilon\_\{i\}^\{\(\\ell\)\}=f\_\{i\}^\{\(\\ell\)\}\-\(\\alpha\_\{i\}f\_\{\\pi\(i\)\}^\{\(\\ell\-\\Delta\)\}\+\\beta\_\{i\}\)\(3\)whereπ​\(i\)\\pi\(i\)denotes the best\-matching feature at layerℓ−Δ\\ell\{\-\}\\Delta\(via decoder weight cosine similarity; Appendix[J](https://arxiv.org/html/2605.23035#A10)\) andΔ=4\\Delta\{=\}4is our layer sampling interval \(robustness acrossΔ∈\{4,8,12\}\\Delta\\in\\\{4,8,12\\\}in Appendix[K](https://arxiv.org/html/2605.23035#A11)\)\.

We align features across layers via decoder weight cosine similarity \(72% of L24 features match at L20 with sim\>0\.5\>0\.5; Appendix[J](https://arxiv.org/html/2605.23035#A10)\)\. At L24, combining raw semantics with prediction errors yieldsr=0\.303r\{=\}0\.303\(Δ​r=\+0\.018\\Delta r\{=\}\{\+\}0\.018,p<0\.01p\{<\}0\.01,d=0\.38d\{=\}0\.38, 95% CI\[0\.005,0\.031\]\[0\.005,0\.031\]; VIF=1\.21=1\.21; MLP replacement:Δ​r=\+0\.014\\Delta r\{=\}\{\+\}0\.014,p=0\.02p\{=\}0\.02\)\. This is preliminary evidence consistent with the brain encoding both*what*is present and*how unexpected*it is\.

SAE semantic features outperform PCA \(r=0\.285r\{=\}0\.285vs\.0\.2510\.251,p<0\.001p\{<\}0\.001\) and achieve comparable performance to QA\-Emb\(Benaraet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib49)\)\(r=0\.278r\{=\}0\.278\); geometric artifact accounts are ruled out by dimensionality\-matched controls \(Appendix[L](https://arxiv.org/html/2605.23035#A12),[AE](https://arxiv.org/html/2605.23035#A31)\)\.

## 6Implications for Theories of Language Processing

#### Cortical topography and the semantic atlas\.

Our primary contribution, the formally validated subcategory×\\timesregion mapping, bridges SAE\-based interpretability and cortical semantic organization\. The convergence \(ρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001; hypergeometricp=0\.007p\{=\}0\.007\) with three independent neuroscience programs supports the hub\-and\-spoke architecture\(Patterson and Lambon Ralph,[2016](https://arxiv.org/html/2605.23035#bib.bib73)\): modality\-specific “spokes” \(concreteness in posterior temporal, spatial in angular gyrus\) converge on abstract hubs \(social/mental in inferior frontal\)\.

#### Feature discovery vs\. predictive processing\.

Our results support the feature\-discovery account\(Antonello and Huth,[2023](https://arxiv.org/html/2605.23035#bib.bib36)\): brains and LLMs converge on semantic representations structured by cortical topography\. Our exploratory prediction\-error analysis \(Δ​r=\+0\.018\\Delta r\{=\}\+0\.018\) provides preliminary evidence consistent with a complementary predictive processing role, though our linear proxy is simplified relative to full hierarchical predictive coding\(Friston,[2010](https://arxiv.org/html/2605.23035#bib.bib33)\)\.

#### Competence, convergence, and the processing gradient\.

Brain–LLM convergence primarily captures shared formal semantic competence\(Mahowaldet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib43); Tuckuteet al\.,[2024a](https://arxiv.org/html/2605.23035#bib.bib45)\), with convergence strongest for concrete/perceptual semantics posteriorly and social/affective semantics frontally\. Behavioral validation across fMRI and reading times provides multi\-modal evidence for cognitively real distinctions\(Levy,[2008](https://arxiv.org/html/2605.23035#bib.bib23)\)\. Our SAE decomposition further reveals a depth\-wise processing gradient \(Appendix[P](https://arxiv.org/html/2605.23035#A16)\): lexical\-feature proportion peaks at L0 \(55%\) and falls monotonically; semantic\-feature proportion peaks at L24 \(41%\); prediction\-correlated\-feature proportion peaks at L44 \(37%\)\. The intermediate\-layer advantage thus reflects peak semantic richness before training\-objective specialization\(McCoyet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib44)\)\.

## 7Conclusion

This work bridges sparse autoencoders from mechanistic interpretability with neural encoding methodology, transforming the question of*why*intermediate LLM layers best predict brain activity into a question of*which*semantic features drive alignment*where*in the cortex\. Our central finding is a systematic cortical topography of brain–LLM convergence: specific semantic feature types \(concreteness, affect, social cognition, spatial relations, and event structure\) map onto specific cortical regions in a pattern that significantly matches predictions derived*a priori*from three independent neuroscience programs \(Spearmanρ=0\.72\\rho\{=\}0\.72,p<0\.001p\{<\}0\.001; hypergeometricp=0\.007p\{=\}0\.007\)\. The topography was more precisely captured by SAE features than by simpler word\-level semantic norms, demonstrating that the fine\-grained contextual representations SAEs extract carry neural information inaccessible to prior decomposition methods\.

Beyond this primary contribution, our SAE decomposition confirms at feature\-level resolution that semantic content dominates brain–LLM alignment, recovering 94% of peak encoding performance\. Behavioral validation across self\-paced reading and eye\-tracking datasets shows that SAE semantic features predict human reading times beyond lexical controls, and an exploratory prediction\-error analysis offers preliminary evidence that the brain may additionally encode unexpected semantic content\.

More broadly, this methodology uses mechanistic interpretability tools both to understand language models and to generate testable neuroscientific hypotheses; natural extensions include temporally resolved data \(ECoG, MEG\), typologically diverse languages, and newer SAE architectures\. All code, configurations, and analysis scripts to reproduce the results are publicly available at[https://github\.com/bettyguo/sae\-brain\-topography](https://github.com/bettyguo/sae-brain-topography)\.

## 8Limitations

\(1\) The five\-way taxonomy assumes mutual exclusivity; GPT\-4 labeling may introduce shared biases\(Huanget al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib11)\); soft categorization mitigates this \(Appendix[F](https://arxiv.org/html/2605.23035#A6)\)\. \(2\) Mean\-ablation is coarse\(Wuet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib14)\); we frame results as “evidence consistent with” a causal role\. \(3\) fMRI’s∼\{\\sim\}2s resolution conflates word\-level and compositional semantics\. \(4\) Three languages, two families, all head\-initial; a generalization check, not cross\-linguistic evidence\(Wilcoxet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib12)\)\. \(5\) Naturalistic stories may bias toward concrete/social content\. \(6\) Our prediction\-error proxy is simplified relative to full hierarchical predictive coding\(Friston,[2010](https://arxiv.org/html/2605.23035#bib.bib33)\)\. \(7\)N=8N\{=\}8limits power \(81% post\-hoc for the observed interaction\)\. \(8\) SAE feature splitting\(Makelovet al\.,[2025](https://arxiv.org/html/2605.23035#bib.bib20)\)and absorption\(Chaninet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib54)\)apply; robustness checks mitigate\. \(9\) Reading\-time and fMRI stimuli differ; future work should use identical stimuli\. \(10\) Alternative subcategorizations produce qualitatively similar patterns \(Appendix[V](https://arxiv.org/html/2605.23035#A22)\)\.

## Acknowledgments

We thank the anonymous reviewers and the area chair of CoNLL 2026 for their thoughtful and constructive feedback, which substantially improved this paper\. We are also grateful to the curators of the publicly available datasets used in this study: the naturalistic language fMRI dataset ofLeBelet al\.\([2023](https://arxiv.org/html/2605.23035#bib.bib48)\), the Natural Stories Corpus\(Boyce and Levy,[2023](https://arxiv.org/html/2605.23035#bib.bib75)\), and the Provo Corpus\(Luke and Christianson,[2017](https://arxiv.org/html/2605.23035#bib.bib55)\), without which this research would not have been possible\.

## References

- R\. Antonello and A\. Huth \(2023\)Predictive coding or just feature discovery? an alternative account of why language models fit brain data\.Neurobiology of Language,pp\. 1–16\.External Links:ISSN 2641\-4368,[Document](https://dx.doi.org/10.1162/nol%5Fa%5F00087)Cited by:[item 3](https://arxiv.org/html/2605.23035#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§1](https://arxiv.org/html/2605.23035#S1.p6.2),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2605.23035#S4.SS2.p1.12),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px2.p1.1)\.
- G\. Aparin, T\. Sadekova, A\. Rukhovich, A\. Yermekova, L\. Kushnareva, V\. Popov, K\. Kuznetsov, and I\. Piontkovskaya \(2026\)AudioSAE: towards understanding of audio\-processing models with sparse AutoEncoders\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3221–3254\.External Links:[Link](https://aclanthology.org/2026.eacl-long.149/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.149),ISBN 979\-8\-89176\-380\-7Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.
- V\. Benara, C\. Singh, J\. X\. Morris, R\. J\. Antonello, I\. Stoica, A\. Huth, and J\. Gao \(2024\)Crafting interpretable embeddings for language neuroscience by asking llms questions\.InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 \- 15, 2024,A\. Globersons, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. M\. Tomczak, and C\. Zhang \(Eds\.\),Cited by:[Appendix AE](https://arxiv.org/html/2605.23035#A31.p1.8),[§5\.4](https://arxiv.org/html/2605.23035#S5.SS4.p4.4)\.
- Y\. Benjamini and Y\. Hochberg \(1995\)Controlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society Series B: Statistical Methodology57\(1\),pp\. 289–300\.External Links:ISSN 1467\-9868,[Document](https://dx.doi.org/10.1111/j.2517-6161.1995.tb02031.x)Cited by:[§4\.4](https://arxiv.org/html/2605.23035#S4.SS4.p2.12)\.
- J\. R\. Binder, L\. L\. Conant, C\. J\. Humphries, L\. Fernandino, S\. B\. Simons, M\. Aguilar, and R\. H\. Desai \(2016\)Toward a brain\-based componential semantic representation\.Cognitive Neuropsychology33\(3–4\),pp\. 130–174\.External Links:ISSN 1464\-0627,[Document](https://dx.doi.org/10.1080/02643294.2016.1147426)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[§3\.6](https://arxiv.org/html/2605.23035#S3.SS6.p5.1)\.
- J\. R\. Binder, R\. H\. Desai, W\. W\. Graves, and L\. L\. Conant \(2009\)Where is the semantic system? a critical review and meta\-analysis of 120 functional neuroimaging studies\.Cerebral Cortex19\(12\),pp\. 2767–2796\.External Links:ISSN 1047\-3211,[Document](https://dx.doi.org/10.1093/cercor/bhp055)Cited by:[Table 16](https://arxiv.org/html/2605.23035#A29.T16),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[Figure 3](https://arxiv.org/html/2605.23035#S4.F3)\.
- V\. Boyce and R\. Levy \(2023\)A\-maze of natural stories: comprehension and surprisal in the maze task\.Glossa Psycholinguistics2\.External Links:[Document](https://dx.doi.org/10.5070/g6011190)Cited by:[Appendix B](https://arxiv.org/html/2605.23035#A2.SS0.SSS0.Px4.p1.1),[Appendix AJ](https://arxiv.org/html/2605.23035#A36.p1.1),[§3\.4](https://arxiv.org/html/2605.23035#S3.SS4.p2.1),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px2.p1.11),[Acknowledgments](https://arxiv.org/html/2605.23035#Sx1.p1.1)\.
- M\. Brysbaert, A\. B\. Warriner, and V\. Kuperman \(2013\)Concreteness ratings for 40 thousand generally known english word lemmas\.Behavior Research Methods46\(3\),pp\. 904–911\.External Links:ISSN 1554\-3528,[Document](https://dx.doi.org/10.3758/s13428-013-0403-5)Cited by:[§5\.2](https://arxiv.org/html/2605.23035#S5.SS2.p1.7),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px1.p1.3)\.
- C\. Caucheteux, A\. Gramfort, and J\. King \(2021\)Disentangling syntax and semantics in the brain with deep networks\.InProceedings of the 38th International Conference on Machine Learning, ICML 2021, 18\-24 July 2021, Virtual Event,M\. Meila and T\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.139,pp\. 1336–1348\.Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2605.23035#S4.SS3.p1.1)\.
- C\. Caucheteux, A\. Gramfort, and J\. King \(2023\)Evidence of a predictive coding hierarchy in the human brain listening to speech\.Nature Human Behaviour7\(3\),pp\. 430–441\.External Links:ISSN 2397\-3374,[Document](https://dx.doi.org/10.1038/s41562-022-01516-2)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§5\.4](https://arxiv.org/html/2605.23035#S5.SS4.p1.1)\.
- C\. Caucheteux and J\. King \(2022\)Brains and algorithms partially converge in natural language processing\.Communications Biology5\(1\)\.External Links:ISSN 2399\-3642,[Document](https://dx.doi.org/10.1038/s42003-022-03036-1)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p2.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- D\. Chanin, J\. Wilken\-Smith, T\. Dulka, H\. Bhatnagar, and J\. Bloom \(2024\)A is for absorption: studying feature splitting and absorption in sparse autoencoders\.arXiv preprintarXiv\.2409\.14507\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2409.14507)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2605.23035#S3.SS2.SSS0.Px1.p1.1),[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- R\. M\. Cichy, K\. Dwivedi, B\. Lahner, A\. Lascelles, P\. Iamshchinina, M\. Graumann, A\. Andonian, N\. A\. R\. Murty, K\. Kay, G\. Roig, and A\. Oliva \(2021\)The algonauts project 2021 challenge: how the human brain makes sense of a world in motion\.arXiv preprintarXiv\.2104\.13714\.Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Clark \(2013\)Whatever next? predictive brains, situated agents, and the future of cognitive science\.Behavioral and Brain Sciences36\(3\),pp\. 181–204\.External Links:ISSN 1469\-1825,[Document](https://dx.doi.org/10.1017/s0140525x12000477)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§5\.4](https://arxiv.org/html/2605.23035#S5.SS4.p1.1)\.
- I\. Covert, S\. M\. Lundberg, and S\. Lee \(2021\)Explaining by removing: A unified framework for model explanation\.J\. Mach\. Learn\. Res\.22,pp\. 209:1–209:90\.Cited by:[Appendix M](https://arxiv.org/html/2605.23035#A13.p1.1),[§3\.5](https://arxiv.org/html/2605.23035#S3.SS5.p2.2),[§4\.3](https://arxiv.org/html/2605.23035#S4.SS3.SSS0.Px1.p1.5)\.
- F\. Deniz, A\. O\. Nunez\-Elizalde, A\. G\. Huth, and J\. L\. Gallant \(2019\)The representation of semantic information across human cerebral cortex during listening versus reading is invariant to stimulus modality\.The Journal of Neuroscience39\(39\),pp\. 7722–7736\.External Links:ISSN 1529\-2401,[Document](https://dx.doi.org/10.1523/jneurosci.0675-19.2019)Cited by:[Table 16](https://arxiv.org/html/2605.23035#A29.T16),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[Figure 3](https://arxiv.org/html/2605.23035#S4.F3)\.
- V\. Diveica, P\. M\. Pexman, and R\. J\. Binney \(2023\)Quantifying social semantics: an inclusive definition of socialness and ratings for 8388 English words\.Behavior Research Methods55,pp\. 461–473\.External Links:[Document](https://dx.doi.org/10.3758/s13428-022-01810-x)Cited by:[§5\.2](https://arxiv.org/html/2605.23035#S5.SS2.p1.7),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px1.p1.3)\.
- A\. Doerig, T\. C\. Kietzmann, E\. Allen, Y\. Wu, T\. Naselaris, K\. Kay, and I\. Charest \(2025\)High\-level visual representations in the human brain are aligned with large language models\.Nature Machine Intelligence7\.External Links:[Document](https://dx.doi.org/10.1038/s42256-025-01072-0),ISSN 25225839Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah \(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2021/framework/index.html)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.
- E\. Fedorenko, A\. A\. Ivanova, and T\. I\. Regev \(2024\)The language network as a natural kind within the broader landscape of the human brain\.Nature Reviews Neuroscience25\(5\),pp\. 289–312\.External Links:ISSN 1471\-0048,[Document](https://dx.doi.org/10.1038/s41583-024-00802-4)Cited by:[Appendix B](https://arxiv.org/html/2605.23035#A2.SS0.SSS0.Px1.p1.5),[Appendix B](https://arxiv.org/html/2605.23035#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§3\.4](https://arxiv.org/html/2605.23035#S3.SS4.p1.5)\.
- K\. Friston \(2005\)A theory of cortical responses\.Philosophical Transactions of the Royal Society B: Biological Sciences360\(1456\),pp\. 815–836\.External Links:ISSN 1471\-2970,[Document](https://dx.doi.org/10.1098/rstb.2005.1622)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§5\.4](https://arxiv.org/html/2605.23035#S5.SS4.p1.1)\.
- K\. Friston \(2010\)The free\-energy principle: a unified brain theory?\.Nature Reviews Neuroscience11\(2\),pp\. 127–138\.External Links:ISSN 1471\-0048,[Document](https://dx.doi.org/10.1038/nrn2787)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px2.p1.1),[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- A\. Geiger, Z\. Wu, C\. Potts, T\. Icard, and N\. D\. Goodman \(2024\)Finding alignments between interpretable causal variables and distributed neural representations\.InCausal Learning and Reasoning, 1\-3 April 2024, Los Angeles, California, USA,F\. Locatello and V\. Didelez \(Eds\.\),Proceedings of Machine Learning Research, Vol\.236,pp\. 160–187\.Cited by:[§3\.7](https://arxiv.org/html/2605.23035#S3.SS7.p1.3)\.
- A\. Goldstein, Z\. Zada, E\. Buchnik, M\. Schain, A\. Price, B\. Aubrey, S\. A\. Nastase, A\. Feder, D\. Emanuel, A\. Cohen, A\. Jansen, H\. Gazula, G\. Choe, A\. Rao, C\. Kim, C\. Casto, L\. Fanda, W\. Doyle, D\. Friedman, P\. Dugan, L\. Melloni, R\. Reichart, S\. Devore, A\. Flinker, L\. Hasenfratz, O\. Levy, A\. Hassidim, M\. Brenner, Y\. Matias, K\. A\. Norman, O\. Devinsky, and U\. Hasson \(2022\)Shared computational principles for language processing in humans and deep language models\.Nature Neuroscience25\(3\),pp\. 369–380\.External Links:ISSN 1546\-1726,[Document](https://dx.doi.org/10.1038/s41593-022-01026-4)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Hale \(2001\)A probabilistic earley parser as a psycholinguistic model\.InLanguage Technologies 2001: The Second Meeting of the North American Chapter of the Association for Computational Linguistics, NAACL 2001, Pittsburgh, PA, USA, June 2\-7, 2001,Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1)\.
- U\. Hasson, S\. A\. Nastase, and A\. Goldstein \(2020\)Direct fit to nature: an evolutionary perspective on biological and artificial neural networks\.Neuron105\(3\),pp\. 416–434\.External Links:[Document](https://dx.doi.org/10.1016/j.neuron.2019.12.002)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1)\.
- M\. N\. Hebart, O\. Contier, L\. Teichmann, A\. H\. Rockter, C\. Y\. Zheng, A\. Kidder, A\. Corriveau, M\. Vaziri\-Pashkam, and C\. I\. Baker \(2023\)THINGS\-data, a multimodal collection of large\-scale datasets for investigating object representations in human brain and behavior\.eLife12\.External Links:ISSN 2050\-084X,[Document](https://dx.doi.org/10.7554/elife.82580)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Heilbron, K\. Armeni, J\. Schoffelen, P\. Hagoort, and F\. P\. de Lange \(2022\)A hierarchy of linguistic predictions during natural language comprehension\.Proceedings of the National Academy of Sciences119\(32\),pp\. e2201968119\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2201968119)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1)\.
- E\. A\. Hosseini, M\. Schrimpf, Y\. Zhang, S\. Bowman, N\. Zaslavsky, and E\. Fedorenko \(2024\)Artificial neural network language models predict human brain responses to language even after a developmentally realistic amount of training\.Neurobiology of Language5\(1\),pp\. 43–63\.External Links:[Document](https://dx.doi.org/10.1162/nol%5Fa%5F00137)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Huang, Z\. Wu, C\. Potts, M\. Geva, and A\. Geiger \(2024\)RAVEL: evaluating interpretability methods on disentangling language model representations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 8669–8687\.External Links:[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.470)Cited by:[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- R\. Huben, H\. Cunningham, L\. R\. Smith, A\. Ewart, and L\. Sharkey \(2024\)Sparse autoencoders find highly interpretable features in language models\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2605.23035#S3.SS2.SSS0.Px1.p1.1)\.
- A\. G\. Huth, W\. A\. de Heer, T\. L\. Griffiths, F\. E\. Theunissen, and J\. L\. Gallant \(2016\)Natural speech reveals the semantic maps that tile human cerebral cortex\.Nature532\(7600\),pp\. 453–458\.External Links:[Document](https://dx.doi.org/10.1038/NATURE17637)Cited by:[Appendix W](https://arxiv.org/html/2605.23035#A23.p3.4),[Table 16](https://arxiv.org/html/2605.23035#A29.T16),[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[Figure 3](https://arxiv.org/html/2605.23035#S4.F3),[§4\.5](https://arxiv.org/html/2605.23035#S4.SS5.p6.6)\.
- S\. Jain and A\. Huth \(2018\)Incorporating context into language encoding models for fmri\.InAdvances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3\-8, 2018, Montréal, Canada,S\. Bengio, H\. M\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),pp\. 6629–6638\.Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Kauf, G\. Tuckute, R\. Levy, J\. Andreas, and E\. Fedorenko \(2024\)Lexical\-semantic content, not syntactic structure, is the main contributor to ANN\-brain similarity of fMRI responses in the language network\.Neurobiology of Language5\(1\),pp\. 7–42\.External Links:[Document](https://dx.doi.org/10.1162/nol%5Fa%5F00116)Cited by:[item 3](https://arxiv.org/html/2605.23035#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2605.23035#S1.p6.2),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2605.23035#S4.SS2.p1.12)\.
- T\. C\. Kietzmann, C\. J\. Spoerer, L\. K\. A\. Sörensen, R\. M\. Cichy, O\. Hauk, and N\. Kriegeskorte \(2019\)Recurrence is required to capture the representational dynamics of the human visual system\.Proceedings of the National Academy of Sciences116\(43\),pp\. 21854–21863\.External Links:ISSN 1091\-6490,[Document](https://dx.doi.org/10.1073/pnas.1905544116)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Kornblith, M\. Norouzi, H\. Lee, and G\. E\. Hinton \(2019\)Similarity of neural network representations revisited\.InProceedings of the 36th International Conference on Machine Learning, ICML 2019, 9\-15 June 2019, Long Beach, California, USA,K\. Chaudhuri and R\. Salakhutdinov \(Eds\.\),Proceedings of Machine Learning Research, Vol\.97,pp\. 3519–3529\.Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1)\.
- N\. Kriegeskorte \(2008\)Representational similarity analysis – connecting the branches of systems neuroscience\.Frontiers in Systems Neuroscience\. Frontiers Media SA\.External Links:ISSN 1662\-5137,[Document](https://dx.doi.org/10.3389/neuro.06.004.2008)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Kumar, T\. R\. Sumers, T\. Yamakoshi, A\. Goldstein, U\. Hasson, K\. A\. Norman, T\. L\. Griffiths, R\. D\. Hawkins, and S\. A\. Nastase \(2024\)Shared functional specialization in transformer\-based language models and the human brain\.Nature Communications15,pp\. 5523\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-49173-5)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- A\. LeBel, L\. Wagner, S\. Jain, A\. Adhikari\-Desai, B\. Gupta, A\. Morgenthal, J\. Tang, L\. Xu, and A\. G\. Huth \(2023\)A natural language fMRI dataset for voxelwise encoding models\.Scientific Data10,pp\. 555\.External Links:[Document](https://dx.doi.org/10.1038/s41597-023-02437-z)Cited by:[Appendix B](https://arxiv.org/html/2605.23035#A2.SS0.SSS0.Px1.p1.5),[Appendix AJ](https://arxiv.org/html/2605.23035#A36.p1.1),[§3\.4](https://arxiv.org/html/2605.23035#S3.SS4.p1.5),[Acknowledgments](https://arxiv.org/html/2605.23035#Sx1.p1.1)\.
- R\. Levy \(2008\)Expectation\-based syntactic comprehension\.Cognition106\(3\),pp\. 1126–1177\.External Links:[Document](https://dx.doi.org/10.1016/j.cognition.2007.05.006)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px3.p1.1)\.
- Y\. Li, E\. J\. Michaud, D\. D\. Baek, J\. Engels, X\. Sun, and M\. Tegmark \(2025\)The geometry of concepts: sparse autoencoder feature structure\.Entropy27\(4\),pp\. 344\.External Links:[Document](https://dx.doi.org/10.3390/E27040344)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.
- J\. Liu, X\. Shi, T\. D\. Nguyen, H\. Zhang, T\. Zhang, W\. Sun, Y\. Li, A\. V\. Vasilakos, G\. Iacca, A\. A\. Khan, A\. Kumar, J\. W\. Cho, A\. Mian, L\. Xie, E\. Cambria, and L\. Wang \(2025\)Neural brain: A neuroscience\-inspired framework for embodied agents\.arXiv preprintarXiv\.2505\.07634\.External Links:[Link](https://doi.org/10.48550/arXiv.2505.07634),[Document](https://dx.doi.org/10.48550/ARXIV.2505.07634)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- S\. G\. Luke and K\. Christianson \(2017\)The provo corpus: a large eye\-tracking corpus with predictability norms\.Behavior Research Methods50\(2\),pp\. 826–833\.External Links:ISSN 1554\-3528,[Document](https://dx.doi.org/10.3758/s13428-017-0908-4)Cited by:[Appendix B](https://arxiv.org/html/2605.23035#A2.SS0.SSS0.Px4.p1.1),[Appendix AJ](https://arxiv.org/html/2605.23035#A36.p1.1),[§3\.4](https://arxiv.org/html/2605.23035#S3.SS4.p2.1),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px2.p1.11),[Acknowledgments](https://arxiv.org/html/2605.23035#Sx1.p1.1)\.
- K\. Mahowald, A\. A\. Ivanova, I\. A\. Blank, N\. Kanwisher, J\. B\. Tenenbaum, and E\. Fedorenko \(2023\)Dissociating language and thought in large language models: a cognitive perspective\.arXiv preprintarXiv\.2301\.06627\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2301.06627)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px6.p1.1),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px3.p1.1)\.
- A\. Makelov, G\. Lange, and N\. Nanda \(2025\)Towards principled evaluations of sparse autoencoders for interpretability and control\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1),[§3\.2](https://arxiv.org/html/2605.23035#S3.SS2.SSS0.Px1.p1.1),[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- R\. Mao, Q\. Liu, X\. Li, E\. Cambria, and A\. Hussain \(2025\)Bridging minds and machines: toward an integration of AI and cognitive science\.arXiv preprintarXiv\.2508\.20674\.External Links:[Link](https://doi.org/10.48550/arXiv.2508.20674),[Document](https://dx.doi.org/10.48550/ARXIV.2508.20674)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p2.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- M\. W\. Mathis, A\. Perez Rotondo, E\. F\. Chang, A\. S\. Tolias, and A\. Mathis \(2024\)Decoding the brain: from neural representations to mechanistic models\.InCell,Vol\.187,pp\. 5814–5832\.External Links:ISSN 0092\-8674,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cell.2024.08.051),[Link](https://www.sciencedirect.com/science/article/pii/S0092867424009802)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- R\. T\. McCoy, S\. Yao, D\. Friedman, M\. D\. Hardy, and T\. L\. Griffiths \(2024\)Embers of autoregression show how large language models are shaped by the problem they are trained to solve\.Proceedings of the National Academy of Sciences121\(41\)\.External Links:ISSN 1091\-6490,[Document](https://dx.doi.org/10.1073/pnas.2322420121)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px6.p1.1),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px3.p1.1)\.
- K\. McRae, G\. S\. Cree, M\. S\. Seidenberg, and C\. Mcnorgan \(2005\)Semantic feature production norms for a large set of living and nonliving things\.Behavior Research Methods37\(4\),pp\. 547–559\.External Links:ISSN 1554\-3528,[Document](https://dx.doi.org/10.3758/bf03192726)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[§3\.6](https://arxiv.org/html/2605.23035#S3.SS6.p5.1)\.
- J\. A\. Michaelov, M\. D\. Bardolph, C\. K\. Van Petten, B\. K\. Bergen, and S\. Coulson \(2024\)Strong prediction: language model surprisal explains multiple n400 effects\.Neurobiology of Language5\(1\),pp\. 107–135\.External Links:ISSN 2641\-4368,[Document](https://dx.doi.org/10.1162/nol%5Fa%5F00105)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- K\. Misra and K\. Mahowald \(2024\)Language models learn rare phenomena from less rare phenomena: the case of the missing aanns\.arXiv preprintarXiv\.2403\.19827\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2403.19827)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px6.p1.1)\.
- T\. M\. Mitchell, S\. V\. Shinkareva, A\. Carlson, K\. Chang, V\. L\. Malave, R\. A\. Mason, and M\. A\. Just \(2008\)Predicting human brain activity associated with the meanings of nouns\.Science320\(5880\),pp\. 1191–1195\.External Links:ISSN 1095\-9203,[Document](https://dx.doi.org/10.1126/science.1152876)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1)\.
- H\. Nili, C\. Wingfield, A\. Walther, L\. Su, W\. D\. Marslen\-Wilson, and N\. Kriegeskorte \(2014\)A toolbox for representational similarity analysis\.PLoS Comput\. Biol\.10\(4\)\.External Links:[Document](https://dx.doi.org/10.1371/JOURNAL.PCBI.1003553)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- S\. R\. Oota, M\. Gupta, and M\. Toneva \(2023\)Joint processing of linguistic properties in brains and language models\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Pasquiou, Y\. Lakretz, J\. T\. Hale, B\. Thirion, and C\. Pallier \(2022\)Neural language models are not born equal to fit brain data, but training helps\.InInternational Conference on Machine Learning, ICML 2022, 17\-23 July 2022, Baltimore, Maryland, USA,K\. Chaudhuri, S\. Jegelka, L\. Song, C\. Szepesvári, G\. Niu, and S\. Sabato \(Eds\.\),Proceedings of Machine Learning Research, Vol\.162,pp\. 17499–17516\.Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Patterson and M\. A\. Lambon Ralph \(2016\)Chapter 61 \- the hub\-and\-spoke hypothesis of semantic memory\.InNeurobiology of Language,G\. Hickok and S\. L\. Small \(Eds\.\),San Diego,pp\. 765–775\.External Links:ISBN 978\-0\-12\-407794\-2,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/B978-0-12-407794-2.00061-4),[Link](https://www.sciencedirect.com/science/article/pii/B9780124077942000614)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px1.p1.4)\.
- F\. Pereira, B\. Lou, B\. Pritchett, S\. Ritter, S\. J\. Gershman, N\. Kanwisher, M\. Botvinick, and E\. Fedorenko \(2018\)Toward a universal decoder of linguistic meaning from brain activation\.Nature Communications9\(1\)\.External Links:ISSN 2041\-1723,[Document](https://dx.doi.org/10.1038/s41467-018-03068-4)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px6.p1.1)\.
- R\. P\. N\. Rao and D\. H\. Ballard \(1999\)Predictive coding in the visual cortex: a functional interpretation of some extra\-classical receptive\-field effects\.Nature Neuroscience2\(1\),pp\. 79–87\.External Links:ISSN 1546\-1726,[Document](https://dx.doi.org/10.1038/4580)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1)\.
- M\. Schrimpf, I\. A\. Blank, G\. Tuckute, C\. Kauf, E\. A\. Hosseini, N\. Kanwisher, J\. B\. Tenenbaum, and E\. Fedorenko \(2021\)The neural architecture of language: integrative modeling converges on predictive processing\.Proceedings of the National Academy of Sciences118\(45\),pp\. e2105646118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2105646118)Cited by:[Appendix W](https://arxiv.org/html/2605.23035#A23.p3.4),[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1),[§4\.5](https://arxiv.org/html/2605.23035#S4.SS5.p6.6)\.
- C\. Shain, C\. Meister, T\. Pimentel, R\. Cotterell, and R\. Levy \(2024\)Large\-scale evidence for logarithmic effects of word predictability on reading time\.Proceedings of the National Academy of Sciences121\(10\),pp\. e2307876121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2307876121)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px1.p1.3)\.
- M\. Tänzer, S\. Ruder, and M\. Rei \(2022\)Memorisation versus generalisation in pre\-trained language models\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Dublin, Ireland,pp\. 7564–7578\.External Links:[Link](https://aclanthology.org/2022.acl-long.521/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.521)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px6.p1.1)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, A\. Tamkin, E\. Durmus, T\. Hume, F\. Mosconi, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. Henighan \(2024\)Scaling monosemanticity: extracting interpretable features from Claude 3 Sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p3.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.
- M\. Toneva and L\. Wehbe \(2019\)Interpreting and improving natural\-language processing \(in machines\) with natural language\-processing \(in the brain\)\.InAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8\-14, 2019, Vancouver, BC, Canada,H\. M\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d’Alché\-Buc, E\. B\. Fox, and R\. Garnett \(Eds\.\),pp\. 14928–14938\.Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p2.1),[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Tuckute, N\. Kanwisher, and E\. Fedorenko \(2024a\)Language in brains, minds, and machines\.Annual Review of Neuroscience47\(1\),pp\. 277–301\.External Links:ISSN 1545\-4126,[Document](https://dx.doi.org/10.1146/annurev-neuro-120623-101142)Cited by:[§6](https://arxiv.org/html/2605.23035#S6.SS0.SSS0.Px3.p1.1)\.
- G\. Tuckute, A\. Sathe, S\. Srikant, M\. Taliaferro, M\. Wang, M\. Schrimpf, K\. Kay, and E\. Fedorenko \(2024b\)Driving and suppressing the human language network using large language models\.Nature Human Behaviour8\(3\),pp\. 544–561\.External Links:[Document](https://dx.doi.org/10.1038/s41562-023-01783-7)Cited by:[§1](https://arxiv.org/html/2605.23035#S1.p1.1),[§1](https://arxiv.org/html/2605.23035#S1.p2.1)\.
- S\. Vafaei, R\. Fukuma, T\. Yanagisawa, H\. Yang, S\. Oshino, N\. Tani, H\. M\. Khoo, H\. Sugano, Y\. Iimura, H\. Suzuki, M\. Nakajima, K\. Tamura, and H\. Kishima \(2026\)Brain\-aligning of semantic vectors improves neural decoding of visual stimuli\.Communications Biology\.External Links:[Document](https://dx.doi.org/10.1038/s42003-025-09482-x),ISSN 23993642Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px2.p1.1)\.
- A\. B\. Warriner, V\. Kuperman, and M\. Brysbaert \(2013\)Norms of valence, arousal, and dominance for 13,915 english lemmas\.Behavior Research Methods45\(4\),pp\. 1191–1207\.External Links:ISSN 1554\-3528,[Document](https://dx.doi.org/10.3758/s13428-012-0314-x)Cited by:[§5\.2](https://arxiv.org/html/2605.23035#S5.SS2.p1.7),[§5\.3](https://arxiv.org/html/2605.23035#S5.SS3.SSS0.Px1.p1.3)\.
- N\. Wasserman, M\. Cosarinsky, Y\. Golbari, A\. Oliva, A\. Torralba, T\. R\. Shaham, and M\. Irani \(2025\)BrainExplore: large\-scale discovery of interpretable visual representations in the human brain\.arXiv preprintarXiv\.2512\.08560\.External Links:[Link](https://arxiv.org/abs/2512.08560)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.
- E\. G\. Wilcox, T\. Pimentel, C\. Meister, R\. Cotterell, and R\. P\. Levy \(2023\)Testing the predictions of surprisal theory in 11 languages\.Trans\. Assoc\. Comput\. Linguistics11,pp\. 1451–1470\.External Links:[Document](https://dx.doi.org/10.1162/TACL%5FA%5F00612)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px5.p1.1),[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- Z\. Wu, A\. Geiger, T\. Icard, C\. Potts, and N\. D\. Goodman \(2023\)Interpretability at scale: identifying causal mechanisms in alpaca\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Cited by:[§3\.7](https://arxiv.org/html/2605.23035#S3.SS7.p1.3),[§5\.1](https://arxiv.org/html/2605.23035#S5.SS1.SSS0.Px1.p1.3),[§8](https://arxiv.org/html/2605.23035#S8.p1.2)\.
- A\. Zeng and J\. L\. Gallant \(2026\)Disentangling superpositions: interpretable brain encoding model with sparse concept atoms\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=3aNvX9TQTo)Cited by:[§2](https://arxiv.org/html/2605.23035#S2.SS0.SSS0.Px4.p1.1)\.

## Appendix ASAE Training Hyperparameters

Table 3:SAE training hyperparameters\.Computing: 4×\\timesNVIDIA A100 80GB \(∼\{\\sim\}240 GPU\-hours total\)\. Encoding models on CPU \(∼\{\\sim\}2 hours/layer/subject\)\. Seed variance: SD<0\.003<0\.003for all metrics\.

## Appendix BDataset Details and Preprocessing

#### Primary fMRI dataset\.

We use the publicly released UTS naturalistic\-language fMRI dataset\(LeBelet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib48)\), comprising 8 native\-English participants who listened to 27 spoken narratives \(The Moth Radio Hour, podcast monologues, audiobook excerpts\) for a total of∼\{\\sim\}6 h of stimulus exposure per subject\. BOLD data were collected at TR=\{=\}2\.0 s\. We follow the dataset’s distributed preprocessing pipeline \(motion correction, slice\-timing correction, registration to MNI152 space\) and apply additional motion\-spike censoring \(framewise displacement\>\{\>\}0\.5 mm\) and high\-pass filtering at1/1281/128Hz to remove low\-frequency drift\. Because the UTS dataset does not include per\-participant functional language\-localizer scans, we identify language\-network voxels using the group\-level*anatomical*parcellation of the canonical language network described byFedorenkoet al\.\([2024](https://arxiv.org/html/2605.23035#bib.bib29)\), intersected with each subject’s native\-space gray\-matter mask to obtain∼\{\\sim\}5,000 bilateral voxels per subject\.

#### Word\-to\-TR alignment\.

Story stimuli have published word\-onset timestamps\. For each layerℓ\\elland feature subsetSS, we extract activations at the position of each word token, convolve the resulting per\-word feature time series with the canonical double\-gamma HRF \(peak=\{=\}5 s, undershoot peak=\{=\}15 s\), and downsample to the fMRI TR by integrating within each TR window\. This yields TR\-aligned feature matrices \(T×\|S\|T\\times\|S\|per subject\) that are then entered into the ridge regression of §[3\.5](https://arxiv.org/html/2605.23035#S3.SS5)\.

#### Generalization datasets \(Chinese and French\)\.

For the cross\-language generalization check, we use publicly available open\-access naturalistic listening fMRI datasets in Mandarin Chinese and French drawn from the cross\-linguistic “Little Prince” tradition of recording native speakers listening to translations of the same narrative\. The Chinese subset comprises 15 native Mandarin speakers and 2 chapters; the French subset comprises 12 native French speakers and 1 chapter\. Both datasets were collected by their original authors under institutional ethics protocols described in the corresponding data papers, with consent procedures and compensation documented therein\. We did not collect any new human data and did not require new IRB approval for this study, which is a secondary analysis of publicly available datasets\. We apply the same anatomical language\-network parcellation \(Fedorenkoet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib29), registered to each subject’s native space\) and the same word\-to\-TR alignment pipeline as for the primary dataset; story\-specific word timestamps are taken from the metadata distributed with each dataset\.

#### Behavioral datasets\.

Self\-paced reading times come from the Natural Stories Corpus\(Boyce and Levy,[2023](https://arxiv.org/html/2605.23035#bib.bib75)\)\(181 participants reading 10 stories one word at a time\); eye\-tracking comes from the Provo Corpus\(Luke and Christianson,[2017](https://arxiv.org/html/2605.23035#bib.bib55)\)\(84 participants reading 55 short passages\)\. All behavioral analyses use only words with valid reading\-time measurements after the corpora’ standard outlier\-trimming procedures\.

## Appendix CFeature Categorization Protocol

GPT\-4 prompt\(verbatim\):*“You are a neurolinguistic annotator\. Below are the top\-20 tokens that maximally activate a specific feature in a neural network, each with 5 tokens of surrounding context\. Based on these activation patterns, \(a\) provide a 1–2 sentence description, and \(b\) classify into exactly one category: SEMANTIC, SYNTACTIC, LEXICAL, PREDICTION, or OTHER\. Respond in JSON\.”*

Category definitions:*Semantic*: activates for tokens sharing a semantic property regardless of position\.*Syntactic*: based on grammatical role\.*Lexical*: based on surface form\.*Prediction*: correlatesr\>0\.5r\>0\.5with tuned lens entropy\.*Other*: no consistent pattern\.

## Appendix DAnnotator Details

Two graduate students \(computational linguistics,≥\\geq2 years experience, compensated $25/hour,∼\{\\sim\}40 hours each, neither an author\) independently categorized 500 features\. Overall disagreement: 14%\. Main confusion: semantic↔\\leftrightarrowprediction \(8% of disagreements\)\.

## Appendix EQualitative Feature Examples

Feature IDCategoryTop\-5 Activating Tokens \(with context\)Human JudgmentL24\-3847Sem\. \(concrete\)“thedogran across the”, “a largehousestood”, “bright redcarparked”, “old woodentablein”, “the talltreeswayed”Clear: physical objectsL24\-9102Sem\. \(affect\)“felt deeplysadabout”, “overwhelmingjoyfilled”, “a sense ofdread”, “wasfuriouswhen”, “terrifiedof the dark”Clear: emotional statesL24\-6538Sem\. \(social\)“shebelievedthat he”, “theywantedto help”, “hethoughtabout her”, “shedecidedto tell”, “theyagreedto meet”Clear: mental statesL24\-1205Syntactic“The manwhocame”, “booksthatwere old”, “ideawhichled to”, “placewherethey met”, “timewhenit rained”Clear: relative pronounsL24\-4421Lexical“tionof the process”, “mentand development”, “nessof the dark”, “ityin modern life”, “mentin the final”Clear: nominal suffixesL24\-8817Prediction“Thepresidentof the”, “Accordingto the new”, “Inconclusion, we find”, “As aresultof”, “Forexample, consider”Clear: high\-predictabilityL24\-7731Sem\. \(ambig\.\)“theoldbridge over”, “ancientRome was known”, “ranquickly to the”, “darkclouds gathered”, “thecoldwinter night”Ambiguous: mixedL24\-2194Pred\./Sem\.“sheneverexpected to”, “itsuddenlybecame clear”, “herarelyspoke of”, “surprisingly, the door”, “unexpectedlyfound”BorderlineTable 4:Representative SAE features at GPT\-2 XL L24 including clear and ambiguous cases\.
## Appendix FSoft \(Probabilistic\) Categorization

GPT\-4 was prompted to provide confidence \(0–1\) per category\. Variance partitioning under soft assignment: semanticΔ​R2=0\.045\\Delta R^\{2\}=0\.045\(vs\. 0\.048 hard\), syntactic = 0\.012 \(vs\. 0\.011\), prediction = 0\.007 \(vs\. 0\.006\)\. Rankings unchanged\.

## Appendix GActivation\-Variance Matching

At L24, semantic features \(41% of features\) account for 52% of total L2 activation variance\. Variance\-matched random ablation selects features until 52% of variance is captured \(∼\{\\sim\}55% of features needed\), yieldingΔ​r2=−0\.06\\Delta r^\{2\}=\-0\.06vs\.−0\.14\-0\.14for semantic\.

## Appendix HOther/Uninterpretable Features

Table 5:Other/uninterpretable features contribute minimally\.
## Appendix ISemantic Prediction\-Error: Full Details

For each of the∼\{\\sim\}6,700 semantic features at L24, we fit a linear regression from the aligned feature at L20\. The improvement from combining \(\+0\.018\+0\.018,p<0\.01p\{<\}0\.01, bootstrapped 95% CI\[0\.005,0\.031\]\[0\.005,0\.031\]\) is robust across subjects \(7/8 showΔ​r\>0\\Delta r\>0\)\. The correlation between raw semantic features and semantic prediction errors isr=0\.42r\{=\}0\.42\(VIF = 1\.21\), indicating they carry partially independent information with low multicollinearity\.

## Appendix JCross\-Layer Feature Alignment

We compute cosine similarity between all pairs of decoder weight columns \(d=1600d\{=\}1600for GPT\-2 XL\)\. 72% of L24 features have a unique best match with sim\>0\.5\>0\.5; 41% exceed0\.70\.7\. Mean similarity for best\-matched pairs: 0\.63 \(SD = 0\.18\)\. Re\-running using aligned features yieldsΔ​r=\+0\.021\\Delta r\{=\}\{\+\}0\.021; restricting to sim\>0\.7\>0\.7yieldsΔ​r=\+0\.024\\Delta r\{=\}\{\+\}0\.024\(p<0\.01p\{<\}0\.01\)\.

## Appendix KLayer Offset Robustness

Table 6:Prediction\-error analysis across layer offsets\.
## Appendix LGeometric Artifact Analysis

Effective dimensionality \(participation ratio\): L24 PR = 847 vs\. L4: 312, L44: 198\. Random projection to 847 dimensions:r=0\.201r\{=\}0\.201vs\. semantic features:r=0\.285r\{=\}0\.285\(p<0\.001p\{<\}0\.001\)\.

## Appendix MShapley Value Decomposition

FollowingCovertet al\.\([2021](https://arxiv.org/html/2605.23035#bib.bib13)\), approximate Shapley values \(1,000 orderings; mean absolute deviation between successive 100\-ordering blocks<0\.001<0\.001\)\. Results: Semantic = 0\.054 \(65%\), Syntactic = 0\.015 \(18%\), Prediction = 0\.009 \(11%\), Lexical = 0\.002 \(2%\), Other = 0\.003 \(4%\)\.

## Appendix NSAE Reconstruction Error

Encoding from reconstruction error \(𝐱−𝐱^\\mathbf\{x\}\-\\hat\{\\mathbf\{x\}\}\): meanr=0\.031r\{=\}0\.031\(not significant after FDR\)\.

## Appendix OPer\-Subject Results

Table 7:Per\-subject results at GPT\-2 XL L24\.
## Appendix PFull Feature Distribution

Table 8:Complete feature\-type distribution across GPT\-2 XL layers\.
## Appendix QPrediction Threshold Sensitivity

Table 9:Sensitivity to prediction\-feature threshold at L24\.
## Appendix RSAE Hyperparameter Robustness

Table 10:Robustness across dictionary sizes and sparsity targets\.
## Appendix SRegion\-Specific Encoding

Table 11:Region\-specific encoding at L24\. Semantic dominance holds across all regions; syntactic contributions are relatively larger in inferior frontal cortex\.
## Appendix TCross\-Linguistic Patching

Table 12:Activation patching across datasets \(Llama\-3\.1\-8B\)\.
## Appendix UCross\-Linguistic Subcategory Proportions

Table 13:Semantic subcategory proportions across languages \(at peak layer, Llama\-3\.1\-8B\)\. Distributions are broadly similar; Chinese shows slightly elevated social/mental features\.
## Appendix VAlternative Subcategorizations

To test robustness of the cortical topography to subcategory definitions, we re\-ran the analysis using two alternative subcategorization schemes:

Binder et al\. \(2016\) 65\-dimensional:We collapsed Binder’s 65 experiential attributes into 7 groups following their factor analysis \(sensory, motor, spatial, temporal, affective, social, cognitive\) and mapped SAE features accordingly\. The subcategory×\\timesregion interaction remains significant \(F​\(24,168\)=2\.91F\(24,168\)\{=\}2\.91,p<0\.001p\{<\}0\.001\) with convergenceρ=0\.61\\rho\{=\}0\.61\(p=0\.002p\{=\}0\.002\)\. The pattern is qualitatively similar but with more distributed activation across regions for the finer\-grained categories\.

McRae et al\. \(2005\) feature norms:Using McRae’s semantic feature production norms to categorize SAE features \(visual, functional, encyclopedic, taxonomic\), we find a significant interaction \(F​\(12,84\)=2\.58F\(12,84\)\{=\}2\.58,p=0\.006p\{=\}0\.006\) with convergenceρ=0\.54\\rho\{=\}0\.54\(p=0\.008p\{=\}0\.008\)\. The coarser McRae categories capture less topographic specificity, consistent with the prediction that finer\-grained categories produce sharper topographic maps\.

## Appendix WPower Analysis

Post\-hoc power analysis for the subcategory×\\timesregion interaction: withN=8N\{=\}8, the observed effect sizef=0\.42f\{=\}0\.42\(computed fromηp2=0\.15\\eta^\{2\}\_\{p\}\{=\}0\.15\), andα=0\.05\\alpha\{=\}0\.05, the estimated power is 0\.81 \(G\*Power, repeated\-measures ANOVA, 5×\\times5 within\-factors, correlation among repeated measures=0\.3=0\.3\)\. This indicates adequate power to detect the observed interaction\.

For individual cells: givenN=8N\{=\}8and the observed cell\-level effect sizes \(r=0\.07r\{=\}0\.07–0\.140\.14\), the power to detect individual cell effects atα=0\.002\\alpha\{=\}0\.002\(FDR\-corrected threshold\) ranges from 0\.15 \(forr=0\.07r\{=\}0\.07\) to 0\.62 \(forr=0\.14r\{=\}0\.14\)\. The pattern of 7/25 FDR\-surviving cells is thus consistent with the power profile: strongly predicted associations with large effect sizes survive, while weaker associations do not\. The joint probability of observing≥\\geq6 of 7 FDR\-surviving cells in predicted locations given these power constraints isp<0\.01p\{<\}0\.01\(simulation\-based\)\.

The observed effect sizes \(r=0\.07r\{=\}0\.07–0\.140\.14\) are comparable to voxelwise encodingrr\-values in comparable studies:Huthet al\.\([2016](https://arxiv.org/html/2605.23035#bib.bib2)\)report median voxelwiser=0\.12r\{=\}0\.12for their semantic atlas;Schrimpfet al\.\([2021](https://arxiv.org/html/2605.23035#bib.bib25)\)report brain\-score values in a similar range for naturalistic stimuli\.

## Appendix XEffective Degrees of Freedom

The∼\{\\sim\}6,700 semantic SAE features exhibit substantial collinearity\. The effective dimensionality, estimated via the participation ratio of the feature covariance matrix \(PR=\(∑iλi\)2/∑iλi2\\text\{PR\}=\(\\sum\_\{i\}\\lambda\_\{i\}\)^\{2\}/\\sum\_\{i\}\\lambda\_\{i\}^\{2\}whereλi\\lambda\_\{i\}are eigenvalues\), is∼\{\\sim\}1,200 at L24\. This effective feature count is well within the regularized regime for ridge regression with∼\{\\sim\}6 hours of fMRI data per subject \(∼\{\\sim\}10,800 TRs\)\. The condition number of the feature matrix with ridge regularization \(λ=103\\lambda\{=\}10^\{3\}, typical selected value\) isκ=47\\kappa\{=\}47, indicating moderate but not problematic collinearity\.

## Appendix YTuned Lens Analysis

Table 14:Tuned lens metrics and brain prediction across GPT\-2 XL layers\. Peak brain prediction corresponds to∼\{\\sim\}45% top\-1 accuracy\.
## Appendix ZDetailed Subcategory×\\timesRegion Heatmap

Post\. Temp\.Ant\. Temp\.Inf\. Front\.Ang\. Gyr\.dmPFCConcreteEventAffectSocialSpatial\.141\*\*\.108\.082\.131\*\*\.079\.118\.105\.098\.109\.085\.095\.128\*\*\.104\.088\.121\*\.087\.122\*\*\.119\*\*\.083\.098\.124\*\*\.089\.071\.135\*\*\.076rrvalue:\.06\.08\.10\.12\.14Figure 4:Subcategory×\\timesregion encoding at L24\. \*\* survives FDR correction \(q<0\.05q\{<\}0\.05, Benjamini\-Hochberg\); \* nominally significant \(p<0\.05p\{<\}0\.05, permutation\)\.
## Appendix AAFeature Categorization Confusion Matrix

GPT\-4↓\\downarrow/ Human→\\rightarrowSemSynLexPredOthSemantic824329Syntactic386218Lexical538417Prediction4217914Other65101762Per\-cat\.κ\\kappa\.78\.83\.80\.74\.58Table 15:Confusion matrix \(%\) for feature categorization at GPT\-2 XL L24 \(n=500n\{=\}500, stratified 100/category\)\.
## Appendix ABGPT\-4 Labeling Bias Audit

Using one LLM \(GPT\-4\) to label features of another LLM \(GPT\-2 XL / Llama\-3\.1\-8B\) risks introducing shared inductive biases that could propagate into downstream conclusions\. We complement the human\-validation analysis \(Appendix[D](https://arxiv.org/html/2605.23035#A4), Table[15](https://arxiv.org/html/2605.23035#A27.T15)\) and soft probabilistic categorization \(Appendix[F](https://arxiv.org/html/2605.23035#A6)\) with three additional bias\-audit checks at GPT\-2 XL L24\.

#### Audit 1: Cross\-LLM relabeling\.

A stratified sample of 200 features \(40 per category\) was independently relabeled using a different LLM family \(Claude 3 Opus\) with the same prompt template \(Appendix[C](https://arxiv.org/html/2605.23035#A3)\)\. Per\-category agreement with GPT\-4 was: semanticκ=0\.87\\kappa\{=\}0\.87; syntacticκ=0\.91\\kappa\{=\}0\.91; lexicalκ=0\.93\\kappa\{=\}0\.93; predictionκ=0\.74\\kappa\{=\}0\.74; otherκ=0\.68\\kappa\{=\}0\.68\(overallκ=0\.84\\kappa\{=\}0\.84\)\. Agreement is substantial\-to\-strong for the four substantive categories and only weak for the residual “other” bucket, consistent with the human\-vs\-GPT\-4 pattern \(Table[15](https://arxiv.org/html/2605.23035#A27.T15)\)\. The strong cross\-LLM agreement on*semantic*labels in particular indicates that the labels of features driving our headline results are not idiosyncratic to GPT\-4\.

#### Audit 2: Label\-perturbation sensitivity\.

We simulate residual labeling error by randomly flipping a fractionppof within\-category labels \(uniformly to one of the other four categories\) and re\-running the variance partitioning of §[4\.2](https://arxiv.org/html/2605.23035#S4.SS2)\. Withp∈\{0\.05,0\.10,0\.20\}p\\in\\\{0\.05,0\.10,0\.20\\\}, semanticΔ​Runique2\\Delta R^\{2\}\_\{\\text\{unique\}\}falls modestly from0\.0480\.048\(clean\) to0\.0440\.044,0\.0410\.041, and0\.0350\.035respectively, while the rank ordering Semantic\>\>Syntactic\>\>Prediction\>\>Lexical is preserved at all noise levels\. The convergence\-test Spearmanρ\\rhodegrades from0\.720\.72to0\.690\.69,0\.660\.66,0\.580\.58, remaining significant \(p<0\.05p\{<\}0\.05via permutation\) up top=0\.20p\{=\}0\.20\. The qualitative conclusions are thus robust to relabeling error well beyond the∼\\sim14% human\-vs\-GPT\-4 disagreement rate\.

#### Audit 3: Confidence\-thresholded subset\.

GPT\-4 was prompted in Pass 1 to emit a per\-category confidence \(0–1; Appendix[F](https://arxiv.org/html/2605.23035#A6)\)\. Restricting analysis to features whose maximum\-category confidence exceeds0\.80\.8\(the top 78% of features at L24\) yields qualitatively identical results: semantic encodingr=0\.288r\{=\}0\.288\(vs\.0\.2850\.285on the full set\), variance partitioning Semantic 54% / Syntactic 12% / Prediction 7% \(vs\. 52% / 12% / 7%\), and a priori convergenceρ=0\.71\\rho\{=\}0\.71,p<0\.001p\{<\}0\.001\(vs\.0\.720\.72\)\. Headline results survive the high\-confidence restriction\.

#### Limitations of the audit\.

These checks address*labeling reliability*\(whether the categorization is reproducible\) but cannot rule out*shared representational bias*\(whether both GPT\-2 XL and the labeling LLM share a systematic blind spot relative to the brain\)\. Such shared bias would affect any LLM\-mediated annotation pipeline; we view feature\-production\-norm\-based subcategorization \(Appendix[V](https://arxiv.org/html/2605.23035#A22)\) as the strongest available LLM\-free robustness check\.

## Appendix ACA Priori Subcategory Predictions

Table 16:A priori predicted subcategory×\\timesregion mappings derived fromBinderet al\.\([2009](https://arxiv.org/html/2605.23035#bib.bib46)\),Huthet al\.\([2016](https://arxiv.org/html/2605.23035#bib.bib2)\), andDenizet al\.\([2019](https://arxiv.org/html/2605.23035#bib.bib47)\)\. Each subcategory’s predicted primary regions are those consistently associated with the corresponding semantic dimension across≥\\geq2 of the three programs\.
## Appendix ADEvidence Sources per A Priori Cell

To make the derivation in §[3\.6](https://arxiv.org/html/2605.23035#S3.SS6)fully transparent, Table[17](https://arxiv.org/html/2605.23035#A30.T17)records, for each of the ten predicted\-primary cells of the5×55\\times 5subcategory×\\timesregion matrix, the specific text passage or figure in each source program supporting that prediction\. Cells were classified as “predicted primary” when≥\\geq2 of the three programs explicitly mapped the subcategory to the region\. The remaining 15 cells of the matrix are unmarked priors\. This procedure produces a binary5×55\\times 5prediction matrix*before*any examination of SAE encoding results, and is the matrix entered into the formal convergence tests of §[4\.5](https://arxiv.org/html/2605.23035#S4.SS5)\.

Table 17:Source\-paper evidence supporting each of the ten “predicted primary” cells in the a priori subcategory×\\timesregion matrix\. Each row corresponds to one predicted\-primary cell of Table[16](https://arxiv.org/html/2605.23035#A29.T16)\. “—” indicates the source program did not explicitly include the corresponding subcategory\-region mapping among its primary findings; we required≥\\geq2 of the three programs to support a cell\.
## Appendix AEAlternative Decompositions and Geometric Controls

SAE semantic features outperform PCA \(r=0\.285r\{=\}0\.285vs\.0\.2510\.251,p<0\.001p\{<\}0\.001,d=1\.12d\{=\}1\.12\) and achieve comparable performance to QA\-Emb\(Benaraet al\.,[2024](https://arxiv.org/html/2605.23035#bib.bib49)\)\(r=0\.278r\{=\}0\.278,p=0\.21p\{=\}0\.21\) with finer\-grained decomposition \(16K features vs\.∼\{\\sim\}100 questions\)\. The SAE advantage is*qualitative*: only SAEs enable the subcategory×\\timesregion analysis at 16K\-feature granularity and the activation patching analyses\.

Effective dimensionality \(participation ratio\) is higher at intermediate layers \(PR = 847 at L24 vs\. 312 at L4, 198 at L44\)\. However, random projection to 847 dimensions achieves onlyr=0\.201r\{=\}0\.201, and dimensionality\-matched random feature subsets achiever=0\.218r\{=\}0\.218\(p<0\.001p\{<\}0\.001vs\. semantic0\.2850\.285\)\. Geometric properties alone do not explain brain alignment\.

## Appendix AFActivation Patching Figure

Sem\.Syn\.Pred\.Rnd\-NRnd\-V\*\*\*\*n\.s\.Δ​r2\\Delta r^\{2\}0\-0\.05\-0\.10\-0\.15Figure 5:Activation patching at L24 with 95% CIs\. Results are evidence consistent with a causal role for semantic features\.
## Appendix AGCross\-Linguistic Generalization

Table 18:Generalization check across datasets and languages\. Rnd\.†: count\-matched random baseline\.
## Appendix AHSAE vs\. Word\-Level Norms: Full Comparison

Table 19:Comparison of cortical topography across decomposition methods\.ρ\\rho: Spearman correlation with a priori predicted matrix\. \*p<0\.05p\{<\}0\.05; \*\*\*p<0\.001p\{<\}0\.001\.
## Appendix AISemantic Prediction\-Error: Summary Table

Table 20:Exploratory semantic prediction\-error analysis at L24\.
## Appendix AJReproducibility Statement

Upon acceptance, we release: \(1\) all SAE weights; \(2\) feature categorization labels \(automated \+ human\); \(3\) GPT\-4 prompts \(Appendix[C](https://arxiv.org/html/2605.23035#A3)\); \(4\) encoding model code; \(5\) activation patching code; \(6\) semantic prediction\-error analysis code; \(7\) cross\-layer alignment code; \(8\) formal convergence test code; \(9\) behavioral validation analysis code; \(10\) a priori subcategory derivation documentation\. All fMRI datasets are publicly available\(LeBelet al\.,[2023](https://arxiv.org/html/2605.23035#bib.bib48)\)\. Reading\-time datasets are publicly available\(Boyce and Levy,[2023](https://arxiv.org/html/2605.23035#bib.bib75); Luke and Christianson,[2017](https://arxiv.org/html/2605.23035#bib.bib55)\)\.

Similar Articles

Brain-LLM Alignment Tracks Training Data, Not Typology

arXiv cs.CL

This paper investigates brain-LLM alignment across English, Chinese, and French using fMRI data and multiple LLMs, finding that training-language dominance and typological distance, not an inherent English advantage, drive alignment patterns.

Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction

arXiv cs.CL

Researchers propose Brain-CLIPLM, a two-stage EEG-to-text decoding framework using contrastive learning for semantic anchor extraction and a retrieval-grounded LLM with Chain-of-Thought reasoning, achieving 67.55% top-5 sentence retrieval accuracy and suggesting EEG-to-text decoding should focus on recovering compressed semantic content rather than full sentence reconstruction.