Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

arXiv cs.LG Papers

Summary

This paper applies TopK Sparse Autoencoders to three EEG foundation models (SleepFM, REVE, LaBraM) to extract interpretable feature dictionaries and introduces a framework for concept steering, revealing representational failures and clinical entanglements.

arXiv:2605.13930v1 Announce Type: new Abstract: EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $\alpha$-band restoration.
Original Article
View Cached Full Text

Cached at: 05/15/26, 06:25 AM

# Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
Source: [https://arxiv.org/html/2605.13930](https://arxiv.org/html/2605.13930)
William Lehn\-Schiøler1,2,3&Magnus Ruud Kjær2&Rahul Thapa4, 5&Magnus Guldberg Pedersen1,3&Anton Storgaard Mosquera1,3&Nick Williams6&Radu Gătej1&Tue Lehn\-Schiøler1&Sándor Beniczky7,8&Sadasivan Puthusserypady2&James Zou4,5&Lars Kai Hansen3

###### Abstract

EEG foundation models achieve state\-of\-the\-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust\. We apply TopK Sparse Autoencoders \(SAEs\) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings\. By grounding these features in a clinical taxonomy \(abnormality, age, sex, and medication\), we benchmark monosemanticity and entanglement across architectures\. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures\. Via concept steering, we introduce a "target vs\. off\-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non\-encoded\. This framework exposes critical representational failures: "wrecking\-ball" interventions that collapse global model performance, and clinical entanglements, such as age–pathology confounding, where it is impossible to suppress one concept without corrupting the other\. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow\-wave suppression andα\\alpha\-band restoration\.

1BrainCapture, Kongens Lyngby, Denmark

2DTU Health Tech, Technical University of Denmark, Kongens Lyngby, Denmark

3DTU Compute, Technical University of Denmark, Kongens Lyngby, Denmark

4Department of Biomedical Data Science, Stanford University, Stanford, CA, USA

5Department of Computer Science, Stanford University, Stanford, CA, USA

6Seer Medical, Melbourne, Australia

7Filadelfia Epilepsy Hospital, Dianalund, Denmark

8University Hospital of Copenhagen, Copenhagen, Denmark

## 1Introduction

EEG foundation models pre\-trained on large corpora achieve strong performance across a range of tasks: sleep staging, pathology detection, and brain\-computer interface decoding\[[28](https://arxiv.org/html/2605.13930#bib.bib2),[8](https://arxiv.org/html/2605.13930#bib.bib4),[16](https://arxiv.org/html/2605.13930#bib.bib5),[19](https://arxiv.org/html/2605.13930#bib.bib6),[20](https://arxiv.org/html/2605.13930#bib.bib3)\]\. Yet the internal computations that produce these predictions remain opaque\. Unlike image models, where attention maps already give a coarse visual story, EEG models operate on multichannel time\-series whose clinically relevant events: spindles, K\-complexes,δ\\deltawaves, epileptiform discharges, are transient, channel\-distributed, and defined by decades of human expert knowledge\[[5](https://arxiv.org/html/2605.13930#bib.bib28)\]\. Understanding*what*an EEG foundation model has learned, and*where*it has learned it, is both a scientific question and a clinical\-trust requirement\.

Mechanistic interpretability\[[10](https://arxiv.org/html/2605.13930#bib.bib13),[11](https://arxiv.org/html/2605.13930#bib.bib12)\]addresses this by reverse\-engineering the functional roles of individual directions inside a network\. The central obstacle issuperposition\[[11](https://arxiv.org/html/2605.13930#bib.bib12)\]: a transformer layer withdddimensions can represent up toN≫dN\\gg dfeatures by co\-activating them sparsely, making individual neurons uninterpretable\. Sparse Autoencoders \(SAEs\)\[[6](https://arxiv.org/html/2605.13930#bib.bib16),[27](https://arxiv.org/html/2605.13930#bib.bib15),[21](https://arxiv.org/html/2605.13930#bib.bib17)\]have emerged as the dominant tool for recovering these features in language models\. A small but growing literature ports the toolkit to medical imaging\[[25](https://arxiv.org/html/2605.13930#bib.bib11),[24](https://arxiv.org/html/2605.13930#bib.bib10)\], protein language models\[[26](https://arxiv.org/html/2605.13930#bib.bib9)\], and transformers trained on neural signals\[[13](https://arxiv.org/html/2605.13930#bib.bib7),[17](https://arxiv.org/html/2605.13930#bib.bib8)\]\. Despite this progress, EEG foundation models remain unaddressed; no prior work has bridged the full path from a frozen EEG encoder to sparse feature dictionaries, clinical concept attribution, and spectrum\-level mechanistic explanation across architectures\.

#### Contributions\.

We make the following contributions:

1. 1\.A Cross\-Architecture Interpretability Pipeline:We introduce a unified framework for interpreting EEG transformers that bridges the gap between high\-dimensional embeddings and clinical physiology\. The pipeline integrates layer\-wise TopK SAEs, concept attribution via Testing with Concept Activation Vectors \(TCAV\), and concept steering, utilizing a single hyperparameter\-selection procedure that remains robust across three distinct architectures: SleepFM, REVE, and LaBraM\. The interactive framework is available at
2. 2\.The Clinical Semanticity Taxonomy:We propose a taxonomy to audit encoder representations by partitioning clinical features into three regimes: Separable \(monosemantic\), Entangled \(polysemantic co\-activations\), and Dead \(semantically uninformative/inactive\)\. This serves as a tool to identify latent clinical biases where labels like age or medication co\-activate with pathology\.
3. 3\.Selectivity via Concept Steering:We formalize a probe\-based selectivity metric; calculating the area between target and off\-target "steering curves" to evaluate the fidelity of model interventions\. This allows us to distinguish between selective concept removal and "wrecking\-ball" interventions, where feature suppression inadvertently collapses the entire embedding space\.
4. 4\.Mechanistic Spectral Explanations:By re\-decoding steered activations through a spectral decoder, we transform abstract latent embedding into human interpretable space\. This provides mechanistic evidence for model attributions, allowing experts to verify that a pathology prediction is driven by physiologically relevant features\.

## 2Background

### 2\.1EEG Foundation Models

We study three architecturally distinct EEG transformer models spanning two different self\-supervised pretraining objectives \(multi\-modal contrastive learning and masked\-token prediction\); these are summarized in Table[1](https://arxiv.org/html/2605.13930#S2.T1)\. All three encoders are subsequently finetuned end\-to\-end on the same binary normal/abnormal classification target which is later used as a clinical concept\. Performance metrics can be found in Table[2](https://arxiv.org/html/2605.13930#S4.T2)\.

### 2\.2TopK Sparse Autoencoders

A TopK SAE\[[23](https://arxiv.org/html/2605.13930#bib.bib18)\]trained on activations𝐚∈ℝd\\mathbf\{a\}\\in\\mathbb\{R\}^\{d\}learns

𝐳=TopK​\(𝐖enc​𝐚−𝝁ℓ𝝈ℓ,k\),𝐚^=𝐖dec​𝐳\+𝐛dec,\\mathbf\{z\}=\\mathrm\{TopK\}\\left\(\\mathbf\{W\}\_\{\\text\{enc\}\}\\,\\frac\{\\mathbf\{a\}\-\\boldsymbol\{\\mu\_\{\\ell\}\}\}\{\\boldsymbol\{\\sigma\_\{\\ell\}\}\},k\\right\),\\quad\\hat\{\\mathbf\{a\}\}=\\mathbf\{W\}\_\{\\text\{dec\}\}\\,\\mathbf\{z\}\+\\mathbf\{b\}\_\{\\text\{dec\}\},\(1\)where𝐳∈ℝN\\mathbf\{z\}\\in\\mathbb\{R\}^\{N\}and𝐖dec∈ℝd×N\\mathbf\{W\}\_\{\\text\{dec\}\}\\in\\mathbb\{R\}^\{d\\times N\}has unit\-norm columns \(decoder directions𝐰i∈ℝd\\mathbf\{w\}\_\{i\}\\in\\mathbb\{R\}^\{d\}\);N≜d⋅EN\\triangleq d\\cdot EandEEis the expansion rate\. The conventional SAE formulation uses anL1L\_\{1\}sparsity penalty\[[7](https://arxiv.org/html/2605.13930#bib.bib14)\], which encourages but does not guarantee sparsity\. We instead use the TopK hard constraint\[[27](https://arxiv.org/html/2605.13930#bib.bib15)\]: pre\-activations are computed for allNNfeatures \(after standardizing the input usingμℓ\\mu\_\{\\ell\}andσℓ\\sigma\_\{\\ell\}, the per\-dimension mean and standard deviation computed over the training\-split activations at layerℓ\\ell\), then all but thekklargest are set to zero\. Exactlykkfeatures fire per token by construction, making sparsity directly controllable\.

Table 1:Encoder overview\.Three architecturally distinct EEG transformers spanning different SSL objectives: multi\-modal contrastive learning \(SleepFM\), masked\-token reconstruction \(REVE\), and masked spectrum prediction \(LaBraM\) are inspected\. SleepFM is pretrained on PSG data \(EEG, ECG, EMG, and respiratory\) with a multi\-modal contrastive objective, while REVE and LaBraM are pretrained on EEG alone\. All three are subsequently finetuned end\-to\-end on the same binary normal/abnormal target\.EncoderArchitectureddLayersToken lengthSSL objectiveModality≈\\approxHoursSleepFM\[[28](https://arxiv.org/html/2605.13930#bib.bib2)\]SetTransformer12831 secMMC \(InfoNCE\)\[[29](https://arxiv.org/html/2605.13930#bib.bib31)\]PSG585,000REVE\[[8](https://arxiv.org/html/2605.13930#bib.bib4)\]Transformer512221 secMAE\[[15](https://arxiv.org/html/2605.13930#bib.bib32)\]EEG60,000LaBraM\[[16](https://arxiv.org/html/2605.13930#bib.bib5)\]BERT\-style200121 secMSP \(VQ / BEiT\)\[[2](https://arxiv.org/html/2605.13930#bib.bib33)\]EEG2,500

### 2\.3TCAV and Concept Attribution

Concept Activation Vectors\[[18](https://arxiv.org/html/2605.13930#bib.bib23)\]define a conceptCCvia a linear classifier separatingXCX\_\{C\}fromX¬CX\_\{\\neg C\}in representation space, yielding a unit\-norm direction𝐯C∈ℝd\\mathbf\{v\}\_\{C\}\\in\\mathbb\{R\}^\{d\}\. We additionally train a linear probe𝐰\\mathbf\{w\}on SAE activations𝐳∈ℝN\\mathbf\{z\}\\in\\mathbb\{R\}^\{N\}, assigning each example a sensitivitySC​\(𝐱\)=𝐰⊤​𝐳​\(𝐱\)S\_\{C\}\(\\mathbf\{x\}\)=\\mathbf\{w\}^\{\\top\}\\mathbf\{z\}\(\\mathbf\{x\}\)\. The TCAV score is the fraction of concept examples with positive sensitivity:

TCAVC=\|\{𝐱∈XC:SC​\(𝐱\)\>0\}\|\|XC\|,\\mathrm\{TCAV\}\_\{C\}=\\frac\{\|\\\{\\mathbf\{x\}\\in X\_\{C\}:S\_\{C\}\(\\mathbf\{x\}\)\>0\\\}\|\}\{\|X\_\{C\}\|\},\(2\)
with significance assessed againstNrand=50N\_\{\\mathrm\{rand\}\}=50random\-label null CAVs\.

## 3Methodology

The pipeline \(Figure[1](https://arxiv.org/html/2605.13930#S3.F1)\) runs in four stages on a frozen encoder\. Stages I–III use established interpretability components \(linear probing\[[1](https://arxiv.org/html/2605.13930#bib.bib35),[3](https://arxiv.org/html/2605.13930#bib.bib36)\], dictionary learning\[[7](https://arxiv.org/html/2605.13930#bib.bib14)\], concept attribution\[[22](https://arxiv.org/html/2605.13930#bib.bib25)\]\); the novel contribution is Stage IV \(concept steering\), which uses the dictionary and the spectral decoder to produce mechanistic, spectrum\-level explanations of any TCAV attribution\.

![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/pipeline.png)Figure 1:Pipeline overview\.Starting from a frozen EEG foundation model: \(Stage I\) A shallow MLP*spectral decoder*translates token embeddings back into a human interpretable space\. \(Stage II\) For each transformer layer, a TopK SAE recovers a sparse, over\-complete feature dictionary from normalized encoder activations\. \(Stage III\) SAE features are mapped to known clinical concepts using TCAV\. \(Stage IV\)*Concept steering*substitutes the top\-nnconcept\-enriched feature activations with the target\-group centroid, decoding the intervention through the SAE and spectral decoder to produce a mechanistically grounded, spectrum\-level explanation\.### 3\.1Dataset and Preprocessing

We use a clinical 27\-lead EEG dataset collected using BrainCapture’s BC\-1 comprising3,0363\{,\}036subjects\[[12](https://arxiv.org/html/2605.13930#bib.bib1)\]\. The cohort is57\.8%57\.8\\%male and includes both pediatric and adult populations\. Recordings averaged33\.833\.8minutes, with30\.2%30\.2\\%\(917917sessions\) identified as abnormal; of these,76\.4%76\.4\\%showed epileptiform discharges\. EEG signals are preprocessed using an evolved version of\[[14](https://arxiv.org/html/2605.13930#bib.bib24)\]\. The pipeline applies band\-pass \(0\.5\-70Hz\) and notch \(50Hz\) filtering before resampling to the encoder’s native frequency:128128Hz for SleepFM and200200Hz for REVE and LaBraM\. Recordings are segmented into non\-overlapping 60\-second windows\. Subjects are partitioned into 80% train/10% validation/10% test sets via scikit\-learn’sGroupShuffleSplitwith random seed 42\.

### 3\.2Stage I: Spectral Decoder

The spectral decoder \(SD\) is a shallow MLP that maps a frozen encoder token embedding𝐭∈ℝd\\mathbf\{t\}\\in\\mathbb\{R\}^\{d\}to amplitude and phase for each frequency bin:

A^ν,\(cos⁡ϕ^ν,sin⁡ϕ^ν\)=SD​\(𝐭\),ν=1,…,F\.\\hat\{A\}\_\{\\nu\},\\;\(\\cos\\hat\{\\phi\}\_\{\\nu\},\\,\\sin\\hat\{\\phi\}\_\{\\nu\}\)=\\mathrm\{SD\}\(\\mathbf\{t\}\),\\quad\\nu=1,\\ldots,F\.\(3\)F=64F=64bins,0–6464Hz, for all encoders \(consistent with the7070Hz low\-pass filter applied in preprocessing; Section[3\.1](https://arxiv.org/html/2605.13930#S3.SS1)\)\. Phase is parameterised as a unit\-vector pair to eliminate the2​π2\\pidiscontinuity\. SleepFM tokens are channel\-averaged; REVE and LaBraM tokens are per\-channel\-per\-second\.

The spectral decoder serves two roles: \(i\) it provides a physiological vocabulary for interpreting any direction𝐰∈ℝd\\mathbf\{w\}\\in\\mathbb\{R\}^\{d\}viaSD​\(𝐰\)\\mathrm\{SD\}\(\\mathbf\{w\}\), and \(ii\) it is the final stage of concept steering \(Section[3\.5](https://arxiv.org/html/2605.13930#S3.SS5)\)\.

### 3\.3Stage II: Layer\-wise SAE Training

For each encoder layerℓ\\ellwe collect all available training\-split token activations, normalise per\-dimension, and train a TopK SAE withE⋅dE\\cdot dfeatures, whereddis the encoder embedding dimension andEEis the*expansion ratio*controlling dictionary size relative to the encoder\. We scale the sparsity budget with the dictionary size,k=k0⋅Ek=k\_\{0\}\\cdot Ewithk0=8k\_\{0\}=8, so that the per\-token active fractionk/N=k0/dk/N=k\_\{0\}/dis constant per encoder across the expansion sweep\. This couples the two hyperparameters into a singleEE\-axis and ensures that comparisons acrossEEare at matched relative sparsity rather than at matched absolutekk\. The reconstruction objective is MSE on normalised activations; dead neurons \(never firing in 500 steps\) are periodically resampled following\[[27](https://arxiv.org/html/2605.13930#bib.bib15)\]\.

### 3\.4Stage III: Concept Attribution via TCAV

For each conceptC∈\{abnormality, age group, sex, medication \(psychiatric\), medication \(ASM\)\}C\\in\\\{\\textit\{abnormality, age group, sex, medication \(psychiatric\), medication \(ASM\)\}\\\}, we obtain𝐯C\\mathbf\{v\}\_\{C\}viaKfold=10K\_\{\\mathrm\{fold\}\}=10\-foldL2L\_\{2\}\-regularised logistic regression on dense layer\-ℓ\\ellactivations𝐚\\mathbf\{a\}, then fit probe𝐰\\mathbf\{w\}on𝐳\\mathbf\{z\}and computeTCAVC\\mathrm\{TCAV\}\_\{C\}per Eq\.[2](https://arxiv.org/html/2605.13930#S2.E2)\. For feature enrichment, we apply a one\-sidedzz\-test with Benjamini–Hochberg correction \(q<0\.05q<0\.05,NNsimultaneous hypotheses\) to identify concept\-enriched SAE features, which form the ranking pool for Stage IV\.

### 3\.5Stage IV: Concept Steering

Concept steering provides interventional evidence for the representational fidelity of the SAE dictionary, evaluated against a model\-agnostic*selectivity*criterion\. Given a target clinical conceptCC\(e\.g\., age\) and a secondary off\-target concept \(e\.g\., abnormality\), the protocol includes:

1. 1\.Rank by CAV alignment\.The concept direction is mapped onto the sparse dictionary by projecting it onto each SAE decoder direction𝐰i∈ℝd\\mathbf\{w\}\_\{i\}\\in\\mathbb\{R\}^\{d\}, giving the alignment\-based ranking rankC​\(i\)=\|𝐯C⋅𝐰i\|\\mathrm\{rank\}\_\{C\}\(i\)=\|\\mathbf\{v\}\_\{C\}\\cdot\\mathbf\{w\}\_\{i\}\|\(4\)
2. 2\.Compute the target\-concept centroid\.We define the centroid𝐜target∈ℝN\\mathbf\{c\}\_\{\\text\{target\}\}\\in\\mathbb\{R\}^\{N\}as the empirical mean of the SAE feature activations across the target datasetXtargetX\_\{\\text\{target\}\}\(e\.g\., the mean of the “normal” pool when steering abnormal→\\tonormal\): 𝐜target=1\|Xtarget\|​∑𝐱∈Xtarget𝐳​\(𝐱\)\\mathbf\{c\}\_\{\\text\{target\}\}=\\frac\{1\}\{\|X\_\{\\text\{target\}\}\|\}\\sum\_\{\\mathbf\{x\}\\in X\_\{\\text\{target\}\}\}\\mathbf\{z\}\(\\mathbf\{x\}\)\(5\)
3. 3\.Clamping features\.For a given intervention fractionf∈\[0,1\]f\\in\[0,1\]and a source sample with SAE activations𝐳\\mathbf\{z\}, we clamp the activations of the topn≜⌊f⋅N⌋n\\triangleq\\lfloor f\\cdot N\\rfloorTCAV\-aligned features to their corresponding scalar values in the target centroid𝐜target\\mathbf\{c\}\_\{\\text\{target\}\}\. This produces the intervened sparse activation𝐳∗​\(f\)\\mathbf\{z\}^\{\*\}\(f\), which is then decoded back into embedding space: 𝐚^∗​\(f\)=𝐖dec​𝐳∗​\(f\)\+𝐛dec\\hat\{\\mathbf\{a\}\}^\{\*\}\(f\)=\\mathbf\{W\}\_\{\\text\{dec\}\}\\,\\mathbf\{z\}^\{\*\}\(f\)\+\\mathbf\{b\}\_\{\\text\{dec\}\}\(6\)

### 3\.6Evaluating "steerability"

1. 4\.Frozen probe evaluation\.We score the decoded embeddings𝐚^∗​\(f\)\\hat\{\\mathbf\{a\}\}^\{\*\}\(f\)using two linear probes \(target and off\-target\)\. Crucially, these probes are fit once on clean \(f=0f=0\) decoded embeddings and held frozen across the sweep; they act solely as measurement instruments to read out the result of the clamp, completely isolated from the substitution loop\.
2. 5\.Excess Selectivity Scalar\.We quantify the quality of the intervention by the integrated area between the off\-target and target performance curves: Δ​\(rank\)=∫01\(AUROCoff\-target​\(f\)−AUROCtarget​\(f\)\)​df\\Delta\(\\text\{rank\}\)=\\int\_\{0\}^\{1\}\\left\(\\mathrm\{AUROC\}\_\{\\text\{off\-target\}\}\(f\)\-\\mathrm\{AUROC\}\_\{\\text\{target\}\}\(f\)\\right\)\\mathrm\{d\}f\(7\) To account for random feature suppression which may degrade one probe faster than the other, we report theexcess selectivityΔ~\\tilde\{\\Delta\}as the difference between the TCAV\-ranked area and the expected area under uniformly random feature permutationsπ\\pi: Δ~=Δ​\(TCAV\)−𝔼π​\[Δ​\(π\)\]\\tilde\{\\Delta\}=\\Delta\(\\text\{TCAV\}\)\-\\mathbb\{E\}\_\{\\pi\}\[\\Delta\(\\pi\)\]\(8\)

A value ofΔ~\>0\\tilde\{\\Delta\}\>0indicates a selective push that successfully erases the target concept while sparing off\-target axes\. Conversely, a value ofΔ~≈0\\tilde\{\\Delta\}\\approx 0diagnoses a "wrecking\-ball" regime: the intervention is no more selective than random noise, implying that the target concept is so entangled with the global embedding structure that it cannot be independently manipulated\. This protocol operationalizes the selectivity criteria used in closed\-form concept erasure\[[4](https://arxiv.org/html/2605.13930#bib.bib21)\]and amnesic probing\[[9](https://arxiv.org/html/2605.13930#bib.bib19)\]for the SAE latent space\. Finally, decoding the intervened embedding𝐚^∗​\(f\)\\hat\{\\mathbf\{a\}\}^\{\*\}\(f\)through the spectral decoder \(Section[3\.2](https://arxiv.org/html/2605.13930#S3.SS2)\) yields domain\-readable spectral signatures of the intervention \(Figure[6](https://arxiv.org/html/2605.13930#S4.F6)\)\.

### 3\.7Choosing the operating point per model

We identify an optimal layerℓ∗\\ell^\{\*\}and SAE expansion ratioE∗E^\{\*\}for Stages II–IV by sweeping all transformer layers andE∈\{1,2,4,8,16,32,64\}E\\in\\\{1,2,4,8,16,32,64\\\}\. To match the active feature fraction across scales, sparsity scales linearly ask=8​Ek=8E\. We select the configuration that maximizes monosemanticity while penalizing dead features:

\(ℓ∗,E∗\)=arg⁡maxℓ,E⁡\(separableℓ,E−deadℓ,E\),\(\\ell^\{\*\},E^\{\*\}\)=\\arg\\max\_\{\\ell,E\}\\left\(\\mathrm\{separable\}\_\{\\ell,E\}\-\\mathrm\{dead\}\_\{\\ell,E\}\\right\),\(9\)
constrained to layers where target concepts achieve model\-level TCAV significance \(p<0\.05p<0\.05\)\. This taxonomy\-driven metric efficiently isolates the model’s most transparent regime \(Figure[3](https://arxiv.org/html/2605.13930#S4.F3)\)\. While this identifies the optimal single coordinate, our computational capacity allows us to fixE=E∗E=E^\{\*\}and explore various layers in subsequent interventional analyses to map the encoder’s full representational trajectory\.

## 4Results

### 4\.1Binary finetune performance and SAE faithfulness

All three encoders separate the binary normal/abnormal target after end\-to\-end finetuning\. We finetune each encoder under 5\-fold subject\-disjoint cross\-validation and report the native classifier head’s test performance, that is, the predictor we would actually deploy\. As a faithfulness check on the SAE, we replace the encoder’s layer\-ℓ∗\\ell^\{\*\}activations with the TopK\-SAE reconstruction and re\-evaluate the same head on the same test windows: classification AUROC is preserved within0\.0170\.017on all three encoders \(Table[2](https://arxiv.org/html/2605.13930#S4.T2)\), confirming that the SAE captures the task\-relevant structure of the representation rather than discarding it\.

Table 2:Layer\-averaged SAE\-faithfulness summary \(5\-fold CV\)\.All values are mean±\\pmstd across 5 subject\-disjoint folds\. Baseline metrics \(Balanced Accuracy, F1\-Score, and AUROC\) are computed for the baseline \(no\-SAE\) linear probe\. Mean AUROC \(SAE\) is the per\-layer test AUROC averaged across all transformer blocks of the encoder \(±\\pmis the mean per\-fold std averaged across layers, so the unit matches the no\-SAE column\)\. MeanΔ\\Deltais the average AUROC drop from baseline across all layers, with the std taken over layers; MaxΔ\\Deltais the worst\-layer drop\. The same 5\-fold protocol underlies Figure[2](https://arxiv.org/html/2605.13930#S4.F2)\.Encoder\# LBalanced Acc\.F1\-ScoreAUROC \(no SAE\)AUROC \(SAE\)MeanΔ\\DeltaMaxΔ\\DeltaSleepFM30\.901±0\.0190\.901\\pm 0\.0190\.866±0\.0470\.866\\pm 0\.0470\.954±0\.0130\.954\\pm 0\.0130\.951±0\.0150\.951\\pm 0\.0150\.003±0\.0020\.003\\pm 0\.0020\.0050\.005REVE220\.890±0\.0180\.890\\pm 0\.0180\.857±0\.0440\.857\\pm 0\.0440\.944±0\.0090\.944\\pm 0\.0090\.944±0\.0090\.944\\pm 0\.0090\.000±0\.0010\.000\\pm 0\.0010\.0060\.006LaBraM120\.875±0\.0190\.875\\pm 0\.0190\.835±0\.0450\.835\\pm 0\.0450\.939±0\.0170\.939\\pm 0\.0170\.933±0\.0200\.933\\pm 0\.0200\.006±0\.0060\.006\\pm 0\.0060\.0170\.017

The faithfulness check generalises beyond the operating layer\. Figure[2](https://arxiv.org/html/2605.13930#S4.F2)sweeps the SAE\-substitution AUROC across every layer of every encoder: SleepFM stays within0\.0050\.005of its no\-SAE baseline at all 3 layers, REVE within0\.0060\.006across all 22 layers, and LaBraM within0\.0170\.017across all 12 layers\. Interestingly, we see an increase in performance over LaBraM’s 12 layers, but from the fourth layer of LaBraM, and for all layers of SleepFM and REVE, the SAE substitution is statistically indistinguishable from the intact encoder regardless of where it is inserted; not just atℓ∗\\ell^\{\*\}\.

![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/layer_sweep_sae_faithfulness.png)Figure 2:SAE\-faithfulness layer sweep\.Test AUROC of a linear probe trained via 5\-fold cross\-validation on mean\-pooled embeddings of each finetuned encoder\. During inference, layer\-ℓ\\ellactivations are replaced by their TopK\-SAE reconstructions asℓ\\ellsweeps through every transformer block\. Shaded bands represent 95% confidence intervals across the CV folds; the dotted horizontal lines indicate the no\-SAE baseline mean\. The SAE substitution consistently remains within the baseline performance envelope across all layers and architectures\.![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/taxonomy_paper_grid.png)Figure 3:Monosemanticity taxonomy across SAE expansion and encoder depth\.Each cell reports the fraction of concept\-enriched SAE features in one of three taxonomy classes \(Separable: monosemantic; Entangled: polysemantic co\-activations; Dead: semantically uninformative/inactive\)\. Columns represent encoders \(SleepFM, LaBraM, REVE\), with x\-axes indexing the encoder layer and y\-axes indexing expansion factorE∈\{1,2,4,8,16,32,64\}E\\in\\\{1,2,4,8,16,32,64\\\}\. For each layer, we highlight the expansion factor that maximizes the difference between separable and dead feature fractions\. The optimum\(ℓ∗,E∗\)\(\\ell^\{\*\},E^\{\*\}\)is golden\. Find an overview of howEEtranslates to dictionary size in Figure[7](https://arxiv.org/html/2605.13930#A1.F7)\.![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/cross_model_steering_concepts.png)Figure 4:Concept encoding strength and steering selectivity\.Top:Encoding strength \(AUROC0\\mathrm\{AUROC\}\_\{0\}\) measured via per\-layer linear probes fit to the clean SAE\-decoded reconstructions\.Bottom:Excess selectivity \(Δ~\\tilde\{\\Delta\}\), quantifying the integrated asymmetry between target and off\-target probe degradation under TCAV\-ranked clamping \(Section[3\.5](https://arxiv.org/html/2605.13930#S3.SS5)\)\. Forabnormalityas target, we useageas off\-target\. For all other targets, we useabnormalityas off\-target\. Together, these metrics map the representational landscape of each model, distinguishing selectively steerable features from highly entangled ones\. Exactly how excess selectivityΔ~\\tilde\{\\Delta\}is computed is illustrated in Figure[5](https://arxiv.org/html/2605.13930#S4.F5)through examples\.
### 4\.2Spectral decoder reconstruction\.

The spectral decoder reliably recovers token\-level amplitude spectra from frozen embeddings across all architectures\. Overall amplitudeR2R^\{2\}reaches0\.816±0\.0040\.816\\pm 0\.004for SleepFM,0\.783±0\.0040\.783\\pm 0\.004for REVE, and0\.772±0\.0040\.772\\pm 0\.004for LaBraM\. Conversely, phase remains unrecoverable across all models \(cosine similarity≤0\.22\\leq 0\.22\), directly reflecting the time\-translation invariance inherent to their self\-supervised pretraining objectives\. All margins represent 95% bootstrap intervals\.

### 4\.3Operating points transfer across encoders

Figure[3](https://arxiv.org/html/2605.13930#S4.F3)shows the per\-encoder operating points\(ℓ∗,E∗\)\(\\ell^\{\*\},E^\{\*\}\)selected by the procedure of Section[3\.7](https://arxiv.org/html/2605.13930#S3.SS7)\. Despite the three encoders differing in depth \(3/12/22 transformer blocks\), embedding dimension \(128/200/512\) and training objective, the recipe converges on a common modal expansion across each network:E∗=8E^\{\*\}=8wins at 15/22 REVE layers, 3/12 LaBraM layers, and 2/3 SleepFM layers\. High separation at lowEEconcentrates at the boundaries: the encoders’ first and last transformer blocks collapse to a smaller dictionary \(LaBraM L1, L3, L12 and REVE L2, L22 atE∗=1E^\{\*\}=1; LaBraM L2, L10, L11 and REVE L2 atE∗=2E^\{\*\}=2\), consistent with the classification block compressing representations into a low\-dimensional task\-aligned subspace\. SleepFM, with only three transformer blocks, does not compress the representations\.

![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/concept_steering_curves_3x3.png)Figure 5:Steering sweeps across the encoding–selectivity landscape\.Nine representative configurations from Figure[4](https://arxiv.org/html/2605.13930#S4.F4)arranged by encoding strength \(rows\) and selectivityΔ~\\tilde\{\\Delta\}\(columns\)\. Panels track target \(red\) and off\-target \(blue:abnormality\) AUROC as clamping fractionffincreases\. Green shading highlights the TCAV\-driven selectivity gain while grey shading highlights the random\-clamping baseline\. The grid contrasts canonical selective steering \(top\-left: target collapses, off\-target is spared\) with the “wrecking\-ball” failure mode \(top two cells in the right column: both concepts degrade through entanglement,Δ~≈0\\tilde\{\\Delta\}\\approx 0\)\. Encoder colors match Figure[4](https://arxiv.org/html/2605.13930#S4.F4)\.
### 4\.4Cross\-model concept selectivity

We apply the substitution\-sweep protocol of Section[3\.5](https://arxiv.org/html/2605.13930#S3.SS5)to every \(encoder, layer, concept\) triplet and report excess selectivityΔ~\\tilde\{\\Delta\}alongside the clean\-baseline encoding strength \(Figure[4](https://arxiv.org/html/2605.13930#S4.F4)\)\. The two rows together discriminate three regimes:

- •Encoded and steerable:High AUROC0\+ highΔ~\\tilde\{\\Delta\}\.
- •Encoded but not steerable:High AUROC0\+Δ~≈0\\tilde\{\\Delta\}\\approx 0\.
- •Weakly encoded:AUROC<00\.7\{\}\_\{0\}<0\.7\.

To better illustrate the underlying probe\-AUROC sweeps that compute excess selectivity \(Δ~\\tilde\{\\Delta\}\), in Figure[5](https://arxiv.org/html/2605.13930#S4.F5), we unpack nine representative cells from the cross\-model summary in Figure[4](https://arxiv.org/html/2605.13930#S4.F4)\. To understand this grid, the axes should be read as two distinct constraints on the intervention:

Rows define the interventional headroom\.The baseline encoding strength dictates the maximum possible impact of the clamp\. If a concept is weakly encoded, there is no signal to erase; the target probe is already near random guessing, rendering the sweep flat and uninformative\.

Columns map the selectivity regimes\.This axis illustrates the continuum from surgical concept removal to unselective collapse\. We categorize this into three distinct behaviors:

- •Strong selectivity:The target probe \(red\) degrades rapidly under TCAV\-ranked clamping while the off\-target probe \(blue\) remains robust\. Crucially, the target degrades much faster than the random\-clamping baseline, yielding a large selectivity area \(green\) while the random\-area expectation \(grey\) is insignificant\.
- •Moderate selectivity:TCAV ranking isolates the concept faster than random noise, but the off\-target probe is dragged down alongside it, indicating partial entanglement\.
- •Weak selectivity:The green area shrinks to or below the grey area\. Target and off\-target representations degrade in lockstep\. The intervention removes signal, but it acts as a global “wrecking ball” rather than a targeted push\.

Ultimately, this grid provides the qualitative, mechanistic proof for the scalarΔ~\\tilde\{\\Delta\}summaries in Figure[4](https://arxiv.org/html/2605.13930#S4.F4)\. It allows the reader to visually verify that a highΔ~\\tilde\{\\Delta\}truly guarantees an asymmetric, concept\-aligned steering effect, rather than a uniform collapse of the layer’s representation\.

Looking at Figures[4](https://arxiv.org/html/2605.13930#S4.F4)and[5](https://arxiv.org/html/2605.13930#S4.F5), examples of encoded and steerable concepts areabnormalityandagein SleepFM L2/L3 andASMin LaBraM L12\. An example of an encoded but non steerable concept isabnormalityacross all representative REVE layers\. Finally, an example of a weakly encoded concept is sex across all encoders; which functions as the implicit negative control\. Two "wrecking\-ball" examples can be found in Figure[5](https://arxiv.org/html/2605.13930#S4.F5)in the two top cells in the right column \(psychiatric medicinein SleepFM L1 andagein LaBraM L4\): both concepts degrade through entanglement\.

### 4\.5A worked spectrum\-level example

Figure[6](https://arxiv.org/html/2605.13930#S4.F6)grounds the abstract selectivity metrics \(Figure[4](https://arxiv.org/html/2605.13930#S4.F4)\) in domain\-readable signatures\. Steering a representative abnormal adult EEG toward a normal target corrects canonical pathological markers: clamping the top TCAV\-aligned features actively collapses the elevated pathologicalδ\\deltaandθ\\thetapower while restoring the suppressedα\\alphapeak\. Byn=164n=164, the intervened spectrum fully merges into the healthy reference band across the entire 0\.5–45 Hz range, providing clinical experts with physical proof of targeted concept manipulation\.

![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/perfect_steering_classification.png)Figure 6:Spectrum\-level concept steering \(abnormal→\\tonormal\)\.SleepFM layer 2 \(E=8E=8\)\. Shading denotes 95% bootstrap CIs on the target mean\.Left:Baseline abnormal source vs\. normal target centroid\.Centre & Right:The decoded spectrum after clamping the topn=104n=104andn=164n=164TCAV\-aligned features, respectively, to the target centroid\.

## 5Discussion and limitations

Decoder limitations\.Our spectral decoder recovers amplitude but cannot reconstruct phase, as current pretraining objectives discard precise temporal morphology in favor of time\-translation invariance\. Consequently, while we can steer broad frequency distributions, we cannot generate exact time\-domain waveforms \(e\.g\., specific spike\-wave morphologies\)\. High\-fidelity time\-domain steering will ultimately require phase\-aware foundation models or significantly more expressive decoders\.

Separability vs\. selectivity\.Using observational monosemanticity, we assume that separable features provide clean interventional handles\. However, separability does not guarantee selectivity: clamping an observationally pure feature can still trigger a collapse\. Future work should evaluate whether optimizing directly for interventional metrics \(Δ~\\tilde\{\\Delta\}\) yields safer operating points\.

Concept taxonomy is investigator\-defined\.Because our steering protocol \(Section[3\.5](https://arxiv.org/html/2605.13930#S3.SS5)\) requires pre\-specified concepts, our evaluation is strictly bounded by available cohort metadata\. The pipeline verifies rather than discovers\. Although inherent to all probe\-based methods, this restricts how comprehensively a label\-poor dataset can audit a model’s latent space\.

## 6Conclusion

We propose a model\-agnostic interpretability framework for auditing the representational health of EEG foundation models\. By integrating Sparse Autoencoders, concept steering, and spectral decoding across three architecturally distinct encoders, the pipeline maps the semantic taxonomy of transformer blocks; isolating where clinical concepts consolidate into selectively steerable features, and where persistent entanglement triggers “wrecking\-ball” failures during intervention\. By decoding latent representations back into domain\-readable spectral signatures, such as focal slowing or alpha\-band suppression, we bridge the gap between abstract embedding spaces and clinical physiology, grounding model attributions in mechanistic, frequency\-level explanations\.

#### Broader Impact\.

This framework is intended as a diagnostic tool to increase transparency and clinical trust in EEG foundation models\. By surfacing representational failures such as age–pathology entanglement, it can help practitioners identify and mitigate latent demographic biases\. As with all medical AI, frameworks like this should be rigorously validated\.

## References

- \[1\]\(2017\)Understanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§3](https://arxiv.org/html/2605.13930#S3.p1.1)\.
- \[2\]H\. Bao, L\. Dong, S\. Piao, and F\. Wei\(2022\)BEiT: BERT pre\-training of image transformers\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.5.6)\.
- \[3\]Y\. Belinkov\(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422)Cited by:[§3](https://arxiv.org/html/2605.13930#S3.p1.1)\.
- \[4\]N\. Belrose, D\. Schneider\-Joseph, S\. Ravfogel, R\. Cotterell, E\. Raff, and S\. Biderman\(2023\)LEACE: perfect linear concept erasure in closed form\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§3\.6](https://arxiv.org/html/2605.13930#S3.SS6.p2.3)\.
- \[5\]S\. Beniczky, H\. Aurlien, J\. C\. Brøgger, L\. J\. Hirsch, D\. L\. Schomer, E\. Trinka,et al\.\(2017\)Standardized computer\-based organized reporting of EEG: SCORE – second version\.Clinical Neurophysiology128\(11\),pp\. 2334–2346\.External Links:[Document](https://dx.doi.org/10.1016/j.clinph.2017.07.418)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1)\.
- \[6\]T\. Brickenet al\.\(2023\)Towards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[7\]H\. Cunninghamet al\.\(2023\)Sparse autoencoders find highly interpretable features in language models\.arXiv preprint arXiv:2309\.08600\.Cited by:[§2\.2](https://arxiv.org/html/2605.13930#S2.SS2.p1.13),[§3](https://arxiv.org/html/2605.13930#S3.p1.1)\.
- \[8\]Y\. El Ouahidi, J\. Lys, P\. Thölke, N\. Farrugia, B\. Pasdeloup, V\. Gripon, K\. Jerbi, and G\. Lioi\(2025\)REVE: a foundation model for EEG: adapting to any setup with large\-scale pretraining on 25,000 subjects\.Advances in Neural Information Processing Systems\.External Links:[Link](https://brain-bzh.github.io/reve/)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1),[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.4.1)\.
- \[9\]Y\. Elazar, S\. Ravfogel, A\. Jacovi, and Y\. Goldberg\(2021\)Amnesic probing: behavioral explanation with amnesic counterfactuals\.Transactions of the Association for Computational Linguistics9,pp\. 160–175\.Cited by:[§3\.6](https://arxiv.org/html/2605.13930#S3.SS6.p2.3)\.
- \[10\]N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly, N\. DasSarma, D\. Drain, D\. Ganguli, Z\. Hatfield\-Dodds, D\. Hernandez, A\. Jones, J\. Kernion, L\. Lovitt, K\. Ndousse, D\. Amodei, T\. Brown, J\. Clark, J\. Kaplan, S\. McCandlish, and C\. Olah\(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2021/framework/index.html)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[11\]N\. Elhageet al\.\(2022\)Toy models of superposition\.Transformer Circuits Thread\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[12\]N\. Enkhtsetseg, W\. Lehn\-Schiøler, A\. Storgaard Mosquera, M\. Guldberg Pedersen, D\. Rice, G\. Wambugu, Nshimiyimana Jules Fidele, M\. Cacic Hribljan, A\. A\. Arbune, S\. Armand Larsen, S\. Beniczky, and F\. J\. Mateen\(2026\)Clinical utility and feasibility of smartphone\-based EEG in kenya: a multicenter observational study\.arXiv preprint arXiv:2605\.08157\.Cited by:[§3\.1](https://arxiv.org/html/2605.13930#S3.SS1.p1.8)\.
- \[13\]L\. Freeman, P\. Shamash, V\. Arora, C\. Barry, T\. Branco, and E\. Dyer\(2025\)Beyond black boxes: enhancing interpretability of transformers trained on neural data\.External Links:2506\.14014Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[14\]A\. Gjølbye, L\. Skerath, W\. Lehn\-Schiøler, N\. Langer, and L\. K\. Hansen\(2024\)SPEED: scalable preprocessing of EEG data for self\-supervised learning\.InProceedings of the 2024 IEEE International Workshop on Machine Learning for Signal Processing,Cited by:[§3\.1](https://arxiv.org/html/2605.13930#S3.SS1.p1.8)\.
- \[15\]K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollár, and R\. Girshick\(2022\)Masked autoencoders are scalable vision learners\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.4.6)\.
- \[16\]W\. Jiang, L\. Zhao, and B\. Lu\(2024\)Large brain model for learning generic representations with tremendous EEG data in BCI\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1),[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.5.1)\.
- \[17\]M\. Kalnāre, S\. Kitharidis, T\. Bäck, and N\. van Stein\(2026\)Mechanistic interpretability for transformer\-based time series classification\.InComputational Intelligence\. IJCCI 2025,Communications in Computer and Information Science, Vol\.2829\.External Links:[Document](https://dx.doi.org/10.1007/978-3-032-15638-9%5F15)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[18\]B\. Kim, M\. Wattenberg, J\. Gilmer, C\. Cai, J\. Wexler, F\. Viegas, and R\. Sayres\(2018\)Interpretability beyond classification accuracy: quantitative testing with concept activation vectors \(TCAV\)\.InProceedings of the 35th International Conference on Machine Learning \(ICML\),Cited by:[§2\.3](https://arxiv.org/html/2605.13930#S2.SS3.p1.7)\.
- \[19\]D\. Kostas, S\. Aroca\-Ouellette, and F\. Rudzicz\(2021\)BENDR: using transformers and a contrastive self\-supervised learning task to learn from massive amounts of EEG data\.Frontiers in Human Neuroscience15\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1)\.
- \[20\]W\. Lehn\-Schiøler, M\. R\. Kjær, P\. Hempel, M\. G\. Pedersen, R\. Thapa, B\. He, N\. Spicher, A\. Brink\-Kjaer, L\. K\. Hansen, and E\. Mignot\(2026\)Pretraining on sleep data improves non\-sleep biosignal tasks\.arXiv preprint arXiv:2605\.02500\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1)\.
- \[21\]T\. Lieberumet al\.\(2024\)Gemma scope: open sparse autoencoders everywhere all at once on Gemma 2\.arXiv preprint arXiv:2408\.05147\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[22\]A\. G\. Madsen, W\. T\. Lehn\-Schiøler, Á\. Jónsdóttir, B\. Arnardóttir, and L\. K\. Hansen\(2023\-09\)Concept\-based explainability for an eeg transformer model\.In2023 IEEE 33rd International Workshop on Machine Learning for Signal Processing \(MLSP\),pp\. 1–6\.External Links:[Link](http://dx.doi.org/10.1109/MLSP55844.2023.10285992),[Document](https://dx.doi.org/10.1109/mlsp55844.2023.10285992)Cited by:[§3](https://arxiv.org/html/2605.13930#S3.p1.1)\.
- \[23\]A\. Makhzani and B\. Frey\(2014\)K\-sparse autoencoders\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2605.13930#S2.SS2.p1.1)\.
- \[24\]K\. K\. Nakka\(2025\)Mammo\-sae: interpreting breast cancer concept learning with sparse autoencoders\.External Links:2507\.15227Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[25\]R\. Renzulli, C\. Lepoutre, E\. Cassano, and M\. Grangetto\(2025\)MedSAE: dissecting medclip representations with sparse autoencoders\.arXiv preprint arXiv:2510\.26411\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[26\]E\. Simon and J\. Zou\(2025\)InterPLM: discovering interpretable features in protein language models via sparse autoencoders\.Nature Methods22\(10\),pp\. 2107–2117\.Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2)\.
- \[27\]A\. Templetonet al\.\(2024\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p2.2),[§2\.2](https://arxiv.org/html/2605.13930#S2.SS2.p1.13),[§3\.3](https://arxiv.org/html/2605.13930#S3.SS3.p1.10)\.
- \[28\]R\. Thapa, M\. R\. Kjaer, B\. He, I\. Covert, H\. Moore IV, U\. Hanif, G\. Ganjoo, M\. B\. Westover, P\. Jennum, A\. Brink\-Kjaer, E\. Mignot, and J\. Zou\(2026\)A multimodal sleep foundation model for disease prediction\.Nature Medicine32,pp\. 752–762\.External Links:[Document](https://dx.doi.org/10.1038/s41591-025-04133-4)Cited by:[§1](https://arxiv.org/html/2605.13930#S1.p1.1),[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.3.1)\.
- \[29\]A\. van den Oord, Y\. Li, and O\. Vinyals\(2018\)Representation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.Cited by:[Table 1](https://arxiv.org/html/2605.13930#S2.T1.2.2.3.6)\.

## Appendix ATechnical appendices and supplementary material

Table 3:Notation reference\.SymbolType / shapeDefinitionFirst usedEncoderddscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Embedding dimension \(128 / 200 / 512 for SleepFM / LaBraM / REVE\)Eq\.[1](https://arxiv.org/html/2605.13930#S2.E1)𝐚\\mathbf\{a\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}Layer activation extracted from the frozen encoderEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)𝐭\\mathbf\{t\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}Token embedding passed to the spectral decoderEq\.[3](https://arxiv.org/html/2605.13930#S3.E3)ℓ\\ellscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Transformer layer indexSec\.[3\.3](https://arxiv.org/html/2605.13930#S3.SS3)ℓ∗\\ell^\{\*\}scalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Optimal layer selected by Eq\.[9](https://arxiv.org/html/2605.13930#S3.E9)Eq\.[9](https://arxiv.org/html/2605.13930#S3.E9)SAE architectureEEscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Expansion ratioSec\.[3\.3](https://arxiv.org/html/2605.13930#S3.SS3)N≜d⋅EN\\triangleq d\\cdot Escalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Dictionary size \(number of SAE features\)Eq\.[1](https://arxiv.org/html/2605.13930#S2.E1)kkscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}TopK sparsity budget;k=k0⋅Ek=k\_\{0\}\\cdot E,k0=8k\_\{0\}=8Eq\.[1](https://arxiv.org/html/2605.13930#S2.E1)k0k\_\{0\}scalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Base sparsity constant;k0=8k\_\{0\}=8, sok/N=k0/dk/N=k\_\{0\}/dis constantSec\.[3\.3](https://arxiv.org/html/2605.13930#S3.SS3)WencW\_\{\\text\{enc\}\}matrix∈ℝN×d\\in\\mathbb\{R\}^\{N\\times d\}SAE encoder weight matrixEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)WdecW\_\{\\text\{dec\}\}matrix∈ℝd×N\\in\\mathbb\{R\}^\{d\\times N\}SAE decoder; columns𝐰i∈ℝd\\mathbf\{w\}\_\{i\}\\in\\mathbb\{R\}^\{d\}are unit\-norm directionsEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)𝐰i\\mathbf\{w\}\_\{i\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}ii\-th column ofWdecW\_\{\\text\{dec\}\}; unit\-norm decoder directionEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)bdecb\_\{\\text\{dec\}\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}SAE decoder biasEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)μℓ,σℓ\\mu\_\{\\ell\},\\,\\sigma\_\{\\ell\}vectors∈ℝd\\in\\mathbb\{R\}^\{d\}Per\-dimension mean and std used to normalise activationsEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)𝐳\\mathbf\{z\}vector∈ℝN\\in\\mathbb\{R\}^\{N\}Sparse SAE feature activation \(exactlykknon\-zero entries\)Eq\.[1](https://arxiv.org/html/2605.13930#S2.E1)𝐚^\\hat\{\\mathbf\{a\}\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}SAE reconstruction of the encoder activationEq\.[1](https://arxiv.org/html/2605.13930#S2.E1)E∗E^\{\*\}scalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Optimal expansion ratio selected by Eq\.[9](https://arxiv.org/html/2605.13930#S3.E9)Eq\.[9](https://arxiv.org/html/2605.13930#S3.E9)Spectral decoderSD\\mathrm\{SD\}ℝd→ℝF×ℝ2​F\\mathbb\{R\}^\{d\}\\\!\\to\\\!\\mathbb\{R\}^\{F\}\\\!\\times\\\!\\mathbb\{R\}^\{2F\}Shallow MLP mapping a token embedding to amplitude and phaseEq\.[3](https://arxiv.org/html/2605.13930#S3.E3)FFscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Number of frequency bins \(F=64F=64\)Eq\.[3](https://arxiv.org/html/2605.13930#S3.E3)A^ν\\hat\{A\}\_\{\\nu\}scalar∈ℝ\\in\\mathbb\{R\}Predicted amplitude at frequency binν\\nuEq\.[3](https://arxiv.org/html/2605.13930#S3.E3)φ^ν\\hat\{\\varphi\}\_\{\\nu\}scalar∈\[0,2​π\)\\in\[0,2\\pi\)Phase at binν\\nu; parameterised as\(cos⁡φ^ν,sin⁡φ^ν\)\(\\cos\\hat\{\\varphi\}\_\{\\nu\},\\sin\\hat\{\\varphi\}\_\{\\nu\}\)Eq\.[3](https://arxiv.org/html/2605.13930#S3.E3)Concept attribution \(TCAV\)𝒞\\mathcal\{C\}labelClinical concept \(abnormality, age group, sex, medication\)Sec\.[3\.4](https://arxiv.org/html/2605.13930#S3.SS4)𝐯C\\mathbf\{v\}\_\{C\}vector∈ℝd\\in\\mathbb\{R\}^\{d\}Concept Activation Vector; unit\-norm direction in activation spaceSec\.[2\.3](https://arxiv.org/html/2605.13930#S2.SS3)SC​\(x\)S\_\{C\}\(x\)scalar∈ℝ\\in\\mathbb\{R\}TCAV sensitivity score for examplexxunder conceptCCEq\.[2](https://arxiv.org/html/2605.13930#S2.E2)TCAVC\\mathrm\{TCAV\}\_\{C\}scalar∈\[0,1\]\\in\[0,1\]Fraction of concept examples withSC​\(x\)\>0S\_\{C\}\(x\)\>0; above 0\.5⇒\\RightarrowencodedEq\.[2](https://arxiv.org/html/2605.13930#S2.E2)NrandN\_\{\\text\{rand\}\}scalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Number of random\-label null CAVs for significance testing;Nrand=50N\_\{\\text\{rand\}\}=50Sec\.[2\.3](https://arxiv.org/html/2605.13930#S2.SS3)Kf​o​l​dK\_\{fold\}scalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Cross\-validation folds for CAV training;Kf​o​l​d=10K\_\{fold\}=10Sec\.[3\.4](https://arxiv.org/html/2605.13930#S3.SS4)Concept steeringrankC​\(i\)\\mathrm\{rank\}\_\{C\}\(i\)functionℤ\+→ℝ\\mathbb\{Z\}^\{\+\}\\\!\\to\\\!\\mathbb\{R\}CAV\-alignment score for featureii:\|𝐯C⋅𝐰i\|\|\\mathbf\{v\}\_\{C\}\\cdot\\mathbf\{w\}\_\{i\}\|Eq\.[4](https://arxiv.org/html/2605.13930#S3.E4)XtargetX\_\{\\text\{target\}\}datasetTarget pool for the steering interventionEq\.[5](https://arxiv.org/html/2605.13930#S3.E5)𝐜target∈ℝN\\mathbf\{c\}\_\{\\text\{target\}\}\\in\\mathbb\{R\}^\{N\}vector∈ℝN\\in\\mathbb\{R\}^\{N\}Target\-concept centroid; mean SAE activations overXtargetX\_\{\\text\{target\}\}Eq\.[5](https://arxiv.org/html/2605.13930#S3.E5)ffscalar∈\[0,1\]\\in\[0,1\]Clamping fractionEq\.[6](https://arxiv.org/html/2605.13930#S3.E6)n≜⌊f⋅N⌋n\\triangleq\\lfloor f\\cdot N\\rfloorscalar∈ℤ\+\\in\\mathbb\{Z\}^\{\+\}Number of features clamped \(also used bare in Section 4\.5 / Fig\. 6\)Eq\.[6](https://arxiv.org/html/2605.13930#S3.E6)𝐳∗​\(f\)\\mathbf\{z\}^\{\*\}\(f\)vector∈ℝN\\in\\mathbb\{R\}^\{N\}Intervened activation after clamping top\-nnfeaturesEq\.[6](https://arxiv.org/html/2605.13930#S3.E6)𝐚^∗​\(f\)\\hat\{\\mathbf\{a\}\}^\{\*\}\(f\)vector∈ℝd\\in\\mathbb\{R\}^\{d\}SAE\-decoded embedding of𝐳∗​\(f\)\\mathbf\{z\}^\{\*\}\(f\)Eq\.[6](https://arxiv.org/html/2605.13930#S3.E6)Δ​\(rank\)\\Delta\(\\text\{rank\}\)scalar∈ℝ\\in\\mathbb\{R\}Raw selectivity areaEq\.[7](https://arxiv.org/html/2605.13930#S3.E7)Δ~\\tilde\{\\Delta\}scalar∈ℝ\\in\\mathbb\{R\}Excess selectivity=Δ​\(TCAV\)−𝔼π​\[Δ​\(π\)\]=\\Delta\(\\text\{TCAV\}\)\-\\mathbb\{E\}\_\{\\pi\}\[\\Delta\(\\pi\)\]Eq\.[8](https://arxiv.org/html/2605.13930#S3.E8)Hyperparameter selectionseparableℓ,E\\text\{separable\}\_\{\\ell,E\}scalar∈\[0,1\]\\in\[0,1\]Fraction of features classified as separable at layerℓ\\ell, expansionEEEq\.[9](https://arxiv.org/html/2605.13930#S3.E9)deadℓ,E\\text\{dead\}\_\{\\ell,E\}scalar∈\[0,1\]\\in\[0,1\]Fraction of features classified as dead at layerℓ\\ell, expansionEEEq\.[9](https://arxiv.org/html/2605.13930#S3.E9)Table 4:Encoder architectures and binary fine\-tuning hyperparameters\. All three encoders are finetuned end\-to\-end on the binary \(normal vs\. abnormal\) label using AdamW with two\-phase training: a head\-only warm\-up phase followed by full unfreezing\. Learning rate is dropped when the encoder is unfrozen and the weight decay is increased\.[SleepFM](https://www.nature.com/articles/s41591-025-04133-4)[REVE\-Base](https://arxiv.org/abs/2510.21585)[LaBraM\-Base](https://arxiv.org/abs/2405.18765)*Where can I find the encoder?*[Github](https://github.com/zou-group/sleepfm-clinical)[Hugging Face](https://huggingface.co/brain-bzh/reve-base)[braindecode](https://braindecode.org/dev/generated/braindecode.models.Labram.html)and[Github](https://github.com/935963004/LaBraM)*Architecture*BackboneSetTransformerViTNeuralTransformerEmbed\. dim128512200Transformer layers32212Attention heads8810Head dim166420MLP ratio162\.66 \(GeGLU\)4\.0Patch size \(samples\)128200200Tokens per window6019×6619\\times 66\(overlapping\)19×6019\\times 60Token receptive field1\.00 s1\.00 s1\.00 sToken stride1\.00 s0\.90 s1\.00 s*Input data*DatasetBinary \(normal/abnormal\)Binary \(normal/abnormal\)Binary \(normal/abnormal\)Sample rate \(Hz\)128200200Window length606060Channels27 \(10\-20\)27 \(10\-20\)19 \(10\-20\)*Finetuning*Poolingtemporal\_poolingmean overC⋅SC\\\!\\cdot\\\!Stokensmean over patch tokensHeadLinear\(128,1\)Linear\(512,1\)Linear\(200,1\)LossBCEBCEBCE*Optimisation*OptimiserAdamWAdamWAdamWTotal epochs121015Head\-only w/u ep\.223Batch size3288LR \(head phase\)10−310^\{\-3\}10−310^\{\-3\}10−310^\{\-3\}LR \(full phase\)10−410^\{\-4\}10−410^\{\-4\}10−410^\{\-4\}WD \(head/full\)0\.01 / 0\.050\.01 / 0\.050\.01 / 0\.05Grad\-norm clip1\.01\.01\.0![Refer to caption](https://arxiv.org/html/2605.13930v1/figures/dictionary_size.png)Figure 7:SAE dictionary size across encoders and expansion rates\.Each cell gives the number of learned SAE features, which equals the encoder’s embedding dimension \(denc=128d\_\{\\text\{enc\}\}=128for SleepFM,200200for LaBraM,512512for REVE\) times the expansion rateEE\. Because REVE is4×4\\timeswider than SleepFM, anE=1E\{=\}1REVE SAE already exceeds the size of anE=4E\{=\}4SleepFM SAE, and theE=64E\{=\}64REVE configuration spans32,76832\{,\}768features\. This asymmetry should be kept in mind when comparing taxonomy fractions \(Fig\.[3](https://arxiv.org/html/2605.13930#S4.F3)\) across encoders at matchedEE: equalEEdoes*not*mean equal capacity\.

Similar Articles

The Identity Trap in EEG Foundation Models: A Diagnostic Audit

arXiv cs.LG

This paper identifies and diagnoses the 'Identity Trap' in EEG foundation models, where high accuracy may stem from subject-identity features rather than genuine clinical biomarkers. It proposes FMScope, a frozen-representation protocol to disentangle these signals, and demonstrates that subject-identity confounding is universal across three models and removable with linear methods.