词性作为SAE潜在空间中的涌现类别
摘要
该论文研究了词性类别在Sparse AutoEncoder潜在空间中的编码方式,发现它们是分布式的,并不与单个潜在变量一一对应。
arXiv:2609.29362v1 Announce Type: new
Abstract: Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
查看缓存全文
缓存时间: 2026/09/25 09:19
# Parts-of-Speech as Emergent Categories in SAE Latent Space Source: [https://arxiv.org/html/2609.29362](https://arxiv.org/html/2609.29362) Alessandro BondielliAffiliation:CoLingLab, Department of Philology, Literature and Linguistics, University of PisaAffiliation:Department of Computer Science, University of Pisa \*Equal contribution\.Correspondence:[alessandro\.bondielli@unipi\.it](mailto:[email protected]),[lucia\.passaro@unipi\.it](mailto:[email protected]),Preprint version of the paper in the Proceedings of EMNLP 2026\.Lucia PassaroAffiliation:CoLingLab, Department of Philology, Literature and Linguistics, University of PisaAffiliation:Department of Computer Science, University of Pisa \*Equal contribution\.Correspondence:[alessandro\.bondielli@unipi\.it](mailto:[email protected]),[lucia\.passaro@unipi\.it](mailto:[email protected]),Preprint version of the paper in the Proceedings of EMNLP 2026\.Serena AuriemmaAlessandro LenciAffiliation:CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa ###### Abstract Sparse AutoEncoders \(SAEs\) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose\. We use part\-of\-speech \(PoS\) categories as a controlled test case to study whether morpho\-syntactic information is encoded by individual latents or by structured groups of features\. We find thatPoSdistinctions are highly recoverable fromSAEactivations, but do not align with one\-to\-one latent / category mappings\. This recoverability is not reducible to lexical memorisation, and Open and ClosedPoSclasses differ substantially\. Categories are supported by compact groups of sparse latents, with substantial variation across tags\. These groups remain stable on held\-out data, while also showing overlap between related categories\. Our results show that SAEs localise morpho\-syntactic information in a distributed and category\-dependent form rather than through atomic grammatical features\.111Code and Data available here:[https://github\.com/colinglab/pos\-sae\-latents](https://github.com/colinglab/pos-sae-latents)\. ## 1Introduction Large language models \(LLMs\) encode a wide range of linguistic regularities in their internal representations, from lexical and syntactic information to more abstract semantic and discourse\-level properties\. Yet, despite substantial progress in probing and representation analysis, it remains unclear how such information is organized internally, for instance whether linguistic categories correspond to localized and interpretable units, or they are instead distributed across many dimensions of the representation space[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.29362#bib.bib10)\. This question has become particularly relevant with the growing use of Sparse AutoEncoders \(SAEs\) as tools for interpreting LLMs[Bricken et al\. \(2023\)](https://arxiv.org/html/2609.29362#bib.bib11);[Cunningham et al\. \(2023\)](https://arxiv.org/html/2609.29362#bib.bib12);[Templeton et al\. \(2024\)](https://arxiv.org/html/2609.29362#bib.bib13)\. SAEs aim to decompose dense model activations into high\-dimensional sparse representations, where individual dimensions \(akalatents\) are expected to capture more interpretable directions of variation\.SAElatents are expected to provide a bridge between low\-level model activations and human\-interpretable features\. This has motivated their use in mechanistic interpretability, where they are often discussed in terms of feature discovery and monosemanticity[Elhage et al\. \(2022\)](https://arxiv.org/html/2609.29362#bib.bib10);[Bricken et al\. \(2023\)](https://arxiv.org/html/2609.29362#bib.bib11);[Templeton et al\. \(2024\)](https://arxiv.org/html/2609.29362#bib.bib13)\. However, the relationship between SAE latents andlinguistic categoriesis still not clear\. In fact, the latter result from the combinations of multiple lexical, morphological, syntactic, and distributional features, which need not correspond to individual isolated latents\. Understanding whether linguistic abstractions are localized or distributed inSAEspaces is thus important for evaluating what kind of interpretabilitySAEs provide[Kantamneni et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib17);[Karvonen et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib18);[Engels et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib19)\. Figure 1:Overview of the workflow\. We testPoSrecoverability from token\-levelSAEactivations, identify and compactPoS\-relevant latent groups, and validate them on held\-out and controlled data\.In this paper, we study this question by targeting part\-of\-speech \(PoS\) categories\.PoStags offer a controlled testbed for analyzing morpho\-syntactic abstraction: They are discrete, independently annotated, and linguistically interpretable, while also differing in frequency, lexical openness, and syntactic function\. For instance, open\-class categories such as nouns and verbs are lexically productive and highly variable, whereas closed\-class categories such as determiners, conjunctions, and pronouns are more restricted and often tied to specific syntactic roles\. This makesPoSa useful setting for testing whetherSAElatents behave as localized linguistic features or instead participate in broader distributed representations\. We analyzeSAEactivations extracted from LLaMA\-3\-8B[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.29362#bib.bib21)on the GUM Corpus Treebank[Zeldes \(2017\)](https://arxiv.org/html/2609.29362#bib.bib1)\. For each token, we encode itsSAErepresentation into a sparse activation vector and study how gold Universal Dependencies \(UD\)PoStags are represented in this latent space\. Our experimental design follows a three\-step interpretability pipeline: i\.\) we usebinary probing classifiersto test whether individualPoSdistinctions are recoverable fromSAEactivations[Belinkov \(2022\)](https://arxiv.org/html/2609.29362#bib.bib4); ii\.\) we usefeature\-salience analysisto rank the latents most relevant to eachPoScategory, andcoverage analysisto estimate how many of these latents are needed to account for most instances of the category; iii\.\) we validate the selected latent groups onheld\-out and controlled data, and test whether their union is sufficient to train a multi\-classPoSclassifier\. This setup allows us to move beyond standard probing accuracy\. A high probing score may show thatPoSinformation is present inSAEactivations, but it does not explain how this information is organized[Hewitt and Liang \(2019\)](https://arxiv.org/html/2609.29362#bib.bib8);[Pimentel et al\. \(2020\)](https://arxiv.org/html/2609.29362#bib.bib9);[Belinkov \(2022\)](https://arxiv.org/html/2609.29362#bib.bib4)\. By combining probing, salience, coverage, compact\-feature classification, and held\-out validation, we can understand whetherPoScategories are associated with individual monosemantic latents or with structured groups of sparse latents\. We also examine whether differentPoScategories are represented by different numbers of latents, which might suggest that theSAErepresentation reflects differences in the linguistic nature of the categories themselves\. We address three research questions:\(i\)RQ1:Are PoS categories explicitly encoded inSAElatent activations?\(ii\)RQ2:What is the organization of thePoScategories encoding in theSAElatent spaces?\(iii\)RQ3:How stable and systematic are these latent representations across linguistic categories, datasets, and evaluation settings? Our study makes two main contributions\. First, we show thatPoScategories are aligned with structured groups of sparse features\. Through feature\-salience and coverage analyses, we quantify the size and organization of these groups, show that it varies substantially across categories, and highlight differences between Open\- and Closed\-classPoSclasses\. Second, we show thatthese latent groups are compact yet effective: Their union preserves strong multi\-classPoSclassification performance, and they remain stable on held\-out data, while still exhibiting overlap across related categories\. ## 2Related Work Probing classifiers have long been used to test what linguistic information neural language models encode in their representations\([Conneau et al\., 2018](https://arxiv.org/html/2609.29362#bib.bib3);[Belinkov, 2022](https://arxiv.org/html/2609.29362#bib.bib4)\)\. Prior work shows that lower layers capture morpho\-syntactic information such asPoS, while higher layers encode more abstract semantic and discourse properties\([Tenney et al\., 2019a](https://arxiv.org/html/2609.29362#bib.bib5);[Tenney et al\., 2019b](https://arxiv.org/html/2609.29362#bib.bib6);[Hewitt and Manning, 2019](https://arxiv.org/html/2609.29362#bib.bib7);[Rogers et al\., 2020](https://arxiv.org/html/2609.29362#bib.bib20)\)\. However, probing accuracy alone is limited: control tasks\([Hewitt and Liang, 2019](https://arxiv.org/html/2609.29362#bib.bib8)\)and information\-theoretic critiques\([Pimentel et al\., 2020](https://arxiv.org/html/2609.29362#bib.bib9)\)show that probes can fit arbitrary mappings, and that recoverability does not imply use\. We share this concern, but shift the focus from*what*information is present to*how*it is organised at the level of individual sparse latents\. A growing body of work studies mechanistic interpretability in LLMs[Sharkey et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib22)\. Within this area,SAEs map dense activations to high\-dimensional sparse vectors whose units are intended to be more monosemantic and interpretable\([Bricken et al\., 2023](https://arxiv.org/html/2609.29362#bib.bib11);[Cunningham et al\., 2023](https://arxiv.org/html/2609.29362#bib.bib12)\)\. Subsequent work has improvedSAEtraining through scaling\([Templeton et al\., 2024](https://arxiv.org/html/2609.29362#bib.bib13)\)and TopK activations\([Gao et al\., 2025](https://arxiv.org/html/2609.29362#bib.bib2)\), and released openSAEsuites for widely used base models\([Lieberum et al\., 2024](https://arxiv.org/html/2609.29362#bib.bib14);[He et al\., 2024](https://arxiv.org/html/2609.29362#bib.bib15)\)\. These studies often identify latents aligned with intuitive concepts, using top\-activating examples or automated natural\-language explanations, but provide limited evidence on how*theoretically motivated linguistic categories*are represented in latent space\. Recent work also questions whetherSAElatents behave as genuinely monosemantic features\.[Kantamneni et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib17)find that probes trained onSAElatents do not consistently outperform simple baselines across113113binary classification tasks, while SAEBench\([Karvonen et al\., 2025](https://arxiv.org/html/2609.29362#bib.bib18)\)shows that gains on standardSAEproxy metrics often do not transfer to downstream performance\. Still,SAEs remain useful tools for probing LM knowledge and behaviour[Dupre la Tour and Mossing \(2025\)](https://arxiv.org/html/2609.29362#bib.bib23);[Fraser\-Taliente et al\. \(2026\)](https://arxiv.org/html/2609.29362#bib.bib24)\. Closer to our work,[Marks et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib16)useSAEfeatures to construct interpretable causal circuits for syntactic phenomena such as subject–verb agreement, suggesting that morpho\-syntactic information is at least partly recoverable fromSAEspace\. At the same time,[Engels et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib19)show that not all language model features are well captured by single linear directions\. Existing linguistic analyses ofSAEs have mainly focused on isolated phenomena, such as subject–verb agreement, or broad properties such as language identity\. We address the open question of whether and how classical morpho\-syntactic categories are encoded by latents, usingPoSas a controlled testbed beyond the binary probing regime explored by prior work\. ## 3Method and Materials ### 3\.1Dataset We conducted our experiments on two datasets: a naturally occurring corpus and a small controlled dataset constructed for targeted evaluation\. The GUM treebank\.For the naturally occurring data, we selected the UD English GUM treebank\([Zeldes, 2017](https://arxiv.org/html/2609.29362#bib.bib1)\), annotated following the Universal Dependencies scheme\.222[https://universaldependencies\.org/](https://universaldependencies.org/)We chose this treebank for its representativeness across diverse textual genres \(academic, blog, legal, news, social, wiki, etc\.\), its medium size \(14,353 sentences, 252,284 tokens\), and its complete coverage of the 17 Universal PoS tags\. The training split was used as the discovery set, while the test split was kept held out and used only to evaluate whether discovered activations remain active on unseen tokens of the correspondingPoScategories\. Controlled dataset\.To complement the naturally occurring data, we constructed a small controlled dataset of 180 lexical items to verify whether latents associated with specificPoStags activate systematically in minimal, grammatically well\-formed sentences\. The dataset focuses primarily on nouns and verbs\. For nouns, we selected 160 items spanning multiple semantic categories \(e\.g\., mammals, birds, flowers, vehicles, etc\.\), evenly split between animate and inanimate referents\. Each noun was instantiated in singular and plural form within neutral templates, including impersonal constructions such asThere is a dogand transitive constructions such asI see the dogandI have a dog\. These templates vary determiner contexts, including indefinite articles, definite articles, and bare plurals\. For verbs, we included 20 high\-frequency verbs compatible with a minimal intransitive template \(I \+ verb, as inI walk\), balanced between 10 regular and 10 irregular past\-tense forms, to limit the impact of morphological idiosyncrasies\. All base sentences were augmented with two variants: one adding an adjacent adjective for noun sentences or adverb for verb sentences, and one appending punctuation to the augmented sentence\. This allows us to assess whether additionalPoStokens introduce their own characteristic activations and whether these interact with those observed in the base sentence\. All sentences were also instantiated in present and past tense to account for potential tense\-driven effects\. ### 3\.2Processing Pipeline In the following, we describe the processing pipeline to obtain SAE latent activations\. ##### Model\. We experiment onLLaMA\-3\-8B\. We employ theEleutherAI/sae\-llama\-3\-8b\-32xpre\-trained model as our SAE\. Both models are available on HuggingFace\. The SAE model is trained and used via the Sparsify library\.333[https://github\.com/EleutherAI/sparsify](https://github.com/EleutherAI/sparsify)The library is designed to follow the SAE implementation described in[Gao et al\. \(2025\)](https://arxiv.org/html/2609.29362#bib.bib2)\. ##### Token\-level Activations Extraction\. To extract token\-level activations, we feed the raw sentence text to the model using its original subword tokenizer, and recover hidden state activations from the residual stream of layer 30 \(last layer before the output\) of the model\. We encode such activations with the SAE to produce the sparse activation vectors for each subword\. We obtain, for each subword, the fraction of SAE latents that fired on that subword, and their activation strength\. Then, we align subword tokens and UD surface forms via character\-span overlap: For each UD token with character span\[tstart,tend\)\[t\_\{\\mathrm\{start\}\},\\,t\_\{\\mathrm\{end\}\}\), all subword tokens whose span\[s,e\)\[s,\\,e\)satisfiess<tends<t\_\{\\mathrm\{end\}\}ande\>tstarte\>t\_\{\\mathrm\{start\}\}are identified as overlapping\. The leftmost such subword token is designated theanchor, and its SAE activations are adopted as the representation of the corresponding UD token\. Note that we chose the leftmost subword because it is the position at which the UD token’s identity first becomes available to the model\. Averaging over subwords may instead dilute category\-bearing activations with continuation\-piece activations\. This yields, for each token, a dense SAE activation vector that is composed of all SAE latents that fired on the token and their activation strength\. ##### Sparse Feature Matrix Construction\. For the probing experiments, we construct a SparseSAE Feature Matrixfrom dense token\-level activations\. To do so, we consider the union of latents active across the entire treebank, which constitutes a subset of the full SAE latent space, spanning4096×32=131,0724096\\times 32=131\{,\}072dimensions \(i\.e\., LLM hidden size×\\timesSAE expansion factor\)\. We therefore project all token representations into the common sparse vector space defined by the latents observed at least once in the Treebank\. Concretely, we construct a feature matrix𝐗∈ℝN×D\\mathbf\{X\}\\in\\mathbb\{R\}^\{N\\times D\}, whereNNis the number of tokens andD=130,246D=130\{,\}246is the number of attested latents, with entryxi,jx\_\{i,j\}set to the activation strength of latentjjon tokenii, and zero otherwise\. The sparse matrix is the input to the probing classifiers\.444[https://huggingface\.co/datasets/colinglab/UD\_English\-GUM\-Latents\_Meta\-Llama\-3\-8B\_L30](https://huggingface.co/datasets/colinglab/UD_English-GUM-Latents_Meta-Llama-3-8B_L30) ## 4Experiments Our experiments are designed to assess not only whetherPoSinformation is recoverable fromSAEactivations, but also how this information is organized in the latent space\. In particular, we structure the analysis around the three research questions introduced in Section[1](https://arxiv.org/html/2609.29362#S1)\. First, we test whether morpho\-syntactic distinctions are explicitly available in the sparse activation space \(RQ1\)\. Second, we investigate the organization ofPOScategories in the latent space \(RQ2\)\. Third, we evaluate whether such organization is stable across splits and evaluation settings, and whether it supports generalPoSclassification \(RQ3\)\. The experimental pipeline proceeds as follows\. We first assess the linear recoverability of eachPoScategory with one\-vs\-rest probing classifiers \(Section[4\.1](https://arxiv.org/html/2609.29362#S4.SS1)\)\. We then localizePoS\-relevant latent groups by combining feature salience, coverage, and compactness analyses \(Section[4\.2](https://arxiv.org/html/2609.29362#S4.SS2)\)\. Finally, we test the robustness of the selected groups on held\-out data \(Section[4\.3](https://arxiv.org/html/2609.29362#S4.SS3)\)\. This design separates recoverability, localization, and stability\. Probing shows whetherPoSinformation is present, while localization and validation assess how such information is organized in the latent space\. ### 4\.1Recoverability ofPoSInformation \(RQ1\) We first test whetherPoSdistinctions are linearly recoverable fromSAEactivations\. For each token in the GUM training split, we use the sparseSAEactivation vector \(cf\. Section[3\.2](https://arxiv.org/html/2609.29362#S3.SS2)\) as input representation and the gold UDPoStag as supervision\. We evaluate the one\-vs\-all setting using 5\-fold cross\-validation on the GUM Treebank Train split\. For eachPoScategory, we train a binary classifier to distinguish tokens with thatPoSfrom all other tokens\. We train an L1\-regularized logistic regression classifier \(C = 0\.1\) using the liblinear solver, with balanced class weighting to account for label imbalance\. The L1 penalty encourages sparse weight vectors, effectively performing feature selection and yielding interpretable models where most coefficients are driven to zero, given the high\-dimensional nature of SAE latent spaces\. This allows us to assess the extent to which individualPoSdistinctions are linearly recoverable fromSAEactivations\. ### 4\.2PoSOrganization in Latent Space \(RQ2\) We next ask how thePoSinformation recovered by the probes can be localized in the space ofSAElatents\. To this end, we use the one\-vs\-rest classifiers introduced in Section[4\.1](https://arxiv.org/html/2609.29362#S4.SS1)not only as predictive models, but also as feature\-salience mechanisms\. ##### Feature salience\. For eachPoScategory, the corresponding logistic regression classifier assigns a coefficientβ\\betato each latent\. Sinceβ\>0\\beta\>0indicates that the activation of a latent increases the probability of the positive class, we rank latents for each category according to their positive coefficients\. We consider only latents withβ\>0\\beta\>0, obtaining for eachPoStag a salience\-ranked list of features that support the classification of that category\. This analysis moves from recoverability to localization: rather than asking whetherPoSinformation is present, we ask which sparse features contribute most to each distinction\. ##### Coverage and compactness\. We then quantify how compact each localized group is\. For eachPoStagcc, letLc\(k\)L^\{\(k\)\}\_\{c\}denote the set of the top\-kklatents in its salience\-ranked list,TcT\_\{c\}the set of gold\-label tokens tagged withcc, andaℓ\(t\)a\_\{\\ell\}\(t\)the activation of latentℓ\\ellon tokentt\. We define the coverage ofLc\(k\)L^\{\(k\)\}\_\{c\}as: Covc\(k\)=1\|Tc\|∑t∈Tc𝟏\[∃ℓ∈Lc\(k\):aℓ\(t\)\>0\]\\mathrm\{Cov\}\_\{c\}\(k\)=\\frac\{1\}\{\|T\_\{c\}\|\}\\sum\_\{t\\in T\_\{c\}\}\\mathbf\{1\}\\left\[\\exists\\ell\\in L^\{\(k\)\}\_\{c\}:a\_\{\\ell\}\(t\)\>0\\right\] Coverage measures the proportion of tokens of categoryccfor which at least one of the top\-kksalient latents is active\. We define the number of latents required to account for categoryccas the smallestkksuch that coverage reaches a target thresholdτ=0\.95\\tau=0\.95: kc⋆=min\{k∈ℕ:Covc\(k\)≥τ\}k^\{\\star\}\_\{c\}=\\min\\\{k\\in\\mathbb\{N\}:\\mathrm\{Cov\}\_\{c\}\(k\)\\geq\\tau\\\} This gives an estimate of the effective size of the latent group associated with eachPoScategory\. We use a per\-class threshold rather than a global top\-kkor coefficient threshold to avoid biasing the comparison due to high imbalance in i\.\) relevant latents for eachPoSand ii\.\) coefficient profiles in open\- vs closed\-classes\. Comparingkc⋆k^\{\\star\}\_\{c\}across tags allows us to test whether different categories are represented with different degrees of compactness, for example whether open\-class categories require broader latent groups than closed\-class categories\. ##### Compact\-feature classification\. Finally, we test whether the localized latent groups are sufficient for jointPoSprediction\. LetCCdenote the set ofPoScategories\. We define the set ofPoS\-relevant latents as the union of the minimal coverage sets: L⋆=⋃c∈CLc\(kc⋆\)L^\{\\star\}=\\bigcup\_\{c\\in C\}L^\{\(k^\{\\star\}\_\{c\}\)\}\_\{c\} We then train a multinomial logistic regression classifier using onlyL⋆L^\{\\star\}as input features\. This provides a stricter test of the localization procedure: if the selected latents capture systematic morpho\-syntactic information, they should support multi\-classPoSclassification with limited degradation\. A substantial drop with respect to the fullSAErepresentation would instead suggest that relevant information remains distributed across additional latents\. ### 4\.3Validation on Held\-Out Data \(RQ3\) We evaluate whether the latent groups identified in Section[4\.2](https://arxiv.org/html/2609.29362#S4.SS2)are stable beyond the data used to select them\. The salience and coverage analyses are performed on the GUM training split, wherePoStags may correlate with lexical identity, frequency, position, or local syntactic patterns\. We therefore test the selected groups in two complementary settings: the held\-out GUM test split and the controlled dataset described in Section[3\.1](https://arxiv.org/html/2609.29362#S3.SS1)\. To assess how specific each minimal latent group is to its target category in held out data we construct a cross\-PoSactivation matrix\. For each pair of categories\(c,c′∈C\)\(c,c^\{\\prime\}\\in C\), we compute the probability that at least one latent inL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}is active on tokens whose gold label iscc: Mc,c′=1\|Tc\|∑t∈Tc𝟏\[∃,ℓ∈L\(kc′⋆\):aℓ\(t\)\>0\]M\_\{c,c^\{\\prime\}\}=\\frac\{1\}\{\|T\_\{c\}\|\}\\sum\_\{t\\in T\_\{c\}\}\\mathbf\{1\}\\left\[\\exists,\\ell\\in L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}:a\_\{\\ell\}\(t\)\>0\\right\] This analysis assesses whether the POS\-discriminative SAE latents are category\-specific or shared across the various syntactic categories\. ### 4\.4Controls and Baselines We also provide a set of controls and baselines that address possible confounds and contextualise the SAE results\. First, we perform a control experiment to test whetherlexical identity\(i\.e\., the specific word form, like the preposition*of*\) can act as a confound for the probing experiments\. To estimate how much of the original performance can be attributed to memorisation, we re\-run the probe but assign each word type a random UPOS label\. A probe relying solely on word identity would fit the control labels as well as the real ones\. Second, we provide several baselines for the probing experiment: i\.\) we use raw embeddings \(layer 0\) and raw activations from layer 30 as features for the probe instead of SAE activations; ii\.\) we provide a random latent subset baseline, where we re\-run the compact\-feature classification, but we keep a percentage of theL⋆L\\star\(0, 25, and 50\) and randomly choose the remaining latents\. ## 5Results ### 5\.1RQ1: Recoverability ofPoSInformation Figure 2:One\-vs\-rest probing performance for eachPoS\. F1 scores are computed with 5\-fold cross\-validation on the GUM training split\.The one\-vs\-rest probing results show thatPoSdistinctions are consistently recoverable fromSAEactivations\. As shown in Figure[2](https://arxiv.org/html/2609.29362#S5.F2), the binary classifiers achieve high F1 scores for most categories across the 5\-fold cross\-validation setting\. This indicates that morpho\-syntactic information is explicitly available in the sparse latent space\. Performance, however, is not uniform across tags\. Closed\-class categories and low\-variability labels \(e\.g\., punctuation\), are easier to recover, while more lexically heterogeneous or less frequent categories show lower scores\. Interestingly, nouns and verbs are the best performing openPoS\. This suggests that recoverability is affected both by the linguistic nature of the category and by its support\. Overall, these results answerRQ1positively:SAEactivations contain information that is predictive ofPoScategories\.At the same time, probing performance alone does not reveal how this information is organized within the SAE latent space\. ### 5\.2RQ2:PoS\-related Latent Groups RQ2asks whether thePoSinformation recovered by the probes is localized in restricted regions of theSAElatent space, and at what granularity\. We report below the results of the three experiments introduced in Section[4\.2](https://arxiv.org/html/2609.29362#S4.SS2)\. ##### Feature salience\. The coefficient\-based salience analysis shows that the binary probes do not rely uniformly on the fullSAElatent space\. For eachPoScategory, only a subset of latents receives positive weight, indicating that the classifier uses information concentrated in category\-specific groups of sparse features\. At the same time, these groups are not single\-latent representations:PoSdistinctions are supported by multiple positively contributing latents, consistent with a distributed, but non\-uniform, organization of morpho\-syntactic information\. Figure[3](https://arxiv.org/html/2609.29362#S5.F3)displays the number of non\-zero coefficients for each one\-vs\-rest classifier\. It emergers quite clearly that Open\-classPoShave generally more non\-zero latents, while Closed\-class and Other\-class have markedly less\.555See Appendix[B\.2](https://arxiv.org/html/2609.29362#A2.SS2)for additional details\. Figure 3:Non\-zero coefficients for each one\-vs\-rest classifier; results are color coded byPoSclass \(Open, Closed, Other\)\. ##### Coverage and compactness\. The coverage analysis further quantifies the effective size of these groups\. As shown in Figure[4](https://arxiv.org/html/2609.29362#S5.F4), the cardinalitykc95k\_\{c\}^\{95\}varies acrossPoScategories\. Some tags reach the 95% coverage threshold with a small number of latents, suggesting compact activation patterns\. Others require broader latent groups, indicating that the corresponding distinction is more diffuse or depends on a wider set of lexical and contextual cues\. This variability shows that localization is category\-dependent and cannot be reduced to a one\-latent\-per\-tag mapping\. Figure 4:Number of salient latents required to reach 95% coverage for eachPoScategory\. For each tagcc,kc95k\_\{c\}^\{95\}denotes the smallest number of highest\-coefficient latents needed to activate on at least 95% of gold tokens of that category\. Lower values indicate compact groups\. ##### Compact\-feature classification\. Finally, we test whether the selected latent groups are sufficient for multi\-classPoSprediction\. The compact feature set includes 498 latents, with 12% of them being shared between 2\+PoS\. The classifier trained on it achieves performance comparable to the classifier trained on the fullSAErepresentation \(Figure[5](https://arxiv.org/html/2609.29362#S5.F5)\)\. The selected latents then preserve most of the information needed for multi\-classPoSdiscrimination, and the salience and coverage analyses recover a compact but effective subset of the latent space\.666See App\.[B\.2](https://arxiv.org/html/2609.29362#A2.SS2)for a sensitivity analysis forCCandτ\\tau\. Figure 5:Row\-normalized confusion matrix of the multi\-classPoSclassifier trained on the compact feature setL⋆L^\{\\star\}\. Each cell reports the percentage of tokens of a goldPoScategory predicted as each class\.Overall, these results answerRQ2by showing thatPoSinformation is localized at the level of structured groups of latents\. These groups are compact for some categories and broader for others, but they are sufficient to support both category\-wise coverage and multi\-class classification\. ### 5\.3RQ3: Validation on held\-out data We evaluate whether the latent groups identified in Section[4\.2](https://arxiv.org/html/2609.29362#S4.SS2)remain stable and systematic beyond the data used to select them, addressingRQ3\. We do not train another classifier, but rather test whether the latent groups identified in the discovery setting remain active on unseen instances of the correspondingPoScategories\. Recall that on held\-out data we compute the probability that at least one latent inL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}is active on tokens whose gold label isccFor each pair of categoriesc,c′∈Cc,c^\{\\prime\}\\in C\. Table 1:Pseudo\-multilabel classification results on the controlled dataset\.On the controlled dataset, which includes only a subset ofPoScategories, we evaluate the selected latent groups as a pseudo\-multilabel classification task\. For each token, we encode its goldPoSas a binary vector over the considered categoriesCC\. Predictions are obtained by applying the indicator function defined in Section[5\.3](https://arxiv.org/html/2609.29362#S5.SS3)to each categoryc′∈Cc^\{\\prime\}\\in C\. Table[1](https://arxiv.org/html/2609.29362#S5.T1)reports micro\-averaged precision, recall, and F1 for each template\. The results show limited precision but very high recall in general, with some template\-dependent variation\. We also note the presence of a set of latents that we associate with a “first\-token” concept and that confound the results\. We verified that by including a prefix with a differentPoSto each sentence, all these activations shift onto this newPoS,777We report an example of this in Appendix[B\.3\.1](https://arxiv.org/html/2609.29362#A2.SS3.SSS1)indicating that latent groups are present in controlled data\. For more generalizability over the wholePoSset, we use the held\-out treebank test set\. Figure[6](https://arxiv.org/html/2609.29362#S5.F6)reports a heatmap with the co\-activation patterns\. By construction, the diagonal entriesMc,cM\_\{c,c\}represents the per\-category recall\. The off\-diagonal entriesMc,c′M\_\{c,c^\{\\prime\}\}\(c≠c′c\\neq c^\{\\prime\}\) measure the rate at which latents selected as salient forc′c^\{\\prime\}nevertheless fire on tokens of a different categorycc, i\.e\., a spurious co\-activation rate\. Low off\-diagonal values indicate that the minimal latent groups are disjoint and category\-specific; high values suggest overlapping representations acrossPoScategories\. Results highlight two main aspects: First, the recall is consistently high, with most values≥0\.95\\geq 0\.95, indicating that the latents identified as salient for a givenPoSremain active on unseen tokens drawn from the same underlying distribution; second, we observe a relatively high variability in off\-diagonal scores, indicating overlaps across categories\. We further validate the latter observation by computing aDistinctiveness\(DD\) score for eachPoS, asD\(c\)=Mc,c∑c′∈CMc,c′\\mathrm\{D\}\(c\)=\\frac\{M\_\{c,c\}\}\{\\sum\_\{c^\{\\prime\}\\in C\}M\_\{c,c^\{\\prime\}\}\}\. Intuitively,D\(c\)=1\\mathrm\{D\}\(c\)=1indicates that the latents inL\(kc⋆\)L^\{\(k^\{\\star\}\_\{c\}\)\}fire exclusively oncctokens; a value of1/\|C\|1/\|C\|\(0\.06 in our case\) corresponds to the chance baseline under uniform co\-activation\. We observe aD\\mathrm\{D\}μ=0\.27\\mu=0\.27\(σ=0\.07\\sigma=0\.07\)\.888See Appendix[B\.3](https://arxiv.org/html/2609.29362#A2.SS3.SSS0.Px2)for the per\-ccresults\. Figure 6:Co\-activations of selected latent groups on the held\-out treebank test set\. Diagonal values correspond to recall for eachPoS; off\-diagonal values correspond to false positive rates on otherPoScategories\.Overall, the results on both held\-out datasets answerRQ3affirmatively for stability, while qualifying the systematicity claim: the identified latent groups remain consistently active on unseen tokens of their target category, but are only partially category\-specific\. More generally, this suggests that,despite the transition to a sparse representation through theSAE, a one\-to\-one correspondence between linguistic categories and groups of latents holds only to a limited extent\. Figure 7:Active L⋆\\starlatents per relevantPoSin each template variant\. Each square is one latent that is active on at least 25% of template’s examples\. Color saturation indicates activation percentage \(darker==more active\)\. ### 5\.4Controls and Baselines The control probe with random labels obtains 0\.54 Accuracy/0\.42 Macro F1, against 0\.88/0\.97 respectively for the real probe\. The confound is not negligible, but the 34–37 point performance gap still supports the conclusion that POS recoverability is not simply reducible to lexical memorisation, and thatlexical identity plays a minor role\. This is consistent with evidence provided in Sec\.[5\.5](https://arxiv.org/html/2609.29362#S5.SS5)\. TheSAE\-based probe has comparable performances with both the raw residual stream probes, both at the embedding layer and at layer 30 \(Table[2](https://arxiv.org/html/2609.29362#S5.T2)\) using∼\\sim8 times less features\. However, note that we do not claim superiority of SAEs as a probing tool\. Rather, we claim and show that \(i\) POS information is recoverable from subsets of SAE latents and that \(ii\) the SAE’s contribution is decomposition and localisation, which the dense probe cannot provide as easily\. Table 2:Performances using different probes\.Finally, we see thatL⋆L^\{\\star\}features arePoSrelevant\.In fact, randomly replacing features from theL⋆L^\{\\star\}set drastically reduces performances\. Table[3](https://arxiv.org/html/2609.29362#S5.T3)shows the results at 0, 25, 50 and 100% overlap withL⋆L^\{\\star\}\. This further demonstrates thatL⋆L^\{\\star\}latents arePoSrelevant\. Table 3:Accuracy and Macro\-F1 across overlap levels, comparing cross\-validation and train/test evaluation\. ### 5\.5Linguistic analysis of SAE activations The results confirm a systematic association betweenSAEactivations andPoSinformation\. Overall, four main patterns emerge\. Closed classes are more stable and compact\.On the training set \(Figure[2](https://arxiv.org/html/2609.29362#S5.F2)\) and on the held\-out treebank test set \(Figure[6](https://arxiv.org/html/2609.29362#S5.F6)\), Closed\-classes consistently achieve higher Recall and F1\-score\. Note that these categories also require fewer latents in the coverage analysis \(Figure[4](https://arxiv.org/html/2609.29362#S5.F4)and Appendix[B\.2](https://arxiv.org/html/2609.29362#A2.SS2), Figure[3](https://arxiv.org/html/2609.29362#S5.F3)\)\. By contrast,xandsymexhibit the highest error rates in both experiments, likely due to their low support in the training data \(Figure[2](https://arxiv.org/html/2609.29362#S5.F2)\)\. The compactness gradient we observe co\-varies with the size and formal variability of each category’s type inventory\. A category with a handful of invariant word forms can be covered by a small latent group with no category\-level abstraction being involved, since a group of form\-specific latents suffices\. The controlled data support this reading: DET collapses to a single activation value where the determiner is invariably the \(Figure[18](https://arxiv.org/html/2609.29362#A2.F18)\), and acquires structure only where the a/an versus bare\-plural alternation introduces formal variation \(Figure[7](https://arxiv.org/html/2609.29362#S5.F7)\)\. The compactness ordering should therefore be read primarily as a gradient in lexical variability rather than as direct evidence of graded abstraction\. Open classes show broader and less selective activation patterns\.Spurious co\-activations are more frequent for openPoSclasses, where the same lemma or morphologically related forms can serve different functions depending on context\. For instance,adjtokens are mainly confused withnoun,adv, andpropn, reflecting attributive noun uses, adjective–adverb overlap, and nominal modification patterns\. Similarly,advshows diffuse co\-activation withadp,noun,adj,sconj, andverb, suggesting that some adverb\-associated latents capture positional or contextual cues rather than adverbial function alone\.propnandnounalso co\-activate, consistent with their shared nominal distribution, whileverbshows overlap withnounin homograph pairs e\.g\.to drink/the drink\. Some off\-diagonal patterns reflect annotation and lexical overlap, as well as syntagmatic properties\.The strongest non\-target activations are not random\.intjshows high spurious activation rates across several categories, plausibly due to annotation conventions that assign heterogeneous forms such aslike,well, orGodtointjin pragmatic contexts\. Similarly,sconjco\-activates withadpandverb, reflecting lexical overlap between subordinating conjunctions and prepositions in English \(e\.g\.,by,after,since\) and broader positional regularities\. These patterns suggest that the selected latents do not encode purely abstractPoSfunction, but also respond to surface form, lemma sharing and orthographic cues\. Moreover, thePosemerging out of the latent space are also defined in terms of their syntactic contexts\. For instance, theadjlatents strongly co\-activate withnounreflecting the nature of adjectives as nominal modifiers\. Controlled examples confirm additive and category\-sensitive activations\.Turning to the controlled dataset, Figure[7](https://arxiv.org/html/2609.29362#S5.F7)confirms that, as new tokens are introduced into the sentence, their associated latents activate consistently within the latent groups characteristic of their PoS, suggesting thatPoS\-specific latent activations are largely additive across tokens\.For example, addingloyaltoThere is a dogtriggers the activation of more latents associated withadj\. We observe that some latents of relatedPoSes are already present even without the corresponding words \(e\.g\.,adjfornounandadvforverb\), but the number of active ones corresponding to that category consistently grows when the word is included in the template\. These findings show thatPoS\-related latent groups are stable and systematic, but not category\-exclusive\. Closed classes tend to yield compact and selective representations, whereas open classes involve broader latent groups that also capture lexical, morphological, and contextual regularities\. ## 6Conclusion We usedPoScategories as a controlled testbed to study how morpho\-syntactic information is organized inSAEs\. We show thatPoSdistinctions are consistently recoverable from sparse activations, and that each category is supported by a compact group of sparse features, whose size varies with the linguistic nature of the category\. For Open\-classPoS, these groups are more diffuse, with a larger number of active latents, while Closed\-class ones have fewer active latents\. A small union of these category\-specific latents preserves strong multi\-class classification performance, and the selected groups remain stable on held\-out treebank data and largely additive on controlled examples\. Our results reveal thatPoSare internally represented in LLMs as emerging sets of localizable but distributed features in latent SAE space\. Moreover, analyses suggest that interpretability claims at the latent level should be evaluated against theoretically grounded category inventories rather than top\-activating examples alone\. At the same time, cross\-category co\-activations show that the identified latents partly track lexical, positional, and annotation\-driven regularities, motivating future work on richer linguistic levels and typologically diverse languages\. ## Limitations Our study focuses onPoScategories as a controlled morpho\-syntactic test case\. While this choice allows us to rely on exhaustive and independently annotated labels,PoStags capture only one level of linguistic abstraction\. Future work should extend the analysis to finer\-grained morphological features, dependency relations, semantic roles, and discourse\-level phenomena, where latent organization may be recoverable in different ways\. We also analyze a single base language model,LLaMA\-3\-8B, and one publicly availableSAE\. The observed patterns may depend on the underlying model, the layer from which activations are extracted, theSAEtraining procedure, and the sparsity regime\. Comparing multiple models, layers, andSAEvariants would be necessary to assess how general these findings are\. Our localization procedure relies on linear probing coefficients as a salience signal\. Although the use ofℓ1\\ell\_\{1\}\-regularization encourages sparse and interpretable solutions, probe coefficients should not be interpreted as direct causal evidence\. The identified latents are predictive ofPoScategories, but further causal interventions would be needed to establish whether they are used by the model for morpho\-syntactic processing\. Finally, our controlled dataset is intentionally small and targets a restricted set of constructions, mainly involving nouns and verbs\. It is useful for validating whether selected latent groups remain active in simple and independently constructed contexts, but it does not cover the full syntactic and lexical variability of English\. Broader controlled datasets would allow a more systematic evaluation of how lexical ambiguity, word order, morphology, and sentence complexity affectPoS\-related latent activations\. ## Acknowledgments This work has been supported by i\.\) the PNRR MUR project[PE0000013\-FAIR](https://fondazione-fair.it/)\(Spoke 1\), funded by the European Commission under the NextGeneration EU programme; ii\.\) the EU EIC project[EMERGE](https://eic-emerge.eu/)\(Grant No\. 101070918\); and iii\.\) The PNRR MUR project FAIR TT\_02 “Innovare la sorveglianza automatizzata delle infezioni del sito chirurgico tramite modelli di elaborazione del linguaggio naturale”\. ## References - Belinkov \(2022\)Y\. BelinkovProbing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Link](https://aclanthology.org/2022.cl-1.7/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p5.1),[§1](https://arxiv.org/html/2609.29362#S1.p6.1),[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Brickenet al\.\(2023\)T\. Bricken, A\. Templeton, J\. Batson, B\. Chen, A\. Jermyn, T\. Conerly, N\. Turner, C\. Anil, C\. Denison, A\. Askell, R\. Lasenby, Y\. Wu, S\. Kravec, N\. Schiefer, T\. Maxwell, N\. Joseph, Z\. Hatfield\-Dodds, A\. Tamkin, K\. Nguyen, B\. McLean, J\. E\. Burke, T\. Hume, S\. Carter, T\. Henighan, and C\. OlahTowards monosemanticity: decomposing language models with dictionary learning\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2023/monosemantic-features/index.html)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p1.1),[§1](https://arxiv.org/html/2609.29362#S1.p2.1),[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Conneauet al\.\(2018\)A\. Conneau, G\. Kruszewski, G\. Lample, L\. Barrault, and M\. BaroniWhat you can cram into a single $&\!\#\* vector: probing sentence embeddings for linguistic properties\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 2126–2136\.External Links:[Link](https://aclanthology.org/P18-1198/),[Document](https://dx.doi.org/10.18653/v1/P18-1198)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Cunninghamet al\.\(2023\)H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. SharkeySparse autoencoders find highly interpretable features in language models\.External Links:2309\.08600,[Link](https://arxiv.org/abs/2309.08600)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p1.1),[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Dupre la Tour and Mossing \(2025\)T\. Dupre la Tour and D\. MossingDebugging misaligned completions with sparse\-autoencoder latent attribution\.Note:OpenAI Alignment Research BlogExternal Links:[Link](https://alignment.openai.com/sae-latent-attribution/)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p3.1)\. - Elhageet al\.\(2022\)N\. Elhage, T\. Hume, C\. Olsson, N\. Schiefer, T\. Henighan, S\. Kravec, Z\. Hatfield\-Dodds, R\. Lasenby, D\. Drain, C\. Chen, R\. Grosse, S\. McCandlish, J\. Kaplan, D\. Amodei, M\. Wattenberg, and C\. OlahToy models of superposition\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2022/toy_model/index.html)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p1.1),[§1](https://arxiv.org/html/2609.29362#S1.p2.1)\. - Engelset al\.\(2025\)J\. Engels, E\. J\. Michaud, I\. Liao, W\. Gurnee, and M\. TegmarkNot all language model features are one\-dimensionally linear\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d63a4AM4hb)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p3.1),[§2](https://arxiv.org/html/2609.29362#S2.p4.1)\. - Fraser\-Talienteet al\.\(2026\)K\. Fraser\-Taliente, S\. Kantamneni, E\. Ong, D\. Mossing, C\. Lu, P\. C\. Bogdan, E\. Ameisen, J\. Chen, D\. Kishylau, A\. Pearce, J\. Tarng, A\. Wu, J\. Wu, Y\. Zhang, D\. M\. Ziegler, E\. Hubinger, J\. Batson, J\. Lindsey, S\. Zimmerman, and S\. MarksNatural language autoencoders produce unsupervised explanations of llm activations\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2026/nla/index.html)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p3.1)\. - Gaoet al\.\(2025\)L\. Gao, T\. D\. la Tour, H\. Tillman, G\. Goh, R\. Troll, A\. Radford, I\. Sutskever, J\. Leike, and J\. WuScaling and evaluating sparse autoencoders\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=tcsZt9ZNKD)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.29362#S3.SS2.SSS0.Px1.p1.1)\. - Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§A\.3](https://arxiv.org/html/2609.29362#A1.SS3.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.29362#S1.p5.1)\. - Heet al\.\(2024\)Z\. He, W\. Shu, X\. Ge, L\. Chen, J\. Wang, Y\. Zhou, F\. Liu, Q\. Guo, X\. Huang, Z\. Wu,et al\.Llama scope: extracting millions of features from llama\-3\.1\-8b with sparse autoencoders\.arXiv preprint arXiv:2410\.20526\.Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Hewitt and Liang \(2019\)J\. Hewitt and P\. LiangDesigning and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2733–2743\.External Links:[Link](https://aclanthology.org/D19-1275/),[Document](https://dx.doi.org/10.18653/v1/D19-1275)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p6.1),[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Hewitt and Manning \(2019\)J\. Hewitt and C\. D\. ManningA structural probe for finding syntax in word representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4129–4138\.External Links:[Link](https://aclanthology.org/N19-1419/),[Document](https://dx.doi.org/10.18653/v1/N19-1419)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Kantamneniet al\.\(2025\)S\. Kantamneni, J\. Engels, S\. Rajamanoharan, M\. Tegmark, and N\. NandaAre sparse autoencoders useful? a case study in sparse probing\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=rNfzT8YkgO)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p3.1),[§2](https://arxiv.org/html/2609.29362#S2.p3.1)\. - Karvonenet al\.\(2025\)A\. Karvonen, C\. Rager, J\. Lin, C\. Tigges, J\. I\. Bloom, D\. Chanin, Y\. Lau, E\. Farrell, C\. S\. McDougall, K\. Ayonrinde, D\. Till, M\. Wearden, A\. Conmy, S\. Marks, and N\. NandaSAEBench: a comprehensive benchmark for sparse autoencoders in language model interpretability\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=qrU3yNfX0d)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p3.1),[§2](https://arxiv.org/html/2609.29362#S2.p3.1)\. - Lieberumet al\.\(2024\)T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramar, A\. Dragan, R\. Shah, and N\. NandaGemma scope: open sparse autoencoders everywhere all at once on gemma 2\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, N\. Kim, J\. Jumelet, H\. Mohebbi, A\. Mueller, and H\. Chen \(Eds\.\),Miami, Florida, US,pp\. 278–300\.External Links:[Link](https://aclanthology.org/2024.blackboxnlp-1.19/),[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.19)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Markset al\.\(2025\)S\. Marks, C\. Rager, E\. J\. Michaud, Y\. Belinkov, D\. Bau, and A\. MuellerSparse feature circuits: discovering and editing interpretable causal graphs in language models\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=I4e82CIDxv)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p4.1)\. - Pimentelet al\.\(2020\)T\. Pimentel, J\. Valvoda, R\. H\. Maudslay, R\. Zmigrod, A\. Williams, and R\. CotterellInformation\-theoretic probing for linguistic structure\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4609–4622\.External Links:[Link](https://aclanthology.org/2020.acl-main.420/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.420)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p6.1),[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Rogerset al\.\(2020\)A\. Rogers, O\. Kovaleva, and A\. RumshiskyA primer in BERTology: what we know about how BERT works\.Transactions of the Association for Computational Linguistics8,pp\. 842–866\.External Links:[Link](https://aclanthology.org/2020.tacl-1.54/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00349)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Sharkeyet al\.\(2025\)L\. Sharkey, B\. Chughtai, J\. Batson, J\. Lindsey, J\. Wu, L\. Bushnaq, N\. Goldowsky\-Dill, S\. Heimersheim, A\. Ortega, J\. Bloom, S\. Biderman, A\. Garriga\-Alonso, A\. Conmy, N\. Nanda, J\. Rumbelow, M\. Wattenberg, N\. Schoots, J\. Miller, E\. J\. Michaud, S\. Casper, M\. Tegmark, W\. Saunders, D\. Bau, E\. Todd, A\. Geiger, M\. Geva, J\. Hoogland, D\. Murfet, and T\. McGrathOpen problems in mechanistic interpretability\.External Links:2501\.16496,[Link](https://arxiv.org/abs/2501.16496)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Templetonet al\.\(2024\)A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, T\. Bricken, B\. Chen, A\. Pearce, C\. Citro, E\. Ameisen, A\. Jones, H\. Cunningham, N\. L\. Turner, C\. McDougall, M\. MacDiarmid, C\. D\. Freeman, T\. R\. Sumers, E\. Rees, J\. Batson, A\. Jermyn, S\. Carter, C\. Olah, and T\. HenighanScaling monosemanticity: extracting interpretable features from claude 3 sonnet\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html)Cited by:[§1](https://arxiv.org/html/2609.29362#S1.p1.1),[§1](https://arxiv.org/html/2609.29362#S1.p2.1),[§2](https://arxiv.org/html/2609.29362#S2.p2.1)\. - Tenneyet al\.\(2019a\)I\. Tenney, D\. Das, and E\. PavlickBERT rediscovers the classical NLP pipeline\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4593–4601\.External Links:[Link](https://aclanthology.org/P19-1452/),[Document](https://dx.doi.org/10.18653/v1/P19-1452)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Tenneyet al\.\(2019b\)I\. Tenney, P\. Xia, B\. Chen, A\. Wang, A\. Poliak, R\. T\. McCoy, N\. Kim, B\. V\. Durme, S\. R\. Bowman, D\. Das, and E\. PavlickWhat do you learn from context? probing for sentence structure in contextualized word representations\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SJzSgnRcKX)Cited by:[§2](https://arxiv.org/html/2609.29362#S2.p1.1)\. - Zeldes \(2017\)A\. ZeldesThe GUM corpus: creating multilayer resources in the classroom\.Language Resources and Evaluation51\(3\),pp\. 581–612\.External Links:[Document](https://dx.doi.org/http%3A//dx.doi.org/10.1007/s10579-016-9343-x)Cited by:[§A\.3](https://arxiv.org/html/2609.29362#A1.SS3.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.29362#S1.p5.1),[§3\.1](https://arxiv.org/html/2609.29362#S3.SS1.p2.1)\. ## Appendix ## Appendix AFurther Details on Implementation Here we provide some further details on several of the implementation choices, including the rationale for selecting the target layer and experimental setup\. ### A\.1Layer Selection We extract hidden state activations of the LLM from the residual stream of the last layer before the output\. The residual stream hookpoint is set after the MLP layer\. This means that the activations we take are very close to the output\. We acknowledge that the choice of the layer might drastically affect the results\. Thus, to verify the stability of our findings, we replicate some of the experiments across a sample of layers\. As for RQ1 \(Recoverability ofPoSinformation, see Section[4\.1](https://arxiv.org/html/2609.29362#S4.SS1)\), we extract hidden representations from the residual stream after the MLP of layerxx, and encode it with the corresponding SAE\. We report in Figure[8](https://arxiv.org/html/2609.29362#A1.F8)the F1\-Scores of the binary probing classifiers across all layers\. The results show that performances are generally stable across layers, with minor variations\. Other classPoSes are the ones showing higher variability, exceptPunct\. We also note a systematic slight dip in performances in the last few layers across allPoSclasses\. This may be attributable to the vicinity to the unembed layer\. Figure 8:Per\-PoSone\-vs\-rest classifier performances across layers\.As for RQ2 \(Organization in Latent Space, see Section[4\.2](https://arxiv.org/html/2609.29362#S4.SS2)and RQ3 \(Validation on held\-out data, see Section[4\.3](https://arxiv.org/html/2609.29362#S4.SS3)\), we report result for Salience \(latent activations count\), Coverage \(, Compact\-feature classification, and distinctiveness for Layers 2 and 15 \(beginning and middle of the model stack\)\. First,the number of non zero activations per POS increases modestly across layers for most POS\(e\.g\., NOUN: 9024→\\rightarrow9935→\\rightarrow10043; ADP: 4955→\\rightarrow5742→\\rightarrow6511\),but stays within the same order of magnitude: open\-class POS in the low\-thousands\-to\-10k range, closed\-class in the hundreds\-to\-low\-thousands, Other in the hundreds\. Pooled mean rises only 19% from layer 2 to layer 30 \(3105→\\rightarrow3697\)\. This reflects gradual densification with depth, not a qualitative reorganization\. Second,L∗L^\{\*\}remains stable stable across depth: 530 \(L2\)→\\rightarrow501 \(L15\)→\\rightarrow498 \(L30\)\. Coverage efficiency \(L⋆L^\{\\star\}over the sum ofkc⋆k^\{\\star\}\_\{c\}\) improving slightly \(84%→\\rightarrow81%→\\rightarrow87%\)\. Per\-category values do redistribute a bit\. For example, ADP falls \(71→\\rightarrow35→\\rightarrow16\), SCONJ has a non monotonic trend \(7→\\rightarrow51→\\rightarrow17\), NOUN rises \(12→\\rightarrow38→\\rightarrow34\), PROPN rises \(22→\\rightarrow47→\\rightarrow60\)\. However, all values remain in the same single\-to\-double\-digit\-tens range at every layer, with no category jumping an order of magnitude\. Third,compact\-feature classification Macro F1 increases only slightly\(0\.77→\\rightarrow0\.79→\\rightarrow0\.81\), and accuracy stays mostly flat \(0\.87→\\rightarrow0\.88→\\rightarrow0\.87\)\. This may be an indication that later layers encode more linearly separable POS distinctions for minority/harder classes\. Finally,distinctiveness remains stable across layers\. We report mean and standard deviation: 0\.15±\\pm0\.03 \(L2\)→\\rightarrow0\.32±\\pm0\.09 \(L15\)→\\rightarrow0\.27±\\pm0\.07 \(L30\)\. These results suggest that the layer choice do not drastically affect the degree to whichPoSes are encoded into the SAE latents, and that the findings of the paper should hold across the whole model\. Results on additional layers at the beginning and middle of the model stack show slight variation, but are also a clear indication that the conclusions from the paper hold also for layers other than the one analyzed\. Individual POS categories reshuffle which latents they rely on across depth, but the total representational budget and per\-category magnitudes remain stable\. This supports our claim of layer\-consistent POS structure with depth\-wise redistribution rather than qualitative change\. ### A\.2Experimental setup All experiments were conducted using a GPU node equipped with 8 A100 80GB GPUs\. Given the model sizes, only one GPU was sufficient to extract hidden representations from a layer and encode it with its corresponding SAE\. The process takes roughly 0\.16 GPU hours per layer\. The probing classifiers were implemented using SciKit\-Learn\. The library does not leverage GPUs, but allows parallelization across CPU cores\. A single 5\-fold cross\-validation training/test run on the GUM Treebank training set requires roughly 4 hours using all CPU cores available\. ### A\.3Artifacts and Intended Use We use three existing artifacts, employed consistently with their intended use and license\. ##### LLaMA\-3\-8B \([Grattafiori et al\., 2024](https://arxiv.org/html/2609.29362#bib.bib21)\)is released by Meta under the Llama 3 Community License, which permits research and academic use\. We use the model exclusively for interpretability analysis, extracting hidden representations without fine\-tuning or redistribution of model weights\. ##### EleutherAI SAE The checkpoint we use fromEleutherAI/sae\-llama\-3\-8b\-32xis a pretrained Sparse Autoencoder released on HuggingFace by EleutherAI for interpretability research on LLaMA\-3\-8B, Our use directly aligns with its intended purpose\. ##### The GUM treebank \([Zeldes, 2017](https://arxiv.org/html/2609.29362#bib.bib1)\)is distributed under Creative Commons licenses \(primarily CC BY 4\.0, with some subcorpora under more restrictive terms\) for research and educational use in computational linguistics\. We use the treebank’s text and Universal Dependencies annotations for probing and evaluation, which is consistent with its intended use\. ##### Artifacts produced\. We release999[https://github\.com/colinglab/pos\-sae\-latents](https://github.com/colinglab/pos-sae-latents), [https://huggingface\.co/datasets/colinglab/UD\_English\-GUM\-Latents\_Meta\-Llama\-3\-8B\_L30](https://huggingface.co/datasets/colinglab/UD_English-GUM-Latents_Meta-Llama-3-8B_L30)the token\-level SAE activations aligned to UD POS tags, the controlled evaluation dataset \(180 items\), and the code for the probing and salience pipeline\. These artifacts are released for research use only, consistent with the access conditions of the underlying resources\. To respect the per\-source licensing of GUM, activation files are keyed to token indices in the original GUM release rather than redistributing the source text\. ## Appendix BAdditional Results Here we present additional results from our experiments across our three RQs\. ### B\.1RQ1: Recoverability ofPoSinformation Figures[9](https://arxiv.org/html/2609.29362#A2.F9),[10](https://arxiv.org/html/2609.29362#A2.F10), and[11](https://arxiv.org/html/2609.29362#A2.F11)show, for eachPoSclass, the top 10 latents ranked by coefficient in the classifier\. Closed\-class categories show the most peaked coefficient distributions, with a single latent dominating indet\(75751,≈5\.5\\approx\\\!5\.5\),CCONJ\(34665,≈4\.3\\approx\\\!4\.3\) andPRON\(6631,≈2\.7\\approx\\\!2\.7\)\. Open\-class categories are flatter: onlyVERBshows a clearly dominant feature \(94414,≈2\.2\\approx\\\!2\.2\);ADJ\(0\.80\),PROPN\(0\.77\) andADV\(0\.73\) have no single salient latent\.PUNCTis the best\-classified category overall, driven by latent 117946 \(≈5\.5\\approx\\\!5\.5\)\. Moreover, Several latents recur across closed\-class categories \(e\.g\., 72975 inAUX/CCONJ/PART; 116300 inCCONJ/NUM/PRON; 86665 inDET/PRON/PUNCT\), which may indicate poly\-functional latents\. Figure 9:Per\-PoSone\-vs\-rest classifier top coefficients\.OpenclassPoSes\.Figure 10:Per\-PoSone\-vs\-rest classifier top coefficients\.ClosedclassPoSes\.Figure 11:Per\-PoSone\-vs\-rest classifier top coefficients\.OtherclassPoSes\. ### B\.2RQ2:PoS\-related Latent Groups Here, we present results related to RQ2\. ##### Feature Salience\. As for the feature salience, Table[4](https://arxiv.org/html/2609.29362#A2.T4)shows the number of non\-zero latents for each one\-vs\-rest classifier\. See also Figure[3](https://arxiv.org/html/2609.29362#S5.F3)in the main paper\. From the Table and Figure, it emerges quite clearly that Open\-classPoShave generally more non\-zero latents, while Closed\-class and Other\-class have markedly less\. The two main exceptions areINTJfor the Open\-class, which has very few, andADPfor the Closed\-class, which is more akin to open ones\. ForADP, this may be attributable to the fact that words that function as adpositions may also be used to mark adverbial clauses\. As forINTJ, they typically express an emotional reaction and are not syntactically related to other accompanying expressions; moreover, their support in the GUM treebank is very low\. Both factors could play a role in poor classification performances \(See Section[5\.2](https://arxiv.org/html/2609.29362#S5.SS2)and below\) and limited number of latents active as features during classification\. Table[5](https://arxiv.org/html/2609.29362#A2.T5)reports mean and standard deviation counts for non\-zero coefficients\. Again, we observe that Other class have the least number of non\-zero coefficients on average, and Open class has the highest number and highest variability\. Table 4:Non\-zero latents and fraction of latent space by POS tag\.Table 5:Number of latents with non\-zero coefficients for classification for each one\-vs\-rest classifier, aggregated byPoSGroup\. Classifiers forOpen\-classPoStags have the highest number of non\-zero coefficients\. ##### Coverage and Compactness\. In Table[6](https://arxiv.org/html/2609.29362#A2.T6)we report, for eachPoS, the ratio between the coverage forL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}and the number of latents with non\-zero coefficients forc′c^\{\\prime\}\. Intuitively, this represent the proportion of latents needed to reach coverage for aPoSwith respect to all salient latents for thePoS\. ClassPOSCoverage over Salience \(%\)Openadj0\.91adv1\.38intj3\.76noun0\.34propn1\.14verb0\.70Closedadp0\.25aux1\.06cconj0\.62det0\.42num1\.20part0\.47pron0\.76sconj0\.45Otherpunct0\.33sym5\.65x8\.95Table 6:Coverage over number of salient features per UPOS, grouped by class\. ##### Compact\-feature classification\. Table[7](https://arxiv.org/html/2609.29362#A2.T7)reports per\-PoSperformances of the multiclass classifier trained with the compact feature setL∗L^\{\*\}\. precisionrecallf1\-scoresupportADJ0\.850\.840\.8511489ADP0\.930\.880\.9116655ADV0\.840\.790\.818556AUX0\.930\.910\.929682CCONJ0\.960\.970\.965854DET0\.950\.930\.9414307INTJ0\.360\.890\.521860NOUN0\.920\.870\.9029289NUM0\.880\.930\.903368PART0\.820\.950\.884314PRON0\.950\.930\.9415197PROPN0\.830\.830\.8310144PUNCT0\.990\.940\.9624563SCONJ0\.660\.860\.752878SYM0\.330\.960\.49281VERB0\.940\.860\.9018642X0\.140\.880\.24331accuracy0\.89177410macro avg0\.780\.890\.81177410weighted avg0\.910\.890\.90177410Table 7:Multiclass classifier performances\.Figure[12](https://arxiv.org/html/2609.29362#A2.F12)shows the heatmap of coefficient association to eachPoSclass in the multilabel classification experiment\. Coefficient are sorted for importance acrossPoSclasses\. Figure 12:Coefficients importance heatmap for eachPoSin the mutliclass classifier trained with 5\-fold cross\-validation on the GUM Training set\.Figure[13](https://arxiv.org/html/2609.29362#A2.F13)provides a sensitivity analysis of the performances of the classifier with respect to values ofCCandτ\\tau\. We report mean Macro\-F1 score and Accuracy in the cross validation setting\. We also provide standard deviation in the form of error bars\. From the plot, it clearly emerges that the classification results are not particularly sensitive neither to theτ\\tauthreshold nor theCCvalue\. As for theτ\\tau, performances increase monotonically, but with a difference of∼\\sim5 points using 3x less features\. As for theCCvalues, performances remain almost identical, and the standard deviation is near zero, indicating no meaningful differences\. Figure 13:Parameter sweep forCCandτ\\taufor the compact feature classification task\. ### B\.3RQ3: Validation on held\-out data ##### Density of Activations perPoS\. In Figure[14](https://arxiv.org/html/2609.29362#A2.F14)we report a KDE plot representing density of number of activations fromL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}overPoSwith tagc′c^\{\\prime\}in the Treebank test set, for allc′∈Cc^\{\\prime\}\\in C\. The Figure shows that in most cases at least 1 to 4% of latents inL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}fire on all tokens ofc′c^\{\\prime\}\. Figure 14:Density plot: number of activations fromL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}overPoSwith tagc′c^\{\\prime\}in the Treebank test set, for allc′∈Cc^\{\\prime\}\\in C\.We further provide indications that most tokens associated with categoryc′c^\{\\prime\}fire latents inL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}\. Over the whole treebank test set, we count the number of cases in which no latents inL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}fire over a token of categoryc′c^\{\\prime\}\. Table[8](https://arxiv.org/html/2609.29362#A2.T8)shows the fraction of tokens with noL\(kc′⋆\)L^\{\(k^\{\\star\}\_\{c^\{\\prime\}\}\)\}for eachc′∈Cc^\{\\prime\}\\in C\. ##### Distinctiveness\. Table[8](https://arxiv.org/html/2609.29362#A2.T8)reports per\-ccresults ofMc,cM\_\{c,c\}and distinctivenessD\(c\)D\(c\)\. We observe thatINTJandXhave the least distinctive activations\. In this case we do not see a clearPoS\-class based trend with regards to distinctiveness\. Table 8:Zero\-activation fraction,MccM\_\{cc\}, and distinctivenessD\(c\)D\(c\)per UPOS category\. ##### PerPosActivation distribution in the controlled dataset\. Figures[15](https://arxiv.org/html/2609.29362#A2.F15)through[18](https://arxiv.org/html/2609.29362#A2.F18)report the KDE distributions of active\-latent percentages perPoStag for the remaining three construction types in the controlled dataset\. In each figure,blue,red, andgreencorrespond to sentences of increasing syntactic complexity: the minimal construction \(e\.g\., \[PRON\] \+ \[AUX\] \+ \[DET\] \+ \[NOUN\]\), the extension with an adjective\[ADJ\]or an\[ADV\]for the verb template, and the further addition of punctuation\[PUNCT\], respectively; dashed vertical lines indicate cases where only a single value is available for a givenPoStag and sentence type\. The distributions for mostPoStags are highly stable across sentence types, with the three curves largely overlapping\. A consistent exception is thenountag: in both transitive constructions \(I have/hadandI see/saw\), theredandgreencurves, corresponding to sentences containing an additional adjective, show a slight shift in the activation distribution relative to thebluecurve\. This suggests that the presence of an adjacent adjective marginally affects noun\-associated latent activations, consistent with co\-activation patterns discussed in Section[5\.5](https://arxiv.org/html/2609.29362#S5.SS5); similarly,auxin Figure[16](https://arxiv.org/html/2609.29362#A2.F16)displays two distinct peaks, reflecting the alternation between the presenthaveand pasthadforms across sentences, confirming latents’ sensitivity to surface form inPoSclasses with relatively low token variability\. By contrast, theverbtag in Figure[17](https://arxiv.org/html/2609.29362#A2.F17)shows only a minor shift across sentence types, suggesting that verb\-associated latents are less sensitive to the surrounding syntactic context than noun\-associated ones\. Figure 15:KDE of active\-latent percentages perPoStag in theI\+ \[VERB\] templates\.Figure 16:KDE of active\-latent percentages perPoStag in theI have/hadtemplates\.Figure 17:KDE of active\-latent percentages perPoStag in theI\+ \[VERB\] templates\.Figure 18:KDE of active\-latent percentages perPoStag in theThere is/aretemplates\. #### B\.3\.1First\-token activations in the controlled dataset We report an example of first\-token activations confounds on the controlled dataset\. All tested cases show the same behavior, but we report only one example for brevity\. We do the following: we prepend “1\.” to all templates, e\.g\., “1\. I saw the cute cat”, recompute activations, and compare pre\-vs\-post addition of theNUM\+PUNCT\. In Figure[19](https://arxiv.org/html/2609.29362#A2.F19)we report the results for the “I see/saw a \[ADJ\] \[NOUN\]”\. We observe that all latents that fire onPRONin the original sentence \(the first token “I”\) shift to the first tokenNUM\. Figure 19:Demonstration of the behavior on first\-token activations\. We show that several activations shift fromPRONin the top sentence \(the first token “I”\) to the first tokenNUMin the bottom sentence\.
相似文章
词性的语义空间
本文使用word2vec嵌入和神经网络来分析词性分类中的固有模糊性,创建一个三维语义空间来可视化典型词和语言类别之间的边界。
WriteSAE:面向循环状态的稀疏自编码器
WriteSAE 引入了第一个稀疏自编码器,能够分解状态空间模型和混合循环语言模型中的矩阵缓存写入,相比现有方法实现了更优的令牌级干预。
单令牌稀疏自编码器特征是否具有因果必要性?层深度与SAE家族效应
本文研究了单令牌稀疏自编码器特征对于模型输出是否具有因果必要性,发现因果角色因SAE家族而异,并且对训练方法(而非激活函数或规模)敏感。
在应稀疏分解时稀疏分解,在应密集吸收时勿密集吸收
论文假设语言模型激活包含一个低秩密集分量,该分量被稀疏自编码器(SAEs)低效表示。通过添加一个线性瓶颈来吸收密集结构,作者减少了密集潜变量,并改进了在Gemma-2-2B上的稀疏探针性能。
用于特征发现与长上下文归因的回合平均SAEs
本文介绍了基于回合平均的稀疏自编码器(SAEs),该编码器基于对话回合的平均激活值运行,能够实现长上下文的高效特征发现与归因图。此外,本文还提出了一种嵌套架构,用于与每个词元的特征联合训练。