Selecting The Most Informative Tokens in Natural Language Autoencoders
Summary
The paper studies which token positions in natural language autoencoders are most informative to explain, finding that chat-structure-based ranking selects relevant explanations better than computational signals, and that explaining only 5% of positions retains nearly all audit success on prompt injection and concealment tasks.
View Cached Full Text
Cached at: 09/30/26, 12:16 PM
Paper page - Selecting The Most Informative Tokens in Natural Language Autoencoders
Source: https://huggingface.co/papers/2609.37040
Abstract
Naturallanguageautoencoderstranslatealanguagemodel’sinternalactivationsintoreadableexplanations.Explainingeverytokenpositioniscostly.Whichpositionsshouldanauditorinspecttounderstandapotentialthreat?Westudythisquestionacross4.7millionexplanationsonpromptinjectionandconcealment.Wecomparesignalsfrommodelcomputationwitharankertrainedonlyonchatstructure.Chatstructureusuallyselectsmorerelevantexplanationsthanthecomputationalsignals,withoutrequiringamodelforwardpassforpositionselection.Onthreeoffourdatasets,explainingjust5%ofpositionsretainsnearlyallofthesuccessratefromexplainingeveryposition,wheresuccessmeansobtaininganexplanationaboutthethreat.Thebenefitvarieswiththeaudittask.Wealsoshowthatpretrainedverbalizersrecoverwordsthatmodelshavelearnedtoconcealthroughfine-tuning,withoutadditionalverbalizertraining.Theseresultsidentifywhereauditorscanconcentrateexplanationgenerationandshowthatusefulexplanationscanextendbeyondthemodelaverbalizerwastrainedtodescribe.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.37040
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.37040 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.37040 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.37040 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Compute Optimal Tokenization (2 minute read)
This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.
Robust Explanations for User Trust in Enterprise NLP Systems
This paper proposes a unified black-box robustness evaluation framework for token-level explanations in enterprise NLP, comparing encoder (BERT, RoBERTa) and decoder (Qwen, Llama) models. It finds decoder LLMs produce substantially more stable explanations, with stability improving with scale, and provides a cost-robustness tradeoff curve for pre-deployment model selection.
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization
This paper proposes using sparse autoencoders to detect out-of-distribution inputs for transformers, including typos and jailbreak prompts, by analyzing spurious concept activations. The method enables a mechanistically grounded fine-tuning strategy to improve LLM robustness.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
This paper demonstrates that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing scalability concerns for dictionary learning. The features are multilingual, multimodal, and include safety-relevant concepts like deception and sycophancy, with causal influence on model outputs.