Selecting The Most Informative Tokens in Natural Language Autoencoders

Hugging Face Daily Papers Papers

Summary

The paper studies which token positions in natural language autoencoders are most informative to explain, finding that chat-structure-based ranking selects relevant explanations better than computational signals, and that explaining only 5% of positions retains nearly all audit success on prompt injection and concealment tasks.

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Original Article
View Cached Full Text

Cached at: 09/30/26, 12:16 PM

Paper page - Selecting The Most Informative Tokens in Natural Language Autoencoders

Source: https://huggingface.co/papers/2609.37040

Abstract

Naturallanguageautoencoderstranslatealanguagemodel’sinternalactivationsintoreadableexplanations.Explainingeverytokenpositioniscostly.Whichpositionsshouldanauditorinspecttounderstandapotentialthreat?Westudythisquestionacross4.7millionexplanationsonpromptinjectionandconcealment.Wecomparesignalsfrommodelcomputationwitharankertrainedonlyonchatstructure.Chatstructureusuallyselectsmorerelevantexplanationsthanthecomputationalsignals,withoutrequiringamodelforwardpassforpositionselection.Onthreeoffourdatasets,explainingjust5%ofpositionsretainsnearlyallofthesuccessratefromexplainingeveryposition,wheresuccessmeansobtaininganexplanationaboutthethreat.Thebenefitvarieswiththeaudittask.Wealsoshowthatpretrainedverbalizersrecoverwordsthatmodelshavelearnedtoconcealthroughfine-tuning,withoutadditionalverbalizertraining.Theseresultsidentifywhereauditorscanconcentrateexplanationgenerationandshowthatusefulexplanationscanextendbeyondthemodelaverbalizerwastrainedtodescribe.

View arXiv pageView PDFGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.37040

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37040 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37040 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37040 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Compute Optimal Tokenization (2 minute read)

TLDR AI

This paper systematically derives compression-aware neural scaling laws by training nearly 1,300 models, demonstrating that the widely used heuristic of 20 tokens per parameter is an artifact of specific tokenizers. The authors propose a tokenizer-agnostic scaling law based on bytes, offering a new framework for compute-efficient training across diverse languages and modalities.

Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

arXiv cs.CL

This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.

Robust Explanations for User Trust in Enterprise NLP Systems

arXiv cs.CL

This paper proposes a unified black-box robustness evaluation framework for token-level explanations in enterprise NLP, comparing encoder (BERT, RoBERTa) and decoder (Qwen, Llama) models. It finds decoder LLMs produce substantially more stable explanations, with stability improving with scale, and provides a cost-robustness tradeoff curve for pre-deployment model selection.

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

arXiv cs.AI

This paper demonstrates that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing scalability concerns for dictionary learning. The features are multilingual, multimodal, and include safety-relevant concepts like deception and sycophancy, with causal influence on model outputs.