Tag
This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.
This paper systematically studies how lie typology, representation depth, probe expressivity, and sparse features impact deception detection in LLMs, finding that detection performance is highly dependent on training data and representation choice.
Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.
This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.
This paper introduces a novel method using sparse autoencoders (SAEs) to learn interpretable features from intermediate network activations for out-of-distribution (OOD) detection, achieving state-of-the-art performance and providing insights into how distribution shifts affect learned representations.
Toy Models of Superposition by Elhage et al. explains why interpretability is hard: models represent more features than dimensions via superposition, leading to polysemantic neurons as compression. This paper spawned the sparse autoencoder research program.
This paper introduces a matched evaluation protocol for sparse feature interventions in language models, showing that the claimed efficiency advantage of SAE-based safety control disappears or reverses when properly comparing against fair dense baselines.
This paper proposes a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder to extract universal features across independently trained BERT models, improving cross-seed feature alignment beyond post-hoc methods.
Introduces Stochastic Token Steering (STS) and Stochastic Block Steering (SBS) for LLM activation steering, which probabilistically gate steering signals per token or per sequence. Shows that steering only 50% of tokens recovers most of the dense-steering effect while preserving fluency, and that the behavioral outcome is rate-limited by cumulative signal dosage.
This paper introduces a concept-based interpretability framework for end-to-end autonomous driving models, using unsupervised dictionary learning via sparse autoencoders to decompose driving behavior into human-interpretable concepts. The framework enables analysis and targeted correction of model decisions, improving overall driving performance.
Proposes using Top-k Sparse Autoencoders to disentangle dense sentence embeddings into human-interpretable concepts, enabling steering of retrieval results without retraining.
This paper introduces sparse autoencoders to resolve superposition in neural networks, improving interpretability and geometric fidelity of latent spaces, and presents GW-map for cross-modal alignment between image representations and single-cell RNA sequencing data.
This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.
This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.
The paper introduces Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring, assigning each feature an intrinsic token name based on nearest token embedding, achieving high alignment in early layers without reducing reconstruction quality.
Introduces PairSAE, a method that adapts sparse autoencoders to interpret pairwise representations in protein co-folding models, enabling the discovery of interpretable features that align with biological annotations and predict binding affinities.
This paper proposes a continual learning method for LLMs that uses pretrained sparse autoencoders (SAEs) to regularize in activation space instead of weight space, achieving better memory efficiency and stronger performance on benchmarks while avoiding catastrophic forgetting without storing previous data.
This paper introduces Qwen3-Instruct SAE, a suite of sparse autoencoders trained on Qwen3 instruction-tuned models, enabling the discovery of millions of interpretable features and demonstrating refusal steering capabilities.
This paper uses Dedicated Feature Crosscoders to localize RL-induced tool-use capability in Qwen2.5-3B to a single steerable feature, achieving +65pp tool-correctness via feature steering and demonstrating capability spillover to frozen base models.
This paper proposes using sparse autoencoders to detect out-of-distribution inputs for transformers, including typos and jailbreak prompts, by analyzing spurious concept activations. The method enables a mechanistically grounded fine-tuning strategy to improve LLM robustness.