sparse-autoencoders

Tag

Cards List
#sparse-autoencoders

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.

0 favorites 0 likes
#sparse-autoencoders

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

arXiv cs.AI ↗ · 2026-07-24 Cached

This paper systematically studies how lie typology, representation depth, probe expressivity, and sparse features impact deception detection in LLMs, finding that detection performance is highly dependent on training data and representation choice.

0 favorites 0 likes
#sparse-autoencoders

@fnruji316625: Agentic interpretability is becoming a research direction of its own. Instead of one-shot labeling, AI agents can: form…

X AI KOLs Timeline ↗ · 2026-07-15 Cached

Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.

0 favorites 0 likes
#sparse-autoencoders

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

arXiv cs.LG ↗ · 2026-07-15 Cached

This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.

0 favorites 0 likes
#sparse-autoencoders

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

arXiv cs.LG ↗ · 2026-07-15 Cached

This paper introduces a novel method using sparse autoencoders (SAEs) to learn interpretable features from intermediate network activations for out-of-distribution (OOD) detection, achieving state-of-the-art performance and providing insights into how distribution shifts affect learned representations.

0 favorites 0 likes
#sparse-autoencoders

@TamazGadaev: day 10/n (series on fundamental texts for AI researchers - not specific papers so much as pieces that hand you a lens f…

X AI KOLs Timeline ↗ · 2026-07-14 Cached

Toy Models of Superposition by Elhage et al. explains why interpretability is hard: models represent more features than dimensions via superposition, leading to polysemantic neurons as compression. This paper spawned the sparse autoencoder research program.

0 favorites 0 likes
#sparse-autoencoders

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

arXiv cs.AI ↗ · 2026-07-14 Cached

This paper introduces a matched evaluation protocol for sparse feature interventions in language models, showing that the claimed efficiency advantage of SAE-based safety control disappears or reverses when properly comparing against fair dense baselines.

0 favorites 0 likes
#sparse-autoencoders

Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

arXiv cs.CL ↗ · 2026-07-10 Cached

This paper proposes a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder to extract universal features across independently trained BERT models, improving cross-seed feature alignment beyond post-hoc methods.

0 favorites 0 likes
#sparse-autoencoders

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

arXiv cs.LG ↗ · 2026-07-08 Cached

Introduces Stochastic Token Steering (STS) and Stochastic Block Steering (SBS) for LLM activation steering, which probabilistically gate steering signals per token or per sequence. Shows that steering only 50% of tokens recovers most of the dense-steering effect while preserving fluency, and that the behavioral outcome is rate-limited by cumulative signal dosage.

0 favorites 0 likes
#sparse-autoencoders

Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models

arXiv cs.AI ↗ · 2026-07-08 Cached

This paper introduces a concept-based interpretability framework for end-to-end autonomous driving models, using unsupervised dictionary learning via sparse autoencoders to decompose driving behavior into human-interpretable concepts. The framework enables analysis and targeted correction of model decisions, improving overall driving performance.

0 favorites 0 likes
#sparse-autoencoders

Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders

arXiv cs.AI ↗ · 2026-07-02 Cached

Proposes using Top-k Sparse Autoencoders to disentangle dense sentence embeddings into human-interpretable concepts, enabling steering of retrieval results without retraining.

0 favorites 0 likes
#sparse-autoencoders

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

arXiv cs.LG ↗ · 2026-07-01 Cached

This paper introduces sparse autoencoders to resolve superposition in neural networks, improving interpretability and geometric fidelity of latent spaces, and presents GW-map for cross-modal alignment between image representations and single-cell RNA sequencing data.

0 favorites 0 likes
#sparse-autoencoders

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.

0 favorites 0 likes
#sparse-autoencoders

Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper introduces turn-averaged sparse autoencoders (SAEs) that operate on average activations across conversational turns, enabling efficient feature discovery and attribution graphs for long contexts. It also proposes a nested architecture for joint training with per-token features.

0 favorites 0 likes
#sparse-autoencoders

VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

arXiv cs.CL ↗ · 2026-06-29 Cached

The paper introduces Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring, assigning each feature an intrinsic token name based on nearest token embedding, achieving high alignment in early layers without reducing reconstruction quality.

0 favorites 0 likes
#sparse-autoencoders

PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

arXiv cs.LG ↗ · 2026-06-29 Cached

Introduces PairSAE, a method that adapts sparse autoencoders to interpret pairwise representations in protein co-folding models, enabling the discovery of interpretable features that align with biological annotations and predict binding affinities.

0 favorites 0 likes
#sparse-autoencoders

From Weights to Features: SAE-Guided Activation Regularization for LLM Continual Learning

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper proposes a continual learning method for LLMs that uses pretrained sparse autoencoders (SAEs) to regularize in activation space instead of weight space, achieving better memory efficiency and stronger performance on benchmarks while avoiding catastrophic forgetting without storing previous data.

0 favorites 0 likes
#sparse-autoencoders

Discovering Millions of Interpretable Features with Sparse Autoencoders

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper introduces Qwen3-Instruct SAE, a suite of sparse autoencoders trained on Qwen3 instruction-tuned models, enabling the discovery of millions of interpretable features and demonstrating refusal steering capabilities.

0 favorites 0 likes
#sparse-autoencoders

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper uses Dedicated Feature Crosscoders to localize RL-induced tool-use capability in Qwen2.5-3B to a single steerable feature, achieving +65pp tool-correctness via feature steering and demonstrating capability spillover to frozen base models.

0 favorites 0 likes
#sparse-autoencoders

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper proposes using sparse autoencoders to detect out-of-distribution inputs for transformers, including typos and jailbreak prompts, by analyzing spurious concept activations. The method enables a mechanistically grounded fine-tuning strategy to improve LLM robustness.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback