sparse-autoencoders

Tag

Cards List
#sparse-autoencoders

Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training

arXiv cs.AI ↗ · 2d ago Cached

This preprint uses adversarial training as a controlled instrument on GPT-2 Small to test whether representational simplicity (SAE decomposability, concentrated attribution) implies causal circuit simplicity. It finds that robust models are more SAE-decomposable, while circuit size is regime-dependent: robustness helps at high faithfulness levels (90-95%) but not below 85%.

0 favorites 0 likes
#sparse-autoencoders

BiasReducer: Adaptive Bias Mitigation for Reward Models

Hugging Face Daily Papers ↗ · 6d ago Cached

BiasReducer is a lightweight framework that edits only the linear reward head of reward models to adaptively mitigate biases toward superficial attributes like response length and confidence, using a sparse-autoencoder-style encoder to detect and rank relevant biases per dataset. Across five reward models it improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming training-based baselines and reducing downstream verbosity and sycophancy.

0 favorites 0 likes
#sparse-autoencoders

Parts-of-Speech as Emergent Categories in SAE Latent Space

arXiv cs.CL ↗ · 2026-09-25 Cached

The paper investigates how parts-of-speech categories are encoded in Sparse AutoEncoder latent spaces, finding that they are distributed and not one-to-one with individual latents.

0 favorites 0 likes
#sparse-autoencoders

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

arXiv cs.CL ↗ · 2026-09-10 Cached

MonoTM is an interpretable topic modeling framework that decouples mixture estimation from feature interpretation using sparse autoencoders, providing semantically meaningful topics beyond traditional word-based representations.

0 favorites 0 likes
#sparse-autoencoders

Enhancing SAE-based Steering via Neighbor Integrated Feature Selection

arXiv cs.AI ↗ · 2026-09-01 Cached

The paper shows that statistical top-k feature selection for SAE-based steering of LLMs is suboptimal and proposes Neighbor Integrated Feature Selection (NIFS) to improve performance by leveraging representation similarity.

0 favorites 0 likes
#sparse-autoencoders

Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper introduces reward-informed sparse autoencoders (RI-SAEs) to use reinforcement learning rewards for interpretability, but finds that the separation between good and bad reasoning is largely driven by solution completeness rather than reasoning quality.

0 favorites 0 likes
#sparse-autoencoders

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

arXiv cs.LG ↗ · 2026-08-26 Cached

This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.

0 favorites 0 likes
#sparse-autoencoders

Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection

arXiv cs.LG ↗ · 2026-08-14 Cached

Prof-K is a probabilistic one-pass filtering algorithm for fast, scalable top-k selection with correctness guarantees, achieving 1.5x–10x speedups over PyTorch topk and RadiK, especially in large-scale small-k regimes.

0 favorites 0 likes
#sparse-autoencoders

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.

0 favorites 0 likes
#sparse-autoencoders

Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv cs.AI ↗ · 2026-08-12 Cached

This paper introduces Feature Nonlocality (FNL), an entropy-based metric for measuring the semantic abstractness of SAE features in LLMs, and demonstrates applications in auditing jailbreak mitigation features and improving MATH-500 accuracy via steering.

0 favorites 0 likes
#sparse-autoencoders

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.

0 favorites 0 likes
#sparse-autoencoders

"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper uses sparse autoencoders to decompose how language models represent the default Assistant, roleplay personas, and story characters, finding that personas retain an Assistant core while differentiating across layers, and story characters lack that core.

0 favorites 0 likes
#sparse-autoencoders

Finding Usable Weight Mechanisms with Tiled SVD

arXiv cs.AI ↗ · 2026-08-10 Cached

This paper proposes extracting mechanism mounts directly from linear weight sites via column-tiled SVD, offering an alternative to proxy dictionaries like sparse autoencoders for mechanistic interpretability. Evaluated on Gemma-2-2B, the method passes all 182 site-layer checks.

0 favorites 0 likes
#sparse-autoencoders

Multimodal Model Diffing for Feature Discovery and Control

Hugging Face Daily Papers ↗ · 2026-08-10 Cached

MMDiff uses multimodal sparse autoencoders to isolate, detect, and control features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.

0 favorites 0 likes
#sparse-autoencoders

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

arXiv cs.LG ↗ · 2026-08-07 Cached

CircuitSteer is a novel framework that uses sparse autoencoders to identify and manipulate multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering compared to single-layer methods like CAA.

0 favorites 0 likes
#sparse-autoencoders

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

arXiv cs.CL ↗ · 2026-08-07 Cached

A systematic empirical study showing that concept directions extracted from one language model can steer other independently trained models when sufficient scale (≥1.7B parameters) is reached, providing functional evidence for the Platonic Representation Hypothesis and highlighting scale thresholds for cross-model interpretability tools.

0 favorites 0 likes
#sparse-autoencoders

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv cs.LG ↗ · 2026-07-31 Cached

ECG-InterpBench is a new benchmark that systematically evaluates the interpretability of ECG foundation model representations using matched-scale sparse autoencoders, covering reconstruction fidelity, clinical concept accessibility, and reproducibility across 450 cells.

0 favorites 0 likes
#sparse-autoencoders

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.

0 favorites 0 likes
#sparse-autoencoders

Reference Feature Atlases for Mechanistic Auditing of Language Models

arXiv cs.AI ↗ · 2026-07-28 Cached

Proposes reference feature atlases for auditing language models, enabling efficient transfer of feature libraries across models and revealing model-specific anomalies.

0 favorites 0 likes
#sparse-autoencoders

Scaling Interpretable Transformers with Parity Bottleneck Layers

arXiv cs.LG ↗ · 2026-07-24 Cached

Introduces the ParityTransformer, a GPT-2-scale architecture with Deep Parity Bottleneck layers that make intermediate representations interpretable-by-design, efficiently enforcing sparsity without the memory costs of traditional wide bottlenecks, and demonstrating competitive performance on sparse probing tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback