Tag
This preprint uses adversarial training as a controlled instrument on GPT-2 Small to test whether representational simplicity (SAE decomposability, concentrated attribution) implies causal circuit simplicity. It finds that robust models are more SAE-decomposable, while circuit size is regime-dependent: robustness helps at high faithfulness levels (90-95%) but not below 85%.
BiasReducer is a lightweight framework that edits only the linear reward head of reward models to adaptively mitigate biases toward superficial attributes like response length and confidence, using a sparse-autoencoder-style encoder to detect and rank relevant biases per dataset. Across five reward models it improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming training-based baselines and reducing downstream verbosity and sycophancy.
The paper investigates how parts-of-speech categories are encoded in Sparse AutoEncoder latent spaces, finding that they are distributed and not one-to-one with individual latents.
MonoTM is an interpretable topic modeling framework that decouples mixture estimation from feature interpretation using sparse autoencoders, providing semantically meaningful topics beyond traditional word-based representations.
The paper shows that statistical top-k feature selection for SAE-based steering of LLMs is suboptimal and proposes Neighbor Integrated Feature Selection (NIFS) to improve performance by leveraging representation similarity.
This paper introduces reward-informed sparse autoencoders (RI-SAEs) to use reinforcement learning rewards for interpretability, but finds that the separation between good and bad reasoning is largely driven by solution completeness rather than reasoning quality.
This research investigates whether multilingual large language models develop shared internal representations for mathematical reasoning across languages, introducing a novel Geometry-Invariant Sparse Autoencoder (GI-SAE) method. It finds that cross-language feature sharing is model-dependent and that geometric similarity does not consistently imply functional interchangeability.
Prof-K is a probabilistic one-pass filtering algorithm for fast, scalable top-k selection with correctness guarantees, achieving 1.5x–10x speedups over PyTorch topk and RadiK, especially in large-scale small-k regimes.
This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.
This paper introduces Feature Nonlocality (FNL), an entropy-based metric for measuring the semantic abstractness of SAE features in LLMs, and demonstrates applications in auditing jailbreak mitigation features and improving MATH-500 accuracy via steering.
This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.
This paper uses sparse autoencoders to decompose how language models represent the default Assistant, roleplay personas, and story characters, finding that personas retain an Assistant core while differentiating across layers, and story characters lack that core.
This paper proposes extracting mechanism mounts directly from linear weight sites via column-tiled SVD, offering an alternative to proxy dictionaries like sparse autoencoders for mechanistic interpretability. Evaluated on Gemma-2-2B, the method passes all 182 site-layer checks.
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
CircuitSteer is a novel framework that uses sparse autoencoders to identify and manipulate multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering compared to single-layer methods like CAA.
A systematic empirical study showing that concept directions extracted from one language model can steer other independently trained models when sufficient scale (≥1.7B parameters) is reached, providing functional evidence for the Platonic Representation Hypothesis and highlighting scale thresholds for cross-model interpretability tools.
ECG-InterpBench is a new benchmark that systematically evaluates the interpretability of ECG foundation model representations using matched-scale sparse autoencoders, covering reconstruction fidelity, clinical concept accessibility, and reproducibility across 450 cells.
This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.
Proposes reference feature atlases for auditing language models, enabling efficient transfer of feature libraries across models and revealing model-specific anomalies.
Introduces the ParityTransformer, a GPT-2-scale architecture with Deep Parity Bottleneck layers that make intermediate representations interpretable-by-design, efficiently enforcing sparsity without the memory costs of traditional wide bottlenecks, and demonstrating competitive performance on sparse probing tasks.