sparse-autoencoders

Tag

Cards List
#sparse-autoencoders

Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection

arXiv cs.LG · yesterday Cached

Prof-K is a probabilistic one-pass filtering algorithm for fast, scalable top-k selection with correctness guarantees, achieving 1.5x–10x speedups over PyTorch topk and RadiK, especially in large-scale small-k regimes.

0 favorites 0 likes
#sparse-autoencoders

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

arXiv cs.CL · yesterday Cached

This paper reports an activation study of Gemma 3 4B IT showing that the model's internal representations distinguish necessary falsehoods (impossibilities) from contingent falsehoods, with impossibility directions orthogonal to truth directions and overlapping with semantic anomaly directions.

0 favorites 0 likes
#sparse-autoencoders

Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv cs.AI · 3d ago Cached

This paper introduces Feature Nonlocality (FNL), an entropy-based metric for measuring the semantic abstractness of SAE features in LLMs, and demonstrates applications in auditing jailbreak mitigation features and improving MATH-500 accuracy via steering.

0 favorites 0 likes
#sparse-autoencoders

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

arXiv cs.CL · 4d ago Cached

This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.

0 favorites 0 likes
#sparse-autoencoders

"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

arXiv cs.CL · 4d ago Cached

This paper uses sparse autoencoders to decompose how language models represent the default Assistant, roleplay personas, and story characters, finding that personas retain an Assistant core while differentiating across layers, and story characters lack that core.

0 favorites 0 likes
#sparse-autoencoders

Finding Usable Weight Mechanisms with Tiled SVD

arXiv cs.AI · 5d ago Cached

This paper proposes extracting mechanism mounts directly from linear weight sites via column-tiled SVD, offering an alternative to proxy dictionaries like sparse autoencoders for mechanistic interpretability. Evaluated on Gemma-2-2B, the method passes all 182 site-layer checks.

0 favorites 0 likes
#sparse-autoencoders

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

arXiv cs.LG · 2026-08-07 Cached

CircuitSteer is a novel framework that uses sparse autoencoders to identify and manipulate multi-layer semantic circuits in LLMs, enabling more robust and fluency-preserving behavioral steering compared to single-layer methods like CAA.

0 favorites 0 likes
#sparse-autoencoders

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

arXiv cs.CL · 2026-08-07 Cached

A systematic empirical study showing that concept directions extracted from one language model can steer other independently trained models when sufficient scale (≥1.7B parameters) is reached, providing functional evidence for the Platonic Representation Hypothesis and highlighting scale thresholds for cross-model interpretability tools.

0 favorites 0 likes
#sparse-autoencoders

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

arXiv cs.LG · 2026-07-31 Cached

ECG-InterpBench is a new benchmark that systematically evaluates the interpretability of ECG foundation model representations using matched-scale sparse autoencoders, covering reconstruction fidelity, clinical concept accessibility, and reproducibility across 450 cells.

0 favorites 0 likes
#sparse-autoencoders

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

arXiv cs.CL · 2026-07-30 Cached

This paper adapts Funder's personality triad framework to LLMs, using sparse autoencoders to discover and validate trait-like internal representations, and demonstrating controllable bidirectional behavioral shifts through feature-level interventions.

0 favorites 0 likes
#sparse-autoencoders

Reference Feature Atlases for Mechanistic Auditing of Language Models

arXiv cs.AI · 2026-07-28 Cached

Proposes reference feature atlases for auditing language models, enabling efficient transfer of feature libraries across models and revealing model-specific anomalies.

0 favorites 0 likes
#sparse-autoencoders

Scaling Interpretable Transformers with Parity Bottleneck Layers

arXiv cs.LG · 2026-07-24 Cached

Introduces the ParityTransformer, a GPT-2-scale architecture with Deep Parity Bottleneck layers that make intermediate representations interpretable-by-design, efficiently enforcing sparsity without the memory costs of traditional wide bottlenecks, and demonstrating competitive performance on sparse probing tasks.

0 favorites 0 likes
#sparse-autoencoders

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

arXiv cs.LG · 2026-07-24 Cached

This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.

0 favorites 0 likes
#sparse-autoencoders

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

arXiv cs.AI · 2026-07-24 Cached

This paper systematically studies how lie typology, representation depth, probe expressivity, and sparse features impact deception detection in LLMs, finding that detection performance is highly dependent on training data and representation choice.

0 favorites 0 likes
#sparse-autoencoders

@fnruji316625: Agentic interpretability is becoming a research direction of its own. Instead of one-shot labeling, AI agents can: form…

X AI KOLs Timeline · 2026-07-15 Cached

Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.

0 favorites 0 likes
#sparse-autoencoders

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

arXiv cs.LG · 2026-07-15 Cached

This paper audits sparse autoencoder features using causal interventions, finding that up to 77% of correlationally recovered features in degraded SAEs and 9% in well-trained ones are causally inert. The authors introduce the sae-causal-audit tool for reproducible evaluation.

0 favorites 0 likes
#sparse-autoencoders

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

arXiv cs.LG · 2026-07-15 Cached

This paper introduces a novel method using sparse autoencoders (SAEs) to learn interpretable features from intermediate network activations for out-of-distribution (OOD) detection, achieving state-of-the-art performance and providing insights into how distribution shifts affect learned representations.

0 favorites 0 likes
#sparse-autoencoders

@TamazGadaev: day 10/n (series on fundamental texts for AI researchers - not specific papers so much as pieces that hand you a lens f…

X AI KOLs Timeline · 2026-07-14 Cached

Toy Models of Superposition by Elhage et al. explains why interpretability is hard: models represent more features than dimensions via superposition, leading to polysemantic neurons as compression. This paper spawned the sparse autoencoder research program.

0 favorites 0 likes
#sparse-autoencoders

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

arXiv cs.AI · 2026-07-14 Cached

This paper introduces a matched evaluation protocol for sparse feature interventions in language models, showing that the claimed efficiency advantage of SAE-based safety control disappears or reverses when properly comparing against fair dense baselines.

0 favorites 0 likes
#sparse-autoencoders

Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

arXiv cs.CL · 2026-07-10 Cached

This paper proposes a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder to extract universal features across independently trained BERT models, improving cross-seed feature alignment beyond post-hoc methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback