sparse-autoencoder

Tag

Cards List
#sparse-autoencoder

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

arXiv cs.AI · 6h ago Cached

This paper demonstrates that linear probes decoding concepts from language model activations do not necessarily identify causally relevant features, and introduces a feature-level diagnostic using sparse autoencoders to separate probe alignment from behavioral drivers.

0 favorites 0 likes
#sparse-autoencoder

Test-Time Unlearning via Sparse Autoencoder

arXiv cs.LG · yesterday Cached

ARIA is a test-time unlearning method for large language models that uses sparse autoencoders to suppress unwanted knowledge during inference without modifying weights, improving the forget-retain trade-off and remaining robust to adversarial attacks.

0 favorites 0 likes
#sparse-autoencoder

Data Attribution of Emergent Misalignment with Persona Features

arXiv cs.CL · 2026-08-12 Cached

This paper investigates emergent misalignment in fine-tuned language models, using SAE-based model diffing to identify persona features that control misalignment and attributing them to pre-training web documents.

0 favorites 0 likes
#sparse-autoencoder

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

arXiv cs.AI · 2026-08-12 Cached

This paper investigates the compressibility of latent-space communication between vision-language model agents by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations. It demonstrates a 128x reduction in transmitted bytes with minimal accuracy loss, revealing that the dense communication channel is highly redundant.

0 favorites 0 likes
#sparse-autoencoder

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

arXiv cs.AI · 2026-07-29 Cached

CADENCE uses sparse autoencoders to decompose ECG foundation model representations into interpretable physiological concepts, significantly improving alignment with clinical phenotypes and waveform morphology.

0 favorites 0 likes
#sparse-autoencoder

Interpreting Brain Responses to Language with Sparse Features from Language Models

arXiv cs.CL · 2026-06-08 Cached

This paper introduces Augmented Sparse Encoding Models to interpret brain responses to language using sparse features from language models, validated on high-field 7T fMRI data. It recovers known neural tuning properties and discovers a new voxel population tuned to people-related content.

0 favorites 0 likes
#sparse-autoencoder

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

Hugging Face Daily Papers · 2026-06-05 Cached

This paper demonstrates that Whisper's hallucination failures on silence, noise, or music can be detected and mitigated purely from internal activations using sparse autoencoders, achieving large reductions in hallucination rate without fine-tuning.

0 favorites 0 likes
#sparse-autoencoder

How Quantization Changes Interpretable Features: A Sparse Autoencoder Analysis of Language Models

arXiv cs.LG · 2026-06-03 Cached

This paper investigates whether interpretable features identified by sparse autoencoders in full-precision language models remain faithful after quantization, finding systematic degradation that behavioral metrics like perplexity can miss.

0 favorites 0 likes
#sparse-autoencoder

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

arXiv cs.AI · 2026-06-02 Cached

Introduces Latent Reward Steering (Lrs), an adaptive inference-time framework that uses sparse autoencoder latent states and a learned reward model to implicitly promote cognitive behaviors like verification and backtracking in reasoning LLMs, improving performance across multiple models and benchmarks.

0 favorites 0 likes
#sparse-autoencoder

I built a tool that shows you what GPT-2 is "thinking" in real-time as it generates 3D graph of concept activations per token [R]

Reddit r/MachineLearning · 2026-05-19

A developer built AXON, a tool that visualizes GPT-2's internal concept activations as a live 3D force graph using Sparse Autoencoders, allowing users to see interpretable features firing before token generation.

0 favorites 0 likes
#sparse-autoencoder

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

arXiv cs.CL · 2026-05-13 Cached

This article introduces Qwen-Scope, a toolkit of Sparse Autoencoders (SAEs) trained on Qwen3 and Qwen3.5 models to enable mechanistic analysis and intervention. It releases 14 groups of SAE weights covering dense and MoE backbones, providing sparse representations for residual-stream activations.

0 favorites 0 likes
#sparse-autoencoder

WriteSAE: Sparse Autoencoders for Recurrent State

Hugging Face Daily Papers · 2026-05-12 Cached

WriteSAE introduces the first sparse autoencoder that decomposes matrix cache writes in state-space and hybrid recurrent language models, enabling superior token-level interventions compared to existing methods.

0 favorites 0 likes
← Back to home

Submit Feedback