feature-analysis

Tag

Cards List
#feature-analysis

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

arXiv cs.CL · 2026-09-03 Cached

This paper introduces Sparse Readout Prism (SRP), a method to analyze language model predictions by decomposing the readout matrix into sparse features, enabling corpus-independent explanations of logit-lens scores and improving interpretability.

0 favorites 0 likes
#feature-analysis

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

arXiv cs.LG · 2026-07-24 Cached

This paper investigates whether single-token sparse autoencoder features are causally necessary for model outputs, finding that causal roles vary across SAE families and are sensitive to training methodology rather than activation function or scale.

0 favorites 0 likes
← Back to home

Submit Feedback