feature-survival

Tag

Cards List
#feature-survival

How Quantization Changes Interpretable Features: A Sparse Autoencoder Analysis of Language Models

arXiv cs.LG · 2026-06-03 Cached

This paper investigates whether interpretable features identified by sparse autoencoders in full-precision language models remain faithful after quantization, finding systematic degradation that behavioral metrics like perplexity can miss.

0 favorites 0 likes
← Back to home

Submit Feedback