Tag
SPICE introduces a generalizable framework for analyzing polysemanticity in neural networks, enabling systematic comparison across CNNs and Transformers and automatically determining concept clusters per neuron.
This paper provides a mathematical analysis of superposition in neural networks, deriving upper and lower bounds on L2 reconstruction loss for simple autoencoders with power activation functions, corroborating empirical findings by Elhage et al.