polysemanticity

Tag

Cards List
#polysemanticity

SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation

arXiv cs.LG ↗ · 2026-09-15 Cached

SPICE introduces a generalizable framework for analyzing polysemanticity in neural networks, enabling systematic comparison across CNNs and Transformers and automatically determining concept clusters per neuron.

0 favorites 0 likes
#polysemanticity

Effects of sparsity and superposition on loss in simple autoencoders

arXiv cs.LG ↗ · 2026-06-18 Cached

This paper provides a mathematical analysis of superposition in neural networks, deriving upper and lower bounds on L2 reconstruction loss for simple autoencoders with power activation functions, corroborating empirical findings by Elhage et al.

0 favorites 0 likes
← Back to home

Submit Feedback