Tag
This paper introduces a quantitative framework for estimating how many near-orthogonal directions a transformer language model's latent space can support, based on the linear representation and superposition hypotheses. The authors define representational capacity as an upper bound on distinguishable features and show it is exponentially sensitive to the allowed deviation from orthogonality, with larger models favoring tighter constraints.