Tag
This thread discusses a finding that embedding models struggle with binding (a core requirement for multi-object understanding) and that solving it requires going beyond the linear representation hypothesis.
This paper introduces a quantitative framework for estimating how many near-orthogonal directions a transformer language model's latent space can support, based on the linear representation and superposition hypotheses. The authors define representational capacity as an upper bound on distinguishable features and show it is exponentially sensitive to the allowed deviation from orthogonality, with larger models favoring tighter constraints.
This paper investigates the Platonic Representation Hypothesis, proposing that alignment arises from linear structure in representations, and introduces a statistical framework of signal, bias, and noise.