Tag
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.
This paper argues that within-class variance in language model representations is not incomplete neural collapse but allocated information storage, and that the allocation obeys an information floor law. Across 14 models, macro-category structure carries only 4–12% of representational variance, while within-token context dominates at 79–91%.