标签
This paper applies Generalizability Theory to agent benchmarks, showing leaderboards rank specialization rather than capability, and proposes a framework (DDR) for sizing reliable deployment evaluations.
本文认为,语言模型表示中的类内方差并非不完全的神经坍缩,而是分配的信息存储,且这种分配服从信息下限定律。在14个模型中,宏观类别结构仅承载4–12%的表示方差,而词元内上下文则占据79–91%。