Tag
This paper introduces the Entropic Bound, a spectral measure of task-intrinsic capacity for transformers, proving that the intrinsic rank of the token-mixing operator provides a tight lower bound on required model capacity. It shows that while a naive transfer from linear attention fails for real attention, an attention-native intrinsic rank restores the full theoretical structure.
This paper presents Unified Neural Scaling Laws (UNSL), a functional form that accurately models and extrapolates deep neural network scaling behaviors as multiple dimensions such as parameters, data, and steps vary simultaneously, improving over previous scaling laws.