Tag
The paper demonstrates that standard neural scaling law derivations fail when data forms group orbits, as intrinsic dimension is undefined, leading to exponential rather than power law scaling in model performance.
This paper provides causal evidence that gradient descent in matrix memories trained on group composition tasks recruits exactly the rank required by the task's algebra, correlating with the minimal faithful representation dimensions of finite groups.