neural-scaling-laws

Tag

Cards List
#neural-scaling-laws

The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers

arXiv cs.LG · 2026-07-28 Cached

This paper introduces the Entropic Bound, a spectral measure of task-intrinsic capacity for transformers, proving that the intrinsic rank of the token-mixing operator provides a tight lower bound on required model capacity. It shows that while a naive transfer from linear attention fails for real attention, an attention-native intrinsic rank restores the full theoretical structure.

0 favorites 0 likes
#neural-scaling-laws

Unified Neural Scaling Laws

arXiv cs.LG · 2026-05-27 Cached

This paper presents Unified Neural Scaling Laws (UNSL), a functional form that accurately models and extrapolates deep neural network scaling behaviors as multiple dimensions such as parameters, data, and steps vary simultaneously, improving over previous scaling laws.

0 favorites 0 likes
← Back to home

Submit Feedback