Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks
Summary
This paper presents an exponential depth hierarchy for ReLU networks in terms of L2 approximation error, demonstrating that deeper networks offer exponentially improved representational power for function approximation.
View Cached Full Text
Cached at: 08/26/26, 09:28 AM
# Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks Source: [https://arxiv.org/abs/2608.23877](https://arxiv.org/abs/2608.23877) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools IArxiv recommender toggle About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
Shallower ReLU Network Representations via Exact Linear Algebra
This paper improves theoretical bounds on the depth of ReLU networks needed to represent the maximum function, showing exact two-hidden-layer representations for up to 10 inputs and improved depth for larger n via exact linear algebra techniques.
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression
This paper proves that the minimax risk for deep variation-norm ReLU regression has quadratic dependence on depth, using local packing arguments and approximation theorems.
Representing MAX functions using two-hidden-layer ReLU networks
This paper explores the theoretical representation of MAX functions using two-hidden-layer ReLU neural networks, providing detailed coefficient lists and identities for exact representations.
A law of robustness for two-layer neural networks with arbitrary weights
This paper proves a conjectured law of robustness for two-layer neural networks with unbounded weights, showing that a network fitting noisy data must have a Lipschitz constant at least of order sqrt(n/m), up to a logarithmic factor, for continuous piecewise-linear activations like ReLU.
Generalized Neurons
The article explores the Universal Approximation Theorem in deep learning, analyzing the representation capacity of individual neurons and neural network layers using ReLU activation functions.