From Approximation to Emergence: A Theory of Deep Learning
Summary
This paper proposes a theoretical framework that bridges approximation theory and emergent phenomena in deep learning, offering new insights into how neural networks learn.
View Cached Full Text
Cached at: 07/03/26, 05:39 AM
# From Approximation to Emergence: A Theory of Deep Learning Source: [https://arxiv.org/abs/2607.01311](https://arxiv.org/abs/2607.01311) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools IArxiv recommender toggle About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
Emergence via Phase Transitions: Mechanism Landscapes and Universal Convergence Across Complex Systems
This paper introduces the Hierarchical Emergence Framework (HEF), which explains how diverse systems such as neural networks and biological evolution converge to similar internal representations through phase transitions in mechanism landscapes under physical and informational constraints. The framework is validated empirically with 111 grokking experiments that confirm universal convergence and identify a critical energy threshold.
Universal Approximation of Nonlinear Operators and Their Derivatives
This paper proves the first universal approximation theorems for nonlinear operators and their derivatives in infinite-dimensional settings, extending classical results to operator learning architectures like DeepONet and PCA-Net.
The Hamilton-Jacobi Theory of Deep Learning
This paper identifies neural network training as a search through Hamilton-Jacobi initial-value problems, showing that residual networks, transformers, and RNNs discretize the same class of viscous Hamilton-Jacobi equations. It derives quantitative consequences including minimax optimal generalization rates, adversarial robustness bounds, and a closed-form influence function.
From Score Approximation to Distribution Approximation in Score-Based Diffusion Models
This paper establishes a rigorous quantitative connection between neural network score function approximation and the resulting distribution approximation in score-based diffusion models, proving that accurate score approximation leads to close distribution approximation in KL divergence, with an explicit bound.
@k_solidified_: https://arxiv.org/abs/2106.10165 All of humanity should read this
This book develops an effective theory for deep neural networks, showing that their predictions are nearly-Gaussian and governed by the depth-to-width ratio, and introduces representation group flow to analyze signal propagation and learning dynamics.