Tag
This paper proposes a theoretical framework that bridges approximation theory and emergent phenomena in deep learning, offering new insights into how neural networks learn.
This paper derives a scaling law for sketched linear contrastive learning under a Gaussian latent-variable model, analyzing how risk decomposes into approximation, optimization, and statistical terms, and provides theoretical guidance for balancing model size, data, and compute in contrastive learning.
Explains the Student's t-distribution correction for small sample confidence intervals, providing a memorizable table for 90% intervals and a rule-of-thumb for estimating standard deviation from two samples.
A philosophical reflection on the problem of approximation in mathematics, finance, and AI, discussing how models are imperfect but useful, and the implications for AI safety and machine consciousness.