Tag
This paper applies Wilsonian renormalization group theory to analyze Transformer attention as a perturbation of the MLP residual-stack fixed point, determining whether attention is relevant or irrelevant based on data correlation length. Experiments on synthetic Markov chains confirm that attention's relevance depends on the spectral structure of the data-generating process, with the first-layer head dominating the transition.
A researcher claims to have established the scientific principles behind deep learning using Renormalization Group theory, moving beyond engineering hacks to potentially pave the way for AGI. This work builds on papers from 2021 in JMLR and Nature Communications.
This paper introduces the RG-Flow Transformer, a model with a renormalization-group inductive bias for analyzing scarce EEG data. It benchmarks against a vanilla transformer on sleep staging from the Sleep-EDF dataset, finding no accuracy advantage but better interpretability through recovery of the spectral exponent.
This book develops an effective theory for deep neural networks, showing that their predictions are nearly-Gaussian and governed by the depth-to-width ratio, and introduces representation group flow to analyze signal propagation and learning dynamics.
Proposes RGNet, a neural network architecture based on renormalization group theory for hierarchical coarse-graining of feature space to address class imbalance and noise in fault diagnosis. Experimental results on the AI4I dataset show RGNet provides interpretable and competitive performance.
This article explores the deep connections between physics and deep learning, analyzes the isomorphism of phenomena such as Scaling Law and emergence with concepts like critical scaling laws and phase transitions in physics, and reviews the current status and prospects of applying physical methodologies in AI.