Tag
This paper presents a new logarithmic-free upper bound for the generalization gap in uniformly stable algorithms and constructs a deterministic learning problem that achieves optimal high-probability dependence, closing a gap in the literature.
This paper investigates scenarios where adding correctly labeled data can harm model performance, introducing insertion-stability and examining the limits of dimension-based theory in machine learning generalization.
This paper presents a layer-wise information-theoretic framework for replay-based continual learning, decomposing the generalization gap into replay-induced representation drift and optimization-dependence terms, with refinements via Wasserstein relaxation and SGLD instantiation.
This paper extends PAC-Bayes theory by decomposing its complexity measure using behavioral equivalence, introducing PAC-Bayes Z-information and exact structural decomposition beyond parameter space.
This paper proposes performing PAC-Bayesian analysis on quotient parameter spaces to remove KL contributions from parameter symmetries, and constructs a geometry-induced prior that approximates the ideal posterior-matched prior, resulting in tighter generalization bounds. Experiments on Fourier regression and Query-Key attention show significant reductions in KL divergence and certificate values.
Researchers from Bridgewater AIA Labs, UIUC, and MIT prove the first non-vacuous generalization bounds for reasoning LLMs trained with RLVR, providing provable accuracy lower bounds on unseen data to guide safe deployment.
This paper theoretically analyzes linear transformers for in-context learning under domain generalization, establishing dimension-independent convergence rates and proposing novel activation and loss designs for linearizing pretrained softmax LLMs.
This paper presents theoretical bounds for uncertainty estimation and generalization in modern deep learning models.
This paper introduces a theoretical framework for quantifying deployment risk when training and deployment distributions differ due to latent regime dynamics modeled as a Markov-switching process, providing exact decomposition and finite-sample bounds.
This paper develops a PAC-Bayesian framework for test-time adaptation that uses MMD-balls as credal sets, providing formal generalization bounds and separating epistemic from aleatoric uncertainty under distribution shift.
This paper proposes an information-theoretic framework for emergent communication in Agentic AI Networking (AgentNet), addressing physical constraints and providing generalization bounds. Experimental validation on hardware prototypes demonstrates improved generalization performance compared to state-of-the-art solutions.