Tag
The article contrasts Noam Chomsky's views on Universal Grammar with modern AI language models, highlighting how statistical learning in GPUs has achieved superior grammar and understanding.
This paper investigates the limits of proper learning and regularization in multiclass learning, resolving open problems by demonstrating that learning cannot always be reduced to proper learning and that regularization has structural constraints.
This paper establishes tight generalization bounds for multi-dimensional hyperparameter tuning in data-driven algorithm design, using real algebraic geometry and a multi-regime lower-bound framework to resolve theoretical gaps.
The paper introduces NS-RIS, a scalable Newton-Schulz retraction-based algorithm for learning hidden quantum Markov models on the Stiefel manifold, providing the first mathematical performance guarantee and empirical evidence that HQMMs can outperform EM-trained HMMs on non-quantum-generated data.
This paper uses a developmental approach to study how neural language models, specifically Transformers, learn statistical patterns from a synthetic grammar, finding that they first acquire global abstract statistics then local dependencies, with over-generalizations early on.
This paper develops a statistical theory for offline reinforcement learning from trajectory-level outcome supervision, proposing the OPAC algorithm and characterizing when such supervision enables efficient learning versus when fundamental barriers arise.
This paper develops a framework to grade the capability to infer in data-driven systems under the European AI Act, using credit scoring as a case study to illustrate where inference occurs and where regulatory clarity is needed.
This paper presents a computational framework to test competing maturational theories of syntactic development in children, specifically comparing bottom-up versus inward accounts using statistical grammar induction.
This paper proposes Online Localized Conformal Prediction (OLCP) to address covariate heterogeneity in online learning and time-series settings. It introduces OLCP-Hedge for bandwidth selection and demonstrates valid long-run coverage with narrower prediction sets compared to existing baselines.