theory

Tag

Cards List
#theory

Why Backdooring Neural Networks is so Easy?

arXiv cs.LG ↗ · 5h ago Cached

This paper derives an exact closed-form analysis showing that feature learning in neural networks makes them more vulnerable to backdoor attacks, with feature-learning regimes requiring only a trigger strength scaling as α ∝ π^{-1/4} versus π^{-1/2} for lazy learners — theoretically explaining why backdooring large networks is surprisingly easy and why linear security audits underestimate the threat.

0 favorites 0 likes
#theory

The syntax and semantics of goals

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper examines the syntax and semantics of goal representations, analyzing their compositionality and role in rational behavior for cognitive science and AI.

0 favorites 0 likes
#theory

The Relation Between Mathematics and Physics by Paul Dirac

Hacker News Top ↗ · 2026-09-13

Paul Dirac explores the deep connection between mathematics and physics, emphasizing their interdependence in scientific frameworks.

0 favorites 0 likes
#theory

@randall_balestr: Will be speaking at Harvard today at 9:45 about world models, Le* family and why we need more theory and mathematics to…

X AI KOLs Following ↗ · 2026-09-11 Cached

The author announces a speaking engagement at Harvard University where they will discuss world models, the Le* family, and the need for more theory and mathematics to advance JEPAs.

0 favorites 0 likes
#theory

A Function-Space Approach to the Statistical Mechanics of Learning Dynamics

arXiv cs.AI ↗ · 2026-09-11 Cached

This paper develops a statistical-mechanical framework for analyzing learning dynamics in deep neural networks by shifting from parameter space to function space, deriving exact error dynamics and fluctuation-induced effects.

0 favorites 0 likes
#theory

Bit Radix Theory — Cognition Between Human–AI and Across All Life (w/ AI Summary)

Reddit r/ArtificialInteligence ↗ · 2026-08-29

The Bit Radix Theory argues that different forms of intelligence, such as human and AI, may have fundamentally distinct cognitive approaches rather than converging, emphasizing complementary strengths. The paper promotes open critical engagement by inviting readers to test the theory using AI tools.

0 favorites 0 likes
#theory

Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper analyzes how finite Newton-Schulz iterations in the Muon optimizer benefit nonsmooth nonconvex optimization by smoothing the polar map, providing convergence guarantees that match best-known bounds.

0 favorites 0 likes
#theory

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

arXiv cs.LG ↗ · 2026-08-27 Cached

The paper proposes 'unit' as an explicit primitive in machine learning, where learning tasks declare persistent individuals, and supervised learning specializes to unit-conditioned response laws with tokenization.

0 favorites 0 likes
#theory

Glm 5.3 flash?

Reddit r/LocalLLaMA ↗ · 2026-08-25

The article speculates on the upcoming release of GLM 5.3 weights and suggests that OxAlpha may be a new GLM model variant.

0 favorites 0 likes
#theory

The Origin of Consciousness (2008)

Hacker News Top ↗ · 2026-08-17 Cached

The article discusses Julian Jaynes' theory that human consciousness is a recent development, emerging around 3,000 years ago due to a shift in brain integration, supported by historical and neurological evidence.

0 favorites 0 likes
#theory

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

arXiv cs.AI ↗ · 2026-08-17 Cached

The paper establishes a theoretical connection between probabilistic Joint-Embedding Predictive Learning (JEPA) and Hidden Markov Models (HMMs), providing a state-space interpretation and introducing Markov-Chain JEPA for enhanced consistency.

0 favorites 0 likes
#theory

Are there any theoretically-guided practices left in machine learning nowadays? [D]

Reddit r/MachineLearning ↗ · 2026-08-14

The article questions whether theoretical principles still guide machine learning practices, highlighting how many once-standard theories have been challenged by empirical evidence.

0 favorites 0 likes
#theory

On the Expressive Power of Transformers

arXiv cs.AI ↗ · 2026-08-14 Cached

A survey paper examining the expressive power of transformers as language recognizers, using concepts and methods from circuit complexity to compare them with classical models of computation.

0 favorites 0 likes
#theory

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper introduces low interaction rank as a unified theoretical framework for multiplicative dual-encoder networks, covering approximation, sample complexity, normalization, and identifiability, with experiments on operator learning and CLIP models.

0 favorites 0 likes
#theory

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper proves that a single normalized nonnegative kernel-attention head requires exponentially many features to solve a simple Min-IP task on three-token sequences, whereas dense softmax attention solves it with constant temperature and m-dimensional scores, highlighting a fundamental expressive-power gap between kernel and full attention.

0 favorites 0 likes
#theory

Structuring the Space of Perspectives

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper reviews the concept of 'perspective' in NLP, proposes a hierarchy of perspective-related concepts along a specificity axis, and demonstrates how this hierarchy can help researchers choose appropriate operationalizations.

0 favorites 0 likes
#theory

Decoupled Descent: Enforcing Exact Train-Test Error Tracking Via AMP Onsager Corrections [R]

Reddit r/MachineLearning ↗ · 2026-08-11

A theory paper introducing Decoupled Descent (DD), a training method that uses approximate message passing Onsager corrections to enforce asymptotic equality between training and test error during gradient descent, potentially enabling better stopping and hyperparameter tuning.

0 favorites 0 likes
#theory

Information Routing across Batch Boundaries: Memory--Batch Tradeoffs in Lipschitz Bandits

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper studies the joint effect of memory width and batch depth in stochastic Lipschitz bandits, characterizing the minimax pseudo-regret tradeoff up to logarithmic factors and showing that state width and update depth are not interchangeable.

0 favorites 0 likes
#theory

The Sample Complexity of Policy Learning with Mu-Resets

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper studies the sample complexity of policy learning under the mu-resets interaction protocol in reinforcement learning, resolving a question about the role of policy realizability and showing horizon dependence is exponential under all-policy concentrability and sqrt-exponential under pushforward concentrability.

0 favorites 0 likes
#theory

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper introduces a constant-aware comparison protocol for average-reward reinforcement learning regret bounds, deriving an explicit finite lower certificate for communicating MDPs and improving published coefficients.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback