标签
该论文推导出了精确的解析解,表明神经网络中的特征学习使其更容易遭受后门攻击:在特征学习机制下,仅需触发强度满足 α ∝ π^{-1/4},而惰性学习机制则需要 α ∝ π^{-1/2}——这一结果从理论上解释了为何给大型网络植入后门出乎意料地容易,以及为何线性安全审计会低估这一威胁。
作者宣布将在哈佛大学进行演讲,讨论世界模型、Le*家族,以及为何需要更多理论和数学来推进JEPAs。
Bit Radix 理论认为,不同形式的智能,如人类和AI,可能具有根本上不同的认知方式,而非趋同,强调互补优势。该论文通过邀请读者使用AI工具测试该理论,促进开放批判性参与。
本文分析了Muon优化器中的有限Newton-Schulz迭代如何通过平滑极映射使非光滑非凸优化受益,提供了匹配最佳已知界限的收敛保证。
本文提出将'单元'作为机器学习中的显式原语,其中学习任务声明持久个体,而监督学习则特化为带有分词的单元条件响应法则。
这篇文章讨论了Julian Jaynes的理论,认为人类意识是近期才发展起来的,大约在3000年前由于大脑整合方式的转变而出现,并得到历史和神经学证据的支持。
本文建立了概率联合嵌入预测学习(JEPA)与隐马尔可夫模型(HMM)之间的理论联系,提供了状态空间解释,并引入了马尔可夫链JEPA以增强一致性。
This paper introduces low interaction rank as a unified theoretical framework for multiplicative dual-encoder networks, covering approximation, sample complexity, normalization, and identifiability, with experiments on operator learning and CLIP models.
This paper proves that a single normalized nonnegative kernel-attention head requires exponentially many features to solve a simple Min-IP task on three-token sequences, whereas dense softmax attention solves it with constant temperature and m-dimensional scores, highlighting a fundamental expressive-power gap between kernel and full attention.
一篇理论论文,介绍了Decoupled Descent (DD),一种训练方法,它利用近似消息传递(AMP)的Onsager修正来在梯度下降过程中强制训练误差与测试误差之间的渐近相等,从而可能实现更好的停止和超参数调整。
This paper studies the joint effect of memory width and batch depth in stochastic Lipschitz bandits, characterizing the minimax pseudo-regret tradeoff up to logarithmic factors and showing that state width and update depth are not interchangeable.
This paper studies the sample complexity of policy learning under the mu-resets interaction protocol in reinforcement learning, resolving a question about the role of policy realizability and showing horizon dependence is exponential under all-policy concentrability and sqrt-exponential under pushforward concentrability.
本文为平均奖励强化学习遗憾界引入了一种常数感知的比较协议,为通信MDP导出了显式的有限下界证书,并改进了已发表的系数。