singular-learning-theory

标签

Cards List
#singular-learning-theory

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

arXiv cs.AI · 昨天 缓存

This paper argues that semantic safety constraints are off-support objects not invariant under the learning problem, explaining phenomena like reward hacking and sandbox escape. It derives consequences for prior design, containment, and formal verification, using a July 2026 OpenAI–Hugging Face incident as a motivating case.

0 人收藏 0 人点赞
#singular-learning-theory

测量死方向:脱离规范对齐的奇异结构解构与分类

arXiv cs.LG · 2026-07-02 缓存

本文提出了一种无需下降和对齐的方法来测量训练后神经网络中的奇异结构。该方法从方向Fisher率中恢复死方向的阶数,将真实奇点与平坦规范对称性区分开来,并展示了该技术在Transformer和卷积层上的应用。

0 人收藏 0 人点赞
#singular-learning-theory

奇异学习理论:人工智能像冰融化一样学习

Reddit r/artificial · 2026-06-12 缓存

奇异学习理论(SLT)使用代数几何来解释为什么神经网络尽管存在退化性却能很好地泛化,引入了实对数规范阈值(RLCT)作为模型复杂度的度量。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈