标签
This paper argues that semantic safety constraints are off-support objects not invariant under the learning problem, explaining phenomena like reward hacking and sandbox escape. It derives consequences for prior design, containment, and formal verification, using a July 2026 OpenAI–Hugging Face incident as a motivating case.
本文提出了一种无需下降和对齐的方法来测量训练后神经网络中的奇异结构。该方法从方向Fisher率中恢复死方向的阶数,将真实奇点与平坦规范对称性区分开来,并展示了该技术在Transformer和卷积层上的应用。
奇异学习理论(SLT)使用代数几何来解释为什么神经网络尽管存在退化性却能很好地泛化,引入了实对数规范阈值(RLCT)作为模型复杂度的度量。