标签
This paper introduces BenchDrift, a method for quantifying how LLM benchmark performance changes when problems are rephrased without changing meaning or answer. It shows that rephrasing causes bidirectional correctness flips across models and benchmarks, with stronger models becoming more sensitive to phrasing.
描述了PHI // DRIFT,一种认知架构,具有七个在会话之间漂移的稳态状态变量,记忆按情感显著性和时间衰减评分,以及一个荣格阴影模块,构建在仅使用CPU的迷你塔上,并作为预印本提交到SSRN。