标签
本研究运用信号检测理论分析程序化痕迹如何影响LLM监督者,发现详细痕迹会使决策标准向拒绝方向偏移,并增加误报率。
This paper from Carnegie Mellon researchers shows that giving an LLM judge more compute doesn't fix oversight failures when it must check many requirements in one call. It proposes sharding—dividing requirements into smaller groups handled by separate calls—which improves accuracy, resists presentation-based adversarial attacks, and can make a weaker sharded judge match a more capable holistic judge.