标签
本文在调用级别测量工具使用智能体的可靠性,揭示了在精确匹配评分下,错误传播参数由评分规则固定,并提出了一种基于状态的条件评分方法来解决此问题。
This paper proves that using error-penalized scoring rules with abstention as a discrete action can kill both the reward gradient and the KL anchor, causing models to collapse toward refusing everything. It proposes a structural repair — training a mandatory confidence report — and validates the mechanism with simulations and language model experiments.