Tag
This paper measures the reliability of tool-using agents at the invocation level, revealing that under exact-match scoring, error propagation parameters are fixed by the scoring rule and proposes a conditional-on-state scoring method to address this.
This paper proves that using error-penalized scoring rules with abstention as a discrete action can kill both the reward gradient and the KL anchor, causing models to collapse toward refusing everything. It proposes a structural repair — training a mandatory confidence report — and validates the mechanism with simulations and language model experiments.