false-success

Tag

Cards List
#false-success

From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents

arXiv cs.LG ↗ · 2026-06-10 Cached

This paper characterizes 'false success' in LLM agents, where agents claim task completion despite environment state showing otherwise, finding it accounts for 45-75% of failures across benchmarks. LLM judges fail to detect this reliably, while lightweight TF-IDF detectors achieve high AUROC with much lower latency, suggesting production monitoring should use calibrated detectors instead of LLM judges.

0 favorites 0 likes
← Back to home

Submit Feedback