Tag
This paper proposes a two-stage approach for early failure alerting in dialogs and LLM-agent trajectories, addressing the challenge of sparse evidence by learning turn-level failure evidence from trajectory labels and using an attention-based predictor with a preference-conditioned stopping policy (α-STOP) to achieve controllable accuracy-earliness trade-offs.