Tag
This paper studies behavioral detection of unfaithful chain-of-thought reasoning in LLMs, finding that answer correctness structures detection performance: on incorrect answers, where most unfaithfulness occurs, behavioral signals are at chance, while on correct answers they offer modest separation.
This paper investigates the two axes of LLM abstention: answer correctness and question answerability. It shows that a single confidence threshold conflates these two failure modes, and proposes a three-class selective acceptance framework with separate budgets. Experiments across five instruction-tuned models reveal that answerability is internally legible but poorly captured by output confidence or self-assessments.