标签
Introduces CROWN-QA, a benchmark for completeness-sensitive negative reasoning in LLMs, showing models struggle to distinguish justified negative answers from insufficient evidence, often over-closing.