Tag
This paper introduces a novel framework for evaluating second-order social norm reasoning in large language models, releases the NormReact dataset, and reveals that current LLMs overpredict punitive social sanctions compared to human judgments.