Tag
This paper introduces a novel framework for evaluating second-order social norm reasoning in large language models, releases the NormReact dataset, and reveals that current LLMs overpredict punitive social sanctions compared to human judgments.
NormAct is a benchmark that evaluates embodied planning agents on hidden social norm compliance, revealing that state-of-the-art MLLMs achieve 67.3% goal achievement but only 26.4% norm compliance, and proposes NormPerceptor to improve task success from 24.2% to 46.7%.