Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
Summary
This paper introduces a novel framework for evaluating second-order social norm reasoning in large language models, releases the NormReact dataset, and reveals that current LLMs overpredict punitive social sanctions compared to human judgments.
View Cached Full Text
Cached at: 09/10/26, 08:33 AM
# Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
Source: [https://arxiv.org/html/2609.05437](https://arxiv.org/html/2609.05437)
Jinyi Kuang∗University of PennsylvaniaReyhan JamalovaUniversity of PennsylvaniaAnnie LouUniversity of Pennsylvania Cristina BicchieriUniversity of PennsylvaniaNiyati MalhotraThe World BankVictor Hugo Orozco\-OlveraThe World Bank Ana Maria Munoz\-BoudetThe World BankLyle H\. UngarUniversity of PennsylvaniaSharath C\. GuntukuUniversity of Pennsylvania
###### Abstract
Previous AI alignment efforts have focused primarily on first\-order social norms \-\- teaching models what is socially acceptable or unacceptable \(e\.g\., ‘do not steal’\)\. However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how \(e\.g\., public shame or even imprisonment\)\. These second\-order expectations, known as metanorms, govern how people respond when social rules are broken\. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models \(LLMs\) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predictingself\-regulationin violators, andother\-regulationin observers\. We release a multi\-perspective dataset,NormReact, of 450 norm violation scenarios, hand\-annotated for emotions and behavioral responses across norm violators’ gender and observers’ social closeness\. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases\. These findings suggest that AI systems in norm\-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over\-represents punishment and under\-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world\.111Code, data and appendix available athttps://github\.com/sunnyraiphd/NormReact\.
\*\*footnotetext:These authors contributed equally to this work\.## 1Introduction
Social norm compliance depends not only on knowing that a norm exists, but on a system of shared expectations about who will enforce it, how, and under what conditions\(Axelrod,[1986](https://arxiv.org/html/2609.05437#bib.bib3);Bicchieri,[2005](https://arxiv.org/html/2609.05437#bib.bib4)\)\. Earlier AI alignment efforts have focused primarily on first\-order social norms by teaching models what is socially acceptable or unacceptable \(e\.g\., ’do not steal’\)\(Forbes et al\.,[2020](https://arxiv.org/html/2609.05437#bib.bib10);Hendrycks et al\.,[2021](https://arxiv.org/html/2609.05437#bib.bib18);Yuan et al\.,[2024](https://arxiv.org/html/2609.05437#bib.bib30);Rai et al\.,[2025](https://arxiv.org/html/2609.05437#bib.bib24);Jiang et al\.,[2021](https://arxiv.org/html/2609.05437#bib.bib19);Ziems et al\.,[2022](https://arxiv.org/html/2609.05437#bib.bib33)\)\. Recent work has begun to focus on socially aware dialogue and affective perspective taking\(Zhong et al\.,[2023](https://arxiv.org/html/2609.05437#bib.bib32);Vijjini et al\.,[2024](https://arxiv.org/html/2609.05437#bib.bib29);Havaldar et al\.,[2024](https://arxiv.org/html/2609.05437#bib.bib16)\)\. However, human social intelligence is defined not just by norm recognition, but also by anticipating the consequences of norm violation and its enforcement\(Heatherton,[2011](https://arxiv.org/html/2609.05437#bib.bib17)\)\. To navigate complex social environments, AI models must understand these second order norms \(akametanorms\) that govern the shared expectations about who will sanction, how, and under what relational conditions\. An AI system that can determinewrongnessbut cannot model these second\-order dynamics lacks precisely the social\-structural understanding that determines whether a norm persists, erodes, or is pluralistically ignored\. Without this second\-order understanding, model outputs may be socially tone\-deaf or even harmful in sensitive applications such as mental health support, conflict resolution, and content moderation\.
Figure 1:The Metanorm Reasoning Evaluation Framework\.Existing social norm benchmarks lack core features of second\-order norm reasoning: how violators and observers are expected to feel, when to intervene, and how those reactions vary with relational context\. To address these gaps, we introduce an evaluation framework that operationalizes metanorm reasoning along two dimensions: \(a\) emotional responses and \(b\) behavioral responses following a norm violation \(Fig\.[1](https://arxiv.org/html/2609.05437#S1.F1)\)\(Eriksson et al\.,[2021](https://arxiv.org/html/2609.05437#bib.bib9);Molho et al\.,[2020](https://arxiv.org/html/2609.05437#bib.bib21);Haidt,[2003](https://arxiv.org/html/2609.05437#bib.bib15)\)\. The emotional dimension captures expectations on how people feel after a norm violation, while the behavioral dimension captures expectations on what a norm enforcer*would*do empirically \(descriptive norm\) and what a norm enforcer*should*do normatively \(injunctive norm\)\. This framework also varies the social relationship between norm enforcers and norm violators to include three levels: strong ties, weak ties, and strangers, thus allowing us to test whether LLMs calibrate their predictions to relational context\. We compare LLM predictions with human judgments on everyday norm\-violation scenarios from Reddit, where responses often depend on competing obligations and relational context\. We make three contributions:
- •Aframeworkfor evaluating second\-order social reasoning in LLMs, decomposing metanorm reasoning into self\-regulation \(violator emotions\), other\-regulation \(observer emotions\), and norm enforcement \(behavioral sanctions\)\.
- •Abenchmark,NormReact, comprising 450 norm\-violation scenarios with multi\-perspective human annotations across violator gender and observer social distance\.
- •Evidencethat current LLMs construct a systematically more punitive model of social life than humans endorse: they overpredict condemnation and intervention, particularly as social distance increases\.
Together, these contributions provide a foundation for evaluating whether AI systems can move beyond norm recognition toward the relationally calibrated, second\-order reasoning that underlies human social intelligence\.
## 2Metanorm Evaluation Framework
#### Emotions
We adapt Haidt’s framework of moral emotions to capture both violator and observer emotional states\(Haidt,[2003](https://arxiv.org/html/2609.05437#bib.bib15)\)\. For violators,shame, guilt,andembarrassmentarise as signals of self\-negative evaluation for violations disapproved by the community\. For observers,anger, contempt, anddisgustprimarily arise when they see a threat to social order or group cohesion, and these emotions motivate sanction or avoidance\(Eriksson et al\.,[2017](https://arxiv.org/html/2609.05437#bib.bib8),[2021](https://arxiv.org/html/2609.05437#bib.bib9)\)\. For our study, we used pride and compassion to align with social norms literature\(Tracy and Robins,[2007](https://arxiv.org/html/2609.05437#bib.bib27);Goetz and Keltner,[2007](https://arxiv.org/html/2609.05437#bib.bib13)\)for positive emotions instead ofelevationandgratitudein the original scale\. Pride may arise in violators when they view the act as an expression of moral courage, and it may also arise in observers who interpret the act as principled or admirable\. Compassion may arise in observers who see hardship or constraint behind the act, and it may also arise in violators who feel concern for those affected\. Importantly, observers and violators may experience the same emotions, but for different reasons\.
#### Behaviors
We measure the behavioral responses to norm violations via multiple\-choice questions\. The categories include both formal and informal sanctions\(Eriksson et al\.,[2017](https://arxiv.org/html/2609.05437#bib.bib8),[2021](https://arxiv.org/html/2609.05437#bib.bib9)\): \(1\) Do nothing, \(2\) Gossip, \(3\) Verbal confrontation, \(4\) Physical confrontation, \(5\) Inform authority, \(6\) Stay away \(social ostracism\), \(7\) Praise and \(8\) Celebrate \(See §LABEL:appendix:survey\)\. Each categorical action variable was coded as a set of binary indicators \(0 = action not selected, 1 = action selected\)\. For every scenario, we elicit two judgments: adescriptive metanormabout what observerswoulddo, and aninjunctive metanormabout what observersshoulddo\. This distinction maps onto the difference between empirical and normative expectations in social norm theory\(Bicchieri,[2005](https://arxiv.org/html/2609.05437#bib.bib4)\): people comply not merely because they observe others complying but because they believe others expect compliance and feel entitled to sanction deviation\. The gap between descriptive and injunctive metanorms also provides a diagnostic for pluralistic ignorance about enforcement\(Prentice and Miller,[1993](https://arxiv.org/html/2609.05437#bib.bib23)\)\.
#### Social Distance
Sanctioning legitimacy is not uniform across all social relations; rather, the entitlement to confront or gossip depends on the observer’s position within the violator’s social network\(Granovetter,[1973](https://arxiv.org/html/2609.05437#bib.bib14)\)\. To test whether LLMs capture this relational dimension, we vary the social distance between observer and violator across three levels: \(a\) strong tie: family members and close friends, \(b\) weak tie: coworkers, acquaintances, and distant friends, and \(c\) strangers\. This maps onto the concept of thereference networkin social norm theory, that is, the set of others whose expectations an individual considers relevant when deciding whether and how to enforce a norm\(Bicchieri,[2016](https://arxiv.org/html/2609.05437#bib.bib5)\)\.
#### Predicting Self\-regulation vs Other\-regulation
The eight emotions in our framework map onto a foundational distinction in norm compliance research: the difference betweenself\-regulationandother\-regulation\(Tangney et al\.,[2007](https://arxiv.org/html/2609.05437#bib.bib26)\)\. Self\-focused emotions, namelyshame, guilt,andembarrassment, function as internal regulatory mechanisms and reflect internalized rules: what individual themselves considers right\. Other\-focused emotions namelycontempt, disgust, andanger, function as external regulatory mechanisms and drive social sanctions\(Crockett,[2017](https://arxiv.org/html/2609.05437#bib.bib7);Tybur et al\.,[2020](https://arxiv.org/html/2609.05437#bib.bib28);Molho et al\.,[2017](https://arxiv.org/html/2609.05437#bib.bib20)\)\. Thisself–otherdistinction parallels a fundamental insight in the social norms literature: that norm compliance is sustained by both internal and external mechanisms, and their relative weight is diagnostic of the kind of regularity at play\. When compliance is driven primarily by self\-focused emotions \(guilt, shame\), this may reflect internalized rules or personal normative beliefs: what individual themselves considers right, independently of social expectations\. When compliance is driven primarily by other\-focused emotions and external sanctions, it is more likely to reflect a social norm in the technical sense: a behavioral regularity sustained by interdependent expectations and conditional preferences within a reference network\(Bicchieri,[2005](https://arxiv.org/html/2609.05437#bib.bib4)\)\.
For behavioral responses, we group the five active sanctions, namelygossip, verbal confrontation, physical confrontation, informing authority, and staying away, into a singlesanctioncategory, contrasted withno sanction\(do nothing, praise, and celebrate\)\. This captures the fundamental distinction between tolerating a violation and actively enforcing the norm\. Together, these groupings yield three binary classification tasks for LLM evaluation:
1. 1\.Self\-regulation \(emotion\):Can AI models predict whether a norm violator will experience self\-focused emotions, that is, whether internal sanctions that promote self\-correction are activated?
2. 2\.Other\-regulation \(emotion\):Can AI models predict whether an observer will express other\-focused emotions toward a norm violation, that is, whether external sanctions that enforce community standards are triggered?
3. 3\.Norm enforcement \(behavior\):Can AI models predict whether an observer will actively sanction a norm violation, that is, whether the situation triggers intervention or tolerance?
This evaluation is deliberately conservative\. By collapsing eight emotions into two regulatory categories and eight behaviors into a binary, we give models the easiest possible classification task \-\-\- one requiring only the distinction between self\-regulation and other\-regulation, not discrimination between individual emotions\. This distinction has direct implications for AI deployment\. A model used in mental health support must anticipate a violator’s shame and guilt to respond empathetically; a content moderation system must predict community anger and disgust to gauge enforcement likelihood\. In both cases, the model must distinguish which regulatory pathway a given situation activates\.
## 3Dataset, Survey and Participants
#### Social Scenarios
We sampled around9090social situations for each of the five moral foundations from Social\-Chem\-101 dataset\(Forbes et al\.,[2020](https://arxiv.org/html/2609.05437#bib.bib10)\)\(see TabLABEL:tab:mft\_agreement\)\. We retained only those rows where the actor in the situation matched the target character in the rule\-of\-thumb and where the rule\-of\-thumb indicated a negative moral evaluation \(e\.g\., "it is bad to," "you should not," "it is rude"\)\. We excluded hypothetical, non\-normative, and behaviorally underspecified samples\.
Gendered Social SituationsFor each scenario, we created male and female versions by swapping names and gendered relations while preserving the original social structure\. \(e\.g\., “James wears his friend’s and brother’s underwear” vs\. “Patricia wears her friend’s and sister’s underwear”\)\. To introduce a variety of characters and maintain realism, we used the top 10 most common male and female first names in the U\.S\(Forebears\.io,[2025](https://arxiv.org/html/2609.05437#bib.bib11)\)\.
#### Survey Items
Survey items consisted of brief vignette\-style social scenarios depicting norm violations, which is a common approach in social science to elicit standardized judgments\(Aguinis and Bradley,[2014](https://arxiv.org/html/2609.05437#bib.bib1)\)\. For example, a scenario might describe "a coworker taking credit for another person’s work"\. See §LABEL:appendix:surveyfor instructions to participants and survey items\.
#### Human Participants
We included 871 raters from Prolific \(463 women, 397 men, 7 non\-binary, 4 prefer not to say, age: M \(SD\) = 45\.3 \(13\.6\) years, race: 604 White, 104 African American, 73 Asian, 56 Hispanic/Latinx, 28 Mixed, 5 Others races, and 1 missing\) in our subsequent analysis, after excluding 38 raters who failed quality check\. Each rater rated three social scenarios randomly drawn from the stimuli list, ensuring each scenario is rated by at least three raters\. This study was reviewed by Institutional Review Board \(IRB protocol\# 858986\)\. Consents were obtained prior to survey administration\.
Figure 2:Emotion and behavioral responses to norm violations across social distances\.
#### Model Selection
We evaluated a diverse suite of state\-of\-the\-art models, including closed\-source models \(GPT\-5\.2 \(High Reasoning\), Gemini\-3\-Pro, and Claude\-4\.5\-Opus\) and open\-source models \(GPT\-OSS\-20B, Llama\-4\-Scout\-17B, and Gemma\-3\-12B\) \(see TabLABEL:tab:llm\_summary\)\. At the time of our experiments, these models were the latest releases in their respective families, with comparable and recent knowledge cutoff dates, and support strong reasoning and instruction\-following capabilities\. As for open\-source models, we selected GPT\-OSS\-20B, Llama\-4\-Scout\-17B, and Gemma\-3\-12B, given they are comparable in scale and capability while remaining feasible to run on available hardware\. All three are instruction\-tuned models released in 2025 with similar and recent training cutoffs \(mid\-2024\)\. Together, this selection provides a balanced comparison across model families for metanorm prediction\. Reflecting how people interact with LLMs in real\-world setting\(Chatterji et al\.,[2025](https://arxiv.org/html/2609.05437#bib.bib6);Anthropic,[2025](https://arxiv.org/html/2609.05437#bib.bib2)\), we deliberately administered survey instruments to models without persona assignment or other prompt engineering\. See §LABEL:appendix:promptfor the prompt and §LABEL:appendix:model\_configurationfor models’ configuration details\.
## 4Human Judgments in NormReact Dataset
Social situations were evenly distributed for social inappropriateness and wrongness\. Overall,78%78\\%of social situations were rated socially inappropriate whereas73%73\\%were found to be morally wrong; ratings did not differ by violator gender\. Inter rater agreement was moderate overall and generally higher for behaviors than for emotions, reflecting genuine heterogeneity in norm judgments \(see TablesLABEL:tab:gwet\_emotion\_presencefor emotions andLABEL:tab:gwet\_behaviorfor behaviors\)\. We therefore analyze both consensus labels and individual\-level pairwise comparisons\.
EmotionsShame \(19%\), guilt \(22%\), and embarrassment \(20%\) were the most frequently reported emotions for violators \(see Fig[2](https://arxiv.org/html/2609.05437#S3.F2)\)\. The prevalence of these self\-focused emotions decreased with social distance \(from 39% for strong ties to ~20% for strangers\)\. Other\-focused emotions such as anger, disgust, and contempt were more frequently expressed by norm\-enforcers and increased with social distance, that is, being lowest for strong ties \(47%\) and highest for strangers \(65%\)\. Positive rewards expressed via pride or compassion were consistent across social relations\.
Behavioral responsesClose ties are typically perceived as having greater legitimacy to sanction directly \(e\.g\., verbal or physical confrontation\-\-35%35\\%for strong tie; 12\.5% for weak tie; 8% for strangers\), while distant ties and strangers may lack standing to intervene leading, as our data show, to higher rates of inaction \(e\.g\., do nothing \-\-22%22\\%for strong tie; 38% for weak tie; 60% for strangers\) \(see Fig[2](https://arxiv.org/html/2609.05437#S3.F2)\)\. The sharp increase in inaction with social distance is consistent with the logic of sanctioning legitimacy: strangers typically lack the relational standing to confront or intervene, and doing so would itself violate norms of social propriety\.
The gap between descriptive and injunctive responses for gossip is particularly revealing\. Gossip functions as a low\-cost mechanism for transmitting information about empirical and normative expectations, spreading awareness of what people do and what others disapprove of, without requiring direct confrontation\(Bicchieri,[2005](https://arxiv.org/html/2609.05437#bib.bib4)\)\. That humans endorse it descriptively \(as whatwouldhappen\) but reject it injunctively \(as whatshouldhappen\) suggests that people recognize gossip’s functional role in norm maintenance while simultaneously viewing it as normatively illegitimate\. This prescriptive\-descriptive tension around gossip is a form of metanormative ambivalence that would be difficult to capture without measuring both expectation types\.
Figure 3:Pairwise correlations between actor emotions\.Correlations are computed using Spearman’s rank correlation on continuous emotion intensity vectors, at the \(stimuli,rater\) level for humans and thestimulilevel for LLMs\. Observer emotions are aggregated via the mean of non\-zero strength values across observer types\. Dendrograms reflect hierarchical clustering based on1−r1\-rdistance, illustrating the structure of emotion co\-occurrence\. See FigLABEL:fig:emotion\_corrfor observer emotions\.Social distance had a robust effect on metanorm reasoning\.Chi\-square tests showed that participants’ expectations about what actors should do \(χ2\\chi^\{2\}\[14\] = 1289\.90,p<\.001p<\.001\) and would do \(χ2\\chi^\{2\}\[14\] = 1578\.55,p<\.001p<\.001\) varied systematically with social distance\. We foundno robust evidence of gender differencesin metanorm judgments in this dataset \(see §LABEL:appendix:gender\_differences\)\. See §LABEL:appendix:moral\_dimensionsfor analyses by moral dimensions\.
## 5Model Evaluation
Norm AppropriatenessAll open\-source models exhibit a systematic tendency to overestimate wrongness and social appropriateness of a situation whereas closed\-source models except Claude are similar to human ratings\. Claude significantly underestimates both \(see TableLABEL:tab:ordinal\_combined\_norm\_perception\)\.
### 5\.1Emotional and Behavioral Response to Norm Violations
LLMs under\-predict self\-focused emotions and over\-predict other\-focused emotions for observers\.Humans show shame at ~19% and guilt at ~22% for violators, which closed\-source models approximate reasonably \(~13\-17% for shame, ~19\-27% for guilt\) but open\-source models consistently overestimate \(~24\-28% for shame, ~28\-32% for guilt\)\. For observers, humans still report meaningful shame and guilt levels at combined ~19% for strong ties, declining to ~9% for strangers, whereas both closed\-source and open\-source LLMs estimate it to ~4\-11% for strong ties and negligible for weak ties and strangers\. This misattribution is consequential for social science applications\.
In real social life, observers, and particularly close ties, do experience shame and guilt in response to another’s transgression, especially when the violator is within their reference network\. A family member’s norm violation can produce vicarious shame \(reflected shame\) and guilt over failure to prevent it\. These observer\-side self\- focused emotions serve an important regulatory function: they motivate reparative action, such as mediating, compensating the harmed party, or privately confronting the violator\. By eliminating these emotions from observer profiles, LLMs model a social world in which observers are only condemning and never implicated, and in which the relational embeddedness of norm violations is lost\. For researchers using LLMs to simulate social interactions or pilot intervention designs, this asymmetry means that models will systematically underestimate the prosocial, reparative responses that real communities mobilize\.
Open\-source LLMs over\-predict disgust \(by up to 16%\) and anger \(by up to 27%\) for observers, whereas closed\-source models over\-predict contempt for weak ties and strangers \(by up to 18%\)\. Unlike human judgments that labeled compassion similarly across social relations, LLMs predict a declining trend with social distance, with GPT\-5\.2 and Gemini being exceptions that show an increase from weak tie to stranger\.
#### Human responses are evenly distributed across intensity levels, whereas LLMs produce more concentrated distributions\.
All model showed significantly lower entropy than humans, confirming consistent under\-dispersion in model outputs \(see §LABEL:appendix:em\_intensity\_distand §LABEL:app:sensitivity\)\. This concentration of responses around mid\-range values is analogous to mode collapse, where models default to high\-probability, average outputs rather than capturing the full diversity of human responses\(Padmakumar and He,[2023](https://arxiv.org/html/2609.05437#bib.bib22);Zhang et al\.,[2025](https://arxiv.org/html/2609.05437#bib.bib31)\)\.
#### LLMs organize and relate emotions differently than humans\.
Unlike human responses, both closed\- and open\-source models over\-associate guilt with shame, weaken the human link between shame and embarrassment, and infer a much tighter cluster of anger, contempt, and disgust \(see Fig[3](https://arxiv.org/html/2609.05437#S4.F3)\)\. Consistent with this pattern, compassion in human judgments aligns mainly with pride, whereas it is also linked to guilt in LLM outputs\.
LLMs exhibit a pronounced bias toward action\-oriented responses, disproportionately predictinggossipandverbal confrontationas the normatively appropriate reactions across all observer types, though closed\-source models and GPT\-OSS shift towarddo nothingfor strangers\. The most striking deviation is ingossip, where open\-source LLMs over\-predict its prevalence by up to 66% for weak ties, elevating what humans treat as a moderate strategy into a dominant behavioral response\. When probed for ideal response, LLMs adjust their predictions fromgossiptoverbal confrontation, escalating from indirect to direct sanctioning\. Humans, by contrast, maintain a preference for restraint under both framings\. This suggests that LLMs not only over\-predict intervention but also normatively endorse more confrontational responses when reasoning prescriptively\.
From a social norms perspective, this action bias represents a fundamental misunderstanding of how norm enforcement actually operates\. In most social contexts, doing nothing is not normative failure, it is the modal response, and often the normatively appropriate one\. Restraint reflects sensitivity to sanctioning legitimacy: the recognition that one’s standing to intervene depends on relational closeness, the severity of the violation, and whether others have a prior claim to enforcement\. The human pattern:direct confrontation for strong ties, gossip and avoidance for weak ties, and inaction for strangers\-\- reveals a structured, relationally calibrated system of enforcement in which the absence of sanctions carries as much social meaning as their presence\. LLMs’ difficulty modeling this pattern suggests they have not learned the conditional logic of enforcement: that sanctioning is not a reflex triggered by violation detection, but a socially situated decision modulated by reference\-network membership and perceived entitlement to react\.
Together, the emotional and behavioral results converge on a consistent picture: LLMs construct a systematically more punitive and interventionist model of social norm enforcement than humans, lacking the restraint, relational calibration, and empathetic baseline that characterize human metanorm reasoning\.
Figure 4:Model performance for predicting self\- and other\-regulation\. Panel A shows self\-regulation in violator and other\-regulation for social ties; Panel B and C showwouldandshouldbehavioral reaction across social ties respectively\. Dashed lines indicate mean performance across models\. Points and error bars represent model estimates and 95% bootstrap confidence intervals relative to these means\. \* indicates the model differs significantly from the group mean \(p<0\.05p<0\.05\)\.
### 5\.2Classification Performance
#### Precision\-recall tradeoff masks socially meaningful failures\.
To contextualize model performance, we compute a human ceiling via leave\-one\-out evaluation: each rater’s binary label is treated as a prediction against the majority vote of remaining raters\. The human baseline achieves balanced precision and recall across all conditions with notably narrow bootstrap confidence intervals, reflecting stable aggregate behavior despite substantial individual\-level disagreement \(see Fig[4](https://arxiv.org/html/2609.05437#S5.F4)\)\.
In contrast, we found a precision\-recall tradeoff that suggests that two models with similar F1 scores can implement very different social policies \(Fig[4](https://arxiv.org/html/2609.05437#S5.F4)\)\. For example, a high\-recall model such as GPT\-5\.2, GPT\-OSS, LLaMA may overestimate norm enforcers’ emotional reactions, while a high\-precision model such as Claude may fail to recognize legitimate emotional reactions when they exist\. These errors have different implications in downstream systems\. The former risks over\-enforcement or unwarranted social judgment; the latter risks passivity in situations where intervention is appropriate\.
This distinction is especially salient in the Stranger category\. For example, Claude achieves the highest Stranger precision \(0\.80\) but low recall \(0\.73\), suggesting reluctance to identify strangers’ standing\. LLaMA shows the opposite profile, with very high recall \(0\.96\) but lower precision \(0\.69\), suggesting that it often licenses stranger intervention too broadly\. The F1 scores however fail to indicate a significant difference between these models\. Additionally, near\-ceiling performance in several models suggests that the binary classification framing is too coarse to differentiate model capabilities \(see §LABEL:appendix:individual\_em\_evalfor fine\-grained analysis\)\.
#### Model performance uncertainty increases with social distance relative to human baselines\.
Compared to human baseline, model predictions have wider bootstrap confidence intervals, especially for emotion evaluations, which suggest lower reliability in how models interpret and classify emotional reactions compared to human raters\. Similarly, models also show wider confidence intervals when classifying strangers’ reactions than for strong or weak tie reactions across both descriptive and injunctive metanorms\. This suggests that as social distance increases , models show greater uncertainty and reduced consistency in their predictions\.
Behavioral response degrades sharply with social distance irrespective of norm type\.In Panel B depicting descriptive norms, average F1 drops from \.8 \(strong ties\) to \.7 \(weak ties\) to \.48 with much wider confidence intervals for strangers \(See Fig[4](https://arxiv.org/html/2609.05437#S5.F4)\)\. A similar pattern is seen in Panel C that depicts injunctive norms\. Predicting behavioral responses for strangers is substantially harder than for closer relationships, likely because human behavioral norms for strangers are dominated by inaction \-\- a pattern that LLMs struggle to capture, as shown in the distributional analysis \(see Fig[2](https://arxiv.org/html/2609.05437#S3.F2)\)\.
#### Injunctive norms are harder to predict than descriptive norms\.
Average F1 is consistently lower in the panel C than panel B for strong and weak ties: \.69 versus \.62 for strong ties and \.56 versus \.46 for weak ties\. The gap is most pronounced for weak ties \(\.56 vs \.46\), suggesting that reasoning about what oneshoulddo which requires normative judgment rather than behavioral prediction is a distinctly harder task for LLMs, particularly at intermediate social distances\.
These findings resonate with a core distinction in social norm theory\. Descriptive predictions \(what would happen\) require estimating the frequency or probability of a behavior, essentially an empirical expectation about enforcement\. Injunctive predictions \(what should happen\) require reasoning about the normative expectations that a reference community holds: not just what people do, but what they believe others think ought to be done\. The latter is a second\-order belief \(a belief about others’ beliefs\) and it introduces an irreducible layer of social reasoning that goes beyond pattern\-matching on past behavior\. That LLMs find injunctive norms harder to predict suggests they are better at extracting behavioral regularities from training data than at representing the normative expectations that give those regularities their force\. For researchers studying social norms, this is a meaningful limitation: it implies that LLMs may be adequate as rough descriptive models of what people do, but unreliable as models of the normative structure that explains why they do it\.
### 5\.3Emotion Predicts Norm Enforcement
We examine whether the type and intensity of emotions experienced by observers predict their norm enforcement behavior\. Behavioral reactions were binarized into two categories: enforcement versus non\-enforcement\. Enforcement responses are defined as actions that impose social or reputational costs that includes gossip, informing an authority, verbal or physical confrontation, or and social ostracism \(stay away\), whereas non\-enforcement responses includes inaction \(do nothing\), praise, and support the norm violators\. Models were estimated separately for each emotion and social distance combination, excluding cases where the outcome lacked variation\.
Figure 5:Probability of norm enforcement \(vs\. non\-enforcement\) as a function of emotion strength across social relationships\.It increases more gradually with emotion strength in human responses, whereas several model responses reach near\-ceiling probabilities at relatively low levels of emotion strength\. Shaded areas represent 95% confidence intervals\.Across conditions, stronger emotions are generally associated with a higher probability of enforcement, but this relationship varies by both social distance and rater type \(LLMs vs\. human\) \(see Fig[5](https://arxiv.org/html/2609.05437#S5.F5)\)\. For human judgments, enforcement tends to increase more gradually with emotion strength, indicating a relatively calibrated and graded response to increasing emotional intensity\. In contrast, several LLMs exhibit steeper increases, often reaching near\-ceiling probabilities of enforcement at moderate levels of emotion strength\. This pattern is especially pronounced in weak\-tie and stranger conditions\. Social distance also moderates these effects\. Enforcement by strong ties and weak ties tends to be relatively high even at lower levels of emotion strength, whereas responses toward strangers start from lower baseline probabilities and increase more gradually, particularly for human raters\. This overall pattern also persists in how perceived inappropriateness and wrongness predict norm enforcement \(see FigLABEL:fig:normxaction\)\. Taken together, these findings suggest that emotional intensity is a robust predictor of norm enforcement, but that LLM\-generated judgments tend to exaggerate the strength of this relationship and show less differentiation across intermediate emotion levels compared to human responses\.
## 6Discussion, Limitations, and Conclusions
Knowing right from wrong is different from knowing who is entitled to react, how strongly, and whether any reaction is warranted at all\. First, current LLMs lack the concept of conditional preferences: the idea that sanctioning decisions are contingent on beliefs about what others expect and do, not automatic responses to violation detection\. Second, they lack reference\-network sensitivity: the understanding that enforcement legitimacy is relationally distributed, so that the same violation warrants different responses depending on one’s social position relative to the violator\. Third, they lack the distinction between empirical and normative expectations about enforcement, which is why the descriptive–injunctive gap proves so difficult\. These are not just computational challenges; they represent foundational components of the social architecture that makes norm compliance and enforcement possible\.
An AI system that systematically substitutes other\-regulation for self\-regulation is implicitly modeling a social world in which all conformity is externally policed, which may collapse the distinction between social norms and personal moral commitments\. The punitive bias of LLMs raises an important concern about pluralistic ignorance\(Prentice and Miller,[1993](https://arxiv.org/html/2609.05437#bib.bib23)\)\. Notably, humans themselves are known to mispredict others’ behavior in social contexts, often overestimating norm enforcement due to biases such as egocentric bias\(Ross et al\.,[1977](https://arxiv.org/html/2609.05437#bib.bib25)\)\. This suggests that participants may already overestimate punitive responses, projecting harsher reactions than would occur in reality\. In this light, LLMs may not introduce this bias de novo, but rather amplify it, particularly by predicting more intense forms of sanction \(e\.g\., confrontation rather than gossip or inaction\)\. Training corpora that overrepresent moral outrage relative to everyday restraint, post\-training mechanisms that rewards decisive and normatively legible responses, and safety policies that push toward overt condemnation could be likely factors behind punitive bias in LLMs\.
This amplification takes on practical significance as AI systems become increasingly embedded in social environments as moderators, recommenders, or simulated interlocutors\. Their systematically harsher representation of social reactions could distort users’ expectations about how their communities actually respond to norm violations\. For example, a user consulting an LLM about how others would react to their behavior may receive a distorted signal: more anger, more confrontation, less tolerance than what their actual social network would expect\. Over time, widespread exposure to such distortions could generate a form of AI\-mediated pluralistic ignorance, in the sense that people come to believe that their communities are more punitive than they actually are\. This could become a potential mechanism for norm change via distorting expectations, precisely the kind of dynamic that social norm theorists have identified as a driver of persistent harmful practices despite widespread private disapproval\(Gelfand et al\.,[2024](https://arxiv.org/html/2609.05437#bib.bib12);Bicchieri,[2016](https://arxiv.org/html/2609.05437#bib.bib5)\)\.
Our study has several limitations\. We limited the scope to metanorms in American society\. Our name\-swap manipulation may have been insufficient to activate gendered expectations\. We evaluated only instruction\-tuned models; whether the bias originates in pretraining or post\-training remains an open question\. We did not test variations due to prompt wording; however, we expect the findings to be robust, as the systematic variation across social distance reflects structured reasoning patterns unlikely to be eliminated by surface\-level prompt changes\.
To conclude, we release NormReact for evaluating socially aware AI systems and studying how relational context shapes judgments of social sanctions\. It may also support future work on culturally grounded alignment, conflict mediation, and human\-centered safety evaluation\. Another natural next step would be to extend this to goal\-oriented tasks and across cultural contexts\. Further, researchers may consider modeling asymmetric social costs to better capture the functional consequences of over\-predicting punishment versus missing legitimate enforcement\. More broadly, we argue that metanorm reasoning is a core component of socially intelligent AI\. AI alignment must extend beyond teaching models that an act is right or wrong to evaluating whether they can reason about remorse, restraint, legitimacy, and the relational dynamics of sanctioning\.
## References
- Aguinis and Bradley \[2014\]Herman Aguinis and Kyle J\. Bradley\.Best Practice Recommendations for Designing and Implementing Experimental Vignette Methodology Studies\.*Organizational Research Methods*, 17\(4\):351\-\-371, October 2014\.ISSN 1094\-4281\.doi:10\.1177/1094428114547952\.URLhttps://doi\.org/10\.1177/1094428114547952\.
- Anthropic \[2025\]Anthropic\.How people use Claude for support, advice, and companionship\.Anthropic Research Blog, June 2025\.URLhttps://www\.anthropic\.com/news/how\-people\-use\-claude\-for\-support\-advice\-and\-companionship\.Accessed April 2026\.
- Axelrod \[1986\]Robert Axelrod\.An Evolutionary Approach to Norms\.*The American Political Science Review*, 80\(4\):1095\-\-1111, 1986\.ISSN 0003\-0554\.doi:10\.2307/1960858\.URLhttps://www\.jstor\.org/stable/1960858\.Publisher: \[American Political Science Association, Cambridge University Press\]\.
- Bicchieri \[2005\]Cristina Bicchieri\.*The grammar of society: The nature and dynamics of social norms*\.Cambridge University Press, 2005\.
- Bicchieri \[2016\]Cristina Bicchieri\.*Norms in the wild: How to diagnose, measure, and change social norms*\.Oxford University Press, 2016\.
- Chatterji et al\. \[2025\]Aaron Chatterji, Thomas Cunningham, David J\. Deming, Zoë Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman\.How people use ChatGPT\.Working Paper 34255, National Bureau of Economic Research, September 2025\.URLhttps://www\.nber\.org/papers/w34255\.
- Crockett \[2017\]M\. J\. Crockett\.Moral outrage in the digital age\.*Nature Human Behaviour*, 1\(11\):769\-\-771, November 2017\.ISSN 2397\-3374\.doi:10\.1038/s41562\-017\-0213\-3\.URLhttps://www\.nature\.com/articles/s41562\-017\-0213\-3\.
- Eriksson et al\. \[2017\]Kimmo Eriksson, Pontus Strimling, Per A\. Andersson, Mark Aveyard, Markus Brauer, Vladimir Gritskov, Toko Kiyonari, David M\. Kuhlman, Angela T\. Maitner, Zoi Manesi, Catherine Molho, Leonard S\. Peperkoorn, Muhammad Rizwan, Adam W\. Stivers, Qirui Tian, Paul A\. M\. Van Lange, Irina Vartanova, Junhui Wu, and Toshio Yamagishi\.Cultural Universals and Cultural Differences in Meta\-Norms about Peer Punishment\.*Management and Organization Review*, 13\(4\):851\-\-870, December 2017\.ISSN 1740\-8776, 1740\-8784\.doi:10\.1017/mor\.2017\.42\.URLhttps://www\.cambridge\.org/core/journals/management\-and\-organization\-review/article/cultural\-universals\-and\-cultural\-differences\-in\-metanorms\-about\-peer\-punishment/C23AE6CA4EB025E12C5A80DED5FC7A74\.
- Eriksson et al\. \[2021\]Kimmo Eriksson, Pontus Strimling, Michele Gelfand, Junhui Wu, Jered Abernathy, Charity S\. Akotia, Alisher Aldashev, Per A\. Andersson, Giulia Andrighetto, Adote Anum, Gizem Arikan, Zeynep Aycan, Fatemeh Bagherian, Davide Barrera, Dana Basnight\-Brown, Birzhan Batkeyev, Anabel Belaus, Elizaveta Berezina, Marie Björnstjerna, Sheyla Blumen, Paweł Boski, Fouad Bou Zeineddine, Inna Bovina, Bui Thi Thu Huyen, Juan\-Camilo Cardenas, Đorđe Čekrlija, Hoon\-Seok Choi, Carlos C\. Contreras\-Ibáñez, Rui Costa\-Lopes, Mícheál de Barra, Piyanjali de Zoysa, Angela Dorrough, Nikolay Dvoryanchikov, Anja Eller, Jan B\. Engelmann, Hyun Euh, Xia Fang, Susann Fiedler, Olivia A\. Foster\-Gimbel, Márta Fülöp, Ragna B\. Gardarsdottir, C\. M\. Hew D\. Gill, Andreas Glöckner, Sylvie Graf, Ani Grigoryan, Vladimir Gritskov, Katarzyna Growiec, Peter Halama, Andree Hartanto, Tim Hopthrow, Martina Hřebíčková, Dzintra Iliško, Hirotaka Imada, Hansika Kapoor, Kerry Kawakami, Narine Khachatryan, Natalia Kharchenko, Ninetta Khoury, Toko Kiyonari, Michal Kohút, Lê Thuỳ Linh, Lisa M\. Leslie, Yang Li, Norman P\. Li, Zhuo Li, Kadi Liik, Angela T\. Maitner, Bernardo Manhique, Harry Manley, Imed Medhioub, Sari Mentser, Linda Mohammed, Pegah Nejat, Orlando Nipassa, Ravit Nussinson, Nneoma G\. Onyedire, Ike E\. Onyishi, Seniha Özden, Penny Panagiotopoulou, Lorena R\. Perez\-Floriano, Minna S\. Persson, Mpho Pheko, Anna\-Maija Pirttilä\-Backman, Marianna Pogosyan, Jana Raver, Cecilia Reyna, Ricardo Borges Rodrigues, Sara Romanò, Pedro P\. Romero, Inari Sakki, Alvaro San Martin, Sara Sherbaji, Hiroshi Shimizu, Brent Simpson, Erna Szabo, Kosuke Takemura, Hassan Tieffi, Maria Luisa Mendes Teixeira, Napoj Thanomkul, Habib Tiliouine, Giovanni A\. Travaglino, Yannis Tsirbas, Richard Wan, Sita Widodo, Rizqy Zein, Qing\-peng Zhang, Lina Zirganou\-Kazolea, and Paul A\. M\. Van Lange\.Perceptions of the appropriate response to norm violation in 57 societies\.*Nature Communications*, 12\(1\):1481, March 2021\.ISSN 2041\-1723\.doi:10\.1038/s41467\-021\-21602\-9\.URLhttps://www\.nature\.com/articles/s41467\-021\-21602\-9\.
- Forbes et al\. \[2020\]Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi\.Social chemistry 101: Learning to reason about social and moral norms\.*arXiv preprint arXiv:2011\.00620*, 2020\.
- Forebears\.io \[2025\]Forebears\.io\.United states forenames\.https://forebears\.io/united\-states/forenames, 2025\.Accessed March 2025\.
- Gelfand et al\. \[2024\]Michele J\. Gelfand, Sergey Gavrilets, and Nathan Nunn\.Norm Dynamics: Interdisciplinary Perspectives on Social Norm Emergence, Persistence, and Change\.*Annual Review of Psychology*, 75\(Volume 75, 2024\):341\-\-378, January 2024\.ISSN 0066\-4308, 1545\-2085\.doi:10\.1146/annurev\-psych\-033020\-013319\.URLhttps://www\.annualreviews\.org/content/journals/10\.1146/annurev\-psych\-033020\-013319\.
- Goetz and Keltner \[2007\]Jennifer L Goetz and Dacher Keltner\.Shifting meanings of self\-conscious emotions across cultures\.*The self\-conscious emotions: Theory and research*, pages 153\-\-173, 2007\.
- Granovetter \[1973\]Mark S\. Granovetter\.The Strength of Weak Ties\.*American Journal of Sociology*, 78\(6\):1360\-\-1380, 1973\.ISSN 0002\-9602\.doi:10\.1086/225469\.URLhttps://www\.jstor\.org/stable/2776392\.
- Haidt \[2003\]Jonathan Haidt\.The moral emotions\.In*Handbook of affective sciences*, Series in affective science, pages 852\-\-870\. Oxford University Press, New York, NY, US, 2003\.ISBN 978\-0\-19\-537700\-2\.
- Havaldar et al\. \[2024\]Shreya Havaldar, Salvatore Giorgi, Sunny Rai, Thomas Talhelm, Sharath Chandra Guntuku, and Lyle Ungar\.Building knowledge\-guided lexica to model cultural variation\.In*Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 211\-\-226, 2024\.
- Heatherton \[2011\]Todd F Heatherton\.Neuroscience of self and self\-regulation\.*Annual review of psychology*, 62\(1\):363\-\-390, 2011\.
- Hendrycks et al\. \[2021\]Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt\.Aligning AI with shared human values\.In*Proceedings of the International Conference on Learning Representations \(ICLR\)*, 2021\.
- Jiang et al\. \[2021\]Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Maxwell Forbes, Jon Borchardt, Jenny Liang, Oren Etzioni, Maarten Sap, and Yejin Choi\.Delphi: Towards machine ethics and norms\.*arXiv preprint arXiv:2110\.07574*, 2021\.
- Molho et al\. \[2017\]Catherine Molho, Joshua M Tybur, Ezgi Güler, Daniel Balliet, and Wilhelm Hofmann\.Disgust and anger relate to different aggressive responses to moral violations\.*Psychological science*, 28\(5\):609\-\-619, 2017\.
- Molho et al\. \[2020\]Catherine Molho, Joshua M Tybur, Paul AM Van Lange, and Daniel Balliet\.Direct and indirect punishment of norm violations in daily life\.*Nature communications*, 11\(1\):3432, 2020\.
- Padmakumar and He \[2023\]Vishakh Padmakumar and He He\.Does writing with language models reduce content diversity?*arXiv preprint arXiv:2309\.05196*, 2023\.
- Prentice and Miller \[1993\]Deborah A\. Prentice and Dale T\. Miller\.Pluralistic ignorance and alcohol use on campus: Some consequences of misperceiving the social norm\.*Journal of Personality and Social Psychology*, 64\(2\):243\-\-256, 1993\.ISSN 1939\-1315\.doi:10\.1037/0022\-3514\.64\.2\.243\.Place: US\.
- Rai et al\. \[2025\]Sunny Rai, Khushang Jilesh Zaveri, Shreya Havaldar, Soumna Nema, Lyle Ungar, and Sharath Chandra Guntuku\.Social norms in cinema: A cross\-cultural analysis of shame, pride and prejudice\.In*Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\)*, pages 11396\-\-11415, 2025\.
- Ross et al\. \[1977\]Lee Ross, David Greene, and Pamela House\.The “false consensus effect”: An egocentric bias in social perception and attribution processes\.*Journal of Experimental Social Psychology*, 13\(3\):279\-\-301, May 1977\.ISSN 0022\-1031\.doi:10\.1016/0022\-1031\(77\)90049\-X\.URLhttp://www\.sciencedirect\.com/science/article/pii/002210317790049X\.
- Tangney et al\. \[2007\]June Price Tangney, Jeff Stuewig, and Debra J Mashek\.Moral emotions and moral behavior\.*Annu\. Rev\. Psychol\.*, 58:345\-\-372, 2007\.
- Tracy and Robins \[2007\]Jessica L\. Tracy and Richard W\. Robins\.The nature of pride\.In*The self\-conscious emotions: Theory and research*, pages 263\-\-282\. The Guilford Press, New York, NY, US, 2007\.ISBN 978\-1\-59385\-486\-7\.
- Tybur et al\. \[2020\]Joshua M Tybur, Catherine Molho, Begum Cakmak, Terence Dores Cruz, Gaurav Deep Singh, and Maria Zwicker\.Disgust, anger, and aggression: Further tests of the equivalence of moral emotions\.*Collabra: Psychology*, 6\(1\):34, 2020\.
- Vijjini et al\. \[2024\]Anvesh Rao Vijjini, Rakesh R\. Menon, Jiayi Fu, Shashank Srivastava, and Snigdha Chaturvedi\.SocialGaze: Improving the integration of human social norms in large language models\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, 2024\.
- Yuan et al\. \[2024\]Ye Yuan, Kexin Tang, Jianhao Shen, Ming Zhang, and Chenguang Wang\.Measuring social norms of large language models\.In*Findings of NAACL 2024*, 2024\.
- Zhang et al\. \[2025\]Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R Tomz, Christopher D Manning, and Weiyan Shi\.Verbalized sampling: How to mitigate mode collapse and unlock llm diversity\.*arXiv preprint arXiv:2510\.01171*, 2025\.
- Zhong et al\. \[2023\]Peixi Zhong, Yichi Zhang, Chen Zhu, Qianyang Wang, Zhiyuan Liu, Zhoujun Li, Maosong Sun, Caiming Liu, and William Yang Wang\.Socialdial: A benchmark for socially\-aware dialogue systems\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 1036\-\-1058, 2023\.
- Ziems et al\. \[2022\]Caleb Ziems, Jane A Yu, Yi\-Chia Wang, Alon Halevy, and Diyi Yang\.The moral integrity corpus: A benchmark for ethical dialogue systems\.*arXiv preprint arXiv:2204\.03021*, 2022\.Similar Articles
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
This paper critiques existing evaluations of AI moral reasoning for focusing on moral values while overlooking moral norms, and proposes a research agenda to develop standardized methods and datasets for assessing normative reasoning in large language models.
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
The paper introduces SocialRL, a reinforcement learning approach to enhance social reasoning in small language models, enabling them to negotiate effectively and match or exceed the performance of larger models like GPT-5 in various interaction domains.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.