Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
Source note: Preprints. Not peer reviewed at time of posting.
arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.
Every model that read this
| Model | Provider | Stage | Score | Conf. | Latency | Prompt | When |
|---|---|---|---|---|---|---|---|
| Llama 3.3 70B | Meta | analysis | -50 | 70% | 3355ms | v1.0.0 / m1.0.1 | 2026-09-10 09:18 |
| Mistral Small 3.1 24B | Mistral AI | consensus | -55 | 80% | 6129ms | v1.0.0 / m1.0.1 | 2026-09-10 09:20 |
| gpt-oss 120B | OpenAI | consensus | -45 | 85% | 17332ms | v1.0.0 / m1.0.1 | 2026-09-10 09:20 |
LLMs may produce distorted social regulation pictures.
The paper identifies a significant distortion in how LLMs perceive and predict social regulation. This distortion could lead to misjudgments in norm-sensitive applications, posing risks to social harmony and fairness. The findings suggest a material concern for human flourishing.
Empirical findings show LLMs systematically exaggerate punitive responses, which could lead to harsher outcomes in applications that rely on accurate social reasoning, posing a materially concerning negative impact on human flourishing.
Evidence extracted
- Current LLMs overpredict negative sanctions
- LLMs portray a harsher social world
SOURCE arXiv cs.AI (tier 1)
↓
DOCUMENT c11e0a25-c88e-4c4e-be8a-7cb6422413ab
https://arxiv.org/abs/2609.05437
↓
EVIDENCE 2 extracted excerpts
↓
MODEL RUN 3 runs, methodology 1.0.1
↓
SCORE -50 (Adverse)
↓
CONFIDENCE 78%