Provenance record

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

arXiv cs.AI (tier 1, academic) 2026-09-10T04:00:00.000Z Original ↗

Source note: Preprints. Not peer reviewed at time of posting.


Discourse valence
-50
Adverse
confidence 78% · 3 items · range -55 to -45
Adverse readingFavourable reading
Consensus of 3 models from different labs. Spread 10 points, agreement high.
Existential and catastrophic riskLoss of control and alignment
Excerpt as ingested

arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.

Every model that read this

ModelProviderStageScoreConf.LatencyPromptWhen
Llama 3.3 70BMetaanalysis -50 70%3355ms v1.0.0 / m1.0.1 2026-09-10 09:18
Mistral Small 3.1 24BMistral AIconsensus -55 80%6129ms v1.0.0 / m1.0.1 2026-09-10 09:20
gpt-oss 120BOpenAIconsensus -45 85%17332ms v1.0.0 / m1.0.1 2026-09-10 09:20
Llama 3.3 70B · reading

LLMs may produce distorted social regulation pictures.

evidence: speculative horizon: n/a
Mistral Small 3.1 24B · reading

The paper identifies a significant distortion in how LLMs perceive and predict social regulation. This distortion could lead to misjudgments in norm-sensitive applications, posing risks to social harmony and fairness. The findings suggest a material concern for human flourishing.

evidence: primary horizon: n/a
gpt-oss 120B · reading

Empirical findings show LLMs systematically exaggerate punitive responses, which could lead to harsher outcomes in applications that rely on accurate social reasoning, posing a materially concerning negative impact on human flourishing.

evidence: primary horizon: n/a capability 0

Evidence extracted

The chain
SOURCE     arXiv cs.AI (tier 1)
   ↓
DOCUMENT   c11e0a25-c88e-4c4e-be8a-7cb6422413ab
           https://arxiv.org/abs/2609.05437
   ↓
EVIDENCE   2 extracted excerpts
   ↓
MODEL RUN  3 runs, methodology 1.0.1
   ↓
SCORE      -50  (Adverse)
   ↓
CONFIDENCE 78%