Provenance record

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms

arXiv cs.CY (Computers and Society) (tier 1, academic) 2026-09-11T04:00:00.000Z Original ↗

Source note: Preprints on societal impact.


Discourse valence
-11
Mixed or uncertain
confidence 76% · 4 items · range -100 to +65
Adverse readingFavourable reading
Consensus of 4 models from different labs. Spread 165 points, agreement low.
Cyber capabilityExistential and catastrophic riskGovernance and regulation
Excerpt as ingested

arXiv:2609.11373v1 Announce Type: new Abstract: Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a

Every model that read this

ModelProviderStageScoreConf.LatencyPromptWhen
Llama 3.3 70BMetaanalysis 0 70%3553ms v1.0.0 / m1.0.1 2026-09-11 05:41
GPT-4.1 miniOpenAIconsensus -10 85%4024ms v1.0.0 / m1.0.1 2026-09-11 05:42
Claude Sonnet 5Anthropicconsensus -100 60%7515ms v1.0.0 / m1.0.1 2026-09-11 05:43
Mistral Small 3.1 24BMistral AIconsensus +65 90%6916ms v1.0.0 / m1.0.1 2026-09-11 05:43
Llama 3.3 70B · reading

Study reveals limitations of AI moderation system.

evidence: primary horizon: n/a
GPT-4.1 mini · reading

The analysis reveals that the Bluesky Moderation Service employs a combination of automated detection and human oversight to moderate harmful content with high precision but low recall, meaning many harmful posts remain undetected. This balance supports content safety but also leaves significant harmful content visible, reflecting moderate societal risk and some mitigation potential through transparency and auditability.

evidence: primary horizon: near capability 40 societal 0 existential 0 economic 10
Claude Sonnet 5 · reading

This is a rigorous empirical audit revealing a real moderation system underperforms on recall despite high precision, a concrete and moderately concerning finding about AI-assisted governance at scale. The transparency enabling this audit is itself a modest positive, tempering the negative implication.

evidence: reported horizon: n/a
Mistral Small 3.1 24B · reading

The study provides a detailed audit of a transparent content moderation system. The high precision and identified harms suggest effective moderation. The low recall indicates room for improvement, but the collaborative approach and transparency are positive for societal impact.

evidence: primary horizon: n/a existential 0

Evidence extracted

The chain
SOURCE     arXiv cs.CY (Computers and Society) (tier 1)
   ↓
DOCUMENT   cd07dedf-b7cd-4f31-b697-5a318d088288
           https://arxiv.org/abs/2609.11373
   ↓
EVIDENCE   2 extracted excerpts
   ↓
MODEL RUN  4 runs, methodology 1.0.1
   ↓
SCORE      -11  (Mixed or uncertain)
   ↓
CONFIDENCE 76%