Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms
Source note: Preprints on societal impact.
arXiv:2609.11373v1 Announce Type: new Abstract: Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies). Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5$\times$ more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content. Our work offers a look into the operational realities of a deployed moderation system, providing a
Every model that read this
| Model | Provider | Stage | Score | Conf. | Latency | Prompt | When |
|---|---|---|---|---|---|---|---|
| Llama 3.3 70B | Meta | analysis | 0 | 70% | 3553ms | v1.0.0 / m1.0.1 | 2026-09-11 05:41 |
| GPT-4.1 mini | OpenAI | consensus | -10 | 85% | 4024ms | v1.0.0 / m1.0.1 | 2026-09-11 05:42 |
| Claude Sonnet 5 | Anthropic | consensus | -100 | 60% | 7515ms | v1.0.0 / m1.0.1 | 2026-09-11 05:43 |
| Mistral Small 3.1 24B | Mistral AI | consensus | +65 | 90% | 6916ms | v1.0.0 / m1.0.1 | 2026-09-11 05:43 |
Study reveals limitations of AI moderation system.
The analysis reveals that the Bluesky Moderation Service employs a combination of automated detection and human oversight to moderate harmful content with high precision but low recall, meaning many harmful posts remain undetected. This balance supports content safety but also leaves significant harmful content visible, reflecting moderate societal risk and some mitigation potential through transparency and auditability.
This is a rigorous empirical audit revealing a real moderation system underperforms on recall despite high precision, a concrete and moderately concerning finding about AI-assisted governance at scale. The transparency enabling this audit is itself a modest positive, tempering the negative implication.
The study provides a detailed audit of a transparent content moderation system. The high precision and identified harms suggest effective moderation. The low recall indicates room for improvement, but the collaborative approach and transparency are positive for societal impact.
Evidence extracted
- Bluesky Moderation Service operates with high precision (0.837)
- Bluesky Moderation Service struggles with low recall (0.222)
SOURCE arXiv cs.CY (Computers and Society) (tier 1)
↓
DOCUMENT cd07dedf-b7cd-4f31-b697-5a318d088288
https://arxiv.org/abs/2609.11373
↓
EVIDENCE 2 extracted excerpts
↓
MODEL RUN 4 runs, methodology 1.0.1
↓
SCORE -11 (Mixed or uncertain)
↓
CONFIDENCE 76%