Responding to Online Harm through Community-Based Content Moderation
- Byun, Eujean
- Advisor(s): Yeo, Lisa
Abstract
Online harassment, hate speech, and misinformation pose persistent challenges for digital platforms. As centralized moderation faces well-documented limits in scale and accuracy, platforms increasingly turn to community-based systems that mobilize ordinary users to identify and correct harmful content. Yet what makes such systems effective remains poorly understood. This dissertation examines community-based content moderation across three levels of analysis: platform, individual, and collective. It draws on archival data from X’s Community Notes (1.4 million notes posted between 2021 and 2025) and a controlled experiment on EatSnap.Love, a simulated social media platform (N = 137).Chapter 3 investigates how platform design parameters shape bystander intervention. Drawing on bystander intervention theory, it theorizes how disapproval channel availability and reporting effort jointly structure participation, and tests this framework using a 2 × 3 × 2 × 2 between-subjects experimental design. Contrary to predictions from offline bystander theory, higher procedural effort increased rather than reduced reporting, suggesting that effort functions as a commitment filter rather than a uniform deterrent. Disapproval channel availability broadened engagement, and rules reminders moderated the selective effect of effort. Chapter 4 examines personality differences using the Big Five framework. Integrating LIWC-22 linguistic analyses of 2,197 Community Notes contributors with self-reported personality data from the experimental sample, the chapter identifies a systematic engagement–quality gap: Openness predicted who actively intervened, but Conscientiousness predicted whose contributions achieved cross-partisan helpful status. Extraversion and Neuroticism were negatively associated with quality, indicating that the traits motivating intervention differ from those producing effective corrections. Chapter 5 analyzes the linguistic properties of corrections through sentiment analysis, LDA topic modelling, and fractional logit regression on 8,941 tweet-note pairs. Notes adopted a verification-oriented vocabulary and were consistently more negative than the tweets they addressed. Crucially, the degree of shared agreement among note authors that the original tweet was misleading dominated all linguistic predictors of helpfulness, suggesting that effective collective correction depends on shared evaluative consensus rather than stylistic features of the correction itself. Together, these findings advance a design-contingent account of community-based moderation in which platform design, individual differences, and shared evaluative judgment operate as selective conditions determining who recognizes harm, who acts effectively, and when collective correction succeeds.