Score
Specifying, estimating, and interpreting statistical interaction (moderation) effects to determine how relationships between variables change across levels of a moderator (e.g., user expertise, team composition) and over time, including tests for changes after specific dates. Involves model specification, interaction terms, and inference procedures for assessing moderation and temporal shifts.
This study addresses a critical gap in content moderation research by moving beyond the prevailing view of moderation as a monolithic intervention and instead examining the joint effects of moderator type, violation context, and linguistic style on user compliance and self-censorship. Grounded in the HAII-TIME framework, the analysis leverages over ten million moderation events from Reddit, integrating probabilistic behavioral classification, ANOVA, OLS regression, and PCA-based linguistic feature analysis. The findings reveal, for the first time, that bot moderators are more effective than human moderators at enhancing compliance while reducing self-censorship. The study further demonstrates that violation severity moderates the efficacy of linguistic strategies and incorporates violation salience into the HAII-TIME model. Among 480 language-context interactions, 33 effects remain significant after FDR correction, offering empirical foundations for context-adaptive moderation systems.
This study addresses the challenge of accurately predicting user behavioral responses to content moderation interventions to inform user-centered policy design. Leveraging pre- and post-intervention behavioral data from 16,800 Reddit users, we propose a novel feature importance assessment framework that integrates quantification learning with greedy feature selection, modeling 753-dimensional features spanning social behavior, linguistic patterns, relational networks, and psychological attributes. Results reveal that response heterogeneity is jointly driven by a small set of cross-task generalizable features—such as historical participation stability—and numerous task-specific features. The model achieves strong performance in predicting changes in activity levels and toxicity, while diversity prediction remains comparatively challenging. These findings empirically characterize the multidimensional and heterogeneous nature of user responses, providing both methodological grounding and empirical support for interpretable, customizable content moderation strategies.
This work addresses the inherent limitations of conventional content moderation—namely, post-hoc evaluation and reactive responses—by formally proposing and modeling a novel task: *intervention-induced user attrition prediction*, which forecasts users’ likelihood of abandoning the platform *prior* to moderator interventions. Leveraging 142 subreddit bans and 13.8 million user behavioral logs from Reddit, we construct a 142-dimensional feature set capturing activity patterns, social connectivity, content toxicity, and linguistic style. Our method integrates XGBoost with BERT-based text encoding for binary classification. Experiments reveal that activity-related features exhibit the highest discriminative power; the best-performing model achieves a micro-F1 score of 0.914 and demonstrates robust cross-community generalization. This study pioneers the “predictive content moderation” paradigm, delivering a deployable tool for pre-intervention impact assessment—thereby substantially mitigating unintended user attrition and collateral enforcement errors.
This study evaluates the causal effects of large-scale deplatforming interventions, exemplified by Reddit’s “Great Ban.” Leveraging longitudinal log data from 34,000 users and 53 million comments, it applies a difference-in-differences (DID) framework—the first such use for quantifying heterogeneous treatment effects in deplatforming. Results reveal: (1) 15.6% of banned users permanently disengage; (2) among retained users, aggregate comment toxicity declines by 4.1%, yet a distinct subgroup exhibits a 70% *increase* in toxicity; (3) this high-toxicity subgroup shows no significant rise in activity or engagement, challenging the common hypothesis that deplatforming intensifies extremist behavior. The paper’s contributions are threefold: it pioneers DID-based estimation of differential deplatforming effects; uncovers non-monotonic toxicity responses; and demonstrates substantial individual-level heterogeneity—refuting uniform-effect assumptions. These findings provide granular, causal evidence to inform platform content governance policies.
This study addresses the inconsistency in causal effect estimates between observational studies and randomized controlled trials (RCTs) by proposing the first unified framework for decomposing causal effect heterogeneity. The framework systematically identifies and quantifies three sources of heterogeneity: differences in covariate distributions, variation in mediating pathways, and shifts in outcome-generating mechanisms. Methodologically, it formally defines effect decomposition across data types (observational vs. experimental), integrating causal inference, sensitivity analysis, and decomposition modeling, while enabling robust parameter estimation under multiple hypotheses. Evaluated through simulation studies and an empirical analysis of the “Moving to Opportunity” experiment, the framework demonstrates improved interpretability, robustness, and policy generalizability in synthesizing evidence from heterogeneous data sources.
Toxic content propagation on online social platforms demands governance mechanisms that balance theoretical rigor with practical feasibility. This paper proposes a toxicity propagation simulation framework based on an extended SEIZ (Susceptible–Exposed–Infected–Zombie) epidemic model. It introduces, for the first time in content moderation simulation, user-level modeling of the Dark Triad personality traits—narcissism, Machiavellianism, and psychopathy—as key determinants of susceptibility and transmission behavior. We design a threshold-driven, configurable, and interpretable personalized moderator that dynamically adjusts intervention intensity according to individual psychological profiles, thereby departing from conventional uniform-intervention paradigms. Experimental results demonstrate that the proposed intelligent moderator significantly suppresses toxicity diffusion: average propagation rate and duration decrease by 47% relative to a baseline moderator. These findings validate the critical advantages of personality-aware moderation strategies in enhancing both intervention efficacy and interpretability.
Existing multidimensional factor models struggle to accommodate the typical 3–5 dimensional latent constructs in psychometrics and lack a unified framework for modeling multiparameter moderation effects. This paper proposes a scalable penalized maximum likelihood estimation method applicable to arbitrarily many factors, enabling— for the first time—the joint estimation of linear and nonlinear moderation effects within high-dimensional models. By incorporating ridge, lasso, and alignment penalties, the approach simultaneously stabilizes parameter estimation, detects partial measurement noninvariance, and enhances interpretability. Leveraging closed-form analytical gradients, the method avoids computationally intensive numerical integration and MCMC sampling, substantially improving computational efficiency. Simulation and empirical studies demonstrate accurate recovery of complex moderation patterns. The proposed method provides a scalable, efficient, and robust new tool for measurement invariance research involving multidimensional constructs.
This study addresses the challenge of inconsistent content moderation in online communities, where volunteer moderators frequently disagree on borderline cases characterized by ambiguous user intent—so-called “gray-area” instances. Leveraging a large-scale dataset of 4.3 million moderation logs spanning five years across 24 Reddit subreddits, this work provides the first quantitative definition and characterization of gray-area cases. Using information-theoretic measures to assess decision difficulty, the analysis reveals that approximately one-seventh of all moderation decisions are contentious, with nearly half involving automated tools. Gray-area cases substantially increase adjudication complexity, underscoring the irreplaceable role of human expert oversight and exposing significant limitations of current language models in handling such nuanced moderation tasks.
This work addresses the lack of standardized, comparable datasets in content moderation research, which has hindered systematic evaluation of the effectiveness and biases of different intervention strategies. To bridge this gap, we introduce TBBT, a large-scale dataset encompassing 25 distinct moderation interventions, 339,000 users, and nearly 39 million messages. The dataset includes three months of standardized metadata and anonymized user behavioral records both before and after each intervention. TBBT enables, for the first time, reproducible, multidimensional, and cross-intervention analyses, substantially improving the consistency and efficiency of moderation impact assessment. It supports a wide range of research scenarios and lays the groundwork for more systematic investigation in the field of content moderation.
This study investigates power-related biases among human moderators in online asymmetric power conflicts and examines how these biases are influenced by AI recommendations. Employing a mixed experimental design grounded in real-world consumer–merchant dispute scenarios, the research constructs a human–AI collaborative moderation framework and systematically identifies multiple forms of bias favoring the more powerful party. The findings reveal that AI assistance generally mitigates most of these biases; however, under specific conditions, it paradoxically exacerbates certain types. This work is the first to delineate the manifestations of moderation bias under power asymmetry and to uncover the moderating role of AI, offering empirical evidence and novel insights for optimizing platform moderation systems and designing fairer human–AI collaborative decision-making mechanisms.