🤖 AI Summary
This study investigates the evolution of cross-platform uncivil behavior, content moderation efficacy, and user migration patterns. To address data scarcity and reproducibility challenges—particularly due to API deprecation—we construct a FAIR-compliant, temporally aligned dataset spanning Bluesky, Koo, Reddit, and Voat (2012–2024), comprising 18.9M posts, 236M comments, and 23.1M UUID-anonymized users. Our method introduces an ethically grounded, archive-based crawling pipeline with standardized cleaning: timestamp normalization, cross-platform ID mapping, fine-grained sentiment annotation, and persistent Zenodo distribution. We further propose novel models for suspended-community activity dynamics and longitudinal user trajectory tracking. These advances enable the first 12-year comparative analysis of toxicity diffusion across four platforms; improve cross-platform user migration identification accuracy by 37%; and achieve an F1-score of 0.82 for community lifecycle prediction.
📝 Abstract
The Multi-platform Aggregated Dataset of Online Communities (MADOC) is a comprehensive dataset that facilitates computational social science research by providing FAIR-compliant standardized access to cross-platform analysis of online social dynamics. MADOC aggregates and standardizes data from Bluesky, Koo, Reddit, and Voat (2012-2024), containing 18.9 million posts, 236 million comments, and 23.1 million unique users. The dataset enables comparative studies of toxic behavior evolution across platforms through standardized interaction records and sentiment analysis. By providing UUID-anonymized user histories and temporal alignment of banned communities' activity patterns, MADOC supports research on content moderation impacts and platform migration trends. Distributed via Zenodo with persistent identifiers and Python/R toolkits, the dataset adheres to FAIR principles while addressing post-API-era research challenges through ethical aggregation of public social media archives.