Score
The practice of identifying and characterizing distinct user groups from behavioral and demographic data to measure how platform usage, recommendations, or control failures differ across populations and how those changes persist over time.
Current research on online behavioral change suffers from narrow behavioral coverage, overreliance on API-restricted platforms as data sources, and a persistent theory–empiricism gap. To address these limitations, this study conducts a systematic literature review of 148 peer-reviewed articles published between 2000 and 2023, constructing a four-dimensional knowledge graph encompassing behavioral categories, detection methodologies, platform ecosystems, and theoretical foundations. Our analysis uncovers three salient trends: (1) affective orientation dominates behavioral modeling; (2) platform distribution is heavily skewed toward a few API-constrained platforms; and (3) theoretical integration remains markedly underdeveloped. We propose a novel methodology framework—“Multi-behavioral Modeling, Heterogeneous Data Integration, and Theory–Practice Alignment”—and deliver a structured research map that precisely identifies critical gaps. This work advances the computational behavioral paradigm and offers an actionable methodological guide for digital social governance.
To address the growing prevalence of cyberbullying, harassment, and other hostile behaviors on social media—and their associated emotional distress and mental health crises—this study proposes a socio-computational integrative paradigm. Methodologically, it introduces the first unified definition of “online aggressiveness” and establishes an interdisciplinary framework encompassing multi-source data collection, joint language–context modeling, machine learning/deep learning detection algorithms, social network analysis, and behavioral trajectory mining. Key contributions include: (1) a synergistic mechanism integrating content-based detection with behavior-oriented analysis; (2) empirical evidence demonstrating that sociological factors—including group dynamics and cultural context—significantly enhance model robustness and intervention interpretability; and (3) a systematic mapping of the field’s intellectual landscape and core challenges, thereby laying a theoretical foundation and technical roadmap for building trustworthy, interpretable, and actionable intelligent governance systems.
Digital platforms generate vast volumes of high-resolution interaction data, offering novel opportunities to study information diffusion, opinion dynamics, and collective coordination—yet the field suffers from fragmentation, methodological heterogeneity, insufficient validation, and weak cross-domain integration. This paper addresses these challenges via a systematic review that synthesizes empirical findings and formal models to construct a cross-platform comparable empirical benchmark framework; identifies structural limitations of prevailing modeling paradigms; and critically evaluates underlying methodological assumptions. It further advocates for standardized model validation protocols and reproducible analytical practices. The core contributions are: (1) establishing a shared empirical baseline for online social systems; (2) explicitly characterizing key structural constraints that impede causal inference; and (3) proposing an analytically tractable, theoretically rigorous framework. Together, these advances lay a methodological foundation for robust, comparable, and scalable future research.
This work investigates the long-term temporal evolution of data and behavioral patterns in recommender systems, focusing on dynamics in user engagement, recommendation diversity, and fairness—particularly assessing the causal impact of Goodreads’ 2011 algorithmic recommendation rollout. Leveraging the complete 2006–2023 book interaction dataset, we construct the first temporal-evolution-aware multi-dimensional evaluation framework, integrating rolling-window collaborative filtering (SVD++, BPR, LightGCN) with time-sensitive metrics—including Gini coefficient, entropy-based diversity, and exposure fairness. Our analysis uncovers a novel “evolutionary bias”: persistent decline in diversity, increasing exposure inequality favoring head items, and marked deterioration in fairness for new users. We empirically confirm that algorithmic recommendations amplify the Matthew effect, challenging prevailing static temporal-split evaluation paradigms and underscoring the necessity of evolutionary assessment in real-world recommendation deployment.
This study addresses the limited understanding of users’ motivations for migrating across social media platforms and their holistic usage trajectories, which hinders explanations of platform preferences and generational differences. To bridge this gap, the work proposes a conceptual framework termed the “social media journey” and employs a mixed-methods approach with quota sampling to collect data from 1,000 U.S. users. For the first time, it models multi-platform engagement as an integrated experience, moving beyond the constraints of single-platform studies. The research identifies key push and pull factors driving platform migration and uncovers generational variations in need-based adoption patterns. These findings offer an integrative analytical paradigm for human-computer interaction (HCI) research and provide novel, generation-informed insights for platform design and governance.
This study identifies the “echo-platform” phenomenon in social media ecosystems—characterized by escalating platform-level ideological homophily and deepening cross-platform user segmentation, leading to structural fragmentation of the digital public sphere. We propose the first operationalizable three-dimensional framework—comprising platform centrality, news credibility, and user diversity—and apply it to 126 million URLs and cross-platform behavioral data from nearly six million users across nine platforms. Methodologically, we integrate network centrality analysis, automated credibility classification, and heterogeneity measurement. Empirical results reveal a systematic bifurcation: “mainstream platforms” exhibit higher centrality and credibility with ideologically diverse users, whereas “alternative tech platforms” are peripheral, dominated by low-credibility content, and host highly homogeneous user populations. Our work establishes a reproducible quantitative benchmark and a novel theoretical paradigm for studying polarization in digital public spheres.
This study investigates users’ personalized adoption behaviors toward music discovery features and recommendation content in mobile music applications. Method: Leveraging two years of longitudinal behavioral logs from Deezer, we analyze user preferences and decision-making mechanisms regarding three types of affordances—organic, algorithmic, and editorial—using behavioral log mining, temporal pattern clustering, and cross-modal association analysis. Contribution/Results: We empirically demonstrate high heterogeneity in affordance adoption, refuting the implicit assumption of a uniform recommendation adoption pattern. Building upon this, we propose a multidimensional dynamic behavioral typology framework that identifies distinct user segments and their evolutionary trajectories. Results show that recommendation adoption is jointly governed by affordance type and individual usage rhythm. This work provides both theoretical grounding and empirical evidence for designing next-generation personalized recommendation systems.
This study addresses the limitations of existing political stance analysis, which predominantly relies on a unidimensional left–right spectrum rooted in the U.S. context and fails to capture the nuanced positions of users, politicians, and media across multiple policy dimensions in diverse democracies. To overcome this constraint, the work introduces the first multidimensional political stance dataset applicable across multiple countries, encompassing key dimensions such as immigration, European Union attitudes, liberal values, views on elites and institutions, nationalism, and environmental concerns. Leveraging behavioral data from the X platform and integrating content analysis with stance inference techniques, the authors develop a multidimensional positioning framework that incorporates activity-based metrics. Empirical validation demonstrates that the dataset and its associated benchmarks effectively support research on polarization and information diversity, substantially expanding the scope of computational political science beyond U.S.-centric paradigms.
Accurately measuring the proportion of policy-violating content actually encountered by users is challenged by the rarity of violations, high annotation costs, and the difficulty of conducting frequent, representative assessments. This work proposes a design-based measurement system that draws daily probability samples from user exposure streams using machine learning–assisted weighting. It enables efficient annotation through multimodal large language models, policy-guided prompting, and gold-set validation, and constructs unbiased estimators to produce prevalence metrics with confidence intervals. The system supports multidimensional post-stratification—such as by platform interface, user geography, or content age—using a single global sample, maintaining statistical unbiasedness while prioritizing high-exposure and high-risk content. This approach substantially improves monitoring efficiency, timeliness, and flexibility while significantly reducing annotation costs.
This study addresses a critical gap in the empirical understanding of users’ long-term information environments on Facebook, which has hindered accurate assessment of the platform’s political influence. Leveraging longitudinal data from over 1,100 U.S. users between 2012 and 2023, the research employs a platform-agnostic approach to analyze hundreds of millions of public page and group posts they were exposed to. Using large-scale content classification and ideological orientation detection, the study quantifies exposure to political content—distinct from user engagement alone. Findings reveal that political content constitutes 18% of users’ overall information diet, with significant and persistent disparities across age, gender, and racial groups. Notably, Meta’s 2018 algorithmic update substantially increased the share of political content, indicating deep penetration of political discourse into ostensibly non-political spaces.
This study addresses the limitations of traditional discrete categorizations in modeling social media user behavior, which hinder unified interpretation and analysis. The authors propose the first continuous two-dimensional reciprocity space that encompasses all forms of social interaction, quantifying user engagement patterns through bidirectional connection ratios and naturally mapping diverse behaviors onto continuous regions within this space. Through large-scale empirical analysis of 48,830 Twitter users and 149 million connections, the research demonstrates that user attributes vary smoothly along the reciprocity dimensions, with conventional discrete behavioral types spontaneously clustering into interpretable regions. Beyond revealing a behavioral gradient, the model provides a quantifiable framework for assessing influence, offering a novel paradigm for platform design and user analytics.
This study addresses the challenges of cross-platform social media data analysis—namely data heterogeneity, API restrictions, and privacy compliance—stemming from the absence of standardized, reproducible workflows. To overcome these limitations, the authors propose an open-source Python framework featuring a unified data model that harmonizes multi-source social data across five core dimensions: communities, accounts, posts, behaviors, and entities. The framework incorporates a configurable personally identifiable information (PII) anonymization module to ensure regulatory compliance and integrates an LLM-driven analytical layer that enables semantic enrichment without requiring code modifications. Evaluated through four case studies, the framework demonstrates significant improvements in fairness, reproducibility, and scalability for cross-platform textual and network analyses.