Score
Designs, implements, and evaluates systems and components that identify false, misleading, or deceptive content and reduce its spread or impact. Work includes building detection models and annotation pipelines, creating evaluation metrics and datasets, and developing mitigation mechanisms such as automated flagging, ranking or removal policies, warnings and counter-messaging interventions, and analyses of propagation dynamics and adversarial robustness.
The proliferation of misinformation in the digital era necessitates robust and interpretable detection methods. This work addresses the problem by modeling information diffusion mechanisms, proposing the first unified classification framework for misinformation detection that integrates both homophilic and heterophilic propagation characteristics. We formalize the problem definition and systematically survey state-of-the-art models, benchmark datasets, and empirical performance limits. Methodologically, we synergize propagation dynamics modeling, graph neural networks, and multimodal natural language processing to uncover cross-platform diffusion patterns. Our key contributions are: (1) the first propagation-mechanism-driven unified classification framework; (2) a standardized benchmarking paradigm; and (3) advocacy for multimodal fusion and novel fact-checking tasks. The resulting reproducible and extensible research map advances misinformation detection toward joint optimization of efficiency, interpretability, and practical utility.
This paper addresses critical challenges in social media misinformation detection—including poor model generalizability and unreliable evaluation—specifically for fake news, spam, and fake accounts. We systematically review and empirically evaluate 36 machine learning and deep learning methods. For the first time, we apply the PROBAST tool to quantify systematic biases across the full lifecycle: data selection, class imbalance mitigation, linguistic preprocessing, and hyperparameter tuning—revealing pervasive bias throughout. We propose an upgraded, real-world-oriented evaluation paradigm: (i) resampling techniques to mitigate dataset bias; (ii) adoption of F1-score and AUROC—rather than accuracy—as primary metrics; and (iii) explicit emphasis on critical preprocessing steps such as negation handling. Experimental results demonstrate that standardized preprocessing and robust evaluation significantly enhance model reliability and cross-platform generalization capability.
This study addresses the urgent need for systematic strategies to identify malicious actors in online social networks by proposing a structured taxonomy that characterizes their behavioral patterns and roles in disinformation campaigns. Integrating domain expert knowledge with insights from academic literature, the framework delineates the mechanisms through which such actors operate. Developed through qualitative analysis, expert collaboration, and case studies, the taxonomy was validated using social media data centered on anti-immigration discourse. The resulting classification system effectively supports researchers and platform operators in detecting and mitigating disinformation, offering both a theoretical foundation and a practical framework for governance interventions.
Automated detection of harmful social media content—such as hate speech, rumors, and extremist text—is vulnerable to adversarial textual perturbations, leading to increased false negatives and poor generalization. To address this, we propose LLM-SGA-ARHOCD: a framework that first leverages large language models to generate and aggregate diverse adversarial samples (LLM-SGA), thereby enhancing attack coverage; it then introduces an Adaptive Robust Hierarchical Online Content Detector (ARHOCD), integrating multi-base model ensembling, Bayesian dynamic weighting, and domain-knowledge-guided collaborative adversarial training. Evaluated on three real-world datasets, our method achieves significant improvements in adversarial robustness (+12.7% on average) and clean-sample accuracy (+3.4% on average), while demonstrating strong cross-attack generalization and high precision. This work establishes a scalable, robust paradigm for secure online content moderation.
This study addresses the growing threat posed by generative large language models (LLMs) in amplifying disinformation within digital ecosystems. To systematically investigate users’ ability to detect AI-generated fake news, the authors develop an end-to-end experimental framework integrating RogueGPT—a novel, controllable disinformation generation engine—and JudgeGPT, an evaluation platform. The framework incorporates multimodal content generation, LLM-assisted detection, cognitive inoculation interventions, and human perception experiments. Findings reveal that while human detection capabilities have improved, a dynamic adversarial interplay persists between generation and detection mechanisms. The proposed strategies demonstrate significant efficacy in mitigating risks associated with AI-generated disinformation, advancing the field beyond theoretical discourse toward empirically grounded countermeasures.
Existing misinformation detection models exhibit poor generalizability against human-crafted disinformation, suffer from non-independent evaluation protocols, and rely on datasets with systemic biases—leading to a severe disconnect between academic research and industrial deployment. Method: We conduct a cross-disciplinary methodology audit of 248 highly cited papers across security, NLP, and computational social science, introducing the first machine learning evaluation framework specifically designed for trust and safety applications. Our audit integrates bibliometric analysis, reproducibility assessment, three representative replication experiments, and systematic evaluation of dataset and code availability. Contribution/Results: We identify critical flaws in data curation practices and evaluation methodologies; empirical results demonstrate substantial performance degradation of fully automated detectors under real-world conditions. This work delivers a practical, actionable evaluation guideline and a concrete research roadmap for trustworthy AI governance.
This study investigates the critical impact of data quality and diversity on the effectiveness and robustness of fake news detection models. Addressing prevalent issues in existing datasets—including inconsistent annotations, significant biases, and missing metadata—we propose a systematic data governance framework. Specifically, we construct the first unified, open-source GitHub repository encompassing 120+ publicly available fake news datasets, enabling standardized integration, bias identification, multi-dimensional metadata annotation, and searchable indexing. Our work provides the first empirical evidence demonstrating that intrinsic data characteristics—rather than model architecture alone—fundamentally determine detection performance, thereby establishing a data-centric research paradigm. The repository has been adopted as a benchmark data entry point by over ten academic institutions and industry teams worldwide, catalyzing a paradigm shift in fake news detection from model-centric to data–model co-evolutionary development.
Current automated detection tools struggle to meet regulatory practice demands due to insufficient transparency, interpretability, and the inability to map findings to specific legal provisions, resulting in a disconnect between academic research and enforcement applications. Through in-depth interviews with nine regulatory practitioners and an analysis integrating regulatory workflows with technical feasibility, this study systematically uncovers, from a regulatory perspective, the practical barriers to deploying automated tools for identifying deceptive designs. The work proposes a human-in-the-loop compliance review framework that is user-need-driven, supports the entire investigative workflow, and aligns both research and regulatory objectives, offering critical guidance for developing automated detection systems that genuinely meet real-world enforcement requirements.
This study addresses a critical gap in disinformation research by shifting focus beyond content veracity to encompass non-content-based manipulative behaviors. The authors introduce the concept of “strategic misrepresentation” and propose a novel four-dimensional analytical framework integrating content, actors, processes, and concealment. This framework uniquely unifies coordinated behavior, procedural manipulation, and obfuscation tactics within a single perspective. By synthesizing machine learning, network science, and visualization techniques, the work develops a cross-modal detection system capable of systematically identifying and evaluating both legitimate and illicit information operations. The approach offers a new taxonomic paradigm and practical toolkit for analyzing cognitive interference in social networks.
This study addresses the lack of systematic validation regarding design choices in multimodal fake news detection by conducting a large-scale empirical investigation. Leveraging pretrained vision-language backbones across multiple benchmark datasets, we perform over 3,375 controlled experiments to systematically evaluate model design decisions and their robustness. Our analysis identifies the core factors influencing model behavior, distills key design guidelines, and reveals typical failure modes. By rigorously examining these architectural and training considerations, this work establishes a solid foundation for developing reliable multimodal fake news detection systems.
研究通过SuriCap平台和CTF式工作坊,分析了60名参与者创建网络入侵检测规则的过程与方法,揭示经验对规则质量影响有限,并指出标记数据的重要性。
Accurately measuring the proportion of policy-violating content actually encountered by users is challenged by the rarity of violations, high annotation costs, and the difficulty of conducting frequent, representative assessments. This work proposes a design-based measurement system that draws daily probability samples from user exposure streams using machine learning–assisted weighting. It enables efficient annotation through multimodal large language models, policy-guided prompting, and gold-set validation, and constructs unbiased estimators to produce prevalence metrics with confidence intervals. The system supports multidimensional post-stratification—such as by platform interface, user geography, or content age—using a single global sample, maintaining statistical unbiasedness while prioritizing high-exposure and high-risk content. This approach substantially improves monitoring efficiency, timeliness, and flexibility while significantly reducing annotation costs.