🤖 AI Summary
To address the information integrity crisis caused by pervasive LLM-generated text in news and academic writing, this work introduces the first cross-domain benchmark for AI-generated text detection. It comprises a large-scale, multi-source dataset of 30,000 samples—generated by diverse models (e.g., GPT, LLaMA) and prompting strategies—and covers high-stakes domains: news and academic writing. Methodologically, we propose a unified binary classification framework that jointly leverages pre-trained language model representations and fine-grained textual features, substantially improving cross-domain generalization. The benchmark has attracted 46 registered teams; the top four achieved high detection accuracy on both subtasks (news and academic). Key contributions include: (1) the first large-scale, multi-source, consistently annotated detection benchmark grounded in real-world applications; (2) empirical validation of cross-domain detection feasibility and an effective technical pathway; and (3) a reproducible evaluation standard and strong baseline for AI content governance.
📝 Abstract
The generation of highly fluent text by Large Language Models (LLMs) poses a significant challenge to information integrity and academic research. In this paper, we introduce the Multi-Domain Detection of AI-Generated Text (M-DAIGT) shared task, which focuses on detecting AI-generated text across multiple domains, particularly in news articles and academic writing. M-DAIGT comprises two binary classification subtasks: News Article Detection (NAD) (Subtask 1) and Academic Writing Detection (AWD) (Subtask 2). To support this task, we developed and released a new large-scale benchmark dataset of 30,000 samples, balanced between human-written and AI-generated texts. The AI-generated content was produced using a variety of modern LLMs (e.g., GPT-4, Claude) and diverse prompting strategies. A total of 46 unique teams registered for the shared task, of which four teams submitted final results. All four teams participated in both Subtask 1 and Subtask 2. We describe the methods employed by these participating teams and briefly discuss future directions for M-DAIGT.