M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text

📅 2025-11-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the information integrity crisis caused by pervasive LLM-generated text in news and academic writing, this work introduces the first cross-domain benchmark for AI-generated text detection. It comprises a large-scale, multi-source dataset of 30,000 samples—generated by diverse models (e.g., GPT, LLaMA) and prompting strategies—and covers high-stakes domains: news and academic writing. Methodologically, we propose a unified binary classification framework that jointly leverages pre-trained language model representations and fine-grained textual features, substantially improving cross-domain generalization. The benchmark has attracted 46 registered teams; the top four achieved high detection accuracy on both subtasks (news and academic). Key contributions include: (1) the first large-scale, multi-source, consistently annotated detection benchmark grounded in real-world applications; (2) empirical validation of cross-domain detection feasibility and an effective technical pathway; and (3) a reproducible evaluation standard and strong baseline for AI content governance.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Application Domains: Misinformation & Fake News

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
The generation of highly fluent text by Large Language Models (LLMs) poses a significant challenge to information integrity and academic research. In this paper, we introduce the Multi-Domain Detection of AI-Generated Text (M-DAIGT) shared task, which focuses on detecting AI-generated text across multiple domains, particularly in news articles and academic writing. M-DAIGT comprises two binary classification subtasks: News Article Detection (NAD) (Subtask 1) and Academic Writing Detection (AWD) (Subtask 2). To support this task, we developed and released a new large-scale benchmark dataset of 30,000 samples, balanced between human-written and AI-generated texts. The AI-generated content was produced using a variety of modern LLMs (e.g., GPT-4, Claude) and diverse prompting strategies. A total of 46 unique teams registered for the shared task, of which four teams submitted final results. All four teams participated in both Subtask 1 and Subtask 2. We describe the methods employed by these participating teams and briefly discuss future directions for M-DAIGT.
Problem

Research questions and friction points this paper is trying to address.

Detecting AI-generated text across multiple domains
Developing binary classification for news and academic writing
Creating a large-scale benchmark dataset for detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-domain detection of AI-generated text
Binary classification for news and academic writing
Large-scale dataset with diverse LLM-generated content
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
Salima Lamsiyah
Salima Lamsiyah
NLP-Machine Learning Researcher, Luxembourg University
NLPMachine LearningDeep LearningTransfer LearningLLM
S
Saad Ezzini
King Fahd University of Petroleum and Minerals, Saudi Arabia
A
Abdelkader El Mahdaouy
Mohammed VI Polytechnic University, Morocco
H
H. Alami
Sidi Mohamed Ben Abdellah University, Morocco
A
Abdessamad Benlahbib
Sidi Mohamed Ben Abdellah University, Morocco
S
Samir El Amrany
University of Luxembourg, Luxembourg
S
Salmane Chafik
Mohammed VI Polytechnic University, Morocco
H
Hicham Hammouchi
University of Luxembourg, Luxembourg