LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text

📅 2025-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing AI-generated text detection datasets suffer from three critical limitations: outdated model coverage, monolingual bias (predominantly English), and insufficient fine-grained annotation—particularly hindering AI fragment localization in human-AI collaborative writing. To address these gaps, we introduce the first large-scale bilingual (English-Russian) dataset of AI-generated text, featuring character-level annotations produced via human–machine collaboration (combining manual verification with automated pre-labeling). It encompasses outputs from diverse contemporary closed- and open-source large language models. The dataset supports both document-level binary classification and precise interval-level localization of AI-generated segments, thereby establishing the first benchmark for multilingual, fine-grained AI-text detection. Empirical evaluation demonstrates substantial improvements in localization accuracy and cross-model generalization under mixed-authorship scenarios, providing a foundational resource for advancing robust, multilingual AI-content identification research.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Humans and AI: AI for Accessibility

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
The widespread use of human-like text from Large Language Models (LLMs) necessitates the development of robust detection systems. However, progress is limited by a critical lack of suitable training data; existing datasets are often generated with outdated models, are predominantly in English, and fail to address the increasingly common scenario of mixed human-AI authorship. Crucially, while some datasets address mixed authorship, none provide the character-level annotations required for the precise localization of AI-generated segments within a text. To address these gaps, we introduce LLMTrace, a new large-scale, bilingual (English and Russian) corpus for AI-generated text detection. Constructed using a diverse range of modern proprietary and open-source LLMs, our dataset is designed to support two key tasks: traditional full-text binary classification (human vs. AI) and the novel task of AI-generated interval detection, facilitated by character-level annotations. We believe LLMTrace will serve as a vital resource for training and evaluating the next generation of more nuanced and practical AI detection models. The project page is available at href{https://sweetdream779.github.io/LLMTrace-info/}{iitolstykh/LLMTrace}.
Problem

Research questions and friction points this paper is trying to address.

Detecting AI-generated text with limited suitable training data
Addressing mixed human-AI authorship scenarios in text analysis
Providing character-level annotations for precise AI segment localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bilingual corpus for AI text detection
Character-level annotations for interval localization
Diverse modern LLMs for dataset construction
🔎 Similar Papers
2024-06-21Journal of Artificial Intelligence ResearchCitations: 6
💼 Related Jobs
No related jobs found.
SALUTEDEV LLC