🤖 AI Summary
This study addresses the limitation of existing AI text detectors in providing fine-grained analysis of mixed human-AI editing, specifically their inability to localize modified tokens or quantify editing intensity. We propose a token-level detection framework that, for the first time, decouples editing scope from intensity, overcoming the constraints of conventional methods that output only holistic labels or scores. Methodologically, through token-level supervised training, source-edited sequence alignment, and conditional probability modeling, our approach independently predicts the editing status and intensity of each token, requiring only the original input text during inference. Experimental results demonstrate that this framework precisely localizes edited tokens and reveals distinct intensity patterns across different editing operations, while maintaining robust classification performance under cross-domain and cross-generator variations.
📝 Abstract
Large language models are increasingly used to edit human-written text rather than generate entire texts from scratch. Conventional AI-text detectors mainly distinguish human-written from fully AI-generated text, while recent methods for AI-edited text typically provide only a text-level label or editing-degree score. We introduce MixDetect, a word-level framework for localizing and quantifying AI editing. MixDetect separately predicts whether each word has been edited and, conditional on editing, how substantial the edit is, allowing editing scope and editing intensity to be estimated separately. During training, source--edited pairs are aligned to construct word-level supervision, while inference requires only the input text. Experiments show that MixDetect accurately localizes AI-edited words, reflects differences in editing intensity, and reveals different scope--intensity patterns across editing degrees and operations. The overall AI editing magnitude increases under additional AI editing, decreases when AI-generated text is edited by humans, and remains nearly unchanged under ordinary human-to-human editing. The aggregated text-level predictions also perform well on binary and ternary AI-text classification and remain effective under domain and generator shifts. These results show that AI editing can be analyzed beyond a single authorship label or editing-degree score by identifying both where AI editing occurs and how substantial the edits are.