🤖 AI Summary
This study addresses the limited cross-domain generalization of detectors for LLM-generated text and the unclear mechanisms underlying self-detection versus cross-generation detection. It systematically evaluates the zero-shot classification capabilities of fifteen LLMs serving as both generators and detectors, leveraging a large-scale benchmark dataset comprising human- and AI-authored texts for comparative experiments. The findings reveal, for the first time, that detection efficacy depends primarily on detector capability rather than generation source. Furthermore, no systematic self-detection advantage is observed, while state-of-the-art models achieve superior detection trade-offs. Notably, models across different generations exhibit significant biases toward either false positives or false negatives, and they rely on inconsistent textual cues to justify their decisions.
📝 Abstract
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.