Robust Detection of LLM-Generated Text under Contamination

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of LLM-generated text detection under adversarial editing and data contamination. By leveraging finite-order Markov processes and the Huber contamination model, it rigorously establishes theoretical bounds for reliable detection. Furthermore, this work proposes a trimmed likelihood ratio test, demonstrating through additive score analysis that the trimming operation effectively prevents the worst-case collapse of detection power to zero. The primary contribution lies in offering a simple yet effective strategy for enhancing the robustness of existing detectors. Extensive experiments across multiple datasets and the RAID benchmark confirm that the proposed method substantially improves detection robustness; notably, it increases the median true positive rate of the LRR detector by 8.3 percentage points.
📝 Abstract
We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.
Problem

Research questions and friction points this paper is trying to address.

LLM-generated text detection
robust detection
Huber contamination
text editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-generated text detection
Huber contamination
clipped likelihood-ratio test
Markov process
robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jiaxun Li
Jiaxun Li
Ph.D. of Statistics, University of Michigan
StatisticsLearning Theory
S
Saptarshi Chakraborty
Department of Statistics, University of Michigan
A
Ambuj Tewari
Department of Statistics, University of Michigan