Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of AI writing detectors to iterative updates of large language models (LLMs), which threatens the reliability of academic integrity screening. Based on 4,000 PNAS abstracts rewritten across 23 LLM versions, this work quantitatively evaluates the performance boundaries of detection under generational model shifts. It reveals a sharp degradation in detector performance at generational boundaries and demonstrates that lexical distributional divergence is the critical determinant of detection transferability, thereby challenging the validity of static benchmark evaluations. Experiments indicate that non-updated detectors exhibit miss rates as high as 96%, while strategies covering all versions incur substantial false-positive risks. Accordingly, this paper proposes a maintenance paradigm requiring detector revalidation upon every LLM release.
📝 Abstract
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
AI-assisted writing detection
Model turnover
Scientific integrity screening
Detector reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM turnover
AI-generated text detection
scientific writing screening
model versioning
research integrity
🔎 Similar Papers
2024-01-30Journal of Data and Information ScienceCitations: 5