LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how large language model (LLM) detection tools, when deployed as interventions, can exert counterintuitive effects on downstream metrics—such as usage frequency and output quality—due to strategic user behavior. Drawing on game-theoretic insights, the authors develop a structural model that captures how users adapt their LLM usage and post-processing strategies in response to detection mechanisms. Empirical validation is conducted using word-frequency data from arXiv abstracts. The work reveals, for the first time, a non-monotonic relationship between detection signals and downstream outcomes: while detectors effectively reduce the targeted attribute, they may paradoxically increase LLM usage and degrade content quality, exhibiting an initial rise followed by a decline. This challenges conventional assumptions about the efficacy of detection-based interventions.
📝 Abstract
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.
Problem

Research questions and friction points this paper is trying to address.

LLM detection
strategic user behavior
downstream impact
output quality
intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM detection
strategic behavior
downstream impact
intervention
output quality
🔎 Similar Papers
No similar papers found.