🤖 AI Summary
This study addresses a critical methodological flaw in existing research that identifies scholars’ adoption of large language models (LLMs) based on the first preprint flagged as LLM-generated. We demonstrate through theoretical analysis and Monte Carlo simulations that this approach mechanically labels months of high research output as “adoption events,” thereby generating spurious positive treatment effects in event studies. Using multiple placebo tests—including random assignment of adoption dates, neutral keyword labeling, and analyses restricted to periods before the release of ChatGPT—we show that this identification strategy significantly overestimates the impact of LLMs on scholarly productivity even in the absence of any true causal effect. Our findings challenge the validity of prior conclusions and expose a fundamental bias in current methods for detecting LLM adoption.
📝 Abstract
Kusumegi et al. (2025) study whether researchers' preprint output rises after adopting large language models (LLMs), dating adoption as the first month in which at least one submitted abstract exceeds an LLM-detection threshold. We show that this treatment-timing rule is mechanically related to output. The probability that at least one paper is flagged in a month is increasing in the number of papers submitted in that month, so detected-adoption months are disproportionately high-output months. An event study centered on first detection can therefore display positive post-event dynamics even when the flagging rule contains no information about true LLM adoption, because the omitted pre-treatment period is selected from months with no prior detection. We demonstrate this in a simulation: with i.i.d. productivity and no causal effect, first-detection timing generates a spurious positive post-treatment path. We also replicate the stacked event study of Kusumegi et al. (2025) and show that three placebo exercises (random paper-level assignment, neutral keyword flags, and a pre-ChatGPT observation window) each produce a similarly positive post-treatment pattern.