🤖 AI Summary
This study investigates the proliferation of contrastive constructions, such as "rather than," in NLP papers generated by large language models (LLMs), which introduces content redundancy and elicits reviewer dissatisfaction. Through text mining, human annotation, preference data analysis, and reward model evaluation, this work systematically examines the causes underlying this surge in LLM-assisted academic writing. Our findings reveal that post-training mechanisms based on human preference data inadvertently induce the overuse of contrastive structures as a side effect, and we quantify their detrimental impact on reader experience. Notably, the frequency of such constructions in 2026 is seven times that of 2019, with the majority judged as low-quality expressions that degrade academic rigor.
📝 Abstract
For better or worse, LLMs are by now used routinely for scientific writing.\footnote{This paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments).} Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emph{rather than impressing them}. We study the construction \emph{rather than} in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emph{annoying} and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emph{Annoying} uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.