🤖 AI Summary
This study addresses how to enhance the capability of large language models (LLMs) to generate high-value review questions that facilitate manuscript revision. Leveraging paired ICLR/NeurIPS datasets and automated text-diff detection, the authors employ GPT-series models to generate edit-inducing questions for paper drafts, evaluating them against human reviewer comments. Notably, the research reveals a counterintuitive phenomenon: extended context processing diminishes the utility of outputs from reasoning models. The findings demonstrate that although automatically generated questions exhibit a relatively low hit rate, they achieve broader coverage and elicit more substantive revisions. Overall, this work validates the effectiveness and application potential of LLM-driven automation in assisting academic writing.
📝 Abstract
We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.