Score
Designs, builds, or evaluates systems and methods that generate concrete, constructive revision suggestions for user-authored text, producing prioritized, actionable guidance and brief inline comments instead of only scores or full rewrites; this includes mechanisms to decide which changes to propose and how to format suggestions for insertion alongside the original content.
Existing research ideation tools emphasize breadth-oriented idea generation but lack support for iterative refinement, elaboration, and evaluation—hindering literature-grounded, deep-reading–driven conceptual evolution. Method: We propose the first literature-driven interactive research ideation system, integrating a composable “idea element” canvas model with a multi-dimensional (problem/solution/evaluation/contribution) co-evolution mechanism. Our approach innovatively incorporates LLM-powered literature-aware feedback generation, graph-structured idea modeling, and interactive multi-path variant exploration. Contribution/Results: Experiments demonstrate a 42% increase in user-generated idea output and significantly enhanced detail elaboration. Seven researchers successfully applied the system across the full ideation pipeline—from initial topic conception to paper outline revision—validating its efficacy in supporting deep, iterative, literature-informed research design.
This work addresses the challenge researchers often face in balancing novelty with effective grounding in existing literature when developing new ideas, as well as the lack of tools that support dynamic interaction between emerging concepts and relevant scholarly works. The paper introduces a novel “literature-driven idea pivoting” mechanism—a closed-loop framework that integrates idea drafting, dynamic literature retrieval, semantic clustering, and generative critical feedback to enable co-evolution of research ideas and the literature space. The system performs context-aware analysis of partial idea content and provides real-time improvement suggestions based on clusters of relevant papers. Experimental results demonstrate that this approach significantly enhances the quality of user-generated ideas and strengthens researchers’ ability to comprehend and leverage the scholarly context effectively.
This study addresses the challenge of ensuring review quality under constrained peer-review resources by automating the assessment of review utility for authors. We first propose a systematic, four-dimensional definition of utility—actionability, evidentiary support & specificity, verifiability, and helpfulness—and construct RevUtil, a large-scale benchmark comprising human-annotated and controllably synthesized review–response pairs. We design an evaluation framework supporting multi-dimensional scoring and rationale generation. Leveraging fine-grained annotations, we adapt open-source language models, achieving human-level or better inter-annotator agreement (vs. GPT-4o) across all dimensions. Empirical analysis reveals that current AI-generated reviews remain substantially less useful than human-written ones. Nevertheless, this work establishes the first reproducible benchmark and methodological foundation for automated utility evaluation of peer reviews.
Current peer review systems—constrained by unimodal (text-only) input, limited contextual grounding, and non-actionable feedback—fail to harness the full potential of large language models (LLMs). This paper introduces the first multimodal (text + figure) peer review simulation system augmented with community knowledge. Leveraging OpenReview data, we design a retrieval-augmented generation (RAG) framework that jointly models visual content and academic semantics to generate high-quality, context-aware review comments. We propose a novel structured format—Action:Objective[#]—to transform feedback into executable, traceable revision tasks. The system is deployed via a web-based interactive interface integrated with academic writing platforms, enabling real-time pre-submission feedback and revision tracking. Experiments demonstrate significant improvements over ablated baselines across comprehensiveness, practicality, and expert agreement, thereby enhancing both review quality and collaborative efficiency.
Peer review often struggles to effectively enhance manuscript quality and impact due to inconsistent feedback. To address this challenge, this work proposes the first automated paper revision framework that integrates citation prediction with content-preserving constraints, leveraging a large language model (LLM) agent system to refine textual expression without altering core scientific content, thereby boosting potential academic impact. The approach employs automated evaluation criteria discovery, controllable text rewriting, and a consistency-preserving mechanism to significantly improve manuscript quality: revised papers exhibit a 19.6% reduction in future citation prediction error and are rated superior to the original by expert reviewers in 79% of cases.
This study addresses the limitations of traditional manual user profiling—high cost and poor scalability—and the unreliability and lack of iterative refinement in existing single-pass large language model (LLM)-based approaches. To overcome these challenges, the authors propose PerGent, a novel method that introduces, for the first time in an industrial setting, a multi-agent collaborative framework comprising three LLM-based agents: a generator, a critic, and a coordinator. By integrating structured and unstructured external data sources such as interviews and surveys, PerGent enables multiple rounds of critique-and-refinement iterations to progressively enhance profile quality. Evaluated in a real-world deployment at Kinaxis, the method achieved a 96.9% expert approval rate, significantly outperforming all baseline methods by not only accurately reproducing expert-derived content but also generating substantial high-value supplementary insights.
This work addresses the limited actionability of current AI-generated peer reviews, which often fail to provide authors with concrete guidance for revision. To tackle this issue, the authors propose RbtAct, a novel approach that formulates paragraph-level review generation under reviewer perspective conditioning as a new task and leverages authors’ rebuttals to reviews as implicit supervision signals. Built upon the Llama-3.1-8B-Instruct model, RbtAct integrates supervised fine-tuning with preference optimization guided by rebuttal pairs, and introduces RMR-75K, a dataset comprising 75K review-rebuttal samples. Experimental results demonstrate that RbtAct significantly enhances the actionability and specificity of generated reviews, as evaluated by both human experts and LLM-as-a-judge metrics, while preserving content relevance and factual consistency.
This study investigates how teachers revise student feedback generated by large language models and the resulting impact on instructional content. Drawing on revision data from 117 teachers across 1,349 AI-generated feedback instances, the research integrates sentence embeddings with a machine learning classifier (AUC = 0.75), quantitative analysis, and qualitative coding to systematically characterize teachers’ actual revision behaviors for the first time. Findings reveal that approximately 80% of AI-generated feedback remains unedited, revised feedback tends to be shorter, only about 10% of teachers frequently modify the content, and educators prefer simplifying feedback with high information density. These results uncover key behavioral patterns in human-AI collaborative feedback processes and offer design implications for minimizing unnecessary editing efforts.