Learning to Ideate for Scientific Impact

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing scientific ideation systems that rely on immediate proxy metrics rather than optimizing for long-term academic impact. We propose leveraging citation influence as a feedback signal to transform delayed scholarly adoption into scalable training signals, alongside a history-based de-looped evaluation protocol to prevent data leakage. Methodologically, we construct an objective-conditioned reward model from large-scale bibliometric data and align large language models through supervised fine-tuning followed by reinforcement learning. Experimental results demonstrate that ideas generated by the RL-aligned model significantly outperform those from baseline and SFT-only models in estimated citation impact. This work establishes an effective paradigm for generating high-impact scientific discoveries by directly optimizing for long-term scholarly influence rather than short-term proxies.
📝 Abstract
Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.
Problem

Research questions and friction points this paper is trying to address.

scientific ideation
large language models
scientific impact
citation-normalized impact
reward alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Ideation
Reinforcement Learning
Reward Model
Citation Impact
Large Language Models