LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of predicting a paper's future impact using multi-source heterogeneous evidence available at publication time by proposing the LLM4Impact framework. This method integrates semantic, graph-structural, and large language model information, introducing a novel context-aware dynamic evidence weighting mechanism. Through continuous prefix injection and an adaptive gating network, it achieves deep fusion of multimodal evidence, while an independent calibration module eliminates cross-domain citation scale discrepancies. Evaluated on large-scale benchmarks, the proposed model reduces RMSE by over 10% compared to existing baselines. These results empirically validate the context-dependency of evidence utility and establish a new paradigm for academic impact prediction.
📝 Abstract
Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.
Problem

Research questions and friction points this paper is trying to address.

Scientific Impact Prediction
Heterogeneous Information
Citation Prediction
Evidence Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Impact Prediction
Heterogeneous Information Integration
Prefix Tuning
Context-Aware Gating Mechanism
Calibration
🔎 Similar Papers
No similar papers found.