gap analysis

Systematic identification and characterization of open research questions, unresolved challenges, and cross-method insights to prioritize future work. Helps synthesize where communities have converged and where critical problems or practical considerations remain.

gapanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

The large language model (LLM) era urgently requires high-quality, authentic, and structured scientific reasoning data to advance trustworthy AI research. Method: This paper pioneers the systematic adoption of OpenReview—a dynamic, expert-curated academic interaction dataset encompassing papers, peer reviews, author rebuttals, meta-reviews, and final decisions—as a scarce source of expert-level alignment data. Our approach emphasizes data governance, standardized benchmark design, an ethical usage framework, and community co-governance—model-agnostic by design. Contributions: (1) We establish OpenReview’s irreplaceable value across three dimensions: scalability of review-based evaluation, authenticity of open scientific benchmarks, and empirical rigor in alignment research; (2) we introduce the first standardized OpenReview usage guidelines and a shared-responsibility agreement; and (3) we catalyze three novel research paradigms—review-augmented reasoning, open scientific benchmarking, and value-aligned AI—thereby laying a robust academic infrastructure for explainable, value-consistent LLMs.

Creating open-ended benchmarks from expert deliberationEnhancing quality, scalability, and accountability of peer reviewSupporting alignment research via expert assessments and values

To address the lack of systematic quality assurance for bibliographic and citation data in the OpenCitations infrastructure, this paper designs and implements an interpretable validation and dynamic quality monitoring framework tailored to the OpenCitations Data Model (OCDM). Methodologically, it integrates a customizable rule engine, SPARQL-based consistency checking, semantic constraint validation, and an incremental quality dashboard, enabling error attribution analysis and quantitative assessment. Key contributions include: (1) the first interpretable validation tool specifically designed for OCDM; and (2) a novel dynamic, sustainable quality tracking mechanism. Experimental evaluation demonstrates that the framework accurately identifies structural and semantic defects in the Matilda dataset and detects, localizes, and quantifies persistent issues—including duplication, incompleteness, and inconsistency—in OpenCitations Meta. The approach significantly enhances data reliability and fills a critical gap in systematic quality assurance for open citation data.

Ensuring quality in bibliographic and citation datasetsMonitoring data quality in open research infrastructuresValidating metadata and citations from diverse sources

This study addresses a significant gap in the digital humanities concerning methodological transparency and open peer review, where data management documentation and open review mechanisms are largely absent. Through a systematic investigation combining content analysis of published works with a survey of publishing practices, this research reveals that only a minimal number of articles provide reusable methodological documentation, and the vast majority of journals and conferences continue to employ traditional blind peer review. As the first comprehensive assessment of open science practices in the digital humanities, this work fills a critical void in the literature and underscores the urgent need for improved transparency and reproducibility in scholarly communication within the field.

Digital Humanitiesmethodological transparencyopen peer review

This work addresses the challenge of efficiently and objectively evaluating the novelty of scholarly submissions in peer review, particularly amidst the rapidly expanding volume of scientific literature. To this end, we propose an intelligent agent system powered by large language models that implements a four-stage pipeline—contribution extraction, semantic retrieval, hierarchical classification coupled with fine-grained full-text comparison, and evidence synthesis—to deliver an end-to-end, traceable, and evidence-based novelty assessment grounded in actual published works. This approach effectively mitigates hallucination risks inherent in large language models. Deployed on over 500 submissions to ICLR 2026, our method accurately identifies relevant prior work omitted by authors, significantly enhancing the fairness, consistency, and interpretability of peer reviews. All evaluation reports have been publicly released.

academic noveltyevidence-based evaluationnovelty assessment

IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback

Oct 05, 2024
KP
Kevin Pu
🏛️ University of Toronto | University of Washington | Allen Institute for AI

Existing research ideation tools emphasize breadth-oriented idea generation but lack support for iterative refinement, elaboration, and evaluation—hindering literature-grounded, deep-reading–driven conceptual evolution. Method: We propose the first literature-driven interactive research ideation system, integrating a composable “idea element” canvas model with a multi-dimensional (problem/solution/evaluation/contribution) co-evolution mechanism. Our approach innovatively incorporates LLM-powered literature-aware feedback generation, graph-structured idea modeling, and interactive multi-path variant exploration. Contribution/Results: Experiments demonstrate a 42% increase in user-generated idea output and significantly enhanced detail elaboration. Seven researchers successfully applied the system across the full ideation pipeline—from initial topic conception to paper outline revision—validating its efficacy in supporting deep, iterative, literature-informed research design.

Bridges gap between broad idea generation and deep refinementEnables evolving and composing idea facets for research developmentSupports iterative refinement of research ideas with literature feedback

Latest Papers

What's happening recently
View more

Traditional citation-based metrics struggle to capture the full scholarly impact of open research infrastructures. Addressing this limitation, this study presents the first domain-agnostic NLP-based scientometric framework, exemplified through a case study of the LXCat platform. The approach integrates chemical entity recognition, extraction of dataset and solver mentions, institutional geolocation mapping, and topic modeling to uncover implicit usage patterns beyond formal citations. Applied to LXCat, the framework systematically reveals evolving data dependencies, latent adoption trends, and thematic shifts within low-temperature plasma research. By moving beyond conventional citation analysis, this work establishes a scalable, data-driven paradigm for evaluating and governing the influence of open scientific infrastructures.

Low Temperature PlasmaLXCatOpen Research Information Infrastructures

This study systematically evaluates the coverage of OpenCitations—an open scholarly infrastructure—for research outputs from six Italian universities and its potential to serve as an alternative to commercial bibliographic databases. By matching persistent identifiers (DOIs, PMIDs, and ISBNs) from publications recorded in institutional IRIS systems against OpenCitations Meta data, and complementing this with metadata reconciliation and citation linkage analysis, the research provides the first multi-institutional, CRIS-level quantification of OpenCitations’ coverage, benchmarked against Scopus and Web of Science. Findings indicate that OpenCitations covers, on average, over 40% of IRIS-indexed publications, demonstrating performance comparable to leading commercial databases overall. However, significant gaps remain in its coverage of monographs and critical editions in the humanities and social sciences, highlighting current limitations and key areas for future enhancement.

bibliographic metadataopen science infrastructuresOpenCitations

This study addresses systemic challenges confronting software engineering research—including reviewer overload, metric-driven incentives, publication distortions, and AI misuse—which have repeatedly resisted isolated reform efforts. For the first time in this domain, the work integrates complex systems theory with theories of change to develop a novel analytical framework that combines ecosystem metaphors with feedback loop analysis. This approach uncovers the intrinsic coupling mechanisms among these interrelated problems, identifies critical leverage points, and elucidates the structural roots underlying the failure of current interventions. Building on these insights, the paper proposes coordinated, multi-level systemic interventions designed to foster a healthier and more sustainable research ecosystem, offering both a theoretical foundation and actionable pathways for transformative change.

incentive structurespublication practicesresearch ecosystem

Existing approaches struggle to assess, at a fine-grained level, how citations in interdisciplinary research substantively integrate ideas from multiple fields. This work proposes a citation-purpose classification framework tailored for interdisciplinary scholarship, combining manual annotation, citation context analysis, and qualitative categorization to construct the first quantifiable system that measures both the depth and significance of citation engagement. Validated on publications at the intersection of natural language processing and computational social science, the framework not only uncovers actual patterns of interdisciplinary citation usage but also provides an actionable metric for evaluating citation quality.

citation engagementcitation purposecomputational approaches

This work addresses the challenge researchers often face in balancing novelty with effective grounding in existing literature when developing new ideas, as well as the lack of tools that support dynamic interaction between emerging concepts and relevant scholarly works. The paper introduces a novel “literature-driven idea pivoting” mechanism—a closed-loop framework that integrates idea drafting, dynamic literature retrieval, semantic clustering, and generative critical feedback to enable co-evolution of research ideas and the literature space. The system performs context-aware analysis of partial idea content and provides real-time improvement suggestions based on clusters of relevant papers. Experimental results demonstrate that this approach significantly enhances the quality of user-generated ideas and strengthens researchers’ ability to comprehend and leverage the scholarly context effectively.

dynamic contextualizationidea refinementliterature landscape

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
YB

Yi Bu

Assistant Professor, Department of Information Management, Peking University
scholarly communicationbibliometricsscience policyscience of science
LW

Lingfei Wu

University of Pittsburgh
science of scienceteam science
YL

Yiling Lin

University of Pittsburgh
science of scienceteam scienceinnovation
CN

Chaoqun Ni

University of Wisconsin-Madison
ScientometricsMetascienceQuantitative Science StudiesScience Policy