systematic literature review

Designs and executes structured, reproducible literature syntheses and their supporting artifacts — including systematic, scoping, and PRISMA‑guided reviews, systematic mapping studies, surveys, and meta‑analyses — by formulating research questions, defining and applying inclusion–exclusion criteria, and recording complete search strategies. Builds and curates study corpora and evidence‑synthesis datasets or benchmarks (labeling verified positives and hard negatives), performs critical appraisal of study methods and quality, synthesizes findings and methodological trends, and identifies gaps and future research directions.

systematicliteraturereview

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review

Apr 25, 2025
RF
Raymond Fok
🏛️ University of Washington | Allen Institute for AI

Large-scale literature reviews face significant challenges in automated deep analysis and synthesis due to insufficient semantic understanding and structural reasoning capabilities. To address this, we propose DimInd, an interactive system introducing a novel hierarchical compression-based structured representation framework. It unifies paper-level understanding, multi-dimensional comparison, conceptual categorization, and narrative synthesis into a traceable, progressive workflow: papers → comparative tables → conceptual taxonomy → narrative review. DimInd integrates prompt engineering, structured information extraction, hierarchical clustering modeling, and interactive visualization, leveraging large language models (LLMs) for end-to-end semantic parsing and organization. In evaluations with 23 researchers, DimInd significantly reduced cognitive load in information extraction and conceptual organization compared to a ChatGPT baseline, while improving review construction efficiency and structural coherence. It is the first system to enable automated, deep, and narratively coherent synthesis for large-scale scholarly corpora.

Facilitates synthesis of large paper collections for literature reviewsProvides structured representations to guide literature understandingReduces manual effort in organizing and analyzing research papers

This study addresses the lack of executable and verifiable knowledge representations in existing meta-analyses, which hinders the traceability and reproducibility of critical analytical decisions. To overcome this limitation, the authors propose Executable Analytical Knowledge Representation (EAKR) and introduce MetaSynDec, an agent-based framework that, for the first time, enables explicit modeling, machine-actionable execution, and closed-loop validation of meta-analytic decisions. The system leverages large language models to generate structured knowledge and validates and executes it through deterministic, schema- and contract-based services. Evaluated across 58 synthesis units, EAKR successfully constructed all units, achieved exact evidence-set consistency in 75% of cases, and produced confidence intervals overlapping with published results in 98.2% of cases—substantially outperforming direct LLM-generated approaches.

analytical knowledge representationevidence synthesisexecutable knowledge

This study addresses the tendency of systematic reviews to overgeneralize by overlooking fine-grained characteristics of included studies, thereby obscuring inter-study relationships and gaps in the literature. To mitigate this limitation, the authors propose an interactive evidence mapping approach that integrates large language models, topic modeling, and visualization techniques to automatically extract themes from heterogeneous review data and construct a dynamically explorable knowledge map. Validation through a scoping review on pedagogical agents in K–12 education demonstrates that this method transcends the constraints of traditional static summaries, substantially enhancing review transparency, effectively uncovering latent patterns and research gaps, and strengthening exploratory analytical capabilities.

evidence synthesisliterature visualizationovergeneralization

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Setting The Table with Intent: Intent-aware Schema Generation and Editing for Literature Review Tables

Jul 18, 2025
VP
Vishakh Padmakumar
🏛️ New York University | AI2 | Northwestern University

The exponential growth of academic literature poses significant challenges for efficiently constructing comparative tables in survey papers. Existing schema generation methods suffer from ambiguous evaluation criteria and limited editability. To address these issues, this paper proposes an intent-aware schema generation and editing framework: (1) it introduces intent modeling to mitigate semantic ambiguity in comparative dimension identification; (2) it designs an editable generation pipeline enabling on-demand customization of comparison dimensions; (3) it constructs the first benchmark dataset tailored for conditional schema generation; and (4) it integrates LLM-based prompt engineering with lightweight fine-tuning, combining one-shot generation and multi-stage editing strategies. Experimental results demonstrate that intent enhancement substantially improves schema reconstruction accuracy, while the editing mechanism further refines output quality. Notably, our lightweight fine-tuned model achieves performance competitive with state-of-the-art prompting-based large language models.

Addressing ambiguity in schema evaluation through synthesized intentsDeveloping refinement methods to improve generated schemasGenerating schemas to organize academic literature collections

Latest Papers

What's happening recently
View more

This study addresses the challenges of synthesizing multi-source heterogeneous evidence—such as academic papers, reports, policies, and media content—which vary widely in quality and structure and entail high manual effort. The reliability of current large language models (LLMs) across individual synthesis subtasks remains unclear. To tackle this, the authors propose the Knowledge Synthesis Review (KSR) framework, decomposing the review process into four stages: screening, extraction, analysis, and synthesis. Using a high-agreement expert gold standard (92.2% agreement, κ=0.80), they conduct task-level evaluations of leading LLMs—including GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro—and introduce a dynamic routing mechanism that automatically selects the best-performing model under human supervision. This model-agnostic, auditable, and transparent approach significantly enhances review efficiency and coverage. Experiments reveal no single model dominates all tasks: Claude Sonnet 4 achieves the highest screening accuracy (82.8%), while GPT-5 attains the best recall (91.8%). Moreover, multi-source synthesis uncovers critical themes—such as worker well-being, small and medium enterprises, and Global South perspectives—often missed in single-source analyses.

evidence synthesisknowledge fragmentationLLM reliability

Existing benchmarks lack real-world annotated data covering the full meta-analysis pipeline, limiting comprehensive evaluation of large language models (LLMs) in systematic scientific reasoning. This work introduces MetaSyn, the first expert-annotated benchmark encompassing research questions, PI/ECO criteria, 140,000 PubMed articles, positive and negative examples, and complete search strategies. It further proposes a stage-attribution evaluation framework. Leveraging RAG variants within a protocol-driven agent architecture, the system performs retrieval, screening, and synthesis under structured PI/ECO guidance. Experiments reveal that while retrieval achieves a 90.9% recall at K=200, end-to-end study inclusion recall drops to 52.7%, exposing a critical bottleneck in LLMs’ ability to accurately screen studies meeting precise PI/ECO criteria.

benchmarkingevidence synthesisLLM agents

Current assessments of the integrity of randomized controlled trials (RCTs) rely heavily on manual processes that are complex, subjective, and prone to inconsistency, thereby compromising the quality of evidence-based guidelines. To address this limitation, this work proposes INSPECT-AI, a novel framework that integrates large language models (LLMs) with a knowledge graph grounded in the RIPE-O ontology (RIPE-KG) to automate integrity evaluation, standardize semantic interpretation, and enable full auditability. The authors constructed a RIPE-KG comprising 95 RCTs annotated by experts across 140 assessment criteria and demonstrated that LLM-augmented evaluation significantly enhances efficiency, inter-rater consistency, and traceability. This approach establishes a transparent, reproducible paradigm for evidence synthesis in systematic reviews and guideline development.

assessment consistencyevidence-based clinical guidelinesrandomised controlled trials

Hot Scholars

RD

Ronnie de Souza Santos

Assistant Professor, University of Calgary
Human Aspects of Software EngineeringSoftware TestingSoftware FairnessSoftware Development
MK

Marcos Kalinowski

Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
ZT

Zeerak Talat

University of Edinburgh
NLPOnline AbuseHate SpeechSTS
MZ

Marcos Zampieri

George Mason University
Computational LinguisticsNatural Language Processing