Score
Structured synthesis of empirical and theoretical literature (including meta-analysis) to categorize models and methods, derive actionable design principles, and quantify prevalence of evaluation failures across studies.
The exponential growth of scientific literature underscores the urgent need for automated meta-analysis (AMA), yet existing approaches critically lack high-level synthesis capabilities—such as heterogeneity and bias assessment—and end-to-end automation. We systematically reviewed 978 publications (2006–2024) and conducted an in-depth analysis of 54 AMA studies, establishing the first comprehensive, lifecycle-spanning evaluation framework. Our analysis reveals that only 2% of studies achieve full workflow automation, with a pronounced application gap between medical (67%) and non-medical domains. To address these limitations, we propose a next-generation AMA paradigm integrating large language models (LLMs) with statistical rigor, synergizing NLP, explainable AI, and PRISMA-compliant methodology. Empirical results confirm that AMA substantially improves efficiency and reproducibility; however, robust, generalizable, fully automated meta-analysis remains unrealized. This work provides both a theoretical foundation and a concrete technical pathway toward overcoming this fundamental bottleneck.
To address the lack of systematic methodologies for writing survey and tutorial papers in communications and networking, this paper—drawing on editorial experience from top-tier journals—proposes, for the first time, a seven-dimensional writing framework. It integrates literature synthesis, critical analysis, case-based pedagogy, and information visualization to establish a novel survey paradigm that balances tutorial utility with forward-looking insight. The framework systematically covers key stages: topic selection strategy, structural organization, diagrammatic design, and future research direction identification—emphasizing case-driven exposition and actionable insights. Empirical evaluation demonstrates that this roadmap substantially lowers the entry barrier for novice researchers, significantly enhancing survey papers’ readability, comprehension, and scholarly impact. It thus provides a reusable, methodologically grounded foundation for domain knowledge integration and dissemination.
This study addresses the tendency of systematic reviews to overgeneralize by overlooking fine-grained characteristics of included studies, thereby obscuring inter-study relationships and gaps in the literature. To mitigate this limitation, the authors propose an interactive evidence mapping approach that integrates large language models, topic modeling, and visualization techniques to automatically extract themes from heterogeneous review data and construct a dynamically explorable knowledge map. Validation through a scoping review on pedagogical agents in K–12 education demonstrates that this method transcends the constraints of traditional static summaries, substantially enhancing review transparency, effectively uncovering latent patterns and research gaps, and strengthening exploratory analytical capabilities.
Large-scale literature reviews face significant challenges in automated deep analysis and synthesis due to insufficient semantic understanding and structural reasoning capabilities. To address this, we propose DimInd, an interactive system introducing a novel hierarchical compression-based structured representation framework. It unifies paper-level understanding, multi-dimensional comparison, conceptual categorization, and narrative synthesis into a traceable, progressive workflow: papers → comparative tables → conceptual taxonomy → narrative review. DimInd integrates prompt engineering, structured information extraction, hierarchical clustering modeling, and interactive visualization, leveraging large language models (LLMs) for end-to-end semantic parsing and organization. In evaluations with 23 researchers, DimInd significantly reduced cognitive load in information extraction and conceptual organization compared to a ChatGPT baseline, while improving review construction efficiency and structural coherence. It is the first system to enable automated, deep, and narratively coherent synthesis for large-scale scholarly corpora.
This paper addresses reliability concerns and human-centered boundaries in the automated application of large language models (LLMs) for systematic literature reviews (SLRs) in Human-Computer Interaction (HCI). Method: Through empirical LLM experiments, human-AI collaborative workflow design, and HCI methodology analysis, it systematically delineates review stages amenable to automation (e.g., initial screening) versus those requiring human agency (e.g., thematic modeling, cross-study inference). Contribution/Results: The study proposes the “human-centered augmentation” ethical framework, rigorously defining LLMs’ capabilities and limitations in research synthesis. It delivers actionable, rigor-preserving guidelines for SLR practitioners and has been selected for a spotlight discussion at CHI ’25.
Scientific literature exhibits high heterogeneity, and manual meta-analyses are inefficient and error-prone. Method: This paper proposes an intelligent agent pipeline tailored for systematic reviews. It introduces a novel, expert-knowledge-guided, multi-stage collaborative agent architecture that integrates structured human–agent dialogue, cross-modal information extraction (from text, tables, and figures), and standardized semantic mapping—enabling one-time domain knowledge injection and end-to-end knowledge-driven processing. Contribution/Results: We present the first reusable, end-to-end evidence structuring framework that automatically transforms unstructured scientific papers into unified, machine-readable, standardized evidence tables. Evaluated on a meta-analysis task for NMC811 lithium-ion battery cathode materials, our pipeline reduces analysis time from months to minutes while ensuring high reproducibility. This significantly advances the automation, scalability, and reliability of large-scale literature synthesis.
Existing research ideation tools emphasize breadth-oriented idea generation but lack support for iterative refinement, elaboration, and evaluation—hindering literature-grounded, deep-reading–driven conceptual evolution. Method: We propose the first literature-driven interactive research ideation system, integrating a composable “idea element” canvas model with a multi-dimensional (problem/solution/evaluation/contribution) co-evolution mechanism. Our approach innovatively incorporates LLM-powered literature-aware feedback generation, graph-structured idea modeling, and interactive multi-path variant exploration. Contribution/Results: Experiments demonstrate a 42% increase in user-generated idea output and significantly enhanced detail elaboration. Seven researchers successfully applied the system across the full ideation pipeline—from initial topic conception to paper outline revision—validating its efficacy in supporting deep, iterative, literature-informed research design.
This work addresses the challenge of quantifying the academic impact of commercial engineering software such as Ansys Granta, which is hindered by inconsistent citation practices and rapidly growing publication volumes. We propose the first reproducible, semi-automated framework that integrates DOI and citation parsing, expert annotation, and a relational database (Ansys Granta MI Enterprise) to transform heterogeneous usage evidence into a structured knowledge base. As of September 2025, the framework has compiled a multi-source literature repository comprising over 1,100 manually verified records, enabling rapid retrieval, systematic review reproduction, and technology landscape scanning. The resulting knowledge base reveals dominant application domains, key contributing institutions, and integration patterns within CAD/CAE/FEM environments, thereby facilitating systematic tracking and analysis of the long-term technical influence of commercial engineering software.
This study addresses the scalability limitations of quantitative evidence synthesis, which remains heavily reliant on manual effort and thus impedes timely access to reliable knowledge in research, medicine, and policy. The authors propose the first end-to-end multi-agent system capable of autonomously conducting full meta-analyses from natural language research questions. The system performs literature retrieval, screening, data extraction, effect size computation, and random-effects meta-analysis, while also supporting heterogeneity assessment and risk-of-bias evaluation. It generates complete, PRISMA-compliant meta-analysis reports without human intervention. Validated on over 28 studies, the system extracted more than 20 quantitative findings and produced pooled effect estimates highly concordant with those derived manually by experts, substantially enhancing the scalability, transparency, and reliability of evidence synthesis.
This study addresses the frequent under-identification and inadequate interpretation of outliers, non-replicable findings, and highly influential studies in meta-analyses, which often compromise the robustness of conclusions. It clarifies conceptual distinctions among these three types of problematic studies and proposes a systematic diagnostic framework that integrates robust statistical methods, graphical diagnostic tools, and advanced modeling techniques accounting for sampling variance dependencies. This approach enables more accurate detection of anomalous studies while leveraging visualization to facilitate interpretation of their potential sources. By synthesizing recent methodological advances, the work offers meta-analysts practical diagnostic strategies and cautious interpretive guidance, substantially enhancing the reliability and transparency of meta-analytic results.
Existing meta-analysis tools often relegate researchers to mere search operators, fragmenting their cognitive coherence and diminishing their epistemic agency. To address this limitation, this work proposes Research IDE—an integrated development environment grounded in the “research-as-code” paradigm. By deeply integrating multi-agent systems into the scholarly writing workflow, Research IDE introduces a novel “hypothesis breakpoint” mechanism that enables researchers to validate hypotheses in real time, invoke prior knowledge, and maintain a closed-loop reasoning process during manuscript composition. Preliminary expert evaluations indicate that this approach effectively preserves users’ cognitive primacy, fosters the emergence of novel insights, and has garnered strong endorsement from domain specialists.
This study addresses the persistent gap between systematic literature reviews (SLRs) in software engineering and their practical uptake in industry, often referred to as the evidence-to-practice translation gap. To bridge this divide, the work introduces the Evidence to Decision (EtD) framework—originally developed in health sciences—into software engineering for the first time. By convening expert panels to conduct structured evaluations of SLR evidence against multidimensional criteria, the approach generates practitioner-oriented evidence briefs and actionable recommendations. This methodology strengthens the mechanism for translating research findings into real-world decisions, offering the first application of EtD in software engineering, identifying key dimensions essential for generating trustworthy recommendations, and highlighting major challenges that must be addressed for broader adoption of the framework.