Score
Designs and executes reproducible, comprehensive literature reviews that catalogue and code publications by theme and year, apply systematic search and inclusion/exclusion criteria, and extract structured data for analysis. Produces thematic or quantitative syntheses of methodological and empirical trends, identifies capability and reliability gaps, and documents protocols and evidence supporting the conclusions.
To address the lack of systematic methodologies for writing survey and tutorial papers in communications and networking, this paper—drawing on editorial experience from top-tier journals—proposes, for the first time, a seven-dimensional writing framework. It integrates literature synthesis, critical analysis, case-based pedagogy, and information visualization to establish a novel survey paradigm that balances tutorial utility with forward-looking insight. The framework systematically covers key stages: topic selection strategy, structural organization, diagrammatic design, and future research direction identification—emphasizing case-driven exposition and actionable insights. Empirical evaluation demonstrates that this roadmap substantially lowers the entry barrier for novice researchers, significantly enhancing survey papers’ readability, comprehension, and scholarly impact. It thus provides a reusable, methodologically grounded foundation for domain knowledge integration and dissemination.
This work addresses the inefficiency and steep learning curve researchers often encounter when trying to map academic papers to their corresponding implementation code. To bridge this gap, the authors propose an automated tool powered by large language models (LLMs) that achieves cross-modal semantic alignment between scholarly texts and source code for the first time. By integrating program analysis techniques, the method automatically identifies code segments that implement specific research ideas described in a paper and generates high-quality traceability mappings. This approach substantially reduces the manual effort required for alignment, enhances the comprehensibility of research software, and improves reproducibility. Preliminary experiments demonstrate the tool’s practicality and effectiveness in real-world scenarios.
This study addresses the time-consuming and inefficient nature of manual code review by conducting a structured systematic literature review (SLR) of 119 publications—the first to propose a comprehensive, task-dimensional taxonomy for automated code review. Methodologically, it integrates machine learning, information retrieval, program analysis, and natural language processing techniques, and empirically evaluates approaches using datasets from GitHub, Gerrit, and other platforms, with metrics including BLEU, F1, and MAP. Key contributions are: (1) a refined classification of 12 automated review tasks and 7 core technical paradigms; (2) a curated inventory of 32 publicly available tools and datasets; (3) identification of critical bottlenecks in data-driven methods—particularly regarding interpretability and cross-project generalizability; and (4) a reproducible evaluation benchmark alongside four concrete directions for future research. The findings are synthesized into a rigorous, structured SLR report.
Scientific software frequently suffers from poor robustness, low maintainability, and weak sustainability. To address these challenges, this work systematically integrates software engineering best practices with domain-specific research requirements, proposing a set of ten high-quality principles for building scientific software across its entire lifecycle. The principles cover critical phases—including project planning, readable coding, version control, automated testing, modular design, reproducibility assurance, performance optimization, and long-term maintenance—and are supported by technical enablers such as automated documentation generation, continuous integration, and performance profiling. Designed to be both broadly applicable and practically actionable, the framework has been empirically validated across multiple scientific domains. Results demonstrate significant improvements in software reliability, reusability, and collaborative efficiency within research communities, thereby enhancing the academic impact of scientific tools and advancing open science and reproducible research ecosystems.
Large-scale literature reviews face significant challenges in automated deep analysis and synthesis due to insufficient semantic understanding and structural reasoning capabilities. To address this, we propose DimInd, an interactive system introducing a novel hierarchical compression-based structured representation framework. It unifies paper-level understanding, multi-dimensional comparison, conceptual categorization, and narrative synthesis into a traceable, progressive workflow: papers → comparative tables → conceptual taxonomy → narrative review. DimInd integrates prompt engineering, structured information extraction, hierarchical clustering modeling, and interactive visualization, leveraging large language models (LLMs) for end-to-end semantic parsing and organization. In evaluations with 23 researchers, DimInd significantly reduced cognitive load in information extraction and conceptual organization compared to a ChatGPT baseline, while improving review construction efficiency and structural coherence. It is the first system to enable automated, deep, and narratively coherent synthesis for large-scale scholarly corpora.
This study addresses the lack of executable and verifiable knowledge representations in existing meta-analyses, which hinders the traceability and reproducibility of critical analytical decisions. To overcome this limitation, the authors propose Executable Analytical Knowledge Representation (EAKR) and introduce MetaSynDec, an agent-based framework that, for the first time, enables explicit modeling, machine-actionable execution, and closed-loop validation of meta-analytic decisions. The system leverages large language models to generate structured knowledge and validates and executes it through deterministic, schema- and contract-based services. Evaluated across 58 synthesis units, EAKR successfully constructed all units, achieved exact evidence-set consistency in 75% of cases, and produced confidence intervals overlapping with published results in 98.2% of cases—substantially outperforming direct LLM-generated approaches.
Existing benchmarks for systematic reviews are limited in scale and disciplinary coverage, hindering robust cross-domain method evaluation and metascientific research. This work introduces Webis-SR4ALL-26, the first large-scale corpus of systematic reviews spanning all scientific fields, comprising 318,710 reviews. Through a multi-stage preprocessing pipeline, the corpus is enriched with linked OpenAlex metadata, reference lists, and structured methodological artifacts—such as search strategies and inclusion/exclusion criteria. Notably, the project pioneers the structured extraction and standardization of executable search strategies. The authors release both the open corpus and associated processing code. Large-scale baseline retrieval experiments conducted on OpenAlex demonstrate the resource’s effectiveness for cross-disciplinary literature retrieval, screening, and metascientific analysis.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This study addresses the tendency of systematic reviews to overgeneralize by overlooking fine-grained characteristics of included studies, thereby obscuring inter-study relationships and gaps in the literature. To mitigate this limitation, the authors propose an interactive evidence mapping approach that integrates large language models, topic modeling, and visualization techniques to automatically extract themes from heterogeneous review data and construct a dynamically explorable knowledge map. Validation through a scoping review on pedagogical agents in K–12 education demonstrates that this method transcends the constraints of traditional static summaries, substantially enhancing review transparency, effectively uncovering latent patterns and research gaps, and strengthening exploratory analytical capabilities.
Traditional peer review relies excessively on authors’ narrative accounts, making it difficult to verify the authenticity of reported results. This work proposes a “code-first” review paradigm in which authors submit executable research artifacts alongside a checklist of claims. An AI-driven review infrastructure automatically provisions execution environments, runs experiments, audits code paths, and precisely maps each claim to empirical evidence, producing a standardized review package for human evaluation. By introducing AI as a core component of the review process, this study pioneers the concepts of claim-evidence contracts, generative review views, and the review package abstraction. It shifts the focus of peer review from narrative persuasion to verifiable, reproducible evidence and presents a comprehensive protocol framework encompassing system architecture, empirical validation, and analysis of governance challenges such as AI bias and prompt injection.