Score
Systematically evaluating methods, benchmarks, and synthesized findings to assess validity, limitations, and gaps, and to contextualize results relative to the broader literature and future research directions.
Accurately evaluating the performance of analytical methods in simulation studies is hindered by underreporting and inconsistent handling of “missingness” issues—such as algorithm failure or non-convergence—that compromise validity and reproducibility. Method: We conducted a large-scale empirical analysis of 482 methodological simulation studies, systematically extracting metadata, applying qualitative coding, and performing case studies—including publication bias correction—to quantify the prevalence and reporting practices of missingness. Contribution/Results: We found that only 23% of studies mentioned missingness and merely 14% described mitigation strategies. Based on these findings, we developed a novel missingness taxonomy tailored to simulation research and proposed actionable principles—including mandatory missingness reporting—alongside a comprehensive, end-to-end practice guideline covering reporting, handling, and replication. Validation confirmed substantial improvements in transparency, comparability, and reproducibility of simulation studies.
This study addresses the tendency of systematic reviews to overgeneralize by overlooking fine-grained characteristics of included studies, thereby obscuring inter-study relationships and gaps in the literature. To mitigate this limitation, the authors propose an interactive evidence mapping approach that integrates large language models, topic modeling, and visualization techniques to automatically extract themes from heterogeneous review data and construct a dynamically explorable knowledge map. Validation through a scoping review on pedagogical agents in K–12 education demonstrates that this method transcends the constraints of traditional static summaries, substantially enhancing review transparency, effectively uncovering latent patterns and research gaps, and strengthening exploratory analytical capabilities.
Scientific papers often report research limitations in vague or imprecise ways, undermining reproducibility and scholarly trust. To address this, we introduce the first end-to-end benchmark specifically designed for limitations—encompassing automatic extraction, generation, and dual-layer evaluation (fine-grained + meta-evaluation). Our contributions include: (1) a limitations-oriented Retrieval-Augmented Generation (RAG) framework; (2) a high-quality, manually annotated dataset integrating papers from major venues (ACL, NeurIPS, PeerJ) with external peer reviews; and (3) a suite of multidimensional automated evaluation metrics alongside a rigorous meta-evaluation protocol. Experiments demonstrate that our approach significantly improves the relevance and verifiability of generated limitations, while the evaluation framework exhibits strong discriminative power and robustness across diverse models and settings. All data, annotations, and code are publicly released to advance AI-assisted research integrity.
Current AI research tools lack evaluation benchmarks that simultaneously account for usability, interpretability, and integration into scientific workflows, making it difficult to assess their practical reliability in academic settings. This work proposes a comprehensive evaluation framework that integrates human-centered dimensions—such as usability and interpretability—with computational metrics. Through a human-AI collaborative approach—including explainable AI (xAI) analysis, source tracing validation, task-oriented testing, and workflow integration observation—the study systematically evaluates AI-powered question-answering and literature review tools on both exploratory and precision-oriented tasks. Findings reveal a core tension: while these tools effectively support initial exploration by providing useful overviews, they exhibit unreliable precision in factual extraction, with xAI highlights often misaligned with actual answers. Similarly, literature tools aid preliminary discovery but suffer from poor reproducibility and low transparency, necessitating rigorous human verification.
This study addresses the lack of automated assessment of methodological quality and risk of bias (RoB) in biomedical literature by introducing RoBBR, the first NLP benchmark specifically designed for RoB evaluation. RoBBR is built upon 500+ peer-reviewed papers and 2,000 expert-annotated instances, covering four fine-grained tasks: study design identification, RoB domain classification, bias type determination, and evidence strength rating. It is the first effort to operationalize established RoB frameworks—such as those from Cochrane—in a rigorously evaluable NLP setting, incorporating a strict content alignment verification protocol. Empirical evaluation reveals that state-of-the-art large language models underperform human experts by 32.7 percentage points in average F1 score across all four tasks, underscoring the substantial challenge of automating methodological appraisal. The benchmark—including its dataset, annotation guidelines, and open-source code—is publicly released to advance trustworthy AI for scientific evidence assessment.
Current evaluation of long-form question answering systems predominantly relies on human pairwise preference judgments, which often fail to capture the nuanced, expert-level assessment of in-depth research report quality. This work systematically examines the applicability and limitations of such meta-evaluation approaches in scientific QA using the ScholarQA-CS2 benchmark. The study finds that pairwise preferences are suitable only for system-level comparisons, whereas metric-level evaluation requires explicit dimension-wise annotations combined with domain-expert review. It identifies subjectivity as a central challenge and proposes a set of meta-evaluation design guidelines aligned with expert expectations, offering practical recommendations for future evaluation frameworks, annotator expertise matching, and reporting practices in deep research-oriented QA systems.
This study addresses the systematic bias in current AI usage assessment methods, which often overlook contextual differences across countries and academic disciplines, leading to inaccurate estimations of AI involvement in scholarly writing. Leveraging large-scale journal publication data from Dimensions, the authors employ a large language model to rewrite human-authored abstracts and establish customized “AI similarity” baselines tailored to specific country–discipline combinations. This approach effectively disentangles inherent disciplinary and national writing styles from genuine AI-generated characteristics. The proposed contextualized benchmark substantially mitigates the distortions introduced by uniform thresholds— which tend to overestimate AI use in certain regions and fields while underestimating it in others—and demonstrates markedly fairer and more accurate evaluation performance for publications projected in 2025.
This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.
This study addresses the lack of high-quality benchmarks for evaluating models’ ability to assess the feasibility of scientific claims. To this end, we introduce a novel benchmark dataset comprising 197 original materials science claims, each annotated by domain experts with a five-point feasibility rating and an open-ended natural language explanation. This benchmark uniquely combines non-literature-derived claims, expert annotation, and structured scoring—a design that substantially mitigates training data contamination risks while increasing task complexity. Using this dataset, we conduct baseline evaluations with GPT-family models, revealing significant limitations in current large language models’ capacity for complex scientific reasoning. Our work establishes a robust foundation for future research in scientifically grounded model evaluation.
This study addresses the challenge in applied microeconomics of effectively synthesizing empirical evidence, predicting effect sizes in new contexts, and correcting for publication bias. It proposes an integrated methodological framework that combines systematic literature review, covariate reweighting for extrapolation, and selection bias correction techniques—applicable even with as few as three prior studies. The approach innovates by offering a transparent and reproducible pipeline for out-of-sample effect prediction and, for the first time, quantifies the extent to which publication bias distorts average treatment effects. Empirical results demonstrate that bias-corrected average effects amount to only 12%–21% of naive unweighted averages, substantially improving predictive accuracy and enhancing the relevance of findings for policy design.