Score
Interpreting experimental outputs, statistical summaries, and comparative evaluations to determine whether observed differences are systematic and significant across runs and conditions, and to translate those findings into practical recommendations (e.g., detector selection and deployment guidelines).
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.
In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.
Classical variance change-point detection methods suffer from p-value bias and inflated Type I error due to data reuse in model selection. Existing post-selection inference (PSI) frameworks are restricted to mean-shift detection and do not extend to variance changes. Method: This paper introduces the first PSI framework for variance change-point detection, proposing two general-purpose constructions for post-selection p-values compatible with diverse algorithms (e.g., piecewise constant modeling) and test forms (e.g., constrained likelihood ratio tests). Leveraging conditional inference, convex optimization, and statistical functional theory, the methods rigorously control Type I error conditional on the selected model path and yield uniformly calibrated p-values. Contribution/Results: We establish theoretical validity of the proposed procedures and demonstrate, via extensive simulations and real-data analyses, their improved statistical power and accurate p-value calibration—overcoming a key limitation of PSI in detecting heteroscedastic structural changes.
Null Hypothesis Significance Testing (NHST) suffers from fundamental limitations, including conflation of statistical and practical significance, sensitivity to sample size, and inability to distinguish “failure to reject” from “acceptance” of the null hypothesis. This paper introduces REACT—a novel hypothesis testing framework that integrates Bayesian logic with frequentist interpretability. REACT employs a dual-threshold decision rule based on confidence intervals for effect sizes and the minimal effect size of interest (MES), enabling, for the first time, joint inference over multiple parameters without multiplicity correction. Crucially, it formally distinguishes “absence of evidence” from “evidence of absence.” Empirical evaluation across multiple real-world datasets demonstrates that REACT substantially enhances scientific robustness and reproducibility of inferences, while maintaining computational and operational complexity comparable to NHST—facilitating straightforward adoption by researchers.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study addresses the common reliance on unrealistic assumptions about average treatment effects in experimental and observational research designs. It proposes a novel paradigm that shifts focus from directly positing average effects to modeling the full distribution of individual treatment effects, from which more plausible assumptions about average effects can be derived. By integrating distributional modeling with cross-disciplinary case studies, the approach demonstrates its validity and utility across diverse fields—including medicine, economics, and psychology—offering researchers a principled, heterogeneity-aware framework for specifying effect sizes grounded in empirical realism rather than idealized assumptions.
This study addresses the lack of decision-oriented evaluation methodologies in current machine translation quality estimation (QE) systems. It introduces receiver operating characteristic (ROC) analysis into QE evaluation for the first time, complementing and validating against conventional metrics. Experimental results demonstrate that ROC analysis not only aligns consistently with existing evaluation outcomes but also yields actionable performance insights. By providing a clearer understanding of trade-offs between true positive and false positive rates across varying decision thresholds, this approach significantly enhances the practical utility of QE assessment and offers robust guidance for deployment decisions in real-world applications.
This study addresses the substantial bias often introduced in meta-analyses when estimating standard deviations solely from the five-number summary—specifically, the minimum, maximum, and median—due to insufficient information, which can compromise inferential reliability. To mitigate this issue, the authors propose a novel estimation method based on a scaled Beta distribution that incorporates data shape characteristics to improve accuracy. A comprehensive sensitivity analysis is systematically conducted to quantify estimation uncertainty. Through extensive simulation studies and real-data applications, the proposed approach demonstrates markedly superior performance over conventional estimators across a variety of underlying distributions. Additionally, the authors provide an interactive web tool to facilitate practical implementation, enabling researchers to readily assess and correct potential bias in standard deviation estimates, thereby enhancing the robustness of meta-analytic findings.
This study addresses the lack of systematic evaluation in outlier handling within meta-analyses, which can render conclusions susceptible to subjective methodological choices. We preregistered and systematically compared four commonly used outlier detection and adjustment methods—including Winsorizing and DFBETAS—across 358 meta-analyses in the behavioral sciences, employing random-effects models with unrestricted weighted least squares estimation. For the first time at scale, we quantified how these approaches influence pooled effect sizes, statistical significance, and the smallest effect size of interest. Results indicate that while outlier treatment exerts minimal impact on average effect estimates (median change ≤ 0.047), it reverses significance judgments in 11.5% of cases and alters effect size interpretations in 15.9%, particularly among marginally significant findings.