Score
Designs and produces quantitative checks and visualizations that evaluate model or algorithm fit, assumptions, convergence, and estimate reliability. Builds diagnostic plots (residual, spectral, convergence, correlation), computes statistics (effective sample size, diagnostic correlations), flags unreliable estimates, and compares diagnostics across datasets to identify bias, misspecification, or other data/model issues.
This study addresses the limitations of traditional residual plot diagnostics—namely, their reliance on subjective human interpretation, low efficiency, and poor scalability—by introducing computer vision techniques for the first time to automate the assessment of residual plots in linear models. The authors develop the R package autovi and an accompanying Shiny-based interactive web application, autovi.web. Their approach leverages deep visual models to quantify the strength of structural signals in residual plots and produces interpretable diagnostic metrics. This methodology significantly enhances the consistency and efficiency of model fit evaluation, offering statisticians and data analysts a robust, objective, and scalable tool for automated diagnostic assessment in statistical modeling.
Diagnostic meta-analyses frequently exhibit spurious associations between disease prevalence and test accuracy, primarily driven by pervasive biases—such as imperfect reference standards and spectrum effect confounding—in primary diagnostic accuracy studies. Method: We formalized the data-generating process using directed acyclic graphs (DAGs), integrated causal inference principles with simulation experiments, and systematically identified the sufficient conditions under which such bias-induced associations arise. Contribution/Results: We demonstrate that spurious prevalence–accuracy associations emerge when the reference standard exhibits misclassification or when unmeasured confounders jointly influence patient spectrum and prevalence distribution. Crucially, we propose and validate the first statistically principled correction method to eliminate this bias—achieving robust disassociation upon appropriate adjustment for these sources. This work establishes a rigorous theoretical framework and practical methodology for bias detection, causal interpretation, and statistical calibration in diagnostic meta-analysis.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This work addresses a critical limitation in existing chart-to-code generation methods, which rely on reference code containing unobservable latent variables for supervision, often leading to model hallucination and over-specification. The study systematically identifies this issue and introduces an observation-aligned supervision framework that restricts training objectives to quantities directly inferable from chart images—such as boxplot statistics, pie chart proportions, and histogram bin weights—ensuring alignment between supervision signals and visual observations. By integrating chart understanding from vision-language models with supervised fine-tuning and data rewriting techniques, the proposed approach significantly improves both the accuracy of observable attribute recovery and code executability on ChartMimic and ChartX benchmarks, demonstrating the pivotal role of observation-aligned supervision in enhancing model performance.
Manual review of unstructured electronic health record (EHR) text to construct reference standards for large-scale database studies is time-consuming and labor-intensive. Method: We propose an NLP-driven, multi-wave adaptive sampling validation framework that integrates NLP-assisted annotation, quantitative bias analysis, and a predefined termination rule based on error convergence—dynamically optimizing both sample selection and stopping timing while preserving measurement accuracy. Results: Empirical evaluation shows that NLP reduces per-record review time by 40%; multi-wave sampling with termination criteria skips 77% of records requiring no manual review, with negligible impact (<0.5 percentage points) on final algorithm performance estimation bias. The framework significantly improves validation efficiency, feasibility, and scalability, offering a reproducible, resource-efficient, and standardized validation pathway for coded-algorithm-based health outcome measurement.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
This study addresses the lack of interpretable, scalable diagnostic tools for extreme value regression models that can identify regions in covariate space where local fit is poor. The authors propose two visualization-based diagnostics—standardized tail plots and normalized residual plots—leveraging the asymptotic distribution of normalized exceedance probabilities to construct sample-size-invariant uncertainty bounds. This enables consistent assessment of both global and local goodness-of-fit. Notably, the approach provides the first framework for local diagnostics in low-dimensional or non-Euclidean covariate domains, supports model comparison across varying sample sizes, and facilitates large-scale model screening. In two real-world applications, the method successfully evaluated thousands of candidate models, yielding actionable modeling recommendations that substantially enhance the reliability and practical utility of extreme value regression models.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.