Score
Model validation is the practice of designing and implementing evaluation protocols and diagnostic analyses that determine whether a statistical or machine-learned model appropriately represents the phenomenon of interest and meets specified performance, reliability, and safety criteria. It includes constructing test sets and resampling schemes, selecting and computing validation metrics, checking assumptions and calibration, estimating generalization error and uncertainty, probing robustness to distribution shifts and data issues, and producing acceptance criteria and failure-mode analyses used to decide model readiness.
This paper addresses the challenge of evaluating the generalizability of data-driven models to target populations. Methodologically, it integrates statistical learning theory with empirical validation paradigms to propose the first systematic, domain-agnostic framework for model validation. The framework establishes three core principles: (1) sufficiency of validation strategies, (2) mandatory disclosure of limitations, and (3) design of performance metrics enabling cross-model comparability. Its key contribution lies in unifying transparency, reproducibility, and methodological rigor within a single, generalizable validation standard—thereby overcoming longstanding fragmentation and inconsistent reporting in current validation practices. Applicable to diverse data-driven models—including machine learning and statistical prediction models—the framework substantially enhances validation reliability, result comparability across studies, and cross-study reproducibility. It provides a foundational methodological basis for clinical deployment and regulatory evaluation of predictive models.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
To address the loss of fidelity in digital twins caused by dynamic physical system evolution—such as maintenance, wear, and human intervention—this paper proposes a model-verification-based continuous validation framework. The framework integrates real-time monitoring with historical data comparison to construct an interpretable validation metric system, incorporates a lightweight anomaly detection mechanism, and introduces a data-driven parameter self-adaptation estimation algorithm for online twin diagnosis and closed-loop model updating. Unlike conventional static calibration methods, our approach enables long-term trustworthiness preservation and autonomous evolution of the digital twin. Evaluated on an industrial quay crane use case, the framework accurately detects system deviations and dynamically refines model parameters, reducing modeling error by 37.2% and improving maintenance response timeliness by 52%. These results demonstrate significant enhancements in the representativeness, robustness, and engineering practicality of digital twins.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
The AI/ML community faces a severe reproducibility crisis, primarily driven by conceptual ambiguity in verification terminology—such as “reproducibility,” “replicability,” and “dependency/independence”—which undermines research credibility and scientific progress. To address this, we propose the first five-dimensional verification taxonomy, systematically defining core concepts—including reproducibility, dependency vs. independent re-executability, and direct vs. conceptual replicability—by clarifying their objectives, prerequisites, and evaluation criteria. Our framework integrates conceptual analysis, terminological standardization, and methodological modeling to yield a structured verification guideline. It enhances experimental rigor in study design, fosters consensus across the research community on verification practices, and significantly improves cross-team result reproducibility and outcome reliability.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
This study addresses the challenge of effectively validating input model specifications in digital twin simulations, where conventional approaches—relying solely on marginal output distributions—often fail to detect misspecified joint input models. To overcome this limitation, the authors propose a novel statistical validation framework based on sub-trajectory conditioning. By repeatedly restarting simulations from observed system states while conditioning on subsets of random inputs, the method constructs conditional output distributions that enable goodness-of-fit testing of the full joint input model. This approach innovatively transcends the constraints of marginal validation and is complemented by diagnostic tools to pinpoint specific input sources responsible for detected discrepancies. Empirical evaluations on M/M/1 and tandem queueing systems demonstrate the framework’s heightened sensitivity and effectiveness, successfully identifying input model misspecifications that traditional methods overlook.
This study investigates the impact of model selection criteria—such as accuracy versus loss—on test performance in neural classifier training, particularly under early stopping with patience. Through systematic empirical evaluation using k-fold cross-validation on standard benchmarks, the work compares multiple validation metrics, including accuracy, cross-entropy, C-Loss, and PolyLoss, under both early stopping and post-hoc full-trajectory selection strategies. The findings reveal that validation loss–based criteria consistently outperform validation accuracy, which exhibits not only inferior performance but also lower stability. More critically, regardless of the selection criterion employed, the chosen models are typically substantially worse than the best test performance observed during training, thereby exposing a fundamental limitation in current model selection paradigms.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the challenge of root-cause localization in automotive software testing, where high-dimensional sensor data generated during hardware-in-the-loop (HIL) simulations render traditional threshold-based methods ineffective. Existing data-driven approaches often require extensive labeled data and lack interpretability, failing to meet ISO 26262 traceability requirements. To overcome these limitations, the authors propose a two-stage diagnostic framework: first, safety requirements are automatically verified on a dSPACE real-time platform to filter anomalous test records; then, sliding windows of sensor signals are abstracted into statistical, relational, and contextual descriptors, which are fed as fixed prompts to an open-source large language model fine-tuned with 4-bit low-rank adaptation (LoRA). Evaluated on six fault types injected into a gasoline engine, the approach achieves 81.6% accuracy with a minimal 2B-parameter model—comparable to larger models—while operating entirely on a single consumer-grade GPU. This work pioneers the use of instruction-tuned large language models for sensor-level automotive fault diagnosis, demonstrating that diagnostic performance hinges more on task-specific adaptation convergence than on model scale, thereby achieving high accuracy, data efficiency, and explainable decision-making.