Score
Designs and implements validation protocols that split and hold out data by time or event to assess models' prospective performance and temporal generalizability, including time-split holdouts, event-level confirmation, and temporal external validation. Builds time-ordered cohorts or prospective test sets, evaluates discrimination and calibration over later periods, and detects or quantifies performance inflation caused by leakage of future information.
This study addresses the issues of information leakage and optimistic bias in performance evaluation arising from improper data splitting during machine learning model validation. Focusing on biomedical contexts, it systematically reviews validation strategies and contrasts flawed versus leakage-free designs through eight controlled experiments. Employing techniques such as nested grouped cross-validation for reproducible simulations, this work proposes deployment-oriented, scenario-specific validation guidelines. Its primary contributions include practical tools—namely decision trees, checklists, and code templates—that operationalize the core principle of aligning independent units with deployment objectives within an auditable evaluation framework. Ultimately, these contributions significantly enhance the reliability and standardization of machine learning model validation practices.
This paper addresses the challenge of evaluating the generalizability of data-driven models to target populations. Methodologically, it integrates statistical learning theory with empirical validation paradigms to propose the first systematic, domain-agnostic framework for model validation. The framework establishes three core principles: (1) sufficiency of validation strategies, (2) mandatory disclosure of limitations, and (3) design of performance metrics enabling cross-model comparability. Its key contribution lies in unifying transparency, reproducibility, and methodological rigor within a single, generalizable validation standard—thereby overcoming longstanding fragmentation and inconsistent reporting in current validation practices. Applicable to diverse data-driven models—including machine learning and statistical prediction models—the framework substantially enhances validation reliability, result comparability across studies, and cross-study reproducibility. It provides a foundational methodological basis for clinical deployment and regulatory evaluation of predictive models.
Model evaluation strategies significantly impact reported accuracy, yet systematic comparisons of cross-validation methods across diverse datasets and algorithms remain limited. Method: This study conducts a comprehensive empirical analysis of hold-out, iterative hold-out, and K-fold cross-validation across two real-world datasets (Framingham Heart Study and COVID-19) and three classifiers (decision tree, naive Bayes, k-nearest neighbors). A parameter sensitivity analysis investigates interactions among test set proportion, random seed, and fold count (K). Contribution/Results: We find that a 10% hold-out split consistently yields higher accuracy than the conventional 20% split across most configurations; iterative hold-out substantially reduces variance; and no universally optimal K exists—its efficacy depends critically on dataset size, feature distribution, and learning algorithm. The study demonstrates that validation strategy selection must jointly account for data characteristics and model properties, providing a reproducible evaluation framework and empirically grounded guidelines for robust model assessment.
Existing generative models for synthetic temporal tabular data often produce temporally inconsistent outputs, such as time reversals, repetitions, or implausible trajectories, yet conventional evaluation methods fail to detect these issues due to their neglect of the temporal dimension. This work proposes the first systematic evaluation framework tailored for temporal tabular data, which dynamically selects assessment criteria based on four key properties: temporal representation, sampling regularity, trajectory dependency, and pattern structure. The framework holistically evaluates timestamp validity, cross-sectional structural coherence, entity-level dynamic evolution, and time-varying relationships. By elevating both utility and privacy assessments from static records to the trajectory level, it reveals—across 13 real-world datasets—that traditional evaluations significantly diverge from temporally aware results, with failure modes closely tied to model architecture, thereby underscoring the necessity of explicitly modeling the temporal axis.
This study addresses the longstanding challenge in time-to-event trials of balancing inferential maturity with practical feasibility during interim monitoring, where conventional event-driven or enrollment-driven approaches suffer from inherent limitations. The authors propose the WCR framework, which reframes interim monitoring as an information–time alignment problem. By fixing the cohort size and calibrating follow-up requirements, WCR enables continued enrollment while synchronizing interim analyses with information maturity, reserving later-enrolled patients for the final analysis. The framework explicitly distinguishes follow-up constraints between landmark survival estimators and proportional hazards models, jointly calibrates design parameters and decision thresholds, and integrates constrained optimization, simulation-based calibration, and Bayesian methods, implemented in the open-source R package WCRBayesDesign. In simulations of rare pediatric oncology trials, WCR substantially improves the stability and interpretability of interim analysis timing while rigorously controlling Type I error and maintaining power, outperforming existing strategies.
本文针对AI决策无法在事后验证的问题,提出了基于执行治理3.0的三种验证时构造方法,并通过实验验证了这些方法的有效性。
Current evaluations of financial large language models (LLMs) rely excessively on static benchmarks and lack comprehensive validation across the full system stack. This work proposes the first LLM full-stack verification framework tailored to financial scenarios, encompassing data, model, retrieval, generation, agent behavior, governance, and deployment layers, advocating for verification as an ongoing engineering practice. By introducing a multi-rater LLM-as-a-judge mechanism integrated with scoring rubrics, consistency checks, and auditability, the framework uncovers system failure modes that static benchmarks fail to capture. The study defines critical failure types and advances novel directions—including system-aware benchmarks, agent trajectory validation, rater alignment protocols, and lifecycle-oriented verification standards—thereby shifting the evaluation paradigm from score-driven metrics toward evidence-based readiness for real-world decision-making.
研究通过配对历史审核方法,评估了不同记忆机制在面对未来更新时的充足性问题,并测试了修复方案的效果。
研究解决了语言模型评估不稳定的问题,通过预注册审计发现请求重复性和一致性未达预期标准,提出设计规则和报告清单以改善测量可靠性。
论文提出了一种新的验证框架,用于解决合成研究中决策行为预测准确性的问题,并通过定义三个公正维度和子群体报告要求来提高验证的代表性和公平性。