Score
Designs and evaluates algorithms and procedures that use the observed dataset to choose how to split data into training and calibration (or validation) subsets, including estimating the optimal calibration proportion and selecting the allocation that optimizes objectives such as predictive interval length or coverage. Builds practical selection routines for split choice and validates their performance on synthetic and real data.
In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.
This work addresses the optimal train-test split for ridge regression in the high-dimensional asymptotic regime where both sample size $m$ and feature dimension $n$ diverge, with $n/m o gamma in (0,1)$. The objective is to maximize “completeness” of model evaluation—i.e., minimize the asymptotic bias between test error and theoretical generalization error. We propose the first rigorous large-sample analytical framework for determining the optimal split ratio, leveraging high-dimensional asymptotics and random matrix theory to optimize the ridge regression error functional. Our analysis reveals that the optimal split depends asymptotically only on $m$ and $n$, and is nearly insensitive to the regularization parameter $alpha$. Moreover, its first two asymptotic expansion terms coincide with those of ordinary linear regression, rendering it practically parameter-free. This yields the first theoretically grounded, optimal data allocation principle for model evaluation in high dimensions.
Observational studies are often prone to bias in causal effect estimation due to unmeasured confounding. To address this issue, this work proposes a novel method that automatically optimizes the allocation ratio between planning and analysis samples using plasmode datasets, thereby overcoming the arbitrariness of conventional manual specifications and extending applicability to high-dimensional outcome settings. The approach integrates plasmode simulation, sample splitting, sensitivity analysis, and high-dimensional modeling, and is implemented in the accompanying OptimalSampling R package. Validation in a study on the effects of secondhand smoke exposure in children demonstrates that the proposed method substantially enhances the robustness of causal estimates against unmeasured confounding.
This work addresses the lack of instance-level uncertainty modeling in machine learning predictions for online algorithm design. Methodologically, it is the first to systematically integrate probabilistic calibration—such as Platt scaling and isotonic regression—as a foundation for uncertainty quantification into classical online problems, including ski-rental and online job scheduling. It introduces a calibration-driven competitive ratio analysis framework that yields prediction-confidence-dependent theoretical guarantees. Theoretically, it establishes a quantitative relationship between calibration quality and competitive ratio performance, proving superiority over conventional uncertainty estimation—particularly under high-variance prediction regimes. Empirically, the proposed algorithms significantly outperform baselines on real-world job scheduling datasets and achieve optimal prediction-dependent performance in the ski-rental problem. Crucially, the theoretical guarantees align closely with empirical results, demonstrating both rigor and practical efficacy.
The NP-hard problem of best-subset selection in high-dimensional linear regression motivates this work. We propose an efficient, suboptimal algorithmic framework that integrates greedy search, regularized path tracking, and cross-validation-based model evaluation to yield a stable and scalable solution pipeline. Compared with mainstream heuristic approaches—including LASSO and orthogonal matching pursuit (OMP)—our method substantially reduces computational cost in ultra-high-dimensional settings (p ≥ 1000) while achieving superior trade-offs between model sparsity and predictive accuracy. Comprehensive benchmark experiments on both synthetic data and diverse real-world datasets demonstrate that our approach consistently attains higher solution quality and greater robustness than state-of-the-art baselines. By bridging efficiency and statistical reliability, the proposed framework establishes a new paradigm for high-dimensional sparse modeling.
This study addresses the optimal split between training and calibration sets in split conformal prediction, aiming to simultaneously guarantee valid coverage probability and minimize prediction interval length under finite-sample settings. For the first time, we derive analytically the optimal split proportion that minimizes interval length in a general regression framework, and elucidate how factors such as model complexity influence this proportion. Combining theoretical analysis with a data-driven strategy, our approach is applicable across diverse models, including linear regression, nonparametric regression, and neural networks. Experimental results on both synthetic and real-world datasets demonstrate that the proposed splitting strategy substantially shortens prediction intervals while rigorously maintaining the prescribed coverage guarantees.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.
本文解决了机器学习系统中因数据聚类导致阈值设定不准确的问题,通过提出一种新的有效样本量计算方法来修正阈值设定。