data-driven split selection

Designs and evaluates algorithms and procedures that use the observed dataset to choose how to split data into training and calibration (or validation) subsets, including estimating the optimal calibration proportion and selecting the allocation that optimizes objectives such as predictive interval length or coverage. Builds practical selection routines for split choice and validates their performance on synthetic and real data.

data-drivensplitselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.

Derives asymptotic distributions and constructs valid confidence intervalsDevelops a theoretical framework for sample splitting in A/B testingValidates results through simulations and provides implementation guidance

Test Set Sizing for the Ridge Regression

Apr 27, 2025
AD
Alexander Dubbs

This work addresses the optimal train-test split for ridge regression in the high-dimensional asymptotic regime where both sample size $m$ and feature dimension $n$ diverge, with $n/m o gamma in (0,1)$. The objective is to maximize “completeness” of model evaluation—i.e., minimize the asymptotic bias between test error and theoretical generalization error. We propose the first rigorous large-sample analytical framework for determining the optimal split ratio, leveraging high-dimensional asymptotics and random matrix theory to optimize the ridge regression error functional. Our analysis reveals that the optimal split depends asymptotically only on $m$ and $n$, and is nearly insensitive to the regularization parameter $alpha$. Moreover, its first two asymptotic expansion terms coincide with those of ordinary linear regression, rendering it practically parameter-free. This yields the first theoretically grounded, optimal data allocation principle for model evaluation in high dimensions.

Analyze asymptotic dependence on parameters m and nDetermine optimal train/test split for ridge regressionMaximize model integrity by minimizing error discrepancy

Observational studies are often prone to bias in causal effect estimation due to unmeasured confounding. To address this issue, this work proposes a novel method that automatically optimizes the allocation ratio between planning and analysis samples using plasmode datasets, thereby overcoming the arbitrariness of conventional manual specifications and extending applicability to high-dimensional outcome settings. The approach integrates plasmode simulation, sample splitting, sensitivity analysis, and high-dimensional modeling, and is implemented in the accompanying OptimalSampling R package. Validation in a study on the effects of secondhand smoke exposure in children demonstrates that the proposed method substantially enhances the robustness of causal estimates against unmeasured confounding.

bias sensitivityobservational studiessample splitting

Algorithms with Calibrated Machine Learning Predictions

Feb 05, 2025
JS
Judy Shen
🏛️ Stanford University

This work addresses the lack of instance-level uncertainty modeling in machine learning predictions for online algorithm design. Methodologically, it is the first to systematically integrate probabilistic calibration—such as Platt scaling and isotonic regression—as a foundation for uncertainty quantification into classical online problems, including ski-rental and online job scheduling. It introduces a calibration-driven competitive ratio analysis framework that yields prediction-confidence-dependent theoretical guarantees. Theoretically, it establishes a quantitative relationship between calibration quality and competitive ratio performance, proving superiority over conventional uncertainty estimation—particularly under high-variance prediction regimes. Empirically, the proposed algorithms significantly outperform baselines on real-world job scheduling datasets and achieve optimal prediction-dependent performance in the ski-rental problem. Crucially, the theoretical guarantees align closely with empirical results, demonstrating both rigor and practical efficacy.

Improves performance in ski rental and job schedulingIncorporates machine learning advice in online algorithmsUses calibration to bridge reliability gap in predictions

Solving the Best Subset Selection Problem via Suboptimal Algorithms

Mar 31, 2025
VS
Vikram Singh
🏛️ University of Central Oklahoma | The University of Alabama

The NP-hard problem of best-subset selection in high-dimensional linear regression motivates this work. We propose an efficient, suboptimal algorithmic framework that integrates greedy search, regularized path tracking, and cross-validation-based model evaluation to yield a stable and scalable solution pipeline. Compared with mainstream heuristic approaches—including LASSO and orthogonal matching pursuit (OMP)—our method substantially reduces computational cost in ultra-high-dimensional settings (p ≥ 1000) while achieving superior trade-offs between model sparsity and predictive accuracy. Comprehensive benchmark experiments on both synthetic data and diverse real-world datasets demonstrate that our approach consistently attains higher solution quality and greater robustness than state-of-the-art baselines. By bridging efficiency and statistical reliability, the proposed framework establishes a new paradigm for high-dimensional sparse modeling.

Addresses computational challenges in best subset selectionCompares performance of new and existing methodsProposes suboptimal algorithms for high-dimensional data

Latest Papers

What's happening recently
View more

This study addresses the optimal split between training and calibration sets in split conformal prediction, aiming to simultaneously guarantee valid coverage probability and minimize prediction interval length under finite-sample settings. For the first time, we derive analytically the optimal split proportion that minimizes interval length in a general regression framework, and elucidate how factors such as model complexity influence this proportion. Combining theoretical analysis with a data-driven strategy, our approach is applicable across diverse models, including linear regression, nonparametric regression, and neural networks. Experimental results on both synthetic and real-world datasets demonstrate that the proposed splitting strategy substantially shortens prediction intervals while rigorously maintaining the prescribed coverage guarantees.

coverage guaranteedata splittingprediction interval

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.

GeneralizabilityImbalanced ClassificationResampling

Hot Scholars

MR

Mohammad Reza Deylam Salehi

Ph.D. candidate at Eurecom, Sorbonne University
Information TheoryCoding TheoryMachine LearningAge of Information
EM

Erxue Min

University of Manchester, Baidu Inc.
Information RetrievalLarge Language Model
JL

Junteng Liu

Hong Kong University of Science and Technology
Machine LearningNature Language Processing
FZ

Fan Zhao

Cue Biopharma
T Cell RecognitionAntigen Processing and PresentationImmunotherapyChemistry & Biochemistry