holdout validation

Using held-out data subsets to estimate generalization and validate discoveries (e.g., persona assignments or expert selection), often via probe-success rules or learned scorers to select best-performing models or profiles.

holdoutvalidation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.

Bayesian inferencebest subset selectionexpert knowledge

Chiseling: Powerful and Valid Subgroup Selection via Interactive Machine Learning

Sep 23, 2025
NC
Nathan Cheng
🏛️ Harvard University | Stanford University

In regression and causal inference, controlled subgroup selection aims to identify subpopulations—within the covariate space—whose average response or treatment effect significantly exceeds a prespecified threshold, while providing rigorous statistical inference guarantees. Existing methods often sacrifice efficiency (e.g., via data splitting), lack flexibility, or fail to ensure valid inference. This paper introduces Chiseling: an interactive machine learning framework that iteratively refines candidate subgroups via targeted contraction. Crucially, contraction directions are determined solely from data outside the current subgroup, ensuring valid hypothesis testing under only finite-moment conditions. The framework seamlessly incorporates domain knowledge and accommodates arbitrary machine learning algorithms, supporting both randomized experiments and observational studies. Simulation and empirical analyses demonstrate that Chiseling achieves substantially improved subgroup detection power and practical utility over existing guaranteed methods—without compromising statistical validity.

Enabling interactive refinement of subgroups while maintaining statistical controlIdentifying subgroups with specific treatment effects using inferential guaranteesOvercoming limitations of existing subgroup selection methods lacking validity

Strategic Hypothesis Testing

Aug 05, 2025
SH
Safwan Hossain
🏛️ Harvard University | MPI for Intelligent Systems

This paper studies strategic hypothesis testing within a principal–agent framework: the agent holds private beliefs about product efficacy and may manipulate submitted data to maximize expected payoff; the principal must design a p-value threshold to balance Type I and Type II error risks. Methodologically, it innovatively integrates game-theoretic reasoning with classical statistical hypothesis testing by imposing incentive-compatibility constraints. The analysis establishes that the optimal p-value threshold exhibits a monotonic, analytically tractable structure in the agent’s strategic behavior, yielding a closed-form solution. Theoretically, it demonstrates that regulators can endogenously mitigate strategic reporting by calibrating the critical p-value, thereby unifying statistical robustness with incentive compatibility. Empirical validation using FDA drug approval data confirms the model’s predictive power, providing regulators with an interpretable and computationally tractable framework for optimizing approval policies.

Balances false positives and negatives with p-value thresholdsExamines strategic agent behavior in hypothesis testingValidates model using drug approval data

Confidence on the Focal: Conformal Prediction with Selection-Conditional Coverage

Mar 06, 2024
YJ
Ying Jin
🏛️ Harvard University | University of Pennsylvania

Under data-driven selection, conventional prediction intervals fail to guarantee marginal coverage for the selected units—compromising reliability for focal samples. Method: We propose the first finite-sample exact coverage framework for post-selection inference, extending Mondrian conformal prediction to multiple test samples and non-equivariant models while accommodating arbitrary permutation-invariant selection rules. Our approach integrates conditional randomization tests, top-K or optimization-driven selection, conformal p-values, and preliminary screening prediction sets to enable efficient computation. Contribution/Results: Evaluated on drug discovery and health risk prediction tasks, our method substantially improves empirical coverage for focal units, ensuring statistically valid inference in real-world decision-making scenarios. This provides the first provably exact finite-sample coverage guarantee for post-selection prediction intervals under general selection mechanisms.

Addresses selection bias in conformal predictionEnsures valid coverage for selected focal unitsGeneralizes to multiple test units and classifiers

Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning

Oct 12, 2023
ML
Michael Lingzhi Li
🏛️ Harvard Business School | Harvard University

In causal subgroup identification, conventional methods suffer from high estimation noise in conditional average treatment effect (CATE) estimation and multiplicity issues arising from two-stage procedures. To address these challenges, this paper proposes the Global Adaptive Treatment Effect Sets (GATES) uniform confidence band method. Grounded in randomized trial design and empirical process theory, GATES provides finite-sample, model-agnostic global statistical guarantees for CATE estimates produced by arbitrary black-box machine learning models—without requiring parametric assumptions or resampling. It enables rigorous, threshold-agnostic identification of credible subgroups exhibiting clinically meaningful treatment effects. Empirically, GATES maintains nominal coverage even in small samples (n = 100), substantially improving the reliability of subgroup inference. Applied to a late-stage prostate cancer clinical trial, it robustly identifies a clinically significant “exceptional responder” subgroup. This work establishes a verifiable, statistically principled paradigm for causal subgroup discovery in precision medicine.

Addressing bias and noise in CATE estimation for subgroup identificationAvoiding modeling assumptions and intensive resampling proceduresProviding statistical guarantees for treatment effect subgroup selection

Latest Papers

What's happening recently
View more

This study addresses the challenge of optimizing data collection to enhance social welfare in policy learning when unobserved heterogeneity is present. Accounting for latent individual differences in policy responses, the authors propose a repeated-measurement design based on proxy variables for latent traits and derive minimax regret bounds for policy rules that either incorporate or omit these latent variables. The theoretical analysis uncovers a novel trade-off between policy class complexity and estimation accuracy, leading to an optimal data collection strategy that allocates resources efficiently between measurement precision and sample size. In a development economics application, incorporating a proxy for entrepreneurs’ managerial ability increases social welfare by 5% and reduces the probability of welfare loss by 50%.

data collectionpolicy learningtreatment effects

This work addresses the challenge of efficiently predicting the accuracy gains of Best-of-N inference without performing full-scale sampling. The authors propose a lightweight prediction method that leverages only three core statistical features derived from a single model pass on the validation set: prompt-level consistency distribution, position of the first correct sample, and variance in generation length, supplemented by entropy. A ridge regression-based predictor is constructed and theoretically grounded through Bootstrap-Lasso stability analysis and concentration bounds on linear approximation residuals. Evaluated across diverse models, training strategies, and tasks, the approach achieves a Spearman correlation of ρ=0.90 with actual Best-of-N performance gains, substantially reducing the computational cost associated with reward model evaluation.

best-of-Ninference-time scalinglanguage models

This study addresses the failure of inference for data-driven subgroup identification in within-sample evaluation due to selection bias—particularly when subgroup boundaries are non-smooth and depend on infinite-dimensional functionals. The authors propose a conditional adaptive perturbation method grounded in a triple robustness theoretical framework, which accommodates any machine learning algorithm, including black-box models, without requiring parametric assumptions or smoothness conditions on subgroup boundaries. The approach jointly optimizes subgroup identification and nuisance parameter estimation rates, enabling fully efficient, unbiased within-sample inference without data splitting. In a reanalysis of the ACTG 175 clinical trial, the method substantially improves estimation stability and statistical efficiency while avoiding the information loss inherent in conventional sample-splitting strategies.

data-dependent objectsin-sample evaluationnonregularity

This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.

generalizationi.i.d.optimization

Hot Scholars

ZB

Zahra Bami

Center for Biostatistics, Epidemiology, and Public Health, Department of Clinical and Biological Sci
BiostatisticsMLImage processingDeep Learning
PP

Paolo Penna

IOG
Game TheoryMachine LearningTheoretical Computer Science