Score
Using held-out data subsets to estimate generalization and validate discoveries (e.g., persona assignments or expert selection), often via probe-success rules or learned scorers to select best-performing models or profiles.
This work proposes a novel approach to integrate multi-source expert prior knowledge into best subset selection. Addressing the limitations of purely data-driven feature selection, the method embeds expert assessments of feature relevance—aggregated in the form of Poisson binomial distributions, pairwise win probabilities, or normalized average ranks—as log-odds penalty terms within a mixed-integer optimization (MIO) objective function via a maximum a posteriori (MAP) framework. This constitutes the first theoretically principled and analytically tractable Bayesian–MIO joint formulation, which naturally reduces to classical best subset regression in the absence of expert information. Theoretical derivations and algorithmic implementation have been completed, with empirical results forthcoming.
In regression and causal inference, controlled subgroup selection aims to identify subpopulations—within the covariate space—whose average response or treatment effect significantly exceeds a prespecified threshold, while providing rigorous statistical inference guarantees. Existing methods often sacrifice efficiency (e.g., via data splitting), lack flexibility, or fail to ensure valid inference. This paper introduces Chiseling: an interactive machine learning framework that iteratively refines candidate subgroups via targeted contraction. Crucially, contraction directions are determined solely from data outside the current subgroup, ensuring valid hypothesis testing under only finite-moment conditions. The framework seamlessly incorporates domain knowledge and accommodates arbitrary machine learning algorithms, supporting both randomized experiments and observational studies. Simulation and empirical analyses demonstrate that Chiseling achieves substantially improved subgroup detection power and practical utility over existing guaranteed methods—without compromising statistical validity.
This paper studies strategic hypothesis testing within a principal–agent framework: the agent holds private beliefs about product efficacy and may manipulate submitted data to maximize expected payoff; the principal must design a p-value threshold to balance Type I and Type II error risks. Methodologically, it innovatively integrates game-theoretic reasoning with classical statistical hypothesis testing by imposing incentive-compatibility constraints. The analysis establishes that the optimal p-value threshold exhibits a monotonic, analytically tractable structure in the agent’s strategic behavior, yielding a closed-form solution. Theoretically, it demonstrates that regulators can endogenously mitigate strategic reporting by calibrating the critical p-value, thereby unifying statistical robustness with incentive compatibility. Empirical validation using FDA drug approval data confirms the model’s predictive power, providing regulators with an interpretable and computationally tractable framework for optimizing approval policies.
Under data-driven selection, conventional prediction intervals fail to guarantee marginal coverage for the selected units—compromising reliability for focal samples. Method: We propose the first finite-sample exact coverage framework for post-selection inference, extending Mondrian conformal prediction to multiple test samples and non-equivariant models while accommodating arbitrary permutation-invariant selection rules. Our approach integrates conditional randomization tests, top-K or optimization-driven selection, conformal p-values, and preliminary screening prediction sets to enable efficient computation. Contribution/Results: Evaluated on drug discovery and health risk prediction tasks, our method substantially improves empirical coverage for focal units, ensuring statistically valid inference in real-world decision-making scenarios. This provides the first provably exact finite-sample coverage guarantee for post-selection prediction intervals under general selection mechanisms.
In causal subgroup identification, conventional methods suffer from high estimation noise in conditional average treatment effect (CATE) estimation and multiplicity issues arising from two-stage procedures. To address these challenges, this paper proposes the Global Adaptive Treatment Effect Sets (GATES) uniform confidence band method. Grounded in randomized trial design and empirical process theory, GATES provides finite-sample, model-agnostic global statistical guarantees for CATE estimates produced by arbitrary black-box machine learning models—without requiring parametric assumptions or resampling. It enables rigorous, threshold-agnostic identification of credible subgroups exhibiting clinically meaningful treatment effects. Empirically, GATES maintains nominal coverage even in small samples (n = 100), substantially improving the reliability of subgroup inference. Applied to a late-stage prostate cancer clinical trial, it robustly identifies a clinically significant “exceptional responder” subgroup. This work establishes a verifiable, statistically principled paradigm for causal subgroup discovery in precision medicine.
This study addresses the challenge of optimizing data collection to enhance social welfare in policy learning when unobserved heterogeneity is present. Accounting for latent individual differences in policy responses, the authors propose a repeated-measurement design based on proxy variables for latent traits and derive minimax regret bounds for policy rules that either incorporate or omit these latent variables. The theoretical analysis uncovers a novel trade-off between policy class complexity and estimation accuracy, leading to an optimal data collection strategy that allocates resources efficiently between measurement precision and sample size. In a development economics application, incorporating a proxy for entrepreneurs’ managerial ability increases social welfare by 5% and reduces the probability of welfare loss by 50%.
This work addresses the challenge of efficiently predicting the accuracy gains of Best-of-N inference without performing full-scale sampling. The authors propose a lightweight prediction method that leverages only three core statistical features derived from a single model pass on the validation set: prompt-level consistency distribution, position of the first correct sample, and variance in generation length, supplemented by entropy. A ridge regression-based predictor is constructed and theoretically grounded through Bootstrap-Lasso stability analysis and concentration bounds on linear approximation residuals. Evaluated across diverse models, training strategies, and tasks, the approach achieves a Spearman correlation of ρ=0.90 with actual Best-of-N performance gains, substantially reducing the computational cost associated with reward model evaluation.
This study addresses the failure of inference for data-driven subgroup identification in within-sample evaluation due to selection bias—particularly when subgroup boundaries are non-smooth and depend on infinite-dimensional functionals. The authors propose a conditional adaptive perturbation method grounded in a triple robustness theoretical framework, which accommodates any machine learning algorithm, including black-box models, without requiring parametric assumptions or smoothness conditions on subgroup boundaries. The approach jointly optimizes subgroup identification and nuisance parameter estimation rates, enabling fully efficient, unbiased within-sample inference without data splitting. In a reanalysis of the ACTG 175 clinical trial, the method substantially improves estimation stability and statistical efficiency while avoiding the information loss inherent in conventional sample-splitting strategies.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.