Score
Designs and implements methods and pipelines to detect, model, and remove systematic variation introduced by processing batches or experimental groups in datasets, including covariate-based correction, normalization/background-adjustment algorithms, and data-alignment procedures. Builds and evaluates these corrections by measuring residual batch signal and their effects on downstream analyses and retrieval/alignment performance.
This study addresses the challenge in genomic diagnostics where batch effects are often confounded with biological signals such as copy number variations (CNVs), leading to false positives or missed detections—particularly when batch labels are unavailable. The authors propose a novel Bayesian approach that requires no prior batch information and instead leverages model evidence to cluster samples: technical artifacts reduce model evidence, whereas genuine biological variation does not. Heterogeneity is identified via a likelihood ratio test in evidence space, calibrated using a parametric bootstrap procedure. This work represents the first application of model evidence to disentangle technical from biological signals. The method demonstrates superior clustering accuracy over correlation- and dimensionality-reduction-based approaches on synthetic data, three clinical targeted sequencing panels (liquid biopsy, BRCA, and thalassemia), and mouse electrophysiology data, while maintaining strict control over false positive rates.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
This work addresses the problem of multiple hypothesis testing for multivariate Gaussian means under arbitrary covariance dependence structures. The authors propose a novel approach that integrates maximum residual descent (MRD) with a multi-stage calibration scheme. By introducing a new representation of residual statistics based on a single active precision matrix, the method achieves covariance-adaptive residualization, substantially reducing computational complexity. It replaces model-dependent thresholds with a simple multi-stage calibration rule, combining generalized stepwise critical values and precision matrix reconstruction techniques. The proposed procedure significantly lowers the normalized misclassification risk across diverse dependence structures, achieving error discovery rate control close to the nominal level, extremely low missed detection rates, near-perfect statistical power, and accurate estimation of the number of true signals.
Traditional stratified randomized experiments fail to effectively leverage covariate information predictive of potential outcomes, resulting in suboptimal statistical efficiency. To address this, we propose an adaptive stratification design that integrates the strengths of stratification and regression adjustment: it employs batch-adaptive stratification, cross-batch rematching, and a stratified estimator to enable post-design, nonparametric covariate adjustment. This approach breaks the conventional paradigm that treats stratification and adjustment as mutually exclusive—achieving robustness to model misspecification while preserving randomization guarantees. Through simulations on synthetic data and real-world political science experiments, our method demonstrates substantial improvements in estimation accuracy and statistical efficiency, particularly when covariates exhibit strong predictive power for outcomes.
This study addresses the bias in Cox proportional hazards model estimates arising from measurement error in AI-extracted covariates, a setting where downstream users only have access to the extracted data and limited calibration summary statistics. Within a multivariate calibration framework, the work provides the first decomposition of Cox model bias into a dominant, calibratable component and higher-order residual terms. Building on this insight, the authors propose a post-processing correction method that relies solely on calibration summary statistics and can be directly applied to outputs from standard Cox regression software. The approach is accompanied by uncertainty-adjusted confidence intervals and sensitivity diagnostic tools. Empirical evaluations on synthetic data demonstrate substantial bias reduction, with near-nominal coverage maintained even under mild violations of the linear calibration assumption. The paper also recommends a minimal set of calibration statistics that data providers should report to enable effective bias correction.
Existing contamination detection methods rely on assumptions such as access to training data, handcrafted statistics, or predefined labels, limiting their applicability in real-world settings. This work proposes a novel approach that dispenses with such assumptions by constructing depth profiles via linear probes in residual streams, introducing a metric termed “excess separability,” and combining label permutation tests, item-wise bootstrapping, and size-matched placebo control sets to detect whether a model has been exposed to test data while controlling for confounding factors. The method achieves a substantially reduced false positive rate—down to 0.02—rejects fragile ablations, and demonstrates effectiveness on real Transformer models: it finds no evidence of contamination across four Pile subsets. All implementation and auditing code is publicly released.
This work addresses the challenge of accurately estimating performance differences and their associated uncertainty when comparing randomly trained models, a task for which conventional methods are often inefficient. The authors propose a novel approach that leverages training logs as covariates to design model-specific covariate adjustment strategies. This method substantially reduces estimation uncertainty while preserving the original mean performance difference by performing posterior adjustment that effectively incorporates dynamic information from the training process. Crucially, it avoids introducing additional noise through careful covariate selection. Empirical evaluation across three model architectures and three datasets demonstrates that judicious use of early-stage training logs significantly enhances the statistical reliability of performance comparisons.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.
This work addresses the degradation of model generalization in supervised learning caused by heterogeneity in training data. To mitigate this issue, the authors propose an input-space adaptive partitioning method grounded in the intrinsic heterogeneity of the data. By introducing a variance-based metric that quantifies the inconsistency in pairwise sample influence, they demonstrate that this variance is maximized under mixture distributions. Leveraging this property, the method automatically partitions the data into homogeneous subsets without requiring prior knowledge, enabling independent training of submodels on each subset. Experiments on EMNIST and synthetic datasets show significant improvements in test accuracy, confirming that the proposed variance metric effectively captures data heterogeneity and offers a novel pathway to enhance model generalization.