Score
Designs and executes controlled experiments that remove, mask, or modify model components, inputs, training conditions, or hyperparameters to estimate their causal and marginal contributions to model behavior and performance. Constructs measurement protocols (including preregistered and controlled designs), analyzes pre/post-ablation differences to rank components by importance, detect diminishing returns, quantify per-layer or per-feature effects, and validate robustness across implementations and settings.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
In large-scale aggregate-unit experiments (e.g., markets), conventional randomized treatment assignment often yields severe baseline imbalance due to extremely few treated units, leading to biased causal estimates. To address this, we systematically integrate the synthetic control method into experimental design, proposing a non-randomized treatment allocation mechanism: dynamically constructing a weighted synthetic control group based on pre-treatment covariates. We further develop配套 components—including counterfactual prediction, distance-driven unit matching, robust variance estimation, and a novel confidence interval construction procedure. Theoretically, our estimator is proven consistent and asymptotically normal. Empirically, it reduces estimation bias by 40–65% relative to standard randomization and substantially improves statistical power. Our core contribution is a new causal inference paradigm for small-N aggregate experiments—rigorous in inference, unbiased under mild assumptions, and highly interpretable.
In causal effect estimation, the absence of standardized hyperparameter tuning evaluation criteria impedes reliable model selection and creates a substantial gap between commonly used metrics and true performance. This paper systematically investigates the interplay between hyperparameter tuning and evaluation, jointly analyzing estimators (T-/X-/R-Learner), base learners (random forests, gradient boosting, neural networks), and evaluation metrics (IPW, DR, PEHE) across four benchmark datasets. Key findings are: (1) thorough hyperparameter tuning eliminates performance differences among mainstream causal estimators; (2) the choice of evaluation strategy exerts greater influence on final performance than either the estimator type or base learner architecture; and (3) existing evaluation metrics underestimate the performance gain from optimal model selection by over 35% on average. These results demonstrate that hyperparameter tuning is the primary determinant of causal estimation accuracy, underscoring an urgent need for more robust, theoretically grounded evaluation paradigms in causal machine learning.
This paper addresses the challenge of estimating causal effects when the target variable cannot be directly intervened upon and the underlying mechanism is complex—nonlinear, high-dimensional, and confounded. We propose the first active experimental design framework tailored for *indirect experiments*. Methodologically, we formulate a bilevel optimization model that integrates kernel-based estimation with adaptive sequential experimental design, yielding an analytically tractable and computationally efficient estimator for upper and lower bounds on the causal effect. Our key contributions are: (1) the first systematic formalization of feasibility conditions for indirect intervention under nonlinear confounding; and (2) dynamic narrowing of the causal bound gap to precisely localize the target query value. Extensive synthetic experiments across diverse settings demonstrate that our method significantly improves causal effect identification accuracy, with faster convergence of bound width compared to state-of-the-art baselines.
Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.
This study addresses the challenge of efficiently estimating causal effects under confounding when experimental budgets are limited. The authors propose a novel approach that integrates instrumental variable regression with Gaussian graphical models, leveraging prior knowledge of partial joint distributions to optimize the allocation between fully observed samples and partially observed data (e.g., only \(X_{12}\)). Under a fixed budget constraint, this method analytically derives the optimal sampling scheme that minimizes the asymptotic variance of the causal effect estimator—a solution not previously available in closed form. Theoretical analysis demonstrates that the proposed allocation significantly reduces both the total budget and the number of complete observations required to detect non-zero causal effects. Empirical validation in automotive analytics and drug discovery underscores the method’s practical utility alongside its theoretical contributions.
This study addresses the challenge of biased effect estimation in online controlled experiments caused by overlapping tests on shared traffic, which hinders accurate assessment of feature interactions. To resolve this, the authors propose Multi-Experiment Analysis (MEA), a method grounded in statistical modeling and causal inference that consistently estimates joint effects under arbitrary partial or full overlap and multi-variant settings—without requiring predefined factorial designs or constrained traffic allocation. MEA uniquely enables, without coordination overhead, the simultaneous modeling of bias-corrected individual effects, joint effects for any combination of variants, and conditional effects. Simulations confirm the estimator’s consistency and nominal confidence interval coverage, and the approach has been successfully deployed in large-scale production systems across multiple real-world business applications.
This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.
This study investigates whether the self-repair capability of frozen small code models in non-retrainable settings stems from repeated exposure to failed code or relies on external executable falsification feedback. To address this, we introduce a falsifiable methodology comprising feedback decomposition, content-controlled placebo design, matched-generation-budget control experiments, and executable auditing. We conduct large-scale evaluations on HumanEval+ and MBPP+ benchmarks using frozen models ranging from 0.5B to 1.5B parameters. Results show that blind resampling solves 18 more tasks than naive retrying; significant repair efficacy occurs only when feedback includes executable counterexamples, whereas pure instructions or content-irrelevant placebos yield no measurable improvement. These findings demonstrate that effective self-repair depends critically on external falsifying information rather than mere self-restatement.
This study addresses widespread misconceptions in the practical application of the synthetic control method, particularly concerning its reliance on covariates, claims of robustness, and prevailing model selection criteria—assertions often lacking empirical validation and potentially undermining causal inference reliability. Through rigorous theoretical analysis and extensive simulation experiments, the paper systematically evaluates these common misunderstandings and compares the performance of standard implementations against alternative approaches. The findings uncover critical pitfalls in current practices and, grounded in empirical evidence, offer concrete recommendations for more principled implementation and interpretation. By doing so, the work provides researchers with a practical guide to significantly enhance the quality and credibility of causal inferences derived from synthetic control methods.