Score
Designing and applying offline and online evaluation protocols and metrics (including simulations, baseline comparisons, and A/B tests) to measure and validate uplift-model performance and to ensure robust platform-wide gains.
This study addresses the detrimental impact of structural biases—arising from selection bias, spillover effects, and unobserved confounding in real-world settings—on the estimation accuracy and evaluation reliability of causal uplift models. To systematically assess model robustness under controlled yet realistic conditions, the authors propose a benchmark framework based on semi-synthetic data that preserves authentic feature dependencies while introducing tunable structural biases. Their analysis reveals a fundamental distinction between targeting and prediction tasks, demonstrating that TARNet exhibits remarkable robustness across diverse bias scenarios. Crucially, they identify the mathematical alignment between evaluation metrics and the Average Treatment Effect (ATE) as a key factor underlying this stability. These findings establish more reliable principles for evaluating and selecting uplift models in practice.
This paper addresses the practical deployment challenge of heterogeneous treatment effects (HTE) in marketing. We propose a causal-driven constrained optimization framework: first estimating conditional average treatment effects (CATE) via uplift modeling, then jointly optimizing target audience selection and incentive policy design under hard constraints—including budget limits and sales decay thresholds—to maximize revenue and retention rate. We introduce the novel “causal uplift + constrained allocation” two-stage paradigm, ensuring both KPI alignment and customer experience preservation, while maintaining interpretability and engineering deployability. Offline evaluation employs uplift AUC, inverse propensity scoring (IPS), and self-normalized IPS (SNIPS). Large-scale online A/B tests demonstrate significant improvements over propensity score matching and static baselines across user retention targeting, campaign incentivization, and spending-threshold assignment—yielding higher revenue, improved task completion rates, and strict adherence to experience constraints.
Existing multivariate treatment effect modeling approaches typically extend binary-treatment frameworks or adopt feature adaptation strategies, yet exhibit poor robustness and substantial estimation bias under complex conditions such as noisy data and confounding between observation and intervention. This paper identifies their fundamental limitation: the failure to model functional dependencies among multiple interventions. To address this, we introduce—*for the first time in this domain*—function approximation theory into multivariate causal inference and propose the Orthogonal Function Adaptation (OFA) framework. OFA constructs an orthogonal basis function family over the intervention space, enabling unbiased and stable approximation of the underlying potential response function. Extensive experiments on synthetic and multi-source real-world datasets demonstrate that OFA consistently outperforms state-of-the-art adaptation methods, reducing average estimation error by 23.6%. Crucially, it maintains superior generalization performance under high noise and strong confounding—establishing a novel paradigm for multivariate intervention causal inference.
This study systematically evaluates whether large language models (LLMs) can substantially enhance the performance of biological novices in high-stakes dual-use computational biology tasks. Through multi-model, multi-benchmark human experiments, the authors compare novices’ accuracy on eight biosafety-related tasks with and without LLM assistance, benchmarking against both internet-assisted experts and standalone model performance. Results demonstrate that LLM support increases novices’ accuracy by a factor of 4.16 relative to unassisted controls and even surpasses internet-assisted experts on three tasks. Notably, 89.6% of participants readily accessed sensitive dual-use information, highlighting both the transformative potential and significant safety risks of human-LLM collaboration. This work provides the first quantitative assessment of LLMs’ real-world augmentation effect for non-experts in high-risk biological contexts.
Traditional A/B testing often suffers from insufficient positive correlation between estimators, leading to an overestimation of inferior algorithms and underestimation of superior ones, thereby increasing selection error rates. This work proposes a novel estimator that enhances selection accuracy by actively inducing positive correlation through a two-stage performance difference estimation on shared data, leveraging a hypothesized intermediate algorithm. The method uniquely integrates the advantages of offline evaluation into A/B testing, combining bias-variance analysis with the derivation of an optimal intermediate algorithm to substantially reduce critical selection errors. Empirical results on real-world datasets demonstrate that the proposed approach achieves comparable algorithm selection accuracy to existing methods using only half the sample size.
This study addresses the validity challenges that arise when evaluating human-AI collaboration in high-stakes decision-making, where conventional causal inference assumptions—particularly those underlying randomized controlled trials (RCTs)—often misalign with the dynamic nature of cutting-edge AI systems, thereby compromising internal, external, and construct validity. Through in-depth interviews with 16 domain experts spanning biosafety, cybersecurity, education, and labor, the research employs qualitative analysis and methodological mapping to systematically identify and structure key validity threats inherent in applying RCTs to frontier AI evaluation. The work further proposes practical mitigation strategies aligned with distinct phases of the AI development lifecycle, delineates the boundaries within which evidence of human enhancement remains valid, and offers actionable methodological guidance for AI governance, deployment, and safety assessment.
Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.
This study evaluates whether state-of-the-art language models substantially enhance non-expert users’ ability to plan high-consequence chemical, biological, radiological, or nuclear (CBRN) misuse compared to using publicly available tools alone. It introduces the Threshold Exceedance Criterion (TEC) framework, which decomposes CBRN capabilities into standardized, independently assessable components and distinguishes between generative and corrective model-assisted effects. Through controlled experiments, expert review, and structured scoring, the work provides the first systematic quantification of model-assisted capability gains. Empirical results confirm a substantiated increase in capability only in the radiological domain, with no significant risk elevation observed in the other CBRN areas. These findings directly inform model mitigation strategies and deployment governance decisions.
This study addresses the lack of standardized randomized controlled trial (RCT) frameworks in artificial intelligence evaluation, which has led to inconsistent research designs and limited reproducibility and comparability of results. Integrating Shadish’s four validity framework with TOP transparency guidelines, and drawing on RCT methodologies from clinical medicine, economics, psychology, and software engineering, this work proposes a structured RCT framework for AI that uniquely places human performance at the core of evaluation. It incorporates causal inference, heterogeneity analysis, and assessments of practical significance, while explicitly addressing AI-specific challenges such as model versioning, human–AI interaction, contamination effects, and fairness. The framework articulates five guiding principles and 33 actionable recommendations, supported by a tiered transparency mechanism that substantially enhances the rigor, reproducibility, and cross-study comparability of AI evaluation research, thereby establishing a foundational methodology for the field.