Score
Designing and running well-powered, preregistered between-subjects online experiments and surveys with controlled stimuli and visualizations, including considerations of device effects and task difficulty, to estimate causal effects on outcomes like credibility, policy support, and workload.
This work addresses the challenge of accurately estimating the causal effect of content features on user engagement in personalized recommendation systems, where observational data are confounded by unmeasured factors and high-dimensional content spaces. The authors propose a novel method that leverages the randomness inherent in content display algorithms to identify causal effects using only a single non-displayed item per user log and the ratio of display probabilities between displayed and non-displayed items. Built upon counterfactual reasoning and importance weighting, the approach yields an unbiased estimator without requiring parametric assumptions about the content space or user preferences. The paper establishes theoretical guarantees for causal identifiability and introduces a new estimator enabling robust inference, thereby providing a reliable data-driven foundation for optimizing personalized content delivery.
In online two-sided markets, estimating the treatment effect on the treated (TATE) for item-side interventions is often biased due to interference among users. To address this, we propose Two-Sided Prioritized Ranking (TSPR), the first experimental design that leverages a recommendation system as a causal intervention vehicle: it dynamically adjusts item ranking priority in search results based on item-level treatment status, explicitly modeling and mitigating cross-user interference while preserving user access completeness and treatment consistency. Integrating causal inference, experimental design, and recommender modeling, TSPR is evaluated via simulation on real-world search impression data from an online travel platform. Results demonstrate that TSPR accurately recovers the true TATE and significantly reduces overestimation error compared to conventional randomization baselines. This work establishes a scalable, low-intrusion paradigm for causal experimentation on two-sided platforms.
In large-scale online platforms with hundreds of millions of users, conventional A/B testing is infeasible for post-hoc policy evaluation under sparse interventions, leading to severe estimation bias. To address this, we propose a two-stage debiased counterfactual estimation framework: (1) a covariate-based nearest-neighbor matching stage to construct high-fidelity control units and mitigate interpolation bias; and (2) a high-dimensional supervised learning stage—using XGBoost or neural networks—to model treatment effects while systematically diagnosing and correcting machine-learning-induced estimation bias. This is the first scalable solution enabling synthetic control methods to operate effectively under ultra-large-scale, sparse-intervention settings. Evaluated across six real-world online experiments, our method significantly improves causal effect estimation accuracy and reduces policy decision error rates by 42%. It has been deployed operationally to support closed-loop decision-making across multiple business units.
This study addresses the risk of inference bias and potential failure of the CUPED method in online A/B testing under complex experimental conditions. The authors systematically investigate five critical issues related to variance reduction with CUPED, and for the first time delineate its applicability boundaries in designs such as multi-arm experiments and two-stage sampling. To overcome these limitations, they propose a robust variance estimation approach tailored to such settings. Through rigorous theoretical analysis and large-scale empirical validation, the proposed method significantly improves inference accuracy. The solution has been successfully deployed in ByteDance’s experimentation platform, demonstrating its practical effectiveness and scalability.
In information provision experiments, standard two-stage least squares (TSLS) and panel estimators systematically underestimate the average partial effect (APE) because their weighting schemes assign higher weights to individuals with stronger first-stage belief updates—thereby over-downweighting weak updaters. This weighting bias arises from the implicit assumption of unweighted identification in conventional instrumental variable (IV) methods, which fails under heterogeneous belief updating. We propose a Bayesian belief-updating–based control function approach that achieves unbiased, unweighted APE identification without relying on update strength. By decoupling estimation from update intensity, our method corrects TSLS’s upward bias toward strong updaters. Applied to a gender wage gap beliefs experiment, our estimator yields an APE 40% larger than TSLS, substantially improving causal inference accuracy. Our key contribution is the first structural identification framework enabling consistent estimation of the unweighted APE.
This study addresses the bias in causal inference within two-sided networked markets, where user interactions violate the Stable Unit Treatment Value Assumption (SUTVA). To mitigate interference effects while maintaining high node coverage, the authors propose two network-aware clustering algorithms—EgoCluster V3 and MultiEgoCluster—that leverage iterative clustering and multi-center strategies. Building upon these clusters, they integrate graph structure to enable theory-driven bias correction for average treatment effect (ATE) estimation and enhance result generalizability. Compared to existing approaches, the proposed methods reduce spillover effects by a factor of three, increase effective sample size by approximately 38%, and double statistical power. The framework has been successfully deployed in LinkedIn’s production environment, supporting high-impact experimentation at scale.
This study addresses the lack of effective auxiliary tools in remote psychotherapy by introducing, for the first time, an AI-driven virtual supporter into real clinical sessions. Employing a two-stage approach, the research first analyzes the role of human supporters and then designs a dual-mode virtual agent capable of operating in both “everyday” and “therapeutic” modes. The system was integrated into the Zoom platform and evaluated empirically. Findings from a user study (N=14) demonstrate that the virtual supporter significantly reduces client anxiety, enhances emotional expression, and increases psychological safety—all without disrupting the therapeutic process—thereby validating its feasibility and innovative potential as a clinical support tool.
This study addresses a critical limitation in using large language models (LLMs) to simulate interventional experiments: such simulations inherently constitute observational studies and are thus susceptible to intervention-induced shifts in user attributes—termed “user drift”—stemming from the observational nature of training data, which biases causal effect estimation. The work formally characterizes this problem for the first time and proposes a novel approach that leverages negative control outcomes to diagnose user drift. To mitigate the resulting bias, it explicitly incorporates key confounders through role-based prompt engineering. Empirical evaluations in both survey-style and multi-turn dialogue settings demonstrate that the proposed method substantially reduces estimation bias and significantly enhances the reliability of causal inference in LLM-simulated experiments across diverse scenarios.
This study addresses the lack of systematic support for creating, managing, and deploying stimulus materials in visualization experiments—a gap that often leads to invalid results or wasted resources. Through semi-structured interviews with 19 visualization researchers, the work systematically examines practices and challenges across the full lifecycle of stimulus materials, from exploration and selection to deployment and analysis, integrating perspectives from user research and human factors engineering. The findings reveal, for the first time, a heavy reliance on manual processes and significant scalability limitations as core pain points in current workflows. Building on these insights, the study identifies key opportunities for improvement, including automated generation and intelligent validation of stimuli, thereby laying the groundwork for future directions such as AI-assisted stimulus design.