experimental counterbalancing

Designing experimental protocols that remove or separate logically irrelevant factors (such as word choice, answer order, or shared tokens) from the phenomena of interest, and ensuring perceptible, unbiased measurement of system or user responses. This involves counterbalancing, randomization, and control strategies to avoid confounds in behavioral and system evaluations.

experimentalcounterbalancing

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Causal Inference in Counterbalanced Within-Subjects Designs

May 06, 2025
JH
Justin Ho
🏛️ Harvard University | University of California, Berkeley

Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.

Challenges in causal inference with counterbalanced within-subjects designsNeed for alternative methods to ensure valid causal inferenceUnverifiable assumptions about symmetric carryover effects in counterbalancing

This work addresses the prevalence of erroneous conclusions in scientific experimental design due to overlooked confounding variables and the absence of formal verification methods meeting the rigorous standards of programming language communities. It presents the first probability-free semantic characterization of the d-separation criterion in causal inference, establishing its equivalence to non-interference semantics from security theory. By integrating graph theory, formal semantics, and program analysis, the authors mechanize this result in the theorem prover Rocq, thereby formally verifying the semantic correctness of d-separation. This foundational contribution enables automated, falsifiable, and formally verifiable modeling of real-world systems, offering a principled basis for assessing the quality of experimental designs.

causalityconfounding variablesd-separation

On the distinction between the per-protocol effect and the effect of the treatment strategy

Aug 27, 2024
IJ
Issa J. Dahabreh
🏛️ Harvard T.H. Chan School of Public Health | Brown University School of Public Health

This paper addresses the causal distinction between the per-protocol effect (PPE) and the treatment strategy effect (TSE) in randomized trials. Using the potential outcomes framework and causal graph models, we rigorously establish that these two effects differ fundamentally in their causal definitions, identifiability conditions, and data requirements—particularly regarding dependence on assignment information—and are generally non-interchangeable. Even under complete randomization, identifying either effect necessitates explicit use of assignment mechanism information, challenging the conventional belief that randomization automatically eliminates confounding. We further derive necessary and sufficient conditions for PPE–TSE equality and formally refute the common practice in observational studies of substituting TSE for PPE—unless strong additional assumptions hold. These results provide a theoretical foundation and practical guidelines for causal interpretation and analysis of clinical trials.

Clarifying role of assignment in defining causal effects of interestDistinguishing per-protocol effects from treatment strategy effects in trialsIdentifying when assignment information is needed for causal estimation

Isolated Causal Effects of Natural Language

Oct 18, 2024
VL
Victoria Lin
🏛️ Carnegie Mellon University

This work addresses the modeling and quantification of *isolated causal effects* in natural language—i.e., the independent causal impact of a targeted linguistic intervention (e.g., factual errors) on reader cognition or behavior, while rigorously controlling for confounding influence from non-focal linguistic components. Method: We formally define “language-isolated causal effect” and propose a novel dual-axis evaluation framework grounded in omitted-variable bias theory: one axis measures the fidelity of non-focal language approximation; the other quantifies sensitivity of effect estimation to approximation error. The framework integrates causal inference, controllable language generation, and semi-synthetic data construction. Contribution/Results: Empirical validation on semi-synthetic and real-world datasets demonstrates that our framework accurately recovers ground-truth causal effects and quantitatively characterizes how modeling imperfections in non-focal language systematically bias causal estimates—establishing both theoretical foundations and practical tools for trustworthy causal analysis in NLP.

Addressing bias from poor non-focal language approximationEstimating causal effects of language changes on reader perceptionsValidating framework for isolated language effect measurement

This work challenges the conventional paradigm of eliminating all cognitive biases, instead investigating how large language models (LLMs) can *consciously leverage* cognitive biases to improve multiple-choice decision-making. Method: We propose the “bias-as-resource” perspective, designing a heuristic modulation strategy and a confidence-driven abstention mechanism; we further introduce BRU—the first bias-aware multiple-choice framework—and a balanced, expert-annotated evaluation dataset for bias analysis. Contribution/Results: Our approach achieves human–model reasoning alignment and bias-directed calibration, significantly improving accuracy, reducing error rates, and enhancing response efficiency—while preserving logical rigor. Crucially, this is the first work to formalize human cognitive biases as *controllable variables* rather than noise, establishing a novel LLM decision-optimization paradigm that jointly ensures practical utility and reliability.

Aligns LLM decisions with human reasoning using BRU datasetExamines cognitive biases in LLM decision-making for MCQsProposes heuristic moderation to reduce errors and improve accuracy

Latest Papers

What's happening recently
View more

Standard A/B testing introduces uninformative noise in regions of policy overlap, leading to high estimation variance and low statistical power. This work conceptualizes the random assignment mechanism as a meta-policy and proposes a novel Δ-off-policy estimation method that leverages the structure of policy overlap to eliminate irrelevant noise, thereby enabling unbiased estimation of the average treatment effect. Theoretically, under common support conditions, the proposed estimator strictly dominates the conventional mean-difference estimator. Empirical results demonstrate substantial reductions in variance and marked improvements in statistical power, highlighting its applicability to recommendation systems, information retrieval, and large language model interfaces.

A/B testingcounterfactual estimationpolicy overlap

This work reveals that in the LLM-as-a-judge framework, preference labels can inadvertently encode non-semantic behavioral biases, thereby compromising the reliability of alignment systems. Specifically, the study demonstrates for the first time that preference labels may function as a covert communication channel at a sub-semantic level, enabling a biased judge model to transmit unintended behavioral signals to a student model through binary preference feedback, which are then reinforced over multiple alignment rounds. Through controlled experiments pairing a neutral student model with a biased judge model, the authors systematically trace the transmission pathways and cumulative effects of these preference signals. The findings show that even when the student model generates semantically unbiased responses, it can still internalize and amplify the judge’s biases via implicit signals—challenging the conventional assumption that preference labels provide purely semantic supervision.

alignmentcovert communicationLLM-as-a-judge

This study addresses bias in causal inference arising from unmeasured confounding in observational studies by proposing a novel research design that integrates expert knowledge. Specifically, it actively identifies potential unmeasured confounders by querying clinical experts about differences in treatment intent between matched patient pairs. This approach systematically incorporates clinicians’ judgments on treatment intent into the confounding detection pipeline for the first time, establishing a theoretical foundation that transcends the limitations of methods relying solely on observed data. Combining propensity score matching, natural language processing of clinical notes, and a semi-synthetic validation framework, the method demonstrates significant unmeasured confounding in electronic health records within an ICU setting. Using clinical notes as proxies for physician knowledge, the authors validate the approach’s efficacy in an environment with known ground-truth causal effects.

causal inferenceconfounder detectionobservational study

This work addresses the challenge of selecting the optimal inference protocol—direct answering, voting, or debate—and achieving efficient routing under a fixed computational budget for open-source large language models. Under a unified generation-length constraint, the study systematically evaluates greedy decoding, three-sample voting, and two-agent critique-and-revise debate on MuSiQue and GSM8K. It proposes a dynamic routing strategy based on pre-inference signals, revealing that voting entropy predicts debate safety but not its necessity, with many beneficial debates occurring when votes are unanimous yet incorrect—highlighting the limitations of lightweight probing methods in identifying samples requiring debate. Experiments employ Llama-3.1-8B and Mistral-3-8B-Instruct models, combined with entropy thresholds, logistic regression, gradient-boosted trees, and self-critique probes. While ideal routing yields up to a 14-point gain, a simple entropy-threshold controller achieves 1.3–1.7 points, with learned methods offering no significant improvement over this baseline.

computational ceilingLLM routingreasoning protocols

This study investigates whether large language model (LLM)-based automated scoring systems are susceptible to construct-irrelevant factors such as spelling errors, textual redundancy, and off-topic content. To this end, a dual-architecture LLM scoring system was developed for automatically evaluating short-answer responses in situational judgment tests, and its robustness was systematically assessed through adversarial perturbation experiments. Results indicate that the system demonstrates strong stability against meaningless padding, spelling errors, and variations in stylistic complexity, while appropriately penalizing repetitive text and off-topic responses. Overall, its scoring performance surpasses that of conventional non-LLM approaches. This work provides the first systematic validation of how LLM-based scoring responds to diverse construct-irrelevant perturbations, revealing its distinctive scoring logic and advantages.

automated essay scoringconstruct-irrelevant factorshallucinations

Hot Scholars

MR

Marek Rutkowski

The University of Sydney
Mathematical FinanceStochastic Processes
AB

Arno Botha

Ph.D, University of Pretoria; North-West University
Credit risk modellingMachine learningMathematical financeRisk management
TV

Tanja Verster

Professor at Centre for BMI Research, Potchefstroom campus, NWU, South Africa
Credit ScoringStatisticsQuantitative Risk ManagementData Science
SB

Silvia Bartolucci

University College London
Complex networksmarket microstructureblockchain technologiesstatistical physics
SA

Sabrina Aufiero

University College London
Systemic RiskStatistical MechanicsMachine LearningComplex Networks