Score
Designs and analyzes experimental stimulus sets and assignment schemes that systematically balance and symmetrize logically irrelevant factors—such as wording, option order, and specific stimulus tokens—across participants or trials so that each combination of factor levels is represented. Builds crossed symmetrization layouts and counterbalancing/randomization procedures to isolate and estimate token, order, and verdict effects and to detect cross‑form incoherence.
Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.
When legal, ethical, or engineering constraints preclude user-level randomization, alternatives such as item-level randomization often yield biased causal estimates due to confounding and carryover effects. This paper proposes Regularized Balanced Switch Design (RBSD), the first balanced switch design family that jointly randomizes over time and items while ensuring theoretical unbiasedness and practical robustness. We rigorously characterize sufficient conditions for RBSD to mitigate carryover effects and derive an unbiased estimator within the potential outcomes framework. Extensive simulations and real-world e-commerce experiments demonstrate that RBSD significantly outperforms conventional item-level randomization and unbalanced switch designs: it reduces average estimation error by 32%–57% without introducing additional bias.
This paper addresses the dual challenges of covariate balance and causal effect estimation in observational studies. We propose Cross-Balancing—a method that decouples feature learning from weight estimation via sample splitting, enabling adaptive selection or learning of balancing features from high-dimensional candidate variables. The approach integrates covariate balancing, high-dimensional variable screening, and plug-in bias correction, achieving both multiple robustness and asymptotic efficiency. Theoretically, the proposed estimator is shown to be consistent, asymptotically normal, and semiparametrically efficient—attaining the semiparametric efficiency bound. Empirically, Cross-Balancing significantly improves estimation accuracy and inferential power while preserving statistical validity and model interpretability.
Visual generative models often exhibit unstable and error-prone responses to complex concepts, yet existing work lacks systematic analysis and effective mitigation strategies. To address this, we propose Online Concept Balancing (OCB), a training-time intervention that introduces a concept-level imbalance-aware loss (IMBA Loss) to dynamically calibrate response biases—without requiring offline data resampling or architectural modifications. Through causal-driven empirical analysis, we identify root causes of concept imbalance and construct Inert-CompBench, a novel benchmark, alongside a multi-source test suite for rigorous evaluation. On three public benchmarks, OCB significantly improves baseline models’ accuracy on complex concepts and compositional robustness, yielding an average gain of +12.7%. The method demonstrates strong generalization across architectures and datasets, and operates as a plug-and-play module with minimal implementation overhead.
To address inadequate covariate balance and low estimation precision of treatment effects in randomized controlled trials with high-dimensional covariates, this paper proposes a novel randomization method based on the Deville–Tillé cube method. By integrating covariate balance constraints with cube sampling, the approach achieves strong balance at the sample level, substantially improving estimation efficiency. We establish, for the first time, the asymptotic statistical theory for both population and sample average treatment effects under the cube method, rigorously deriving an upper bound on imbalance error of $O(sqrt{p/n})$; we prove that this bound dominates mainstream methods when $p/n o 0$. Monte Carlo simulations demonstrate that, under 100-dimensional covariates, the proposed method reduces the variance of treatment effect estimation by over 40%, markedly enhancing causal inference accuracy.
This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.
This study addresses the intersection of psychometrics and fairness research in artificial intelligence and machine learning (AI/ML), aiming to clarify the conceptual distinction between “equality” and “fairness” and to highlight the pivotal role of causal reasoning in fairness evaluation. By systematically mapping the full psychometric pipeline onto AI/ML fairness frameworks and integrating conceptual analysis with interdisciplinary comparison, the work develops a unified theoretical perspective. It not only elucidates key similarities and differences in how fairness is understood across these fields but also underscores the necessity of causal considerations for achieving genuine fairness. The resulting framework provides a clear conceptual foundation and theoretical direction for future interdisciplinary research on fairness.
This study investigates how alignment fine-tuning renders large language models susceptible to spurious input cues—such as flattery or fabricated examples—leading to inconsistent or erroneous responses. Through representation probing, cross-dataset transfer, and causal interventions across five model families, the authors systematically identify alignment-induced cue-based biases. Their analysis demonstrates that these biases predominantly arise during the alignment phase rather than pretraining and reveals that distinct biases occupy independent, separable subspaces within the model representations. Moreover, by inversely manipulating the direction of these bias subspaces, the study shows that biased errors can be substantially corrected without compromising the model’s ability to produce correct outputs, thereby confirming both the malleability and representational disentanglement of alignment-induced biases.
This study introduces, for the first time, the psychological conflict-task paradigm into the domain of large language models (LLMs) to investigate the underlying mechanisms driving congruency effects in verbal conflict tasks. By designing purely textual congruent and incongruent prompt conditions, the work reveals the competitive dynamics between a model’s default color associations and explicitly stated rules. Through causal attribution, attention analysis, ablation studies, and fine-tuning interventions, the research replicates robust congruency effects across Gemma-2-2B and multiple Pythia models. It further demonstrates that short-range and long-range attention mechanisms preferentially govern processing under congruent and incongruent conditions, respectively, thereby validating a competition between default weight-embedded mappings and contextually specified rule-based mappings.