Score
Designs, builds, and evaluates methods, models, and evaluation procedures that produce and assess explanations grounded in causal relations rather than mere correlations. This includes developing causal reasoning and explainability algorithms, intervention-oriented explanation workflows, and metrics or analyses to differentiate, validate, and align model causal claims with established causal knowledge.
This paper addresses the practical challenge of deploying causal inference in real-world decision-making by proposing the first unified causal decision-making framework that integrates causal discovery, effect identification, and policy optimization. Methodologically, it systematically constructs an end-to-end technical pipeline encompassing causal structure learning (e.g., PC, GES), causal effect estimation (e.g., do-calculus, double machine learning), and causal policy learning (e.g., causal reinforcement learning, counterfactual evaluation), implemented in an open-source Python library—Causal-Decision-Making. Key contributions include: (1) establishing a reproducible, modular causal decision-making methodology; (2) providing a practitioner-oriented implementation framework with application guidelines; and (3) empirically validating effectiveness across diverse domains—including healthcare, economics, and recommender systems. This work bridges the gap between theoretical causal modeling and actionable decision support in real-world settings.
This work addresses the lack of a systematic approach to composing and ordering do-calculus rules, which hinders efficient exploration of the space of equivalent interventional queries. The paper introduces, for the first time, a derivation graph structure that formally captures the application and composition logic of do-calculus rules, systematically representing equivalence relations between observational and interventional probabilities under the do-calculus framework. Building upon this representation, the authors devise a streamlined identification procedure requiring at most four simplification steps. This approach not only reveals the intrinsic organizational structure underlying do-calculus reasoning but also enables the generation of multiple equivalent estimands for the same causal quantity, substantially improving estimation efficiency and facilitating practical applications of do-calculus.
Randomized controlled trials (RCTs) are often infeasible in software engineering, hindering rigorous causal assessment of tools, processes, or guidelines on development outcomes (e.g., efficiency, quality, user experience). Method: We propose a statistical causal inference methodology grounded in observational data, integrating the potential outcomes framework, propensity score matching, and difference-in-differences to systematically address confounding bias and selection bias. Contribution/Results: This work pioneers the systematic application of formal causal inference paradigms to requirements engineering and software practice research, tailoring analytical workflows and evaluation criteria to the characteristics of software engineering data. Empirical validation demonstrates that our approach substantially improves internal validity and reproducibility of causal conclusions in non-experimental settings. By enabling robust, evidence-based causal claims from real-world development data, it strengthens the empirical foundation for translating research findings into industrial practice.
Machine learning evaluation is frequently compromised by benchmark bias, data leakage, and undetected failure modes—undermining result reliability, reproducibility, and inferential validity. To address these issues, we propose a novel evaluation paradigm grounded in causal modeling: explicitly formalizing evaluative assumptions, constructing testable causal hypotheses, and guiding the design of robust benchmarks. Innovatively, we introduce Causal Abstraction Topologies (CATs)—a class of canonical structures from causal graphs—into large language model (LLM) reasoning assessment for the first time, enabling unified diagnosis of the “evaluation monster” problem. Through multiple empirical case studies, we demonstrate that our framework precisely delineates method applicability boundaries, clarifies ambiguous evaluation conclusions, and enables new evaluation pathways that are reproducible, interpretable, and extensible.
This study addresses the challenge of identifying causal graphs and estimating causal effects from observational data. We propose the first unified analytical framework that horizontally integrates major causal discovery paradigms—including constraint-based methods (e.g., PC), score-based methods (e.g., GES), functional causal models (e.g., LiNGAM, ANM, CAM, NOTEARS), and neural causal learning—while rigorously characterizing their identifiability conditions and practical applicability boundaries. Our contribution comprises: (1) a comprehensive knowledge graph covering 12 algorithmic families, 8 open-source toolkits, and applications across six domains (e.g., healthcare, economics, ecology); (2) standardized benchmark datasets, reproducible evaluation protocols, and practitioner-oriented guidelines; and (3) paradigm-level unification, formal identification boundary analysis, and an end-to-end resource ecosystem for real-world causal discovery deployment.
Causal Loop Diagram (CLD) construction in system dynamics suffers from low efficiency and high entry barriers for novices. Method: This paper proposes the first stepwise prompt engineering framework tailored for CLD generation, leveraging large language models (LLMs) to automatically map textual dynamic hypotheses into structured CLDs. The approach integrates chain-of-thought reasoning, role-guided prompting, and domain-specific constraints, representing CLDs as standard directed graphs; it is fine-tuned and evaluated on a textbook-based system dynamics dataset. Contribution/Results: Experiments show that the automatically generated CLDs achieve 89% agreement with expert-built diagrams on simple dynamic structures, substantially reducing modeling time. This work establishes the first end-to-end, accurate, interpretable, and domain-aligned natural-language-to-CLD generation pipeline, empirically validating the feasibility and practical utility of LLMs in automating system modeling.
This work addresses pervasive causal challenges in large language model (LLM) development and evaluation—such as shifts in data domains, annotator preference biases, and routing decision confounds—that undermine conventional predictive approaches due to unmeasured confounding and distributional shifts. To overcome these limitations, the paper introduces the first systematic causal inference framework spanning the entire LLM lifecycle, encompassing pretraining, alignment, routing, agent workflows, and evaluation. By integrating techniques from causal identification, counterfactual estimation, and interventional analysis, this framework replaces fragile purely predictive modeling with robust causal reasoning. The proposed approach substantially enhances the robustness and interpretability of LLM development, establishing a novel paradigm and a comprehensive toolkit for reliable, scientifically grounded foundation model research.
This work proposes CausalSE, a novel framework that systematically integrates structural causal models (SCMs) with propensity score matching to rigorously identify the true causal effects of interventions—such as prompt engineering—on large language model code generation performance. Addressing a critical limitation in traditional software engineering empirical studies, which often rely on statistical associations vulnerable to confounding bias, this study introduces Pearl’s causal inference paradigm into the field. Empirical evaluation on the Galeras dataset reveals that while conventional association-based analyses suggest complex prompts improve performance, causal analysis under CausalSE finds no significant treatment effect, thereby exposing false-positive conclusions arising from unaccounted confounders. The paper further provides a reproducible methodology for causal inference in software engineering contexts.
This work addresses the lack of a unified validation framework for evaluating whether high-level causal abstractions faithfully reflect underlying mechanisms. The authors construct a benchmark encompassing ten classes of complex systems—spanning discrete/continuous and static/dynamic types—and systematically assess over thirty metrics under a common causal abstraction framework to distinguish valid from invalid abstractions. They propose a novel continuous measure, Causal Abstraction Error (CAE), which passes discriminative tests across all systems and converges with only 30 interventions. Additionally, they introduce a fidelity test for unmapped variables that integrates observational, functional, information-theoretic, and causal criteria. Experiments demonstrate that causal metrics constrained solely by faithfulness reliably discriminate abstraction validity, with CAE exhibiting both superior performance and computational efficiency.
This work addresses the challenge that causal analysis methods, due to their conceptual complexity and limited validation on real-world data, remain difficult for domain experts to use effectively. To bridge this gap, we propose ORCA—the first end-to-end, interactive causal analysis collaborator designed specifically for non-expert users. ORCA employs a multi-agent architecture to jointly interpret user intent and supports the full causal analysis pipeline, including causal discovery, effect estimation, interpretability analysis, and root cause diagnosis, with adjustable levels of automation ranging from fully automatic to highly manual intervention. The system automatically generates structured reports, visualizations, and performance comparisons, substantially lowering the barrier to entry. Experimental evaluations across multiple real-world scenarios demonstrate that ORCA significantly enhances the efficiency, accuracy, and accessibility of causal analysis.