Score
Design and build algorithms, pipelines, and simulation procedures that synthesize counterfactual datasets and scenarios—including feasible intervention-based samples, paired counterfactuals, hard-negative examples, density-guided or density-aware samples, and simulation/DSGE-derived trajectories—to create off-path state coverage and fill missing regions of a data distribution. These systems generate counterfactual training data and placeholders, produce paired views and negatives, and are used to analyze or realign model decision-token read-outs and intermediate belief–answer agreement for testing, training augmentation, and diagnostic evaluation.
This work proposes a novel counterfactual training mechanism that integrates the objective of counterfactual generation directly into the model training process, addressing the limitations of existing methods which often produce explanations lacking data plausibility and feature manipulability—key requirements for supporting real-world decision-making. By incorporating priors on feature mutability, modeling the underlying data distribution, and employing constrained optimization, the approach guides the model to learn representations that yield both plausible and actionable counterfactual explanations. Experimental results demonstrate that the proposed method not only generates high-quality counterfactuals but also significantly enhances the model’s adversarial robustness, thereby offering dual benefits in interpretability and reliability.
Counterfactual prediction under evolving intervention policies or hypothetical decision scenarios remains challenging due to unobservable potential outcomes, hindering model identifiability, evaluation, and generalization. Method: We propose the first systematic theoretical framework addressing this challenge—comprising (i) identifiability conditions for counterfactual prediction models, (ii) a performance evaluation system targeting loss, AUC, and calibration, and (iii) robust hyperparameter selection under model misspecification. Our approach integrates causal inference principles, doubly robust estimation, and loss-driven evaluation metric design. Contribution/Results: Validated via simulation studies and a real-world clinical application—cardiovascular risk prediction in statin-naïve populations—the framework significantly improves out-of-distribution generalization and clinical decision reliability in counterfactual settings.
This work formally defines and systematically investigates data poisoning attacks against counterfactual explanations (CEs), revealing their mechanisms for significantly inflating algorithmic redress costs at instance-, subgroup-, and population-levels. We propose theoretically grounded, provably correct poisoning strategies and conduct adversarial injection experiments on state-of-the-art CE generators—including DiCE and CFProto—complemented by a multi-level cost measurement framework. Results demonstrate that current SOTA CE methods exhibit systemic vulnerabilities: poisoning induces substantial increases in counterfactual path length or complete failure of redress generation. Our contribution is twofold: (1) it provides a critical security alert regarding the robustness of explainable AI systems; and (2) it establishes the first principled analytical paradigm for data poisoning targeting counterfactual explanations—thereby advancing rigorous robustness evaluation and defense research for trustworthy AI.
This work addresses the problem of counterfactual image generation methods producing edits that violate intrinsic causal logic within images. To this end, we propose the first unified benchmark framework for systematically evaluating causal consistency and visual fidelity—without requiring ground-truth labels. Methodologically, we integrate structural causal models (SCMs) with hierarchical variational autoencoders (Hierarchical VAEs), establishing a multi-model, multi-dataset, and multi-causal-graph evaluation paradigm. We further introduce customized metrics, including causal consistency, to quantify alignment with underlying causal mechanisms. Experimental results demonstrate that Hierarchical VAEs significantly outperform GAN- and flow-based baselines on both natural and medical imaging domains, highlighting their generalizability across modalities. The framework is released as an open-source, extensible Python benchmark package, enabling community-wide reproducibility, validation, and extension.
Existing causal diagnostic models struggle with effective local deployment due to the absence of interpretable reasoning pathways, scarcity of real-world critical-case samples, and poor cross-scenario transferability. To address these challenges, this work proposes a counterfactual data flywheel mechanism that leverages family-level physical or clinical rules to automatically generate high-quality training scenarios. By iteratively identifying reasoning failure points in a student model through staged analysis and applying bounded repair combined with stage-wise localized reinforcement learning, the approach optimizes the diagnostic reasoning chain. Remarkably, it achieves high-fidelity causal inference using only synthetic data—without requiring expert annotations—and establishes a closed-loop synergy between data generation and model refinement. Evaluated on industrial system monitoring and medical diagnosis tasks, the method improves strict-path accuracy by 11.6 and 5.5 percentage points, respectively, substantially outperforming the strongest baselines and proprietary models.
This study addresses the challenge of generating counterfactual explanations from real-world inputs that often contain missing values—a scenario underexplored in existing literature. The work presents the first systematic investigation into how input incompleteness affects counterfactual generation, evaluating multiple state-of-the-art algorithms under simulated missingness and comparing the performance of robust versus non-robust approaches. Results reveal that while robust methods exhibit marginally better validity, all current techniques struggle to consistently produce counterfactuals that are both valid and plausible. These findings underscore a critical limitation in the current state of counterfactual explanation methods and highlight the urgent need for novel approaches specifically designed to handle incomplete input data.
This work addresses the limitations of existing counterfactual distribution estimation methods, which often ignore the intrinsic connections between observed and counterfactual distributions, leading to substantial bias and poor generation quality. To overcome these issues, the authors propose a deconfounded flow matching framework that explicitly models the tight relationships in support sets, tail behaviors, and confounding-invariant features between the two distributions. The key innovations include a semiparametrically efficient estimator based on influence function correction and the first application of minimum energy flows to high-dimensional counterfactual modeling, which simplifies the flow objective and enhances training stability. Experimental results demonstrate that the proposed method significantly outperforms current debiasing approaches and effectively mitigates the failure modes commonly observed in high-dimensional flow-based counterfactual generators.