Score
Design and implement algorithms and pipelines that synthesize plausible alternative motion or event sequences (counterfactual trajectories) conditioned on observed or hypothetical scenarios, including methods for generating counterfactual paths, counterfactual augmentation, and optimizing trajectories in latent space. Use these generated trajectories to create structured safety pairings, annotate failure modes, and derive corrective or risk‑mitigating actions for dataset augmentation, labeling, and downstream analysis.
Existing counterfactual models lack a unified graphical modeling framework for spatiotemporal multi-unit interactions, resulting in theoretical fragmentation and limited practical applicability. To address this, we propose the first graph-based causal counterfactual paradigm that jointly incorporates spatial dependence and temporal dynamics. Our method integrates the potential outcomes framework with structural causal models (SCMs) to establish a unified spatiotemporal graph causal inference framework. Technically, we innovatively couple graph neural networks (GNNs), SCMs, and spatiotemporal graph modeling techniques to explicitly characterize the spatiotemporal interaction mechanisms among multiple agents. This framework achieves theoretical unification of causal paradigms while substantially enhancing interpretability, generalizability, and computational traceability of counterfactual reasoning in complex scenarios. It provides a foundational tool for causal analysis in multi-agent systems, enabling principled, scalable, and transparent inference over dynamic, interdependent entities.
This work addresses the challenge of effectively modeling counterfactual outcomes following action interventions in interactive video world models. The authors propose a noise-coupled dual-branch unrolling mechanism that, building upon a shared state prefix and exogenous noise, bifurcates only the action stream after the intervention point to enable precise counterfactual generation. By explicitly recovering exogenous noise from self-generated trajectories, the method circumvents the traditional difficulties of approximate inversion and reformulates the principle of minimal change into a verifiable spatiotemporal locality metric. Grounded in Pearl’s causal framework, the approach integrates state branching, noise coupling, and computable causal descendant regions to construct a discriminator-free counterfactual evaluation system, thereby providing reinforcement learning with reliable reward signals.
Autonomous robotic manipulation struggles to recover from failures, hindered by the high cost of real-world data, the sim-to-real gap, and the absence of executable trajectory-level recovery strategies. This work proposes Dream2Fix, a novel framework that leverages a generative world model to synthesize counterfactual failure-recovery trajectory pairs directly from real-world successful demonstrations, eliminating reliance on simulators. To ensure physical plausibility, the synthesized data undergoes structured validation based on task validity, visual consistency, and kinematic safety. A vision-language model is then fine-tuned to enable end-to-end mapping from visual anomalies to precise recovery actions. Using a high-fidelity dataset comprising over 120,000 samples, the approach boosts real-robot recovery accuracy from 19.7% to 81.3%, achieving zero-shot closed-loop failure recovery for the first time.
Existing counterfactual explanation methods struggle to ensure temporal plausibility of time-series data (e.g., activity traces) under dynamic sequential constraints, often violating real-world domain knowledge. To address this, we propose the first counterfactual generation framework explicitly constrained by Linear Temporal Logic over finite traces (LTLf). Our method encodes LTLf specifications into optimizable constraints via Büchi automata and integrates process trace modeling with a constraint-driven genetic algorithm to perform temporally consistent counterfactual search. All generated explanations provably satisfy the given LTLf constraints—achieving 100% compliance—thereby significantly improving explanation validity and practical utility in temporal interpretability tasks such as process mining. The core contribution lies in the first deep integration of formal temporal logic into the counterfactual generation pipeline, effectively bridging formal verification and interpretable AI.
Existing methods for generating safety-critical driving scenarios struggle to balance realism and adversarial effectiveness due to their lack of explicit modeling of inter-agent interaction dependencies. This work proposes a closed-loop, generative bird’s-eye-view (BEV) world model grounded in structured counterfactual reasoning. By leveraging causal interaction graphs, the approach identifies causally critical agents and applies minimal interventions to steer risk propagation through naturalistic interactions. The method integrates causal adversarial agent identification, conflict-type classification, and phase-adaptive counterfactual guidance to overcome the realism–adversariality trade-off. Experiments demonstrate significant improvements: on nuScenes, the long-horizon collision rate increases from 12.3% to 22.7% while achieving superior trajectory realism (ADE of 1.88 vs. 2.09); on nuPlan, it attains state-of-the-art zero-shot realism.
Existing counterfactual explanation methods typically generate a single explanation path, neglecting path diversity and thereby limiting interpretability controllability and user agency. To address this, we propose the “Explanatory Multiverse” framework—the first to formally model the complete space of feasible counterfactual paths as a structured vector space. This enables geometric navigation and comparative analysis of path topologies (e.g., branching, divergence, convergence). We introduce *opportunity potential*, a unified scalar metric quantifying path-level properties—such as plausibility, sparsity, and effort—allowing users to select explanations interactively based on intrinsic path characteristics rather than endpoint differences alone. Our method integrates vector-space representation learning, geometric counterfactual path analysis, and graph neural networks. Evaluated across six tabular and image datasets, it significantly improves explanation diversity, controllability, and user-directedness over state-of-the-art baselines.