Score
Design and train models that encode complex treatments — including time-varying signals such as videos — into low-dimensional latent representations or embeddings that can be decoded or used to generate observed data while preserving causal content across time. Build generative latent-summary models and video-treatment embeddings that support reconstruction, counterfactual reasoning, and downstream causal estimation.
Deep generative models (DGMs) suffer from poor generalization and limited interpretability, while causal inference struggles to model high-dimensional, complex data-generating processes (DGPs). Method: This work proposes a bidirectional synergy paradigm integrating causality and DGMs. It introduces the first unified analytical framework unifying structural causal models (SCMs), latent-variable modeling, and variational inference—enabling DGMs to support causal discovery and leveraging causal constraints to enhance generative generalization. It further systematically investigates causal intervention and mechanism modeling in large language models (LLMs). Contributions: (1) Establishes three foundational research directions: causal-embedding generative models, generative causal identification, and LLM causality; (2) clarifies theoretical boundaries and distills ten open challenges; and (3) provides a systematic roadmap toward next-generation generative AI that is interpretable, generalizable, and amenable to causal intervention.
This study addresses causal inference over time-varying visual features in videos to assess their impact on viewer responses. The authors propose a novel approach that integrates deep generative models with longitudinal neural networks to nonparametrically identify and estimate potential outcome trajectories under dynamic stochastic interventions. Key contributions include framing video features as treatment variables within a causal inference framework for the first time, constructing the first video benchmark dataset with ground-truth causal effects, and enabling fine-grained analysis of how specific visual features at particular time points influence outcomes. The method is validated in the Super Mario Bros. environment and applied to 2020 U.S. presidential campaign advertisements, revealing that increased on-screen appearance probability of candidates significantly enhances viewer evaluations.
Deep generative models exhibit strong modeling capabilities but suffer from implicit, uninterpretable representations lacking causal semantics—limiting their applicability in high-stakes domains such as scientific discovery and fairness auditing. To address this, we propose a novel paradigm integrating causal representation learning (CRL) with deep generative modeling. Our approach systematically incorporates statistical identifiability theory and explicit causal structural constraints—including causal graphs and disentangled latent variables—to ensure both generative fidelity and semantic interpretability. Methodologically, we unify factor analysis, nonparametric deep generative architectures (e.g., VAE/GAN variants), and causal inference frameworks, rigorously deriving sufficient conditions for identifiable causal representations. Experiments demonstrate significant improvements in representation disentanglement, cross-domain transferability, and downstream causal reasoning tasks. This work establishes both theoretical foundations and practical guidelines for interpretable, causally grounded generative AI.
To address the low compression ratio in video generation caused by over-reliance on pixel-level reconstruction in embedding learning, this paper proposes a visually plausible reconstruction paradigm, establishing an encoder–generator joint framework that prioritizes semantic fidelity over exact pixel-wise reproduction. We innovatively introduce diffusion Transformers (DiTs) into latent-space decoding, design a lightweight latent-conditioning module, and enable end-to-end co-optimization of compression and generation. Our method achieves up to 32× temporal compression—8× higher than the state of the art—while preserving downstream text-to-video generation quality. It significantly reduces GPU memory consumption and training/inference overhead. Extensive experiments demonstrate that our paradigm consistently balances high compression ratios with high visual fidelity across multiple benchmarks, offering a novel pathway toward efficient video generation modeling.
This paper addresses core challenges in cross-domain adaptation of high-dimensional time-series data (e.g., videos): difficulty in transferring temporal dependencies, absence of direct causal edges among observed variables, and non-identifiability of high-dimensional causal structures. We propose a novel paradigm based on low-dimensional latent variable modeling to capture invariant causal mechanisms. First, we establish a theoretically grounded framework for uniquely identifying latent causal mechanisms. Second, we design a dual alignment mechanism—enforcing both intra-domain and inter-domain consistency—under sparsity constraints on latent variables, ensuring identifiability and domain invariance of the causal structure. Third, we integrate variational inference, latent causal discovery, and historical-information-driven sparse graph learning. Extensive experiments on eight benchmark datasets demonstrate significant improvements in cross-domain time-series classification and forecasting. The code is publicly available and validated on real-world scenarios.
Causal inference with high-dimensional unstructured textual covariates—e.g., clinical notes and patient feedback—remains challenging due to the difficulty of identifying valid treatment representations. Method: We propose a novel paradigm for causal effect identification that directly leverages frozen, intrinsic semantic representations (e.g., sentiment, topic) from large language models (LLMs), specifically Llama 3, as treatment features—bypassing data-driven learning of implicit causal representations and enabling perception-driven causal modeling and text reuse. Contribution/Results: We establish nonparametric identifiability of the average treatment effect (ATE) under this framework and prove asymptotic optimality under double machine learning. Experiments on synthetic and real-world datasets demonstrate significant improvements in estimation accuracy and computational efficiency, while robustly mitigating violations of the overlap assumption.
This work addresses the challenge that deep generative models are highly sensitive to unobserved confounders, which hinders accurate modeling of interventional distributions. To overcome this limitation, the authors propose CauVaDE, a method grounded in normalized augmented structural causal models (SCMs). CauVaDE compresses unobserved confounders into discrete latent variables with finite support, while continuous variations are captured by independent noise terms, implemented via a mixture variational autoencoder. Theoretically, the model class is shown to be dense—under both observational and interventional Wasserstein distances—with respect to augmented SCMs compatible with the underlying causal graph. Combined with entropy regularization, CauVaDE enables the generation of diverse counterfactual samples within feasible intervention regions. Empirical evaluations on image benchmarks demonstrate that CauVaDE produces high-quality interventional samples, achieving superior Fréchet Inception Distance (FID) scores compared to unconfounded reference models.
This study addresses the challenge of estimating dynamic causal effects from unstructured data—such as text or images—where existing methods fall short due to their treatment of such data as static observations. The authors propose the first statistical framework that integrates generative artificial intelligence (GenAI) with marginal structural models. By leveraging internal representations extracted from GenAI and jointly learning deconfounders for time-varying treatment features across sequences, the method enables asymptotically efficient estimation of dynamic causal effects. It further supports causal inference on temporal attributes of treatments, such as their position in a sequence, and provides valid confidence intervals. Simulations demonstrate the estimator’s accuracy and nominal coverage, while an analysis of real-world protest-related texts reveals pronounced sensitivity of treatment effects to their sequential position.
Existing methods struggle to generate temporally coherent, long-horizon dynamic 3D content under multimodal 3D representations while supporting topological changes. This work proposes MORPHOS, a novel framework that introduces, for the first time, a unified 4D implicit representation—termed Temporally Structured Latent Variables (T-SLAT)—to jointly model 3D Gaussians, meshes, and radiance fields. Dynamic geometry and appearance are generated frame-by-frame through an autoregressive causal attention mechanism. To mitigate error accumulation over time, the method incorporates a temporal structure enhancement strategy. Extensive experiments demonstrate that MORPHOS achieves state-of-the-art performance in appearance generation across multiple benchmarks, delivers accurate geometric reconstruction, and exhibits strong cross-representation generalization and robustness in long-sequence generation.
Existing causal video generation models rely solely on supervision from the current frame, which fails to preserve the long-term consistency of identity, layout, and motion, resulting in a representational planning gap. This work proposes a novel training strategy that introduces non-causal “foresight” supervision: a frozen non-causal encoder globally processes the full video rollout, and a lightweight predictor distills the stop-gradient foresight targets back into the causal state representations. Crucially, this approach enables future-frame supervision of causal representations without altering the inference architecture or violating causality constraints, thereby effectively bridging the planning gap. On VBench, it improves the overall 5-second generation score from 83.8 to 84.6; for 30-second ultra-long generation, subject and background consistency scores rise from 84.9 and 90.2 to 88.5 and 91.9, respectively.
This study addresses the challenges of jointly modeling instantaneous and lagged causal relationships and handling non-stationarity in time series data by proposing the iCReN framework. Leveraging non-stationarity, the method establishes identifiability conditions for latent states through contrastive learning combined with discrete or continuous auxiliary variables. This approach yields the first unified identifiability theory for both instantaneous and lagged causal structures, thereby overcoming key limitations of existing methods. Experimental results demonstrate that the proposed framework accurately recovers underlying causal structures on synthetic datasets and substantially enhances downstream predictive performance on real-world data.