๐ค AI Summary
To address the dual challenges of time-varying causal discovery and future-value forecasting in multivariate co-evolving data streams, this paper proposes ModePlaitโthe first streaming incremental learning framework that unifies causal discovery and prediction. Its core contributions are: (1) an adaptive dynamic mode transition detection mechanism that accurately captures nonstationary, time-varying causal structures; and (2) a synergistic integration of adaptive sliding windows with lightweight parameter updates, enabling efficient, scalable stream processing over unbounded data. Evaluated on both synthetic and real-world datasets, ModePlait outperforms state-of-the-art methods by improving causal structure identification accuracy by 12.6% and reducing multi-step prediction error by 9.3%, thereby achieving a superior balance among accuracy, computational efficiency, and dynamic adaptability.
๐ Abstract
Given an extensive, semi-infinite collection of multivariate coevolving data sequences (e.g., sensor/web activity streams) whose observations influence each other, how can we discover the time-changing cause-and-effect relationships in co-evolving data streams? How efficiently can we reveal dynamical patterns that allow us to forecast future values? In this paper, we present a novel streaming method, ModePlait, which is designed for modeling such causal relationships (i.e., time-evolving causality) in multivariate co-evolving data streams and forecasting their future values. The solution relies on characteristics of the causal relationships that evolve over time in accordance with the dynamic changes of exogenous variables. ModePlait has the following properties: (a) Effective: it discovers the time-evolving causality in multivariate co-evolving data streams by detecting the transitions of distinct dynamical patterns adaptively. (b) Accurate: it enables both the discovery of time-evolving causality and the forecasting of future values in a streaming fashion. (c) Scalable: our algorithm does not depend on data stream length and thus is applicable to very large sequences. Extensive experiments on both synthetic and real-world datasets demonstrate that our proposed model outperforms state-of-the-art methods in terms of discovering the time-evolving causality as well as forecasting.