Score
Design and implement flow-based probabilistic models that learn time-varying cross-sectional (panel) distributions and continuous-time flow fields, including training procedures that pool information across irregular observation times. Build matching and inference components to map between marginals at different timepoints so you can generate synthetic longitudinal subjects and impute missing or sparsely observed entries.
Longitudinal data are often challenging to model due to irregular observation times, sparsity, and limited sample sizes. To address these issues, this work proposes the Panel Flow Matching (PFM) framework, which introduces continuous-time flow matching to longitudinal analysis for the first time, enabling a unified treatment of density estimation, data imputation, generation, and classification without requiring prior dimensionality reduction. PFM integrates forward flow matching with backward kernel fitting to effectively capture complex time-varying distributional structures. Experiments on both simulated data and a real-world vaginal microbiome pregnancy cohort demonstrate that PFM substantially outperforms existing methods, achieving higher accuracy in preterm birth prediction and revealing distinct dynamic distributional differences between case and control groups.
This paper addresses probabilistic forecasting for multivariate time series by proposing FlowTime, which frames future trajectory generation as efficient sampling from a conditional distribution. Methodologically, it is the first to integrate flow matching with autoregressive conditional density decomposition, leveraging learnable neural ODE flows for simulation-free, lightweight covariate conditioning—naturally supporting multimodality and uncertainty calibration. Key contributions include: (1) introducing autoregressive flow matching to eliminate reliance on iterative simulation; and (2) jointly optimizing extrapolation capability, parameter efficiency, and quantile calibration. Extensive experiments on multiple dynamical systems and real-world benchmarks demonstrate that FlowTime significantly outperforms state-of-the-art probabilistic forecasting models, achieving superior predictive accuracy and statistical consistency.
Existing diffusion-based time-series modeling approaches rely on fixed Gaussian priors, which inadequately capture temporal dynamic structures and thus limit generation quality. To address this, we propose TSFlow—the first conditional probabilistic forecasting model that integrates Conditional Gaussian Processes (CGPs) into the flow matching framework. TSFlow constructs data-dependent dynamic priors via optimal transport paths and introduces a novel conditional prior sampling mechanism, enabling the generative process to better reflect intrinsic temporal correlations and uncertainty. Crucially, it supports high-fidelity unconditional generation and flexible conditional forecasting without modifying the training objective. Evaluated on eight real-world datasets, TSFlow achieves significant improvements in unconditional generation metrics over state-of-the-art methods and attains either SOTA or leading performance across multiple probabilistic forecasting benchmarks.
Modeling irregular, high-dimensional longitudinal trajectories with sparse temporal sampling remains challenging. Method: This paper proposes Interpolative Multi-Marginal Flow Matching (IMMFM), which constructs a continuous-time flow matching objective via piecewise quadratic interpolation paths—departing from conventional pairwise transformation modeling. IMMFM jointly learns the drift term and a data-driven diffusion coefficient under multi-timepoint consistency constraints, enabling robust continuous stochastic dynamic modeling. Theoretical analysis ensures training stability and supports personalized trajectory generation. Contribution/Results: Evaluated on both synthetic benchmarks and real-world neuroimaging longitudinal datasets, IMMFM achieves significant improvements over state-of-the-art methods in prediction accuracy and downstream task performance (e.g., disease progression modeling and biomarker estimation), demonstrating superior capability in capturing complex, subject-specific temporal dynamics under sparse sampling conditions.
Long-term forecasting of irregularly spaced temporal event sequences remains challenging for existing autoregressive models, which suffer from error accumulation and myopic prediction horizons. To address this, we propose EventFlow, the first non-autoregressive generative model for continuous-time event modeling that introduces flow matching to directly learn the joint distribution over event times—enabling likelihood-free, end-to-end joint modeling. EventFlow parameterizes the velocity field via neural ordinary differential equations (neural ODEs), integrating event-time embeddings and temporal encodings to support efficient sampling, stable training, and both unconditional and conditional generation. On multiple standard benchmarks, EventFlow achieves predictive accuracy competitive with or superior to state-of-the-art autoregressive methods (e.g., THP, DyRep), while accelerating sampling by 3–5× and demonstrating enhanced training robustness.
This study addresses the challenge of irregular observations in electronic health records (EHRs), which complicates continuous patient state modeling and history-dependent prediction. To overcome this, we propose EHRFlow, a framework that extends multi-marginal flow matching to a conditional generative paradigm for the first time. By integrating Transformer-based medical history encoding, EHRFlow enables future dynamics to explicitly depend on prior clinical trajectories, facilitating continuous state prediction at arbitrary time points. Evaluated on million-scale real-world datasets, our approach significantly improves Top-5 code prediction accuracy. Furthermore, through counterfactual simulation, it effectively estimates intervention effects, providing a reliable tool for precision medicine.
This study addresses the challenge of statistical inference for multimodal geometric distributions in complex biological systems, particularly when density functions are defined only up to a normalizing constant. Building upon variational inference, this work integrates black-box variational inference, multiple importance sampling ELBO, and flow matching techniques to construct a highly expressive inference framework tailored for spatial transcriptomics. Key contributions include proposing CoLN, an unnormalized target density; revealing and refuting the conventional belief regarding performance gains of mixture models in variational inference; and developing a hybrid approach that combines multi-marginal flow matching with variational interpolation. Ultimately, this framework substantially enhances both the efficiency of approximate inference and the analytical capability for multimodal biological data, such as three-dimensional spatial transcriptomics.
This study addresses the challenge of cross-subject variability in time-series domain adaptation by tackling distribution transport between Markov process trajectories. To this end, it proposes a conditional flow matching-based algorithm that learns a source-to-target trajectory transport map while preserving the underlying Markov structure. Theoretically, the authors derive finite-sample error bounds, establish consistency under mixing-time assumptions, and determine an irreducible lower bound on sample complexity. Empirically, the method demonstrates its effectiveness on the THINGS-EEG2 image retrieval task, improving Top-5 accuracy from 43.87% to 49.05%, representing an 11.82% relative gain. These results validate the efficacy of structure-preserving trajectory transport for mitigating inter-subject distribution shifts in temporal domain adaptation.
This work addresses the challenge of modeling irregular multivariate time series with asynchronous observations and non-uniform sampling. Existing approaches either lose continuous-time semantics through discretization or incur high computational costs via ordinary differential equation solvers. To overcome these limitations, the authors propose WrapFlow, a novel framework that introduces continuous-time tokenization and a gap-aware mechanism to directly encode raw events while explicitly modeling unobserved intervals. WrapFlow leverages a standard Transformer to capture long-range dependencies and employs a residual flow matching training paradigm that avoids numerical simulation, enabling efficient continuous-time prediction. Evaluated on multiple real-world datasets, WrapFlow achieves state-of-the-art performance, generating high-quality continuous forecasts with only a small fixed number of rollback steps.
This study addresses the challenges of autoregressive exposure bias and inefficient diffusion-based inference in multivariate time series forecasting by proposing a non-autoregressive framework that integrates vector quantization (VQ) with prototype prior flow matching. The core innovation lies in pioneering the use of VQ codebooks as structured priors for rectified flow matching, replacing generic Gaussian noise, while leveraging a Diffusion Transformer (DiT) architecture for modeling. This design effectively eliminates the training-inference mismatch and accelerates convergence. Experimental results demonstrate that the proposed method achieves superior predictive accuracy and highly efficient inference across mainstream benchmark datasets.