Score
Designs and implements conditional normalizing-flow models that transform a simple base distribution into complex target conditionals given inputs, while structuring and sharing latent variables or flow parameters across multiple related channels or outputs. Builds and analyzes architectures and training procedures for efficient conditional generation and amortized posterior sampling that exploit this shared latent/parameter structure.
Existing conditional flow-based generative models employ generic unimodal noise priors, leading to unnecessarily long source-to-target mapping paths and inefficient sampling. To address this, we propose the **Conditional Centralized Prior (CCP)**: a conditional encoder maps textual or other prompts to modality-specific centers in data space, from which class-adaptive Gaussian priors are constructed—marking the first dynamic prior customization within the flow matching framework. CCP significantly accelerates training convergence and reduces sampling steps while achieving superior performance across FID, KID, and CLIP Score versus baselines, balancing generation quality and efficiency. Our core contribution lies in departing from conventional fixed-prior paradigms by integrating prior design into the conditional modeling process, thereby enhancing both the geometric plausibility and computational efficiency of flow-based generative models.
This work addresses the topological mismatch between standard normal latent variables and complex data distributions, which hinders the training efficiency and generative performance of normalizing flows. To mitigate this issue, the paper introduces, for the first time, a mixture of probabilistic principal component analyzers (MPPCA) as a learnable low-rank latent prior within the normalizing flow framework. This formulation effectively alleviates topological obstructions, simplifies the flow transformation architecture, and enables efficient initialization. The model is trained end-to-end by integrating the expectation-maximization (EM) algorithm with KL divergence minimization. Empirical evaluations on both tabular and image datasets demonstrate that the proposed approach significantly outperforms baseline methods, achieving faster convergence and superior sample quality.
Existing expert prior elicitation methods struggle to model complex dependency structures and flexibly specify joint distributions. Method: We propose the first end-to-end, nonparametric joint prior learning framework based on normalizing flows. It transforms expert heuristic judgments into a differentiable density estimation task, employs deep normalizing flows to capture high-dimensional nonlinear dependencies, and integrates simulation-based inference for likelihood-free prior calibration. Contribution/Results: This work is the first to systematically introduce normalizing flows into expert elicitation, unifying support for both parametric and nonparametric, as well as independent and joint prior modeling; it further introduces a multi-stage diagnostic evaluation pipeline. Four simulation experiments demonstrate substantial improvements in prior density fidelity and expert interpretability, establishing a more powerful and transparent paradigm for Bayesian prior learning.
Normalized flows (NFs) remain underexploited for density estimation and generative modeling due to architectural complexity and limited scalability. This paper proposes TarFlow—a scalable NF architecture built upon a direction-alternating autoregressive Transformer that directly models pixel-level distributions within image patches. To enhance robustness and sample quality, we introduce Gaussian noise injection during training, post-training denoising, and a unified conditional/unconditional guidance mechanism. TarFlow is the first single-flow model to significantly surpass prior state-of-the-art methods on standard image likelihood estimation benchmarks, while simultaneously achieving sample fidelity and diversity on par with diffusion models. The implementation is publicly available.
Normalized flow generative models suffer from interpolation paths deviating from the data manifold, primarily due to norm drift induced by Gaussian base distributions in latent space. To address this, we propose a norm-constrained base distribution reconstruction framework—introducing Dirichlet and von Mises–Fisher distributions into normalized flows for the first time. These distributions explicitly constrain latent variables to the unit simplex or unit hypersphere, respectively, ensuring geometrically consistent interpolation trajectories. Our method requires no architectural modifications to the flow network and provides an interpretable, unambiguous interpolation criterion, effectively overcoming interpolation distortion inherent to the Gaussian assumption. Experiments demonstrate consistent improvements over baselines across all major evaluation metrics: bits/dim, Fréchet Inception Distance (FID), and Kernel Inception Distance (KID). Interpolation quality is significantly enhanced while strictly preserving original generation performance.
Existing flow matching approaches struggle to jointly model forward generation and reverse classification of multivariate data, lacking consistency in conditional inference. This work proposes a Joint Flow Matching (JFM) framework that assigns symmetric roles to variables at temporal endpoints, thereby constructing a shared joint distribution such that forward and backward integrations naturally correspond to conditional forms of the same joint distribution. JFM is the first method to enable consistent bidirectional conditional inference within continuous normalizing flows, inherently supporting confidence calibration without post-processing and providing an interpretable foundation for discriminative–generative tasks. Experiments demonstrate that JFM achieves competitive classification accuracy on conditional datasets, generates samples highly consistent with the classifier, and yields natively calibrated confidence scores.
Existing stochastic process models often struggle to perform efficient conditional sampling under nonlinear observations, non-Gaussian likelihoods, or global constraints, typically requiring bespoke and complex algorithms. This work proposes the first universal conditioning framework that requires no training and avoids neural network approximations. The approach represents stochastic processes via deterministic mappings of tractable latent innovation variables, recasting conditional sampling as an inference problem in latent space. Exact solutions are achieved through backward-time stochastic differential equations (SDEs). The method accommodates a broad class of heterogeneous processes—including spatial priors, nonlinear dynamics, stochastic partial differential equations, extreme-value processes, and discrete-state processes—and enables high-fidelity conditional sampling on a single CPU within seconds, achieving both theoretical exactness and computational efficiency.
In conditional generative modeling, existing diffusion and flow-matching approaches map standard Gaussian noise to the conditional data distribution, tightly coupling conditional injection with optimal transport—resulting in high model complexity and inefficient training. To address this, we propose Conditional-Aware Reparameterization Flow (CAR-Flow), a flow-matching framework that introduces lightweight, learnable, condition-aware shifts—either to the source (noise) or target (data) distribution—to dynamically reposition them, thereby explicitly decoupling conditional injection from transport path learning. This reparameterization incurs negligible parameter overhead (+0.6%) while substantially shortening the required probability transport distance. On ImageNet-256, CAR-Flow reduces the FID of SiT-XL/2 from 2.07 to 1.68, demonstrating simultaneous improvements in both training efficiency and generation quality.
Existing deep multivariate models are typically tailored to specific tasks, resulting in limited generalization capability. This work proposes a universal modeling framework that parameterizes the conditional distribution of each variable given all others using deep neural networks and represents the joint distribution through a Markov chain kernel. The model is trained by maximizing the likelihood under the stationary distribution of this kernel. By design, the approach eliminates the need for task-specific architectural modifications and inherently supports arbitrary downstream tasks as well as diverse semi-supervised learning scenarios. Consequently, it not only enhances model generalization but also significantly improves the efficiency of leveraging unlabeled data.
This study addresses the challenge of preserving the full conditional distribution of predictors given a response variable in dimension reduction. To this end, it proposes a likelihood-based sufficient dimension reduction (SDR) framework that introduces conditional normalizing flows to the SDR literature for the first time. The method jointly learns a linear projection and a flexible conditional density by maximizing the conditional log-likelihood, employing monotonic rational quadratic spline flows to model complex conditional distributions. The approach is grounded in an interpretable mutual information objective and complemented by a neural Gaussian SDR variant as an auxiliary model. Theoretical analysis establishes Fisher consistency, and empirical evaluations across diverse simulation settings and the UTKFace age prediction task demonstrate accurate recovery of the central subspace, significantly outperforming existing SDR methods and neural Gaussian baselines.