flow-based gradient estimation

Designs and implements learned proposal distributions using normalizing flows that produce tractable likelihoods and importance weights and can be conditioned on current observations for amortized sampling. Builds and trains end-to-end flow-based gradient estimators that provide low-variance, importance-weighted gradient estimates for models with non-conjugate factors and for use inside message-passing or stochastic optimization routines.

flow-basedgradientestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Importance-Weighted Non-IID Sampling for Flow Matching Models

Nov 21, 2025
XL
Xinshuang Liu
🏛️ University of California, San Diego

Under limited sampling budgets, expectation estimation in flow matching models suffers from high variance due to rare, high-impact events under independent sampling. This work proposes an unbiased importance-weighted non-i.i.d. sampling framework—the first to integrate importance weighting into the flow matching generative process. Our method learns a residual velocity field guided by the score function, jointly reconstructing the target marginal distribution and estimating sample importance weights via diversity regularization. A score-based regularization term further enforces moderate separation of samples in high-density regions, mitigating off-manifold drift. Experiments demonstrate that the approach preserves estimator unbiasedness while significantly improving sample diversity and quality. Consequently, it yields more accurate and robust expectation estimates, enhancing both interpretability and reliability of flow matching model outputs.

Balancing diversity and quality in non-IID sampling while maintaining unbiased estimationEstimating expectations of flow-matching model outputs under limited sampling budgetsReducing high variance in estimates caused by rare high-impact outcomes

Expert-elicitation method for non-parametric joint priors using normalizing flows

Nov 24, 2024
FB
F. Bockting
🏛️ TU Dortmund University | Rensselaer Polytechnic Institute

Existing expert prior elicitation methods struggle to model complex dependency structures and flexibly specify joint distributions. Method: We propose the first end-to-end, nonparametric joint prior learning framework based on normalizing flows. It transforms expert heuristic judgments into a differentiable density estimation task, employs deep normalizing flows to capture high-dimensional nonlinear dependencies, and integrates simulation-based inference for likelihood-free prior calibration. Contribution/Results: This work is the first to systematically introduce normalizing flows into expert elicitation, unifying support for both parametric and nonparametric, as well as independent and joint prior modeling; it further introduces a multi-stage diagnostic evaluation pipeline. Four simulation experiments demonstrate substantial improvements in prior density fidelity and expert interpretability, establishing a more powerful and transparent paradigm for Bayesian prior learning.

Develop expert-elicitation method for non-parametric joint priorsEvaluate method via simulations and diagnostic pipelineUse normalizing flows to model complex prior distributions

Importance Corrected Neural JKO Sampling

Jul 29, 2024
JH
Johannes Hertrich
🏛️ University College London | Freie Universität Berlin

This work addresses efficient sampling from unnormalized probability density functions. Methodologically, it introduces a novel framework integrating continuous normalizing flows (CNFs) with importance-weighted rejection resampling. The CNF training is formulated as an iterative JKO variational scheme, with theoretical guarantees of convergence to the Wasserstein gradient flow velocity field. Crucially, the framework proposes an alternating mechanism between local flow steps and non-local rejection steps, where rejection proposals are generated adaptively by the model itself—eliminating reliance on hand-crafted proposal distributions and mitigating the susceptibility of conventional Wasserstein gradient flows to local minima and slow convergence in multimodal settings. The approach yields i.i.d. samples and supports implicit density evaluation. Experiments demonstrate substantial improvements in accuracy over state-of-the-art methods on high-dimensional multimodal benchmarks, confirming both effectiveness and robustness.

Overcoming slow WGF convergence for multimodal distributionsReducing reverse KL loss while generating iid samplesSampling from unnormalized densities using CNFs and rejection-resampling

Normalized flows (NFs) remain underexploited for density estimation and generative modeling due to architectural complexity and limited scalability. This paper proposes TarFlow—a scalable NF architecture built upon a direction-alternating autoregressive Transformer that directly models pixel-level distributions within image patches. To enhance robustness and sample quality, we introduce Gaussian noise injection during training, post-training denoising, and a unified conditional/unconditional guidance mechanism. TarFlow is the first single-flow model to significantly surpass prior state-of-the-art methods on standard image likelihood estimation benchmarks, while simultaneously achieving sample fidelity and diversity on par with diffusion models. The implementation is publicly available.

Achieving state-of-the-art results with Transformer-based NF architectureEnhancing Normalizing Flows for better generative modelingImproving sample quality in likelihood-based image generation

Principled Interpolation in Normalizing Flows

Oct 22, 2020
SG
Samuel G. Fadel
🏛️ University of Campinas | Leuphana University | Norwegian University of Science and Technology

Normalized flow generative models suffer from interpolation paths deviating from the data manifold, primarily due to norm drift induced by Gaussian base distributions in latent space. To address this, we propose a norm-constrained base distribution reconstruction framework—introducing Dirichlet and von Mises–Fisher distributions into normalized flows for the first time. These distributions explicitly constrain latent variables to the unit simplex or unit hypersphere, respectively, ensuring geometrically consistent interpolation trajectories. Our method requires no architectural modifications to the flow network and provides an interpretable, unambiguous interpolation criterion, effectively overcoming interpolation distortion inherent to the Gaussian assumption. Experiments demonstrate consistent improvements over baselines across all major evaluation metrics: bits/dim, Fréchet Inception Distance (FID), and Kernel Inception Distance (KID). Interpolation quality is significantly enhanced while strictly preserving original generation performance.

Addressing side effects of linear interpolation pathsEnabling principled interpolation through base distribution changesImproving interpolation in normalizing flow generative models

Latest Papers

What's happening recently
View more

This work proposes a novel approach that integrates normalizing flows with stratified sampling to estimate expectations without relying on restrictive (semi-)parametric distributional assumptions, such as Gaussian or Gaussian mixture models, which can introduce substantial bias when misspecified. By leveraging the expressive power of neural networks, the method flexibly captures complex, unknown data distributions, thereby overcoming the limitations of traditional parametric frameworks. Empirical evaluations demonstrate that the proposed estimator significantly reduces Monte Carlo uncertainty in high-dimensional settings—specifically in 30- and 128-dimensional problems—and achieves marked improvements in both accuracy and stability compared to conventional Monte Carlo estimators and Gaussian mixture model-based approaches.

estimation uncertaintyexpectation estimationnonparametric distribution

This work addresses the challenge of efficiently and accurately estimating the divergence of the probability flow ordinary differential equation (PF-ODE) in diffusion and flow-based generative models, where existing approaches are either computationally expensive or suffer from high variance. The authors propose StAD, a novel method that, for the first time, incorporates the Langevin–Stein operator into divergence distillation, enabling accurate learning of the PF-ODE divergence without explicit Jacobian computation. They theoretically show that the learned vector field belongs to the Stein class under suitable conditions. By combining function approximation with regularization techniques, StAD significantly reduces estimation variance and accelerates likelihood evaluation on benchmarks such as CIFAR-10 and ImageNet, while demonstrating broad applicability across diverse generative modeling frameworks.

diffusion modelsdivergence computationflow-based models

Existing methods struggle to perform full-path statistical inference for gradient flow optimization trajectories, particularly lacking valid uncertainty quantification when stopping times are data-dependent or the path diverges. This work establishes a time-uniform statistical inference theory for gradient flows by modeling the deviation between empirical and population gradient flows as a continuous Gaussian process indexed over the non-negative real line. It presents the first uniform central limit theorem applicable across the entire optimization trajectory. Building on this foundation, the paper introduces an algorithm-aware covariance estimator that requires neither matrix inversion, resampling, nor data splitting, and which converges uniformly over time. The resulting confidence bands achieve asymptotically valid coverage, offering a theoretically rigorous and practically useful tool for path-level uncertainty quantification in gradient-based algorithms.

empirical risk minimizationgradient flowsstatistical inference

Diffusion models pose a challenge for direct application of policy gradient–based reinforcement learning methods due to their intractable likelihood, and existing research lacks a systematic analysis of how likelihood estimation affects optimization. This work presents the first disentanglement of three key components in reinforcement learning for diffusion models: the policy gradient objective, the likelihood estimator, and the sampling strategy. The study reveals that the final-sample likelihood estimate based on the evidence lower bound (ELBO) is the dominant factor governing optimization efficacy, underscoring the centrality of likelihood estimation over reliance on loss function design. Experiments on SD 3.5 Medium demonstrate that the proposed approach improves the GenEval score from 0.24 to 0.95, achieves 4.6× higher training efficiency than FlowGRPO and 2× that of the state-of-the-art DiffusionNFT, and exhibits no reward hacking behavior.

Diffusion ModelsLikelihood EstimationPolicy Gradient

Traditional normalizing flows struggle to capture the heavy-tailed nature of financial returns, leading to biased estimates of Value-at-Risk (VaR) and Expected Shortfall (ES). This work proposes Lévy-Flow, the first framework to integrate Lévy-driven heavy-tailed distributions—specifically Variance Gamma (VG) and Normal-Inverse Gaussian (NIG)—into normalizing flows. The model explicitly captures tail behavior while preserving exact likelihood computation and enabling efficient reparameterized sampling. Theoretically, it is shown that the proposed flow maintains the tail index under asymptotically linear transformations, which motivates the design of an Identity-tail Neural Spline Flow to faithfully preserve the base distribution’s tail shape. Empirical results on S&P 500 daily returns demonstrate that the VG flow reduces test negative log-likelihood by 69% compared to Gaussian flows and achieves well-calibrated 95% VaR, while the NIG flow yields the most accurate ES estimates.

density estimationExpected Shortfallfinancial risk management

Hot Scholars

YL

Yuan Liu

Assistant Professor, Hong Kong University of Science and Technology (HKUST)
Computer Graphics3D Computer VisionPhotogrammetryRemote Sensing
MS

Michael Schapira

Professor of Computer Science, The Hebrew University of Jerusalem
NetworkingComputer NetworksMachine LearningAlgorithmic Game Theory
YY

Yibo Yan

East China Normal University
High-dimensional Statistics
SS

Sina Sharifi

Johns Hopkins University
Machine LearningOptimization
AC

Adriane Chapman

Professor of Computer Science, University of Southampton
Data Management