wasserstein gan

Designs and trains adversarial generator–critic models that match a model’s output distribution to a target distribution by having a critic network approximate the Wasserstein (Earth-Mover) distance and provide a scalar loss used to update the generator. This includes the algorithms, architectures, and regularization/optimization techniques for stable Wasserstein-based distribution matching (adversarial distribution matching).

wassersteingan

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Wasserstein distributional adversarial training for deep neural networks

Feb 13, 2025
XB
Xingjian Bai
🏛️ Massachusetts Institute of Technology | Imperial College London | University of Oxford

This work addresses the challenge of enhancing deep neural network robustness under concurrent distributional shift and adversarial perturbations. We propose a Wasserstein-distance-based distributed adversarial training method that, for the first time, incorporates sensitivity analysis from Wasserstein distributionally robust optimization (DRO) into the adversarial training framework—extending TRADES to jointly optimize both distributional and pointwise robustness. Our approach requires no additional data, enabling efficient fine-tuning of pre-trained models using only the standard 50k CIFAR-10 training samples. Experiments on RobustBench demonstrate that our method significantly improves Wasserstein distributional robustness while strictly preserving the strong pointwise robustness of baseline models. Notably, it delivers consistent performance gains even when applied to large-scale models pre-trained on millions of synthetic samples, using only small-scale real data. This establishes a practical, data-efficient pathway toward simultaneously strengthening both types of robustness without architectural or data augmentation overhead.

Enhance robustness against distributional adversarial attacksExtend TRADES method for pointwise attacksFine-tune pre-trained models for improved performance

Partial Distribution Matching via Partial Wasserstein Adversarial Networks

Sep 16, 2024
ZW
Zi-Ming Wang
🏛️ University of Chalm

This work addresses the Partial Distribution Matching (PDM) problem—robustly aligning salient subsets of two probability distributions without requiring full distributional alignment. To formalize PDM, we establish, for the first time, the Kantorovich–Rubinstein duality theory for the partial Wasserstein-1 distance. Leveraging this theoretical foundation, we propose PWAN, an adversarial network framework that enables efficient, differentiable partial matching via gradient-based optimization. Our method integrates partial Wasserstein metrics, adversarial training, and bilevel optimization. Evaluated on 3D point-set registration and partial-domain adaptation, PWAN achieves state-of-the-art or competitive performance, significantly improving matching robustness and generalization. It overcomes the strong assumption of complete distribution alignment inherent in conventional optimal transport–based methods.

Application in point set registration and partial domain adaptation tasksPartial matching of probability distributions for robust alignmentTheoretical derivation of Kantorovich-Rubinstein duality for partial Wasserstein discrepancy

This work addresses the challenge of minimizing the Wasserstein-2 (W₂) distance in unsupervised generative modeling. We propose the first explicit, distribution-dependent ordinary differential equation (ODE) characterizing the W₂ gradient flow, and theoretically prove that its time-marginal distributions converge rigorously to the target data distribution. Methodologically, we design a persistent Euler discretization algorithm that couples with the gradient flow structure, circumventing the instability inherent in conventional adversarial training. Our key contributions are: (1) the first explicit ODE formulation of the W₂ gradient flow; and (2) a novel persistence training mechanism that enhances discretization fidelity and convergence robustness. Experiments demonstrate that our approach significantly outperforms WGAN on both high- and low-dimensional benchmarks; moreover, increasing persistence strength further improves generation quality and training stability.

Minimizes Wasserstein-2 loss for unsupervised learningOutperforms WGANs with persistent training in experimentsUses ODE gradient flow to converge to data distribution

Unifying Distributionally Robust Optimization via Optimal Transport Theory

Aug 10, 2023
JB
J. Blanchet
🏛️ Stanford University | EPFL | University of British Columbia | University of Chicago

This paper addresses the conceptual and methodological divide between φ-divergence-based (likelihood-ratio-centric) and Wasserstein-based (outcome-space-centric) paradigms in distributionally robust optimization (DRO). To bridge this gap, we propose the first unified DRO framework. Methodologically, we integrate optimal transport theory with conditional moment constraints to construct a novel DRO model capable of simultaneously perturbing both likelihood ratios and outcome distributions. Via Lagrangian duality analysis, we derive a computationally tractable closed-form dual reformulation, whose equivalent problem is solvable in polynomial time. Theoretical contributions include: (i) establishing a rigorous strong duality theorem under conditional moment constraints; and (ii) introducing a new modeling paradigm for optimal transport that explicitly incorporates such constraints. Empirical results demonstrate that the proposed framework significantly enhances generalization and robustness under distributional shifts.

Bridges divergence-based and Wasserstein-based ambiguity modeling approachesEnables joint perturbation of likelihood ratios and outcomes via generalized couplingUnifies distributionally robust optimization with optimal transport theory

Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss

Oct 29, 2024
JM
José Manuel de Frutos
🏛️ Universidad Carlos III de Madrid | Instituto de Investigación Sanitaria Gregorio Marañón (IiSGM)

Traditional implicit generative models (e.g., GANs) suffer from training instability, mode collapse, and inaccurate tail characterization when modeling heavy-tailed, high-dimensional multivariate distributions. To address these challenges, this paper proposes Pareto-ISL—a novel implicit score learning framework. Its core contributions are: (1) the first integration of generalized Pareto noise into implicit score learning (ISL), explicitly capturing heavy-tailed behavior; and (2) a random-projection-based multidimensional ISL loss, extending ISL beyond univariate settings to scalable high-dimensional implicit modeling. Experiments demonstrate that Pareto-ISL accurately reproduces both central and tail regions of multivariate heavy-tailed distributions, significantly mitigates mode collapse, exhibits robustness to hyperparameter choices, and scales linearly in computational complexity with dimensionality.

Addresses unstable training and mode dropping in traditional generative modelsExtends invariant statistical loss to handle heavy-tailed multivariate distributionsOvercomes computational challenges of high-dimensional data with random projections

Latest Papers

What's happening recently
View more

This work addresses the issue of “Fréchet hacking” in generative model optimization, where Fréchet distance losses based on static pretrained feature spaces yield deceptively high scores despite degraded visual quality and poor cross-feature alignment. To mitigate this, the authors propose the adversarial Fréchet distance (AdvFD) loss, which introduces adversarial learning into Fréchet distance optimization for the first time. AdvFD constructs a learnable, adaptive feature space that dynamically enhances distribution discrepancy measurement and incorporates a real-feature whitening mechanism to suppress feature amplification and stabilize training. The method consistently improves both visual fidelity and distribution alignment in single-step generator post-training across various model scales and backbone architectures, including JiT and pMF.

Fréchet distanceFréchet hackinggenerator post-training

This study addresses the challenge of evaluating worst-case risk in distributionally robust optimization (DRO) under non-convex losses by reconstructing the adversarial training framework through the lens of optimal transport geometry. It reveals that standard approaches incur computational waste by violating cyclic monotonicity, and accordingly proposes a multi-start particle ascent algorithm alongside an input convex neural network-based parameterization for adversarial mappings. Experimental results demonstrate that the proposed method significantly enhances model robustness and generalization under distribution shifts across regression, classification, and control tasks.

Adversarial TrainingCyclical MonotonicityDistributionally Robust Optimization

Recently, Deng et al. (2026) proposed Generative Modeling via Drifting (GMD), a novel framework for generative tasks. This note presents an analysis of GMD through the lens of Wasserstein Gradient Flows (WGF), i.e., the path of steepest descent for a functional in the space of probability measures, equipped with the geometry of optimal transport. Unlike previous WGF-based contributions, GMD can be thought of as directly targeting a fixed point of a specific WGF flow. We demonstrate three main results: first, that one algorithm proposed by Deng et al. (2026) corresponds to finding the limiting point of a WGF on the KL divergence, with Parzen smoothing on the densities. Second, that the algorithm actually implemented by Deng et al. (2026) corresponds to a different procedure, which bears some resemblance to the fixed point of a WGF on the Sinkhorn divergence, but lacks certain desirable properties of the latter. Third, the same same idea can be extended to the limiting point of other WGFs, including the Maximum Mean Discrepancy (MMD), the sliced Wasserstein distance, and GAN critic functions.

Fixed PointGenerative Modeling via DriftingOptimal Transport

This work addresses a critical inconsistency in existing conditional flow matching (CFM) approaches for distributional reinforcement learning, where arbitrary source–target pairings yield losses misaligned with the Wasserstein distance, thereby violating the contraction property of the Bellman operator. To resolve this, the authors propose FlowIQN, which constructs quantile-aligned, monotonic optimal transport couplings by sorting source samples and Bellman targets within each minibatch, ensuring flow trajectories consistent with the Wasserstein metric. FlowIQN provides the first explicit Wasserstein projection guarantee for flow-matching distributional critics and incorporates a shortcut inference model to enhance computational efficiency. Empirical results demonstrate that FlowIQN significantly improves the Wasserstein accuracy of return distributions and achieves strong performance across multiple offline reinforcement learning benchmarks under various policy extraction settings, offering both theoretical rigor and practical effectiveness.

Conditional Flow MatchingDistributional Reinforcement Learningmetric mismatch

This study addresses the absence of a unified theoretical framework for uncertainty estimation and generalization analysis in distributionally robust linear regression. By leveraging the Wasserstein distance, the proposed approach employs quadratic reformulation to unify square-root Lasso and adversarial training, revealing the equivalence between large and small ambiguity set solutions while establishing noise-insensitive pivotal properties. The primary contribution lies in deriving non-asymptotic error bounds that achieve convergence rates of O(n^{-1/2}) under general conditions and O(n^{-1}) under sparsity assumptions. Numerical experiments further demonstrate that the method successfully combines theoretical rigor with computationally efficient solving capabilities.

Adversarial trainingDistributionally robust optimizationLinear regression

Hot Scholars

SZ

Shaoting Zhang

Shanghai AI Lab; SenseTime Research
Medical Image AnalysisComputer VisionFoundation Models
XZ

Xianhao Zhou

University of Electronic Science and Technology of China
computer vision
JW

Jianghao Wu

Monash University
Medical Image AnalysisComputer VisionNatural Language Processing