dual-discriminator training

Design and train generative adversarial networks that use two or more discriminators—e.g., a general adversarial discriminator for realism and a separate activity discriminator for semantic or label-level consistency—by specifying and combining distinct loss terms and update schedules. Build and analyze training procedures, gradient-weighting, and discriminator architectures to stabilize adversarial learning and better align generated sequences or outputs with desired activity-level objectives.

dual-discriminatortraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of theoretical guarantees for generalization in existing discriminator-guided generative models. Building upon the strong duality of f-divergences, we propose a universal discriminator-guided refinement framework that enhances the generalization capability of any generative model, including diffusion models. We provide the first theoretical proof that this guidance mechanism provably reduces the generalization gap and establish a quantitative relationship between the gap reduction and the Rademacher complexity of the discriminator class. Furthermore, our framework offers a unified theoretical explanation for the empirical success of recent score-based diffusion methods. While maintaining broad applicability, the proposed approach delivers rigorous theoretical justification for widely used yet previously heuristic refinement strategies.

discriminator guidancef-divergencesgeneralization

Generalized Dual Discriminator GANs

Jul 23, 2025
PN
Penukonda Naga Chandana
🏛️ International Institute of Information Technology, Hyderabad

To address mode collapse in generative adversarial networks (GANs), this paper proposes the dual-discriminator α-GAN and establishes a generalized dual-discriminator GAN framework based on arbitrary convex functions defined on the positive reals. Theoretically, we prove that its optimization objective is equivalent to a weighted sum of an $f$-divergence and its inverse, thereby embedding the α-loss into the dual-discriminator architecture for the first time. This unifies and substantially extends the theoretical foundations of existing dual-discriminator approaches. Training proceeds via min-max collaborative optimization with tunable loss functions, ensuring stability. Experiments on 2D synthetic data demonstrate that the proposed method significantly improves training stability and generation diversity, consistently outperforming baseline models across multiple quantitative metrics.

Combining dual discriminators with tunable α-loss functionGeneralizing dual discriminator approach to arbitrary f-divergencesMitigating mode collapse in GANs using dual discriminators

Generative Feature Training of Thin 2-Layer Networks

Nov 11, 2024
JH
Johannes Hertrich
🏛️ University College London | Technische Universitat Chemnitz

To address the susceptibility of gradient-based optimization to local minima in two-layer neural networks under few-shot function approximation, this paper proposes a generative latent-feature initialization method. Specifically, it employs a deep generative model to learn a prior distribution over hidden-layer weights—replacing conventional random initialization—and integrates latent-space gradient fine-tuning with noise-robust ℓ₂ regularization; the output layer is solved in closed form via linear least squares. The approach preserves model lightness (i.e., few hidden units) while substantially improving generalization accuracy and training stability. Numerical experiments demonstrate consistent superiority over standard initialization schemes across multiple few-data benchmarks. This work establishes a new paradigm for reliable function approximation with shallow networks in data-scarce regimes.

Approximating functions with 2-layer neural networks efficientlyOvercoming local minima in gradient-based trainingUsing generative models for improved weight initialization

Investigating Generalization Behaviours of Generative Flow Networks

Feb 07, 2024
LA
Lazar Atanackovic
🏛️ University of Toronto | Vector Institute | Valence Labs

This study systematically investigates the generalization mechanisms of Generative Flow Networks (GFlowNets) in discrete spaces. To rigorously test generalization hypotheses, we introduce a graph-structured benchmark environment with tunable reward difficulty, enabling exact computation of the true distribution $p(x)$ and precise quantification of generalization error. Methodologically, we establish a unified training–evaluation framework to conduct offline and off-policy generalization analysis, as well as experiments on implicit reward robustness. Our key contributions are threefold: (1) We empirically demonstrate that GFlowNet generalization arises from implicit structural priors encoded in the learned flow function—not from explicit regularization; (2) We reveal high sensitivity to training distribution shift, yet strong robustness to perturbations in implicit reward signals; (3) We challenge the prevailing assumption that GFlowNets inherently generalize better than alternatives, providing new empirical foundations and design insights for generalization theory in discrete generative modeling.

Explores GFlowNets' sensitivity to offline and off-policy training.Investigates generalization behaviors of Generative Flow Networks (GFlowNets).Tests hypothesized mechanisms of GFlowNets' generalization using a graph-based benchmark.

Stability and Generalization in Free Adversarial Training

Apr 13, 2024
XC
Xiwei Cheng
🏛️ The Chinese University of Hong Kong | Purdue University

This work investigates the intrinsic relationship between generalization performance and optimization stability in Free Adversarial Training (FreeAT). Addressing the large generalization gap and poor training stability inherent in standard adversarial training, we first establish—within the algorithmic stability framework—that FreeAT achieves a tighter generalization error bound by jointly optimizing perturbations and model parameters. We further reveal that its synchronous min-max optimization mechanism is critical for narrowing the train-test accuracy gap. Theoretical analysis demonstrates that FreeAT’s generalization upper bound is significantly lower than that of standard adversarial training. Empirical evaluation confirms that, under identical iteration budgets, FreeAT consistently reduces the generalization gap by 15–22% across diverse benchmarks. Our implementation is publicly available.

Adversarial TrainingGeneralizationStability

Latest Papers

What's happening recently
View more

Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training

Oct 15, 2025
YW
Yisen Wang
🏛️ Peking University | The University of Hong Kong

Adversarial training (AT) suffers from degraded natural accuracy and poor cross-attack robustness generalization. To address these challenges, we propose a multi-task collaborative generalization framework that decomposes global robust learning into multiple specialized subtasks, each optimized by a dedicated base learner. Knowledge fusion and co-evolution between base learners and the global model are achieved via dynamic parameter interpolation, gradient synchronization, and periodic task reallocation—without incurring additional inference overhead. The framework supports three lightweight variants for diverse deployment scenarios. Theoretical analysis shows that our method reduces generalization error. Empirically, it achieves average improvements of 2.1% in natural accuracy and 3.7% in robust accuracy across multiple adversarial attacks, significantly outperforming baselines including PGD-AT and TRADES. This work advances the development of general-purpose robust classifiers.

Addressing trade-offs between accuracy and adversarial robustnessImproving robustness transfer across different attack typesMitigating natural accuracy degradation in adversarial training

This work addresses the instability in semi-supervised GAN training caused by the inherent conflict between maximizing classification accuracy and enhancing the discriminator’s ability to distinguish real from fake samples. To resolve this, the study introduces a multi-objective evolutionary algorithm into the training framework for the first time, formulating discriminator optimization as a multi-objective problem. By leveraging population-based evolutionary strategies and Pareto dominance, the method preserves a diverse set of non-dominated solutions without resorting to scalar loss aggregation. Evaluated on label-limited MNIST, the approach significantly improves training stability, and its elite variant achieves state-of-the-art classification accuracy, outperforming baseline models such as SSL-GAN and CE-SSL-GAN.

Classification accuracyDiscriminator trainingMulti-objective optimization

This work presents the first systematic survey and critical analysis of neural architecture search (NAS) methods tailored for generative adversarial networks (GANs), addressing the inefficiency and instability of manual GAN design, which often struggles to balance performance and generalization. The study organizes existing approaches through a structured comparison based on search strategies, evaluation metrics, and empirical performance. It advocates moving beyond conventional Inception Score (IS) and Fréchet Inception Distance (FID) toward more robust evaluation frameworks and diverse datasets. The analysis further highlights the complementary strengths of evolutionary algorithms and gradient-based methods across different scenarios. By clarifying the current limitations and untapped potential of NAS-GAN methodologies, this work establishes a foundation for future research and advances the standardization and performance of automated GAN architecture design.

Architecture OptimizationAutomated DesignGAN Performance

Adversarial Flow Models

Nov 27, 2025
SL
Shanchuan Lin
🏛️ ByteDance Seed

This work addresses the inherent trade-offs among generation efficiency, training stability, and mapping interpretability in generative modeling by unifying adversarial generative networks (GANs) and normalizing flows. We propose the Adversarial Flow Model (AFM), which directly learns a deterministic optimal transport map from latent noise space to data space—eliminating stochasticity and iterative refinement, and enabling one-step (1-NFE) end-to-end generation. AFM incorporates a deep repeated architecture and an adversarial training objective to ensure convergence and gradient stability in deep networks. On ImageNet-256, the 1-NFE variant achieves a Fréchet Inception Distance (FID) of 2.38; deeper variants with 56 and 112 layers attain FIDs of 2.08 and 1.94, respectively—surpassing Consistency Models XL/2 and leading multi-step methods. To our knowledge, AFM is the first framework to achieve state-of-the-art performance in single-step generative modeling.

Achieves high performance with reduced training and model capacityEnables one-step or few-step generation without intermediate timestepsUnifies adversarial and flow models for stable generation

Existing methods struggle to generate conditional risk scenarios aligned with downstream risk objectives. This work proposes a Generative Adversarial Regression (GAR) framework that extends the elicitable regression properties of risk functionals—such as Value-at-Risk (VaR) and Expected Shortfall (ES)—from point prediction to generative modeling. GAR jointly optimizes a conditional generator and a strategy-aware discriminator through a minimax adversarial mechanism, ensuring risk consistency across a broad class of strategies without requiring a pre-specified set of candidate policies. The framework robustly produces scenarios that faithfully align with the true underlying risk distribution. Empirical experiments on S&P 500 data demonstrate that GAR-generated scenarios significantly outperform unconditional, econometric, and direct prediction baselines in terms of downstream risk accuracy, while maintaining stability under adversarial strategies.

adversarial robustnessconditional risk scenariosgenerative modeling

Hot Scholars

CG

Can Gao

Shenzhen University
Machine Learning
YG

Yihong Gong

Xi'an Jiaotong University
Multimedia content analysisMachine learningPattern recognition
PM

Patryk Marszałek

Master of Science student, Jagiellonian University in Krakow
machine learning