Optimal Mixture-of-Experts Model Averaging for Conditional Generative Models

📅 2026-07-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance instability of conditional generative models across diverse tasks or inputs by proposing a novel model averaging framework that requires no explicit density estimation and relies solely on conditional samples. The approach leverages Maximum Mean Discrepancy (MMD) to measure distances between conditional distributions and introduces two fusion mechanisms: static weighting (StaticMA) and input-adaptive weighting (MoEMA). Notably, this study provides the first theoretical guarantees for MoEMA, establishing its asymptotic optimality and consistency of the learned weighting function. By incorporating softmax-gated neural networks and representation mappings, the method generalizes effectively to unstructured data. Empirical evaluations demonstrate that MoEMA consistently outperforms existing baselines across multimodal domains, including tabular, image, and text data.
📝 Abstract
Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and applied domains. As these models proliferate, practitioners often face multiple plausible generators whose performance can vary with the task, data, or input condition. We propose an optimal model averaging framework for conditional generative models, allowing candidate generators to be combined even when they are accessible only through conditional samples without tractable densities. Specifically, we use a sample-based maximum mean discrepancy between conditional distributions, which first leads to a static model averaging method, StaticMA, assigning fixed weights to different candidates. In addition, we develop MoEMA (mixture-of-experts model averaging), an input-adaptive method that parameterizes covariate-dependent weights through a softmax neural-network gate. We establish in-sample and out-of-sample asymptotic optimality for the proposed methods, together with consistency of the estimated adaptive weight function under regularity conditions. The framework applies directly to Euclidean responses and extends to unstructured data by combining our formulation with fixed representation maps. Across a broad set of simulations and real-data studies spanning tabular, image, and text modalities, MoEMA generally improves over competing baselines, demonstrating the effectiveness of our proposed methods.
Problem

Research questions and friction points this paper is trying to address.

conditional generative models
model averaging
mixture-of-experts
maximum mean discrepancy
adaptive weighting
Innovation

Methods, ideas, or system contributions that make the work stand out.

model averaging
conditional generative models
mixture-of-experts
maximum mean discrepancy
adaptive weighting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shijin Gong
School of Management, University of Science and Technology of China
B
Baihua He
School of Management, University of Science and Technology of China
X
Xinyu Zhang
Academy of Mathematics and Systems Science, Chinese Academy of Sciences; School of Management, University of Science and Technology of China