๐ค AI Summary
This study addresses the limitation of existing model marketplaces that treat models as indivisible units, thereby hindering the exploitation of complementary strengths among heterogeneous experts. To this end, it pioneers the elevation of Mixture-of-Experts (MoE) from a network architecture to a market coordination mechanism. Methodologically, we construct a gating-network-based scheduling framework to deliver composite services, proposing a cost-aware gating strategy alongside a revenue allocation rule that accounts for execution costs. Furthermore, neural and tree-based ensemble algorithms are integrated to estimate latency costs. Experimental results demonstrate that the proposed framework achieves the highest average welfare and lower expected costs across multiple benchmarks, while requiring significantly less computational overhead for allocation compared to Shapley value-based methods.
๐ Abstract
Existing model marketplaces typically trade and select individual models as indivisible units, limiting their ability to exploit complementarities among heterogeneous experts. This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism. In this framework, brokers use gating networks to coordinate multiple heterogeneous experts and deliver a composite model service. We formalize the market participants, service workflow, expert cost structure, and a welfare objective that combines predictive utility with heterogeneous execution costs. We then derive a cost-aware gating mechanism and market-aware training objective, and introduce a cost-adjusted revenue allocation rule that distributes residual revenue according to realized expert participation and execution cost. We also establish basic theoretical properties of the allocation rule, including budget balance, participation monotonicity, and cost sensitivity. Experiments over five random seeds on fifteen tabular and image benchmarks use independently trained and frozen neural and tree-based experts together with latency-derived execution costs. MoE Market achieves the highest mean welfare on all fifteen datasets and a lower mean expected cost than Standard MoE in every case, while maintaining competitive predictive performance. The allocation experiments further demonstrate systematic sensitivity to expert participation and cost, together with substantially lower computational overhead than exact Shapley allocation. These results suggest that MoE can serve as a market-level coordination principle for collaborative, cost-aware, and economically grounded model marketplaces.