Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective

📅 2026-01-09
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how the Mixture-of-Experts (MoE) architecture reshapes the geometry of learned functions and representations through its routing mechanism. Viewing MoE as a soft clustering in function space, we propose a Dual Jacobian-PCA spectral probe to jointly analyze local functional curvature and representation geometry around each expert. From a spectral-geometric perspective, we reveal for the first time that MoE significantly reduces local curvature and enhances the effective rank of representations via soft partitioning, while routing sharpness governs the trade-off between expert specialization and diversity. Experiments on WikiText—employing exact Jacobian computation, weighted PCA, and controlled MLP-MoE and Transformer models—demonstrate that MoE accelerates the decay of singular value spectra, suppresses dominant singular values, and yields expert representations characterized by high effective rank and low inter-expert overlap within a low-rank structure.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Learning to SearchComputer Vision: Representation Learning for Vision

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
📝 Abstract
Mixture-of-Experts (MoE) architectures are commonly motivated by efficiency and conditional computation, but their effect on the geometry of learned functions and representations remains poorly characterized. In this work, we study MoEs through a geometric lens, interpreting routing as a form of soft partitioning of the representation space into overlapping local charts. We introduce a Dual Jacobian-PCA Spectral Geometry probe. It analyzes local function geometry via Jacobian singular-value spectra and representation geometry via weighted PCA of routed hidden states. Using a controlled MLP-MoE setting that permits exact Jacobian computation, we compare dense, Top-k, and fully-soft routing architectures under matched capacity. Across random seeds, we observe that MoE routing consistently reduces local sensitivity, with expert-local Jacobians exhibiting smaller leading singular values and faster spectral decay than dense baselines. At the same time, weighted PCA reveals that expert-local representations distribute variance across a larger number of principal directions, indicating higher effective rank under identical input distributions. We further find that average expert Jacobians are nearly orthogonal, suggesting a decomposition of the transformation into low-overlap expert-specific subspaces rather than scaled variants of a shared map. We analyze how routing sharpness modulates these effects, showing that Top-k routing produces lower-rank, more concentrated expert-local structure, while fully-soft routing yields broader, higher-rank representations. Together, these results support a geometric interpretation of MoEs as soft partitionings of function space that flatten local curvature while redistributing representation variance.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
function geometry
representation geometry
routing mechanism
spectral analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
soft clustering
Jacobian spectrum
weighted PCA
spectral geometry
🔎 Similar Papers
2023-03-10Journal of ClassificationCitations: 10
2024-09-01arXiv.orgCitations: 4