domain-aware feature extraction

Designs and implements feature extraction systems that produce representations conditioned on or partitioned by input domains, typically by composing multiple expert extractors with routing or gating mechanisms (mixture-of-experts) to capture domain-specific structure while enabling shared capacity. Engineers and evaluates the extractor architecture, routing policies, and training/regularization procedures to reduce inter-domain representation interference and stabilize latent representations across domains.

domain-awarefeatureextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications

Mar 10, 2025
SM
Siyuan Mu
🏛️ Sichuan Agricultural University | University of Houston

This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.

Addresses computational resource challenges in large AI modelsImproves handling of diverse and complex datasets with MoESummarizes advancements and applications of Mixture-of-Experts models

Must-Read Papers

Most classic and influential ideas
View more

Tight Clusters Make Specialized Experts

Feb 21, 2025
SK
Stefan K. Nielsen
🏛️ FPT Software AI Center | National University of Singapore

In high-dimensional sparse Mixture-of-Experts (MoE) models, conventional routers struggle to discern latent token clustering structures, resulting in slow convergence, poor robustness against data corruption, and degraded representation learning. To address this, we propose the Adaptive Clustering (AC) router: it employs a learnable feature-weighting mapping to project tokens into a latent space conducive to expert separation; and introduces, for the first time, an expert-level compactness-aware dynamic feature weighting mechanism—enabling each expert to specialize within semantically coherent subspaces. The method integrates clustering optimization theory, adaptive feature scaling, and token-level routing reparameterization. Evaluated on language modeling and image recognition tasks, the AC router achieves significant improvements in convergence speed, robustness, and overall performance, effectively mitigating routing failure induced by clustering unidentifiability in high-dimensional spaces.

Enhances clustering for better expert specializationImproves robustness and performance in diverse tasksOptimizes token-expert routing in MoE models

This work addresses the limitation of existing domain generalization methods that enforce global invariance, thereby overlooking transferable features shared only among subsets of source domains and constraining representational capacity. To overcome this, the authors propose a subset-shared invariance hypothesis and introduce a mixture-of-experts architecture to learn localized invariances: each expert aligns representations within a specific subset of domains, while a routing mechanism dynamically combines expert outputs. The approach jointly optimizes feature representations, selective alignment objectives, and a confidence-balanced routing strategy, further enhanced by a training scheme that encourages diverse expert specialization. Evaluated on the DomainBed benchmark, the method achieves substantial improvements in out-of-domain generalization and demonstrates heightened robustness under increased domain heterogeneity.

domain generalizationdomain heterogeneityinvariance

This work challenges the prevailing assumption that Mixture-of-Experts (MoE) models achieve domain specialization through sparse routing, introducing the COMMITTEEAUDIT framework to systematically analyze expert-level routing behavior. Through quantitative and qualitative evaluation of multiple representative MoE models on the MMLU benchmark, we uncover the existence of persistent “standing committees”—a small subset of experts that consistently dominate routing weights across domains and layers, regardless of routing budget constraints. These core experts anchor structural and syntactic reasoning, while peripheral experts handle only narrow, domain-specific knowledge. Our findings reveal that the actual degree of specialization in MoE models is substantially lower than commonly assumed and suggest that current load-balancing training objectives may conflict with the model’s intrinsic optimization dynamics.

domain specializationexpert utilizationMixture of Experts

This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.

expert specializationinference optimizationMixture of Experts

Latest Papers

What's happening recently
View more

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.

Diffusion TransformersRouting CollapseSelective Deadlock

This study addresses the limited interpretability of expert routing mechanisms in Mixture-of-Experts (MoE) models and the shortcomings of the prevailing "single-domain specialization" assumption by proposing a novel "superimposed specialization" paradigm. To this end, we introduce RouterInterp, a method that integrates sparse autoencoders, feature prediction, and natural language generation to elucidate the intrinsic relationship between sparse features and routing decisions while automatically producing natural language explanations. This approach enables precise mapping from broad domains to fine-grained features. Evaluated on the gpt-oss-20b model, RouterInterp achieves an approximate 65% improvement in detection accuracy over conventional baselines, substantially enhancing the interpretability of MoE architectures.

Expert RoutingInterpretabilityMixture of Experts

This study addresses the substantial computational overhead and inefficiency encountered when adapting multimodal Mixture-of-Experts (MoE) models to specific domains. To this end, it proposes ExpertLens, a data-free framework that exploits the semantic modularity inherent in MoE sparsity. By decoding pretrained router weights, the method precisely identifies domain-relevant experts and applies selective fine-tuning exclusively to their associated parameters. This work highlights the semantic modularity value of sparse architectures. Experimental results demonstrate that updating merely 21.7%–47.0% of the parameters achieves performance on par with full fine-tuning while accelerating training up to fourfold, significantly outperforming mainstream baselines such as LoRA.

Domain ExpertsEfficient AdaptationMultimodal Mixture-of-Experts

Hot Scholars

BT

Bijun Tang

Presidential Postdoctoral Fellow, Nanyang Technological University
2D MaterialsPhase EngineeringAI for Materials Science
YD

Yanchen Deng

Nanyang Technological University
Multiagent systemsArtificial intelligenceConstraint programmingCombinatorial optimization
XW

Xinrun Wang

Singapore Management University
Reinforcement LearningMulti-Agent SystemsDistributed Foundation Agents
PY

Penghui Yang

CCDS, Nanyang Technological University
Machine Learning
HY

Hsuan-Yu Fan

MS student, Computer Science, National Cheng Kung University
Multimodal Models