Score
Design and implement gating mechanisms that route inputs to modality-specific processing components or experts based on which sensors or data modalities are available or meet quality thresholds. Build or analyze systems that selectively enable, skip, or scale computation of modality experts to allow graceful degradation and efficient use of resources without requiring model retraining.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
Sparse Mixture-of-Experts (SMoE) models suffer from performance degradation and poor generalization under multimodal missing-input scenarios. Method: We propose Conf-SMoE, a confidence-guided two-stage imputation and gating architecture. We first theoretically identify and analyze the “expert collapse” phenomenon—where routing collapses to a subset of experts—and design a novel decoupled gating mechanism that separately computes routing scores and task-specific confidence, eliminating the need for load-balancing loss to mitigate collapse. Combined with modality-adaptive imputation, Conf-SMoE enables robust handling of missing inputs. Results: Extensive experiments across four real-world datasets and three missing patterns demonstrate that Conf-SMoE significantly improves accuracy and robustness, achieves more stable multimodal fusion, and generalizes strongly to arbitrary modality missingness. Furthermore, Gaussian/Laplace gating consistency analysis provides theoretical interpretability of the gating behavior.
This work addresses two key limitations in mixture-of-experts (MoE) models: the lack of theoretical connection between MoE routing and self-attention, and the low sample efficiency of linear gating. We propose quadratic gating—replacing conventional linear routing with a quadratic function—and establish, for the first time, its rigorous equivalence to self-attention. Leveraging this equivalence, we derive principled design criteria for optimal quadratic gating and expert functions, leading to a novel high-performance attention mechanism. Theoretically, via statistical learning analysis, we prove that quadratic gating substantially enhances the expressivity and parameter/sample efficiency of expert selection. Empirically, our MoE variant outperforms linear-gating baselines across multiple tasks; the new attention mechanism surpasses state-of-the-art methods—including FlashAttention and Multi-Head Attention—while exhibiting strong alignment between theoretical predictions and empirical results. The framework thus achieves both interpretability and practical efficacy.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
To address strong parameter coupling, slow convergence, and insufficient expert specialization caused by Softmax gating in Hierarchical Mixture of Experts (HMoE) models, this paper proposes a two-level Laplace gating mechanism—introducing Laplace-distributed gating functions uniformly at both the top- and bottom-level expert selection stages of HMoE. This design theoretically decouples parameter update paths across experts, thereby accelerating convergence and enhancing expert specialization. Theoretical analysis demonstrates superior gradient propagation properties compared to Softmax-based gating. Extensive experiments on multimodal understanding, image classification, and latent domain discovery tasks consistently outperform the Softmax-HMoE baseline, validating the proposed method’s effectiveness in modeling complex inputs and adapting to diverse downstream tasks.
This work addresses the lack of a unified computational interpretation for neural policy gating mechanisms. We propose GateMod, a theoretically grounded gating framework that couples task structure with neural circuit dynamics via the principle of free-energy minimization. GateMod comprises two core components: GateFlow—a continuous-time energy-flow model—and GateNet—a soft-competitive recurrent network—enabling emergent gating for skill composition and behavioral planning. We formally prove GateMod’s global exponential convergence and robustness under perturbations. Empirically, GateMod achieves significant performance gains over state-of-the-art methods in multi-agent cooperative tasks and human multi-armed bandit experiments. Crucially, it provides the first quantitative demonstration of how task structure modulates gating behavior through neural energy dynamics. By offering a computationally precise and empirically testable account, GateMod establishes a principled theoretical foundation for understanding strategy selection in prefrontal–basal ganglia circuits.
This study addresses the absence of routing mechanisms and the challenge of incomplete multimodal inputs in integrating Spiking Neural Networks (SNNs) with Mixture-of-Experts (MoE) models by proposing the SpikeMoE framework. Methodologically, it introduces a hippocampus-inspired spiking k-WTA routing mechanism that leverages lateral inhibition and refractory periods to achieve neuron-level dynamic expert selection. Additionally, a two-stage missing-modality modeling module is constructed to enhance robustness. This work achieves state-of-the-art SNN performance across visual, linguistic, and multimodal benchmarks, rivaling Artificial Neural Networks (ANNs) while substantially improving energy efficiency. These results validate the inherent advantages of sparse computation within neuromorphic architectures.
This work proposes the first cost-aware routing framework for supervised fine-tuning data acquisition that integrates statistical gating with an adversarial adjudication mechanism to efficiently identify high-value corpora while avoiding costly misacquisitions. The approach evaluates candidate samples along three axes—diversity, utility, and redundancy—using low-cost statistical estimates for initial filtering and triggering a multi-agent debate between proponent and opponent advocates only when confidence is insufficient. Evaluated through quality assessments with confidence intervals and controlled synthetic benchmarks, the system achieves 0.90 accuracy and 0.83 F₁ score across twelve datasets at a unit cost of just $0.017, substantially outperforming always-verify strategies. Moreover, it provides the first quantitative evidence of stance bias (52% stance reversal) and oppositional advantage (80% win rate) in LLM-based adjudication.
This work addresses the challenge scientists face in efficiently transforming raw sensor data streams into actionable insights across edge-cloud infrastructures, hindered by the need for cross-domain expertise to manage heterogeneous systems and emerging platforms such as DPUs, which impedes rapid prototyping. To overcome this barrier, the authors propose a novel paradigm that integrates pattern-based workflow engineering with AI-assisted development. Implemented on the FABRIC testbed using the Pegasus workflow system and exemplified by the Orcasound hydrophone workflow, this approach enables swift construction of applications for air quality, seismic, and soil moisture monitoring. The framework supports modular extensibility and edge deployment, substantially lowering the barrier for non-expert users to iteratively develop distributed applications. Empirical validation across multiple use cases demonstrates its effectiveness in enhancing development efficiency, accelerating prototyping cycles, and accumulating practical deployment experience.
为解决多模态传感器间虚假关联问题,提出了一种基于证据门控的正则化方法EGR,通过引入每帧和每个传感器的任务相关性信号来提高VLA策略的鲁棒性。
This study addresses the observation that the generalization capability of language models during pretraining does not improve monotonically, but instead oscillates frequently between rote memorization and intelligent reasoning. To investigate this, we construct an evaluation suite to identify and define the "mode jumping" phenomenon, modeling it as a circuit competition problem under capacity constraints. We propose a theoretical framework for capacity allocation, wherein data windows govern circuit competition, and integrate intermediate checkpoint selection with pretraining data selection strategies to monitor and control generalization dynamics. Our findings challenge the conventional assumption of stable model maturation by demonstrating that intermediate checkpoints can exhibit superior reasoning and alignment capabilities compared to the final model. Furthermore, we show that strategic data selection effectively stabilizes the generalization process throughout pretraining.