Score
Designs and implements gating and routing mechanisms that condition on a task identifier or task context to selectively activate, suppress, or route model submodules, parameters, slots, or feature pathways for each task. Builds or analyzes task-conditioned policies (dual gating, gated routing, task routing) that determine which components are used per task to achieve task-specific specialization and reduce cross-task interference.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
This study investigates the precise representational locus of behaviorally relevant states in large language models, moving beyond conventional approaches that rely solely on prompt-based interventions followed by successful post-training. Through controlled routing tasks—augmented with support data selection, held-out query evaluation, and necessity/sufficiency ablation experiments—the work distinguishes, for the first time under comparable conditions, between fixed interface reuse and prompt-position shifting, demonstrating that the former provides stronger evidence for state reuse. Methodologically, it introduces interface-matching controls, zero-retraining compiled transfer, trainable prompt-slot optimization, and a generation–reasoning branch analysis. The approach achieves precise state transfer on GPT-2 in Triop and arithmetic tasks, with fixed interfaces recovering most routing accuracy, and further validates cross-architectural consistency in Qwen models.
Multimodal multitask prediction in clinical settings faces challenges from sample-level data heterogeneity—such as co-occurring structured rating scales and unstructured clinical text—and varying task interdependencies, compounded by pervasive missing values. Method: We propose the first sample-adaptive routing framework for unified multimodal multitask learning. Built upon a mixture-of-experts architecture, it jointly learns modality-specific processing paths (for raw/fused textual and numerical features) and task-sharing strategies (dynamically assigning shared or task-specific prediction heads), enabling personalized information flow. The model is trained end-to-end to yield interpretable awareness of modality importance and task relationships. Results: Evaluated on synthetic data and real-world psychotherapy transcripts, our method significantly outperforms fixed multitask and single-task baselines in predicting depression and anxiety symptom severity. It demonstrates improved personalization and cost-effectiveness for clinical interventions, validating its practical utility in mental health applications.
This work addresses the lack of a reproducible evaluation framework for meta-decision strategies—such as task decomposition and tool invocation—in existing agent systems. We introduce MetaRoute-Bench, the first open benchmark enabling fine-grained analysis of meta-decision routing, comprising 180 synthetic tasks, 8 distinct strategies, and 30 random seeds per configuration. Evaluation employs offline seeded execution and multidimensional metrics—including success rate, cost, and latency—to ensure fair comparison. Experiments demonstrate that task-aware compositional strategies achieve a significantly higher success rate (79.4%) compared to static strategies (76.7%), single-step routing (67.4%), and direct answering (52.9%), with only marginal increases in cost (4.7%) and latency (6.4%). Ablation studies further confirm the critical contributions of compositional operations and verification mechanisms. Code and execution trajectories are publicly released.
This study investigates the relationship between safety behaviors and expert routing mechanisms in aligned Mixture-of-Experts (MoE) large language models. The authors find that a model’s safety capabilities can be concentrated in a small subset of experts and are largely independent of the routing policy, rather than being driven by dedicated refusal-oriented experts. To leverage this insight, they propose the Router-Agnostic Safety-critical Expert Tuning (RASET) framework, which integrates a contrastive routing sensitivity criterion with parameter-efficient fine-tuning to precisely identify and optimize safety-critical experts without altering the original routing behavior. Experiments demonstrate that RASET significantly steers model safety outputs with minimal semantic interference, revealing for the first time the existence of localized, manipulable expert-level safety mechanisms—and their potential vulnerabilities—within MoE architectures.
This work addresses the challenge of task interference in unified image generation and editing models based on dense diffusion Transformers, where shared parameters struggle to reconcile conflicting objectives such as localized editing and subject-driven generation. To mitigate this, the authors propose a task-aware Mixture-of-Experts (MoE) routing mechanism that leverages hierarchical task semantic annotations and prediction alignment regularization to guide the gating network in dispatching experts according to high-level semantic intent. This approach transforms the gate from a task-agnostic executor into a semantics-driven scheduler, enabling semantically grounded expert specialization. While preserving sparse activation, the method significantly outperforms dense baselines, achieving notable improvements in generation fidelity, editing controllability, and expert specialization.
This study addresses the challenges of storage location conflicts—between model weights and context—and performance instability caused by holistic trajectory mixing in self-improving GUI agents. To overcome these limitations, we propose a fine-grained component routing mechanism based on experience attributes. Specifically, this method decomposes experiences into four component types, including locators and programs, and dynamically routes each to its optimal storage location according to recurrence and state-conditionality rules. As the first component-level routing approach, this work reveals optimal storage strategies for distinct components and establishes generalizable routing principles. Experimental results demonstrate that the proposed method outperforms whole-trajectory baselines by an average of 3.5 points across multiple benchmarks, while the derived routing rules exhibit strong generalization to unseen models.
This work addresses the limitations of existing large language model–based multi-agent systems, which struggle with complex, long-horizon tasks due to a lack of parameter-level heterogeneity. The authors propose a task-oriented multi-agent framework that decomposes tasks into dependency-aware directed acyclic graphs and assigns each agent a (role, subtask) pair. By introducing dynamic mixture-of-LoRA experts coupled with semantic routing, the framework achieves role–subtask conditional specialization for the first time at both structural and parametric levels. To stabilize cooperative training under sparse rewards, they further design a hierarchical grouped relative policy optimization algorithm alongside a two-level credit assignment mechanism. Experiments demonstrate significant improvements in both overall and step-level performance across three backbone models on code generation benchmarks, with the learned specialization generalizing effectively to unseen tasks and domains.
This study addresses the inexecutability of task abstractions caused by spatial layout constraints under fixed controllers. We propose an affine-constraint-based layout repair method that compiles task requirements into affine constraints and modifies only continuous coordinates, thereby preserving the original event logic and controller. Furthermore, we introduce a most-violated-row update algorithm coupled with a quadratic projection mechanism to decouple the representation layer from the optimizer, enabling conditional bounded certification. Experimental results demonstrate that the proposed approach successfully certifies all layouts across three tasks, with its effectiveness further validated through 300 paired rollback tests.
This work addresses the challenge that language models often struggle to distinguish evidence types and perform auditable, editable causal reasoning when answering interventional questions. The authors propose a mechanism library approach with type supervision that, under a frozen fine-tuning protocol, induces discrete mechanism slots segregated by evidence type, achieving for the first time a functional disentanglement between routing and answer readout. The resulting architecture incurs negligible performance cost (quality loss of only 0.0082 nats), is exactly invertible, and fully auditable. Evaluated on the CausalWorld benchmark, the method demonstrates cross-scale efficacy (22.6M/125M parameters), with clear separation between routing and readout (output discrepancy ≤3.4×10⁻⁶), flawless execution across 250 single edits and 1,000 stacked rollbacks, and rigorous validation via preregistered, machine-verifiable criteria.
This work addresses a critical limitation in existing large language model (LLM) routing strategies, which rely solely on model-level labels to determine substitutions while neglecting the influence of a model’s role within multi-call workflows and its deployment context on actual performance. To remedy this, the authors propose a role-conditioned substitution principle that decouples replacement decisions into “whether to replace” and “effect evaluation” via predicate-action decomposition. Through a controlled solve-merge-validate pipeline, they systematically investigate the conditional dependencies governing effective substitutions. Empirical results from multi-call LLM workflow experiments—including input-matching interventions, allocation ablations, and cross-model comparisons across Qwen and GPT families—demonstrate that substitution efficacy is jointly determined by the model’s procedural role and deployment context. Notably, in mixed Qwen/GPT configurations, sparse, role-aware replacements reduce RMSE from 4.818 to 1.538, significantly outperforming indiscriminate full-model upgrades.