Dynamic Expert Pruning for Multi-Agent Systems

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of static pruning for Mixture-of-Experts (MoE) models in heterogeneous multi-agent systems, where fixed architectures struggle to accommodate dynamic task demands. To this end, we propose a dynamic expert pruning method that employs a lightweight predictor to generate bespoke masks in real time based on system prompts, enabling on-demand expert activation. Notably, this work introduces the first mechanism capable of producing dynamic masks via a single forward pass without requiring offline calibration. Experimental results demonstrate that our approach outperforms static baselines in accuracy across varying model scales and unseen workflows. By significantly reducing the number of retained experts while effectively controlling performance degradation, the proposed method substantially enhances inference serving sparsity and deployment efficiency.
📝 Abstract
Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing methods are static --- a single mask, calibrated offline, is applied to the model for every subsequent request. This assumption can fail when the workload is heterogeneous, most prominently in multi-agent systems, where one backbone serves many tasks and roles at once: our analysis shows that different tasks and roles recruit different experts, while static methods assign one fixed subset to all of them. We therefore propose Dynamic Expert Pruning (DEP), which rests on a finding we establish here: an agent's system and task prompts are by themselves sufficient to identify the experts that agent and its task require, since that text already describes what the agent will do. A lightweight predictor, trained once on workflow transcripts, turns those prompts into a specialized per-request mask in a single forward pass, with no per-configuration calibration. Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, and generalizes to workflows unseen in training without retraining. Its margin over those baselines is largest when few experts are retained, suggesting that the role specialization inherent to multi-agent systems permits sparser serving than static pruning allows.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Expert Pruning
Multi-Agent Systems
Dynamic Pruning
Memory Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Expert Pruning
Mixture-of-Experts
Multi-Agent Systems
Prompt-based Routing
Model Compression