Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference

📅 2025-09-09
📈 Citations: 0
Influential: 0
📄 PDF

career value

219K/year
🤖 AI Summary
Routing queries to heterogeneous AI execution units—such as large language models (LLMs) and domain-specific agents—remains challenging due to their disparate capabilities and dynamic operational contexts. Method: This paper proposes MoMA, a framework that unifies the hybrid capability space of LLMs and agents via fine-grained intent recognition and context-aware finite-state machine modeling. It introduces a dynamic masking mechanism for adaptive, scalable routing decisions, integrates multi-LLM architectures for capability modeling and task matching, and constructs a dedicated benchmark dataset for rigorous evaluation. Contribution/Results: Experiments demonstrate that MoMA significantly improves routing accuracy while maintaining inference quality and reducing service costs. It achieves both high efficiency and scalability, establishing a novel paradigm for intelligent scheduling in multi-agent systems.

Technology Category

Application Category

📝 Abstract
The rapid advancement of large language models (LLMs) and domain-specific AI agents has greatly expanded the ecosystem of AI-powered services. User queries, however, are highly diverse and often span multiple domains and task types, resulting in a complex and heterogeneous landscape. This diversity presents a fundamental routing challenge: how to accurately direct each query to an appropriate execution unit while optimizing both performance and efficiency. To address this, we propose MoMA (Mixture of Models and Agents), a generalized routing framework that integrates both LLM and agent-based routing. Built upon a deep understanding of model and agent capabilities, MoMA effectively handles diverse queries through precise intent recognition and adaptive routing strategies, achieving an optimal balance between efficiency and cost. Specifically, we construct a detailed training dataset to profile the capabilities of various LLMs under different routing model structures, identifying the most suitable tasks for each LLM. During inference, queries are dynamically routed to the LLM with the best cost-performance efficiency. We also introduce an efficient agent selection strategy based on a context-aware state machine and dynamic masking. Experimental results demonstrate that the MoMA router offers superior cost-efficiency and scalability compared to existing approaches.
Problem

Research questions and friction points this paper is trying to address.

Routing diverse user queries to appropriate AI execution units
Optimizing performance and efficiency in AI service ecosystems
Balancing cost and effectiveness through adaptive routing strategies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generalized routing framework integrating LLM and agents
Dynamic query routing based on cost-performance efficiency
Agent selection via context-aware state machine and masking
🔎 Similar Papers
X
Xiyu Guo
JIUTIAN Team, China Mobile Research Institute, Beijing, China
S
Shan Wang
JIUTIAN Team, China Mobile Research Institute, Beijing, China
C
Chunfang Ji
JIUTIAN Team, China Mobile Research Institute, Beijing, China
Xuefeng Zhao
Xuefeng Zhao
Professor, Dalian university of technology
structural health monitoringLife-Cycle EngineeringLife-Cycle SHM
W
Wenhao Xi
JIUTIAN Team, China Mobile Research Institute, Beijing, China
Y
Yaoyao Liu
JIUTIAN Team, China Mobile Research Institute, Beijing, China
Q
Qinglan Li
JIUTIAN Team, China Mobile Research Institute, Beijing, China
C
Chao Deng
JIUTIAN Team, China Mobile Research Institute, Beijing, China
Junlan Feng
Junlan Feng
Chief Scientist at China Mobile Research
Natural LanguageMachine LearningSpeech ProcessingData Mining