MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing benchmarks overlook the dynamic evolution of tool interfaces and functionalities in Model Context Protocol (MCP) servers, limiting their ability to assess large language model (LLM) agents’ adaptability in realistically changing environments. This work proposes MCPEvol-Bench, the first benchmark incorporating a tool evolution perspective. Grounded in large-scale empirical analysis, it introduces 11 mutation operators to simulate real-world tool evolution across 123 MCP servers and systematically evaluates 12 prominent LLMs across multiple versions. Results reveal that GPT-5.4 and Claude-Sonnet-4-6 experience performance drops of 13.7% and 14.4%, respectively, under evolving conditions, accompanied by significantly increased planning and reasoning errors. These findings expose the fragility of current LLM-based workflows in dynamic tool environments.
πŸ“ Abstract
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents' tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
Problem

Research questions and friction points this paper is trying to address.

MCP servers
tool evolution
LLM agents
dynamic environments
adaptability
Innovation

Methods, ideas, or system contributions that make the work stand out.

MCPEvol-Bench
MCP server evolution
LLM agent adaptability
tool mutation operators
dynamic tool environments
πŸ”Ž Similar Papers
No similar papers found.
H
Huanxi Liu
College of Computer Science and Technology, National University of Defense Technology
K
Kun Hu
College of Computer Science and Technology, National University of Defense Technology
J
Jiaqi Liao
College of Computer Science and Technology, National University of Defense Technology
Qiang Wang
Qiang Wang
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen
GPU ComputingEnergy Efficient ComputingParallel and Distributed SystemsSpatial Intelligence
P
Pengfei Qian
College of Computer Science and Technology, National University of Defense Technology
Y
YuanZhao Zhai
College of Computer Science and Technology, National University of Defense Technology
D
Dawei Feng
College of Computer Science and Technology, National University of Defense Technology
B
Bo Ding
College of Computer Science and Technology, National University of Defense Technology
H
Huaimin Wang
College of Computer Science and Technology, National University of Defense Technology