🤖 AI Summary
This study addresses the limitation that Model Context Protocol (MCP) server error messages are designed for human users, thereby preventing tool-calling AI agents from extracting actionable recovery information. By analyzing 150 MCP servers, we reveal a negative correlation between human-oriented error messaging and agent capabilities. Leveraging the Berkeley Function-Calling Leaderboard (BFCL) benchmark suite with OpenAI models, this work proposes remediation strategies that specify alternative tools or eliminate invalid steps within error responses. Experimental results demonstrate that these optimized error messages significantly improve task recovery rates in credential expiration scenarios from 45% to 84%, while achieving 88% in rate-limiting scenarios. Furthermore, the proposed approach substantially mitigates performance degradation observed in high-capability models, highlighting the importance of agent-aware error design in enhancing the robustness of tool-augmented AI systems.
📝 Abstract
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's"Wait before retrying."left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.