🤖 AI Summary
This study addresses the limitation of existing recommender system benchmarks, which predominantly assume explicit user intents and simplified tool environments, thereby inadequately evaluating tool orchestration capabilities under ambiguous intents. To this end, this work proposes the first recommendation-specific tool orchestration benchmark tailored for ambiguous user intents. Built upon the Model Context Protocol (MCP), the benchmark constructs over a thousand tasks through a synthesis–obfuscation–evaluation pipeline, encompassing invocation patterns ranging from single-step to complex hybrid multi-step calls. Evaluation is conducted by combining rule-based verification with LLM-assisted scoring. Experimental results demonstrate that large language models exhibit significant bottlenecks in semantic parameter grounding and multi-step evidence integration, confirming that tool orchestration remains a core challenge for agentic recommendation systems.
📝 Abstract
Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, existing benchmarks often assume explicit user intent, simplified tool environments, or isolated function calls, leaving realistic tool orchestration for recommendation underexplored. To bridge this gap, we propose RecToolBench, a Model Context Protocol (MCP)-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. RecToolBench contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel tool calls, sequential tool chains, and hybrid tool orchestration. We construct RecToolBench with a scalable synthesize--fuzzify--judge pipeline that generates executable fuzzy recommendation tasks, and evaluates agent trajectories using rule-based execution checks and rubric-based LLM evaluation. Experiments on representative LLMs show that syntactically valid tool calls do not guarantee successful recommendations. Models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases. Our results identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems. Our data and code are available at https://github.com/ShawnChenn/RecToolBench.