VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks inadequately evaluate LLM agents’ capabilities in real-world scenarios involving large-scale information processing, heterogeneous tool invocation, and dynamic user interaction. To address this, we propose VitaBench—the first benchmark targeting life-oriented complex interactions, covering domains such as dining and transportation, with 66 diverse tools and 400 multi-turn, cross-scenario tasks. Methodologically, we introduce a domain-agnostic task composition framework and a sliding-window evaluator grounded in explicit scoring criteria, enabling robust assessment of heterogeneous solution paths and stochastic interactions. VitaBench integrates multi-tool orchestration, spatiotemporal scheduling reasoning, user intent tracking, and proactive clarification mechanisms, all validated through high-fidelity simulation based on real user requests. Experimental results reveal that state-of-the-art models achieve only 30% success rate on cross-scenario tasks and under 50% on single-scenario ones, exposing critical limitations in practical agent deployment.

Technology Category

Multiagent Systems: Modeling other AgentsPlanning, Routing, and Scheduling: Planning with Language ModelsCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic searchEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
📝 Abstract
As LLM-based agents are increasingly deployed in real-life scenarios, existing benchmarks fail to capture their inherent complexity of handling extensive information, leveraging diverse resources, and managing dynamic user interactions. To address this gap, we introduce VitaBench, a challenging benchmark that evaluates agents on versatile interactive tasks grounded in real-world settings. Drawing from daily applications in food delivery, in-store consumption, and online travel services, VitaBench presents agents with the most complex life-serving simulation environment to date, comprising 66 tools. Through a framework that eliminates domain-specific policies, we enable flexible composition of these scenarios and tools, yielding 100 cross-scenario tasks (main results) and 300 single-scenario tasks. Each task is derived from multiple real user requests and requires agents to reason across temporal and spatial dimensions, utilize complex tool sets, proactively clarify ambiguous instructions, and track shifting user intent throughout multi-turn conversations. Moreover, we propose a rubric-based sliding window evaluator, enabling robust assessment of diverse solution pathways in complex environments and stochastic interactions. Our comprehensive evaluation reveals that even the most advanced models achieve only 30% success rate on cross-scenario tasks, and less than 50% success rate on others. Overall, we believe VitaBench will serve as a valuable resource for advancing the development of AI agents in practical real-world applications. The code, dataset, and leaderboard are available at https://vitabench.github.io/
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM agents on complex real-world interactive tasks
Assessing agent performance across dynamic multi-turn conversations
Measuring success in cross-domain scenarios with diverse tools
Innovation

Methods, ideas, or system contributions that make the work stand out.

VitaBench benchmark tests versatile interactive real-world tasks
Framework enables flexible composition of scenarios and tools
Rubric-based sliding window evaluator assesses diverse solution pathways
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wei He
Y
Yueqing Sun
H
Hongyan Hao
X
Xueyuan Hao
Z
Zhikang Xia
Q
Qi Gu
Chengcheng Han
Chengcheng Han
Meituan | East China Normal University
NLPKG
D
Dengchang Zhao
H
Hui Su
K
Kefeng Zhang
M
Man Gao
X
Xi Su
X
Xiaodong Cai
X
Xunliang Cai
Y
Yu Yang
Y
Yunke Zhao