🤖 AI Summary
Existing benchmarks struggle to evaluate personal AI assistants’ ability to handle user states, permissions, and long-term multi-turn interactions within realistic, unified service environments. To address this gap, this work proposes PAUSE—a user-centered, state-aware evaluation benchmark that introduces, for the first time, a user-centric simulated interaction mechanism coupled with a multi-paradigm assessment framework. PAUSE supports behavioral metrics for open-ended tasks and state validation for constrained tasks, integrating user state modeling, multi-service authorization reasoning, semantic and trajectory-level evaluation, deterministic state verification, and synthetic data generation. Experiments reveal that even state-of-the-art models achieve less than 70% task completion on state-reasoning and configuration-aware tasks, exposing systematic failure modes and demonstrating PAUSE’s effectiveness and challenge in evaluating real-world assistant capabilities.
📝 Abstract
Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.