PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to evaluate personal AI assistants’ ability to handle user states, permissions, and long-term multi-turn interactions within realistic, unified service environments. To address this gap, this work proposes PAUSE—a user-centered, state-aware evaluation benchmark that introduces, for the first time, a user-centric simulated interaction mechanism coupled with a multi-paradigm assessment framework. PAUSE supports behavioral metrics for open-ended tasks and state validation for constrained tasks, integrating user state modeling, multi-service authorization reasoning, semantic and trajectory-level evaluation, deterministic state verification, and synthetic data generation. Experiments reveal that even state-of-the-art models achieve less than 70% task completion on state-reasoning and configuration-aware tasks, exposing systematic failure modes and demonstrating PAUSE’s effectiveness and challenge in evaluating real-world assistant capabilities.
📝 Abstract
Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.
Problem

Research questions and friction points this paper is trying to address.

personal AI assistants
user-centric evaluation
unified service environments
stateful reasoning
multi-turn interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

user-centric benchmark
stateful reasoning
service-integrated environments
multi-turn interaction
synthetic task generation
🔎 Similar Papers
No similar papers found.
H
Haoyu Chen
Electrical and Computer Engineering, University of Alberta, Edmonton, Alberta, Canada
X
Xirui Shi
Electrical and Computer Engineering, University of Alberta, Edmonton, Alberta, Canada
Y
Yuyao Wang
Electrical and Computer Engineering, University of Alberta, Edmonton, Alberta, Canada
J
Jerry Chen
Electrical and Computer Engineering, University of Alberta, Edmonton, Alberta, Canada
Di Niu
Di Niu
Professor, University of Alberta
Deep learningDistributed systemsParallel ComputingComputer VisionNLP