Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of proactive LLM agents misinterpreting contexts, increasing review costs, and undermining user trust by proposing the 3T principle, which jointly optimizes Task capability, Time allocation, and Trust. Methodologically, it introduces the first 3T joint optimization framework, constructs a five-dimensional design space with a systematic model, and develops Proactivity-Gym, a simulation benchmark integrating multi-day scenarios and personalized simulated users. The research reveals significant performance gaps among existing models across 3T metrics and identifies biases in LLM-based evaluation. Furthermore, human subject experiments validate that joint optimization is critical for sustaining user trust.
📝 Abstract
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.
Problem

Research questions and friction points this paper is trying to address.

proactive agents
large language models
user trust
temporal allocation
task capability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proactive Agents
3T Principles
PROACTIVITY-GYM
Simulation-based Evaluation
Trust Alignment