🤖 AI Summary
This study addresses the problem of "empty promises," wherein AI agents fail to fulfill commitments due to runtime constraints. It formally defines empty promise semantics and introduces three failure modes alongside their anchoring conditions. Methodologically, this work pioneers a configuration-based static detection mechanism capable of determining promise validity without relying on subsequent execution trajectories. Furthermore, it establishes a multi-environment comparative experimental framework and measurement protocol to systematically evaluate promise feasibility under diverse persistence configurations. The findings reveal the decisive influence of tooling and runtime configurations on promise fulfillment, providing quantitative evaluation criteria for assessing agent reliability.
📝 Abstract
A chatbot that says "I will remind you tomorrow" will not run again until the user writes. We call such a promise an empty commitment: a promise of an action after the current turn that nothing in the agent's tools or runtime can carry out. Unlike a broken promise, its emptiness follows from the agent's configuration alone; no later trajectory is needed. We define empty commitments on top of commitment semantics, with three failure types, an anchoring condition for promises that a tool could make real, and a response-level outcome taxonomy. We then describe a measurement protocol: follow-up requests run in five setups that add one persistence affordance at a time, with the environment either left implicit or stated.