🤖 AI Summary
This study addresses the challenge that self-descriptions of AI companions and user experience ratings fail to verify the authenticity of their underlying mechanisms. We propose the first actionable auditing framework that cross-references agent self-reports, user judgments, and system implementation logs. By integrating behavioral analysis, user feedback, and log inspection for multi-source data triangulation, this framework systematically evaluates the evidential support for each claimed capability. Empirical findings reveal that certain unexecuted memory layers still receive high user ratings, exposing a significant disconnect between fluent self-descriptions and practically ineffective mechanisms. This work establishes a methodological foundation for transparently evaluating the capability boundaries of AI systems, underscoring the critical role of evidence visibility in validating their actual functionality.
📝 Abstract
Companion agents describe themselves: they remember, they understand their users, the relationship has changed them. We argue that such accounts, and the experience ratings that seem to confirm them, are checkable by users only where the evidence is theirs: in the agent's behavior, or in themselves. Where the evidence lives in the machinery, fluent self-description and moderately positive ratings do not establish that the mechanisms behind them ran. We demonstrate an audit procedure that sets an agent's self-description against its users' judgements and its implementation records, reporting each claim as supported, contradicted, or unresolved, and apply it to Lita, a proactive companion we built and deployed for a month with nine colleagues. Participants endorsed stylistic claims, withheld endorsement from relational ones, and rated memory at or above midpoint, while two of three memory layers had never executed their accumulation step. Memory-bearing agents should report what their self-descriptions cannot establish.