🤖 AI Summary
This study addresses the absence of unified evaluation standards for cross-platform personal intelligence and its susceptibility to over-personalization by proposing the first comprehensive, all-platform benchmark. Leveraging millions of anonymized interactions, we construct a temporally indexed digital world that integrates psychological and sociolinguistic theories. By combining large language models (LLMs) with multi-source heterogeneous data fusion techniques, our approach achieves holistic, cross-contextual user understanding, translating implicit signals into explicit intents. Furthermore, we establish an integrated evaluation framework encompassing personalization, recommendation, proactive behavior, and temporal reasoning. Our experiments validate both the effectiveness and alignment of LLM-driven agents operating atop existing recommendation infrastructures, offering a robust foundation for advancing personalized intelligent systems across diverse platforms.
📝 Abstract
Personal intelligence is becoming a central frontier for user-facing AI agents. To be helpful in everyday life, agents must understand users across the digital contexts where their preferences, intents, habits, social relationships, and needs unfold over time. Today's systems can personalize within individual apps or tasks, but personal intelligence as a whole remains under-measured: how agents build cross-context user understanding, support steerable recommendation systems, act proactively across platforms, and avoid over-personalization. We introduce PersonaMem-v3, a real-world-grounded benchmark and evaluation harness for omni-platform personal intelligence. PersonaMem-v3 is seeded from more than one million anonymized real-world engagement histories, most of which are implicit signals, and uses them to construct time-indexed user digital worlds across social media, chatbot, calendar, and AI-companion with preference evolvement over time. The benchmark brings personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning into one framework, anchored in psychology, social-linguistics, and user-behavior theories. It evaluates whether AI agents can infer holistic user understanding from cross-platform evidence, personalize responses, rerank recommendations on social media, follow user steering through natural language, and hold back when personalization would be inappropriate, repetitive, outdated, or unnecessary. PersonaMem-v3 points toward LLM-powered personal intelligent agents that work with existing scalable recommendation infrastructure while making personalization more interactive, agentic, and aligned with how real users experience their digital lives.