FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language model (LLM) agents can effectively maintain and dynamically update personalized user models during extended interactions, with a particular focus on their capacity for event-driven preference adaptation in high-stakes domains such as financial advising. To this end, we introduce the first benchmark that integrates real investor trajectories with theoretically grounded shock events, evaluating LLM performance across 2,994 questions spanning 276 personas through controlled narrative generation, automated filtering, and diverse memory mechanisms—including summarization and retrieval. Results reveal that current systems achieve an overall accuracy of only ~0.47, with multiple-choice accuracy falling below 39%, and simple retrieval often outperforming specialized memory modules—highlighting a critical gap in integrating post-event preference shifts. Our work establishes a longitudinal evaluation paradigm centered on post-event changes, offering a new benchmark and key insights for personalized memory research.
📝 Abstract
Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.
Problem

Research questions and friction points this paper is trying to address.

personalized memory
LLM agents
event-driven adaptation
user modeling
longitudinal evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

personalized memory
event-grounded benchmark
LLM agents
preference adaptation
longitudinal user modeling
🔎 Similar Papers