Separating Memory and Workflow Effects in Predicting Individual Answers

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the confounding effects of memory construction and workflow selection on individual answer prediction in personalized language agents. To disentangle these factors, we propose OwnWords, a framework that decouples memory content from generation strategies by retrieving users’ original utterances via BM25 and generating responses within a single inference call. We systematically compare specific memories, trait descriptions, and various fusion strategies for predicting unseen answers. Empirical results demonstrate that OwnWords significantly outperforms conventional written-memory baselines on external test sets, with unaltered source records yielding optimal performance; however, no substantial improvement is observed on precise multiple-choice tasks. This work establishes a new paradigm for designing memory mechanisms in personalized agents.
πŸ“ Abstract
Personalized language agents choose both what to remember about a person and how to use that memory. We separate these choices when predicting unseen answers to known interview questions. On 1,768 tasks from 188 people, a concrete memory built from a verified interview prefix outscores a trait description by 0.0158 (95% whole-person interval [0.0044, 0.0271]). Crossing both memories with one-shot generation and three-answer fusion, fusion lowers concrete-memory scores by 0.0123 ([-0.0189, -0.0056]); prompted and trained selectors do not detectably beat a random candidate. One call on the longer, unrewritten source record outscores every memory condition. Under a limited context budget, OwnWords retrieves the person's sentences with BM25 and answers in one call. It outperforms the written memory on 500 people outside the benchmark (+0.0127, [+0.0037, +0.0217]; an earlier held-out test was inconclusive) and across four budgets on 300 people (mean +0.0218, [+0.0138, +0.0298]), with the latter result repeated on 114 people. It does not detectably outperform recency truncation. These results compare evidence-construction procedures; they do not isolate the effect of verbatim wording. On Twin-2K-500, OwnWords predicts ordinal survey answers more closely than the written memory, but does not improve exact-choice accuracy and lowers it in one of two samples. Interview scores use a model-based content rubric without human ratings, and the original benchmark's participants were seen during development. These results characterize the tested procedures, not a general human-prediction ceiling.
Problem

Research questions and friction points this paper is trying to address.

personalized language agents
memory effects
workflow effects
individual answer prediction
context budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Personalized language agents
Memory-workflow separation
BM25 retrieval
Context budget
OwnWords
πŸ’Ό Related Jobs
No related jobs found.
T
Tianzhu Qin
University of Cambridge
L
Leo Yang Yang
Hong Kong Baptist University
L
Lee Wei Jun
Cookiy Labs
K
Kun Chen
Cookiy Labs
Ramit Debnath
Ramit Debnath
Assistant Professor and Deputy Director, Centre for Human-Inspired AI, University of Cambridge
Climate ActionComputational social scienceAI designAI for SustainabilityEnvironment
D
Davin Youchao Dong
Cookiy Labs