🤖 AI Summary
This study addresses the persistent memory challenges faced by voice assistants in multi-user, cross-session scenarios, specifically regarding speaker identification, historical state restoration, and instruction revision. To tackle these issues, this work proposes a persistent memory system that explicitly models "Who-What-When" information by structuring cross-session history into event records. A 3W joint scoring mechanism is designed to integrate semantic, voiceprint, and temporal features for precise retrieval. Furthermore, intermediate representations from the dialogue backbone are reused to eliminate redundant re-encoding, substantially reducing latency. The SpokenTrace diagnostic benchmark is also introduced for evaluation. Experimental results demonstrate an end-to-end accuracy of 85.08%, with EM@3 improving from 49.01% to 82.10%, while retrieval latency decreases dramatically from 578.42ms to 7.03ms.
📝 Abstract
Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.