PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent memory challenges faced by voice assistants in multi-user, cross-session scenarios, specifically regarding speaker identification, historical state restoration, and instruction revision. To tackle these issues, this work proposes a persistent memory system that explicitly models "Who-What-When" information by structuring cross-session history into event records. A 3W joint scoring mechanism is designed to integrate semantic, voiceprint, and temporal features for precise retrieval. Furthermore, intermediate representations from the dialogue backbone are reused to eliminate redundant re-encoding, substantially reducing latency. The SpokenTrace diagnostic benchmark is also introduced for evaluation. Experimental results demonstrate an end-to-end accuracy of 85.08%, with EM@3 improving from 49.01% to 82.10%, while retrieval latency decreases dramatically from 578.42ms to 7.03ms.
📝 Abstract
Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.
Problem

Research questions and friction points this paper is trying to address.

spoken dialogue
cross-session memory
speaker identification
full-duplex interaction
multi-speaker
Innovation

Methods, ideas, or system contributions that make the work stand out.

Persistent Memory
Full-Duplex Spoken Dialogue
3W Joint Scoring
Latency Optimization
Diagnostic Benchmark
🔎 Similar Papers
No similar papers found.
A
Achira Lin
Tsinghua University
S
Siyuan Hou
Tsinghua University
W
Wenyi Yu
Tsinghua University
X
Xinnian Zhao
Tsinghua University
Haoyu Niu
Haoyu Niu
Fudan University
Data privacy
W
Wang Geng
Huawei Technologies Ltd.
L
Longshuai Xiao
Huawei Technologies Ltd.
S
Shihai Xiao
Huawei Technologies Ltd.
M
Mangsuo Zhao
Tsinghua University
Chao Zhang
Chao Zhang
Tsinghua University
software and system securityAI for securityblockchaindata security