RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of simultaneously achieving content controllability and low latency in full-duplex spoken dialogue systems by proposing PersonaPlex, a retrieval-based speech playback framework. The method leverages internal hidden states as retrieval queries to precisely invoke pre-recorded audio segments for multi-turn conversations via semantic matching. By streamlining critical network layers, replacing generative modules with a lightweight retrieval head, and integrating an efficient turn-taking mechanism with end-to-end speech modeling, PersonaPlex effectively balances response speed and content controllability. Experimental results demonstrate that the system achieves a median latency of merely 383ms, yielding a 3–7× speedup over cascaded systems of comparable quality. Furthermore, user preference evaluations indicate that PersonaPlex significantly outperforms fast cascaded baselines while approaching the performance of high-quality, slower systems.
📝 Abstract
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and playing pre-recorded lines. Using probing, we identify the layer and frame at which the upcoming response becomes recoverable, and use this hidden state as the retrieval query. RePlay retains only the layers up to that point and replaces text and speech generation with lightweight turn-taking and retrieval heads. In simulated multi-turn interviews, RePlay reaches a median latency of 383 ms, 3 to 7 times lower than ASR-LLM cascades of comparable dialogue quality, at the cost of lower exact-line accuracy. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade with a small LLM (p = 0.008), and showed a non-significant preference (46% vs. 21%) over a slower cascade with a stronger LLM.
Problem

Research questions and friction points this paper is trying to address.

spoken dialogue system
multi-turn conversation
pre-recorded lines
low latency
full-duplex
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Based Voice Playback
Multi-Turn Spoken Dialogue
Hidden State Probing
Lightweight Architecture
Low Latency
🔎 Similar Papers