Controlling Backchannels in Streamable Full-duplex Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unnatural interactions in full-duplex spoken dialogue models caused by the lack of explicit modeling of feedback signals. To this end, it proposes a lightweight feedback head mechanism based on hidden state probing. Employing a fully dual-stream architecture, the method utilizes threshold-triggered forced decoding to precisely predict feedback timing, thereby achieving controllable generation. The core innovation lies in leveraging internal model representations to anticipate human conversational rhythms, with cross-model-scale generalization empirically validated. Experimental results demonstrate that the proposed approach significantly optimizes both the frequency and timing of feedback. Furthermore, human evaluations indicate that the generated responses achieve quality comparable to authentic human reactions.
📝 Abstract
Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Full-duplex dialogue
Backchannel modeling
Force-decoding
Hidden state probing
Lightweight prediction head
🔎 Similar Papers
No similar papers found.