Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
This study investigates the synchronization mechanisms and turn-taking coordination required for human-like interaction in full-duplex spoken dialogue systems. It introduces the concept of neural coupling to this domain for the first time, simulating conversations between two pretrained Moshi models and measuring cross-lagged representational synchrony using Centered Kernel Alignment (CKA). A causal LSTM is employed to extract turn-prediction signals from delayed activations. Experimental results demonstrate that, under noise-free conditions, the internal states of the models exhibit strong near-zero-lag synchrony and encode turn-taking cues in advance, enabling predictive anticipation of speaker transitions. The prediction performance degrades with increasing acoustic noise, underscoring the critical role of neural synchrony in robust interactive dialogue.