🤖 AI Summary
This study addresses the challenge of heterogeneous and annotation-limited turn-taking control in real-time dialogue systems by proposing XTurnix, a unified framework built upon AI-state-based two-state binary decision-making. Employing a compact text architecture, the method formulates turn-taking as a causal decision process that predicts a single control token conditioned on the complete dialogue history. Efficient learning is achieved through self-supervised pre-training on 5.5 million samples followed by fine-tuning with multi-turn synthetic data. Experimental evaluations demonstrate that XTurnix achieves state-of-the-art or comparable performance across multiple public benchmarks, while attaining an accuracy of 89.06% on a newly constructed benchmark, significantly outperforming existing baseline methods.
📝 Abstract
General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them difficult to use as a unified causal controller with comprehensive context. We propose XTurnix, a compact text-based model that formulates turn control as two binary decisions conditioned on the AI's current listening or speaking state and predicts a single control token from the complete dialogue history. XTurnix is pretrained on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, then fine-tuned on synthetic multi-turn examples with a flatter distribution across the four state-action labels. We evaluate XTurnix on four public benchmarks and a balanced self-curated benchmark. Across the public benchmarks, XTurnix achieves the best results on all SemanticVAD and LiveKit splits, ties the native Smart-Turn model on Smart-Turn Bench, and achieves the highest incomplete-turn accuracy on Easy-Turn. On the self-curated benchmark, it reaches 89.06% accuracy, more than 20 percentage points above the strongest third-party baseline at 68.75%, while maintaining F1 scores between 84.21% and 90.63% across all four categories. These results demonstrate unified listening- and speaking-state turn control in a single compact model. Code is available at https://github.com/xcc-zach/xturnix, with an interactive demo at https://huggingface.co/spaces/xcczach/xturnix-demo.