XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of heterogeneous and annotation-limited turn-taking control in real-time dialogue systems by proposing XTurnix, a unified framework built upon AI-state-based two-state binary decision-making. Employing a compact text architecture, the method formulates turn-taking as a causal decision process that predicts a single control token conditioned on the complete dialogue history. Efficient learning is achieved through self-supervised pre-training on 5.5 million samples followed by fine-tuning with multi-turn synthetic data. Experimental evaluations demonstrate that XTurnix achieves state-of-the-art or comparable performance across multiple public benchmarks, while attaining an accuracy of 89.06% on a newly constructed benchmark, significantly outperforming existing baseline methods.
📝 Abstract
General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them difficult to use as a unified causal controller with comprehensive context. We propose XTurnix, a compact text-based model that formulates turn control as two binary decisions conditioned on the AI's current listening or speaking state and predicts a single control token from the complete dialogue history. XTurnix is pretrained on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, then fine-tuned on synthetic multi-turn examples with a flatter distribution across the four state-action labels. We evaluate XTurnix on four public benchmarks and a balanced self-curated benchmark. Across the public benchmarks, XTurnix achieves the best results on all SemanticVAD and LiveKit splits, ties the native Smart-Turn model on Smart-Turn Bench, and achieves the highest incomplete-turn accuracy on Easy-Turn. On the self-curated benchmark, it reaches 89.06% accuracy, more than 20 percentage points above the strongest third-party baseline at 68.75%, while maintaining F1 scores between 84.21% and 90.63% across all four categories. These results demonstrate unified listening- and speaking-state turn control in a single compact model. Code is available at https://github.com/xcc-zach/xturnix, with an interactive demo at https://huggingface.co/spaces/xcczach/xturnix-demo.
Problem

Research questions and friction points this paper is trying to address.

turn-taking
real-time dialogue systems
turn control
turn detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Turn Control
Self-Supervised Learning
Binary Decisions
Dialogue Systems
Causal Controller
🔎 Similar Papers
2024-03-14IEEE International Conference on Robotics and AutomationCitations: 3
Z
Zhanxun Liu
MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University
Y
Yifan Duan
MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University
H
Hengtao Wu
MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University
C
Chen Yang
Shanghai Innovation Institute
Q
Qinyuan Cheng
Shanghai Innovation Institute
K
Kun Wang
SenseTime Group Inc.
Xingyu Zeng
Xingyu Zeng
Shenzhen University of Advanced Technology
Computer VisionDeep Learning
X
Xipeng Qiu
Shanghai Innovation Institute
Chaochao Lu
Chaochao Lu
Shanghai AI Laboratory
Causal AI
Xie Chen
Xie Chen
Shanghai Jiao Tong University <- Microsoft <- Cambridge University
Machine LearningSpeech RecognitionSpeech SynthesisSpeech&Audio Processing