🤖 AI Summary
Existing dialogue datasets generally lack fine-grained annotations of multi-topic evolution and natural topic transitions, hindering dynamic topic identification and modeling in long conversations. To address this, we propose a controlled dialogue collection paradigm that explicitly supports multi-turn topic emergence and dynamic switching—novel in its design. Leveraging a custom-built instant messaging platform, a structured elicitation protocol, and an intent-aware conversational topic annotation scheme, we construct the first high-quality, topic-analyzed long-dialogue corpus. This corpus features explicit temporal structure, precisely annotated topic boundaries, and fine-grained topic transition types (e.g., shift, continuation, elaboration). It fills a critical empirical data gap in spoken dialogue topic structure research and establishes a robust foundation for topic identification, tracking, and computational modeling.
📝 Abstract
Dialogue is at the core of human behaviour and being able to identify the topic at hand is crucial to take part in conversation. Yet, there are few accounts of the topical organisation in casual dialogue and of how people recognise the current topic in the literature. Moreover, analysing topics in dialogue requires conversations long enough to contain several topics and types of topic shifts. Such data is complicated to collect and annotate. In this paper we present a dialogue collection experiment which aims to build a corpus suitable for topical analysis. We will carry out the collection with a messaging tool we developed.