🤖 AI Summary
Current dialogue systems struggle to effectively recognize Other-Initiated Repair (OIR)—a critical conversational mechanism for resolving communication breakdowns. This work proposes a multimodal approach that, for the first time, systematically incorporates fine-grained visual cues grounded in conversation analysis theory—such as gaze shifts, facial expressions, body posture, and gestures—into OIR detection, integrating them with textual and acoustic modalities within a unified model. Experimental results on two cross-linguistic, cross-domain dialogue corpora demonstrate that the inclusion of visual signals substantially enhances model performance, confirming their consistent and generalizable contribution to multimodal OIR recognition.
📝 Abstract
Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. In the case of conversational agents, it is essential to accurately identify these repair initiation strategies to address communication breakdowns efficiently. While conversational analysis studies have shown that OIR initiation is accompanied by both verbal and non-verbal signals such as gaze shifts, facial expressions, body postures, and hand gestures, existing computational approaches rely mainly on text and audio. This paper introduces a novel multimodal model for OIR detection and classification, incorporating a set of visual features drawn from conversation analysis. We evaluate our approach on two corpora with distinct languages and interaction settings. Results demonstrate that visual information consistently improves performance over text and audio baselines, and provide insights into cross-modal feature contributions across two corpora.