🤖 AI Summary
This study addresses the scarcity of annotated dialogue interaction resources in clinical cognitive assessment, which constrains research on physician–patient interactions. We construct the first de-identified corpus encompassing 56 dialogue acts, providing the initial quantification of interaction structures within this context. Using LLaMA-3.1 as the backbone, we establish benchmarks through instruction tuning and explanation-augmented training. Experimental results demonstrate that instruction tuning substantially improves fine-grained classification accuracy and patient response matching. Nevertheless, the model continues to struggle with distinguishing closely related dialogue acts, highlighting inherent challenges in fine-grained intent recognition within clinical settings. Overall, this work validates the effectiveness of transferring such fine-tuning strategies to clinical applications while identifying key directions for future improvement.
📝 Abstract
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.