π€ AI Summary
This study addresses the limitations of large language models in psychiatric interviewing and diagnostic reasoning by proposing a novel dual-agent framework that decouples free-form clinical interviews from expert-level diagnostic review. The architecture leverages ICD-11 retrieval, working note maintenance, and PsyCPG-based patient simulation integrated with tool calling to facilitate dynamic dialogue guidance and evidence-based assessment. Experimental results demonstrate that this approach significantly outperforms direct prompting baselines in diagnostic agreement rates. Furthermore, in blind evaluations, the systemβs hypothesis generation achieves 79.6% concordance with human psychologists. These findings indicate that the proposed interview-reasoning separation effectively enhances the reliability of AI-assisted psychiatric diagnosis.
π Abstract
Large language models show promise in clinical reasoning, but psychiatric interviewing requires guiding an evolving conversation. Their ability to carry out this interactive assessment remains less studied. We present PsyCIDRA, a dual-agent framework linking free-form psychiatric interviewing with diagnostic reasoning for expert review. Its interviewer agent uses tools to maintain working notes, load expert-written skills, and retrieve ICD-11 references to guide inquiry. Its diagnostic reasoning agent then receives the completed interview transcript and reports hypotheses alongside supporting, conflicting, and missing evidence, withholding a final hypothesis when none is sufficiently supported. Using patient profiles generated with PsyCPG, we first evaluate PsyCIDRA in simulation. Across four models on 53 evaluation cases, it achieves higher diagnostic agreement than direct prompting. On 81 held-out simulated cases, rank-1 accuracy is 60.5% versus 51.9%. In a blinded study of 101 human participants in separate arms, PsyCIDRA agrees with psychologists on whether to propose a diagnostic hypothesis in 79.6% of cases, compared with 65.4% for direct prompting. Together, these findings support the potential of LLM agents to assist psychiatric assessment through free-form dialogue. By examining diagnostic reasoning, interview quality, and safety together, this study contributes to understanding the capabilities and limitations of psychiatric interview agents.