🤖 AI Summary
This study addresses the scalability challenges in psychotherapist training, which heavily relies on expert supervision. To overcome this limitation, the authors propose the first multimodal generative AI simulation environment supporting both speech and text to enable a closed-loop interaction among patient, trainee, and supervisor within cognitive behavioral therapy (CBT). Built upon DSM-5-TR diagnostic criteria, the system integrates multimodal large language models, speech synthesis, affective modeling, and structured CBT protocols to generate 2,100 complete training sessions. Experimental results demonstrate that simulated patients accurately exhibit disorder-relevant emotional features; voice-to-voice models achieve supervisor ratings closest to human performance; and five out of seven models improve diagnostic accuracy following supervisor feedback, with larger models showing greater precision in symptom recognition.
📝 Abstract
Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a multimodal voice- and text-based simulation environment for deliberate practice, used to generate 2,100 complete Cognitive Behavioural Therapy (CBT) training sessions. Each session links a DSM-5-TR-grounded patient (with major depressive, generalised anxiety or borderline personality disorder), a therapist-in-training and an expert supervisor. As an initial implementation, we adopted CBT because its structured procedures and competency-based supervision facilitate standardised simulation and evaluation. Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy. Simulated patients expressed disorder-congruent emotional profiles, which trainee therapists mirrored as in real human counselling. The quality of supervision differed across LLMs: while most models overestimated trainees' competences, native speech-to-speech was closest to human scores. Supervisors' feedback led to better diagnoses in simulated psychotherapists in 5 out of 7 LLMs, and symptom identification accuracy increased with model size. This work shows that simulation of deliberate practice is possible for CBT training, although patient fidelity, calibration of supervisors, and harmful feedback should be evaluated together.