🤖 AI Summary
This study addresses the inefficiency in psychological triage caused by vague email subject lines in online mental health counseling. It proposes a hierarchical evaluation framework to classify and rank within-category six-word German subject lines generated by eleven large language models, integrating assessments from both professional counselors and AI scorers to systematically evaluate their practical utility in mental health contexts. The work presents the first comparative analysis of closed-source and open-source models on German-language mental health tasks, examining performance and ethical trade-offs, and employs Krippendorff’s α and Spearman’s ρ to quantify inter-rater agreement and score correlations. Results demonstrate that German-specific fine-tuning substantially enhances model performance, offering empirical support for deploying AI systems in mental health settings that balance privacy and efficacy.
📝 Abstract
Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's $\alpha$, Spearman's $\rho$, Pearson's $r$ and Kendall's $\tau$. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.