🤖 AI Summary
This paper addresses the lack of systematic evaluation in selecting large language model (LLM) adaptation methods for mental health text analysis. We conduct the first horizontal comparison—within a unified experimental framework—of three prominent strategies: prompt engineering, retrieval-augmented generation (RAG), and supervised fine-tuning, evaluated on emotion classification and psychological disorder identification. Using LLaMA-3 across two public mental health datasets, fine-tuning achieves 91% emotion classification accuracy and 80% disorder detection accuracy; prompt engineering and RAG yield lower accuracies (40–68%) but incur significantly lower deployment costs. Our core contribution is a principled characterization of the trade-offs among performance, computational cost, and adaptability across these methods, culminating in the first clinical-deployment-oriented LLM adaptation selection guideline for mental health AI. This work provides a reproducible, empirically grounded methodology to inform real-world implementation.
📝 Abstract
This study presents a systematic comparison of three approaches for the analysis of mental health text using large language models (LLMs): prompt engineering, retrieval augmented generation (RAG), and fine-tuning. Using LLaMA 3, we evaluate these approaches on emotion classification and mental health condition detection tasks across two datasets. Fine-tuning achieves the highest accuracy (91% for emotion classification, 80% for mental health conditions) but requires substantial computational resources and large training sets, while prompt engineering and RAG offer more flexible deployment with moderate performance (40-68% accuracy). Our findings provide practical insights for implementing LLM-based solutions in mental health applications, highlighting the trade-offs between accuracy, computational requirements, and deployment flexibility.