A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG

📅 2025-03-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the lack of systematic evaluation in selecting large language model (LLM) adaptation methods for mental health text analysis. We conduct the first horizontal comparison—within a unified experimental framework—of three prominent strategies: prompt engineering, retrieval-augmented generation (RAG), and supervised fine-tuning, evaluated on emotion classification and psychological disorder identification. Using LLaMA-3 across two public mental health datasets, fine-tuning achieves 91% emotion classification accuracy and 80% disorder detection accuracy; prompt engineering and RAG yield lower accuracies (40–68%) but incur significantly lower deployment costs. Our core contribution is a principled characterization of the trade-offs among performance, computational cost, and adaptability across these methods, culminating in the first clinical-deployment-oriented LLM adaptation selection guideline for mental health AI. This work provides a reproducible, empirically grounded methodology to inform real-world implementation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Adaptive Behavior

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
This study presents a systematic comparison of three approaches for the analysis of mental health text using large language models (LLMs): prompt engineering, retrieval augmented generation (RAG), and fine-tuning. Using LLaMA 3, we evaluate these approaches on emotion classification and mental health condition detection tasks across two datasets. Fine-tuning achieves the highest accuracy (91% for emotion classification, 80% for mental health conditions) but requires substantial computational resources and large training sets, while prompt engineering and RAG offer more flexible deployment with moderate performance (40-68% accuracy). Our findings provide practical insights for implementing LLM-based solutions in mental health applications, highlighting the trade-offs between accuracy, computational requirements, and deployment flexibility.
Problem

Research questions and friction points this paper is trying to address.

Compare LLM approaches for mental health text analysis
Evaluate accuracy and resource needs for emotion classification
Assess trade-offs between performance and deployment flexibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuning achieves highest accuracy but resource-intensive
Prompt engineering offers flexible deployment with moderate performance
RAG balances performance and computational efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arshia Kermani
Department of Computer Science, Texas State University
V
Veronica Perez-Rosas
Department of Computer Science, Texas State University
Vangelis Metsis
Vangelis Metsis
Texas State University
Machine LearningComputer VisionPervasive ComputingAffective ComputingSmart Health