Comparative Analysis of Large Language Models in Generating Telugu Responses for Maternal Health Queries

📅 2026-03-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study presents the first systematic evaluation of mainstream large language models—ChatGPT-4o, Gemini, and Perplexity—in answering maternal health questions in Telugu, a low-resource language. The authors construct a bilingual question-answering dataset and assess model performance through a combination of BERTScore for semantic similarity and expert human evaluations by obstetricians across dimensions including accuracy, fluency, and relevance. Results indicate that Gemini achieves the strongest overall performance in Telugu generation, Perplexity performs well when prompted directly in Telugu, and ChatGPT-4o still exhibits room for improvement. The findings underscore the critical influence of both model selection and prompt language on the quality of maternal health information delivery in low-resource linguistic settings, thereby addressing a significant gap in systematic research on this topic.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Knowledge Representation Languages

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Large Language Models (LLMs) have been progressively exhibiting there capabilities in various areas of research. The performance of the LLMs in acute maternal healthcare area, predominantly in low resource languages like Telugu, Hindi, Tamil, Urdu etc are still unstudied. This study presents how ChatGPT-4o, GeminiAI, and Perplexity AI respond to pregnancy related questions asked in different languages. A bilingual dataset is used to obtain results by applying the semantic similarity metrics (BERT Score) and expert assessments from expertise gynecologists. Multiple parameters like accuracy, fluency, relevance, coherence and completeness are taken into consideration by the gynecologists to rate the responses generated by the LLMs. Gemini excels in other LLMs in terms of producing accurate and coherent pregnancy relevant responses in Telugu, while Perplexity demonstrated well when the prompts were in Telugu. ChatGPT's performance can be improved. The results states that both selecting an LLM and prompting language plays a crucial role in retrieving the information. Altogether, we emphasize for the improvement of LLMs assistance in regional languages for healthcare purposes.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Maternal Health
Low-resource Languages
Telugu
Healthcare AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Maternal Health
Low-resource Languages
Telugu
BERT Score
💼 Related Jobs
No related jobs found.
A
Anagani Bhanusree
National Institute of Technology, Warangal, India
S
Sai Divya Vissamsetty
National Institute of Technology, Warangal, India
K
K VenkataKrishna Rao
National Institute of Technology, Warangal, India
Rimjhim
Rimjhim
University of Bocconi
Computational Social ScienceSocial ComputingGeographic Information Systems