Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts

๐Ÿ“… 2026-10-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of culturally situated reasoning in existing evaluations of medical large language models, particularly within Indian maternal and infant health contexts. We propose MH-INDIC, a framework that assesses the cultural alignment of maternalโ€“infant interactions in North India through ten-dimensional metrics. Methodologically, MH-INDIC decouples population-level cultural alignment from individual behavioral variation, conducting systematic evaluations via a 26-item questionnaire, multi-model comparisons, and human-grounded prompt engineering. Our findings reveal that all evaluated models exhibit lower individual variability than human baselines and demonstrate insufficient sensitivity to demographic attributes. Furthermore, incorporating human-grounded conditions significantly enhances both dialogue quality and cultural congruence.
๐Ÿ“ Abstract
Existing evaluation methods for healthcare LLMs primarily assess factual correctness,safety, and fluency, while providing limited insight into whether generated interactions reflect culturally situated healthcare reasoning. This limitation is particularly important in maternal health, where care decisions are shaped by social and relational norms. We introduce MH-INDIC, a culturally grounded evaluation framework for maternal-health interactions in urban and semi-urban North Indian contexts that operationalises cultural behaviour through ten dimensions of maternal-health reasoning. Using a 26-item survey administered to 102 pregnant and postpartum women from urban and semi-urban North India, we evaluate ten LLMs. We distinguish population level cultural alignment from profile-level behavioural variation. Although several models approximate the human population-level distribution, all evaluated systems exhibit substantially lower variation across demographic and household profiles than the human cohort, revealing a gap between aggregate alignment and profile-conditioned sensitivity. As a downstream application of MH-INDIC, we use the strongest-aligned proprietary and open-source models to generate culturally conditioned maternal-health dialogues under zero-shot, self-conditioned, and human-grounded prompting. Human-grounded conditioning produces stronger profile alignment and dialogue quality ratings, suggesting that measured cultural profiles can improve the cultural grounding of generated interactions
Problem

Research questions and friction points this paper is trying to address.

Cultural Alignment
Maternal Health
Large Language Models
Evaluation Framework
Healthcare Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cultural Alignment
Evaluation Framework
Maternal Health LLM
Profile-level Variation
Human-grounded Prompting
๐Ÿ’ผ Related Jobs
No related jobs found.
U
Umaira Izhar
Indraprastha Institute of Information Technology Delhi
G
Gunjan Arora
Indraprastha Institute of Information Technology Delhi
Pushpendra Singh
Pushpendra Singh
IIIT Delhi
Human Computer Interaction