TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the frequent neglect of traditional medical systems—such as Tibetan medicine—by large language models in healthcare applications, which engenders cultural bias and ontological drift. To counter this, the authors introduce TreeProbe, the first culturally grounded evaluation benchmark for Tibetan medicine, constructed from the classical Tibetan medical text *The Four Tantras* (specifically its *Medical Tree Metaphor* section). TreeProbe encompasses 467 diseases and 10 subtasks, uniquely organizing evaluation content within the indigenous Tibetan medical knowledge framework. Through expert-annotated data, multidimensional subtask design, and model bias analysis, the work reveals systematic ontological drift in mainstream large language models when reasoning about Tibetan medicine, distinguishing whether such drift aligns more closely with biomedicine or Traditional Chinese Medicine. TreeProbe thus provides a critical diagnostic tool for developing linguistically inclusive and cognitively equitable medical AI systems.
📝 Abstract
Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional medical knowledge systems. Tibetan medicine, one of the world's four major traditional medical systems, has an independent and highly structured theoretical framework. When models lack grounded understanding of Tibetan medicine, they may fall back on dominant epistemic systems and distort the native knowledge structure during reasoning. However, quantitative tools for evaluating cultural bias in Tibetan medicine remain largely absent. To address this gap, we introduce TreeProbe, the first cultural-bias benchmark organized around the native Tree of Medicine framework in Tibetan medicine. It contains 4,719 expert-adjudicated items covering 467 diseases and 10 subtasks along the three roots. Experiments on representative LLMs show that current models remain limited in native Tibetan medical contexts and exhibit systematic external ontology drift. Further analysis reveals that models diverge in whether they drift toward biomedical or TCM reasoning, shaped by pretraining data composition and surface resemblance between TCM and Tibetan medicine. TreeProbe provides a diagnostic benchmark for developing medical AI systems that are both linguistically inclusive and epistemically fair. Code and data are available in an anonymous repository at https://anonymous.4open.science/r/TreeProbe/.
Problem

Research questions and friction points this paper is trying to address.

cultural bias
Tibetan medicine
large language models
epistemic fairness
traditional medical systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

cultural bias
Tibetan medicine
LLM evaluation
epistemic fairness
Tree of Medicine