TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of large-scale, expert-aligned evaluation benchmarks in Traditional Chinese Medicine (TCM) by constructing an open benchmark comprising 38,000 licensing examination questions and 15,000 physician responses, systematically evaluating 29 models. The findings reveal that pretraining corpora exert a substantially greater influence on TCM capabilities than model scale, with Chinese-pretrained models significantly outperforming Western counterparts of comparable size. Although certain models surpass physician majority-vote accuracy, notable discrepancies emerge between models and physicians regarding perceived question difficulty. By providing the first large-scale TCM evaluation framework paired with real-world expert references, this work highlights the critical role of linguistic background in shaping the performance of domain-specific large language models.
📝 Abstract
Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rarely paired with a human reference. We present TCMQA, an open benchmark of 38,279 questions from Chinese TCM licensing examinations, paired with 15,151 responses from 101 licensed practitioners. We evaluate 29 instruction-tuned models from 9 families, spanning 0.27B to 14.8B parameters. Accuracy ranges over 59 points, and no model approaches saturation. Pretraining data predicts TCM ability far better than scale: a 12B Western-pretrained model reaches 39.6%, while a Chinese-pretrained model an eighth its size reaches 60.8%. Nine models exceed the practitioner majority vote of 64.9%, the best by 21.8 points, and all nine come from that same Chinese-pretrained family. Yet difficulty does not transfer between models and practitioners: accuracy is flat across practitioner-rated difficulty, item-level agreement is near zero for all 29 models, and on $8.4\%$ of items the practitioners are correct where the leading model is wrong. We release the corpus, the practitioner responses, the harness, and per-item model outputs at https://huggingface.co/datasets/TechTCM/TCMQA.
Problem

Research questions and friction points this paper is trying to address.

Traditional Chinese Medicine
Medical Benchmark
Large Language Models
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Traditional Chinese Medicine Benchmark
Large Language Model Evaluation
Licensed-Practitioner Reference
Pretraining Data Distribution
Instruction-Tuned Models
🔎 Similar Papers
No similar papers found.