TCMClinicalReason-Bench: Can Language Models Reason from Pathogenesis to Prescription over Real-World Clinical Cases?

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of evidence grounding and the cascading propagation of errors in clinical reasoning for Traditional Chinese Medicine (TCM) within large language models. To this end, it constructs a full-pipeline evaluation benchmark derived from real-world, multicenter electronic health records. Methodologically, the work proposes an evidence-constrained scoring mechanism that decouples component-level quality from cross-module logical consistency to precisely localize reasoning bottlenecks, while integrating automated judge models with expert blind reviews to ensure evaluation reliability. The findings reveal that general-purpose models outperform domain-specific counterparts and identify significant deficiencies in prescription generation. These insights provide critical empirical foundations for enhancing the robustness of TCM-oriented AI reasoning systems.
📝 Abstract
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinicalReason-Bench using 2,000 multicenter electronic health record cases to distinguish case-grounded responses from fluent but unsupported diagnostic and therapeutic conclusions. Five general-purpose and two TCM-specific LLMs were evaluated in zero-shot settings. An evidence-constrained rubric assessed seven diagnostic and therapeutic components and three cross-block relations, allowing case-supported alternatives. Qwen3.7-Plus with TCM retrieval served as the automated judge, alongside parallel blinded ratings by five senior TCM clinicians on a 600-case subset. Structural completeness was nearly saturated (99.3-100.0%), but normalized content scores ranged from 40.7% to 54.1%. The five general-purpose models averaged 50.0%, versus 41.3% for the two smaller TCM-specific models. Cross-block logic consistency ranged from 60.8% to 66.8% and correlated moderately with content across cases (Pearson's r = 0.515-0.656). Deficits were greatest in prescription generation, prescription analysis, and symptom-guided modification. In judge stress testing on 100 independent cases, perturbation detection rates across the three relations were 57%, 56%, and 31%, with contradictions detected more reliably than omissions. Separating component quality from cross-block consistency localizes failures missed by endpoint and completeness metrics and identifies where clinician oversight remains necessary.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Traditional Chinese Medicine
Clinical Reasoning
Evidence Grounding
Evaluation Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clinical Reasoning Benchmark
Evidence-Constrained Evaluation
Cross-Block Logic Consistency
Automated LLM Judge
Traditional Chinese Medicine
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jirui Dai
the Department of Computer Science, Johns Hopkins University, Baltimore, USA
C
Chenkai Zhang
the School of Information Engineering, Huzhou University, Huzhou, China
Y
Yan Jia
Graduate School, the Beijing University of Chinese Medicine, Beijing, China
Y
Yukai Wang
the School of Information Engineering, Huzhou University, Huzhou, China
R
Ruiyang He
Dongzhimen Hospital, Beijing University of Chinese Medicine, Beijing, China
C
Changyong Luo
the Infectious disease department, Dongfang Hospital, Beijing University of Chinese Medicine, Beijing, China
Z
Zhi Liu
the School of Pharmacy, Nanjing University of Chinese Medicine, Nanjing, China; the Infectious disease department, Dongfang Hospital, Beijing University of Chinese Medicine, Beijing, China