🤖 AI Summary
This study addresses the lack of evidence grounding and the cascading propagation of errors in clinical reasoning for Traditional Chinese Medicine (TCM) within large language models. To this end, it constructs a full-pipeline evaluation benchmark derived from real-world, multicenter electronic health records. Methodologically, the work proposes an evidence-constrained scoring mechanism that decouples component-level quality from cross-module logical consistency to precisely localize reasoning bottlenecks, while integrating automated judge models with expert blind reviews to ensure evaluation reliability. The findings reveal that general-purpose models outperform domain-specific counterparts and identify significant deficiencies in prescription generation. These insights provide critical empirical foundations for enhancing the robustness of TCM-oriented AI reasoning systems.
📝 Abstract
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinicalReason-Bench using 2,000 multicenter electronic health record cases to distinguish case-grounded responses from fluent but unsupported diagnostic and therapeutic conclusions. Five general-purpose and two TCM-specific LLMs were evaluated in zero-shot settings. An evidence-constrained rubric assessed seven diagnostic and therapeutic components and three cross-block relations, allowing case-supported alternatives. Qwen3.7-Plus with TCM retrieval served as the automated judge, alongside parallel blinded ratings by five senior TCM clinicians on a 600-case subset. Structural completeness was nearly saturated (99.3-100.0%), but normalized content scores ranged from 40.7% to 54.1%. The five general-purpose models averaged 50.0%, versus 41.3% for the two smaller TCM-specific models. Cross-block logic consistency ranged from 60.8% to 66.8% and correlated moderately with content across cases (Pearson's r = 0.515-0.656). Deficits were greatest in prescription generation, prescription analysis, and symptom-guided modification. In judge stress testing on 100 independent cases, perturbation detection rates across the three relations were 57%, 56%, and 31%, with contradictions detected more reliably than omissions. Separating component quality from cross-block consistency localizes failures missed by endpoint and completeness metrics and identifies where clinician oversight remains necessary.