TCMClinicalReason-Bench: Can Language Models Reason from Pathogenesis to Prescription over Real-World Clinical Cases?
This study addresses the lack of evidence grounding and the cascading propagation of errors in clinical reasoning for Traditional Chinese Medicine (TCM) within large language models. To this end, it constructs a full-pipeline evaluation benchmark derived from real-world, multicenter electronic health records. Methodologically, the work proposes an evidence-constrained scoring mechanism that decouples component-level quality from cross-module logical consistency to precisely localize reasoning bottlenecks, while integrating automated judge models with expert blind reviews to ensure evaluation reliability. The findings reveal that general-purpose models outperform domain-specific counterparts and identify significant deficiencies in prescription generation. These insights provide critical empirical foundations for enhancing the robustness of TCM-oriented AI reasoning systems.