Score
Translating clinical standards and evidence-based principles into computable rules and process-level criteria that guide model reasoning and decision support. This includes encoding standards of care and mechanistic links as formal rules, prioritizing research gaps, and producing practical guidance for AI adoption and assurance.
This study addresses the critical challenge of systematically aligning large language models’ reasoning capabilities with real-world clinical demands to enhance their reliability and applicability in healthcare settings. The authors propose the first analytical framework that integrates Miller’s pyramid of clinical competence with deductive, inductive, and abductive reasoning paradigms, enabling the construction of a benchmark dataset spanning five levels of medical reasoning. Through multidimensional evaluation of 18 state-of-the-art models, the study reveals that specialized models excel in diagnostic tasks, whereas general-purpose models demonstrate superior performance in clinical decision support and patient–physician communication. The findings also highlight persistent challenges in hallucination control, data scarcity, and practical deployment in clinical workflows.
This paper addresses critical limitations of large language models (LLMs) in healthcare—namely, weak clinical reasoning, poor interpretability, unmitigated bias, absence of patient safety mechanisms, and challenges in multimodal data integration. To this end, we propose an evolutionary reasoning paradigm for clinical decision support. Our methodology integrates chain-of-thought prompting, a healthcare-specialized multi-agent architecture, reinforcement learning–based trustworthy reasoning optimization (e.g., inspired by DeepSeek-R1), multimodal clinical data fusion, and a domain-customized evaluation framework. Key contributions include: (1) the first systematic roadmap for advancing LLM reasoning capabilities in clinical settings; (2) identification of structural deficiencies in existing frameworks regarding clinical trustworthiness and safety alignment; and (3) a synergistic technical pathway for enhancing interpretability and mitigating bias, significantly improving model robustness and decision reliability in real-world clinical applications—thereby providing both theoretical foundations and practical paradigms for high-stakes clinical AI deployment.
Current medical AI deployment faces a fundamental gap among explainability theory, clinical requirements, and regulatory expectations, compounded by the absence of practical guidance for preclinical evaluation readiness. This study bridges these domains by systematically integrating eXplainable AI (XAI) theory, clinical practice needs, and regulatory frameworks—introducing two foundational preclinical development principles: “Transparency-by-Design” and “Actionability-by-Design” to establish a shared interdisciplinary language. Methodologically, we unify model calibration, uncertainty quantification, and robustness engineering to enable case-level interpretability and full system-behavior traceability. Our contribution is a rigorously defined, empirically verifiable technical boundary and actionable implementation guidelines that significantly reduce the preparation time for clinical evaluation. This work establishes a methodological foundation for compliant, trustworthy, and clinically integrated AI deployment in healthcare. (149 words)
Clinical decision support systems must balance accuracy with auditability, yet existing formal methods struggle to verify the epistemic appropriateness of evidence underlying clinical rules. This work proposes a domain-specific language (DSL) grounded in design-by-contract principles, introducing meta-predicates and an epistemic type system to categorize evidence by purpose, knowledge domain, scale, and acquisition method, while statically constraining the types of evidence permissible within rules. By uniquely integrating meta-predicates with an epistemic type system, the approach enables pre-deployment verification of evidence appropriateness and per-sample audit trails without compromising rule readability. Evaluation on the AnFiSA platform using 5.6 million genomic variants demonstrates that the framework effectively identifies epistemic errors in both AI-generated and human-authored rules, offering a trustworthy, independently auditable foundation for clinical AI decision support.
This work addresses the challenge that current clinical practice guidelines (CPGs), typically represented as free-text documents, are ill-suited for explicitly modeling their underlying decision logic in language model training or retrieval. To overcome this limitation, the study introduces a novel approach that first converts CPGs into executable, programmatic decision structures and then leverages these to generate factual and counterfactual question-answer pairs, thereby constructing structured supervision signals for fine-tuning large medical language models. This enables the models to internalize guideline-driven clinical reasoning rather than merely memorizing surface-level text. Experimental results demonstrate an average relative accuracy improvement of 10.28% across four clinical reasoning benchmarks. Furthermore, clinician evaluations confirm that the model’s generated explanations exhibit significantly higher fidelity, validity, completeness, and clarity compared to baseline methods.
AI deployment in medicine faces a “translation gap,” primarily due to a technology-centric paradigm fundamentally misaligned with clinicians’ diagnostic reasoning and decision-making practices. Method: This study proposes a socio-technical co-support framework anchored in physicians’ cognitive processes and clinical workflows, introducing a novel clinical-cognition–oriented AI support paradigm that prioritizes real-world utility over context-agnostic benchmark performance. Integrating medical anthropology, cognitive science, and explainable AI (XAI), we design a data-driven tool architecture aligned with clinical reasoning habits and operational constraints. Contribution/Results: We establish a new evaluation framework for AI in healthcare—centered on clinical adaptability, explainability, and human-AI collaboration—thereby providing a systematic theoretical and practical guide for developing trustworthy, clinically integrated AI systems.
Despite their promise, large language models (LLMs) face a critical “algorithm-to-application” gap in clinical deployment, hindering real-world integration into electronic health record (EHR) systems. Method: Drawing on empirical EHR deployment experience, we propose the first systematic framework for implementing generative AI agents in clinical settings—centered on sociotechnical implementation tasks, which constitute over 80% of deployment effort. The framework addresses five core challenges: EHR data integration, clinical validation of model trustworthiness, economic sustainability, adaptive management of model and system drift, and multi-stakeholder governance. It synergistically integrates LLMs, prompt engineering, FHIR-compliant EHR interfaces, continuous performance monitoring, and governance protocols. Contribution/Results: We deployed irAE-Agent—a clinical AI agent for automated identification of immune-related adverse events—demonstrating feasibility and robustness. Evaluation by 20 clinical and technical experts confirms the framework significantly enhances translatability from pilot to routine clinical service.
This work addresses the critical challenges of safety and reliability in clinical AI systems, which often suffer from fragile prototype architectures and a lack of holistic governance, leading to accountability gaps. To overcome these limitations, we propose “Maria,” a production-grade clinical AI platform that innovatively treats AI agents as modular units, integrating Clean Architecture with an event-driven design. The platform embeds a Human-in-the-Loop governance mechanism to serve as a continuous source of feedback for iterative improvement. Through autonomous MLOps lifecycle management, Maria ensures system maintainability, auditability, scalability, and effective human oversight. This study establishes a highly reliable and auditable reference architecture for clinical AI, offering a reusable engineering paradigm for deploying trustworthy AI systems in high-stakes domains.
This study addresses the ambiguity surrounding the definition and scope of AI agents in healthcare and the lack of effective evaluation frameworks for clinical translation. Through a systematic evidence mapping of 557 studies, it establishes a clear conceptual boundary for medical AI agents and synthesizes their architectural designs and implementation strategies across key clinical tasks—including medical question answering, medical image interpretation, and electronic health record analysis. The work proposes a novel clinical evaluation framework emphasizing auditability, interoperability, and prospective validation. It reveals that current research predominantly relies on retrospective data and public benchmarks, while largely neglecting systematic assessments of safety, reliability, and impact on clinical workflows, thereby offering critical guidance for the future development and translational deployment of healthcare AI agents.
Current clinical AI systems often lack proper calibration under uncertainty and suffer from opaque decision-making, limiting their reliability in high-stakes medical judgments. This work proposes MedMSA, a novel framework that uniquely integrates large language models with formal probabilistic graphical models. By retrieving relevant prior knowledge and constructing a verifiable probabilistic reasoning mechanism, MedMSA enables calibrated and interpretable quantification of uncertainty. The approach generates a differential diagnosis list weighted by uncertainty estimates, preserving clinical utility while ensuring formal transparency in the reasoning process. This advancement lays a foundational pathway toward safe, trustworthy AI-assisted clinical decision support systems.
This study challenges the conventional paradigm in medical AI that equates “correctness” solely with performance metrics, focusing on the automated classification of plasma cells in bone marrow smears for multiple myeloma. It systematically examines core issues including label instability, model interpretability, clinically relevant evaluation, and human–AI responsibility allocation. Integrating medical image analysis, explainable AI, clinical assessment methodologies, and ethical frameworks for human–AI interaction, the work reconceptualizes “correctness” as a multidimensional construct encompassing data quality, interpretability, evaluation rigor, and accountability. The authors propose a theoretical framework tailored to dynamic clinical environments, uncover practical risks such as automation bias, and underscore the necessity of jointly considering technical performance and ethical practice to chart a responsible pathway for deploying medical AI systems.
This work addresses the challenge that large language models (LLMs) struggle to directly execute diagnostic rules from clinical practice guidelines, typically leveraging guideline texts only indirectly. To bridge this gap, the authors propose GuideSkill, a novel framework that compiles guidelines into executable diagnostic skill functions, establishing a model-agnostic external reasoning layer that outputs ranked diagnostic support scores. The approach comprises two components: zero-shot skills initialized from guidelines (GuideSkill-Zero) and case-driven evolutionary optimization (GuideSkill-Evo), which jointly integrate LLM-based differential diagnosis with skill-based scoring. Evaluated across four benchmarks and four backbone LLMs, GuideSkill-Zero improves macro-accuracy by 13.45% on average over baselines, while GuideSkill-Evo achieves an 18.49% gain over direct reasoning, increases gold-label skill coverage from 56.5% to 99.5%, and surpasses the strongest parameter-finetuned baseline without any backbone model fine-tuning.