Score
Designs and implements evaluation frameworks, annotation protocols, and analytic pipelines to assess the quality and correctness of clinical reasoning outputs by comparing generated reasoning traces and final answers against expert clinician judgments. Builds metrics and human-in-the-loop processes to quantify reasoning-to-output mismatch, surface recurring failure patterns, and produce actionable evaluations for model or workflow improvement.
This study addresses the critical challenge of systematically aligning large language models’ reasoning capabilities with real-world clinical demands to enhance their reliability and applicability in healthcare settings. The authors propose the first analytical framework that integrates Miller’s pyramid of clinical competence with deductive, inductive, and abductive reasoning paradigms, enabling the construction of a benchmark dataset spanning five levels of medical reasoning. Through multidimensional evaluation of 18 state-of-the-art models, the study reveals that specialized models excel in diagnostic tasks, whereas general-purpose models demonstrate superior performance in clinical decision support and patient–physician communication. The findings also highlight persistent challenges in hallucination control, data scarcity, and practical deployment in clinical workflows.
While current large language models demonstrate accuracy in clinical diagnosis, it remains unclear whether their reasoning follows stable, structured clinical logic. This work proposes the Clinical Reasoning Graph framework—a structured graph representation grounded in a clinical ontology comprising five node types and seven edge types—and leverages natural language processing and graph similarity metrics to extract and analyze 750 diagnostic trajectories. The study reveals that graph similarity between correct and incorrect diagnoses is nearly identical (0.488 vs. 0.484), and reasoning structures show no significant consistency across similar cases, indicating a lack of schematic-level stability in cross-case reasoning. Although structured reflection prompts improve feature analysis, they do not enhance structural consistency. These findings underscore the need for process-level evaluation to complement conventional outcome-based accuracy and offer a novel paradigm for explainability in clinical AI.
This work systematically audits three major commonsense reasoning benchmarks—SocialIQa, FauxPas-EAI, and ToMi—and exposes critical flaws in item design and evaluation methodology: current automatic scoring overemphasizes superficial output formatting, making evaluations vulnerable to spurious format-based cues and unable to reliably assess LLMs’ genuine reasoning capabilities. Method: We propose a new evaluation paradigm centered on “reasoning process consistency,” prioritizing logically robust, information-grounded inference over surface-level correctness. To support this, we release a human-verified, re-annotated clean dataset and a diagnostic toolkit. Contribution/Results: Through multi-round diagnostic evaluations across GPT-3/3.5/4/o1 and LLaMA 3.1, we demonstrate that apparent performance gains largely stem from input perturbations rather than substantive reasoning improvements. Our framework establishes a foundation for interpretable, reproducible, process-oriented reasoning evaluation—advancing both theoretical understanding and practical assessment rigor.
This work addresses the vulnerability of large language models (LLMs) to “evaluation hallucination” in clinical reasoning—where fluent yet erroneous explanations mask diagnostic inaccuracies. To this end, the authors propose CLExEval, a novel framework that integrates 5,600 physician annotations and 200 reasoning trajectories to qualitatively analyze LLM reasoning in rare disease diagnosis. Through progressive information masking, human-in-the-loop evaluation, and LLM-as-a-Judge validation, the study identifies three failure modes: verbosity bias under information scarcity, a hidden knowledge paradox stemming from failed expert knowledge retrieval, and misalignment between internal reasoning and final output. Experiments reveal that GPT-4o-mini’s accuracy drops sharply from 95.0% to 32.5% under limited information, and 68.6% of correct reasoning chains fail to translate into accurate final answers, highlighting a significant overestimation of clinical reliability by automated evaluation metrics.
This study addresses the challenge that large language models often produce correct diagnostic conclusions through opaque or flawed reasoning, failing to meet the high standards of interpretability and reliability required in clinical decision-making. To tackle this issue, the authors introduce the Toulmin model of argumentation into clinical diagnosis for the first time and propose a Curriculum Goal-Conditioned Learning (CGCL) framework. This approach employs a three-stage progressive training strategy to guide models in constructing structured, verifiable diagnostic arguments. Integrated with the T-Eval evaluation framework, the method achieves diagnostic accuracy and reasoning quality comparable to reinforcement learning baselines while significantly enhancing reasoning transparency, reliability, and training stability.
This work addresses the limitation of conventional evaluation methods that rely solely on outcome accuracy, which often fail to distinguish between large language models with differing reasoning capabilities but similar accuracy. To overcome this, the authors propose a novel evaluation framework centered on high-confidence reasoning trajectories, introducing the Filtered Reasoning Score (FRS)—a metric that assesses only the model’s top-K% most confident reasoning paths. FRS integrates multiple dimensions of reasoning quality, including faithfulness, coherence, utility, and factuality. Experimental results demonstrate that FRS effectively captures nuanced differences in reasoning ability among models that appear comparable under standard accuracy metrics, and models achieving higher FRS consistently exhibit stronger generalization performance across diverse reasoning benchmarks.
This work addresses the limitations of existing clinical reasoning agents, which rely on manually curated tool libraries with high maintenance costs, and zero-shot code generation that often yields inefficient or unreliable reasoning chains under institutional policy constraints. To overcome these challenges, the authors propose the first composable skill framework for automated construction and evaluation in clinical reasoning. The approach formalizes natural language clinical guidelines into verified Python skills through an offline automated pipeline and introduces CodeClinic—a benchmark built on MIMIC-IV that encompasses longitudinal ICU monitoring and compositional information retrieval tasks. Experimental results demonstrate that, compared to zero-shot generation, this method maintains reasoning consistency while reducing token consumption per query by up to 40%, substantially enhancing skill reusability and reliability.
Existing diagram question answering (Diagram QA) datasets lack structured visual evidence attribution annotations, and their annotation tools are tightly coupled to specific data formats, limiting reusability. This work proposes a lightweight, reviewer-in-the-loop framework that decouples interface logic from data structure through meta-schema abstraction and dataset-specific adapters. It introduces question-answer–conditioned evidence region selection and a human-in-the-loop verification mechanism, enabling automatic generation and interactive refinement of missing questions or candidate regions. Evaluated across six Diagram QA datasets, the approach achieves 85.39% precision and 75.30% recall (micro-averaged), substantially reducing manual annotation costs while maintaining high attribution consistency. The code and demonstration system are publicly released.
This work proposes a novel evaluation task designed to assess AI systems’ ability to integrate continuous visual perception, temporal structure reconstruction, and clinical workflow knowledge in the context of clinical skill assessment. Specifically, the system must reorder shuffled clinical keyframes into their correct temporal sequence and generate expert-verifiable reasoning explanations. To support this, the authors introduce a benchmark dataset comprising 200 test instances across three emergency medical procedures and employ multidimensional metrics—including task accuracy, pairwise accuracy, and BERTScore—for comprehensive evaluation. Analysis of 90 submissions from seven teams reveals that current models still face significant challenges in jointly leveraging visual evidence, temporal logic, and domain-specific knowledge. This study formalizes this reasoning task for the first time, establishing a new benchmark for multimodal understanding in clinical settings.
This work addresses the challenge of verifying fluent yet potentially unreliable multi-step reasoning generated by AI in high-stakes domains. The authors propose the first reference-free reasoning evaluation framework, which decomposes reasoning trajectories into segments, annotates local premise-conclusion relationships using natural language inference (NLI), and constructs a hypergraph to represent their logical structure. By applying deterministic backward AND-OR search over this hypergraph, the method assigns audit labels to each segment, quantifying the strength of its internal logical support. Innovatively integrating NLI with hypergraph-based representation, the approach emphasizes the compositional nature of inferential relations within reasoning chains, eschewing reliance on final answers or LLM-based judges. Evaluated on mathematical (Hard2Verify) and clinical reasoning (UroReason) tasks, the framework significantly outperforms LLM judges, particularly excelling at identifying logically weak yet linguistically fluent reasoning segments in clinical contexts.