Score
Designing and conducting evaluations focused on real users and stakeholders to measure experience, agency and interpretability (e.g., BCI users), produce auditable case-level explanations for clinicians, and ensure specifications are transparent and robust across domains.
Current user evaluations of eXplainable Artificial Intelligence (XAI) in healthcare lack systematic frameworks and practical guidelines, leading to insufficient validation of trustworthiness and usability. Method: We conducted a systematic literature review across multiple databases (e.g., PubMed, ACM Digital Library), coding and analyzing 82 user evaluation studies conducted in clinical or medical contexts. Contribution/Results: We introduce the first atomic Explanatory Experience Attributes Framework tailored specifically for healthcare XAI, uncovering intrinsic relationships among explanatory properties. Building on this, we propose a sensitivity-based evaluation guideline grounded in both system characteristics and clinical context. Our contributions include an updated, validated evaluation framework; interdisciplinary, practice-oriented implementation protocols; and an analysis of emerging empirical trends. Collectively, these advances significantly enhance the verifiability of XAI systems’ trustworthiness and usability in real-world clinical settings.
Current medical AI deployment faces a fundamental gap among explainability theory, clinical requirements, and regulatory expectations, compounded by the absence of practical guidance for preclinical evaluation readiness. This study bridges these domains by systematically integrating eXplainable AI (XAI) theory, clinical practice needs, and regulatory frameworks—introducing two foundational preclinical development principles: “Transparency-by-Design” and “Actionability-by-Design” to establish a shared interdisciplinary language. Methodologically, we unify model calibration, uncertainty quantification, and robustness engineering to enable case-level interpretability and full system-behavior traceability. Our contribution is a rigorously defined, empirically verifiable technical boundary and actionable implementation guidelines that significantly reduce the preparation time for clinical evaluation. This work establishes a methodological foundation for compliant, trustworthy, and clinically integrated AI deployment in healthcare. (149 words)
Medical AI explainability assessments often lack clinical relevance, hindering real-world deployment. Method: This study proposes the first clinically grounded, three-dimensional explainability framework—comprising comprehensibility, trustworthiness, and usability—and implements it in a prototype system for postpartum depression risk prediction. The framework was developed through systematic literature review and expert consensus, followed by human-AI co-design and integration of explainable machine learning models into an interactive web-based interface. Empirical evaluation involved 20 clinicians using a novel, internally validated 13-item System Explainability Scale (SES; Cronbach’s α = 0.84, ρ = 0.81). Results: Clinicians rated the system highly across all dimensions (usability: 4.71; trustworthiness: 4.53; comprehensibility: 4.51; overall explainability: 4.56 on a 5-point scale), confirming the framework’s efficacy in mitigating explainability barriers in clinical AI adoption and establishing a new paradigm for standardized, clinically aligned explainability assessment.
This study addresses the challenge of effectively integrating AI-powered clinical perception tools into real-world healthcare workflows, focusing on balancing clinical utility, user acceptance, and system trustworthiness. Through in-depth interviews and inductive thematic analysis with 20 AI healthcare tool developers, we conducted a qualitative investigation to identify key design barriers impeding clinical adoption. Our analysis yields a novel conceptual framing—“developers as ethical stewards”—emphasizing their proactive role in embedding clinical knowledge and ethical responsibility throughout technical implementation. Based on this, we distill four interdisciplinary design priorities: (1) transparent, customizable decision logic; (2) clearly defined human–AI role boundaries; (3) progressive, context-aware user training pathways; and (4) clinically grounded explainability. The findings provide both a theoretical framework and actionable guidelines for enhancing clinical integration, trustworthiness, and responsible innovation of AI tools in healthcare settings.
Low clinical adoption of AI in healthcare stems primarily from a misalignment between technical explainability and real-world clinical needs. Method: We conducted an empirical, context-sensitive study involving 20 frontline U.S. clinicians, employing qualitative usability testing and reflexive thematic analysis, integrated with human factors engineering and clinical workflow modeling. Contribution/Results: We propose the first structured, actionable operational definition of AI explainability for healthcare, alongside a customizable design framework. This framework systematically characterizes clinicians’ prioritized preferences across three dimensions: *explanatory content* (e.g., salient variables, uncertainty quantification), *presentation modality* (e.g., visualizations, natural-language summaries), and *temporal embedding* (i.e., integration before, during, or after clinical decision-making). By grounding explainability requirements directly in clinical practice, our work bridges the gap between algorithmic transparency and clinical utility, providing an evidence-based foundation and reusable methodology for designing trustworthy, deployable medical AI systems.
This study addresses critical risks—including privacy breaches, algorithmic bias, and patient–clinician relationship alienation—arising from the deployment of computer perception technologies in healthcare. Through a multi-stakeholder empirical investigation, we conducted semi-structured in-depth interviews with 102 participants across patients, clinicians, and technology developers. Thematic analysis, employing dual-coder annotation and consensus-based adjudication, systematically identified seven key implementation concerns. Our principal contribution is a human-centered design framework anchored in “humanization,” featuring a novel “personalized roadmap” mechanism that integrates algorithmic inference with individual lived experience. The findings clarify persistent barriers to sustainable clinical adoption and yield actionable implementation pathways. This work provides an integrated, evidence-informed guidance framework for policymakers, AI developers, and healthcare practitioners seeking ethically robust, clinically viable integration of computer perception systems.
This work addresses the common lack of continuous evaluation and governance mechanisms in deployed clinical AI systems, which hinders dynamic performance optimization. The authors propose the first end-to-end continuous governance framework tailored for clinical AI, integrating standards-driven validation, A/B testing for controlled version updates, real-time performance monitoring, fault tolerance, and deep integration with electronic health records (EHRs) to establish a closed-loop synergy between engineering iteration and clinical feedback. Applied to Hyperscribe—a speech-to-structured-clinical-note system—the framework achieved substantial improvements over seven iterative cycles: median clinician rating increased from 84% to 95%, negative user feedback decreased from 79% to 30%, median audio processing latency was 8.1 seconds, and task completion rate reached 99.6%, collectively enhancing system reliability and user satisfaction.
This work addresses the risks of generative interfaces in high-stakes settings—such as hallucinations, semantic distortion, bias, and accessibility barriers—that arise from insufficient human oversight and undermine user understanding and control. To mitigate these issues, the paper proposes a “supervision-by-design” architecture that deeply integrates human judgment into the generative pipeline. This framework employs automated risk detection across dimensions including readability, semantic fidelity, factual consistency, and accessibility compliance, coupled with explicit UI controls and a tiered escalation protocol that triggers mandatory human review upon violations. By combining human-in-the-loop (HITL) interventions with human-on-the-loop (HOTL) continuous monitoring, the approach establishes a scalable, verifiable governance loop. The resulting system significantly enhances transparency, reliability, and inclusivity, thereby strengthening accountability and user agency in high-risk human-AI collaborative decision-making.
This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.
Current medical research agents lack domain-specific evaluation mechanisms that rigorously assess scientific validity, methodological soundness, reproducibility, and boundary safety. This work proposes MedSkillAudit—the first skill auditing framework tailored for medical research agents—which employs a hierarchical, structured pipeline to evaluate skill readiness prior to deployment. The framework incorporates expert double-blind scoring (0–100), tiered release recommendations, and high-risk flags, and quantifies agreement between the system and human experts using ICC(2,1) and weighted Cohen’s kappa. Evaluated on 75 skills, the system achieved an ICC of 0.449, surpassing inter-human rater agreement (ICC = 0.300) and demonstrating closer alignment with consensus scores (SD = 9.5 vs. 12.4), thereby validating its effectiveness and reliability.
Expert-role-playing language models frequently fail to consistently disclose their AI identity in high-stakes professional settings, leading users to misjudge their capabilities and incur safety risks. Method: Through a controlled behavioral audit involving 19,200 interactions across 16 open-source models (4B–671B parameters), we employed Bayesian validation and Rogan–Gladen correction to quantify disclosure rates. Contribution/Results: Identity disclosure rates ranged narrowly from 2.8% to 73.6%, exhibiting only weak correlation with parameter count (ΔR² = 0.359); instead, training methodology predominantly governed disclosure behavior. Crucially, inference-time optimizations—such as chain-of-thought or self-refinement—systematically reduced transparency, revealing a novel “reverse Gell-Mann forgetting” effect. The findings indicate that enhancing AI self-disclosure necessitates fundamental reconfiguration of training paradigms—not merely scaling model size or refining inference strategies.