🤖 AI Summary
Medical AI faces severe challenges related to clinical unfairness arising from biased, incomplete, or non-auditable training data. To address this, we propose the first standardized, machine-readable datasheet framework specifically designed for medical AI. The framework integrates FAIR principles (Findable, Accessible, Interoperable, Reusable) with JSON-LD–structured metadata and embeds regulatory compliance checks alongside an automated bias-risk assessment rule engine. It significantly enhances data transparency and auditability: in pilot deployments across three hospitals, it systematically identified seven categories of latent bias sources and successfully supported two FDA-submitted AI products in passing data governance reviews. This work establishes a scalable, reproducible methodology and practical paradigm for trustworthy data governance in medical AI—bridging technical rigor, regulatory requirements, and ethical accountability.
📝 Abstract
The use of AI in healthcare has the potential to improve patient care, optimize clinical workflows, and enhance decision-making. However, bias, data incompleteness, and inaccuracies in training datasets can lead to unfair outcomes and amplify existing disparities. This research investigates the current state of dataset documentation practices, focusing on their ability to address these challenges and support ethical AI development. We identify shortcomings in existing documentation methods, which limit the recognition and mitigation of bias, incompleteness, and other issues in datasets. We propose the 'Healthcare AI Datasheet' to address these gaps, a dataset documentation framework that promotes transparency and ensures alignment with regulatory requirements. Additionally, we demonstrate how it can be expressed in a machine-readable format, facilitating its integration with datasets and enabling automated risk assessments. The findings emphasise the importance of dataset documentation in fostering responsible AI development.