Score
Designs and implements data standards, schemas, vocabularies, and governance processes for clinical and healthcare data, and builds the validation, curation, and access controls that ensure datasets comply with those standards. Develops pipelines and mappings to harmonize, integrate, and securely handle medical data across systems, and analyzes data quality, interoperability, and regulatory compliance.
This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.
This study addresses the challenge of unreliable AI models and diminished clinical trust stemming from opaque data quality reporting in the secondary use of electronic health records (EHRs). To this end, the authors propose the first comprehensive framework for transparent data quality reporting across the entire EHR lifecycle. The framework innovatively distinguishes between data producers and consumers, explicitly defines five critical phases, and maps established data quality dimensions to specific workflow stages. Through iterative stakeholder and process analysis, a structured reporting mechanism is developed and validated on real-world datasets, demonstrating its ability to effectively trace the origins of data quality issues. The approach significantly enhances data interpretability, fitness-for-use, and governance efficacy, thereby providing a robust foundation for trustworthy AI development and clinical research.
This study addresses the lack of standardization and explicit semantic representation in clinical data, which hinders interoperability and reproducibility in machine learning. To overcome this limitation, the authors propose the first integration of the MEDS clinical event model with Semantic Web technologies, resulting in MEDS-OWL—a lightweight OWL ontology comprising 13 classes, 10 object properties, 20 data properties, and 24 axioms. They further develop the meds2rdf tool to automatically transform MEDS data into FAIR-aligned RDF graphs. The approach leverages SHACL constraints for validation and RDF graph representations to provide a reusable semantic layer that enables semantic enrichment, cross-system interoperability, and graph-based analytics of clinical data. The methodology is successfully validated on a synthetic dataset capturing care pathways of patients with ruptured aneurysms.
To address the challenges of fragmented data interoperability and semantic heterogeneity across healthcare systems—resulting in insufficient service continuity—this paper proposes an ontology-driven Common Semantic Standardized Data Model (CSSDM). The CSSDM integrates the ISO 13940 ContSys ontology with the FHIR standard, replacing conventional interface-based interoperability paradigms. Leveraging a semi-automated ETL pipeline, it enables semantic alignment, dynamic cross-source linking, and knowledge graph construction from heterogeneous health data. This supports secure, trustworthy, cross-institutional and cross-context (particularly home-based care) interoperability. The model fosters collaborative ecosystem development among health information system vendors and cloud service providers. Empirical evaluation demonstrates significant improvements in the quality of structured and semantically enriched health data sharing, continuity of care services, and clinical decision-making reliability.
Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.
Healthcare faces significant challenges in cross-institutional federated analytics due to strong data heterogeneity and insufficient standardization, severely limiting interoperability. To address this, we propose I-ETL, a decentralized federated analytics framework that enables privacy-preserving collaborative integration of heterogeneous multi-center data—including phenotypic, clinical, imaging, and genomic modalities. Our key innovation is advancing interoperability to the ETL source: we design two generic, extensible conceptual models and integrate metadata modeling with standardization techniques to systematically embed cross-institutional semantic alignment directly into the ETL pipeline—first such approach in the literature. Experiments demonstrate that I-ETL substantially improves unified data representation across sources and enhances cross-center collaboration efficiency. By establishing an interoperable, privacy-aware data foundation, I-ETL enables high-quality, trustworthy federated learning across institutional boundaries.
This study addresses the challenges posed by the high heterogeneity of healthcare data and the lack of effective metadata management, which often degrade conventional data lakes into “data swamps,” impeding data interoperability and machine learning (ML) readiness. To overcome these limitations, the authors propose a dual-hybrid semantic data lake architecture that synergistically integrates the dynamic modeling capabilities of knowledge graphs with the metadata generation power of large language models (LLMs). A human-in-the-loop validation mechanism is incorporated to enable automated metadata annotation and high-level semantic alignment. This approach establishes, for the first time, semantic linkages within a data lake explicitly oriented toward ML operability, substantially enhancing the discoverability and computability of heterogeneous medical data while supporting intelligent recommendation of suitable ML methods.
This work addresses the challenge of cross-institutional incompatibility in electronic health records (EHRs) caused by the absence of standardized data representations, which hinders large-scale machine learning applications. To overcome this limitation, the authors propose an open-source metadata repository grounded in the ISO/IEC 11179-3 standard, employing a “middle-out” standardization strategy and a microservices architecture. The system automatically catalogs EHR data elements and their value domains, supports both local Linux deployment and cloud hosting, and integrates modern authentication mechanisms. It features a user-friendly interface that enables error-free metadata registration and facilitates the visual discovery of interoperable features across heterogeneous databases. Validation through use cases such as rare disease patient identification demonstrates the system’s effectiveness in enhancing metadata management and enabling cross-institutional interoperability.
This study addresses the inefficiency and error-proneness of information extraction in medical form filling by proposing CLAIRE, a hybrid workflow. The method adopts a "schema-grounded, verification-first" architecture that automates form completion through field state discovery, source-to-field mapping, deterministic validation, and audit trails. Large language models from the Qwen series are restricted to assisting with mapping tasks without authorization for critical operations, while a bounded error-correction mechanism ensures data rigor. Benchmark evaluations demonstrate that CLAIRE achieves both a success rate and an accuracy of 1.000, substantially reducing the burden of manual review and enhancing processing throughput.
Interoperability between openEHR and HL7 FHIR is hindered by fundamental differences in their data modeling paradigms. Method: This paper proposes a formal, extensible bidirectional transformation framework. It introduces the first domain-specific language (DSL) tailored for openEHR–FHIR mapping, implements a three-layer architecture—semantic alignment, structural mapping, and API adaptation—and delivers an open-source execution engine (openFHIR) alongside a reusable mapping library. Contribution/Results: The framework enables systematic, semantics-driven mapping between international openEHR archetypes and FHIR profiles, supporting both community-driven standardization and local customization. Evaluated across seven clinical domains, it successfully maps 24 openEHR archetypes to 15 FHIR profiles, achieving a 65% mapping reuse rate. This significantly reduces ETL development effort while improving consistency and efficiency in cross-standard clinical data exchange.