Score
Designs and implements reproducible pipelines, mappings, and crosswalks that align and standardize multiple heterogeneous data sources into a consistent integrated dataset, including unit and coding conversions, variable-definition harmonization, and construction of dataset-level and value-level mappings. Documents and enforces mapping rules, reference choices, and provenance (making reference application parameterizable and auditable), resolves ambiguous category matches, applies country- or source-specific weighting or adjustments, and produces harmonized longitudinal or historical datasets ready for analysis.
To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.
This work addresses the lack of end-to-end public benchmarks for data integration by introducing MaDI-Bench, the first comprehensive benchmark that spans the entire relational table integration pipeline—including schema matching, value normalization, entity blocking, entity matching, and data fusion. To mitigate benchmark saturation, MaDI-Bench incorporates a set of foundational cross-domain tasks along with a mechanism for generating extensible task variants. The benchmark supports both step-wise and end-to-end evaluation of system performance, validated through diverse pipelines ranging from manual and optimal combinations to large language model (LLM)-based approaches. All resources are publicly released to foster reproducible and holistic assessment of data integration systems.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses the challenge of collaborative analysis across multi-institutional electronic health record (EHR) systems, hindered by data heterogeneity, semantic inconsistency, and stringent privacy constraints. We propose a unified two-module framework that enables privacy-preserving EHR harmonization across institutions and disparate data models—without sharing raw individual-level data. The framework integrates standardized clinical coding mapping with machine learning–driven representation learning. Accompanied by open-source software and step-by-step implementation tutorials, it supports end-to-end translational research. Empirical validation across multiple real-world healthcare systems demonstrates substantial improvements in data interoperability and reusability, enabling the construction of high-quality, research-ready EHR datasets. The approach exhibits strong generalizability, scalability, and clinical deployability.
This work addresses key limitations in current pattern matching evaluation—namely, insufficient diversity in benchmark datasets and the absence of interactive validation mechanisms. The authors propose an interactive visualization system that supports human-in-the-loop collaboration by integrating automated matching, coordinated multiple views, interactive heatmaps, hierarchical navigation, and large language model–generated explanations. A standardized interface enables plug-and-play integration of new matchers, while a novel dynamic benchmarking mechanism iteratively refines ground truth labels through expert validation feedback, facilitating continuous algorithmic improvement and cross-domain comparison. The system has been validated in two scenarios: data reconciliation and developer-in-the-loop benchmarking, demonstrating its effectiveness in efficiently constructing high-quality annotated datasets and enabling real-time performance evaluation of matching algorithms.
Although scientific data increasingly adhere to the FAIR principles and employ standardized identifiers, practical interoperability remains hindered by heterogeneity in identifier systems and data models. This work proposes and implements two synergistic tools—Babel and ORION—to bridge this gap. Babel constructs clusters of equivalent identifiers through mapping-based clustering and exposes them via a high-performance quantitative API, while ORION standardizes heterogeneous knowledge bases by aligning them to a community-governed common data model. Together, they systematically address the longstanding disconnect between the FAIR “Interoperable” principle and its real-world implementation. The integration of these tools has enabled the construction of a fully interoperable knowledge base, substantially enhancing cross-resource data integration and query capabilities. The resulting framework is publicly available.
This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.
This study addresses the challenge of effectively integrating structured data with unstructured text, a longstanding barrier in data management. It presents the first systematic argument for the necessity of textual data integration and introduces a unified framework that synergistically combines natural language processing, knowledge extraction, and traditional data integration techniques. By leveraging semantic alignment, the framework achieves deep integration between textual content and structured schemas, thereby tackling key challenges inherent in heterogeneous data integration. The work comprehensively surveys existing methodologies and outstanding issues, establishing a theoretical foundation for the emerging field of textual data integration and offering clear guidance for future research and practical implementation.
This work addresses the limitations of traditional database migration approaches, which are often constrained to specific source-target model pairs and struggle to support general-purpose migration in heterogeneous, multi-model environments. To overcome this, the authors propose a model-driven migration framework based on a unified data model called U-Schema. By mapping diverse data models into a common intermediate representation, the framework drastically reduces the number of required transformation pathways and enables cross-paradigm migrations. It employs traceable metadata to decouple schema transformation from data migration, thereby preserving semantic consistency while enhancing structural fidelity and query behavior equivalence. Empirical evaluation—conducted in a relational-to-document database migration scenario using both synthetic datasets and the Northwind benchmark—demonstrates the approach’s effectiveness and scalability across varying data sizes.