Score
Design and implement methods and pipelines to identify and merge records that represent the same real-world entities across two or more datasets, including blocking/candidate retrieval, construction of matching keys, similarity-scoring functions, multi-signal and probabilistic weighting schemes, and computation of composite match scores with confidence tiers. Validate linkage quality through manual review and quantitative metrics and produce cleaned, documented linked datasets suitable for downstream analysis.
This paper addresses heterogeneous entity matching (HEM), a core challenge in data integration, arising from structural, syntactic, schema, and semantic heterogeneity. To tackle the resulting modeling difficulties, we propose the first unified classification framework that jointly accounts for representational and semantic heterogeneity, grounded in the FAIR principles to expose fundamental limitations of existing methods under semantic inconsistency. Through a systematic literature review, taxonomy-driven modeling, and cross-model experimental evaluation, we empirically demonstrate the robustness deficiencies of mainstream entity matching models in semantically heterogeneous settings. Our analysis identifies key research directions—including multimodal fusion, human-in-the-loop approaches, joint modeling with large language models and knowledge graphs—as critical for advancing HEM. The work establishes a theoretical foundation, provides a standardized evaluation benchmark, and outlines a principled technical roadmap for future HEM research.
This study addresses the entity linking problem in multi-source heterogeneous data lacking unique identifiers. We propose a probabilistic record linkage method that balances accuracy and scalability. Methodologically, we introduce the Stochastic EM algorithm into latent-variable generative models for the first time, explicitly modeling dependencies among link decisions and enforcing one-to-one constraints, while enabling robust linking under variable-quality fields. Our approach innovatively supports dynamic precision–efficiency trade-offs, effectively handling real-world challenges such as information evolution, data entry errors, and low-quality attributes. Extensive evaluation on large-scale real-world healthcare data demonstrates high linkage accuracy; simulation experiments confirm strong robustness to noise and missing values. The open-source R package FlexRL has been released and deployed in production environments.
This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.
In entity resolution, traditional deterministic blocking methods suffer from low recall and poor precision when unique identifiers are unavailable, primarily due to field errors or omissions. To address this, we propose a fault-tolerant blocking framework that integrates approximate nearest neighbor (ANN) search with graph-based algorithms. Specifically, we introduce the first approach that jointly leverages Locality-Sensitive Hashing (LSH) and Hierarchical Navigable Small World (HNSW) indexing within connected-component analysis to generate candidate pairs with high recall and low redundancy. The framework unifies support for diverse vector embeddings and similarity metrics via standardized interfaces. Evaluated on official benchmark datasets, our method achieves 12–28% higher recall compared to conventional blocking techniques while reducing the number of pairwise comparisons by over 90%, thereby significantly improving both efficiency and accuracy.
In heterogeneous data integration, schema matching and entity resolution are significantly affected by domain-specific characteristics, data scale, missingness rates, and attribute overlap—yet existing graph-based methods struggle to jointly leverage structural and semantic information. Method: This paper proposes a context-aware graph embedding framework that unifies tabular structure, column-level textual descriptions, and external knowledge, employing graph neural networks for joint encoding and embedding learning of multi-source heterogeneous data. Contribution/Results: A key innovation is the context-enhancement mechanism, which systematically uncovers how data characteristics influence matching performance and empirically demonstrates that contextual modeling substantially improves robustness and accuracy—especially under challenging conditions such as high missingness rates and prevalent numeric columns. Extensive experiments across multiple domain-specific benchmark datasets show that our method consistently outperforms state-of-the-art graph-based baselines.
This study addresses the critical need for quantifying dataset similarity in model generalization, transfer learning, simulation calibration, and two-sample testing. We systematically survey 118 similarity quantification methods and propose the first ten-dimensional classification framework, organizing approaches into seven technical categories: statistical distances (e.g., Wasserstein distance, Maximum Mean Discrepancy), kernel-based methods, information-theoretic measures, dimensionality-reduction embeddings, permutation tests, generative-model-based discriminators, and Gaussian process likelihood ratios. We develop a multi-dimensional evaluation system balancing theoretical guarantees, interpretability, and practical applicability, yielding a structured recommendation matrix aligned with task requirements and data characteristics. Furthermore, we introduce the first open-source, interactive tool for method selection—enabling real-time filtering and parameter configuration—to significantly enhance both selection efficiency and deployment suitability.
This work addresses the joint schema-value matching problem between pivot tables and relational tables in data lakes, which demands semantic consistency, value compatibility, and generalization under anonymized data. To this end, we propose PiLLar, a novel framework that, for the first time, formulates the task as a large language model (LLM)-guided Monte Carlo Tree Search (MCTS), enabling unsupervised, training-free cross-domain adaptation. We provide a dynamic theoretical error analysis that guarantees asymptotic convergence. Furthermore, we construct PTbench, the first real-world benchmark for this problem. Experimental results show that PiLLar achieves an average matching accuracy of 87.94% on PTbench, significantly outperforming existing methods and demonstrating its effectiveness and strong generalization capability.
This work addresses the lack of end-to-end public benchmarks for data integration by introducing MaDI-Bench, the first comprehensive benchmark that spans the entire relational table integration pipeline—including schema matching, value normalization, entity blocking, entity matching, and data fusion. To mitigate benchmark saturation, MaDI-Bench incorporates a set of foundational cross-domain tasks along with a mechanism for generating extensible task variants. The benchmark supports both step-wise and end-to-end evaluation of system performance, validated through diverse pipelines ranging from manual and optimal combinations to large language model (LLM)-based approaches. All resources are publicly released to foster reproducible and holistic assessment of data integration systems.
This study addresses the high cost and lack of systematic design in manual review of candidate record pairs, which hinders the simultaneous optimization of accuracy, representativeness, and uncertainty coverage under limited budgets. The authors model review as finite-population sampling based on fine-grained stratification and introduce a novel multidimensional stratification framework that integrates match-score intervals, comparison patterns, record-level ambiguity, and demographic groups. Ambiguity is quantified using match probability bands—derived from deciles of model scores—and conditional candidate perplexity. A tunable allocation mechanism is achieved through within-band error tolerances and global budget scaling. Experiments demonstrate that by reviewing only 7% of samples (compared to 23% in the baseline), the framework preserves high-segment matching accuracy and ambiguity distribution, confirming its effectiveness and flexibility under resource constraints.
This work addresses the high computational and communication overhead that hinders scalability in two-party private record linkage (PPRL) at large scale. To overcome this limitation, the authors propose a “filter-then-link” framework that introduces a lightweight privacy-preserving record screening (PPRS) phase prior to PPRL, enabling efficient assessment of collaboration value. The key innovation lies in an Oblivious Attribute/Feature Alignment protocol that supports approximate matching and pattern awareness, thereby transcending the symmetric-function constraint inherent in circuit-based private set intersection (PSI). Building upon circuit PSI, they further develop an Appraisal system that integrates secure shuffling and PSI techniques to achieve highly efficient screening. Experimental results demonstrate that, under identical constraints, their approach handles up to 850× more records and operates 165× faster than the state-of-the-art SFour system, substantially accelerating the identification of high-value collaboration partners.