dataset reconciliation

Designs and implements pipelines, mappings, and validation procedures to merge multiple datasets or corpora into a single consistent collection. This includes aligning heterogeneous schema fields, resolving differing record granularities, deduplicating and normalizing entries, and producing unified representations (e.g., consolidated TSVs).

datasetreconciliation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Towards Scalable Schema Mapping using Large Language Models

May 30, 2025
CB
Christopher Buss
🏛️ Oregon State University | Portland State University

To address key challenges in multi-source data integration—including poor scalability and high maintenance costs of manual schema mapping, as well as output inconsistency, limited expressiveness (e.g., inability to support GLaV), and excessive invocation overhead in LLM-based approaches—this paper proposes a lightweight and robust LLM-augmented schema mapping framework. Methodologically: (i) sampling-based majority voting is introduced to mitigate LLM output volatility; (ii) language-enhanced prompting is designed to natively support highly expressive mappings (e.g., GLaV); and (iii) structured metadata-driven pre-filtering significantly reduces redundant LLM invocations. Experiments across multiple benchmarks demonstrate substantial improvements in mapping accuracy and robustness, with over 40% reduction in LLM calls and efficient generation of mappings for schemas containing up to hundreds of attributes. The framework establishes a new paradigm for scalable, low-maintenance data integration.

Addressing scalability challenges in data integration systemsImproving inconsistent outputs in LLM-based schema mappingReducing computational costs of repeated LLM calls

LLMATCH: A Unified Schema Matching Framework with Large Language Models

Jul 14, 2025
SW
Sha Wang
🏛️ Singapore Management University | Singapore University of Technology and Design | PayPal | DBS Bank

To address low accuracy and poor scalability in complex multi-table schema matching for enterprise data integration, this paper proposes LLMatch, a unified framework. Methodologically, LLMatch introduces (1) a novel two-stage semantic optimization strategy: the Rollup module aggregates semantically related columns to enhance generalization, while the Drilldown module enables fine-grained mapping reconstruction; (2) an LLM-driven column-level alignment method integrating schema preprocessing, candidate table filtering, and hierarchical semantic reduction; and (3) SchemaNet—the first benchmark dataset designed for realistic, complex schema-matching scenarios. Experimental results demonstrate that LLMatch significantly improves matching accuracy on multi-table tasks and substantially enhances engineers’ efficiency and debuggability in practical data integration workflows.

Addresses lack of benchmarks for real-world multi-table schema challengesDevelops a unified framework for complex multi-table schema matchingIntroduces a two-stage optimization strategy for semantic column alignment

This study addresses the challenge of effectively integrating structured data with unstructured text, a longstanding barrier in data management. It presents the first systematic argument for the necessity of textual data integration and introduces a unified framework that synergistically combines natural language processing, knowledge extraction, and traditional data integration techniques. By leveraging semantic alignment, the framework achieves deep integration between textual content and structured schemas, thereby tackling key challenges inherent in heterogeneous data integration. The work comprehensively surveys existing methodologies and outstanding issues, establishing a theoretical foundation for the emerging field of textual data integration and offering clear guidance for future research and practical implementation.

Data IntegrationHeterogeneous DataStructured Data

Semantic drift in enterprise data pipelines—caused by multilingual transformations—decouples metadata from downstream data semantics, undermining reproducibility, governance, and performance of RAG and text-to-SQL applications. To address this, we propose a fine-grained schema lineage extraction method leveraging multilingual parsing, chain-of-thought prompting (optimized for 1.3B–32B models), and human-in-the-loop evaluation. We introduce SLiCE (Schema Lineage Composite Evaluation), the first benchmark framework tailored for multilingual script lineage, alongside a high-quality dataset of 1,700 real-world annotated samples. Experiments show that open-weight 32B models match GPT-4’s lineage accuracy under standard prompting, demonstrating cost-effective lineage extraction. Our core contributions are: (1) a systematic formalization of semantic-faithful lineage; (2) the first open, multilingual schema lineage benchmark with rigorous annotations; and (3) a lightweight, efficient extraction paradigm enabling scalable, accurate lineage inference.

Addressing semantic drift in data reproducibility and governanceEvaluating lineage quality with structural and semantic metricsExtracting fine-grained schema lineage from multilingual enterprise pipelines

This work proposes the first fully large language model–driven, end-to-end data integration framework that eliminates the need for manual configuration, which traditionally incurs high costs and low efficiency. The system autonomously generates a complete integration pipeline encompassing schema mapping, value normalization, entity matching, and conflict resolution without human intervention. Evaluated on three real-world domains—gaming, music, and enterprise data—the GPT-5.2–based framework achieves integration performance comparable to or surpassing that of handcrafted systems. Notably, it accomplishes this at a remarkably low cost of approximately $10 per execution, substantially reducing human labor and operational overhead.

data integrationend-to-end automationhuman effort reduction

Latest Papers

What's happening recently
View more

This work addresses the lack of end-to-end public benchmarks for data integration by introducing MaDI-Bench, the first comprehensive benchmark that spans the entire relational table integration pipeline—including schema matching, value normalization, entity blocking, entity matching, and data fusion. To mitigate benchmark saturation, MaDI-Bench incorporates a set of foundational cross-domain tasks along with a mechanism for generating extensible task variants. The benchmark supports both step-wise and end-to-end evaluation of system performance, validated through diverse pipelines ranging from manual and optimal combinations to large language model (LLM)-based approaches. All resources are publicly released to foster reproducible and holistic assessment of data integration systems.

data fusiondata integrationend-to-end benchmark

This work addresses the challenge of automatic schema integration for multi-source heterogeneous tables, particularly denormalized tables describing multiple entity types. The authors propose SINT-Flow, a framework comprising five composable operators powered by large language models (e.g., GPT-5.2, Qwen-3.6-27B), orchestrated by a workflow engine to enable end-to-end automation. SINT-Flow supports entity decomposition, attribute identification, and schema mapping for denormalized tables. Innovatively integrating a self-consistency strategy and a review-loop mechanism, it achieves the first effective handling of multi-entity denormalized tables. Evaluated on the newly constructed SINT-Bench benchmark, the approach attains over 96% F1 score in entity type detection, 85% in attribute detection, and 83% in schema mapping, demonstrating its significant effectiveness.

denormalized tablesentity-type decompositionrelational tables

This study addresses the challenges posed by the high heterogeneity of healthcare data and the lack of effective metadata management, which often degrade conventional data lakes into “data swamps,” impeding data interoperability and machine learning (ML) readiness. To overcome these limitations, the authors propose a dual-hybrid semantic data lake architecture that synergistically integrates the dynamic modeling capabilities of knowledge graphs with the metadata generation power of large language models (LLMs). A human-in-the-loop validation mechanism is incorporated to enable automated metadata annotation and high-level semantic alignment. This approach establishes, for the first time, semantic linkages within a data lake explicitly oriented toward ML operability, substantially enhancing the discoverability and computability of heterogeneous medical data while supporting intelligent recommendation of suitable ML methods.

data lakeheterogeneous datainteroperability

Automatically generating YAML configuration files that are both structurally valid and compliant with multiple continuous integration (CI) service specifications remains a significant challenge, and the capabilities of current large language models (LLMs) on this task are not well understood. This work introduces DOC2CI, the first cross-CI benchmark dataset comprising 3,363 document–YAML pairs, and systematically evaluates 14 open-source models alongside GPT-series models. A novel failure taxonomy is proposed to uncover the root causes of model discrepancies, and this study provides the first empirical evidence that document similarity and structural validity constitute distinct optimization objectives. Experiments reveal that even the largest models achieve an Exact Match rate below 3.1%; while 97% of generated outputs are syntactically parseable, only 71% conform to the target service schema. Schema-guided post-hoc repair without additional training boosts structural validity to 94%, whereas fine-tuning improves document similarity at the expense of standalone structural correctness.

configuration generationContinuous IntegrationLLM

Hot Scholars

TW

Tim Wittenborg

Research Assistant, L3S, Leibniz University Hannover
Knowledge ManagementKnowledge Representationdigital SustainabilitySystems Engineering
FJ

Feng Jiang

Shenzhen University of Advanced Technology
Discourse ParsingLarge-scale Language ModelDialogue System
CR

Chu-Ren Huang

Chair Professor, The Hong Kong Polytechnic University
computational linguisticscorpus linguisticsChinese linguisticslexical semantics
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion