LLMATCH: A Unified Schema Matching Framework with Large Language Models

📅 2025-07-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address low accuracy and poor scalability in complex multi-table schema matching for enterprise data integration, this paper proposes LLMatch, a unified framework. Methodologically, LLMatch introduces (1) a novel two-stage semantic optimization strategy: the Rollup module aggregates semantically related columns to enhance generalization, while the Drilldown module enables fine-grained mapping reconstruction; (2) an LLM-driven column-level alignment method integrating schema preprocessing, candidate table filtering, and hierarchical semantic reduction; and (3) SchemaNet—the first benchmark dataset designed for realistic, complex schema-matching scenarios. Experimental results demonstrate that LLMatch significantly improves matching accuracy on multi-table tasks and substantially enhances engineers’ efficiency and debuggability in practical data integration workflows.

Technology Category

Search and Optimization: Distributed SearchMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Schema matching is a foundational task in enterprise data integration, aiming to align disparate data sources. While traditional methods handle simple one-to-one table mappings, they often struggle with complex multi-table schema matching in real-world applications. We present LLMatch, a unified and modular schema matching framework. LLMatch decomposes schema matching into three distinct stages: schema preparation, table-candidate selection, and column-level alignment, enabling component-level evaluation and future-proof compatibility. It includes a novel two-stage optimization strategy: a Rollup module that consolidates semantically related columns into higher-order concepts, followed by a Drilldown module that re-expands these concepts for fine-grained column mapping. To address the scarcity of complex semantic matching benchmarks, we introduce SchemaNet, a benchmark derived from real-world schema pairs across three enterprise domains, designed to capture the challenges of multi-table schema alignment in practical settings. Experiments demonstrate that LLMatch significantly improves matching accuracy in complex schema matching settings and substantially boosts engineer productivity in real-world data integration.
Problem

Research questions and friction points this paper is trying to address.

Develops a unified framework for complex multi-table schema matching
Introduces a two-stage optimization strategy for semantic column alignment
Addresses lack of benchmarks for real-world multi-table schema challenges
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modular framework with three-stage schema matching
Two-stage optimization: Rollup and Drilldown modules
SchemaNet benchmark for complex semantic matching
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Sha Wang
Sha Wang
Singapore Management University
Y
Yuchen Li
Singapore Management University
H
Hanhua Xiao
Singapore Management University
B
Bing Tian Dai
Singapore Management University
Roy Ka-Wei Lee
Roy Ka-Wei Lee
Singapore University of Technology and Design
Trust and SafetySocial ComputingComputational Social ScienceNatural Language Processing
Yanfei Dong
Yanfei Dong
PayPal
L
Lambert Deng
DBS Bank