RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL

📅 2025-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address scalability bottlenecks in Text-to-SQL for enterprise-scale databases, this paper proposes a domain-agnostic, retrieval-augmented schema linking framework. Methodologically, it introduces a modular semantic indexing architecture that decouples database schemas and metadata into fine-grained semantic units for independent vectorization; further, it designs a table-level prioritized identification mechanism coupled with column-level contextual fusion, integrating chunked indexing, semantic retrieval, recall re-ranking, and context budget control. Crucially, the approach fully leverages intrinsic semantic cues from raw metadata without requiring domain-specific fine-tuning, enabling high-precision schema matching. Experimental results demonstrate that our method consistently outperforms mainstream baselines across heterogeneous, multi-source data catalogs—achieving both high recall and high accuracy. Its plug-and-play compatibility facilitates seamless enterprise deployment.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Data Mining & Knowledge Management: Intelligent Query ProcessingSearch and Optimization: Distributed Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphs
📝 Abstract
Despite advances in large language model (LLM)-based natural language interfaces for databases, scaling to enterprise-level data catalogs remains an under-explored challenge. Prior works addressing this challenge rely on domain-specific fine-tuning - complicating deployment - and fail to leverage important semantic context contained within database metadata. To address these limitations, we introduce a component-based retrieval architecture that decomposes database schemas and metadata into discrete semantic units, each separately indexed for targeted retrieval. Our approach prioritizes effective table identification while leveraging column-level information, ensuring the total number of retrieved tables remains within a manageable context budget. Experiments demonstrate that our method maintains high recall and accuracy, with our system outperforming baselines over massive databases with varying structure and available metadata. Our solution enables practical text-to-SQL systems deployable across diverse enterprise settings without specialized fine-tuning, addressing a critical scalability gap in natural language database interfaces.
Problem

Research questions and friction points this paper is trying to address.

Scaling text-to-SQL to enterprise databases with massive schemas
Leveraging database metadata for semantic context in retrieval
Maintaining accuracy without domain-specific fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Component-based retrieval architecture for schemas
Semantic indexing of database metadata units
Prioritizes table identification within context budget