🤖 AI Summary
This work addresses the critical limitation of general-purpose Text-to-SQL models, which achieve less than 50% accuracy on production-grade financial databases containing opaque keys. To overcome this, we propose FLINT, a system that introduces a lookup agent to resolve concept mapping challenges and combines expert template retrieval with dynamic parsing for precise SQL generation. Furthermore, FLINT optimizes structured schema linking through foreign key chain traversal pruning. Built upon a large language model and domain-specific agent architecture, FLINT significantly outperforms existing state-of-the-art baselines on a production dataset comprising 359 queries. The system has been successfully deployed in real-world financial data retrieval services, offering an effective semantic parsing solution for complex industrial databases.
📝 Abstract
General-purpose Text-to-SQL systems achieve strong performance on academic benchmarks like Spider and BIRD, where schemas are relatively shallow and column values are often human readable. In production financial databases, where concepts are stored as opaque integer keys rather than human-readable strings, these methods fall below 50%, as even simple queries require multiple joins and filter predicates reference opaque IDs. We present Financial LINking Text-to-SQL (FLINT), a domain-specialized Text-to-SQL system that closes this gap through three key components: (1) a lookup agent that dynamically resolves natural-language concepts to question-specific reference table constraints, (2) embedding-based retrieval of structurally similar query templates from a compact, expert-authored bank, and (3) schema linking that prunes a large table schema to the relevant subset by traversing foreign-key chains, rather than relying on name similarity alone. We evaluate on two datasets totaling 359 questions over production financial schemas. FLINT outperforms various state-of-the-art baselines using the same LLM. The system is deployed in production as part of a financial data retrieval service.