🤖 AI Summary
This study addresses the significant accuracy degradation of Text-to-SQL systems across diverse database dialects by proposing relational algebra query plans as an intermediate representation. The method leverages large language models to generate dialect-agnostic query plans, which are subsequently translated into executable SQL via a deterministic compiler. Additionally, a MetricName indicator is introduced to eliminate semantic ambiguities. Experimental results demonstrate that this framework effectively restores cross-dialect portability across thirteen models, achieving superior performance compared to conventional direct SQL supervision methods after fine-tuning.
📝 Abstract
Text-to-SQL systems are typically trained and evaluated on a single dialect (SQLite), yet production deployments span PostgreSQL, MySQL, ClickHouse, and beyond. We show that this single-dialect assumption leads to a substantial drop in cross-dialect accuracy for every model we tested. The drop persists across scale, architecture, and even purpose-built text-to-SQL systems. We argue that the fix is to change the generation target: instead of asking an LLM to emit dialect-specific SQL, we have it emit a dialect-agnostic relational algebra query plan, which a deterministic compiler then renders into SQL for any supported backend. Across thirteen models from 3B to frontier scale, this restores cross-dialect portability nearly uniformly, at a small cost in peak accuracy on the model's home dialect for capable prompted models and none once fine-tuned on plans; under matched fine-tuning, plan supervision yields a stronger model than SQL supervision. We also introduce MetricName, a question-aware result-set comparator needed to evaluate fairly across dialects, where existing metrics confound semantic errors with benign cross-dialect variation. More broadly, the result is a reminder that a generation target chosen for execution is not necessarily the one that maximizes generation quality.