🤖 AI Summary
This study addresses the lack of evaluation benchmarks for large language models (LLMs) handling multi-syntax and temporal-intent queries over time-series databases. We propose the first multi-syntax time-series Text-to-SQL benchmark, constructed via a human-AI collaborative workflow augmented with AI-assisted pipelines. The benchmark comprises 6,125 high-quality question-answer pairs spanning 97 databases, 23 SQL syntax variants, and cross-domain temporal intents. Based on this resource, we conduct a systematic evaluation of several state-of-the-art LLMs. Results demonstrate that the best-performing model achieves an execution accuracy of only 48.98%, substantially lagging behind the human baseline of 87.34%. These findings highlight critical challenges such as syntactic heterogeneity in time-series Text-to-SQL tasks and provide a robust foundation for advancing future research in this domain.
📝 Abstract
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.