TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of evaluation benchmarks for large language models (LLMs) handling multi-syntax and temporal-intent queries over time-series databases. We propose the first multi-syntax time-series Text-to-SQL benchmark, constructed via a human-AI collaborative workflow augmented with AI-assisted pipelines. The benchmark comprises 6,125 high-quality question-answer pairs spanning 97 databases, 23 SQL syntax variants, and cross-domain temporal intents. Based on this resource, we conduct a systematic evaluation of several state-of-the-art LLMs. Results demonstrate that the best-performing model achieves an execution accuracy of only 48.98%, substantially lagging behind the human baseline of 87.34%. These findings highlight critical challenges such as syntactic heterogeneity in time-series Text-to-SQL tasks and provide a robust foundation for advancing future research in this domain.
📝 Abstract
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.
Problem

Research questions and friction points this paper is trying to address.

Time-Series Databases
Text-to-Query
Benchmark
Large Language Models
Query Syntax
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-Series Databases
Text-to-Query
Multi-Syntax Benchmark
Large Language Models
Schema Linking
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
F
Fei Lyu
College of Computer Science and Electronic Engineering, Hunan University
Z
Zhiyi Peng
College of Computer Science and Electronic Engineering, Hunan University
J
Jiaming Liu
College of Computer Science and Electronic Engineering, Hunan University
Yixuan Yang
Yixuan Yang
PhD Candidate, University of Warwick | SUSTech
3D Computer VisionPoint Cloud3D ReconstructionEmbodied AI
Changjian Chen
Changjian Chen
Associate Professor, Hunan University
Interactive Machine LearningData-Centric AI
Zhuo Tang
Zhuo Tang
Central South University
J
Jiapeng Zhang
College of Computer Science and Electronic Engineering, Hunan University
Kenli Li
Kenli Li
Cheung Kong Professor, Hunan University
High-performance ComputingParallel and Distributed ProcessingAI and Big Data