Score
Designs and implements SQL queries, views, and data-transformation pipelines to extract, aggregate, and analyze structured relational data (including BigQuery or other SQL dialects) and to produce analytical datasets, dashboards, and reports; often integrates SQL with Python or Excel for further analysis and processing. Analyzes query logs and query performance to mine usage patterns, validate and compute metrics, and optimize queries and analytics workflows.
Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.
SQL/PGQ standards enable hybrid relational and graph query processing, yet experiments reveal substantial performance disparities—even for semantically equivalent SQL and PGQ queries—due to syntax-driven, model-isolated optimization in existing systems, lacking cross-model synergy. This paper proposes the first holistic unified optimization framework that transcends syntactic boundaries between SQL and PGQ, instead leveraging query semantics and data characteristics to automatically select optimal execution strategies. Key technical innovations include: (i) cross-model query normalization via equivalence-aware analysis; (ii) cost-aware operator fusion; and (iii) adaptive execution plan generation. Experimental evaluation demonstrates that the framework significantly narrows performance gaps among functionally identical queries, achieving an average 3.2× speedup. It thus provides critical enabling technology for efficient, standardized SQL/PGQ deployment.
Text-to-SQL systems for enterprise data lakes aim to empower non-technical users (e.g., product and operations staff) to autonomously derive data insights via natural language. Addressing key enterprise challenges—dynamic schemas, high query complexity, and deep contextual dependencies—this work proposes: (1) a multi-source dynamic knowledge graph integrating database metadata, query logs, and domain documentation to enable schema-aware table recommendation and semantic retrieval; (2) a Text-to-SQL agent with automated syntactic error correction and retrieval-augmented generation (RAG) to enhance robustness on complex queries; and (3) an intelligent chat interface supporting multi-turn interaction, rich UI feedback, and query debugging. Deployed at LinkedIn, the system achieved over 300 weekly active users. On an internal benchmark, 53% of responses were correct or semantically equivalent. Ablation studies confirm significant performance gains from each component.
本文提出一种将线性Datalog程序编译成递归SQL查询的方法,通过中间语言Midlog转换,利用现有数据库引擎执行程序分析,显著提高性能。
In modern data analytics, entities and their attribute relationships often span multiple granularities, complicating critical attribute derivation and target entity retrieval. Existing OLAP operators, window functions, and aggregation constructs suffer from limitations in composability, formal expressiveness, and runtime performance. To address this, we propose Multi-Relational Algebra (MRA), a novel algebraic framework grounded in the “slice”—a semantic unit comprising a region of tuples and an associated feature table—thereby transcending conventional single-table or single-column constraints. MRA supports dynamic heterogeneous schema modeling and cross-schema composition, and introduces a formal algebraic system, a slice-based computational model, a unified logical execution engine, and a query optimization framework tailored for data insight discovery. The system has been deployed in production, supporting millions of daily operations and effectively handling complex analytical tasks that resist modeling under traditional relational paradigms.
本文探讨了使用列式关系引擎和图查询语言处理大规模图分析问题,证明其性能优于原生图引擎,并指出节点/边模型在关系表中是冗余的。
This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.
Existing Text-to-SQL systems excel in static, single-turn settings but struggle with realistic multi-turn interactions where user intent dynamically evolves—common in financial and business analytics—and lack dedicated evaluation benchmarks. To address this, we propose DySQL-Bench, the first benchmark for dynamic interactive Text-to-SQL. It employs a structured schema-guided LLM-based task generation pipeline, featuring two-stage automated synthesis, interaction filtering, and expert validation to guarantee 100% executable and semantically correct SQL. A novel LLM-simulated user enables closed-loop interactive evaluation. DySQL-Bench spans 13 domains and comprises 1,072 high-quality multi-turn tasks. Experimental results show that GPT-4o achieves only 58.34% overall accuracy and 23.81% Pass@5, underscoring both the challenge of dynamic interaction and the benchmark’s utility for advancing robust, interactive semantic parsing.
This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.
This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.