sql data analysis

Designs and implements SQL queries, views, and data-transformation pipelines to extract, aggregate, and analyze structured relational data (including BigQuery or other SQL dialects) and to produce analytical datasets, dashboards, and reports; often integrates SQL with Python or Excel for further analysis and processing. Analyzes query logs and query performance to mine usage patterns, validate and compute metrics, and optimize queries and analytics workflows.

sqldataanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.94
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$180K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models

Dec 10, 2024
JZ
Jun-Peng Zhu
🏛️ East China Normal University | PingCAP

Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.

Automate SQL-based cross-domain exploratory data analysis.Enhance data visualization through text-to-SQL and text-to-visualization.Improve cross-domain generalization and user intent clarity in EDA.

Towards Cross-Model Efficiency in SQL/PGQ

May 12, 2025
HR
Hadar Rotschield
🏛️ Hebrew University

SQL/PGQ standards enable hybrid relational and graph query processing, yet experiments reveal substantial performance disparities—even for semantically equivalent SQL and PGQ queries—due to syntax-driven, model-isolated optimization in existing systems, lacking cross-model synergy. This paper proposes the first holistic unified optimization framework that transcends syntactic boundaries between SQL and PGQ, instead leveraging query semantics and data characteristics to automatically select optimal execution strategies. Key technical innovations include: (i) cross-model query normalization via equivalence-aware analysis; (ii) cost-aware operator fusion; and (iii) adaptive execution plan generation. Experimental evaluation demonstrates that the framework significantly narrows performance gaps among functionally identical queries, achieving an average 3.2× speedup. It thus provides critical enabling technology for efficient, standardized SQL/PGQ deployment.

Need for holistic cross-model optimization approachPerformance gaps between SQL and graph query modelsSeparate optimizations for SQL and graph queries

Text-to-SQL for Enterprise Data Analytics

Jul 18, 2025
AC
Albert Chen
🏛️ LinkedIn

Text-to-SQL systems for enterprise data lakes aim to empower non-technical users (e.g., product and operations staff) to autonomously derive data insights via natural language. Addressing key enterprise challenges—dynamic schemas, high query complexity, and deep contextual dependencies—this work proposes: (1) a multi-source dynamic knowledge graph integrating database metadata, query logs, and domain documentation to enable schema-aware table recommendation and semantic retrieval; (2) a Text-to-SQL agent with automated syntactic error correction and retrieval-augmented generation (RAG) to enhance robustness on complex queries; and (3) an intelligent chat interface supporting multi-turn interaction, rich UI feedback, and query debugging. Deployed at LinkedIn, the system achieved over 300 weekly active users. On an internal benchmark, 53% of responses were correct or semantically equivalent. Ablation studies confirm significant performance gains from each component.

Building a practical enterprise Text-to-SQL solution for data analyticsCreating a knowledge graph to capture dynamic data semanticsDeveloping an interactive chatbot for self-serve data insights

Multi-Relational Algebra and Its Applications to Data Insights

Nov 08, 2023
XW
Xi Wu
🏛️ Google | Boston University | UW-Madison | Microsoft | Celonis

In modern data analytics, entities and their attribute relationships often span multiple granularities, complicating critical attribute derivation and target entity retrieval. Existing OLAP operators, window functions, and aggregation constructs suffer from limitations in composability, formal expressiveness, and runtime performance. To address this, we propose Multi-Relational Algebra (MRA), a novel algebraic framework grounded in the “slice”—a semantic unit comprising a region of tuples and an associated feature table—thereby transcending conventional single-table or single-column constraints. MRA supports dynamic heterogeneous schema modeling and cross-schema composition, and introduces a formal algebraic system, a slice-based computational model, a unified logical execution engine, and a query optimization framework tailored for data insight discovery. The system has been deployed in production, supporting millions of daily operations and effectively handling complex analytical tasks that resist modeling under traditional relational paradigms.

Handling multi-granular data analytics challengesLimitations in existing relational algebra constructsNeed for modular multi-granular analysis framework

Latest Papers

What's happening recently
View more

This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.

auditabilitydata analysisgovernance

Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration

Oct 30, 2025
LS
Linzhuang Sun
🏛️ University of Chinese Academy of Sciences | Peking University | Tsinghua University

Existing Text-to-SQL systems excel in static, single-turn settings but struggle with realistic multi-turn interactions where user intent dynamically evolves—common in financial and business analytics—and lack dedicated evaluation benchmarks. To address this, we propose DySQL-Bench, the first benchmark for dynamic interactive Text-to-SQL. It employs a structured schema-guided LLM-based task generation pipeline, featuring two-stage automated synthesis, interaction filtering, and expert validation to guarantee 100% executable and semantically correct SQL. A novel LLM-simulated user enables closed-loop interactive evaluation. DySQL-Bench spans 13 domains and comprises 1,072 high-quality multi-turn tasks. Experimental results show that GPT-4o achieves only 58.34% overall accuracy and 23.81% Pass@5, underscoring both the challenge of dynamic interaction and the benchmark’s utility for advancing robust, interactive semantic parsing.

Addressing dynamic multi-turn SQL generation for evolving user intentsBenchmarking SQL systems under iterative query refinement scenariosEvaluating model adaptation in real-world interactive database exploration

This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.

large language modelsmodel pipelineNatural Language to SQL

This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.

ANALYZE operatorcube queryingdata analysis

Hot Scholars

XZ

Xuanhe Zhou

Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence
GL

Guoliang Li

Professor, Tsinghua University
DatabaseBig DataCrowdsourcingData Cleaning & Integration
CL

Cuiping Li

Renmin University of China
Databasebig data analysis and mining
CP

Conor Power

UC Berkeley
DatabasesDistributed SystemsCloud Computing
MR

Manuel Rigger

National University of Singapore
Software EngineeringSystemsDatabasesProgramming Languages