Score
Designs and implements end-to-end processes for extracting insight from structured and unstructured data, including data acquisition, cleaning, feature engineering, exploratory and statistical analysis, and model development. Builds and evaluates predictive and inferential models, reproducible data pipelines, visualizations, experiments, and reports, and prepares solutions for validation, monitoring, and deployment to support decision-making.
Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.
Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.
This work addresses the challenge scientists face in efficiently transforming raw sensor data streams into actionable insights across edge-cloud infrastructures, hindered by the need for cross-domain expertise to manage heterogeneous systems and emerging platforms such as DPUs, which impedes rapid prototyping. To overcome this barrier, the authors propose a novel paradigm that integrates pattern-based workflow engineering with AI-assisted development. Implemented on the FABRIC testbed using the Pegasus workflow system and exemplified by the Orcasound hydrophone workflow, this approach enables swift construction of applications for air quality, seismic, and soil moisture monitoring. The framework supports modular extensibility and edge deployment, substantially lowering the barrier for non-expert users to iteratively develop distributed applications. Empirical validation across multiple use cases demonstrates its effectiveness in enhancing development efficiency, accelerating prototyping cycles, and accumulating practical deployment experience.
This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.