data analysis tools

Designs, builds, configures, and evaluates software components, scripts, libraries, and user interfaces (for example, notebooks, dashboards, and command‑line tools) that ingest, clean, transform, aggregate, and visualize structured and unstructured data to enable exploratory, statistical, and machine‑learning analyses. Implements reproducible data pipelines and workflows, integrates heterogeneous data sources, optimizes performance and scalability, and develops tests, metrics, and documentation to validate data quality and analytical correctness.

dataanalysistools

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.72
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.

analysis reasoningassumptionsdata analysis

DataLens: ML-Oriented Interactive Tabular Data Quality Dashboard

Jan 28, 2025
MA
Mohamed Abdelaal
🏛️ Software AG | TU Darmstadt

Existing data management tools suffer from limited automation, poor interactivity, and insufficient integration with ML workflows, compromising data quality and hindering analytical and modeling performance. To address this, we propose an interactive, ML-oriented tabular data quality dashboard that establishes an adaptive, human-in-the-loop + ML-driven data cleaning闭环. Our approach integrates data profiling, multi-strategy error detection and repair—including statistical analysis, rule-based engines, and supervised/semi-supervised models—while supporting expert rule validation and labeling. Cleaning strategies are iteratively refined using downstream model performance feedback. Furthermore, we unify DataSheets, MLflow, and Delta Lake to ensure reproducibility, traceability, and versioning of the cleaning pipeline. Experiments across multiple benchmark datasets demonstrate significant improvements: error identification rate and repair accuracy increase notably, downstream ML models achieve an average 7.2% accuracy gain, and cleaning time decreases by 40%.

Data ManagementData QualityMachine Learning Integration

Flowco: Rethinking Data Analysis in the Age of LLMs

Apr 18, 2025
SN
Stephen N. Freund
🏛️ Williams College | University of California, Los Angeles | University of Massachusetts Amherst | Amazon Web Services

In data science practice, while large language models (LLMs) can automatically generate analytical code, they lack support for fine-grained control, intermediate result validation, and iterative optimization—compromising analysis controllability, verifiability, and reproducibility. To address this, we propose Flowco: a hybrid, proactive visual dataflow programming framework that uniquely embeds LLMs throughout the entire analytical workflow—spanning code generation, debugging, validation, and iteration. Flowco integrates visual dataflow modeling, LLM-augmented collaborative reasoning, and traceable execution graphs. A user study demonstrates that Flowco significantly improves novices’ efficiency in constructing, debugging, and optimizing analytical tasks, while preserving usability and simultaneously ensuring controllability, verifiability, and reproducibility. By unifying human-in-the-loop interaction with LLM intelligence in a structured, auditable environment, Flowco establishes a novel paradigm for democratizing data science in the LLM era.

Enabling non-experts to conduct data analyses using LLMsProviding fine-grained control and verification in analysis stepsSupporting iterative refinement of data analysis workflows

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Existing large language model (LLM)-driven data analysis tools are often confined to isolated subtasks and struggle to support end-to-end executable analytical workflows. This work proposes an autonomous, sandboxed, and auditable end-to-end system that leverages LLMs for action planning, iteratively generating structured operations, executing code in a secure environment, and integrating streaming traceability with intermediate result previews. By unifying a structured action backend, sandboxed execution, and an interactive visual interface—features integrated here for the first time—the system enables users to drive complete analytical workflows using only natural language. Users can inspect, modify, and export the entire process and its outputs directly within a web browser, ensuring full reproducibility, editability, and transparency throughout the analytical pipeline.

action tracedata analysisend-to-end workflow

This work addresses the significant limitations of spreadsheet-based analysis in reproducibility, auditability, version control, and automation. It proposes a migration pathway from Excel to research-grade analytical workflows by leveraging Python’s pandas library as a bridge. The study introduces an innovative set of Excel-to-pandas mapping rules, categorizes nine canonical workflow patterns, and compiles a catalog of common failure modes. Seven end-to-end real-world examples demonstrate the approach in practice. By retaining Excel as a familiar interface for input and output while integrating version control, automated refreshing, and seamless incorporation of statistical and machine learning methods, the proposed framework enables governed, reproducible, and auditable tabular data analysis.

auditabilitydata analysisgovernance

This work addresses the challenge of irreproducibility in data analysis scripts, which often stems from implicit assumptions—such as specific package versions, expected data formats, or undocumented manual interventions. The paper proposes a static analysis approach tailored to data analysis workflows that, for the first time, unifies diverse implicit assumptions into inferable constraint models. By leveraging customized program analysis and example-driven modeling, the authors develop a prototype system capable of automatically identifying these hidden assumptions, extracting executable preconditions, and generating verifiable constraints. The resulting framework supports runtime validation and automatic documentation generation, substantially enhancing script executability, reproducibility, and interpretability.

code constraintsdata analysisimplicit assumptions

Hot Scholars

KL

Kai Li

University of Chinese Academy of Sciences & City University of Hong Kong
Computer VisionMultimodal Language ModelRemote Sensing
SK

Sumanta Kundu

Postdoctoral Researcher, SISSA, Trieste
Statistical physicsPhase transitionPolymers & knotsActive polymers
EO

Enzo Orlandini

Department of Physics and Astronomy, Universita' di Padova
Statistical physicsSoft MatterBiophysics