data analysis

Designs and implements processes to collect, clean, transform, and explore datasets; builds statistical summaries, visualizations, hypothesis tests, and predictive or descriptive models to quantify patterns and estimate uncertainty. Communicates findings, documents methods and assumptions, and evaluates data quality and model performance to support decision making.

dataanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.84
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.

analysis reasoningassumptionsdata analysis

A Tale of Two Models: Understanding Data Workers' Internal and External Representations of Complex Data

Jan 16, 2025
CS
Connor Scully-Allison
🏛️ University of Utah | Davidson College | Lawrence Livermore National Lab

This study investigates the misalignment between data workers’ implicit cognitive models of complex hierarchical data (e.g., nested tables) and the explicit data models encoded in analysis code, and how such misalignment negatively impacts analytical efficiency and accuracy. Method: Through semi-structured interviews, cognitive sketching, and reflexive thematic coding with 10 collaborative data practitioners, we systematically identify divergent, coexisting cognitive models within teams. Contribution/Results: We introduce the novel concept of “parallel risk”—a form of collaborative breakdown arising from persistent cognitive misalignment between data model designers and end users. All participants exhibited internal representations inconsistent with the true data structure, leading to systematic reasoning errors. Based on these findings, we derive human-centered design principles and intervention strategies for analytical tools that promote cognitive alignment. This work establishes a theoretical foundation and practical framework for improving usability in data engineering and visualization systems.

Analytical CommunicationComplex Data StructuresData Representation

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

Designing a Data Science simulation with MERITS: A Primer

Mar 13, 2024
CF
Corrine F. Elliott
🏛️ University of California, Berkeley | University of Michigan | University of Regensburg | University of California, San Francisco

Long-standing deficiencies in standardized, high-quality criteria for data science simulation studies have led to inconsistent design practices, poor reproducibility, and limited external validity. To address this, we propose MERITS—a simulation quality framework for trustworthy data science—systematically defining six core dimensions: Modularity, Efficiency, Realism, Stability, Intuitiveness, and Transparency. MERITS is the first to operationalize the PCS (Predictability-Computability-Stability) theory into concrete design principles and innovatively introduces a “cooking metaphor” to structure simulations as executable “recipes.” The framework includes 13 actionable design guidelines and is validated through empirical reconstruction of existing studies. Designed for cross-disciplinary applicability, MERITS has been successfully applied to diagnostic reconstructions of prior work, yielding substantial improvements in interpretability, reproducibility, and external validity.

Defining high-quality Data Science simulation standardsProposing MERITS framework for simulation study designProviding guidelines for trustworthy Data Science practices

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Hot Scholars

JH

Jennifer Hu

Johns Hopkins University
Computational linguisticsCognitive sciencePragmaticsMachine learning
QA

Qingyao Ai

Associate Professor, Dept. of CS&T, Tsinghua University
Information RetrievalMachine Learning
WG

Wanling Gao

Institute Of Computing Technology Chinese Academy Of Sciences
Big data and AI benchmarking,Computer architecture
YD

Yue Duan

Singapore Management University
System SecuritySoftware EngineeringBlockchain securityAI security