Score
Designs and conducts analyses and visualizations that extract, validate, and summarize patterns, trends, anomalies, and relationships from structured or unstructured data. Builds metrics, statistical tests, and explanatory narratives that translate raw data into interpretable observations and recommendations for stakeholders.
Narrative-driven data exploration faces core challenges—including contextual discontinuity across views, difficulty in tracing analytical reasoning paths, and insufficient externalization of intermediate interpretations. Method: We conducted a qualitative empirical study with 48 participants, combining in-depth interviews and task-based observations, to code and thematically analyze multi-stage dynamic analytical behaviors. Contribution/Results: The study systematically identifies three critical impediments and derives three design principles for supporting narrative evolution in visual analytics: (1) enforcing cross-view contextual consistency, (2) explicitly tracking reasoning trajectories, and (3) structurally externalizing intermediate interpretations. These principles are operationalized into concrete interaction mechanisms and practical guidelines. The work advances visual analytics systems from static chart presentation toward next-generation tools that actively support dynamic, iterative narrative construction.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
This study addresses the lack of systematic understanding regarding the creation, application, and organizational impact of data visualization style guides. Through interviews with nine authors from journalism, government, and industry, complemented by a cross-case analysis of 26 published guides, the paper proposes the PRISM socio-technical framework to elucidate their operational logic across four dimensions: Purpose, Rules and mechanisms, Institutional enforcers, and Situational flexibility. The findings reveal an inherent tension between standardization and adaptability, demonstrating that publicly available guides represent only partial manifestations of more comprehensive internal systems. By unpacking how these guides function in practice, the research offers both theoretical grounding and novel practical insights for the future development of visualization design standards.
This study investigates the misalignment between data workers’ implicit cognitive models of complex hierarchical data (e.g., nested tables) and the explicit data models encoded in analysis code, and how such misalignment negatively impacts analytical efficiency and accuracy. Method: Through semi-structured interviews, cognitive sketching, and reflexive thematic coding with 10 collaborative data practitioners, we systematically identify divergent, coexisting cognitive models within teams. Contribution/Results: We introduce the novel concept of “parallel risk”—a form of collaborative breakdown arising from persistent cognitive misalignment between data model designers and end users. All participants exhibited internal representations inconsistent with the true data structure, leading to systematic reasoning errors. Based on these findings, we derive human-centered design principles and intervention strategies for analytical tools that promote cognitive alignment. This work establishes a theoretical foundation and practical framework for improving usability in data engineering and visualization systems.
Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.
This work addresses the limited capacity of existing online recruitment platforms to analyze multidimensional attribute relationships and hierarchical market structures. We propose an interactive visual analytics system designed for both job seekers and HR professionals, which, through a coordinated multiple-view design, enables the first integrated exploration of cross-hierarchical, multidimensional hiring data—spanning from macro-level industry trends to micro-level emerging roles. The system integrates data aggregation, pattern recognition, and multi-scale visualization techniques to support comprehensive labor market analysis. Case studies demonstrate its effectiveness in revealing regional salary distributions, characterizing industry evolution trajectories, and accurately identifying high-demand emerging positions.
This work addresses the limitations of existing exploratory data analysis systems, which struggle to effectively detect outliers at the data fact level and suffer from inconsistent metrics due to heterogeneous analytical scopes. To overcome these challenges, we propose FOX, a visual analytics system that first groups data facts into consistent analytical scopes and then integrates distributional and pattern-based features to establish a unified anomaly scoring mechanism. FOX further introduces an interactive paradigm combining a matrix overview with multiple coordinated detail views to facilitate comprehensive exploration. Through case studies on two public datasets and a user study involving twelve participants, we demonstrate that FOX effectively supports the detection, in-depth analysis, and interpretable exploration of fact-level anomalies.
Existing visual analytics workflows are predominantly described in unstructured textual form, hindering systematic comparison, reuse, and practical guidance. This work proposes ATWL, a formal, declarative language for modeling visual analytics workflows through a modular ontology grounded in eight artifact types and standardized intents. For the first time, this approach enables structured, machine-interpretable representations of such workflows. Leveraging large language models, the authors automatically extract workflows from academic papers to construct a reusable repository comprising 17 annotated instances. Empirical evaluation demonstrates that ATWL effectively uncovers cross-workflow structural patterns and yields more compact, structured, and extensible analytical recommendations than original narrative descriptions, thereby facilitating efficient in-context reuse and adaptation.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.