Score
Designs and performs systematic analyses and visualizations of datasets (tabular or large-scale) to discover patterns, trends, anomalies, distributional properties, central tendency and dispersion, stationarity and dependence, ordinal/ranking structure, and latent factor structure. Implements summary statistics and statistical tests, exploratory factor methods, data provenance modeling and tracing, dataset difficulty/diversity metrics, and scalable data processing or generation pipelines to synthesize insights across tables and produce informative visualizations and reports.
This study investigates the misalignment between data workers’ implicit cognitive models of complex hierarchical data (e.g., nested tables) and the explicit data models encoded in analysis code, and how such misalignment negatively impacts analytical efficiency and accuracy. Method: Through semi-structured interviews, cognitive sketching, and reflexive thematic coding with 10 collaborative data practitioners, we systematically identify divergent, coexisting cognitive models within teams. Contribution/Results: We introduce the novel concept of “parallel risk”—a form of collaborative breakdown arising from persistent cognitive misalignment between data model designers and end users. All participants exhibited internal representations inconsistent with the true data structure, leading to systematic reasoning errors. Based on these findings, we derive human-centered design principles and intervention strategies for analytical tools that promote cognitive alignment. This work establishes a theoretical foundation and practical framework for improving usability in data engineering and visualization systems.
Data visualization instruction in statistics and data science education suffers from a disciplinary disconnect—courses are often offered outside statistics departments, emphasizing narrative and design while neglecting core statistical thinking, particularly statistical inference. Method: This study conducts the first systematic nationwide survey of visualization curricula at top U.S. universities, building a comprehensive database covering 62 institutions through course syllabus analysis, structured faculty surveys, and educational empirical research. Contribution/Results: (1) It documents, for the first time, the disciplinary affiliations and curricular biases of university-level visualization courses; (2) it proposes a novel pedagogical framework centered on statistical thinking, formalizing three inference-oriented teaching principles; and (3) it develops and validates scalable, implementation-ready instructional materials and case studies, offering a concrete, evidence-informed pathway to integrate statistical reasoning into visualization education and advance substantive convergence between statistics pedagogy and visualization practice.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.
This work addresses the limitations of existing exploratory data analysis systems, which struggle to effectively detect outliers at the data fact level and suffer from inconsistent metrics due to heterogeneous analytical scopes. To overcome these challenges, we propose FOX, a visual analytics system that first groups data facts into consistent analytical scopes and then integrates distributional and pattern-based features to establish a unified anomaly scoring mechanism. FOX further introduces an interactive paradigm combining a matrix overview with multiple coordinated detail views to facilitate comprehensive exploration. Through case studies on two public datasets and a user study involving twelve participants, we demonstrate that FOX effectively supports the detection, in-depth analysis, and interpretable exploration of fact-level anomalies.
This study addresses the cognitive bias in scatterplots where “data-induced grouping”—arising from the interplay between data values and visual encoding—leads users to misinterpret spatial arrangements as meaningful patterns. Through two user studies, the authors systematically demonstrate the prevalence of this phenomenon, develop the first perceptual model capable of predicting whether users perceive a given set of points as a coherent group, and propose a visualization intervention strategy that integrates user perception with data reordering. Notably, the model effectively captures users’ tendency to group points based on trends even in nominal data contexts. Applied to visualization diagnosis and optimization, this approach significantly enhances the accuracy and reliability of graphical representations.
Traditional “Table 1” formats inadequately convey details of numerical variables when presenting baseline characteristics across groups, hindering intuitive comparison. This work proposes “snapshot plots,” which transform summary tables into parallel univariate visualizations with consistent color encoding—a specialized form of hammock plots—to substantially enhance intergroup comparability and information density. Built upon the principles of parallel coordinates, the approach uniformly handles mixed variable types and is implemented in Python, accompanied by an interactive web application. Validation on two real-world “Table 1” examples demonstrates that snapshot plots significantly improve readability and facilitate cross-group comparisons.
Narrative-driven data exploration faces core challenges—including contextual discontinuity across views, difficulty in tracing analytical reasoning paths, and insufficient externalization of intermediate interpretations. Method: We conducted a qualitative empirical study with 48 participants, combining in-depth interviews and task-based observations, to code and thematically analyze multi-stage dynamic analytical behaviors. Contribution/Results: The study systematically identifies three critical impediments and derives three design principles for supporting narrative evolution in visual analytics: (1) enforcing cross-view contextual consistency, (2) explicitly tracking reasoning trajectories, and (3) structurally externalizing intermediate interpretations. These principles are operationalized into concrete interaction mechanisms and practical guidelines. The work advances visual analytics systems from static chart presentation toward next-generation tools that actively support dynamic, iterative narrative construction.
This work addresses the limited capacity of existing online recruitment platforms to analyze multidimensional attribute relationships and hierarchical market structures. We propose an interactive visual analytics system designed for both job seekers and HR professionals, which, through a coordinated multiple-view design, enables the first integrated exploration of cross-hierarchical, multidimensional hiring data—spanning from macro-level industry trends to micro-level emerging roles. The system integrates data aggregation, pattern recognition, and multi-scale visualization techniques to support comprehensive labor market analysis. Case studies demonstrate its effectiveness in revealing regional salary distributions, characterizing industry evolution trajectories, and accurately identifying high-demand emerging positions.