statistical analysis

Using descriptive statistics, correlation measures, and inferential tests to quantify differences, distributions, and relationships in data; used to measure gaps (e.g., sim-to-real), characterize distributions and switching patterns, and validate the impact of interventions across tasks and domains.

statisticalanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Assessing Inference Methods

Dec 18, 2019
BF
Bruno Ferman
🏛️ Sao Paulo School of Economics - FGV

This study addresses the uncontrolled false positive rates and misleading inferences arising from commonly used simulation methods in shift-share designs. We systematically evaluate prevailing inferential approaches in empirical research through a suite of multilevel simulation experiments. By comparing Monte Carlo analysis with counterfactual data-generating mechanisms, we uncover non-monotonic trade-offs among fidelity, sensitivity, and risk of misdirection across simulation designs. We propose a novel “progressive-fidelity simulation framework,” demonstrating that low-fidelity simulations suffice to expose fundamental inferential flaws, whereas high-fidelity simulations detect subtle, previously overlooked biases—substantially improving detection power. The framework balances interpretability and computational efficiency, offering a reproducible and scalable paradigm for assessing the robustness of causal inference methods.

Analyzing trade-offs in simulation-based inference assessmentsEvaluating reliability of inference methods for false-positive controlProposing alternatives to misleading shift-share design evaluations

How can the gap between methodological research and statistical practice be bridged? This paper proposes a novel paradigm—“translational simulation studies”—designed to enable applied statisticians to scientifically select and evaluate statistical methods for specific contexts using simulation results, without requiring advanced programming skills or substantial time investment. Methodologically, we develop a Shiny-based interactive simulation platform complemented by a modular, extensible R code framework, allowing users to flexibly specify parameters and scenarios. Our key contribution lies in directly packaging methodological findings into rigorously validated, ready-to-use tools that balance statistical rigor, flexibility, and usability. We demonstrate feasibility and practical utility through two empirical applications: power evaluation in clinical trials and measurement error analysis in regression models. Results show that this approach significantly enhances the efficiency and accessibility of translating methodological advances into real-world statistical practice.

Addressing the disconnect between methodological simulation studies and applied statistical practiceOvercoming challenges in adapting published simulation results to real-world applicationsProviding accessible simulation tools for applied statisticians with limited programming expertise

Standardized Descriptive Index for Measuring Deviation and Uncertainty in Psychometric Indicators

Dec 24, 2025
MD
Mark Dominique Dalipe Muñoz
🏛️ Iloilo Science and Technology University

Current psychometric practice relies on separate descriptive statistics (mean and standard deviation) to assess item quality, lacking a standardized diagnostic tool that integrates both to quantify raw deviation from scale midpoints and its uncertainty—especially problematic in small-sample settings. Method: We propose a Standardized Projected Deviation Index (SPDI), derived from Cohen’s *d*, which unifies the magnitude and variability of an item’s raw deviation from the scale midpoint into a single, bounded, scale-invariant, and bias-controlled quality metric. Results: Through theoretical derivation and small-sample simulation studies, we demonstrate that SPDI is interpretable, invariant across items, and robust under limited data. It provides empirically grounded, actionable thresholds for identifying formative indicator redundancy and evaluating reflective indicator consistency—thereby enabling objective, quantitative item-level diagnostics in both exploratory and confirmatory measurement contexts.

Develop a standardized index for item deviation in psychometricsEstablish thresholds for redundancy and consistency in indicatorsMeasure item quality by combining mean and standard deviation

From'What-is'to'What-if'in Human-Factor Analysis: A Post-Occupancy Evaluation Case

Nov 28, 2025
XC
Xia Chen
🏛️ Technische Universität München | University of California, Berkeley | Leibniz Universität Hannover

Traditional human factors analysis relies on correlation-based testing, which only addresses descriptive “what is” questions and fails to resolve causal “if–then” queries. It remains vulnerable to confounding and collider variables, leading to biased decision-making. This paper proposes a paradigm shift from descriptive analysis to causal inference. Leveraging post-occupancy evaluation data from built environments, we integrate causal discovery algorithms, structural equation modeling, and counterfactual intervention analysis to construct directed causal networks among variables—explicitly distinguishing descriptive from interventional queries. Our key contribution is the first systematic application of a rigorous causal inference framework to the human factors domain, enabling robust identification of intervention priorities and hierarchical causal pathways. Empirical validation demonstrates that this approach significantly enhances the accuracy and scientific rigor of human factors–driven system optimization decisions, establishing a novel, interpretable, and actionable causal analysis paradigm for human factors engineering.

Applies causal inference to avoid bias from confounding variablesDistinguishes descriptive from causal questions in human-factor analysisUses post-occupancy data to reveal intervention effects for optimization

This study addresses the analysis of nominal single-choice questionnaire data by proposing a novel archetypal analysis method that extends archetypal analysis—traditionally limited to continuous variables—to nominal variable settings for the first time. The approach represents each individual as a convex combination of actual extreme response patterns, termed archetypes, thereby effectively identifying both typical and boundary cases. Unlike conventional archetypal analysis and its probabilistic variants, which are ill-suited for nominal data, the proposed method explicitly models the convex geometric structure inherent in categorical responses. Experimental results on the German Credit dataset demonstrate that the method substantially enhances the interpretability and structural insight into nominal data, offering a new paradigm for questionnaire data analysis.

archetypal casesarchetypoid analysismultiple-choice questions

Latest Papers

What's happening recently
View more

This study addresses a critical limitation in existing design-based simulations used to evaluate inference methods, which often overstate bias induced by spatial correlation due to unrealistic data-generating mechanisms. In particular, share-shift designs that fix outcomes and resample shocks conflate true treatment effects with error dependence structures, leading to misleading assessments. To remedy this, the paper proposes an improved simulation framework that more accurately models error dependence and avoids spurious entanglement between treatment effects and error terms, thereby better approximating real-world data-generating processes. Integrating resampling techniques with share-shift analysis, the proposed approach substantially enhances the reliability of inference evaluation across multiple empirical applications, underscoring the essential role of aligning simulation designs with genuine underlying mechanisms for valid inference assessment.

data-generating processdesign-based simulationsinference validity

This study addresses the lack of systematic understanding regarding the effectiveness and usage practices of univariate distribution visualizations across diverse tasks and user groups. Through a mixed-methods approach—combining a click-based selection experiment and survey with 215 participants alongside in-depth interviews with five visualization practitioners—the work systematically evaluates the accuracy, user preferences, and common misinterpretations associated with boxplots, violin plots, jittered scatterplots, and histograms in typical analytical tasks. For the first time, it integrates task performance, subjective preference, and real-world practice, revealing a frequent mismatch between chart familiarity and task accuracy, thereby challenging the assumption that commonly used or conventional visualizations are inherently optimal. The findings demonstrate significant performance differences among chart types in low-level tasks, with widely adopted histograms and boxplots not consistently outperforming alternatives.

chart effectivenesstask performanceunivariate distribution

This study addresses the computational complexity associated with calculating quantiles of the inverse normal distribution, Student’s t-distribution, and outlier rejection criteria in hypothesis testing. To overcome the reliance on table lookups or iterative numerical methods, the paper proposes concise and highly accurate analytical approximations formulated as closed-form expressions. These approximations significantly reduce computational overhead while maintaining precision sufficient for practical statistical applications. The resulting method offers substantial gains in computational efficiency, making it particularly well-suited for resource-constrained environments or scenarios requiring rapid statistical inference. By bridging theoretical rigor with practical utility, the approach delivers both methodological insight and real-world applicability.

computational simplificationhypothesis testingoutlier rejection

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Official statistics often exhibit complex structures across geographic and subpopulation dimensions that traditional tabular formats struggle to convey effectively, thereby hindering policymakers’ comprehension and application. This study introduces linked micromaps as a visualization framework that systematically integrates descriptive statistics, multivariate relationships, ranking structures, and spatiotemporal heterogeneity to enable intuitive exploration of high-dimensional official data. The approach substantially enhances the interpretability and readability of statistical information, uncovering latent patterns while also offering new avenues for subsequent modeling and uncertainty quantification. By doing so, it expands the potential of linked micromaps in public policy analysis and social science research.

data visualizationexploratory analysisgeographic variation

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain