feature correlation analysis

Designs and implements analyses, estimators, and visualizations that quantify, test, and interpret associations among features or variables, using rank-based measures (Spearman/rank correlations), Pearson-type measures, spatial-correlation models, and estimators for ordinal data (polychoric). This includes computing correlation coefficients with standard errors and confidence intervals, hypothesis tests, comparisons across metrics or subsets, modeling latent thresholds for ordinal observations, and diagnostics to localize and assess robustness of correlations.

featurecorrelationanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.83
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of a unified computational, management, and visualization framework for characterizing diverse pairwise associations—such as linear correlation, nonlinear dependence, and Simpson’s paradox—between numeric and categorical variables. Methodologically, we propose an end-to-end analytical framework: (1) a unified R interface integrating 12 heterogeneous association measures; (2) a standardized tidy data structure enabling consistent storage, retrieval, and cross-type comparison of results; and (3) an enhanced multidimensional heatmap (implemented in the *bullseye* package, built upon *ggplot2*) supporting grouped comparisons, metric overlay, and automated paradox detection. Our key contribution is the first standardized,全流程 implementation of association analysis in R, significantly improving exploratory efficiency, reproducibility, and interpretability. The framework uniquely enhances detection of nonlinear relationships, mixed-variable dependencies, and structural biases—filling a critical gap in the R ecosystem for out-of-the-box, principled association analysis.

Enhancing traditional heatmaps with richer relationship visualizationsManaging and visualizing multiple pairwise correlation scoresProviding uniform interface for diverse association calculations

A multivariate spatial model for ordinal survey-based data

Jul 28, 2025
MÁ
Miguel Ángel Beltrán-Sánchez
🏛️ University of Valencia

This study addresses the challenge of spatially joint modeling of multivariate ordinal data across multiple health domains in population surveys. We propose a novel multivariate ordinal spatial model that integrates individual-level covariate effects with hierarchical spatial dependence structures. Methodologically, the model combines a latent multivariate Gaussian process with an ordinal regression framework to simultaneously capture intra-domain variable correlations, individual heterogeneity, geographic spatial clustering, and cross-variable spatial covariance. Applied to mental health indicators from the 2022 Valencia (Spain) Health Survey, the model substantially improves spatial pattern estimation accuracy for all outcomes and identifies geographically coherent clusters exhibiting cross-indicator associations of public health relevance. To our knowledge, this is the first scalable and interpretable statistical framework enabling spatially coherent analysis of high-dimensional ordinal survey data.

Analyzing mental health indicators with geographical dependenciesJointly estimating individual effects and regional patternsModeling correlated ordinal health survey responses spatially

This study addresses the common misuse of Pearson correlation for mixed variable types—such as binary, ordinal, and nominal—which often introduces bias in traditional correlation analyses. To resolve this, the authors introduce smartcor (for R) and pysmartcor (for Python), the first toolkits to systematically support all ten possible combinations of variable types. These packages employ automatic variable-type detection and a rule-based engine to intelligently select the optimal correlation or association method from a repertoire of fourteen, while also providing interpretable justifications for each choice. Monte Carlo simulations demonstrate substantially improved selection accuracy, and real-world case studies reveal meaningful discrepancies between type-aware analyses and naive Pearson correlations, thereby enhancing the reliability of statistical inference.

correlationmethod selectionmixed data

This study addresses the lack of intuitive visualization methods for ordinal regression results, which has hindered their application in fields such as visualization and human-computer interaction. To bridge this gap, the paper proposes, for the first time, the use of modified complementary cumulative distribution function (mCCDF) plots to visualize outputs from cumulative link ordinal regression models. This approach not only fills a critical void in the clear and direct representation of ordinal regression outcomes but also effectively conveys key conclusions consistent with those erroneously derived when treating ordinal variables as continuous. By doing so, the method substantially enhances the interpretability and communicability of model results, offering a principled yet accessible visual framework for practitioners and researchers alike.

CCDF plotsLikert dataordinal regression

Robust Estimation of Polychoric Correlation

Jul 26, 2024
MW
Max Welz
🏛️ Erasmus University Rotterdam | University of Zurich | Harvard University

Polychoric correlation estimation for ordinal rating data is highly sensitive to violations of the underlying normality assumption—e.g., due to careless responding causing local model misspecification. To address this, we propose a novel robust estimator grounded in a weighted score function framework, integrating influence function theory and M-estimation principles. Crucially, it achieves robust polychoric estimation without requiring prior specification of the type or proportion of misspecification. The estimator retains consistency and asymptotic normality under standard regularity conditions, while avoiding iterative optimization and auxiliary regularization. Simulation studies and empirical analysis using the Big Five Personality Inventory demonstrate that the proposed estimator substantially mitigates bias induced by careless responding: estimated correlations differ significantly from maximum-likelihood (ML) estimates, and the method further facilitates detection of aberrant respondents.

Handling careless respondents in polychoric correlation analysisMinimizing divergence between observed and theoretical frequencies robustlyRobust estimation of polychoric correlation under model misspecification

Latest Papers

What's happening recently
View more

This work proposes a unified framework for generating multivariate discrete data with user-specified correlation structures and marginal distributions belonging to the generalized Poisson, negative binomial, or binomial families—three classes that existing methods struggle to handle effectively within a single model. By integrating iterative conditional sampling, correlation correction, probability integral transformation, and discretization strategies, the proposed algorithm accurately matches both target marginal distributions and prescribed correlation matrices. Extensive experiments across four simulation scenarios and three real-world datasets demonstrate that the method efficiently produces synthetic multivariate discrete data conforming to the desired statistical properties. The approach is particularly well-suited for simulation and modeling tasks in fields such as biology, medicine, and social sciences, where realistic discrete multivariate data generation is essential.

binomial distributioncorrelated data generationgeneralized Poisson distribution

This study addresses regression modeling for ordinal outcomes in psychology and social behavioral sciences by systematically comparing the inferential performance of five model classes—proportional odds, partial proportional odds (category-specific), location-shift, location-scale, and misapplied linear models—under diverse data-generating mechanisms within a unified Monte Carlo simulation framework. Emphasizing parameter estimation bias, Type I error control, and statistical power, this work provides the first comprehensive quantification of the robustness and reliability of these approaches. The proportional odds model consistently demonstrates low bias, well-controlled Type I error rates, and high power across most scenarios, maintaining stable performance even under high skewness or large effect sizes, thereby establishing it as the recommended default method for analyzing ordinal outcomes.

inferential propertiesmodel selectionordinal outcomes

This study addresses the widespread yet often inappropriate use of parametric or nonparametric statistical methods in Human-Computer Interaction (HCI) research when analyzing ordinal-scale data, such as Likert-type responses, without adequately considering the underlying assumptions about data structure. The paper systematically critiques the limitations of current analytical practices and, for the first time in the HCI field, advocates for and promotes the adoption of cumulative link models (CLMs) and their mixed-effects extensions (CLMMs) as more principled approaches to modeling ordinal outcomes. By exposing critical flaws in conventional methods’ assumptions and providing reproducible R-based examples alongside open-source datasets, this work establishes a rigorous statistical framework tailored to ordinal data, thereby substantially enhancing analytical validity and advancing methodological standardization in HCI research.

HCILikert scalesmethodological assumptions

Traditional Pearson correlation coefficients cannot capture the heterogeneity in association strength that varies with covariates. This work proposes a likelihood-based regression framework that models the correlation coefficient as a function of covariates, accommodating both bivariate normal and Bernoulli responses. Parameter estimation is carried out via the Newton–Raphson algorithm, and statistical inference is facilitated through bootstrapping. The method is implemented for the first time in regcorr, a lightweight R package requiring no compilation and minimal dependencies, now available on CRAN. The package supports reproducible analysis and includes practical guidance, substantially enhancing the flexibility and feasibility of modeling covariate-dependent correlation structures.

correlation heterogeneitycovariate-dependent correlationPearson correlation coefficient

Measures and Models of Non-Monotonic Dependence

Dec 11, 2025
AJ
Alexander J. McNeil
🏛️ University of York | McGill University | University College Dublin

Traditional Spearman’s ρ fails to capture non-monotonic dependence structures and is sensitive to marginal distributions. Method: We propose a margin-free, unit-square-based generalized Spearman correlation coefficient, constructed within the Hilbert space of square-integrable functions. Marginal invariance is achieved via uniformity-preserving transformations and copula modeling; a novel randomized inverse transformation generates extremal singular copulas, enabling a parametric copula family that continuously interpolates non-monotonic dependence strength and supports symmetry detection. Contributions/Results: Leveraging orthogonal expansions in Legendre polynomials and cosine bases, we derive tight analytical bounds. The sample estimator is shown to be uniformly consistent and asymptotically normal. Empirical evaluations demonstrate superior performance in exploratory dependence analysis, symmetry identification, and non-monotonic density modeling compared to existing measures.

Defines a margin-free measure for non-monotonic bivariate associationInvestigates properties and bounds of generalized Spearman correlationProposes sample analogues and demonstrates applications in dependence modeling

Hot Scholars

XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
ID

Isabella Degen

EPSRC Doctoral Impact Fellow, University of Bristol
AI validationmachine learningunsupervised learningtime series
BP

Barbara Plank

Professor, LMU Munich, Visiting Prof ITU Copenhagen
Natural Language ProcessingComputational LinguisticsMachine LearningTransfer Learning