Refactor Analysis: Predictive Evaluations of Factor Models and Dimensionality

📅 2026-03-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional unidimensional factor models, which rely on correlation matrices and fail to capture their actual explanatory and predictive power for the original response matrix. The authors propose Refactor analysis, which translates a unidimensional solution into a rank-1 approximation of the raw response matrix via dual association plots, and integrates Verifactor analysis with row–column double cross-validation to enhance generalization. For the first time, unidimensionality assessment is shifted from fitting correlation matrices to directly predicting observed data, exposing a disconnect between conventional fit indices and data recoverability. The quadrant correlation coefficient (q′) is reintroduced as a robust alternative. Experiments on 200 public binary datasets demonstrate that q′ substantially outperforms standard correlation-based methods in reconstruction accuracy and sample stability, whereas traditional fit indices—despite high intercorrelations—prove poor predictors of actual recovery performance.

Technology Category

Cognitive Modeling & Cognitive Systems: AnalogyMachine Learning: Calibration & Uncertainty QuantificationKnowledge Representation and Reasoning: Qualitative Reasoning

Application Category

User Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Unidimensional factor models justify some of the most consequential summaries in science -- single scores, single ranks, and single leaderboards -- yet unidimensionality is usually assessed indirectly by fitting and evaluating models on images of the data (e.g., correlation matrices) rather than on the response matrix itself. We introduce Refactor analysis, a data-first evaluation paradigm that converts a one-factor solution into a rank-1 prediction of the original matrix by estimating both respondent- and item-side structure from dual association images. We further introduce Verifactor analysis, which evaluates the same construction under bi-cross-validated (BCV) row-column partitions for improved generalization. In simulations where the data-generating mechanism is truly rank-1 and correlational, Refactor metrics align with classical unidimensionality indices, validating the approach. However, across 200 public dichotomous datasets, traditional fit and unidimensionality measures, though highly intercorrelated, are weakly related to data recoverability, especially out of sample. This gap exposes a methodological vulnerability: excellent image-based fit can coexist with poor data-level explanatory power. Finally, treating the association measure itself as a testable hypothesis, we compare $φ$, tetrachoric, and quadrant correlation, $q^\prime$, an important reintroduction. Quadrant correlation emerges as a simple, interpretable, and remarkably robust alternative, yielding consistently stronger reconstruction and more stable behavior under sample-size variation than commonly used correlations. Together, Refactor and Verifactor shift unidimensionality assessment from "does a one-factor model fit the correlation matrix?" to the question that matters for measurement and benchmarking: does a one-factor dependence structure recover and generalize the observed responses?
Problem

Research questions and friction points this paper is trying to address.

unidimensionality
factor models
data recoverability
correlation matrices
response matrix
Innovation

Methods, ideas, or system contributions that make the work stand out.

Refactor analysis
Verifactor analysis
unidimensionality
quadrant correlation
bi-cross-validation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michael Hardy
Stanford University, Stanford, CA