🤖 AI Summary
This study addresses the data separation phenomenon in categorical response models, which frequently renders maximum likelihood estimates nonexistent or non-unique, thereby severely compromising the reliability of statistical inference. To resolve this issue, we develop divoRce, an R package grounded in structural vector theory that integrates computational geometry, linear programming, and rational arithmetic solvers. Supporting multiple link functions and third-party extensions, this toolkit enables existence testing, type classification, and identification of the variables responsible for separation. This work contributes the most comprehensive exact diagnostic suite currently available in the literature, encompassing virtually all categorical models. We recommend incorporating divoRce into standard analytical workflows to ensure modeling robustness and safeguard the validity of downstream inferential procedures.
📝 Abstract
In statistical models for categorical responses, including binary, nominal and ordinal responses, a phenomenon called ``separation'' may occur which can render maximum likelihood estimate nonexistent and/or nonunique. Based on Sablica et al. (2026)'s general characterization of separation phenomena via structure vectors, in this article we discuss computational, numerical and practical aspects of separation phenomena in different software packages, including consequences, remedies and ways of diagnosing them. We further introduce the R package divoRce which provides the most comprehensive software suite to diagnose separation in existence. It allows to check for separation, to distinguish between types of separation, characterize the recession cone and identify variables and observations contributing to separation for almost any categorical response model in the literature, including models with baseline-category link, cumulative link, adjacent-category link and sequential link specifications. Diagnostics are based on methods for computational geometry and linear programming and can be done in floating-point or rational arithmetic with different solvers. The functionality is easily extendible by wrappers and S3 methods to accommodate implementations in third-party packages or of new categorical response models. We discuss our implementation in detail and show our software in action in a wide range of applications. We recommend to incorporate our checks for separation and existence of the MLE as a diagnostic routine in every categorical data analysis both in third-party software implementations and as part of a standard model diagnostic workflow.