Score
A statistical classification and modeling technique using a logistic link to estimate relationships between predictors and binary or categorical outcomes, applied to quantify associations and control for covariates in observational data.
Traditional logistic regression models are constrained by fixed asymptotes at 0 and 1, limiting their ability to capture complex relationships between responses and covariates as well as dependencies among binary outcomes. This work proposes a composite logistic regression model that constructs a more flexible mean response structure by combining multiple logistic functions. The approach retains model interpretability while overcoming the restrictive asymptotic boundaries inherent in standard logistic regression. By naturally incorporating covariates and effectively modeling correlated binary responses, the proposed method substantially extends the applicability and expressive capacity of logistic regression in practical settings.
Conventional statistical software frequently encounters infeasibility issues in estimating cumulative link model (CLM) parameters, and standard CLMs lack flexibility in handling missing responses, longitudinal binary outcomes, and non-proportional odds structures. Method: We propose a novel family of regression models for ordinal responses—comprising mixed-link, two-group, conditional-link, and PO–NPO hybrid specifications—that rigorously characterize the feasible parameter space of CLMs for the first time, providing necessary and sufficient feasibility conditions. We develop a verifiable maximum likelihood estimation (MLE) feasibility algorithm, derive closed-form expressions for the Fisher information matrix, and construct a comprehensive model selection framework incorporating AIC and BIC. Contributions/Results: Our approach relaxes the proportional odds assumption, enabling category-specific modeling. Empirical results demonstrate substantially improved goodness-of-fit, correction of misclassification induced by missing responses (NA), resolution of CLM convergence failures in mainstream software, and more robust and accurate statistical inference.
This paper addresses prediction bias arising from inconsistent covariate classification—e.g., race/ethnicity—between the evidence-generation (study) and decision-making (deployment) stages in evidence translation. We formally introduce the “Differentially Classified Covariates” (DCC) framework, the first of its kind. Leveraging causal inference and partial identification theory, we develop a nonparametric framework for deriving prediction bounds on the conditional probability $P(y mid x)$, characterizing its identifiability limits and conditions under which bounds shrink. Our analysis shows that DCC universally widens prediction intervals; effective tightening requires strong assumptions—such as known classification mechanisms or observable proxy variables. Empirically, we demonstrate that racial misclassification in clinical risk prediction can double predictive uncertainty, exposing a previously overlooked risk of evidentiary failure in public policy and healthcare practice.
In spatial logistic regression, incorporating random effects to account for spatial dependence shifts coefficient interpretation from population-averaged to subject-specific, thereby forfeiting marginal interpretability. To address this, we propose a bridge-process-based spatial logistic regression model that embeds spatially structured random effects without compromising marginal interpretability. This bridge process is the first spatial random-effects formulation that simultaneously preserves both marginal and conditional interpretations, and admits a scale-mixture-of-normals representation with favorable theoretical properties. Using Bayesian inference and an efficient MCMC algorithm, our model achieves superior predictive accuracy, computational efficiency, and interpretability in simulation studies and analysis of Gambian childhood malaria data. The framework establishes a new paradigm for modeling spatial binary data—rigorous from a statistical standpoint while retaining practical, policy-relevant interpretability.
Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.
This study addresses the problem of assessing whether observed data are “sufficiently close” to a binary generalized linear model—such as logistic regression—with fully categorical covariates, rather than requiring exact model fit. To this end, the authors propose a formal equivalence testing framework based on minimum distance methodology. The approach leverages both asymptotic theory and bootstrap procedures to compute critical values, thereby filling a critical gap left by conventional goodness-of-fit tests, which are ill-suited for evaluating practical equivalence. Through extensive simulation studies and analyses of two real-world datasets, the proposed method demonstrates strong finite-sample performance and practical utility, offering a robust tool for model adequacy assessment in applied settings.
This study addresses the challenge of conducting valid statistical inference in high-dimensional logistic regression when observations exhibit higher-order tensor dependence structures induced by a Markov random field. Existing methods are limited either to pairwise interactions or to estimation consistency without inferential guarantees. To overcome these limitations, this work proposes a two-stage inference framework: first, a regularized maximum pseudo-likelihood estimator is constructed, and then a bias-corrected estimator is derived to achieve asymptotic normality, thereby enabling confidence interval construction and hypothesis testing. This approach represents the first method capable of delivering valid inference under high-dimensional tensor network dependence. Theoretical analysis establishes both consistency and asymptotic normality of the proposed estimator, and numerical experiments demonstrate its strong finite-sample performance.
This study addresses the challenge of comparing the importance of distinct groups of predictors in classification problems by proposing a statistical inference framework based on Categorical Gini Correlation (CGC). The work establishes, for the first time, a unified theoretical foundation for testing differences in CGC between predictor groups of arbitrary dimensionality, heterogeneity, and potential dependence structures. It rigorously proves the asymptotic normality of the test statistic under the null hypothesis and its consistency under the alternative. Effective inference is achieved through a nonparametric bootstrap procedure. Extensive simulations and real-data analyses—including applications to breast cancer diagnosis and human activity recognition—demonstrate that the proposed framework achieves high statistical power and offers substantial practical utility.
This work proposes a nonparametric Bayesian clustering approach for multivariate categorical data that explicitly incorporates graphical models to account for heterogeneous dependence structures across clusters. Unlike conventional methods that assume conditional independence of variables within each cluster, the proposed framework employs a Dirichlet process mixture of categorical graphical models to partition individuals into groups that are homogeneous not only in marginal distributions but also in their underlying dependency structures and associated parameters. Full Bayesian inference is performed via Markov chain Monte Carlo (MCMC) to enable posterior analysis. To the best of our knowledge, this is the first method to explicitly integrate graphical models into the clustering of categorical data, thereby effectively capturing inter-group differences in dependence patterns. Experiments on simulated data as well as real-world genomic and voting records demonstrate that the approach significantly outperforms existing methods that ignore such structural dependencies.
This study addresses the limitations of existing circular logistic regression methods, which are confined to symmetric link functions and struggle to effectively model the relationship between circular predictors and binary or binomial responses. The authors propose a generalized linear framework that incorporates circular covariates into the linear predictor via sine and cosine transformations and, for the first time, systematically evaluates the performance of both symmetric and asymmetric link functions. Through Monte Carlo simulations based on the von Mises distribution—assessed using AIC, deviance, and empirical analyses of meteorological and seismic data—the study reveals that link function choice significantly impacts model performance when circular data are dispersed and responses are imbalanced. Under high concentration, symmetric links demonstrate greater robustness, whereas asymmetric links tend to be unstable. This work extends the theoretical boundaries of circular–binary response modeling and offers practical guidance for applied researchers.