🤖 AI Summary
This work addresses the challenge that existing regression methods struggle to handle arbitrary ordinal data—whether continuous or discrete—without imposing restrictive assumptions such as predefined transformation forms or distributional specifications, which limits their ability to model mean-variance relationships flexibly. To overcome this, the authors propose a monotonic transformation linear regression framework based on an extended rank likelihood, treating the mean-variance relationship as a nuisance parameter and thereby avoiding explicit specification of data type or transformation. Parameter estimation is carried out via Bayesian inference with Gibbs sampling, and conformal calibration is integrated to construct predictive intervals. Theoretical analysis and experiments demonstrate that the approach incurs no asymptotic information loss in both continuous and binary extremes and maintains valid marginal frequentist coverage even under model misspecification, achieving both estimation accuracy and prediction reliability.
📝 Abstract
The accuracy of inference from a regression model depends largely on how well the model represents the relationship between the mean and variance of the outcomes. As this relationship is rarely of direct interest, it is natural to treat it as a nuisance parameter, rather than attempt to estimate it. We take this approach in the context of a monotonically transformed linear regression model using a pseudo-likelihood based on an extended notion of ranks. This approach can accommodate a wide range of mean-variance relationships and any ordinal data type, including continuous and discrete ordered data, and requires no estimation or prior specification of the transformation, or decision to treat an outcome as continuous or discrete. We show that the extended rank likelihood incurs no asymptotic information loss at the two extremes of continuous and binary data, and that rank-based prediction intervals can obtain approximate coverage control conditional on the features. Bayesian parameter estimates and prediction intervals are available via a simple Gibbs sampling algorithm. For settings where the model is in doubt, conformal calibration of the Bayesian predictive distribution provides intervals with guaranteed marginal frequentist coverage.