🤖 AI Summary
This work addresses the computational intractability of traditional Bayesian density regression within the logistic Gaussian process framework, which requires evaluating expensive normalization constants that hinder scalability. To overcome this limitation, the authors propose a generalized Bayesian approach based on the Hyvärinen score, employing a loss function that depends solely on the derivatives of the density with respect to the response variable, thereby entirely circumventing the need for normalization constants. By integrating sparse inducing points with variational inference, the method achieves efficient and scalable nonparametric density regression. Notably, this is the first application of the Hyvärinen score to logistic Gaussian processes, offering a principled balance between model interpretability and computational efficiency. The approach demonstrates strong empirical performance and scalability on both synthetic data and large-scale real-world datasets, including German weather records comprising over 150,000 observations.
📝 Abstract
Density regression extends conventional parametric regression by allowing the entire distribution of the response to vary flexibly with covariates rather than just low-order moments. In the Bayesian setting, logistic Gaussian process (GP) priors have been widely used for density estimation and extend naturally to density regression. The prior can be centred on a base density model, with the nonparametric component providing an interpretable correction that is useful for model criticism. However, logistic GP density regression models have seen limited use, since they require computation of a normalizing constant for every observation, typically via numerical integration. We address this difficulty by proposing a generalized Bayesian approach using a loss function based on the Hyvarinen score. The Hyvarinen score depends only on derivatives of the log density with respect to the response, eliminating the need to compute normalizing constants. Since GP computations remain expensive, we also employ sparse inducing point approximations and variational inference to develop a scalable approach. We demonstrate the method on one simulated and two real datasets, including a German weather dataset with more than 150,000 observations.