🤖 AI Summary
This study addresses the inefficiency and limited robustness of multi-population functional generalized linear models arising from their reliance on response distribution assumptions and lack of cross-population information borrowing. To overcome these limitations, this work proposes a semiparametric framework based on density ratio models. Methodologically, it introduces a novel multi-population linking mechanism independent of baseline distributions, integrating maximum empirical likelihood estimation, penalized B-spline approximation, and kernel-smoothed cross-validation to efficiently model both functional and scalar predictors. Simulation studies demonstrate that the proposed method substantially improves estimation efficiency through cross-population information sharing while maintaining strong robustness under distributional misspecification. Its practical utility is further validated through an application to Kansas soybean yield analysis.
📝 Abstract
Functional generalized linear models provide a flexible framework for relating a scalar response to functional and scalar predictors, but their conventional formulation requires specification of the response distribution and is typically developed for a single population. We propose a semiparametric functional generalized linear model for multi-population data that leaves the baseline response distributions unspecified and links them across populations through a density ratio model. The proposed framework accommodates both functional and scalar predictors while borrowing information across related populations. We develop a maximum empirical likelihood estimation procedure and use penalized B-spline approximations to estimate the functional coefficients. A cross-validation procedure based on a kernel-smoothed empirical likelihood density estimator is introduced to select the smoothing parameters. Simulation studies show that borrowing information through the density ratio structure can improve estimation efficiency relative to fitting each population separately, and that the proposed semiparametric method remains competitive when a parametric Gaussian model is correctly specified while providing substantial robustness when it is misspecified. We illustrate the method using county-level soybean yield data from Kansas, with daily temperature curves and irrigation levels as predictors.