Score
Statistical modeling using linear predictors linked to response distributions in the exponential family via a link function, applied to tasks like penalized estimation of R_t, quantifying environmental covariate effects, and joint models combining different prediction objectives.
This paper addresses regression modeling with spherical responses and mixed covariates—both linear and spherical. We propose an anisotropic link function based on an extended Möbius transformation, which enables flexible, interpretable directional scaling and generalizes conventional spherical regression links. Crucially, the link ensures orthogonality—in the Fisher information sense—between regression parameters and shape parameters of the error distribution. Coupled with elliptically symmetric error distributions (e.g., Kent and scaled von Mises–Fisher distributions), we employ parallel transport to identify the symmetry axis and reparameterize the model for numerically stable maximum likelihood estimation. The methodology is validated on real-world data, and an accompanying R package facilitates practical implementation. Our approach substantially enhances the flexibility, interpretability, and statistical robustness of spherical regression models.
Multi-source record linkage often introduces non-ignorable errors—such as missed links or false matches—that induce bias in secondary analyses. Conventional methods rely on the strong non-informative linkage assumption, i.e., that linkage errors are independent of analysis covariates—a condition frequently violated in practice. This paper proposes a two-component mixture model framework that relaxes this assumption by allowing the linkage mechanism to depend on observable covariates. The method enables error correction in secondary analyses where only linked data are available and no details about the linkage process are known. By jointly modeling linkage uncertainty and regression analysis, we derive consistent estimators and provide a practical EM algorithm for implementation. Simulation studies and empirical applications demonstrate that the proposed approach substantially improves parameter estimation accuracy and inferential robustness under non-ignorable linkage errors, thereby broadening the scope of valid analyses for linked data.
Conventional statistical software frequently encounters infeasibility issues in estimating cumulative link model (CLM) parameters, and standard CLMs lack flexibility in handling missing responses, longitudinal binary outcomes, and non-proportional odds structures. Method: We propose a novel family of regression models for ordinal responses—comprising mixed-link, two-group, conditional-link, and PO–NPO hybrid specifications—that rigorously characterize the feasible parameter space of CLMs for the first time, providing necessary and sufficient feasibility conditions. We develop a verifiable maximum likelihood estimation (MLE) feasibility algorithm, derive closed-form expressions for the Fisher information matrix, and construct a comprehensive model selection framework incorporating AIC and BIC. Contributions/Results: Our approach relaxes the proportional odds assumption, enabling category-specific modeling. Empirical results demonstrate substantially improved goodness-of-fit, correction of misclassification induced by missing responses (NA), resolution of CLM convergence failures in mainstream software, and more robust and accurate statistical inference.
Causal structure identification in heterogeneous environments typically requires multiple environments and often fails to yield unique solutions under single-environment settings. Method: We propose a causal parent identification criterion grounded in Pearson risk invariance and conditional likelihood maximization, enabling unique causal structure recovery from a single observational environment—without strong distributional assumptions on non-target variables. Our approach applies to Poisson, logistic, and nonlinear additive models within the generalized linear model framework. Contribution/Results: We establish theoretical guarantees for unique identifiability under mild regularity conditions. Algorithmically, we employ a stepwise greedy search, implemented in the open-source R package *causalreg*. Extensive simulations and real-data analyses demonstrate that our method significantly outperforms multi-environment baselines in single-environment scenarios—particularly in low-sample-size and low-environment-diversity settings—thereby establishing a novel paradigm for causal inference under data-scarce conditions.
This paper addresses partial identification of target coefficients in linear regression when the outcome variable and a subset of covariates reside in two separately collected, non-linkable datasets—without imposing exclusion restrictions. To overcome the limitation of conventional methods that rely on strong exogeneity assumptions, we first constructively characterize the sharp identification set under no exclusion constraints. We then propose a computationally efficient estimator for its bounds, based on moment inequalities and convex optimization. The estimator is analytically tractable, asymptotically normal, and exhibits robust finite-sample performance. Theoretically and empirically, our approach substantially extends the scope of prediction and causal inference in settings with missing individual-level linkage across data sources, offering a novel paradigm for modeling heterogeneous, multi-source data.
This study addresses regression problems involving functional predictors and multivariate responses by proposing a novel coefficient function decomposition method that explicitly leverages the interdependencies among response variables. By integrating the functional predictor structure with response correlations, the approach introduces a joint smoothing-and-sparsity penalty strategy that enhances both curve selection and estimation accuracy across settings ranging from small to large-scale scenarios—with up to thousands of functional predictors. Theoretical analysis and extensive numerical experiments demonstrate that the proposed method substantially outperforms existing alternatives. An efficient implementation is provided in the R package FRegSigCom, enabling scalable and high-dimensional functional regression modeling.
This work proposes a latent-process-driven generalized linear model based on the two-parameter exponential family, overcoming the limitations of existing approaches that rely on the single-parameter exponential family assumption and thus struggle to jointly model count, binary, real-valued, and positive continuous time series. By introducing multiplicative latent process effects, the framework accommodates distributions such as Gamma, relaxes constraints on link functions, and establishes asymptotic normality for likelihood-based estimators under marginalization of the latent process. Key innovations include the first prediction and forecasting methodology for this class of models, method-of-moments estimation for unknown discrete parameters, and a corrected information matrix. Empirical validation on German measles incidence and paleoclimatic varve data demonstrates the model’s flexibility and practical utility in estimation, inference, and prediction.
This study addresses the limitations of traditional regression models, which employ linear predictor structures and struggle to effectively model circular response variables exhibiting periodicity. The authors propose a novel Bayesian regression framework specifically designed for concentrated circular responses, introducing—for the first time—a probabilistic model that inherently preserves circular characteristics. This framework naturally extends to joint modeling of multiple circular responses and accommodates both linear and circular covariates alongside various random effects. Efficient Bayesian inference is achieved through integrated nested Laplace approximation (INLA), substantially overcoming the modeling constraints imposed by conventional approaches to circular data. Extensive simulations and real-data analyses demonstrate the method’s superior accuracy and practical utility.
This study addresses the failure of generalized additive models (GAMs) in extrapolation under covariate distributional shifts or extreme values by proposing a novel framework that integrates multivariate extreme value theory (EVT). The approach employs GAMs in the central region of the covariate space while introducing an EVT-driven asymptotic model in extreme regions, unified through a link function that establishes a coherent linear structure on latent variables or continuous responses. This work presents the first systematic embedding of extreme value theory into the GAM regression framework, substantially enhancing robustness and predictive accuracy under extreme covariate extrapolation. Empirical validation on European wildfire prediction—using environmental and meteorological covariates—demonstrates the method’s superior performance under extreme climatic conditions.
This study addresses the lack of a unified and efficient estimation framework for complex causal inference problems involving time-varying treatments, mediation effects, and censored data. Building on the Riesz representation theorem, the authors propose a general recursive Riesz representer framework that integrates targeted minimum loss-based estimation (TMLE) with semiparametric efficiency theory to construct efficient estimators for nested linear functionals in a unified manner. The approach substantially simplifies the construction of estimators for a broad class of causal parameters while guaranteeing asymptotic efficiency and robustness. Numerical experiments demonstrate its favorable performance, and an open-source software implementation is provided. The method is successfully applied to a reanalysis of data from an HIV vaccine efficacy trial.