Score
Designs, estimates, and evaluates statistical and machine‑learning regression models that quantify relationships between predictors and outcomes (continuous, binary, counts, etc.), including parametric forms (linear, generalized linear models, polynomial, log‑log, power‑law), semi- and nonparametric forms (generalized additive models, local regression, splines, Gaussian process regression, nonparametric regression), and regularized approaches (penalized regression). Tasks include specifying functional forms, estimating covariate‑dependent weights and propensity scores, decomposing main and interaction effects, controlling observed covariates, testing factor significance, estimating cross‑sectional effect sizes, and fitting models across samples.
Students in intelligent computing course sequences (AI, data mining, machine learning, pattern recognition) exhibit heterogeneous mathematical backgrounds, hindering unified instruction in regression analysis. Method: This project develops a self-contained, dependency-free regression pedagogy grounded solely in undergraduate-level calculus, linear algebra, and probability theory. It systematically integrates classical statistical modeling (e.g., least squares, ridge regression, LASSO, kernel methods) with modern machine learning paradigms (gradient-based optimization, neural networks), unifying conceptual treatment of model formulation, loss function design, parameter estimation, and regularization principles. Instruction leverages reproducible code, intuitive visualizations, and real-world case studies to lower cognitive barriers. Contribution/Results: Empirical validation confirms that students can achieve seamless progression from linear models to deep regression without external references. The framework establishes a rigorous methodological foundation for advanced AI coursework.
This work proposes a unified inference framework based on parametric programming to address the bias in statistical inference for regression coefficients following Lasso variable selection in generalized linear models. By locally linearizing the maximum likelihood estimator, the method constructs a linear relationship between pseudo-responses and covariates, thereby extending parametric programming—previously limited to Gaussian settings—to non-Gaussian response distributions. This extension enables valid post-selection inference across a range of exponential family models, including logistic and Poisson regression. The approach maintains computational efficiency while substantially improving inferential accuracy. Simulation studies demonstrate that, in non-Gaussian settings, the proposed method effectively corrects the naive inference that ignores the selection process and achieves higher statistical efficiency compared to existing approaches such as the polyhedral method.
Modeling nonlinear covariate effects and inter-response dependence simultaneously in multivariate survival data remains challenging. Method: We propose a flexible semiparametric single-index Cox-type model that employs a piecewise-constant baseline hazard, B-spline functions to capture nonlinear covariate effects, and a copula function to explicitly characterize the dependence structure among multivariate survival outcomes; estimation is conducted within a full-likelihood framework unifying parametric and nonparametric components. Contribution/Results: This work is the first to integrate the single-index structure, spline-based smoothing, and copula-based dependence modeling into a unified framework, enabling subject-specific survival and hazard function prediction. Simulation studies and application to the Busselton Health Study demonstrate substantial improvements in accuracy of nonlinear effect estimation and joint model fit, while maintaining statistical efficiency and interpretability.
To address statistical efficiency loss and standard error bias arising from the conventional two-stage approach—first estimating marginal distributions nonparametrically/semiparametrically, then fitting a Gaussian copula—in modeling multivariate non-normal data, this paper proposes an integrated likelihood framework that jointly estimates marginal distributions and Gaussian copula parameters. Key contributions include: (i) the first formal definition of four classes of nonparametric normal log-likelihood functions; (ii) identification and exploitation of the biconvex structure of the objective function, enabling a convex approximation optimization strategy; and (iii) derivation of exact score functions via the Genz algorithm, facilitating first-order optimization. The method substantially enhances robustness of transformation-based discriminant analysis for limit-of-detection biomarker data and improves asymptotic efficiency and standard error accuracy in semiparametric polychoric correlation estimation.
This study addresses the limitations of traditional multivariate response regression methods, which often neglect inter-variable dependencies and struggle to simultaneously achieve smooth fitting and dimensionality reduction. The authors propose a novel unified framework that, for the first time, integrates P-spline smoothing into reduced-rank regression by employing B-spline basis expansions with penalized coefficients. This approach jointly models the correlations among multiple responses and their nonlinear trends. An efficient block-relaxation algorithm is developed for parameter estimation, while biplots and partial dependence plots are incorporated to enhance model interpretability. Extensive simulations and analyses of three real-world datasets demonstrate that the proposed method substantially improves both fitting smoothness and the ability to elucidate the underlying multivariate response structure.
This study addresses the challenge of jointly modeling the probability of zero outcomes and the heterogeneous, nonlinear effects in the positive component of semi-continuous data—characterized by a substantial mass at zero and a continuous positive part—along with their complex dependence structure. To this end, we propose the first copula-based semiparametric two-part quantile regression framework. The approach separately models the occurrence of zeros and the magnitude of positive values via quantile regression and flexibly captures their nonlinear dependence across quantiles using a copula function. Theoretical analysis establishes large-sample asymptotic properties, and simulations demonstrate superior performance over existing methods under high zero-inflation and nonlinear scenarios. An empirical application to healthcare data reveals heterogeneous and nonlinear effects of social deprivation on uncompensated and charitable care burdens.
This study addresses the bias in nonparametric regression when covariates are estimated in a first stage, as commonly arises in analyses of heterogeneous treatment effects—such as returns to education—where ignoring estimation error leads to inconsistent inference. The authors propose a general debiasing framework that is agnostic to the first-stage estimation method, offering two approaches: one based on influence functions leveraging pathwise differentiability, and another directly correcting plug-in bias. The second stage can be implemented via local smoothing or sieve methods. Under suitable regularity conditions, the proposed estimators achieve the “oracle” convergence rate, substantially outperforming conventional plug-in estimators. Simulations confirm favorable finite-sample performance, and an application to NLSY97 data reveals that college attainment has the largest effect in reducing unemployment among individuals least likely to complete college, corroborating existing findings on causal heterogeneity.
This study addresses the challenge of jointly modeling population-level trajectories and individual heterogeneity in nonlinear mixed-effects models by proposing an efficient estimation approach that integrates penalized splines with subject-specific transformations. The population trajectory is represented via penalized splines, while Laplace approximation is employed to handle integrals involving random effects. Leveraging automatic differentiation within the Template Model Builder (TMB) framework, the method enables joint optimization of smoothing parameters and variance components, achieving substantial computational gains without compromising statistical accuracy. Simulation studies demonstrate superior performance over existing methods, and the approach is successfully applied to longitudinal data on height growth in infants during the first two years after birth.
This study addresses the challenge of variable selection in linear regression by proposing a data-driven approach based on artificial neural networks (ANNs). The method leverages statistical quantities derived from ordinary least squares (OLS) estimation to automatically assess variable significance, marking the first end-to-end application of ANNs within the OLS framework for variable selection. This yields a scalable and intelligent modeling paradigm. Empirical evaluations demonstrate that the proposed approach consistently achieves superior variable selection accuracy across diverse sample sizes and error variance configurations, outperforming conventional techniques such as forward/backward selection, AIC, BIC, and LASSO. Its practical utility is further validated on the WHO life expectancy dataset. The authors publicly release a pretrained ANN model capable of handling up to 100 predictors.
This study addresses the challenges of tuning and evaluation ambiguity in semiparametric and high-dimensional models arising from redundant components introduced by kernel smoothing or basis expansions. The authors propose an induced replication framework that leverages the principle of parametric inference, transforming model assessment into an in-sample prediction error problem by exploiting known replication mechanisms embedded within the model. Building upon Fisher’s concepts of sufficiency and conditional sufficiency, this approach replaces conventional out-of-sample prediction and is applicable to proportional hazards models, time-varying Poisson processes, and confidence set construction for sparse regression. Both theoretical analysis and numerical experiments demonstrate that the method accurately controls nominal error rates under correct model specification and exhibits high sensitivity to semiparametric misspecification.