Score
Designs, implements, and analyzes supervised regression models that predict multivariate outputs by constraining the coefficient or latent-factor representation to be low-rank, including forms that enforce orthogonal latent factors and permit adaptive selection of rank. This work encompasses building estimators and algorithms for multi-output mapping, deriving identifiability and recovery guarantees, and producing interpretable, structured coefficient or factor representations for downstream analysis.
This work proposes a novel approach that integrates principal component–guided sparse regularization (pcLasso) into the reduced-rank regression framework, addressing a key limitation of existing methods which struggle to simultaneously exploit the principal component structure and group structure of predictors while effectively biasing regression coefficients toward high-variance principal component directions. By explicitly incorporating predictor principal component orientations, group information, and inter-response correlations, the proposed method overcomes constraints inherent in conventional models, enhancing both predictive accuracy and model interpretability. Extensive numerical simulations and real-data analyses demonstrate the substantial advantages of this approach in terms of prediction performance and explanatory power.
Multivariate calibration in multi-output probabilistic regression suffers from ambiguous definitions and practical implementation challenges. Method: This paper proposes a generic pre-ranking regularization framework that (i) employs pre-ranking functions for active calibration—not merely diagnostic assessment—and introduces a joint regularized loss integrating marginal and multivariate calibration; (ii) devises a PCA-based pre-ranking method to automatically identify calibration-sensitive principal directions in predictive distributions; and (iii) incorporates a PIT uniformity deviation penalty, compatible with highest-density-region calibration and copula calibration. The framework is plug-and-play, seamlessly integrated into any probabilistic model’s training objective. Contribution/Results: Evaluated on 18 real-world multi-output datasets, the framework significantly improves multivariate calibration across diverse pre-ranking functions without compromising predictive accuracy.
This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.
This paper addresses partial identification of target coefficients in linear regression when the outcome variable and a subset of covariates reside in two separately collected, non-linkable datasets—without imposing exclusion restrictions. To overcome the limitation of conventional methods that rely on strong exogeneity assumptions, we first constructively characterize the sharp identification set under no exclusion constraints. We then propose a computationally efficient estimator for its bounds, based on moment inequalities and convex optimization. The estimator is analytically tractable, asymptotically normal, and exhibits robust finite-sample performance. Theoretically and empirically, our approach substantially extends the scope of prediction and causal inference in settings with missing individual-level linkage across data sources, offering a novel paradigm for modeling heterogeneous, multi-source data.
Existing conditional latent factor models lack a unified, robust estimation framework under high-dimensional settings. Method: This paper proposes a general estimation framework based on nuclear norm regularization, enabling joint convex optimization for multiple model classes—including conditional principal component analysis and factor-augmented regression. Theoretical analysis establishes statistical consistency and convergence rates under high-dimensional asymptotics. Contribution/Results: We introduce a novel homogeneity constraint that substantially improves out-of-sample prediction accuracy; develop a scalable algorithm integrated with a data-driven cross-validation procedure for hyperparameter selection. Empirically, the method is applied to predicting U.S. stock cross-sectional returns, achieving significantly higher forecasting precision than state-of-the-art alternatives. Additionally, we derive several new asymptotic inference results, including valid confidence intervals for estimated factors and loadings under high-dimensional dependence.
This study addresses the lack of a unified formulation for scalar, multivariate, and functional regression models, which obscures their intrinsic connections. By leveraging an integral operator defined with respect to general measures, the authors propose a unified framework that subsumes all three regression types as special cases of the same operator under different input and output measures. This framework reveals classical regression forms as measure-dependent manifestations of a single operator, clarifies discretized modeling as operator estimation under specific measures, and explains the efficacy of vectorized multivariate regression in linear settings. Theoretically, the authors prove that discrete representations correspond exactly to operator evaluations under discrete measures and converge to the continuous case as the discretization grid refines; moreover, this estimator is equivalent to standard multivariate regression and inherits its classical statistical properties.
Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.
This study addresses the non-identifiability inherent in multi-stream distribution models, where multiple streams act on a common distribution, leading to flat likelihood surfaces and unstable estimation. To resolve this, the authors introduce structured regularization into the framework for the first time, proposing a penalized maximum likelihood estimator with stream-specific heterogeneous penalties that integrate non-smooth regularizers such as adaptive Lasso and elastic net. An efficient proximal gradient algorithm is developed for optimization. Theoretical analysis demonstrates that the proposed method restores model identifiability, ensures strict convexity of the objective function for a unique solution, and establishes oracle properties for the adaptive Lasso component. Simulations and an empirical analysis of NHANES data on asthma and lead exposure show that the method substantially reduces estimation error, controls false discovery rates, and yields stable, interpretable estimates of protective effects.
This study addresses the limitations of traditional multivariate response regression methods, which often neglect inter-variable dependencies and struggle to simultaneously achieve smooth fitting and dimensionality reduction. The authors propose a novel unified framework that, for the first time, integrates P-spline smoothing into reduced-rank regression by employing B-spline basis expansions with penalized coefficients. This approach jointly models the correlations among multiple responses and their nonlinear trends. An efficient block-relaxation algorithm is developed for parameter estimation, while biplots and partial dependence plots are incorporated to enhance model interpretability. Extensive simulations and analyses of three real-world datasets demonstrate that the proposed method substantially improves both fitting smoothness and the ability to elucidate the underlying multivariate response structure.