π€ AI Summary
In economics and finance, sparse linear models often exhibit high dimensionality, strong correlations, and latent factor structures; yet target samples are scarce and highly susceptible to model misspecification, rendering conventional estimators unreliable.
Method: We propose a transfer learning framework leveraging heterogeneous auxiliary data from multiple sources. First, we design a data-driven source detection algorithm to automatically identify informative auxiliary datasets and mitigate negative transfer. Second, we develop a hypothesis testing framework for factor model applicability and enable joint confidence interval inference for regression coefficients. Third, via non-asymptotic analysis, we derive ββ/ββ estimation error bounds to ensure robustness.
Contribution/Results: Theoretically, our estimator achieves the optimal convergence rate. Simulation studies and empirical applications demonstrate substantial improvements in estimation accuracy and statistical reliability under cross-dataset heterogeneity.
π Abstract
In this paper, we study transfer learning for high-dimensional factor-augmented sparse linear models, motivated by applications in economics and finance where strongly correlated predictors and latent factor structures pose major challenges for reliable estimation. Our framework simultaneously mitigates the impact of high correlation and removes the additional contributions of latent factors, thereby reducing potential model misspecification in conventional linear modeling. In such settings, the target dataset is often limited, but multiple heterogeneous auxiliary sources may provide additional information. We develop transfer learning procedures that effectively leverage these auxiliary datasets to improve estimation accuracy, and establish non-asymptotic $ell_1$- and $ell_2$-error bounds for the proposed estimators. To prevent negative transfer, we introduce a data-driven source detection algorithm capable of identifying informative auxiliary datasets and prove its consistency. In addition, we provide a hypothesis testing framework for assessing the adequacy of the factor model, together with a procedure for constructing simultaneous confidence intervals for the regression coefficients of interest. Numerical studies demonstrate that our methods achieve substantial gains in estimation accuracy and remain robust under heterogeneity across datasets. Overall, our framework offers a theoretical foundation and a practically scalable solution for incorporating heterogeneous auxiliary information in settings with highly correlated features and latent factor structures.