🤖 AI Summary
This study addresses the challenges of transfer bias and efficiency trade-offs in high-dimensional multi-output regression arising from target-domain data scarcity and source-domain heterogeneity. To this end, we propose a joint sparse transfer framework that integrates cross-response shared structures with source-target similarity, yielding two complementary estimators for fusion and debiasing. Technically, the approach combines joint sparsity modeling, minimax lower bound analysis, and convex set projection. Theoretically, our analysis reveals the fundamental trade-off between information gain and selection cost, establishing error bounds that are tight up to logarithmic factors. Empirical evaluations on single-cell RNA sequencing data validate the effectiveness of the proposed method.
📝 Abstract
Multitask linear models can improve estimation and prediction by exploiting structure shared across responses, bridging taskwise fitting and complete pooling. In many applications, however, the objective is estimation in a data-limited target domain, while data-rich but heterogeneous source domains are available. Borrowing from these sources can improve efficiency but introduce bias. We develop a joint-sparse transfer-learning framework for high-dimensional multi-output regression that combines shared predictor structure across responses with source-target similarity. The framework yields two complementary estimators: a fused estimator that aggregates jointly fitted domain-specific coefficients and a target-based debiased estimator that adjusts for source-induced shifts. Our error bounds show how transfer increases the available information and sharing predictors across responses reduces selection costs. They also reveal a tradeoff: the fused estimator benefits from larger sources but may retain bias if source shifts point in similar directions, whereas debiasing trades this bias for additional estimation error governed by the smaller target sample. Comparison with a minimax lower bound identifies regimes where the bounds match up to logarithmic factors, where matching remains unresolved, and where projection onto a target-based convex set closes the gap. Simulations and an analysis of single-cell RNA and surface-protein profiles across cell types support the theory.