🤖 AI Summary
This work addresses the challenge in multi-task linear regression where some tasks may be contaminated and the smallest eigenvalue of the task covariance matrix can approach zero. To tackle this, the authors propose an estimator based on matrix-weighted norm regularization, which introduces a relative balance condition that does not rely on a lower bound for the eigenvalues of individual task covariance matrices. This approach adaptively leverages task similarity, is robust to arbitrary outlier tasks, and avoids negative transfer. Theoretical analysis shows that under moderate balance, the prediction mean squared error achieves the best-known rate, and the overall MSE is minimax optimal up to logarithmic factors. Moreover, when tasks are unrelated or exhibit weak balance, the method performs no worse than learning each task independently.
📝 Abstract
We study the multi-task linear regression problem in the presence of contaminated tasks. We address the setting where the unknown parameters of a majority of tasks are close in the $\ell_2$-norm, while a fraction of tasks are arbitrary outliers. Existing theoretical frameworks for this problem rely heavily on the assumption that the empirical second moment of each task has a minimum eigenvalue bounded away from zero (order $Ω(1)$). Crucially, this assumption fails in many high-dimensional scenarios, rendering prior guarantees vacuous. To overcome this limitation, we propose an estimator based on matrix-weighted norm regularization. We also introduce a relative balancedness condition, quantified by a balancedness constant, that compares each task's second moment with the average inlier geometry and relaxes the need for taskwise second-moment lower bounds. In favorable regimes with moderate balancedness, our prediction MSE bounds match the rate of Duan and Wang (2023) under substantially weaker spectral assumptions; the resulting task-overall MSE is minimax optimal up to logarithmic factors. Furthermore, we demonstrate that our estimator enjoys a safety guarantee: when the relevant balancedness constant is large or infinite, or when tasks are unrelated, the method performs no worse than independent task learning. Consequently, our methodology achieves simultaneous adaptivity to task similarity, robustness to outliers, and safety outside favorable transfer regimes.