🤖 AI Summary
This work addresses the challenge of linear regression when both covariates and responses are missing, in the presence of abundant unlabeled data. The authors propose a unified semi-supervised estimation framework that accommodates both structured and unstructured missing patterns. For the non-sparse setting, they employ a weighted imputation strategy, while in the high-dimensional sparse regime, they develop an improved Dantzig selector. Notably, this study establishes, for the first time, matching minimax lower and upper bounds for both supervised and semi-supervised learning under two distinct missingness mechanisms, thereby rigorously quantifying the benefit of unlabeled data in improving convergence rates. The proposed estimators achieve minimax optimal rates, and theoretical findings are corroborated through simulations and semi-synthetic experiments based on California housing data.
📝 Abstract
Missing values in datasets are common in applied statistics. For regression problems, theoretical work thus far has largely considered the issue of missing covariates as distinct from missing responses. However, in practice, many datasets have both forms of missingness. Motivated by this gap, we study linear regression with a labelled dataset containing missing covariates, potentially alongside an unlabelled dataset. We consider both structured (blockwise-missing) and unstructured missingness patterns, along with sparse and non-sparse regression parameters. For the non-sparse case, we provide an estimator based on imputing the missing data combined with a reweighting step. For the high-dimensional sparse case, we use a modified version of the Dantzig selector. We provide non-asymptotic upper bounds on the risk of both procedures. These are matched by several new minimax lower bounds, demonstrating the rate optimality of our estimators. Notably, even when the linear model is well-specified, our results characterise substantial differences in the minimax rates when unlabelled data is present relative to the fully supervised setting. Particular consequences of our sparse and non-sparse results include the first matching upper and lower bounds on the minimax rate for the supervised setting when either unstructured or structured missingness is present. Our theory is coupled with extensive simulations and a semi-synthetic application to the California housing dataset.