Semi-supervised linear regression with missing covariates

📅 2026-02-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of linear regression when both covariates and responses are missing, in the presence of abundant unlabeled data. The authors propose a unified semi-supervised estimation framework that accommodates both structured and unstructured missing patterns. For the non-sparse setting, they employ a weighted imputation strategy, while in the high-dimensional sparse regime, they develop an improved Dantzig selector. Notably, this study establishes, for the first time, matching minimax lower and upper bounds for both supervised and semi-supervised learning under two distinct missingness mechanisms, thereby rigorously quantifying the benefit of unlabeled data in improving convergence rates. The proposed estimators achieve minimax optimal rates, and theoretical findings are corroborated through simulations and semi-synthetic experiments based on California housing data.

Technology Category

Machine Learning: Semi-Supervised LearningReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Non-convex Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Missing values in datasets are common in applied statistics. For regression problems, theoretical work thus far has largely considered the issue of missing covariates as distinct from missing responses. However, in practice, many datasets have both forms of missingness. Motivated by this gap, we study linear regression with a labelled dataset containing missing covariates, potentially alongside an unlabelled dataset. We consider both structured (blockwise-missing) and unstructured missingness patterns, along with sparse and non-sparse regression parameters. For the non-sparse case, we provide an estimator based on imputing the missing data combined with a reweighting step. For the high-dimensional sparse case, we use a modified version of the Dantzig selector. We provide non-asymptotic upper bounds on the risk of both procedures. These are matched by several new minimax lower bounds, demonstrating the rate optimality of our estimators. Notably, even when the linear model is well-specified, our results characterise substantial differences in the minimax rates when unlabelled data is present relative to the fully supervised setting. Particular consequences of our sparse and non-sparse results include the first matching upper and lower bounds on the minimax rate for the supervised setting when either unstructured or structured missingness is present. Our theory is coupled with extensive simulations and a semi-synthetic application to the California housing dataset.
Problem

Research questions and friction points this paper is trying to address.

semi-supervised learning
missing covariates
linear regression
minimax rate
high-dimensional statistics
Innovation

Methods, ideas, or system contributions that make the work stand out.

semi-supervised learning
missing covariates
minimax optimality
Dantzig selector
non-asymptotic bounds
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Benedict M. Risebrow
Department of Statistics, University of Warwick
T
Thomas B. Berrett
Department of Statistics, University of Warwick