Score
Designs and implements algorithms or models that reconstruct missing entries of a partially observed matrix by approximating the full matrix with low-rank factors or low-rank regularizers. Builds imputations and reconstruction procedures that produce accurate estimates from sparse observations and that can be evaluated or optimized under high annotation sparsity.
This paper provides a systematic review of matrix completion, addressing its theoretical foundations, algorithmic frameworks, and empirical evaluation. It focuses on passive versus adaptive sampling paradigms, unifying classical approaches—including singular value thresholding and nuclear norm minimization—with state-of-the-art adaptive strategies, and releases open-source, reproducible implementations. Through controlled synthetic experiments, it quantitatively demonstrates—under a unified benchmark—for the first time that simple adaptive sampling schemes (e.g., uncertainty- or gradient-based) improve reconstruction accuracy by 12.7%–23.4% over random sampling in low-SNR or highly sparse observation regimes, substantially narrowing the gap between theoretical bounds and practical performance. The primary contributions are: (i) establishing a closed-loop “theory–algorithm–evaluation” framework; (ii) rigorously characterizing the practical advantage boundary of adaptive sampling; and (iii) providing both methodological guidance and an empirical benchmark for developing efficient, robust matrix completion methods.
This work addresses the practical efficacy of low-rank matrix completion (LRMC) under data-dependent sampling—common in sensing, recommendation, and sequential decision-making—challenging the standard assumption that sampling is independent of the underlying matrix. Through the first systematic empirical evaluation of mainstream algorithms (e.g., Soft-Impute, AltMin, nuclear norm minimization) on both synthetic and real-world datasets, we demonstrate substantial performance degradation under value-dependent truncated sampling. Crucially, we identify coupling between sampling bias direction and intrinsic matrix structure—particularly singular vector alignment—as the primary cause of failure, and derive an interpretable performance decay law grounded in this coupling. Our study bridges a critical gap between theoretical assumptions and practical deployment, providing an empirical foundation and diagnostic framework for designing robust LRMC methods in realistic, non-i.i.d. sampling scenarios.
This work proposes a distribution-free penalized regression framework to jointly model covariate effects and low-rank latent structure in sparse, noisy observations that exhibit row/column covariates and structural dependencies. The method integrates Lasso, ridge-type kernel regularization, and low-rank decomposition, and is efficiently solved via a scalable alternating least squares algorithm that flexibly accommodates prior similarity information. Theoretically, non-asymptotic error bounds are established, while computationally the approach achieves substantial reductions in complexity. Empirical evaluations on both synthetic and real-world data demonstrate that the proposed method attains prediction accuracy comparable to existing sophisticated approaches at significantly lower computational cost. The accompanying algorithm is publicly available as the R package IMR.
This work addresses the matrix completion problem for highly sparse, ill-conditioned matrices with extreme aspect ratios. We propose Column-Selection Matrix Completion (CSMC), the first framework to jointly model column subset selection (CSS) and low-rank matrix completion within a unified optimization paradigm. CSMC employs a two-stage algorithmic framework—tailored to problems of varying scale—that alternately optimizes column selection and nuclear norm minimization via convex relaxation (SDP/ADMM). We provide rigorous theoretical convergence analysis with probabilistic guarantees. Empirically, CSMC achieves state-of-the-art accuracy on recommendation and image inpainting benchmarks while significantly reducing computational time. Synthetic experiments further demonstrate its robustness and efficiency under high missing rates, large matrix dimensions, and low intrinsic rank—outperforming existing methods in both scalability and reconstruction fidelity.
This work addresses the failure of conventional matrix completion methods under ultra-sparse sampling regimes where the number of observations per row, denoted \( C \), falls below the matrix rank. The authors propose a novel approach that constructs a low-variance estimator of the second-moment matrix via a frequency-normalized unbiased estimator and optimizes a rank-\( r \) factor model using gradient descent to effectively recover the mean second-moment matrix \( T \) and its row space. Theoretically, the method is shown to approximate the global optimum in the neighborhood of a local solution, with sample complexity scaling linearly in the ambient dimension \( d \). Experiments demonstrate an 88% reduction in bias on MovieLens and 59% and 38% lower recovery errors for \( T \) and \( M \), respectively, on Amazon review data at a sparsity level of \( 10^{-7} \), marking the first successful estimation of the row space in the extreme setting where \( C < \text{rank} \).
研究高维数据缺失下的参数估计问题,通过统计与计算复杂性分析,揭示均值和协方差估计存在统计-计算差距,而线性回归则可通过高效算法接近信息论下界。
This work addresses the challenge of feature imputation in real-world machine learning scenarios where missing observations complicate data completion, particularly due to the tension between scalability and structural consistency. The authors propose a novel imputation mechanism based on iterative exploration in the latent space of a generative model, which actively searches for optimal positions within the two-dimensional latent manifold constructed by G-NeuroDAVIS to accurately fill missing values, accompanied by theoretical convergence guarantees. Evaluated across diverse image datasets and missingness patterns, the method consistently outperforms existing approaches, achieving superior performance in terms of RMSE, PSNR, and SSIM metrics. Furthermore, it significantly enhances downstream classification and clustering tasks, with statistical significance confirmed via Wilcoxon signed-rank tests.
This work addresses the challenge of linear regression when both covariates and responses are missing, in the presence of abundant unlabeled data. The authors propose a unified semi-supervised estimation framework that accommodates both structured and unstructured missing patterns. For the non-sparse setting, they employ a weighted imputation strategy, while in the high-dimensional sparse regime, they develop an improved Dantzig selector. Notably, this study establishes, for the first time, matching minimax lower and upper bounds for both supervised and semi-supervised learning under two distinct missingness mechanisms, thereby rigorously quantifying the benefit of unlabeled data in improving convergence rates. The proposed estimators achieve minimax optimal rates, and theoretical findings are corroborated through simulations and semi-synthetic experiments based on California housing data.
This work addresses key limitations in high-dimensional matrix estimation—namely, insufficient integration of row/column side information, restricted nonlinear modeling capabilities, and inadequate noise handling—by proposing a four-component decomposition framework. The target matrix is decoupled into a row-column nonlinear interaction term, row-specific and column-specific main effects, and a low-rank residual structure. Each component is estimated via sieve basis projection and nuclear norm regularization, then aggregated to yield a flexible estimator that accommodates nonlinear interactions, leverages unilateral features, and explicitly models noise. The method is applicable under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms, including block-wise missingness in causal panel settings. Theoretical analysis establishes convergence rates under varying strengths of auxiliary information, while simulations and an empirical study on tobacco sales demonstrate substantial improvements over conventional low-rank and spectral methods in matrix completion and treatment effect estimation.