Hybrid Methods for Robust Tabular Data Imputation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing computational cost and imputation accuracy for missing values in tabular data by proposing NuclearForest and SoftForest, a hybrid imputation framework. The approach integrates nuclear norm minimization with adaptive step-size singular value thresholding (SVT) or SoftImpute for low-rank initialization, followed by a single-pass, non-iterative random forest refinement, thereby effectively overcoming the bottlenecks of conventional iterative optimization. Experimental results demonstrate that the proposed framework achieves imputation accuracy comparable to state-of-the-art methods across diverse missingness mechanisms while accelerating computation by 5.81 to 9.52 times relative to MissForest. Furthermore, it robustly handles mixed-type variables, successfully reconciling computational efficiency with high predictive accuracy.
📝 Abstract
Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstream inference. In this work, we propose two hybrid imputation methods called NuclearForest and SoftForest, which combine nuclear-norm-based low-rank initialization using Singular Value Thresholding (SVT) and SoftImpute, respectively, with a non-iterative Random Forest refinement. For the SVT-based component, we further introduce an adaptive step-size rule, prove adaptive step-size bounds, and establish convergence for the corresponding zero-initialized iteration. The low-rank initialization provides a structured warm start that captures the global covariance patterns in the data, while the subsequent Random Forest step recovers residual nonlinear signals encoding local dependencies. We conduct an extensive benchmark on diverse datasets from different application domains, comparing the proposed methods with seven established imputation methods under the Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms across varying missingness rates. Our results demonstrate that NuclearForest and SoftForest match or exceed the imputation fidelity of state-of-the-art iterative methods such as MissForest, while significantly reducing computational cost. In particular, they achieve speedups of approximately 5.81 times and 9.52 times over MissForest by replacing iterative cycles with a single refinement step. Our approach effectively exploits the low-rank structure of real-world tabular data and accommodates mixed-type variables, providing an efficient and robust solution for data imputation in bioinformatics, economics, and beyond.
Problem

Research questions and friction points this paper is trying to address.

Tabular Data Imputation
Missing Data
Low-rank Structure
Computational Efficiency
Mixed-type Variables
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tabular Data Imputation
Hybrid Methods
Low-rank Initialization
Random Forest
Singular Value Thresholding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinwei Li
Department of Data Science, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany
M
Michelle Bruch
Department of Data Science, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany
Daniel Tenbrinck
Daniel Tenbrinck
Department of Data Science, FAU Erlangen- Nürnberg
Graph methodsMachine LearningImage Processing