🤖 AI Summary
This work addresses the challenge of missing data imputation in high-dimensional multiway arrays by proposing a nonparametric, training-free method that preserves interpretability. The approach formulates imputation as a regression problem in a reproducing kernel Hilbert space, incorporating Tensor Train manifold constraints and a Hadamard over-parameterization structure to enable efficient sparse representation. It jointly optimizes tensor coefficients and kernel covariance matrices on a Riemannian product manifold. A novel automatic kernel hyperparameter selection mechanism—eliminating the need for cross-validation—is introduced by integrating fixed-rank Tensor Train decomposition with positive-definite matrix manifold optimization. Experiments on fMRI data imputation and dynamic graph edge stream recovery demonstrate significantly superior accuracy compared to state-of-the-art tensor-based, Bayesian, and neural network methods.
📝 Abstract
Kernel regression with tensor trains and Hadamard overparameterization (KReTTaH) is introduced as a training-data-free, interpretable, and nonparametric framework for multi-way data imputation. The imputation problem is reformulated as regression in reproducing kernel Hilbert spaces (RKHS), where the tensor regression coefficients are explicitly constrained to lie on fixed-rank tensor-train (TT) manifolds and structured via Hadamard overparameterization to promote sparsity and high representational efficiency. Rather than relying on costly cross-validation, KReTTaH jointly optimizes the TT coefficient tensors and the kernel covariance matrices within a Riemannian product-manifold framework -- the former on fixed-rank TT manifolds, the latter on the manifold of positive-definite matrices -- thereby enabling automated kernel-hyperparameter selection. Numerical tests on two challenging applications -- imputation of high-dimensional functional magnetic resonance imaging (fMRI) data and recovery of missing edge flows in dynamic graphs -- demonstrate that KReTTaH consistently outperforms state-of-the-art tensor-, Bayesian-, and neural-network-based baselines in terms of modeling accuracy.