decompose streaming tensors

Designs and implements algorithms that incrementally compute and maintain tensor-train (TT) decompositions for streaming multiway arrays, including online alternating-least-squares (ALS) and single-sweep core-update rules. Focuses on deriving and engineering exact online core updates and implementations that run with low memory and low latency, scale linearly with TT rank, and preserve high reconstruction accuracy.

decomposestreamingtensors

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the trade-off between accuracy and memory efficiency in tensor decomposition of streaming high-dimensional data, where existing online methods suffer from insufficient accuracy while batch approaches incur prohibitive memory costs. To overcome this, the authors propose an alternating least squares–based online Tensor Train (TT) decomposition algorithm that incrementally enforces orthogonality constraints in a sequential manner. This enables deterministic updates of core tensors in a single pass, guaranteeing monotonic decrease of the local objective function and temporal smoothness. The algorithm reduces computational complexity from quadratic to linear with respect to the TT rank. Experimental results demonstrate that the proposed method achieves superior reconstruction accuracy and perceptual video quality compared to state-of-the-art online techniques, while offering several orders of magnitude speedup over deep learning–based approaches, making it well-suited for low-latency real-time applications.

computational efficiencyonline algorithmorthogonalization

Fast and Accurate SVD-Type Updating in Streaming Data

Sep 02, 2025
JJ
Johannes J. Brust
🏛️ Arizona State University | Stanford University

To address the high computational cost of SVD approximation updates in streaming data and the scalability limitations of existing incremental/truncated SVD methods at large truncation ranks, this paper proposes a low-rank update framework based on Bidirectional Diagonal Decomposition (BDD). Our method enables efficient rank-$r$ updates via three key innovations: (1) a compact Householder transformation reducing memory usage by 50%; (2) Givens rotations enabling $O(r^2)$-complexity rank-$r$ updates; and (3) a hybrid sparse-plus-low-rank separation strategy for accurate and scalable matrix approximation. Experiments on recommendation systems and network subspace tracking demonstrate that our approach significantly outperforms LAPACK SVD and state-of-the-art incremental SVD methods—achieving superior accuracy, real-time performance even at high truncation ranks, and balanced throughput–precision trade-offs.

Efficient SVD updating for streaming low-rank dataReducing computational cost in high-throughput matrix updatesScalable algorithms for large truncation rank scenarios

Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation

Mar 11, 2024
JL
Jan Laukemann
🏛️ Friedrich-Alexander-Universität Erlangen-Nüernberg | Intel Labs | University of Oregon | Laboratory for Physical Sciences

This work addresses efficient decomposition of high-dimensional sparse tensors—common in healthcare and cybersecurity—on modern parallel processors, overcoming restrictive assumptions about mode structure or sparsity distribution inherent in conventional compressed formats. We propose ALTO, an adaptive linearization tensor representation that is agnostic to both mode structure and sparsity distribution. Built upon ALTO, we design a parallel decomposition algorithm featuring low synchronization overhead and high data reuse, augmented by dynamic performance modeling and scheduling heuristics for automatic hardware adaptation. Leveraging cache- and memory-aware optimizations on Intel Xeon Scalable platforms, experiments demonstrate that ALTO achieves over 10× speedup versus the best structure-agnostic format and a 5.1× geometric mean speedup versus the best structure-aware format, while incurring only 25% of the latter’s storage overhead.

Efficient decomposition of high-dimensional sparse tensorsOvercoming irregular shapes and data distributions in sparse tensorsReducing memory footprint and synchronization overhead in tensor computations

High-order tensor-to-tensor (ToT) regression suffers from the “curse of dimensionality,” leading to explosive storage requirements, prohibitive computational cost, and a gap between theoretical analysis and practical implementation. Method: This paper establishes, for the first time, a statistical theory for ToT regression under tensor train (TT) decomposition. We propose two provably convergent algorithms—iterative hard thresholding (IHT) and Riemannian gradient descent (RGD)—both equipped with theoretical guarantees under a restricted isometry property (RIP) condition. To enhance efficiency, we integrate TT-SVD and spectral initialization strategies. Contributions/Results: We derive tight upper bounds on estimation error and matching minimax lower bounds, revealing polynomial dependence on the total order of input/output tensors. Under RIP, both IHT and RGD achieve linear convergence, and their final estimation accuracy attains the minimax optimal rate. The proposed initialization schemes significantly reduce sample complexity and memory footprint, enabling scalable and efficient ToT regression.

Analyzes error bounds for tensor-on-tensor regression with tensor train decompositionEstablishes linear convergence rates under restricted isometry property conditionsProposes optimization algorithms for efficient tensor regression solutions

High-dimensional grid data represented in the tensor train (TT) format often suffer from rank explosion due to global unfolding, hindering efficient interpolation and compression. This work proposes a low-rank TT local interpolation framework that starts from a coarse-grid TT representation and constructs a fine-grid TT with uniformly bounded tail ranks through multiscale local refinement. The method achieves, for the first time, an ℓ² error bound independent of the total number of cores, exponential compression rates at fixed accuracy, and logarithmic computational complexity with respect to the number of grid points. Its efficacy is demonstrated on 1D/2D/3D tasks—including airfoil mask embedding, image super-resolution, and synthetic turbulent noise—and it enables direct generation of fractal noise fields with logarithmic complexity.

grid refinementhigh-dimensional datalow-rank interpolation

Latest Papers

What's happening recently
View more

This work proposes a novel approach to overcoming the computational complexity bottleneck in matrix multiplication by explicitly exploiting the intrinsic structural properties of tensor decompositions. By designing tensor decompositions with specialized algebraic structures and integrating techniques from algebraic complexity theory with numerical optimization, the study achieves a reduction in the exponent for 6×6 matrix multiplication from 2.8075 to 2.8016, while maintaining a reasonable leading constant. Notably, this result yields an effective exponent below the theoretical lower bound implied by conventional tensor rank considerations and significantly enhances practical algorithmic efficiency. The findings establish a new structured design paradigm for fast matrix multiplication algorithms, offering both theoretical advancement and practical relevance.

algorithmcomputational complexityexponent

Existing methods struggle to effectively address the challenge of adaptive forecasting for streaming matrix-valued time series in time-varying environments. This work proposes a novel adaptive tensor regression framework that extends adaptive filtering to matrix-valued streaming data for the first time, encompassing both Matrix-to-Matrix (MoM) and Tensor-to-Matrix (ToM) modeling paradigms. The framework leverages high-order tensor representations and incorporates low-dimensional structural priors—such as sparsity, low-rankness, and their joint structures—and employs stochastic gradient descent for efficient online learning. Theoretical analysis provides finite-time recovery guarantees, while experiments demonstrate that the ToM model achieves lower steady-state error, stronger denoising capability, and superior tracking performance in dynamic environments compared to MoM.

adaptive forecastingmatrix-valued time seriesstreaming data

This work addresses the curse of dimensionality that severely hampers high-dimensional tensor computations, where conventional methods struggle to balance computational efficiency and memory consumption. Building upon the Tensor Train (TT) or Matrix Product State (MPS) representation, the paper introduces specialized algebraic algorithms for efficiently performing vector and matrix addition, Hadamard products, and MPO–MPS multiplication. The proposed approach achieves a superior trade-off among computational complexity, memory footprint, and numerical accuracy compared to existing techniques, thereby substantially enhancing the practicality of high-dimensional tensor operations. This advancement is particularly beneficial for memory-constrained applications such as system identification and dynamic programming.

algebraic operationsHadamard producthigh-dimensional tensors

This work addresses the exponential storage growth bottleneck encountered when directly extending tensor singular value decomposition (T-SVD) to higher-order tensors, which hinders efficient processing of multi-dimensional data with tubal structures. The authors propose the Tubal Tensor Train (TTT) decomposition, which uniquely integrates t-product algebra with the low-rank structure of tensor trains. By constructing a tensor network using two third-order boundary cores and multiple fourth-order internal cores, TTT achieves linear storage complexity with respect to the number of modes. The method employs a TTT-SVD sequential fixed-rank construction algorithm and an optimization strategy based on Alternating Two-Core Updates (ATCU) in the Fourier domain. Experiments on image/video compression, tensor completion, and hyperspectral imaging demonstrate its efficacy, while theoretical analysis provides error bounds analogous to those of TT-SVD.

storage complexityt-producttensor network

This work addresses the memory-bandwidth-bound nature of the sparse MTTKRP (spMTTKRP) operation in sparse tensor decomposition, which suffers from low efficiency on general-purpose processors. To overcome this limitation, the study introduces processing-in-memory (PIM) technology to accelerate spMTTKRP for the first time, leveraging the UPMEM PIM architecture. The authors devise an efficient sparse tensor tiling strategy, a customized numerical format, and specialized compute kernels tailored to the PIM platform, along with a CPU-PIM heterogeneous collaboration mechanism. Experimental results demonstrate that the pure PIM implementation achieves a 2.37× speedup over the state-of-the-art CPU baseline, while the heterogeneous approach yields a 2.64× improvement, both exhibiting substantially higher resource utilization efficiency compared to conventional CPU and GPU implementations.

memory-boundProcessing-In-Memorysparse tensors

Hot Scholars

YS

Yasushi Sakurai

The Institute of Scientific and Industrial Research,Osaka University
Data MiningTime SeriesDatabases
NG

Noah G. Singer

Ph.D. student, Carnegie Mellon University
Theoretical computer science
BL

Bethany Lusch

Argonne National Lab
machine learningoptimizationscientific computingdata science
LH

Ligong Han

Red Hat AI, MIT-IBM, Rutgers University
generative modelcomputer visiondeep learning