batched dst poisson solving

Designs and implements solvers that compute solutions of Poisson-like linear systems by applying discrete sine transforms to many independent right‑hand sides in batched form, including construction of batched DST kernels and a batched DST-based linear solver. Builds GPU‑accelerated implementations compatible with ML array frameworks that avoid expensive global array transposes in higher dimensions, and integrates or analyzes these batched DST solvers with low‑rank or cross‑approximation techniques while assessing performance and numerical stability.

batcheddstpoissonsolving

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of efficient solutions for computing a large batch of small-scale singular value decompositions (SVDs) on GPUs. The authors propose a GPU-accelerated batched SVD solver based on the one-sided Jacobi algorithm, co-designed with hardware architecture to exploit fine-grained parallelism, optimize memory access patterns, and support multiple floating-point precisions. Implemented on both NVIDIA and AMD GPU platforms, the solver demonstrates exceptional robustness and scalability across diverse matrix shapes, conditioning numbers, and precision configurations. Experimental results show that the proposed method significantly outperforms existing vendor-provided libraries and open-source solvers in terms of computational performance while maintaining numerical reliability.

batch SVDGPU computinghigh-performance computing

This work addresses the computational bottleneck in solving high-dimensional, large-scale low-rank Poisson equations by proposing an efficient low-rank solver implemented in PyTorch. The method integrates the Cross-DEIM framework with statistical leverage score–driven adaptive index selection and discrete sine transforms (DST), marking the first incorporation of a machine learning framework into low-rank PDE solvers. GPU-accelerated parallelism is achieved through batched FFTs without global transposition, enabling significant performance gains. Evaluated on A100 GPUs and AMD EPYC CPUs, the solver demonstrates substantial speedups and successfully tackles problem scales previously deemed infeasible, offering both algorithmic novelty and practical engineering value.

Cross ApproximationGPU AccelerationLow-Rank Solver

Harnessing Batched BLAS/LAPACK Kernels on GPUs for Parallel Solutions of Block Tridiagonal Systems

Sep 03, 2025
DJ
David Jin
🏛️ Massachusetts Institute of Technology | Argonne National Laboratory

This work addresses symmetric positive-definite block-tridiagonal linear systems arising in time-dependent estimation and optimal control. We propose a GPU-accelerated parallel solver based on recursive Schur complement reduction, which hierarchically decomposes the original problem into independent, batch-processable subproblems. To maximize GPU throughput, we design customized batched BLAS/LAPACK kernels and optimize task partitioning and memory access patterns via CUDA. The resulting open-source, cross-platform solver—TBD-GPU—demonstrates substantial speedups over state-of-the-art CPU-based sparse solvers (e.g., CHOLMOD, HSL MA57) across multiple benchmark problems, while matching the performance of NVIDIA cuDSS. These results validate the effectiveness and advancement of structure-aware batched computation for solving structured dense systems on GPUs.

Implementing parallel GPU factorization using batched BLAS/LAPACK kernelsOptimizing performance for time-dependent estimation and control problemsSolving block-tridiagonal symmetric positive definite linear systems

This work addresses the challenge of efficiently solving numerous related linear programming (LP) subproblems arising in mixed-integer programming (MIP), particularly in contexts such as strong branching and bound tightening, where traditional methods fail to exploit GPU parallelism effectively. The authors propose a GPU-oriented batched first-order optimization method, reformulating the primal–dual hybrid gradient algorithm into matrix–matrix operations to significantly enhance parallel efficiency. This study represents the first systematic application of batched first-order methods to LP subproblems within MIP solvers, advocating that GPUs should perform core computational tasks rather than merely assist CPU-based heuristics. The approach promotes deeper co-design between MIP algorithms and GPU architectures. Experimental results demonstrate that the proposed method outperforms conventional simplex solvers under specific problem scales and hardware configurations.

batched linear programsGPU accelerationmixed-integer programming

Nearest Neighbors GParareal: Improving Scalability of Gaussian Processes for Parallel-in-Time Solvers

May 20, 2024
GG
Guglielmo Gattiglio
🏛️ University of Warwick | University of St. Gallen

To address the poor scalability of Gaussian process (GP)-driven parallel-in-time (PinT) methods—such as GParareal—in high-dimensional systems and large-scale parallel environments, this paper proposes nnGParareal. The key innovation is the first integration of nearest-neighbor sampling into the GP-PinT framework, combined with adaptive sample selection and neighbor-based data compression. This reduces the GP model complexity from O(N³) to O(N log N) while preserving numerical accuracy. Theoretical analysis establishes rigorous error bounds and guarantees on parallel speedup. Extensive experiments across nine representative dynamical systems—including stiff, chaotic, and high-dimensional cases—demonstrate that nnGParareal significantly outperforms both GParareal and classical Parareal: it achieves faster convergence, higher parallel efficiency, and superior strong scaling behavior.

Enhancing speed and automation in long-time integrationImproving scalability of Gaussian Processes in PinT solversReducing model complexity for high-dimensional systems

Latest Papers

What's happening recently
View more

This work addresses the challenge of scaling the Shampoo optimizer to large models due to its prohibitive computational overhead. The authors propose an efficient distributed implementation that stacks preconditioning blocks into a 3D tensor to enhance GPU utilization and accelerates the computation of matrix inverse square roots by combining Newton–Schulz iteration with Chebyshev polynomial approximation. Furthermore, they provide the first systematic analysis of how matrix scaling affects convergence behavior. Experimental results demonstrate that the proposed method achieves up to a 4.83× speedup in optimizer step time while maintaining optimal validation perplexity, with the Newton–Schulz variant yielding the best performance.

computational slowdowninverse matrix rootsoptimizer efficiency

This work addresses the high memory and computational complexity typically associated with solving three-dimensional partial differential equations on Cartesian grids. By exploiting tensor-product structure, the proposed method decomposes the 3D operator into one-dimensional banded kernels aligned with coordinate axes, thereby avoiding explicit assembly of the global matrix and enabling a matrix-free solution strategy. Within a unified framework that integrates diverse numerical approaches—including Kronecker product algebra, compact finite differences, isogeometric analysis, and direct diagonalization—the study systematically identifies three key techniques: multi-right-hand-side reshaping, sum factorization, and pencil-style MPI decomposition. These innovations collectively enhance hardware affinity and parallel scalability, reducing algorithmic complexity to O(N) and storage requirements to O(Nₓ + Nᵧ + N_z), thus enabling efficient large-scale 3D PDE simulations.

3D operatorsCartesian PDE solversKronecker-product

This work addresses the underutilization of CPU resources in modern heterogeneous high-performance computing (HPC) systems when solving large-scale symmetric positive-definite linear systems using GPU-only approaches. Leveraging the SYCL programming model, the authors present the first heterogeneous implementations of the conjugate gradient (CG) method and Cholesky decomposition that operate across multi-vendor CPU-GPU platforms, including NVIDIA, AMD, and Intel architectures. Experimental results demonstrate that the heterogeneous CG solver achieves up to 32% speedup over GPU-only execution on large matrices, while the heterogeneous Cholesky decomposition attains a 29% acceleration. Furthermore, across diverse hardware vendors, the Cholesky solver consistently delivers at least a 12% performance improvement, significantly enhancing both computational efficiency and portability.

CPU-GPU collaborationGPU accelerationheterogeneous computing

This work addresses the challenge of solving generalized linear models with cardinality constraints, where traditional branch-and-bound methods struggle to exploit GPU parallelism due to discrete variables, combinatorial structures, and nonlinear objectives. The paper introduces the first CPU-GPU cooperative branch-and-bound framework that enables efficient GPU batch processing. By incorporating node padding, lightweight custom CUDA kernels, and heterogeneous scheduling, the framework achieves batched parallel evaluation of irregular search nodes. Empirical results demonstrate 10–100× speedups on challenging instances while attaining zero optimality gap. Moreover, the approach uniquely supports exhaustive enumeration of the full Rashomon set, thereby enabling rigorous variable importance analysis and multi-criteria model selection.

branch and boundcardinality-constrained optimizationcombinatorial optimization

Hot Scholars

QX

Qiao Xiao

Eindhoven University of Technology
Deep LearningAI EfficiencySparse Neural Networks
ZZ

Zhihua Zhang

Professor of Computer Science, Shanghai Jiao Tong University
Artificial IntelligenceMachine Learning
EM

Elena Mocanu

Assistant Professor, University of Twente
Machine LearningSparse Neural Networks
SW

Shunxin Wang

University of Twente
computer visiondeep learning
MP

Mykola Pechenizkiy

Eindhoven University of Technology
data miningpredictive analyticsfairnesstransparency