implement tensor parallelism

Designs and implements parallel execution and data-mapping strategies for multi-device tensor computations and tensor-core hardware, including tensor model parallelism and combined tensor-and-model parallel schemes. Builds and tunes tensor kernels and operators using tensorization, tiling, decomposition, fusion and other kernel-level optimizations, programs and optimizes tensor-core codepaths, and analyzes or simulates tensor-network performance to validate and optimize throughput and resource use.

implementtensorparallelism

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Can Tensor Cores Benefit Memory-Bound Kernels? (No!)

Feb 24, 2025
LZ
Lingqi Zhang
🏛️ RIKEN | University of California, Riverside | Argonne National Laboratory

This work investigates whether Tensor Cores deliver practical acceleration for memory-bound kernels—such as STREAM Scale, SpMV, and stencil computations—challenging recent studies that overestimate their performance in such scenarios. Method: The authors adopt a dual approach: (i) theoretically deriving the upper bound of double-precision speedup under GPU microarchitectural constraints (e.g., memory bandwidth and instruction scheduling overhead), and (ii) empirically validating across V100, A100, and H100 GPUs using both CUDA and WMMA APIs on representative memory-bound kernels. Contribution/Results: The analysis reveals a strict theoretical speedup ceiling of 1.33× for double-precision operations; all empirical measurements fall at or below this bound. The study refutes the efficacy of Tensor Cores in memory-bottlenecked workloads, attributing prior overestimations to neglect of bandwidth saturation and scheduling latency. It establishes the first principled theoretical foundation for Tensor Core applicability boundaries, providing a critical criterion for heterogeneous resource scheduling in GPU-accelerated computing.

Comparison with CUDA cores in memory-bound tasksTensor cores efficacy in memory-bound kernelsTheoretical and empirical analysis of performance

Pushing Tensor Accelerators Beyond MatMul in a User-Schedulable Language

Dec 02, 2025
YZ
Yihong Zhang
🏛️ University of Washington | Adobe

Tensor accelerators (e.g., NVIDIA Tensor Cores) are increasingly prevalent in CPUs and GPUs, yet their programmability remains limited: existing kernel libraries target only traditional ML and scientific computing workloads, failing to support non-ML linear matrix transform workloads such as image processing. This paper proposes an equivalence-saturation–based flexible tensor instruction selection mechanism, enabling general-purpose, schedulable compilation for tensor hardware. Integrating with the Halide domain-specific language and compiler, our approach retains full compatibility with existing scheduling primitives while substantially broadening the programmability of tensor accelerators. Evaluated on an NVIDIA RTX 4070, our framework achieves a 6.1× speedup on image processing pipelines—including downsampling—demonstrating, for the first time, systematic performance acceleration of tensor hardware in non-ML domains.

Demonstrating performance improvements for image processing pipelines via tensor hardwareEnabling tensor accelerators for diverse applications beyond traditional MatMul operationsOvercoming programming difficulties of tensor accelerators using compiler-based techniques

This work addresses the performance limitations of data-intensive applications in edge computing, which often suffer from poor memory locality or excessive memory footprint. To overcome these challenges, the authors propose a hardware-software co-design that integrates a tensor memory engine into the general-purpose CPU datapath. This engine enables runtime dynamic reconfiguration of memory layouts to significantly enhance cache locality, without requiring application code modifications, offloading computation to memory, or incurring memory overhead. By preserving the CPU’s role as the primary compute unit, the approach effectively improves overall system performance. Notably, this study presents the first practical implementation of runtime memory layout reconfiguration on commercial SoC/FPGA platforms, demonstrating both innovation and real-world applicability.

cache localitydata-intensive applicationsmemory layout

Although stencil computations are traditionally considered memory-bound, they exhibit significant acceleration on Tensor Cores—an apparent contradiction that demands explanation. This work proposes the first applicability analysis framework for stencil computations on Tensor Cores, leveraging systematic performance modeling to quantify the computational redundancy introduced by time tiling and its impact on arithmetic intensity. Based on this analysis, the study derives precise criteria for achieving effective acceleration and identifies the corresponding “sweet spot” region in the design space. Furthermore, it reveals that sparse Tensor Cores can substantially expand this design space. Experimental validation on state-of-the-art implementations such as DRStencil and EBISU confirms the model’s accuracy, offering theoretical guidance for the efficient optimization of stencil computations on NVIDIA GPUs.

arithmetic intensitymemory-boundperformance contradiction

Latest Papers

What's happening recently
View more

Existing scientific computing codes are difficult to efficiently port to specialized architectures such as AMD AI Engine, often requiring extensive manual refactoring. This work proposes a tensor abstraction–based compilation approach that automatically elevates generic loops to tensor semantics by parsing lightweight OpenMP annotations, and constructs an end-to-end compilation pipeline to map computations onto the AI Engine execution model. The method significantly reduces programming complexity through minimal OpenMP directives and enables CPU–NPU cooperative scheduling. Experimental results on six scientific and AI kernel benchmarks show that the NPU achieves higher energy efficiency than a multi-core CPU at float32 precision; for two kernels, cooperative execution yields a 40% performance improvement and 15% energy reduction.

AI Enginescode portinghardware acceleration

This work investigates whether handcrafted PTX kernels can outperform the WMMA API for mixed-precision GEMM on NVIDIA L4 GPUs (Ada architecture). We design double-buffered GEMM kernels leveraging PTX instructions such as cp.async, ldmatrix, and mma.sync, and conduct a systematic comparison against WMMA implementations across FP16, INT8, and INT4 precisions, supported by in-depth hardware analysis using Nsight Compute. For the first time on the L4 platform, we quantify the performance limits achievable with hand-optimized PTX, revealing that memory behavior—not Tensor Core utilization—dominates large-matrix performance. Our experiments demonstrate speedups of 1.4×–1.8× for INT8 and 2.9×–4.3× for INT4 over WMMA; compared to the FP16 WMMA baseline at N=8192, peak improvements reach 34.4× for INT8 and 98.7× for INT4.

GEMMmulti-precisionPTX

This work addresses the limited flexibility in distributed programming for large language model scaling and the inefficiency of existing tensor compilers in handling the complex memory hierarchies of heterogeneous clusters. To overcome these challenges, the authors propose a scalable block-level compiler featuring a novel three-tier hierarchical abstraction—Core, Device, and Task—that uniformly supports diverse parallelization strategies, automatically optimizes intra- and inter-node communication, and enables efficient code generation across both NVIDIA and AMD platforms. When integrated into vLLM, the compiler achieves 5%–30% end-to-end inference speedup and over 10% improvement in training model FLOPs utilization (MFU), translating to approximately 500,000 GPU hours saved per month. The system has been deployed in enterprise settings, delivering over 20% inference performance gains.

distributed programminglarge language modelsmemory hierarchy

This work addresses the limited scalability of tensor parallelism in large model online inference, where non-scalable overheads hinder near-linear cluster performance scaling. The authors propose Albireo, a system that eliminates such bottlenecks without modifying model architecture by overlapping scheduling with computation, employing sequence-parallel sampling, and optimizing KV cache management. Albireo further introduces the concept of “empirically optimal tensor parallelism degree” to guide parallelism strategy selection. Experimental results demonstrate that, compared to vLLM, Albireo achieves up to 1.9× higher throughput, 48% lower latency, 28% improved GPU utilization, and 54% reduced energy consumption, with a twofold throughput gain observed in production environments.

Amdahl's LawGPU clusterLLM inference

Hot Scholars

AM

Alejandro Mata Ali

Quantum Team Coordinator, ITCL/Lecturer of MIAX, BME/Teacher
Quantum Computingtensor networksapplied mathematics
ZJ

Ziheng Jiang

Research Scientist, ByteDance
SystemsMachine Learning
HL

Haibin Lin

Bytedance
Machine Learning SystemsNatural Language Processing
QZ

Qibin Zhao

RIKEN AIP
Machine LearningTensor DecompositionTensor Networks
TD

Tri Dao

Princeton University, Together AI
Machine learningSystems