row sharding

Designs and implements schemes that partition rows of weight matrices or activation tensors across compute devices (1-D/rowwise sharding and batching-aware variants), including data placement, communication, and inference-time assembly, to reduce per-device memory and runtime while preserving model outputs without retraining.

rowsharding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$254K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the GPU memory bottleneck in Transformer models caused by high parameter and activation memory demands during training and inference. The authors propose a novel parallelism strategy that integrates tensor parallelism (TP) and sequence parallelism (SP) along the same device axis, enabling each device to simultaneously shard both model weights and input sequences. By leveraging broadcast-based weight sharding with key-value exchange in attention layers and ring-based weight passing with local accumulation in gated MLPs, the method achieves dual compression of both parameter and activation memory. This approach significantly reduces per-device memory consumption, outperforming conventional TP, SP, and their hybrid variants. It demonstrates superior hardware adaptability and scaling efficiency under long-context and memory-constrained settings, while seamlessly integrating with pipeline and expert parallelism.

memory-efficientmodel trainingsequence parallelism

This work addresses the performance limitations of data-intensive applications in edge computing, which often suffer from poor memory locality or excessive memory footprint. To overcome these challenges, the authors propose a hardware-software co-design that integrates a tensor memory engine into the general-purpose CPU datapath. This engine enables runtime dynamic reconfiguration of memory layouts to significantly enhance cache locality, without requiring application code modifications, offloading computation to memory, or incurring memory overhead. By preserving the CPU’s role as the primary compute unit, the approach effectively improves overall system performance. Notably, this study presents the first practical implementation of runtime memory layout reconfiguration on commercial SoC/FPGA platforms, demonstrating both innovation and real-world applicability.

cache localitydata-intensive applicationsmemory layout

Model Parallelism With Subnetwork Data Parallelism

Jul 11, 2025
VS
Vaibhav Singh
🏛️ Mila - Quebec AI Institute | Concordia University | ISIR – Sorbonne Universite

To address high per-node memory pressure and substantial intra-node communication overhead in distributed pre-training of large-scale models, this paper proposes Subnet Data Parallel (SDP): each worker node independently trains a structured, compact subnetwork, eliminating activation transmission across pipeline stages. SDP integrates stochastic block dropping with width-wise subnetwork construction to ensure uniform parameter coverage and gradient alignment across distributed workers. Its communication bandwidth requirement is comparable to or lower than that of standard all-reduce operations. Experiments demonstrate that SDP reduces GPU memory consumption by 20–40% without compromising model accuracy, while preserving convergence properties and significantly improving training efficiency for large models.

Achieves lower memory usage without performance lossAvoids inter-node activation communication with structured subnetworksReduces memory demands in distributed pre-training of large models

Existing distributed matrix multiplication algorithms support only limited partitioning schemes; mismatched configurations necessitate redundant data redistribution, incurring substantial communication overhead. Method: We propose a universal one-sided algorithm that unifies arbitrary tiling strategies and replication factors via slice-index arithmetic, eliminating dependence on specialized algorithm implementations. Leveraging a high-order C++ PGAS framework, it integrates GPU-to-GPU direct communication and intra-node high-speed interconnects, dynamically generating and reordering local compute tasks to maximize computation-communication overlap. Contribution/Results: This is the first implementation to natively support all data distribution patterns within a single codebase, significantly improving system flexibility and maintainability. Experimental evaluation demonstrates performance competitive with PyTorch DTensor across diverse tiling and replication configurations, validating its efficiency and generality for AI training and scientific computing workloads.

Develops universal algorithm for all distributed matrix partitioningsEliminates operand redistribution costs in distributed matrix multiplicationUses slicing technique to compute overlapping tile multiplications

Fast inference with Kronecker-sparse matrices

May 23, 2024
AG
Antoine Gonon
🏛️ Univ Lyon | EnsL | UCBL | CNRS | Inria | valeo.ai

To address the high memory movement overhead (≈50% of execution time spent on tensor reordering) and low energy efficiency of Kronecker-sparse (KS) matrix multiplication on GPUs, this work proposes the first KS-structure-aware tiling memory access strategy, co-optimizing GPU’s multi-level memory hierarchy to significantly reduce redundant data reads and writes. We further introduce the first unified benchmarking framework jointly evaluating energy consumption and execution time for KS-sparse operators. Our CUDA-based kernel achieves a median 1.4× speedup and 15% energy reduction across mainstream KS problem sizes. It has been successfully integrated into Transformer inference pipelines, demonstrating practical deployment value. The core innovations lie in (i) a KS-structure-guided tiling design that exploits inherent sparsity and Kronecker factorization patterns, and (ii) a hardware–software co-designed energy-efficiency optimization paradigm tailored to KS computation.

Enhancing speed and energy efficiency in deep learning modelsImproving GPU kernel efficiency for Kronecker-sparse matricesReducing high data movement costs in KS matrix multiplication

Latest Papers

What's happening recently
View more

This work addresses the challenge of running large language models on consumer-grade hardware, where limited GPU memory forces existing systems to rely on coarse-grained CPU offloading strategies that fail to account for intra-layer tensor heterogeneity and dynamic hardware load. To overcome these limitations, we propose ATSInfer, the first system to enable tensor-level fine-grained offloading scheduling. ATSInfer integrates static placement with load-aware dynamic data transfer and introduces an asynchronous CPU-GPU coordination mechanism to efficiently orchestrate storage, data movement, and computation resources. Experimental results demonstrate that ATSInfer achieves up to 1.94× and 3.29× higher throughput during the prefill and decoding phases, respectively, compared to state-of-the-art approaches, while significantly improving GPU utilization and PCIe bandwidth efficiency.

consumer deviceshybrid CPU-GPU inferenceLLM offloading

This work addresses the conflict between conventional page-granularity interleaved memory layouts and the locality demands of GEMM operations on chiplet-based GPUs, which incurs substantial remote HBM access overhead. To resolve this, the authors propose Chiplet-Contiguous Layout—a hardware- and OS-agnostic global memory organization that colocates each chiplet’s local data contiguously, thereby enabling locality-aware scheduling while remaining compatible with standard 4KB page management. This approach is the first to unify data locality optimization for GEMM with page-granular memory management in both large language model inference and training. Experiments demonstrate that, compared to 4KB interleaved layouts, the proposed method reduces remote HBM traffic by 24.7× and 19.2× on Qwen-3 30B and Llama-3.1 70B, respectively, and further achieves 4.1× and 2.1× reductions over coarse-grained locality-aware placements.

chiplet GPUGEMMlocality-aware placement

This work addresses the challenges of poor scalability and accuracy degradation in scientific machine learning when handling extremely high-resolution data, particularly due to the absence of a general-purpose parallelization framework supporting sub-unit batch sizes per device. The authors propose ShardTensor, a novel domain-parallel paradigm that shards tensors along spatial domains, thereby decoupling data dimensions from hardware constraints. This approach enables, for the first time, general-purpose parallel training and inference with sub-unit batch sizes. By supporting multidimensional parallelism and integrating dynamic computation–communication load balancing, ShardTensor simultaneously achieves strong scaling—reducing latency—and weak scaling—enabling larger-scale data processing—thereby significantly enhancing the scalability and efficiency of high-fidelity scientific computing tasks.

domain parallelismextreme-resolution datainput data parallelization

Hot Scholars

FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
YS

Yujun Shen

Ant Group
Generative ModelingComputer VisionDeep Learning
CM

Chenglong Ma

Fudan University; Shanghai Innovation Institute
multi-modal modelsgenerative modelsmedical image analysis
HW

Hongqiu Wang

Hong Kong University of Science and Technology (Guangzhou)
AI for healthcareLabel-efficient learningMulti-modal learningFairness
SY

Semih Yavuz

Salesforce AI Research
large language modelsnatural language processingdeep learningmachine learning