Score
Designs and implements schemes that partition rows of weight matrices or activation tensors across compute devices (1-D/rowwise sharding and batching-aware variants), including data placement, communication, and inference-time assembly, to reduce per-device memory and runtime while preserving model outputs without retraining.
This work addresses the GPU memory bottleneck in Transformer models caused by high parameter and activation memory demands during training and inference. The authors propose a novel parallelism strategy that integrates tensor parallelism (TP) and sequence parallelism (SP) along the same device axis, enabling each device to simultaneously shard both model weights and input sequences. By leveraging broadcast-based weight sharding with key-value exchange in attention layers and ring-based weight passing with local accumulation in gated MLPs, the method achieves dual compression of both parameter and activation memory. This approach significantly reduces per-device memory consumption, outperforming conventional TP, SP, and their hybrid variants. It demonstrates superior hardware adaptability and scaling efficiency under long-context and memory-constrained settings, while seamlessly integrating with pipeline and expert parallelism.
This work addresses the performance limitations of data-intensive applications in edge computing, which often suffer from poor memory locality or excessive memory footprint. To overcome these challenges, the authors propose a hardware-software co-design that integrates a tensor memory engine into the general-purpose CPU datapath. This engine enables runtime dynamic reconfiguration of memory layouts to significantly enhance cache locality, without requiring application code modifications, offloading computation to memory, or incurring memory overhead. By preserving the CPU’s role as the primary compute unit, the approach effectively improves overall system performance. Notably, this study presents the first practical implementation of runtime memory layout reconfiguration on commercial SoC/FPGA platforms, demonstrating both innovation and real-world applicability.
To address high per-node memory pressure and substantial intra-node communication overhead in distributed pre-training of large-scale models, this paper proposes Subnet Data Parallel (SDP): each worker node independently trains a structured, compact subnetwork, eliminating activation transmission across pipeline stages. SDP integrates stochastic block dropping with width-wise subnetwork construction to ensure uniform parameter coverage and gradient alignment across distributed workers. Its communication bandwidth requirement is comparable to or lower than that of standard all-reduce operations. Experiments demonstrate that SDP reduces GPU memory consumption by 20–40% without compromising model accuracy, while preserving convergence properties and significantly improving training efficiency for large models.
Existing distributed matrix multiplication algorithms support only limited partitioning schemes; mismatched configurations necessitate redundant data redistribution, incurring substantial communication overhead. Method: We propose a universal one-sided algorithm that unifies arbitrary tiling strategies and replication factors via slice-index arithmetic, eliminating dependence on specialized algorithm implementations. Leveraging a high-order C++ PGAS framework, it integrates GPU-to-GPU direct communication and intra-node high-speed interconnects, dynamically generating and reordering local compute tasks to maximize computation-communication overlap. Contribution/Results: This is the first implementation to natively support all data distribution patterns within a single codebase, significantly improving system flexibility and maintainability. Experimental evaluation demonstrates performance competitive with PyTorch DTensor across diverse tiling and replication configurations, validating its efficiency and generality for AI training and scientific computing workloads.
To address the high memory movement overhead (≈50% of execution time spent on tensor reordering) and low energy efficiency of Kronecker-sparse (KS) matrix multiplication on GPUs, this work proposes the first KS-structure-aware tiling memory access strategy, co-optimizing GPU’s multi-level memory hierarchy to significantly reduce redundant data reads and writes. We further introduce the first unified benchmarking framework jointly evaluating energy consumption and execution time for KS-sparse operators. Our CUDA-based kernel achieves a median 1.4× speedup and 15% energy reduction across mainstream KS problem sizes. It has been successfully integrated into Transformer inference pipelines, demonstrating practical deployment value. The core innovations lie in (i) a KS-structure-guided tiling design that exploits inherent sparsity and Kronecker factorization patterns, and (ii) a hardware–software co-designed energy-efficiency optimization paradigm tailored to KS computation.
This work addresses the challenge of running large language models on consumer-grade hardware, where limited GPU memory forces existing systems to rely on coarse-grained CPU offloading strategies that fail to account for intra-layer tensor heterogeneity and dynamic hardware load. To overcome these limitations, we propose ATSInfer, the first system to enable tensor-level fine-grained offloading scheduling. ATSInfer integrates static placement with load-aware dynamic data transfer and introduces an asynchronous CPU-GPU coordination mechanism to efficiently orchestrate storage, data movement, and computation resources. Experimental results demonstrate that ATSInfer achieves up to 1.94× and 3.29× higher throughput during the prefill and decoding phases, respectively, compared to state-of-the-art approaches, while significantly improving GPU utilization and PCIe bandwidth efficiency.
为解决大规模AI模型训练资源分配问题,提出ShardMeter,通过分析模型和硬件特性预测性能,帮助优化分布式训练配置。
This work addresses the conflict between conventional page-granularity interleaved memory layouts and the locality demands of GEMM operations on chiplet-based GPUs, which incurs substantial remote HBM access overhead. To resolve this, the authors propose Chiplet-Contiguous Layout—a hardware- and OS-agnostic global memory organization that colocates each chiplet’s local data contiguously, thereby enabling locality-aware scheduling while remaining compatible with standard 4KB page management. This approach is the first to unify data locality optimization for GEMM with page-granular memory management in both large language model inference and training. Experiments demonstrate that, compared to 4KB interleaved layouts, the proposed method reduces remote HBM traffic by 24.7× and 19.2× on Qwen-3 30B and Llama-3.1 70B, respectively, and further achieves 4.1× and 2.1× reductions over coarse-grained locality-aware placements.
This work addresses the challenges of poor scalability and accuracy degradation in scientific machine learning when handling extremely high-resolution data, particularly due to the absence of a general-purpose parallelization framework supporting sub-unit batch sizes per device. The authors propose ShardTensor, a novel domain-parallel paradigm that shards tensors along spatial domains, thereby decoupling data dimensions from hardware constraints. This approach enables, for the first time, general-purpose parallel training and inference with sub-unit batch sizes. By supporting multidimensional parallelism and integrating dynamic computation–communication load balancing, ShardTensor simultaneously achieves strong scaling—reducing latency—and weak scaling—enabling larger-scale data processing—thereby significantly enhancing the scalability and efficiency of high-fidelity scientific computing tasks.
该研究针对移动异构推理任务,提出一种分区感知调度方法,通过在线迭代搜索框架优化DAG调度,实现低延迟和高效执行。