Score
Designs, implements, and evaluates partitioning and placement schemes that split data, indexes, tensors, model parameters, activations, and gradients across storage and compute nodes; this includes creating sharding strategies (data/index/model/tensor/gradient), auto-sharding and memory-aware sharding algorithms, caching and replication policies, and integration or protocols for sharded training systems such as FSDP. Engineers working in this skill build sharding-aware indexing and caching layers, sharding/partitioning orchestration and synchronization mechanisms, and performance, consistency, and resource-usage analyses for those distributed setups.
该研究通过结合联邦学习与分片数据并行方法,解决了大规模计算中通信开销大的问题,提高了模型训练效率和质量。
This work addresses the critical limitation of existing serverless federated learning systems, which struggle to train large models due to strict memory constraints imposed by serverless functions. To overcome this barrier, the authors propose GradsSharding, a novel approach that shards gradient tensors and processes them in parallel across serverless functions, with each function aggregating only its assigned shard. This method achieves mathematically equivalent results to conventional tree-based aggregation while attaining constant memory consumption per function—decoupled from the number of participating clients—for the first time. Implemented on AWS Lambda, the sharded FedAvg algorithm demonstrates effectiveness across model sizes ranging from 43 MB to 5 GB, reduces training costs by 2.7× on VGG-16, and stands as the only serverless federated learning solution capable of surpassing the 10 GB memory ceiling.
Existing Fully Sharded Data Parallel (FSDP) systems are constrained by fixed sharding formats, which hinder efficient support for structure-aware training techniques—such as block-wise quantization—and non-elementwise optimizers like Shampoo and Muon, thereby limiting scalability to ten-thousand-GPU clusters. This work proposes RaggedShard, a flexible sharding mechanism coupled with a structure-aware scheduling algorithm, which natively supports these advanced methods within FSDP while maintaining minimal code intrusion. Experimental results demonstrate that the proposed approach achieves 5%–66% higher throughput and reduces memory consumption by 16%–30% on large-scale GPU clusters, successfully enabling efficient ten-thousand-GPU-scale training.
To address the high overhead of dynamic data repartitioning in multi-core HPC systems under time-varying workloads, this paper proposes a lightweight, hierarchical partitioning method jointly driven by geometric and statistical principles. The method integrates space-filling curve ordering, greedy knapsack-based load balancing, and hierarchical data decomposition to support efficient dynamic partitioning of 2D/3D structured grids, point sets, and general graphs. It introduces, for the first time, an adaptive repartitioning mechanism guided by real-time feedback on data distribution, substantially reducing computational and communication overhead in frequently updated scenarios. Implemented via a hybrid parallel programming model (MPI + OpenMP) on modern many-core architectures, experimental results demonstrate a 3.2–5.7× speedup in partitioning time and a load imbalance ratio below 3.1%. This approach provides timely, low-overhead data partitioning support for parallel algorithms in large-scale scientific computing.
To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.
为解决大规模AI模型训练资源分配问题,提出ShardMeter,通过分析模型和硬件特性预测性能,帮助优化分布式训练配置。
This study addresses the challenges of fragmented academic computing resources and cross-facility large model pre-training by proposing a distributed training framework that integrates multiple global supercomputing centers. Methodologically, building upon the DiLoCo dual-loop architecture, it introduces elastic Nesterov outer-step optimization, a DARL heartbeat-based data leasing protocol, and a privilege-free queue-aware placement mechanism to enable resilient aggregation and efficient scheduling of intercontinental resources. Experimental results demonstrate that the system achieves fault-tolerant pre-training with zero data loss while reducing overhead to 3.1%, significantly shortening training cycles. This work provides an efficient and viable new paradigm for decentralized, large-scale LLM training.
为解决分布式训练中的性能问题,提出HyperParallel-FSDP方法,通过双模式DTensor执行、拓扑感知FSDP及布局驱动的分布式Muon优化模型并行计算效率。
This work addresses the challenges of poor scalability and accuracy degradation in scientific machine learning when handling extremely high-resolution data, particularly due to the absence of a general-purpose parallelization framework supporting sub-unit batch sizes per device. The authors propose ShardTensor, a novel domain-parallel paradigm that shards tensors along spatial domains, thereby decoupling data dimensions from hardware constraints. This approach enables, for the first time, general-purpose parallel training and inference with sub-unit batch sizes. By supporting multidimensional parallelism and integrating dynamic computation–communication load balancing, ShardTensor simultaneously achieves strong scaling—reducing latency—and weak scaling—enabling larger-scale data processing—thereby significantly enhancing the scalability and efficiency of high-fidelity scientific computing tasks.
Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.