database sharding

Designs, implements, and evaluates partitioning and placement schemes that split data, indexes, tensors, model parameters, activations, and gradients across storage and compute nodes; this includes creating sharding strategies (data/index/model/tensor/gradient), auto-sharding and memory-aware sharding algorithms, caching and replication policies, and integration or protocols for sharded training systems such as FSDP. Engineers working in this skill build sharding-aware indexing and caching layers, sharding/partitioning orchestration and synchronization mechanisms, and performance, consistency, and resource-usage analyses for those distributed setups.

databasesharding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.83
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the critical limitation of existing serverless federated learning systems, which struggle to train large models due to strict memory constraints imposed by serverless functions. To overcome this barrier, the authors propose GradsSharding, a novel approach that shards gradient tensors and processes them in parallel across serverless functions, with each function aggregating only its assigned shard. This method achieves mathematically equivalent results to conventional tree-based aggregation while attaining constant memory consumption per function—decoupled from the number of participating clients—for the first time. Implemented on AWS Lambda, the sharded FedAvg algorithm demonstrates effectiveness across model sizes ranging from 43 MB to 5 GB, reduces training costs by 2.7× on VGG-16, and stands as the only serverless federated learning solution capable of surpassing the 10 GB memory ceiling.

Federated LearningGradient AggregationMemory Constraint

Existing Fully Sharded Data Parallel (FSDP) systems are constrained by fixed sharding formats, which hinder efficient support for structure-aware training techniques—such as block-wise quantization—and non-elementwise optimizers like Shampoo and Muon, thereby limiting scalability to ten-thousand-GPU clusters. This work proposes RaggedShard, a flexible sharding mechanism coupled with a structure-aware scheduling algorithm, which natively supports these advanced methods within FSDP while maintaining minimal code intrusion. Experimental results demonstrate that the proposed approach achieves 5%–66% higher throughput and reduces memory consumption by 16%–30% on large-scale GPU clusters, successfully enabling efficient ten-thousand-GPU-scale training.

FSDPlarge-scale trainingnon-element-wise optimizers

A Distributed Partitioning Software and its Applications

Mar 04, 2025
AS
Aparna Sasidharan
🏛️ University of Illinois

To address the high overhead of dynamic data repartitioning in multi-core HPC systems under time-varying workloads, this paper proposes a lightweight, hierarchical partitioning method jointly driven by geometric and statistical principles. The method integrates space-filling curve ordering, greedy knapsack-based load balancing, and hierarchical data decomposition to support efficient dynamic partitioning of 2D/3D structured grids, point sets, and general graphs. It introduces, for the first time, an adaptive repartitioning mechanism guided by real-time feedback on data distribution, substantially reducing computational and communication overhead in frequently updated scenarios. Implemented via a hybrid parallel programming model (MPI + OpenMP) on modern many-core architectures, experimental results demonstrate a 3.2–5.7× speedup in partitioning time and a load imbalance ratio below 3.1%. This approach provides timely, low-overhead data partitioning support for parallel algorithms in large-scale scientific computing.

Applies geometric and statistical methods for hierarchical data decomposition.Develops software for efficient data partitioning on many-core HPC machines.Optimizes dynamic applications with time-varying load distributions.

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

Latest Papers

What's happening recently
View more

This study addresses the challenges of fragmented academic computing resources and cross-facility large model pre-training by proposing a distributed training framework that integrates multiple global supercomputing centers. Methodologically, building upon the DiLoCo dual-loop architecture, it introduces elastic Nesterov outer-step optimization, a DARL heartbeat-based data leasing protocol, and a privilege-free queue-aware placement mechanism to enable resilient aggregation and efficient scheduling of intercontinental resources. Experimental results demonstrate that the system achieves fault-tolerant pre-training with zero data loss while reducing overhead to 3.1%, significantly shortening training cycles. This work provides an efficient and viable new paradigm for decentralized, large-scale LLM training.

Cross-facility pre-trainingDistributed LLM trainingFragmented compute allocations

This work addresses the challenges of poor scalability and accuracy degradation in scientific machine learning when handling extremely high-resolution data, particularly due to the absence of a general-purpose parallelization framework supporting sub-unit batch sizes per device. The authors propose ShardTensor, a novel domain-parallel paradigm that shards tensors along spatial domains, thereby decoupling data dimensions from hardware constraints. This approach enables, for the first time, general-purpose parallel training and inference with sub-unit batch sizes. By supporting multidimensional parallelism and integrating dynamic computation–communication load balancing, ShardTensor simultaneously achieves strong scaling—reducing latency—and weak scaling—enabling larger-scale data processing—thereby significantly enhancing the scalability and efficiency of high-fidelity scientific computing tasks.

domain parallelismextreme-resolution datainput data parallelization

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

Hot Scholars

FW

Feiyi Wang

Distinguished Research Scientist & Group Leader, Analytics and AI Methods at Scale, NCCS/ORNL
HPCAI for Science at Scale
TD

Tri Dao

Princeton University, Together AI
Machine learningSystems
HL

Haibin Lin

Bytedance
Machine Learning SystemsNatural Language Processing
YP

Yanghua Peng

ByteDance Inc.
Large Language ModelsMachine Learning SystemsGPU Scheduling
GL

Guoliang Li

Professor, Tsinghua University
DatabaseBig DataCrowdsourcingData Cleaning & Integration