perform memory accounting

Design and build methods and tooling to track, model, and predict memory consumption across processes or devices in distributed systems, producing fine‑grained per‑rank or per‑node memory usage estimates. Use those estimates to assign shard ownership, plan tensor partitioning and placement, and evaluate partitioning/replication strategies to balance peak memory, reduce fragmentation, and enable efficient scaling.

performmemoryaccounting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of enhancing productivity in supercomputing clusters and informing the design of exascale systems by analyzing job scheduling logs, GPU trace data, and domain-specific metadata from the Titan supercomputer. It systematically investigates the relationship between requested and actual resource utilization and its temporal evolution. Employing correlation analysis, clustering, and neural networks, the work presents the first comprehensive characterization of seasonal patterns in HPC resource usage and develops a transferable model for predicting resource utilization. By identifying key user behavior patterns, the research substantially improves the accuracy of forecasting future resource demands, thereby providing empirical foundations for optimizing configuration and planning of high-performance computing systems.

HPCresource utilizationTitan supercomputer

This work addresses the asymmetric cost of memory allocation in distributed clusters, where under-allocation risks task failures while over-allocation leads to resource waste. To balance these competing concerns, the authors propose a memory provisioning strategy that integrates conditional quantile regression with a multiplicative safety factor. By ensembling LightGBM and XGBoost to predict high-quantile memory demands, the method explicitly trades off the risks of under- and over-allocation and characterizes their Pareto frontier. Evaluated on a real-world SAP build-task dataset, the approach reduces the fraction of memory-insufficient tasks from 4.17% to 2.89% while significantly cutting the average over-allocation rate from 148% to 44.51%.

distributed systemsmemory allocationoverallocation

Longitudinal Analysis of GPU Workloads on Perlmutter

Feb 25, 2025
OC
Onur Cankur
🏛️ University of Maryland | Lawrence Berkeley National Laboratory

This study addresses the inefficiency of GPU workloads on the Perlmutter supercomputing platform by conducting the first systematic spatiotemporal analysis based on hardware performance counters. Methodologically, it leverages LDMS to collect fine-grained GPU core and memory performance metrics, and introduces two novel metrics—*burstiness* and *temporal imbalance*—to quantify temporal utilization patterns; these are combined with spatial imbalance measures to comparatively analyze usage characteristics of ML and traditional HPC workloads. Results reveal significant spatiotemporal GPU resource imbalance in both workload types: substantial inter-GPU load variance (spatial) and highly volatile utilization over time (temporal), leading to considerable resource underutilization. The findings provide empirical evidence and a quantifiable evaluation framework to guide HPC scheduler optimization, heterogeneous architecture design, and ML–HPC convergence strategies.

Analyzes GPU hardware counters on PerlmutterCompares machine learning and HPC job efficienciesInvestigates spatial and temporal GPU workload imbalances

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

gpu_tracker: Python package for tracking and profiling GPU utilization in both desktop and high-performance computing environments

Apr 01, 2024
ED
Erik D. Huckvale
🏛️ University of Kentucky | Institute for Biomedical Informatics

Fine-grained, cross-platform real-time monitoring of GPU resources—particularly GPU memory peak usage and computational utilization—remains unsupported in Unix/Linux environments. Method: This paper introduces the first lightweight, dependency-free Python tool leveraging the NVIDIA Management Library (NVML) API. It employs multithreading and process-hooking techniques to enable low-overhead (average 0.3%) background sampling and precise peak capture of CPU/GPU utilization and system/GPU memory consumption. Contribution/Results: The tool unifies analysis across desktop and HPC environments with high accuracy (GPU memory peak error <2%). It enables job-level GPU resource profiling—the first such capability for fine-grained, runtime GPU characterization in HPC settings—thereby addressing a critical gap in production-grade GPU observability. The implementation is open-source and has been integrated into multiple scientific computing pipelines.

Monitors maximum RAM usage on motherboard and GPUProvides real-time resource profiling with minimal overheadTracks GPU and CPU utilization in HPC environments

Latest Papers

What's happening recently
View more

This work addresses the low efficiency of GPU resource and power utilization in heterogeneous high-performance computing (HPC) systems, as well as the lack of accurate predictive methods. To tackle these challenges, the authors propose a two-stage prediction framework that uniquely integrates Slurm job logs with fine-grained NVIDIA Data Center GPU Manager (DCGM) metrics. Leveraging only job submission features, the framework achieves high-accuracy predictions of application-level average power consumption, peak GPU utilization, and memory utilization. Experimental results demonstrate prediction accuracies of 97% for peak GPU utilization and 92% for runtime power consumption, significantly enhancing scheduling efficiency and power management capabilities in HPC environments. The study also validates the effectiveness of DCGM metrics in characterizing application behavior.

GPU resource predictionGPU utilizationheterogeneous HPC systems

This study addresses the challenge of inaccurate energy consumption estimation for distributed batch-processing applications like Apache Spark in cloud environments, where node-level hardware energy counters are typically inaccessible. Focusing on Apache Spark deployed on Kubernetes, the work presents the first systematic comparison between resource-utilization-based energy models and ground-truth measurements from Intel RAPL across both AWS bare-metal instances and on-premises clusters. It investigates the impact of CPU and memory utilization signals on estimation accuracy and introduces external monitoring to enhance model fidelity. Experimental results demonstrate that incorporating external monitoring significantly mitigates energy underestimation—reducing the error from −29.58% to −24.41% on AWS and from −24.00% to −16.22% in the local cluster—thereby validating its effectiveness in improving energy estimation accuracy.

cloud computingdistributed data processingenergy estimation

This study addresses the uncertain practical efficacy of dynamic resource management—particularly MPI variability—in high-performance computing (HPC) environments. To bridge this gap, the authors propose a methodology based on replaying real job logs to faithfully reproduce actual workloads on a 125-node partition of the MareNostrum 5 supercomputer, thereby offering the first validation of MPI variability in a real-world HPC setting. By introducing a parallel-efficiency-oriented variability strategy integrated with resource scheduling and performance monitoring mechanisms, the experiments demonstrate a 27% reduction in execution time for variable workloads without compromising baseline job performance or overall resource utilization. These results provide compelling evidence of the approach’s effectiveness and practical value in authentic user scenarios.

Dynamic Resource ManagementHPCMPI malleability

Hot Scholars

TZ

Tieying Zhang

Research Scientist at Bytedance
AI for SystemsSystems for AI
WL

Wenhao Li

Xiamen University
LMSysEfficient LLM
JC

Jonghyun Choi

Associate Professor, Electrical and Computer Engineering, Seoul National University
Computer VisionMachine LearningContinual LearningEmbodied AI
LX

Le Xu

Bytedance
Computer ScienceDistributed SystemsCloud ComputingStream Processing
HZ

Hao Zhang

Tsinghua University
Computer visionComputer graphicsRobotics