ml workloads

Designs, implements, and evaluates machine learning workloads including training and inference pipelines, data preprocessing and feature extraction, hyperparameter searches, batch and streaming jobs, and model serving. Analyzes and optimizes their execution characteristics—throughput, latency, resource utilization, scalability, reliability, and cost—and addresses scheduling, orchestration, checkpointing, monitoring, and other operational aspects of the model lifecycle.

mlworkloads

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

Machine Learning-Driven Predictive Resource Management in Complex Science Workflows

Sep 14, 2025
TC
Tasnuva Chowdhury
🏛️ Brookhaven National Laboratory | University of Massachusetts, Amherst | University of Pittsburgh | Carnegie Mellon University | Oak Ridge National Laboratory | SLAC National Accelerator Laboratory

Accurately estimating resource requirements for scientific workflows remains challenging due to diverse analytical scenarios, varying user expertise, and highly heterogeneous computing platforms. To address this, we propose an end-to-end machine learning framework that directly learns CPU, memory, and runtime requirements for each workflow step from historical task execution profiles—eliminating reliance on domain-specific heuristics or time-consuming two-stage trial runs. Integrated into the PanDA workflow management system, our framework enables proactive, fine-grained dynamic resource pre-allocation. Experimental evaluation in large-scale scientific computing environments—including the Large Hadron Collider (LHC)—demonstrates that our approach significantly outperforms baseline methods: average resource waste is reduced by 32%, scheduling latency decreases by 27%, and heterogeneous resource utilization and workflow execution stability are substantially improved.

Enabling optimal resource allocation using machine learningOvercoming inaccurate initial resource estimation challengesPredicting resource needs for complex scientific workflows

To address the I/O bottleneck in ML training—causing low GPU utilization (often <50%)—this paper proposes a data-driven approach for I/O performance prediction and storage configuration optimization. We conduct systematic benchmarking across 141 configurations spanning diverse storage backends (NVMe SSDs, network-attached storage, in-memory filesystems), data formats, and access patterns. Leveraging key features—including batch size and throughput—we train an XGBoost regression model achieving an R² of 0.991 and a mean absolute error of only 11.8%. The model enables minute-scale recommendation of optimal storage configurations, accelerating configuration search by over an order of magnitude compared to empirical trial-and-error. Our core contribution is the first general-purpose, ML-training-pipeline-aware I/O performance prediction framework, which significantly improves GPU utilization and end-to-end training efficiency. All code and benchmark data are fully open-sourced, ensuring strong reproducibility and extensibility.

Predicts I/O performance for ML training pipelinesRecommends optimal storage configurations to reduce idle timeUses data-driven modeling to replace trial-and-error with predictions

MLOps Monitoring at Scale for Digital Platforms

Apr 23, 2025
YJ
Yu Jeffrey Hu
🏛️ Purdue University | Essec Business School | Maastricht University

Massive, dynamic data streams in digital platforms render conventional ML monitoring methods ineffective or prohibitively costly in manual effort, forcing enterprises to downgrade to simpler models. Method: This paper proposes the Machine Learning Monitoring Agent (MLMA) framework, introducing a test-driven, automated retraining mechanism based on data-adaptive reference loss batches—designed to enable efficient closed-loop operations while preserving human-in-the-loop collaborative governance. The approach integrates design science principles, dynamic reference loss computation, key metric visualization, and human–AI collaborative workflows. Contribution/Results: Evaluated on a large-scale instant-delivery platform, MLMA supports concurrent monitoring of hundreds of models, significantly reduces manual intervention frequency, and sustains long-term online model performance stability. Its core contribution lies in unifying dynamic data adaptation, automated trigger logic, and human–AI collaboration—thereby overcoming critical technical bottlenecks in real-time monitoring and adaptive maintenance of large-scale ML systems.

Automating re-training to maintain model performance at scaleMonitoring ML models in large unstable data streamsReducing labor-intensive MLOps supervision in digital platforms

Sizey: Memory-Efficient Execution of Scientific Workflow Tasks

Jul 23, 2024
JB
Jonathan Bader
🏛️ Technische Universität Berlin | Humboldt-Universität zu Berlin | University of Glasgow

Scientific workflow tasks exhibit heterogeneous inputs and types, making memory demand prediction challenging—leading to resource over-allocation, reduced cluster throughput, and increased task failure risk. To address this, we propose an online dynamic memory prediction method featuring novel parallel online multi-model training and adaptive model selection. We introduce Resource Allocation Quality (RAQ) as a new metric for evaluating allocation efficacy and support continuous runtime retraining and real-time optimization. Our approach integrates ensemble machine learning, online learning, fine-grained resource monitoring, and workflow runtime analysis. Evaluated on six real-world nf-core workflows, our method reduces median memory waste by 24.68% compared to the state-of-the-art baseline, significantly improving resource utilization and system throughput.

Memory-efficient execution of workflowsOnline memory prediction for tasksReduction in memory waste

Latest Papers

What's happening recently
View more

This work addresses the lack of a universal, flexible, and cluster-agnostic workload representation in existing distributed machine learning systems, which hinders efficient design space exploration. To overcome this limitation, the paper introduces Flint, a novel framework that leverages the intermediate representation of machine learning compilers to extract workload graphs for clusters of arbitrary scale—without requiring actual hardware execution. By decoupling workload modeling from underlying hardware specifics and validating accuracy through execution traces, Flint ensures both fidelity and portability. Experimental results demonstrate that Flint effectively enables flexible and efficient design space exploration while substantially reducing evaluation overhead.

compiler intermediate representationdesign space explorationdistributed machine learning

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

This work addresses customs clearance delays in global trade caused by ambiguous product descriptions and frequent updates to Harmonized System (HS) codes. To tackle this challenge, the authors propose a serverless MLOps framework that leverages event-driven pipelines and managed services to enable end-to-end, model-agnostic machine learning lifecycle management. The architecture supports automatic scaling, reproducible training, auditable deployment, and automated A/B testing, ensuring secure and seamless model transitions. By integrating custom text embeddings with models such as Text-CNN, the system achieves 98% accuracy on real-world HS code prediction tasks, meeting stringent service-level agreement (SLA) requirements. This approach significantly reduces long-term operational costs and establishes an efficient, cost-effective, and reproducible deployment paradigm for industrial-scale machine learning systems.

Harmonized System Code PredictionIndustrial Machine LearningMLOps

Data Virtualization for Machine Learning

Jul 23, 2025
SK
Saiful Khan
🏛️ Rutherford Appleton Laboratory Science and Technology Facilities Council (STFC) | University of Oxford

Machine learning teams face significant challenges in multi-workflow concurrent environments, including redundant intermediate data storage, inefficient cross-pipeline sharing, and high collaboration overhead. To address these issues, this paper proposes and implements a data virtualization service architecture tailored for ML workflows. The architecture adopts a service-oriented design, integrating distributed data management with dynamic metadata mapping to enable logical abstraction, on-demand loading, and unified access to heterogeneous intermediate data. Compared to conventional materialized storage approaches, it reduces storage overhead by an average of 62% (measured empirically) and substantially decreases inter-team collaboration latency. The system has been deployed in production, stably supporting six ML applications and over thirty concurrent workflows, demonstrating linear scalability. This work establishes a lightweight, elastic, and reusable data virtualization paradigm for large-scale ML infrastructure.

Handling large amounts of intermediate data storageManaging multiple concurrent ML workflows efficientlyReducing time from data wrangling to model deployment