Score
Designs and constructs analytical, simulation, or statistical models that predict and explain performance metrics (throughput, latency, utilization, response time, scalability) of systems, processes, or algorithms. Builds workload characterizations, calibrates models from measurements, and performs validation and sensitivity or trade‑off analyses to support optimization and capacity planning.
This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
Scientific workflows on clusters suffer from inefficient resource scheduling, high energy consumption, and unpredictable costs due to inaccurate manual performance estimation. To address this, we propose an automated, task-level performance prediction method—estimating both execution time and memory consumption—by integrating machine learning (regression and ensemble models), fine-grained feature engineering, runtime performance modeling, and workflow semantic analysis. Our approach enables cross-platform, multi-objective (including carbon-aware) generalization. We present the first systematic survey and horizontal evaluation of mainstream prediction paradigms, identifying key limitations in dynamism, transferability, and multi-objective coordination, while charting their evolutionary trajectory. We establish a unified benchmarking framework and validate our method on real-world workflows (e.g., CyberShake, SIPHT), achieving 32–47% lower prediction error. This enables resource managers to perform precise scheduling, energy-efficient operation, carbon-aware optimization, and accurate cost estimation—thereby improving cluster resource utilization and scheduling efficiency.
In dynamic edge computing, inaccurate performance prediction arises from application co-location and node heterogeneity. Method: This paper proposes a lightweight, customized performance prediction framework for real-time scheduling. It automatically identifies critical performance metrics from historical monitoring data and jointly optimizes multiple machine learning models to achieve Pareto-optimality between prediction accuracy and inference latency. Furthermore, it customizes model selection per server’s heterogeneous characteristics to adapt to dynamic coexistence scenarios. Contribution/Results: Experimental evaluation demonstrates up to 90% prediction accuracy and inference latency below 1% of end-to-end round-trip time (RTT), significantly improving resource scheduling efficiency and system predictability for electron microscopy workflows in edge environments.
Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.
Cloud performance is influenced by multi-scale, time-varying factors, and existing decomposition methods struggle to effectively capture their intermittent behavior and complex periodic patterns. This work proposes both a hybrid (expert-informed) and a fully automated time series decomposition approach that, for the first time, stably extracts multi-scale trend and seasonal components from a single performance trace. These components are directly leveraged for Serverless function performance prediction and AWS resource scheduling. The proposed method significantly outperforms baseline approaches, achieving prediction MAPE as low as 1.8% (hybrid) and 2.1% (fully automated), reducing latency variability on AWS by over 60%, and decreasing peak latency by up to 10%, thereby offering high-precision decision support for cloud resource provisioning.
This study addresses the limitations of traditional approaches in efficiency and applicability for performance prediction of data flows in wired networks. It systematically reviews the decades-long evolution of network performance modeling, encompassing discrete-event simulation, queueing theory, network calculus, machine learning, and hybrid methods. The work innovatively proposes a unified taxonomy of modeling paradigms, revealing a paradigm shift from analytical and simulation-based techniques toward data-driven deep learning. It further provides a detailed analysis of how these approaches differ in evaluation objectives, underlying assumptions, and comparability. By clarifying the strengths, limitations, and appropriate application scenarios of each methodology, this research establishes a comprehensive reference framework to guide future advances in network performance modeling.
This work addresses the absence of a systematic, traceable, and reproducible framework for reporting the performance of mathematical libraries—a gap that hinders accurate performance evaluation and resource planning for scientific applications on high-performance computing (HPC) systems. To this end, the paper introduces LAAB, the first framework explicitly designed around four core principles: traceability, compatibility, reliability, and accessibility. LAAB establishes an end-to-end reproducible performance evaluation pipeline through standardized benchmarking protocols, comprehensive metadata management, execution environment tracking, and advanced performance analysis techniques. The framework substantially enhances the accuracy and interoperability of mathematical library performance reporting, thereby providing a robust foundation for performance prediction and resource scheduling in scientific computing.
This study addresses the challenges of control design in complex industrial processes characterized by multivariable coupled dynamics by proposing an automated control strategy generation framework that integrates large language models (LLMs) with Bayesian optimization. The approach decomposes control design into structured code generation steps, ensuring physical consistency through execution-based validation and feedback-driven repair. It pioneers the automatic synthesis of decentralized PI controller architectures and their tuning environments directly from dynamic process models. Evaluated on a nonlinear gas preheater benchmark, the generated control schemes—subsequently refined via Bayesian optimization—achieve a 26.5% improvement in closed-loop performance and significantly enhance the transient response of pressure loops, thereby demonstrating the method’s effectiveness and novelty.